A Video Super-Resolution Reconstruction Method Based on Dual Metric Feature Fusion

By using the dual-metric feature fusion method in the video super-resolution reconstruction algorithm, the correlation of feature vectors of adjacent video frames is calculated and used for global feature fusion, the problem of ignoring global correlation and channel difference in the process of feature fusion in the prior art is solved, and a more realistic and clear super-resolution image reconstruction is achieved.

CN115797177BActive Publication Date: 2025-06-03XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211538504.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-02
Publication Date
2025-06-03
Estimated Expiration
2042-12-02

AI Technical Summary

Technical Problem

The existing video super-resolution reconstruction algorithm based on deep learning ignores the global correlation of adjacent video frames and the differences in image features in different channels during feature fusion, resulting in distortion in the reconstruction video.

Method used

The video super-resolution reconstruction method based on dual-metric feature fusion is used to calculate the correlation between the feature vectors of two adjacent frames through the cosine similarity metric and the Tanimoto similarity metric, and the correlation is used as the weight of the feature fusion of different channels to perform global feature fusion of adjacent frames.

Benefits of technology

By utilizing the global correlation of adjacent video frames and the feature differences of different channels for feature fusion, the feature fusion effect is improved, making the reconstructed super-resolution image more realistic and clear.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115797177B_ABST
    Figure CN115797177B_ABST
Patent Text Reader

Abstract

The present invention provides a video super-resolution reconstruction method based on dual metric feature fusion. By obtaining a training data set for video super-resolution, the pre-constructed video super-resolution reconstruction model is trained. The trained model can perform super-resolution reconstruction on existing videos. In the feature fusion module of the present invention, the cosine similarity metric and the Tanimoto similarity metric are used to calculate the cosine of the angle between the feature vectors of two adjacent frames and their distance respectively, and then the correlation between the feature vectors of two adjacent frames is established based on this dual metric; using the correlation as the weight for feature fusion of different channels, global feature fusion of adjacent frames is performed. Compared with the prior art, the present invention improves the feature fusion effect by utilizing the correlation between adjacent frames of the video and combining the global features of different channels, and further makes the reconstructed super-resolution image more realistic.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of digital image processing, and particularly relates to a video super-resolution reconstruction method based on dual metric feature fusion. Background Art

[0002] Video is an important carrier for transmitting digital information. With the continuous upgrade of hardware display devices and the continuous update and iteration of intelligent products, people are increasingly dissatisfied with low-resolution video images, and the demand for high-definition video images is increasing. How to obtain clearer video images with limited resources has become a hot research issue.

[0003] The goal of the video super-resolution (VSR) reconstruction method is to reconstruct a low-resolution (LR) video into a corresponding high-resolution (HR) video, improve the quality of the video, make the reconstructed video image clearer, and the image texture detail information more abundant; video super-resolution reconstruction is an ill-posed problem with great scientific research value; at the same time, this method can be widely applied to many fields such as video surveillance, remote sensing military, medical imaging, and ultra-high-definition film and television.

[0004] In the prior art, the commonly used video reconstruction method is the video super-resolution reconstruction algorithm based on deep learning. This method uses a network including a feature extraction module, an alignment module, a feature fusion module, and a reconstruction module to reconstruct the video resolution. However, the feature fusion module in the existing video super-resolution reconstruction algorithm based on deep learning relies on the local correlation of corresponding pixel pairs between adjacent frames for feature fusion, and the video sequence is a continuous process, ignoring the global correlation of adjacent frames of the video. Moreover, in the feature fusion process, the difference of image features in different channels is ignored. Therefore, the video super-resolution reconstruction algorithm of the existing technology will cause distortion in the reconstructed video. Summary of the Invention

[0005] In order to solve the above problems existing in the prior art, the present invention provides a video super-resolution reconstruction method based on dual metric feature fusion. The technical problems to be solved by the present invention are realized through the following technical solutions:

[0006] A video super-resolution reconstruction method based on dual metric feature fusion provided by the present invention includes:

[0007] Step 1: Obtain a training data set for the video super-resolution model, and perform downsampling processing on the training data set to obtain an image sequence training set;

[0008] Step 2: Use the image sequence training set as the input of a pre-constructed super-resolution reconstruction model;

[0009] Among them, the super-resolution model includes: a feature extraction module, an alignment module, a dual metric feature fusion module, and a reconstruction module. The feature extraction module is used to extract the shallow image features of the input image sequence; the alignment module is used to align the shallow image features of adjacent frames with the shallow image features of the current frame to obtain adjacent frame alignment feature vectors; the dual metric feature fusion module is used to calculate the cosine of the angle and the distance between the feature vectors of two adjacent frames respectively by using the cosine similarity metric and the Tanimoto similarity metric, establish the correlation between the feature vectors of two adjacent frames based on the dual metric, and use the correlation as the weight for fusing the feature vectors of different channels to perform global feature fusion of different channels of adjacent frames; the reconstruction module is used to reconstruct the video frames to be reconstructed in the input video sequence according to the global fusion features;

[0010] Step 3: Iteratively train the super-resolution model based on the input video training set. During each iterative training process, the super-resolution model obtains a reconstructed image through the feature extraction module, the alignment module, the dual metric feature fusion module, and the reconstruction module, and uses the loss function to realize network backpropagation to update the parameters of the super-resolution model;

[0011] Step 4: Determine whether the iterative training of the super-resolution model reaches the iteration termination condition. If so, stop the training to obtain a trained super-resolution model;

[0012] Step 5: Use the trained super-resolution model to perform super-resolution reconstruction on the existing video.

[0013] Advantages of the present invention:

[0014] The present invention provides a video super-resolution reconstruction method based on dual metric feature fusion. By obtaining a training data set for video super-resolution to train a pre-constructed video super-resolution reconstruction model, the trained model can perform super-resolution reconstruction on existing videos. In the feature fusion module of the present invention, the cosine similarity metric and the Tanimoto similarity metric are used to calculate the cosine of the angle and the distance between the feature vectors of two adjacent frames respectively, and then the correlation between the feature vectors of two adjacent frames is established based on this dual metric; the correlation is used as the weight for feature fusion of different channels to perform global feature fusion of adjacent frames. Compared with the prior art, the present invention improves the feature fusion effect by utilizing the correlation of adjacent frames of the video and combining the global features of different channels, and thus makes the reconstructed super-resolution image more realistic.

[0015] The following will further elaborate on the present invention in conjunction with the accompanying drawings and embodiments. Description of the Drawings

[0016] Figure 1 is a schematic flowchart of a video super-resolution reconstruction method based on dual metric feature fusion according to the present invention;

[0017] Figure 2 is a schematic diagram of the network framework of the video super-resolution reconstruction model provided by the present invention;

[0018] Figure 3 is a schematic diagram of the alignment module provided by the present invention;

[0019] Figure 4 is a schematic diagram of the cross-scale information alignment block provided by the present invention;

[0020] Figure 5 is a schematic diagram of the cross-scale dilated residual convolution block provided by the present invention;

[0021] Figure 6 is a schematic diagram of the cosine metric and Tanimoto metric feature fusion unit provided by the present invention;

[0022] Figure 7 is a schematic diagram of the temporal attention unit provided by the present invention;

[0023] Figure 8 is a schematic diagram of the spatial attention unit provided by the present invention;

[0024] Figure 9 is a comparison diagram of the visual effects of 4× multiple super-resolution reconstruction of the "Calendar" dataset in Vid4. Detailed Embodiments

[0025] The following further describes the present invention in detail with reference to specific embodiments, but the implementation manners of the present invention are not limited thereto.

[0026] As Figure 1 shown, a video super-resolution reconstruction method based on dual metric feature fusion provided by the present invention includes:

[0027] Step 1: Obtain a training dataset for video super-resolution, and perform downsampling processing on the training dataset to obtain an image sequence training set;

[0028] Specifically, Step 1 of the present invention includes:

[0029] Step 11: Obtain a training dataset for video super-resolution from a publicly available network training database;

[0030] Step 12: Through bicubic interpolation downsampling of the training dataset by a determined multiple, construct a low-resolution dataset corresponding to the high-resolution image;

[0031] Step 13: For any current frame in the low-resolution dataset, use a sequence of 7 consecutive low-resolution images, which includes the current frame, the previous 3 frames of the current frame, and the next 3 frames of the current frame, as one input unit;

[0032] Among them, the 4th frame in the sequence of 7 consecutive low-resolution images is the current frame image to be reconstructed;

[0033] Step 14: Perform data augmentation operations on the low-resolution image sequence of each input unit to obtain an image sequence training set.

[0034] In each network training of the present invention, 8 units are input, and after performing data augmentation operations such as slicing, rotation, and flipping on the images, they are used as network inputs.

[0035] Step 2: Use the image sequence training set as the input of a pre-constructed super-resolution reconstruction model;

[0036] Reference Figure 2 , the super-resolution model of the present invention includes: a feature extraction module, an alignment module, a dual metric feature fusion module, and a reconstruction module. The feature extraction module is used to extract shallow image features of the input image sequence; the alignment module is used to align the shallow image features of adjacent frames with the shallow image features of the current frame to obtain adjacent frame alignment feature vectors; the dual metric feature fusion module is used to calculate the cosine of the angle and the distance between the feature vectors of two adjacent frames respectively using cosine similarity metric and Tanimoto similarity metric, establish the correlation between the feature vectors of two adjacent frames based on the dual metric, and use the correlation as the weight for fusing the feature vectors of different channels to perform global feature fusion of different channels of adjacent frames; the reconstruction module is used to reconstruct the video frames to be reconstructed in the input video sequence according to the global fusion features;

[0037] In each 7-frame continuous image unit input to the network of the present invention, the 4th frame image is the current frame image to be reconstructed, and the remaining images are adjacent frame images.

[0038] Specifically, the construction process of the video super-resolution reconstruction model of the present invention includes:

[0039] Step 21: Construct a feature extraction module;

[0040] The feature extraction module includes a convolutional layer with a convolutional kernel size of 3×3, which is used to promote the input data from the channel dimension 3 to shallow image features of 64; 5 cascaded residual blocks, each residual block is connected in series with two convolutional layers with a convolutional kernel size of 3×3, and the features are activated by the Relu activation function in the middle, and the input and output dimensions are 64;

[0041] Step 22: Construct an alignment module;

[0042] As Figure 3 shown, the alignment module of the present invention includes 5 cross-scale information alignment blocks. Each cross-scale information alignment block includes a convolutional layer with a convolutional kernel size of 3×3, a cross-scale dilated residual block, a convolutional layer with a convolutional kernel size of 3×3, and a deformable convolutional layer;

[0043] Align the shallow image features of 6 adjacent frames in step b with the shallow features of the current frame respectively to solve the problem of inaccurate displacement estimation caused by large motion or multiple motion directions of adjacent frames in traditional methods, so as to obtain 6 adjacent frame alignment feature vectors with a dimension of 64. Concatenate all adjacent frame alignment feature vectors and the current frame feature with a dimension of 64, and input them into the dual metric feature fusion module;

[0044] Refer to Figure 4 shown, the cross-scale dilated residual block in the alignment module includes 10 convolutional layers. The first to fourth convolutional layers are cascaded convolutional layers with a convolutional kernel size of 3×3. The fifth to ninth layers are 5 parallel convolutional layers with dilation rates increasing from 1 to 5 in sequence. The tenth layer is a convolutional layer module with a convolutional kernel size of 1×1;

[0045] Among them, the feature extraction module outputs shallow image features with a dimension of 64 to the cross-scale dilated residual block. After extracting features through the cross-scale dilated residual block, an offset parameter is generated through a 3×3 convolutional layer, and the image features are aligned using 1 layer of deformable convolution. By cascading 5 cross-scale information alignment blocks, the shallow image features of adjacent frames are progressively aligned using the current frame.

[0046] Step 23: Construct a dual metric feature fusion module;

[0047] The dual metric feature fusion module of the present invention includes six cosine metric and Tanimoto metric feature fusion units, a temporal attention unit, and a spatial attention unit. One cosine metric and Tanimoto metric feature fusion unit includes a total of seven convolutional layers. The input is the feature vectors of adjacent frames and the current frame feature after alignment. The first convolutional layer is a convolutional layer with a convolutional kernel size of 3×3. The second convolutional layer is a convolutional layer with a convolutional kernel size of 3×3. The first convolutional layer is cascaded with two convolutional layers with a convolutional kernel size of 1×1, the third and fourth layers. The second convolutional layer is cascaded with 3 convolutional layers with a convolutional kernel size of 1×1, the third, fifth, and sixth layers. The seventh convolutional layer is a convolutional layer with a convolutional kernel size of 3×3;

[0048] Currently, the feature fusion module of video super-resolution reconstruction algorithms based on deep learning relies on the local correlation of corresponding pixel pairs between adjacent frames, while ignoring the global correlation of adjacent frames in the video. Moreover, during the feature fusion process, the differences in image features across different channels are ignored. In this feature fusion module, the cosine similarity metric and the Tanimoto similarity metric, namely the dual metrics, are used to calculate the cosine of the angle between the feature vectors of two adjacent frames and their distance respectively. Then, based on this dual metric, the correlation between the feature vectors of two adjacent frames is established, and the correlation is used as the weight for feature fusion across different channels to perform global feature fusion of adjacent frames.

[0049] Reference Figure 5 As shown, the input to the first convolutional layer in each group of cosine metric and Tanimoto metric feature fusion units is the feature vector x of adjacent frames, and the input to the second convolutional layer is the feature vector y of the current frame. The features output by the first convolutional layer and the features output by the second convolutional layer are jointly input to the third convolutional layer, and features a and b with a dimension of 1 are generated respectively through the third convolutional layer. The feature vector x of adjacent frames generates a feature X with a dimension of 32 through the first convolutional layer and the fourth convolutional layer. The feature vector y of the current frame passes through the second convolutional layer and then through the fifth and sixth convolutional layers respectively to obtain a feature V with a dimension of 32 y , Y; The distance correlation f of the feature vector is calculated through the Tanimoto similarity metric 1 , and then calculated through the softmax function and multiplied pixel by pixel (⊙) with the feature V y to obtain the Tanimoto metric fusion feature; The features X and Y are calculated and multiplied through the cosine similarity metric to obtain the cosine of the angle correlation f between the feature vectors of two adjacent frames 2 ; The cosine of the angle correlation f 2 is calculated through the softmax function and multiplied pixel by pixel with the feature V y to obtain the cosine metric fusion feature; The Tanimoto metric fusion feature and the cosine metric fusion feature are added together and input to the seventh convolutional layer to increase the feature dimension to 64, and then added to the feature vector x of adjacent frames to obtain the spatio-temporal information fusion feature, and the spatio-temporal information fusion feature is input to the temporal attention unit.

[0050] The following formula is used to calculate the spatio-temporal information fusion feature obtained through the cosine metric and Tanimoto metric feature fusion unit:

[0051]

[0052] where i represents the position of the processed pixel point, n represents all pixel point positions on y, and z i represents x iOutput of the point, S(·) is the softmax function operation, ρ is the adaptive weight coefficient with an initial value of 4, and f 1 (·) is the cosine metric relation function, and f 2 (·) is the Tanimoto metric relation function;

[0053] The cosine similarity metric and the Tanimoto similarity metric are calculated using the following formula:

[0054]

[0055]

[0056] The spatio-temporal information fusion feature performs differential feature fusion through the time attention unit and the space attention unit based on the temporal relationship between adjacent frames and the spatial relationship of different channels of the image frame.

[0057] Reference Figure 6 As shown, the time attention unit includes two embedding layers and one fusion layer. Each embedding layer is a convolutional layer with a convolutional kernel size of 3×3 and two layers, and the input and output dimensions are both 64. The PReLU activation function is used to activate the features after the first convolutional layer; the fusion layer is a convolutional layer with a convolutional kernel size of 3×3, the input dimension is 448, and the output dimension is 64;

[0058] The input of the time attention unit is 6 spatio-temporal information fusion features and the shallow image features of the current frame output by the feature extraction module. The 7 input features are spliced and fused through the splicing function to obtain 7-frame continuous video features {LR t-3 ,..., LR t+3};

[0059] The time attention unit respectively extracts embedding features from the 7-frame continuous video frame features {LR t-3 ,..., LR t+3} and the current frame feature LR t through the embedding layer. The extracted embedding features are shown in formula (4):

[0060]

[0061] where f emb1 , f emb2 represents the embedding function;

[0062] The local feature similarity degree is obtained by performing the embedding feature dot product calculation on the corresponding elements of the current frame and the reference frame, and then the sigmoid activation function is used for activation to obtain the time attention feature map, as shown in the following formula (5):

[0063]

[0064] Multiply the temporal attention feature map with the initial input of the unit pixel by pixel to obtain the weighted features of the corresponding consecutive video frames, as shown in the following formula (6):

[0065] F t-i = LR t-i ⊙ T t-i , |i| ≤ 3 (6)

[0066] Concatenate the weighted features {F t-3 ,..., F t+3} along the channel dimension, and then pass through the fusion layer to obtain the spatio-temporal fusion feature F fusion , as shown in the following formula (7):

[0067] F fusion = f fusion {F t-3 ,..., F t+3} (7)

[0068] where f fusion represents the fusion function;

[0069] Input the spatio-temporal fusion feature into the spatial attention unit.

[0070] The spatial attention unit calculates the weights of different channels in the feature space by using the image feature information of each channel dimension and performs feature fusion, reshaping the features of each channel dimension of the image features to enhance the representation ability of features in different dimensions of the feature space.

[0071] As shown in Figure 7 , the spatial attention unit consists of 4 convolutional layers, and the convolutional kernel size of the 4 convolutional layers is 1×1;

[0072] The input image feature of the first convolutional layer in the spatial attention unit is F fusion ∈ R C×H×W ; where C is the feature channel dimension, H and W are the spatial positions of the feature points, and the image feature x is used to represent F fusion , and it is defined as where N represents all the spatial points that the x i feature points can appear in the spatial position of size H×W;

[0073] The image feature x is transformed into a feature map W q x of H×W×1 through a convolutional layer with a convolutional kernel size of 1×1, and a spatial attention map of size H×W×1 is obtained by using the softmax function, as shown in the following formula (8):

[0074]

[0075] Among them, W v , W q represents the weight matrix, and c is the spatial attention map;

[0076] Multiply the feature input x by the spatial attention map c to obtain a feature of size C×1×1, and pass it through W v to obtain the feature information of each channel feature. The feature information is then adjusted for the channel parameters by concatenating two convolutional layers with a kernel size of 1×1 and activating the feature through the sigmoid activation function to obtain the channel feature descriptor sigmoid(f mlp (c i ));

[0077] Multiply the channel feature descriptor by the input feature x, add the multiplication result to the input feature x, and obtain the global feature after fusing the consecutive video frames output by the dual metric feature fusion module, as shown in the following formula (9):

[0078] F′ fusion = sigmoid(f mlp (c i ))·F fusion + F fusion (9)

[0079] Among them, F′ fusion is the output of the dual metric feature fusion module, and f mlp (·) represents a function calculated through two convolutional layers with a kernel size of 1×1 and a PRelu activation function in the middle. The input and output dimensions of the first convolutional layer are 64 and 128 respectively, and the input and output dimensions of the second convolutional layer are 128 and 64 respectively; F′ fusion realizes the global feature fusion of consecutive video frames, laying a foundation for the next image reconstruction, thereby improving the overall performance of the algorithm.

[0080] Input the global feature into the reconstruction module.

[0081] Step 24: Construct a reconstruction module;

[0082] Among them, the reconstruction module is used to aggregate multi-scale information using 20 layers of densely connected residual blocks, deepen the network depth, upsample the output features using sub-pixel convolution, and finally output reconstruction features with a dimension of 3 through a convolutional layer with a kernel size of 1×1. Add the reconstruction features to the image features obtained by bicubic interpolation upsampling of the current frame feature vector before the input shallow feature extraction module to obtain the reconstructed high-resolution image.

[0083] Step 3: Iteratively train the super-resolution model based on the input video training set. During each iterative training process, the super-resolution model obtains a reconstructed image through a feature extraction module, an alignment module, a dual metric feature fusion module, and a reconstruction module, and uses a loss function to implement network backpropagation to update the parameters of the super-resolution model;

[0084] In this step, the network parameters are trained and updated to optimize the network. After the training rounds are completed, the training ends and proceeds to the next step, which specifically includes the following:

[0085] (1) Data input: Use the Dataloader class in pytorch to collect data and input it into the model. The input batch size is 8, the slice size is 48×48, and 4 threads are used to read the data in parallel;

[0086] (2) Design the network optimizer and hyperparameters: Use the Adam optimizer as the network model optimizer, β 1 = 0.9, β 2 = 0.999, ε = 10 -8 , set the number of training rounds to 150, save the model parameters every 20 training rounds, the initial learning rate is 0.0002, and the learning rate decay strategy is that at the 30th round and the 60th round, the learning rate is halved, and starting from the 60th round, the learning rate is halved every 20 rounds;

[0087] (3) Design the loss function: Use the L1 loss function as the network loss function and use the L2 loss function to fine-tune the network;

[0088] Step 4: Determine whether the iterative training of the super-resolution model reaches the iteration termination condition. If so, stop the training and obtain the trained super-resolution model;

[0089] Among them, the iteration termination condition includes reaching the number of iterations or the result of the loss function no longer decreasing.

[0090] Step 5: Use the trained super-resolution model to perform super-resolution reconstruction on the existing video.

[0091] The following further describes the effect of the present invention in combination with simulation experiments.

[0092] In terms of the operating environment and facilities, experiments are carried out on a high-performance computer with Tesla P100-PCIE-16GB, and the python environment is configured as pytorch1.0.0 and CUDA10.0.

[0093] The experiment uses the Vimeo-90K dataset commonly used in the VSR field. The dataset includes 64,612 high-resolution video sequences with a resolution of 448×256. By performing bicubic interpolation downsampling on the Vimeo-90K dataset by a factor of 4, a low-resolution dataset corresponding to the high-resolution image dataset is constructed; the Vid4 dataset is used as the standard test set, and the Vid4 dataset contains a total of 4 different video sequences, namely: "walk", "city", "foliage" and "calendar" datasets.

[0094] Simulation experiment method: The present invention and 6 existing deep learning-based super-resolution methods use the objective image quality evaluation metrics Peak Signal to Noise Ratio (PSNR) and Structural Similarity Index (SSIM) to compare the advantages and disadvantages of the algorithms.

[0095] Simulation experiment content:

[0096] Compare the present invention with 6 existing super-resolution methods to obtain the PSNR and SSIM metrics, as shown in Table 1:

[0097] Table 1 Results of objective metrics of PSNR (dB) and SSIM for super-resolution reconstruction at 4× magnification for each method on Vid4

[0098]

[0099] Reference Figure 7 , Figure 7 is the visual effect comparison chart of 4× magnification super-resolution reconstruction for the "Calendar" dataset in Vid4. From Figure 7It can be seen that the feature fusion module of the video super-resolution reconstruction algorithm based on deep learning currently relies on the local correlation of corresponding pixel pairs between adjacent frames, while ignoring the global correlation of adjacent frames in the video. Moreover, during the feature fusion process, the differences in image features in different channels are ignored. In the feature fusion module of the present invention, the cosine similarity measure and the Tanimoto similarity measure, that is, the dual measure, are used to calculate the cosine of the angle between the feature vectors of two adjacent frames and their distance respectively. Then, based on this dual measure, the correlation between the feature vectors of two adjacent frames is established. Subsequently, the correlation is used as the weight for feature fusion in different channels to perform global feature fusion of adjacent frames, laying a foundation for the next step of image reconstruction, thereby improving the overall performance of the algorithm. Experiments show that when the module proposed in this paper is applied to the video super-resolution reconstruction algorithm, the reconstructed high-resolution frames are subjectively clearer and have better visual effects. Objectively, in terms of the reconstruction performance metrics on the commonly used video super-resolution test set Vid4, it is superior to most current mainstream algorithms.

[0100] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality" means two or more unless otherwise specifically defined.

[0101] Although the present application has been described in conjunction with various embodiments herein, however, in the process of implementing the claimed present application, those skilled in the art can understand and achieve other variations of the disclosed embodiments by viewing the accompanying drawings, the disclosed content, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "one" does not exclude a plurality of cases.

[0102] The above content is a further detailed description of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.

Claims

1. A video super-resolution reconstruction method based on dual metric feature fusion, characterized in that, it includes: Step 1: Obtain the training dataset of the video super-resolution model, and perform downsampling on the training dataset to obtain an image sequence training set; Step 2: Use the image sequence training set as the input of a pre-constructed super-resolution reconstruction model; Among them, the super-resolution model includes: a feature extraction module, an alignment module, a dual metric feature fusion module, and a reconstruction module. The feature extraction module is used to extract the shallow image features of the input image sequence; the alignment module is used to align the shallow image features of adjacent frames with the shallow image features of the current frame to obtain adjacent frame alignment feature vectors; the dual metric feature fusion module is used to calculate the cosine of the angle and the distance between the feature vectors of two adjacent frames respectively using cosine similarity metric and Tanimoto similarity metric, establish the correlation between the feature vectors of two adjacent frames based on the dual metric, and use the correlation as the weight for fusing the feature vectors of different channels to perform global feature fusion of different channels of adjacent frames; the reconstruction module is used to reconstruct the video frame to be reconstructed in the input video sequence according to the global fusion features; Step 3: Iteratively train the super-resolution model based on the input video training set. During each iterative training process, the super-resolution model obtains a reconstructed image through the feature extraction module, the alignment module, the dual metric feature fusion module, and the reconstruction module, and uses the loss function to implement network backpropagation to update the parameters of the super-resolution model; Step 4: Determine whether the iterative training of the super-resolution model reaches the iteration termination condition. If so, stop training to obtain a trained super-resolution model; Step 5: Use the trained super-resolution model to perform super-resolution reconstruction on existing videos.

2. The video super-resolution reconstruction method based on dual metric feature fusion according to claim 1, characterized in that, Step 1 includes: Step 11: Obtain the training dataset of video super-resolution from the publicly available training database on the network; Step 12: By performing bicubic interpolation downsampling on the training dataset by a certain multiple, a low-resolution dataset corresponding to the high-resolution image is constructed; Step 13: For any current frame in the low-resolution dataset, use the current frame, the first 3 frames of the current frame, and the last 3 frames of the current frame, a total of 7 consecutive low-resolution image sequences as one input unit; Among them, the 4th frame in the 7 consecutive low-resolution image sequences is the current frame image to be reconstructed; Step 14: Perform data augmentation operations on the low-resolution image sequences of each input unit to obtain an image sequence training set.

3. The video super-resolution reconstruction method based on dual metric feature fusion according to claim 2, characterized in that, The construction process of the video super-resolution reconstruction model includes: Step 21: Construct a feature extraction module; The feature extraction module includes a convolutional layer with a convolution kernel size of 3×3, which is used to enhance the input data from a channel dimension of 3 to a shallow image feature with 64 channels; 5 cascaded residual blocks, and each residual block concatenates two convolutional layers with a convolution kernel size of 3×3, and the features are activated by a Relu activation function in the middle, and the input and output dimensions are 64; Step 22: Construct an alignment module; Among them, the alignment module includes 5 cross-scale information alignment blocks, and each cross-scale information alignment block includes a convolutional layer with a convolution kernel size of 3×3, a cross-scale dilated residual block, a convolutional layer with a convolution kernel size of 3×3, and a deformable convolutional layer; Step 23: Construct a dual metric feature fusion module; Among them, the dual metric feature fusion module includes six cosine metric and Tanimoto metric feature fusion units, a temporal attention unit, and a spatial attention unit. One cosine metric and Tanimoto metric feature fusion unit includes a total of seven convolutional layers. The input is the feature vectors of adjacent frames and the current frame feature vector after alignment. The first convolutional layer is a convolutional layer with a convolution kernel size of 3×3, the second convolutional layer is a convolutional layer with a convolution kernel size of 3×3, the first convolutional layer concatenates two convolutional layers with a convolution kernel size of 1×1 as the third and fourth layers, the second convolutional layer concatenates 3 convolutional layers with a convolution kernel size of 1×1 as the third, fifth, and sixth layers, and the seventh convolutional layer is a convolutional layer with a convolution kernel size of 3×3; Step 24: Construct a reconstruction module; Among them, the reconstruction module is used to aggregate multi-scale information using 20 layers of densely connected residual blocks to deepen the network depth. The output features are upsampled using sub-pixel convolution, and finally, a reconstruction feature with a dimension of 3 is output through a convolutional layer with a convolution kernel size of 1×1. The reconstruction feature is added to the image feature obtained by bilinearly interpolating and upsampling the current frame feature vector before the input shallow feature extraction module to obtain a reconstructed high-resolution image.

4. A video super-resolution reconstruction method based on dual metric feature fusion according to claim 3, characterized in that the cross-scale dilated residual block in the alignment module includes 10 convolutional layers. The first to fourth convolutional layers are cascaded convolutional layers with a convolution kernel size of 3×3. The fifth to ninth layers are 5 parallel convolutional layers with dilation rates increasing from 1 to 5 in sequence. The tenth layer is a convolutional layer module with a convolution kernel size of 1×1; Among them, the shallow image feature with an output dimension of 64 from the feature extraction module is input to the cross-scale dilated residual block. After extracting features through the cross-scale dilated residual block, offset parameters are generated through a 3×3 convolutional layer, and the image features are aligned using 1 layer of deformable convolution. Through cascading 5 cross-scale information alignment blocks, the shallow image features of adjacent frames are progressively aligned using the current frame.

5. A video super-resolution reconstruction method based on dual metric feature fusion according to claim 3, characterized in that The input of the first convolutional layer in each group of cosine metric and Tanimoto metric feature fusion units is the feature vector x of adjacent frames, and the input of the second convolutional layer is the feature vector y of the current frame; the features output by the first convolutional layer and the features output by the second convolutional layer are jointly input into the third convolutional layer, and features a and b with a dimension of 1 are respectively generated through the third convolutional layer; the feature vector x of adjacent frames generates a feature X with a dimension of 32 through the first convolutional layer and the fourth convolutional layer; the feature vector y of the current frame passes through the second convolutional layer, and then respectively passes through the fifth and sixth convolutional layers to obtain features V with a dimension of 32 y , Y; the distance correlation f of the feature vector is calculated through the Tanimoto similarity metric for features a and b 1 , and then calculated through the softmax function and multiplied pixel by pixel with feature V y to obtain the Tanimoto metric fusion feature; the feature X and Y are calculated and multiplied by cosine similarity measure to obtain the cosine correlation f of the feature vectors of two adjacent frames 2 ; the cosine correlation f 2 is calculated by the softmax function and multiplied pixel by pixel with the feature V y to obtain the cosine metric fusion feature; The Tanimoto metric fusion feature and the cosine metric fusion feature are added together, and then input into the seventh convolutional layer to increase the feature dimension to 64. Then, it is added to the feature vector x of the adjacent frame to obtain the spatio-temporal information fusion feature, and the spatio-temporal information fusion feature is input into the temporal attention unit.

6. A video super-resolution reconstruction method based on dual metric feature fusion according to claim 5, wherein, the cosine metric and Tanimoto metric feature fusion unit is used for: calculating the spatio-temporal information fusion feature obtained through the cosine metric and Tanimoto metric feature fusion unit by the following formula: Among them, i represents the position of the processed pixel point, n represents all pixel point positions on y, and z i represents the output of the x i point, S(·) is the softmax function operation, ρ is the adaptive weight coefficient with an initial value of 4, and f 1 (·) is the cosine metric relationship function, and f 2 (·) is the Tanimoto metric relationship function; calculating the cosine similarity metric and the Tanimoto similarity metric by the following formula:

7. A video super-resolution reconstruction method based on dual metric feature fusion according to claim 5, wherein, the temporal attention unit includes two embedding layers and a fusion layer, where each embedding layer is a convolutional layer with a convolutional kernel size of 3×3 and two layers, and the input and output dimensions are both 64. After the first convolutional layer, the PReLU activation function is used to activate the features; the fusion layer is a convolutional layer with a convolutional kernel size of 3×3, the input dimension is 448, and the output dimension is 64; The input of the temporal attention unit is six spatio-temporal information fusion features and the shallow image features of the current frame output by the feature extraction module. The seven input features are spliced and fused through a splicing function to obtain seven-frame continuous video features {LR t-3 ,...,R t+3}; The temporal attention unit, through the embedding layer, respectively extracts the embedding features from 7 consecutive video frame features {LR t-3 ,..., R t+3} and the current frame feature LR t . The extracted embedding features are shown in Equation (4): where f emb1 , f emb2 represents an embedding function; calculating the local feature similarity degree by performing an embedding feature dot product on the corresponding elements of the current frame and the reference frame, and then using the sigmoid activation function to activate it to obtain the temporal attention feature map, as shown in the following formula (5): multiplying the temporal attention feature map by the initial input of the unit pixel by pixel to obtain the weighted features of the corresponding consecutive video frames, as shown in the following formula (6): F t-i = LR t-i ⊙T t-i where |i| ≤ 3 (6) The weighted features {F t-3 ,..., F t+3} of consecutive video frames are concatenated along the channel dimension and then passed through a fusion layer to obtain the spatio-temporal fusion feature F fusion , as shown in Equation (7) below: F fusion = f fusion {F t-3 ,..., F t+3}} (7) where f fusion represents a fusion function; Input the spatio-temporal fusion feature into the spatial attention unit.

8. A video super-resolution reconstruction method based on dual metric feature fusion according to claim 7, wherein, the spatial attention unit is composed of 4 convolutional layers, and the convolutional kernel size of the 4 convolutional layers is 1×1; The input image feature of the first convolutional layer in the spatial attention unit is F fusion ∈R C×H×W ; where C is the feature channel dimension, H and W are the spatial positions of the feature points, and the image feature x is used to represent F fusion , define where N represents x i All spatial points that the feature points can appear in the spatial positions of size H×W The image feature x is transformed into a feature map W of H×W×1 through a convolutional layer with a convolutional kernel size of 1×1 q x, and a spatial attention map of size H×W×1 is calculated using the softmax function, as shown in the following equation (8): Among them, W v , W q represents the weight matrix, and c is the spatial attention map; Multiply the feature input x by the spatial attention map c to obtain a feature of size C×1×1, and pass it through W v Obtain the feature information of each channel feature. The feature information is then passed through two convolutional layers with a kernel size of 1×1 in series to adjust the channel parameters and activate the features through the sigmoid activation function to obtain the channel feature descriptor sigmoid(f mlp (c i )); multiplying the channel feature descriptor by the input feature x, and adding the multiplication result to the input feature x to obtain the global feature after fusion of the consecutive video frames output by the dual metric feature fusion module, as shown in the following formula (9): F′ fusion = sigmoid(f mlp (c i ) · F fusion + F fusion (9) Among them, F′ fusion is the output of the dual metric feature fusion module, and f mlp (·) represents a function calculated by a 1×1 convolutional layer with a PRelu activation function in the middle of two layers; Input the global feature into the reconstruction module.

Citation Information

Patent Citations

  • Video super-resolution reconstruction method and device

    CN106254722A

  • Frame level feature aggregation method for video target detection

    CN109993095A