Video frame image enhancement method based on mixed residual dense block
By adopting the image enhancement method of mixed residual dense blocks in the infrared image sequence super-resolution network, the problems of high computing complexity, large parameters and insufficient generalization capabilities in the prior art are solved, and efficient infrared image super-resolution processing in resource-constrained environments are realized.
Patent Information
- Application Number
- CN202510242593.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-06-17
AI Technical Summary
The existing infrared image sequence super-resolution network has the limitations of high computational complexity, large model parameters, lack of efficient feature extraction modules for small infrared targets, and training strategies, which leads to difficulty in application in resource-constrained environments and insufficient model generalization capabilities.
Using a video frame image enhancement method based on mixed residual dense blocks, an image enhancement model of infrared video frames is constructed through feature mapping, feature grouping and fusion, feature enhancement, feature alignment and feature reconstruction modules, to reduce the computational complexity, and to efficiently extract infrared image features under small parameters.
It effectively reduces the complexity of the model, makes its calculation complexity stable when inputting multiple frames, and can efficiently extract infrared image features under small parameters, improving the generalization ability of the model and its application ability in resource-constrained environments.
Smart Images

Figure CN120163750A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image enhancement, and particularly to a video frame image enhancement method based on a hybrid residual dense block. Background Art
[0002] There are usually two methods to obtain high-resolution infrared images: one is to directly obtain high-resolution images by expanding the size of the infrared sensor array; the other is to design an image processing algorithm to improve the resolution of the original infrared image. Given the limitations of sensor technology and the high cost of infrared sensors, developing a low-cost and highly reliable infrared image super-resolution algorithm has become a more practical and feasible solution.
[0003] Generally, the target in an infrared image occupies a very small proportion of the entire image area (usually less than 0.12%), and lacks color and detailed structural information (such as contours, shapes, and textures, etc.). This makes the information available for super-resolution processing in a single-frame image extremely limited. Therefore, a super-resolution method based on an infrared image sequence has been proposed, which improves the super-resolution performance by utilizing the correlation information between multiple frames of the sequence. However, the existing infrared image sequence super-resolution networks still have the following deficiencies:
[0004] (1) Lack of an efficient feature extraction module for infrared small target image sequences. Currently, although many image feature extraction modules have been proposed in the field of image super-resolution and video super-resolution, these modules do not consider the imaging characteristics of medium and long-wave infrared small target images themselves. Designing an efficient feature extraction module for infrared small target images is an important issue in current research.
[0005] (2) High computational complexity: Existing models usually have high computational complexity and a large number of model parameters. At the same time, the complexity of the model increases linearly with the number of frames of the input image sequence. This significantly limits its practical application in resource-constrained environments.
[0006] (3) Limitations of the training strategy: Due to the incompleteness of the dataset, existing methods mainly rely on a limited infrared image sequence dataset for model training. The scenes of these datasets often have a high degree of repetition, resulting in insufficient generalization ability of the model. Summary of the Invention
[0007] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a video frame image enhancement method based on a hybrid residual dense block, whose computational complexity does not increase significantly with the increase in the number of input frames, and can efficiently extract features in infrared images with a small number of parameters.
[0008] The purpose of the present invention is achieved by the following technical solutions: A video frame image enhancement method based on a hybrid residual dense block, comprising the following steps:
[0009] Build an image enhancement model for infrared video frames. The image enhancement model for infrared video frames includes a feature mapping module, a feature grouping and fusion module, a feature enhancement module, a feature alignment module, a feature reconstruction module, and an image output module;
[0010] Among them, the feature mapping module uses a residual dense block common in infrared super-resolution, and maps consecutive infrared video frames into multiple frames of features through convolution operations.
[0011] Among them, the feature grouping and fusion module groups and fuses the input multiple frames of features using a residual dense block common in infrared super-resolution according to the time distance of the input multiple frames of features:
[0012] First, select the first frame of features from the input multiple frames of features, and then calculate the time distance between each frame of the input multiple frames of features and the first frame of features;
[0013] Preset multiple consecutive and non-overlapping time distance intervals, and map each frame of the multiple frames of features into different time distance intervals according to the time distance from the first frame of features. The features within each time distance interval are used as a group of features;
[0014] For each group of features, fuse them respectively using a residual dense block common in infrared super-resolution.
[0015] Among them, the feature enhancement module is used to enhance each group of fused features respectively; the feature enhancement module uses a central difference hybrid residual dense block obtained by improving the residual dense block.
[0016] Among them, the feature alignment module is used to align each group of features and splice them along the channel dimension to achieve feature aggregation;
[0017] The mathematical expression form of the feature alignment module is:
[0018] F a = A(F n , F ref ),
[0019] Among them, F n represents the features of the current frame, F ref represents the features of the reference frame. The core goal of the feature alignment module is to use the reference frame F ref to transform the features of the current frame F n so that it obtains aligned features F ref that are spatially consistent with F a ; where the features of the current frame refer to the features of the frame being subjected to image enhancement; the features of the reference frame refer to the features of other frames in the group except the current frame.
[0020] The process of implementing feature alignment by the feature alignment module includes:
[0021] Perform feature mapping on the reference frame F ref to obtain its key features
[0022] Perform feature mapping on the current frame feature F n to obtain query features and value features respectively
[0023] Use a sliding window operation to perform block processing on the above three features, and perform local attention operations based on the feature blocks.
[0024] Among them, the feature reconstruction module is used to perform feature reconstruction based on the aggregated features to obtain a high-resolution reconstructed image and output it externally by the image output module;
[0025] The feature reconstruction module adopts a lightweight hybrid residual dense block obtained by improving the residual dense block.
[0026] Furthermore, the construction methods of the lightweight hybrid residual dense block and the central difference hybrid residual dense block are as follows:
[0027] The residual dense block consists of multiple convolutional layers with a convolutional kernel size of 3 and a convolutional layer with a convolutional kernel size of 1;
[0028] Based on the residual dense block, a convolutional layer with a convolutional kernel size of 1 is added between each convolutional operation for feature dimensionality reduction;
[0029] To expand the receptive field of feature extraction, a channel self-attention mechanism is added between the residual operation and the dense operation to enable the module to have the ability to sense global information, thereby obtaining a lightweight hybrid residual dense block;
[0030] On the basis of the lightweight hybrid residual dense block, replacing the convolutional operation with a central difference convolution results in a central difference hybrid residual dense block.
[0031] Use consecutive infrared video frames to construct a sample set for the image enhancement model, and train the image enhancement model through the sample set to obtain a trained image enhancement model;
[0032] When constructing the sample set for the image enhancement model using consecutive infrared video frames, the constructed sample set contains multiple samples;
[0033] The sample features of each sample use consecutive infrared video frames at the first resolution, and the label of each sample in the sample set uses consecutive infrared video frames at the second resolution as the label, and the first resolution is less than the second resolution;
[0034] In each sample, for the infrared video frames where the sample features and the sample labels are of the same object, the number of frames of the sample features is the same as that of the labels, and each infrared video frame in the sample features has a corresponding infrared video frame in the labels;
[0035] After being trained, the image enhancement model is used to enhance continuous infrared video frames of a first resolution into continuous infrared video frames of a second resolution, thereby improving the resolution.
[0036] The trained image enhancement model is used to perform enhancement processing on the continuous infrared video frames to be processed, and an enhanced result is obtained.
[0037] Furthermore, when performing small target detection, the trained image enhancement model can be used to perform enhancement processing on the image to be enhanced, and then the enhanced image is sent to the small target detection model for detection.
[0038] The beneficial effects of the present invention are as follows: First, the features of each image sequence are extracted; then, the features of different frames are grouped and fused through time distance, and feature enhancement is performed on the video to be super-resolved based on each group. This way of grouped enhancement effectively reduces the complexity of the model, so that the computational complexity of the model does not increase significantly with the increase in the number of input frames; in order to further improve the performance of the model, a highly efficient hybrid residual dense block for infrared images is proposed, which can efficiently extract features in infrared images with a small number of parameters; the present invention proposes a training strategy based on auxiliary data, which enables researchers to use image data as an auxiliary data set for training, thereby improving the generalization ability of the model. Description of the Drawings
[0039] Figure 1 is the flowchart of the method of the present invention;
[0040] Figure 2 is the overall architecture diagram of the model;
[0041] Figure 3 is the structural comparison diagram of different residual dense blocks;
[0042] Figure 4 is the super-resolution result comparison diagram on the SATID dataset;
[0043] Figure 5 is the super-resolution result comparison diagram on the Anti-uav dataset;
[0044] Figure 6 is the super-resolution result comparison diagram on the Hui dataset;
[0045] Figure 7Schematic diagram of the AUROC curve of the detection algorithm before and after super-resolution;
[0046] Figure 8 Schematic diagram of the comparison of the confidence maps of the detection algorithm before and after super-resolution on the SAITD dataset;
[0047] Figure 9 Schematic diagram of the comparison of the confidence maps of the detection algorithm before and after super-resolution on the Hui dataset;
[0048] Figure 10 Schematic diagram of the comparison of the confidence maps of the detection algorithm before and after super-resolution on the Anti-UAV dataset. Detailed implementation manners
[0049] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings, but the protection scope of the present invention is not limited to the following description.
[0050] As Figure 1 shown, a method for enhancing video frame images based on a hybrid residual dense block includes the following steps:
[0051] Construct an image enhancement model for infrared video frames, where the image enhancement model for infrared video frames includes a feature mapping module, a feature grouping and fusion module, a feature enhancement module, a feature alignment module, a feature reconstruction module, and an image output module;
[0052] The overall architecture of the model is as Figure 2 shown. First, use convolution operations to map consecutive infrared video frames (multiple consecutive infrared video images) into multiple frames of features; then, according to the time distance between each feature and the feature of the first frame, use the channel attention mechanism and convolution operations to group and fuse the features; next, use a feature enhancement module with parameter sharing to enhance the fused features. Among them, the feature enhancement module is composed of the proposed central difference hybrid residual dense group; then, use the feature alignment module based on the attention mechanism to align each group of features and splice them along the channel dimension to achieve feature aggregation; finally, based on the aggregated features, perform feature reconstruction and high-resolution image reconstruction.
[0053] In a multi-frame image sequence, due to the possible inconsistency of the spatial positions of objects, the target features cannot be aligned spatially. To solve this problem, the present invention proposes a feature alignment module based on the attention mechanism. The mathematical expression form of this module is:
[0054] F a = A(F n , F ref ),
[0055] where F n represents the feature of the current frame, and F refRepresents the features of the reference frame. The core objective of the feature alignment module is to utilize the reference frame F ref to transform the current frame features F n so that it obtains aligned features F ref that are spatially consistent with F a . In this study, an architecture based on Swin-Transformer is adopted to achieve feature alignment. The specific process is as follows: First, perform feature mapping on the reference frame F ref to obtain its key features Then, perform feature mapping on the current frame features F n to obtain query features and value features Next, use a sliding window operation to perform block processing on the above three types of features, and perform local attention operations based on the feature blocks. To achieve effective local modeling, the window size of the sliding window is set to 8. Considering that the displacement of small target objects in adjacent frames of infrared remote sensing images is limited, local attention is adopted to achieve feature alignment. Compared with global feature alignment, local feature alignment has a smaller computational complexity. In the feature alignment stage of this model, we use two consecutive attention mechanisms to achieve feature alignment.
[0056] The residual dense block is a feature extraction module commonly used in the field of image super-resolution. As Figure 3 shown, the traditional residual dense block consists of multiple convolutional layers with a convolutional kernel size of 3 and a convolutional layer with a convolutional kernel size of 1. Among them, the convolutional layers with a convolutional kernel size of 3 are cascaded in a dense connection manner, making full use of the feature information extracted by each layer. The convolutional layer with a convolutional kernel size of 1 is used to perform dimensionality reduction on the dense features, thereby reducing the computational complexity. This module combines the advantages of residual operations and dense connections, and can not only achieve efficient feature reuse, but also effectively retain the detailed information of the image. Based on the residual dense block, a residual dense group can be obtained. First, multiple residual dense blocks are concatenated to extract multiple features; then, the multiple extracted features are stacked; finally, a convolutional operation with a convolutional kernel size of 1 is used to perform dimensionality reduction on the stacked features.
[0057] In infrared images, considering the characteristics of small targets and wide field of view in infrared images, extracting the edge detail information of the image can better highlight the targets. Based on this motivation, Ying et al. proposed the central difference residual dense block. The central difference convolution is used to replace the ordinary convolution. And the first residual dense block in the residual dense group is replaced by the central difference residual dense block, thereby obtaining the central difference residual dense group.
[0058] The above-mentioned residual dense group and the central difference residual dense group have good feature extraction effects. However, there are still the following two deficiencies: (1) When the number of convolutional layers is large, dense operations will introduce too many parameters and computational complexity. (2) The receptive field of the module is limited, making it difficult to fully utilize the context information in the image.
[0059] To address these two deficiencies, the present invention proposes a lightweight hybrid residual dense block and a central difference hybrid residual dense block respectively. For the lightweight hybrid residual dense block, we add a convolution with a kernel size of 1 between each convolution operation for feature dimensionality reduction. Define the number of input channels of a certain residual block as C1, and the growth rate of the model as G. For the Nth dense connection operation in the module, the number of parameters of this convolution operation is
[0060] P1 = 9(C1 + NG)G,
[0061] For the improved convolution operation, the number of parameters is
[0062] P2 = (C1 + NG)G + 9G 2 ,
[0063] After subtracting the two, the reduction in the number of model parameters is:
[0064] P1 - P2 = 8C1G + 8NG 2 -9G 2 .
[0065] In the residual dense block, G is generally set to 16, 32, or 64, and C1 is generally a value greater than 64. It can be seen that as the value of N increases, the value of P1 - P2 also increases continuously. At the same time, to expand the receptive field of feature extraction, a channel self-attention mechanism is added between the residual operation and the dense operation to enable the module to have the ability to perceive global information. For the hybrid central residual dense block, we replace the convolution operation with a central difference convolution on the basis of the lightweight hybrid residual dense block.
[0066] For this model, the number of hybrid residual dense blocks in the hybrid residual dense group used in the feature enhancement stage is 4, the growth rate is 32, and the number of convolutional layers in the residual dense block is 6. For the two lightweight residual dense groups in the feature recovery stage, the number of residual dense blocks is 4 and 6 respectively, the growth rate is 32 for both, and the number of convolutional layers is 4 and 8 respectively.
[0067] Multi-scenario Hybrid Training Strategy: To enhance the generalization of the model, we added part of the data from the infrared small target detection datasets IRSTD-1k, NUDT-SIRST, and IRSatVideo-LEO as supplementary datasets to the training set. Among them, to make the image data adaptable to this framework, we regarded the image data as continuous video frames with unchanged scenarios.
[0068] Loss Function: To ensure the effectiveness of the model, multiple loss functions are used for joint supervision of the results, and the final loss function can be expressed as:
[0069] L = L1 + 0.01 * L fft
[0070] Among them, L1 refers to the MAE (Mean Absolute Error) loss, and L fft represents the frequency domain loss. When calculating the L1 loss, the Fourier transform needs to be performed on the model output and the label, and then the L1 loss between the two is calculated.
[0071] Construct a sample set for the image enhancement model using continuous infrared video frames, train the image enhancement model through the sample set, and obtain a trained image enhancement model;
[0072] Use the trained image enhancement model to enhance the continuous infrared video frames to be processed, and obtain the enhanced result.
[0073] In the embodiments of this application, the present application is further described in combination with specific experimental cases:
[0074] 1. Training Details and Evaluation Metrics
[0075] (1) Training Details: Use the video segments numbered 1 - 50 in the SAITD dataset for testing, and the remaining 300 video segments for training. At the same time, to verify the generalization of the model, the Hui and Anti-UAV datasets are used for testing. During testing, the first 100 frames of all datasets are taken. To enhance the generalization of the model, we added part of the data from the infrared small target detection datasets IRSTD-1k, NUDT-SIRST, and IRSatVideo-LEO as supplementary datasets to the training set.
[0076] Using the Adam optimizer, a total of 100 epochs were trained. For the first 80 epochs, only 10,000 randomly selected data from the SAITD dataset were used for training. For the last 20 epochs, supplementary datasets were added for mixed training. For the learning rate, in the first 50 epochs, the learning rate decayed cosine-like from 0.001 to 0.00005. For the 51st - 80th epochs, the learning rate decayed cosine-like from 0.0002 to 0.00005. For the 81st - 100th epochs, the learning rate decayed cosine-like from 0.0002 to 0.00005. For all epochs, the Adam optimizer was used and the batch size was set to 8.
[0077] (2) Evaluation metrics: The model was evaluated from multiple aspects. For objective evaluation, PSNR and SSIM commonly used in the field of image restoration were adopted for evaluation. The SNR, an indicator related to small target detection, was also used to reflect the restoration effect of the image on the target area. For subjective evaluation, some randomly selected results were visually compared. In addition, representative small target detection algorithms were selected to perform target detection on the images before and after super-resolution respectively, in order to prove the impact of image super-resolution (image enhancement, that is, enhancing low-resolution images to high-resolution images) on the downstream task of small target detection. At the same time, we also analyzed the number of model parameters and the computational complexity.
[0078] 2. Objective and subjective comparison results
[0079] (1) Objective comparison results. The objective indicators of each super-resolution method on the test set are given in Table 1. It can be seen that this model has achieved competitive results on all datasets. On the SAITD dataset, the PSNR and SSIM indicators of this model are comparable to those of MocoPnet and much higher than those of models such as TDAN. At the same time, this model has the highest SNR indicator. On the Hui dataset and the Anti-UAV dataset, this model has the highest PSNR indicator. Among them, on the Anti-UAV dataset, this model is the only one with a PSNR exceeding 32 dB.
[0080] Table 1 Objective evaluation indicators of each model on three test sets
[0081]
[0082] (2) Subjective comparison results. To show the subjective effects of the images restored by the model, we randomly selected images from the three datasets respectively for visualization. The visualization results are as Figures 4 to 6As shown, it can be seen that the target in the original image is relatively blurred and the contour is difficult to distinguish. After super-resolution, the object is clearly visible and has higher recognition. At the same time, our method and MocoPnet both have good visualization effects. However, our method has fewer model parameters and computational complexity.
[0083] 3. Verification of the effectiveness of the model in infrared small target detection
[0084] In this part, we first constructed a target detection test set, and then selected three representative small target detection methods to compare the target detection results before and after super-resolution of the test set. In terms of dataset construction, an image sequence was randomly selected from each video segment of the test sets SAITD, Hui, and Anti-UAV datasets to form the target detection test set. The constructed test set included a total of 93 image sequences in different scenarios. In the selection of detection methods, three methods, Tophat, ILCM, and RLCM, were selected.
[0085] The AUROC curves of the detection results of the three methods, Tophat, ILCM, and RLCM, on the datasets before and after super-resolution are as Figure 7 shown. It can be seen that the method proposed in the present invention effectively improves the accuracy of small target detection. At the same time, we also visualized the confidence maps of some detection results, as Figures 8 to 10 shown. It can be seen that after using the method of this application, the confidence of each target detection algorithm at the image target position has been significantly improved, which makes the target easier to be detected.
[0086] 4. Ablation study
[0087] The effectiveness of each module of the model was verified. For this purpose, the following variants of the model were considered
[0088] Model 1: Remove the additional dataset and only use the SAITD dataset for training;
[0089] Model 2: On the basis of Model 1, replace all basic modules with residual dense groups;
[0090] Model 3: On the basis of Model 1, replace all basic modules with central difference residual dense groups
[0091] Model 4: On the basis of Model 1, replace the feature alignment module with the feature alignment module in MocoPnet.
[0092] Table 2 Performance comparison of each variant of the model
[0093]
[0094] By comparing the results of this model with those of Model 1, we can see that after adding additional training data, the performance of the model on Hui and Anti-UAV data has improved to a certain extent, indicating that adding additional training data can increase the generalization of the model; by comparing the results of Model 1 with those of Model 2 and Model 3, we can see that the basic module proposed in this model is lightweight and efficient; similarly, by comparing the results of Model 1 with those of Model 4, we can see that the feature alignment module proposed in this model also has better performance.
[0095] At the same time, we also analyzed the performance, parameter count, and computational complexity of the model under different frame numbers, and selected the best comparison method MocoPnet in the above table for comparison. The comparison results are shown in Table 3. Among them, all models were trained only using the SAITD dataset. It can be seen that as the number of model frames increases, the performance of the model increases, which shows that multi-frame input helps to supplement more information. In addition, the computational degree and parameter count of our model are less than half of MocoPnet, which fully demonstrates the superiority of the proposed method.
[0096] Table 3 Performance and complexity comparison between this model and MocoPnet under different frame number inputs
[0097]
[0098]
[0099] In summary, the present invention constructs a progressive medium- and long-wave infrared video enhancement framework suitable for infrared video quality enhancement based on modules such as feature alignment network and hybrid residual dense group based on attention mechanism. The feature alignment module based on attention mechanism first extracts the features of the image sequence, and groups and fuses the features of different frames through time information so as to enhance the features within each group, thereby effectively reducing the computational complexity of the model and keeping its complexity stable when processing multiple frames of images. The small parameter-based efficient hybrid residual dense block for infrared images can efficiently extract key features at a lower parameter amount. In addition, the training strategy based on auxiliary data improves the generalization ability of the model, making it better adapted to different application scenarios. We conducted extensive ablation experiments and compared them with current mainstream methods. The performance of the proposed network on the public SIATD, Hui and Anti-UAV datasets shows that the method is significantly better than the pure model-driven method and the pure data-driven network. And the model also effectively verifies the advanced nature of our algorithm in infrared small target detection.
Claims
1. A video frame image enhancement method based on hybrid residual dense blocks, characterized by: The following steps are involved: Constructing an image enhancement model of infrared video frames, wherein the image enhancement model of infrared video frames includes a feature mapping module, a feature grouping and fusion module, a feature enhancement module, a feature alignment module, a feature reconstruction module and an image output module; A sample set of an image enhancement model is constructed using continuous infrared video frames, and the image enhancement model is trained using the sample set to obtain a trained image enhancement model; The trained image enhancement model is used to enhance the continuous infrared video frames to obtain the enhanced results.
2. The video frame image enhancement method based on hybrid residual dense blocks according to claim 1, characterized in that: When constructing a sample set of the image enhancement model using continuous infrared video frames, the constructed sample set includes multiple samples; The sample feature of each sample adopts continuous infrared video frames of a first resolution, and the label of each sample in the sample set adopts continuous infrared video frames of a second resolution as a label, and the first resolution is smaller than the second resolution; In each sample, the sample feature and the sample label are infrared video frames of the same object, the sample feature and the label have the same number of frames, and each infrared video frame in the sample feature has a corresponding infrared video frame in the label; After being trained, the image enhancement model is used to enhance continuous infrared video frames of a first resolution into continuous infrared video frames of a second resolution, thereby improving the resolution.
3. The video frame image enhancement method based on hybrid residual dense blocks according to claim 1, characterized in that: The feature mapping module adopts a residual dense block commonly used in infrared super-resolution and maps continuous infrared video frames into multi-frame features through convolution operations.
4. The method for video frame image enhancement based on hybrid residual dense blocks according to claim 3, characterized in that: The feature grouping and fusion module groups and fuses the input multi-frame features according to the temporal distance of the input multi-frame features using the residual dense block commonly used in infrared super-resolution: First, the first frame feature is selected from the input multi-frame features, and then the time distance between each frame of the input multi-frame features and the first frame feature is calculated; Preset multiple continuous and non-intersecting time distance intervals, map each frame of the multi-frame features to different time distance intervals according to the time distance between each frame and the first frame feature, and take the features in each time distance interval as a group of features; For each set of features, the residual dense block commonly used in infrared super-resolution is used for fusion.
5. The method for video frame image enhancement based on hybrid residual dense blocks according to claim 4, characterized in that: The feature enhancement module is used to enhance each set of fused features respectively; The feature enhancement module adopts a central difference mixed residual dense block obtained by improving the residual dense block.
6. The method for video frame image enhancement based on hybrid residual dense blocks according to claim 1, characterized in that: The feature alignment module is used to align each set of features and splice them along the channel dimension to achieve feature aggregation; The mathematical expression of the feature alignment module is: F a =A(F n ,F ref )), Among them, F n Represents the features of the current frame, F ref The core goal of the feature alignment module is to use the reference frame F ref For the current frame feature F n Transform it so that it is the same as F ref Spatially consistent alignment feature F a ; The features of the current frame refer to the features of the frame on which image enhancement is being performed; the features of the reference frame refer to the features of the other frames in the group except the current frame.
7. The method for video frame image enhancement based on hybrid residual dense blocks according to claim 6, characterized in that: The feature alignment module, The process of achieving feature alignment includes: For the reference frame F ref Perform feature mapping to obtain its key features For the current frame feature F n Perform feature mapping to obtain query features and value characteristics The above three features are processed in blocks using a sliding window operation, and local attention operations are performed based on the feature blocks.
8. The method for video frame image enhancement based on hybrid residual dense blocks according to claim 1, characterized in that: The feature reconstruction module is used to perform feature reconstruction based on the aggregated features to obtain a high-resolution reconstructed image which is outputted externally by the image output module; The feature reconstruction module adopts a lightweight mixed residual dense block obtained by improving the residual dense block.
9. The method for video frame image enhancement based on hybrid residual dense blocks according to claim 5 or 8, characterized in that: The lightweight mixed residual dense block and the central difference mixed residual dense block are constructed as follows: The residual dense block consists of a plurality of convolutional layers with a convolutional kernel size of 3 and a convolutional layer with a convolutional kernel size of 1; Based on the residual dense block, a convolution with a kernel size of 1 is first added between each convolution operation to perform feature dimensionality reduction; In order to expand the receptive field of feature extraction, a channel self-attention mechanism is added between the residual operation and the dense operation, so that the module has the ability to perceive global information, thereby obtaining a lightweight mixed residual dense block; On the basis of the lightweight mixed residual dense block, the convolution operation is replaced by the central difference convolution, and the central difference mixed residual dense block is obtained.