Multi-image super-resolution reconstruction method, device, equipment, medium and product

By introducing the inter-frame and inter-window attention mechanisms into the multi-image super-resolution reconstruction model and combining it with the VMamba model for feature fusion and reconstruction, the problems of loss of reconstructed image details and high computational complexity in traditional methods are solved, achieving higher quality image reconstruction.

CN120707390AActive Publication Date: 2025-09-26BEIHANG UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511211710.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-28
Publication Date
2025-09-26
Estimated Expiration
2045-08-28

AI Technical Summary

Technical Problem

Traditional image super-resolution reconstruction methods have problems such as loss of reconstructed image details and high computational complexity in multi-image super-resolution reconstruction, resulting in unsatisfactory reconstructed image quality.

Method used

A multi-image super-resolution reconstruction model is adopted, combined with the inter-frame staggered attention mechanism and the inter-window staggered attention mechanism with overlapping areas, and the VMamba model is used for feature fusion and reconstruction, replacing the Swin Transformer Block to improve image quality and reduce computational complexity.

Benefits of technology

The quality of the reconstructed image is improved, the problem of detail loss is solved, and the computational complexity of the model is reduced, achieving more efficient image reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120707390A_ABST
    Figure CN120707390A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-image super-resolution reconstruction method and device, equipment, a medium and a product, and relates to the field of image super-resolution reconstruction, and the method comprises the steps: building a multi-image super-resolution reconstruction model based on an inter-frame interlaced attention mechanism, an inter-window interlaced attention mechanism with an overlapping region, and a VMama model; the inter-frame staggered attention mechanism comprises N + 1 inter-frame attention mechanisms, and the output ends of the first N inter-frame attention mechanisms are all connected with the input end of the (N + 1) th inter-frame attention mechanism; the inter-window staggered attention mechanism with the overlapped area is that Swin Transform Block is replaced by ST Block, and the ST Block comprises a first convolution and a second convolution; the first input end of the second convolution is connected with the output end of the first convolution; the second input end of the second convolution is connected with the second input end of the first convolution. The method can improve the quality of the reconstructed image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image super-resolution reconstruction, and in particular to a multi-image super-resolution reconstruction method, apparatus, device, medium and product. Background Art

[0002] Image super-resolution reconstruction is a computer vision technique that improves low-resolution images to high-resolution ones. It has extensive applications in medical imaging, satellite remote sensing, and image enhancement. Compared to single-image super-resolution reconstruction, multi-image super-resolution reconstruction captures multiple frames of the same scene and leverages sub-pixel displacement information within these frames, resulting in superior super-resolution reconstruction.

[0003] Traditional image super-resolution reconstruction methods use convolutional neural networks as their foundation to build multi-image super-resolution reconstruction models. However, the resulting reconstructed images lose detail, resulting in suboptimal reconstructed image quality. Traditional image super-resolution reconstruction methods use Transformer-based multi-image super-resolution models that can more fully utilize global information. However, as the number of images and image resolution increase, the model's computational complexity increases significantly, making it difficult to train and resulting in suboptimal reconstructed image quality.

[0004] In summary, there is a need for a multi-image super-resolution reconstruction method that can improve the quality of reconstructed images. Summary of the Invention

[0005] The purpose of this application is to provide a multi-image super-resolution reconstruction method, device, equipment, medium and product, which can improve the quality of reconstructed images.

[0006] To achieve the above objectives, this application provides the following solutions.

[0007] In the first aspect, the present application provides a multi-image super-resolution reconstruction method, including: constructing a multi-image super-resolution reconstruction model; the multi-image super-resolution reconstruction model includes: a feature extraction module, a feature alignment module and a feature fusion and reconstruction module connected in sequence; the feature extraction module is constructed based on an inter-frame staggered attention mechanism and an inter-window staggered attention mechanism with overlapping areas; the inter-frame staggered attention mechanism includes N+1 inter-frame attention mechanisms, wherein the input ends of the first N inter-frame attention mechanisms are used to input the image to be reconstructed, and the output ends of the first N inter-frame attention mechanisms are all connected to the input end of the N+1th inter-frame attention mechanism; the inter-window staggered attention mechanism with overlapping areas is to replace the SwinTransformer Block in the Swin Transformer with an ST Block, and the ST Block includes a first convolution and a second convolution; the first input end of the second convolution is connected to the output end of the first convolution; the second input end of the second convolution is connected to the second input end of the first convolution, which is the first input end of the ST Block; the first input end of the first convolution is the second input end of the ST Block; the feature fusion and reconstruction module is constructed based on the VMamba model.

[0008] Multiple frames of images to be reconstructed are input into the multi-image super-resolution reconstruction model to obtain reconstructed images.

[0009] Optionally, the multi-image super-resolution reconstruction model further includes: a first upsampling layer, the input end of the first upsampling layer is used to input multiple frames of images to be reconstructed, and the output end of the first upsampling layer is connected to the second input end of the feature fusion and reconstruction module.

[0010] Optionally, the feature extraction module specifically includes: an optical flow prediction network and a composite interlaced attention subsystem; the input end of the optical flow prediction network and the input end of the composite interlaced attention subsystem are both used to input multiple frames of images to be reconstructed, and the output end of the optical flow prediction network and the output end of the composite interlaced attention subsystem are both connected to the input end of the feature alignment module.

[0011] The composite interleaved attention subsystem includes: multiple composite interleaved attention modules connected in sequence; the composite interleaved attention module includes: a shallow neural network, a first inter-frame interleaved attention mechanism, an inter-window interleaved attention mechanism with overlapping areas, a first merging operation, a first residual network and a second inter-frame interleaved attention mechanism.

[0012] The input end of the shallow neural network is the input end of the composite interleaved attention module; the output end of the shallow neural network is respectively connected to the input end of the first inter-frame interleaved attention mechanism and the input end of the inter-window interleaved attention mechanism with overlapping areas, the output end of the shallow neural network, the output end of the first inter-frame interleaved attention mechanism and the output end of the inter-window interleaved attention mechanism with overlapping areas are all connected to the input end of the first merging operation, the output end of the first merging operation is connected to the input end of the first residual network, and the output end of the first residual network is connected to the input end of the second inter-frame interleaved attention mechanism; the output end of the second inter-frame interleaved attention mechanism is the output end of the composite interleaved attention module.

[0013] Optionally, the first residual network includes: a second merging operation and a first layer normalization unit and a multilayer perceptron connected in sequence.

[0014] The input end of the first layer normalization unit is connected to the output end of the first merging operation, the input end of the second merging operation is connected to the output end of the multilayer perceptron and the output end of the first merging operation respectively, and the output end of the second merging operation is connected to the input end of the second inter-frame staggered attention mechanism.

[0015] Optionally, the feature fusion and reconstruction module specifically includes: a residual VMamba subsystem and a third merging operation; the input end of the residual VMamba subsystem is connected to the output end of the feature alignment module; the output end of the residual VMamba subsystem and the output end of the first upsampling layer are both connected to the input end of the third merging operation.

[0016] The residual VMamba subsystem includes: multiple residual VMamba modules connected in sequence; the residual VMamba module includes: a second residual network, a third residual network and a second upsampling layer connected in sequence; the second residual network is constructed based on the VMamba model; the input end of the second residual network is the input end of the residual VMamba module, and the output end of the second upsampling layer is the output end of the residual VMamba module.

[0017] Optionally, the third residual network includes: a first skip connection operation, a fourth merging operation, and a second layer normalization unit, a first convolutional layer, and a channel attention network connected in sequence; the channel attention network has a compression mechanism.

[0018] The input end of the second layer normalization unit and the input end of the first skip connection operation are both connected to the output end of the second residual network; the output end of the first skip connection operation and the output end of the channel attention network are both connected to the input end of the fourth merging operation, and the output end of the fourth merging operation is connected to the input end of the second upsampling layer.

[0019] In the second aspect, the present application provides a multi-image super-resolution reconstruction device, including: a model construction module for constructing a multi-image super-resolution reconstruction model; the multi-image super-resolution reconstruction model includes: a feature extraction module, a feature alignment module and a feature fusion and reconstruction module connected in sequence; the feature extraction module is constructed based on an inter-frame staggered attention mechanism and an inter-window staggered attention mechanism with overlapping areas; the inter-frame staggered attention mechanism includes N+1 inter-frame attention mechanisms, wherein the input ends of the first N inter-frame attention mechanisms are used to input the image to be reconstructed, and the output ends of the first N inter-frame attention mechanisms are all connected to the input end of the N+1th inter-frame attention mechanism; the inter-window staggered attention mechanism with overlapping areas is to replace the Swin Transformer Block in the Swin Transformer with an ST Block, and the ST Block includes a first convolution and a second convolution; the first input end of the second convolution is connected to the output end of the first convolution; the second input end of the second convolution is connected to the second input end of the first convolution, which is the first input end of the ST Block; the first input end of the first convolution is the second input end of the STBlock; the feature fusion and reconstruction module is constructed based on the VMamba model.

[0020] The reconstruction module is used to input multiple frames of images to be reconstructed into the multi-image super-resolution reconstruction model to obtain reconstructed images.

[0021] In a third aspect, the present application provides a computer device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement any of the multi-image super-resolution reconstruction methods described above.

[0022] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the multi-image super-resolution reconstruction methods described above.

[0023] In a fifth aspect, the present application provides a computer program product, comprising a computer program, which, when executed by a processor, implements any of the multi-image super-resolution reconstruction methods described above.

[0024] According to the specific embodiments provided by the present application, the present application has the following technical effects: the present application provides a multi-image super-resolution reconstruction method, apparatus, equipment, medium and product. The traditional image super-resolution reconstruction method constructs a multi-image super-resolution reconstruction model based on a convolutional neural network. Since the traditional convolution kernel of the convolutional neural network is difficult to capture the global features of the image and it is difficult to model the long-range dependencies between multiple frames, the cross-scale information interaction is insufficient, resulting in the loss of reconstruction details. The feature extraction module in the multi-image super-resolution reconstruction model of the present application is based on the inter-frame staggered attention mechanism and the inter-window staggered attention mechanism with overlapping areas. The inter-frame staggered attention mechanism is set to include N+1 inter-frame attention mechanisms. The input ends of the first N inter-frame attention mechanisms are used to input the image to be reconstructed, and the output ends of the first N inter-frame attention mechanisms are all connected to the input end of the N+1 inter-frame attention mechanism. Because this residual-like network structure can pass more information to the next module, it can achieve the effect of better preservation of useful information. The SwinTransformer Block in the Swin Transformer is replaced with ST Block obtains an inter-window staggered attention mechanism with overlapping areas, so that the inter-window staggered attention mechanism with overlapping areas can pay attention to features outside the window of interest, so the effect of paying attention to offset features can be achieved, and the inter-frame attention mechanism and the inter-window staggered attention mechanism with overlapping areas are used in combination to fully explore the features of each frame image that are conducive to multi-image super-resolution reconstruction. The VMamba model can fully capture the long-distance dependency of features, and the feature fusion and reconstruction module is constructed based on the VMamba model, which can make full use of the multi-frame image features extracted by the feature extraction module. Therefore, the multi-image super-resolution reconstruction model provided in this application can solve the above-mentioned problem of loss of reconstruction details, and the VMamba model is a linear structure that can maintain linear complexity. Compared with the model built based on Transformer, the computational complexity will be reduced, thereby solving the problem of difficulty in training, so this application can improve the quality of the reconstructed image. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0026] Figure 1 This is a flowchart of a multi-image super-resolution reconstruction method provided in one embodiment of the present application.

[0027] Figure 2 Schematic diagram of the inter-frame staggered attention mechanism structure.

[0028] Figure 3 Schematic diagram of the inter-window staggered attention mechanism with overlapping areas.

[0029] Figure 4 Schematic diagram of the structure of the multi-image super-resolution reconstruction model.

[0030] Figure 5 Schematic diagram of the workflow of the composite interleaved attention module.

[0031] Figure 6 Schematic diagram of the feature alignment module.

[0032] Figure 7 Schematic diagram of the workflow of the residual VMamba module.

[0033] Figure 8 This is a super-resolution effect diagram of the multi-image super-resolution reconstruction model provided in this application on an artificially synthesized dataset.

[0034] Figure 9 This is a diagram of the super-resolution effect of the multi-image super-resolution reconstruction method provided in this application on a real-world dataset.

[0035] Figure 10 A schematic diagram of the structure of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0036] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0037] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0038] In recent years, the multi-image super-resolution reconstruction method based on deep learning has achieved more accurate and detailed super-resolution effects by learning the mapping relationship between high-resolution and low-resolution images and learning the implicit reconstruction process through an end-to-end model. Based on this, the present application provides a multi-image super-resolution reconstruction method. In an exemplary embodiment, Figure 1 As shown, the multi-image super-resolution reconstruction method includes steps 201 and 202.

[0039] Step 201: Construct a multi-image super-resolution reconstruction model; Figure 4As shown, the multi-image super-resolution reconstruction model includes: a feature extraction module, a feature alignment module and a feature fusion and reconstruction module connected in sequence; the feature extraction module is constructed based on the inter-frame staggered attention mechanism and the inter-window staggered attention mechanism with overlapping areas; as shown Figure 2 As shown in , the inter-frame staggered attention mechanism includes N+1 inter-frame attention mechanisms, where the input ends of the first N inter-frame attention mechanisms are used to input the image to be reconstructed, and the output ends of the first N inter-frame attention mechanisms are connected to the input end of the N+1th inter-frame attention mechanism; the inter-window staggered attention mechanism with overlapping areas is to replace the Swin Transformer Block in the Swin Transformer with the ST Block, as shown in Figure 3 As shown, the ST Block includes a first convolution and a second convolution; the first input of the second convolution is connected to the output of the first convolution; the second input of the second convolution is connected to the second input of the first convolution, which is the first input of the STBlock, used to input the K / V matrix with overlapping areas; the first input of the first convolution is the second input of the ST Block, used to input the Q sampling area; the feature fusion and reconstruction module is constructed based on the VMamba model. The principle of the inter-window staggered attention mechanism with overlapping areas is as follows Figure 3 As shown in Figure 2, by sampling the K / V matrices with overlapping areas, the Q sampling area is allowed to establish connections with areas outside it, providing attention to unaligned features between multiple frames.

[0040] Step 202: Input multiple frames of images to be reconstructed into the multi-image super-resolution reconstruction model to obtain reconstructed images.

[0041] In another exemplary embodiment of the present application, the ST Block further includes: a layer normalization and a fully connected layer connected in sequence, and the layer normalization is connected to the output end of the second convolution.

[0042] In another exemplary embodiment of the present application, the feature extraction module specifically includes: an optical flow prediction network and a composite interleaved attention subsystem; the input end of the optical flow prediction network and the input end of the composite interleaved attention subsystem are both used to input multiple frames of images to be reconstructed, and the output end of the optical flow prediction network and the output end of the composite interleaved attention subsystem are both connected to the input end of the feature alignment module. The composite interleaved attention subsystem includes: a plurality of composite interleaved attention modules connected in sequence; the composite interleaved attention module includes: a shallow neural network, a first inter-frame interleaved attention mechanism, an inter-window interleaved attention mechanism with overlapping areas, a first merging operation, a first residual network, and a second inter-frame interleaved attention mechanism. The input end of the shallow neural network is the input end of the composite interleaved attention module; the output end of the shallow neural network is respectively connected to the input end of the first inter-frame interleaved attention mechanism and the input end of the inter-window interleaved attention mechanism with overlapping areas, the output end of the shallow neural network, the output end of the first inter-frame interleaved attention mechanism and the output end of the inter-window interleaved attention mechanism with overlapping areas are all connected to the input end of the first merging operation, the output end of the first merging operation is connected to the input end of the first residual network, and the output end of the first residual network is connected to the input end of the second inter-frame interleaved attention mechanism; the output end of the second inter-frame interleaved attention mechanism is the output end of the composite interleaved attention module.

[0043] In another exemplary embodiment of the present application, the pre-trained SpyNet network is used as the optical flow prediction network to perform optical flow prediction and obtain the optical flow of the three dimensions of the i-th frame image. .

[0044] In another exemplary embodiment of the present application, the feature extraction module also includes two convolutional layers, and multiple frames of images to be reconstructed are input into the optical flow prediction network through one of the convolutional layers, and multiple frames of images to be reconstructed are input into the composite interleaved attention subsystem through another convolutional layer.

[0045] In another exemplary embodiment of the present application, the composite staggered attention module further includes: a third-layer normalization unit, and the output end of the shallow neural network is respectively connected to the input end of the first inter-frame staggered attention mechanism, the input end of the inter-window staggered attention mechanism with overlapping areas, and the input end of the first merging operation through the third-layer normalization unit.

[0046] In another exemplary embodiment of the present application, the shallow neural network is the second convolutional layer.

[0047] In another exemplary embodiment of the present application, the first residual network includes: a second merging operation and a first layer normalization unit and a multilayer perceptron connected in sequence. The input end of the first layer normalization unit is connected to the output end of the first merging operation, the input end of the second merging operation is connected to the output end of the multilayer perceptron and the output end of the first merging operation respectively, and the output end of the second merging operation is connected to the input end of the second inter-frame staggered attention mechanism.

[0048] The composite interleaved attention module first uses a shallow neural network to perform preliminary feature extraction on each input frame. It then applies the first inter-frame interleaved attention mechanism and the inter-window interleaved attention mechanism with overlapping areas to the shallow features obtained, and then merges the obtained features. It then uses the first layer normalization unit and multi-layer perceptron, and adopts the residual network to obtain new features. Finally, the second inter-frame interleaved attention mechanism is used to obtain the final features of this module, such as Figure 5 As shown in the figure, the working process of the composite interleaved attention module is as follows: for the input The frame size is First, a convolutional neural network (the second convolutional layer) is used to extract shallow features from each low-resolution image to obtain a dimension of Features . Perform layer normalization on shallow features , and then parallel large window Swin Transdormer operation LWT () (i.e., inter-window interleaved attention mechanism with overlapping areas) and inter-frame interleaved attention operation CFA () to obtain the feature and , the formula is and .

[0049] The features 、 and Combined, the formula is , Represents the control weight of the output feature of the CFA operation, and then performs layer normalization, and then inputs the feature into a multi-layer perceptron. The formula is expressed as .

[0050] Obtained features Features obtained by parallel operation After merging, perform another inter-frame attention operation to obtain features , the formula is .

[0051] In another exemplary embodiment of the present application, the feature alignment module is an optical flow guided alignment module for aligning multiple frames of images in the feature domain. Figure 6 As shown in the figure, this process mainly uses the Flow-Guided Deformable Convolutional Network (FGDCN) guided by optical flow. The specific working process is: the optical flow obtained by the optical flow prediction network is used to correct the image features of each frame. The formula is expressed as ,in, 、 、 and They respectively represent the image features after correction of the i-th frame image, the optical flow prediction network, the features of the i-th frame image output by the composite interleaved attention subsystem, and the optical flow of the i-th frame image obtained by the optical flow prediction network.

[0052] Then, the corrected features Features of the reference image (the reference image is one of the multiple frames and is manually selected) , the optical flow of the corresponding frame Merge, using convolutional neural networks Get offset feature , the formula is .

[0053] Finally, the deformable convolutional neural network DCN () is used to align the image features of each frame. The formula is expressed as , Represents the image features of the i-th frame after alignment.

[0054] like Figure 4 As shown, in another exemplary embodiment of the present application, the multi-image super-resolution reconstruction model further includes: The first upsampling layer, wherein the input end of the first upsampling layer is used to input multiple frames of images to be reconstructed, and the output end of the first upsampling layer is connected to the second input end of the feature fusion and reconstruction module.

[0055] In another exemplary embodiment of the present application, the multi-image super-resolution reconstruction model further includes: a third convolutional layer, and multiple frames of images to be reconstructed are input into the first upsampling layer through the third convolutional layer.

[0056] In another exemplary embodiment of the present application, the feature fusion and reconstruction module specifically includes: a residual VMamba subsystem and a third merging operation; the input end of the residual VMamba subsystem is connected to the output end of the feature alignment module; the output end of the residual VMamba subsystem and the output end of the first upsampling layer are both connected to the input end of the third merging operation.

[0057] The residual VMamba subsystem comprises multiple sequentially connected residual VMamba modules. Each residual VMamba module comprises a second residual network, a third residual network, and a second upsampling layer, all connected in sequence. The second residual network is constructed based on the VMamba model. The input of the second residual network serves as the input of the residual VMamba module, and the output of the second upsampling layer serves as the output of the residual VMamba module. The VMamba model is derived from the Vision State-Space Model (VSSM) applied to deep learning. The VMamba model processes features using a combination of linear layers and SiLU activation layers, as well as linear layers, depthwise separable convolutions, SiLU activation layers, 2D-SSM layers, and linear layers. The features obtained from these two approaches are then combined and output through a linear layer. The second upsampling layer is a pixel shuffle.

[0058] In another exemplary embodiment of the present application, the third residual network includes: a first skip connection operation, a fourth merging operation, and a second layer normalization unit, a first convolutional layer, and a channel attention network connected in sequence; the channel attention network has a compression mechanism. The input of the second layer normalization unit and the input of the first skip connection operation are both connected to the output of the second residual network; the output of the first skip connection operation and the output of the channel attention network are both connected to the input of the fourth merging operation, and the output of the fourth merging operation is connected to the input of the second upsampling layer.

[0059] In another exemplary embodiment of the present application, the second residual network includes: a second skip connection operation, a fifth merging operation, and a fourth-layer normalization unit and a VMamba model connected in sequence; the input end of the fourth-layer normalization unit and the input end of the second skip connection operation are the input end of the second residual network; the output end of the second skip connection operation and the output end of the VMamba model are connected to the input end of the fifth merging operation.

[0060] The residual VMamba module of this application first organizes the features output by the feature alignment module through the fourth layer normalization unit, and then uses the VMamba model to process the features in parallel. The features are then merged and passed through the linear layer as the output features of the VMamba model. The features are then passed into the second layer normalization unit, the first convolutional layer, and a channel attention network with compression to achieve feature fusion. Finally, the fused features are used to achieve super-resolution reconstruction through PixelShuffle, as shown in the following example. Figure 7 As shown, the specific working process of the residual VMamba module is: first, Features after frame image alignment Arranged as Form. Features After layer normalization, it is input into the VMamba model. This process adopts the form of residual network, through a learnable parameter Control the residual ratio, the formula is expressed as: , Indicates the second skip connection operation. Inside the VMamba model, two parallel branches are used to process features. One branch uses a linear layer. and SiLU activation function , the formula is: ; The other branch uses linear layers and depth-wise separable convolution , SiLU activation function, two-dimensional state space model And layer normalization, the formula is expressed as: The output features F1 and F2 are combined with the features of the residual control, and the formula is expressed as: .

[0061] Will and Merge , the merged features The final feature fusion is completed through the channel attention network with compression mechanism. Specifically, the layer normalization operation, convolution operation Conv() and channel attention network CA() with compression mechanism are performed in sequence to obtain the fused feature map ,in, Represents the first skip connection operation. Finally, the fused feature map is reconstructed at high resolution through PixelShuffle to obtain the reconstructed result.

[0062] In another exemplary embodiment of the present application, the multi-image super-resolution reconstruction model provided by the present application is based on deep learning technology and adopts a supervised learning training method to achieve end-to-end multi-image super-resolution reconstruction.

[0063] In order to demonstrate the effectiveness of the multi-image super-resolution reconstruction model provided in this application, an embodiment is provided for testing the multi-image super-resolution reconstruction model provided in this application and multiple well-known multi-image super-resolution reconstruction models using a test dataset. Specifically, the test dataset is a public dataset provided by the NTIRE2022 Burst Super-Resolution Challenge. This dataset provides 300 artificially synthesized low-resolution images and 882 real-world low-resolution image blocks for testing. Well-known multi-image super-resolution reconstruction models are currently more advanced methods, specifically including: HighResNet model, MFIR model, DBSR model, BIPNet model, BSRT model, EBSR model, Burstormer model and AFCNet model.

[0064] The test results provided in this embodiment are shown in Table 1. Compared with these known models, when training the multi-image super-resolution reconstruction model provided in this application, the image block size used is , the image patch size for training other known models is The image block size of the training model has a positive correlation with the effect of the model, but the multi-image super-resolution reconstruction model provided by this application still has higher evaluation indicators than other methods. The super-resolution effect of the multi-image super-resolution reconstruction model provided by this application on the artificial synthetic dataset is as follows: Figure 8 As shown in the figure, the super-resolution reconstruction effect on the real-world dataset is as follows Figure 9 shown.

[0065] Table 1 Results of the known multi-image super-resolution reconstruction model and the multi-image super-resolution reconstruction model provided by this application on the NTIRE2022 Burst Super-Resolution Challenge dataset

[0066] The related art uses convolutional neural network structure modeling, and traditional convolutional neural networks are used in the feature extraction process. A feature-enhanced pyramid-structured deformable convolutional neural network is used in the feature alignment process. In the reconstruction process, features at different levels are connected over long distances to improve the utilization of global information in the super-resolution reconstruction process. Due to the inadequacy of the convolution kernel in the convolutional neural network for global feature extraction, the image quality of high-resolution reconstruction is limited. This application solves the above problem by setting an inter-frame interleaved attention mechanism and an inter-window interleaved attention mechanism with overlapping areas, which can fully exploit the features of each frame image that are conducive to multi-image super-resolution reconstruction.

[0067] In the related art, Swin Transformer is used in the feature extraction and reconstruction process to obtain long-distance dependencies between global and local information. In the feature alignment process, a deformable convolutional neural network guided by optical flow has a problem of insufficient feature alignment due to the fixed-size window of the deformable convolution. This application solves the above problem by using the VMamba model with composite sampling in the feature fusion and reconstruction module to further focus on the features around the pixels of interest, thereby improving the reconstruction effect.

[0068] A novel Transformer structure is used in related technologies to more effectively align multi-frame features by designing a hierarchical multi-scale structure. In order to increase the connection between inter-frame information, this method also designs a cyclic sampling module based on a convolutional neural network for multi-frame images. This type of method cyclically samples every two frames of multi-frame image features, but still cannot fully pay attention to the global information problem of multiple frames. This application solves the above problem by setting an inter-frame staggered attention mechanism to sample multi-frame image features, thereby improving the reconstruction effect.

[0069] Based on the same inventive concept, embodiments of the present application also provide a multi-image super-resolution reconstruction device for implementing the multi-image super-resolution reconstruction method described above. The solution provided by this device is similar to the solution described in the method described above. Therefore, the specific limitations in one or more embodiments of the multi-image super-resolution reconstruction device provided below can be found in the above-mentioned limitations on the multi-image super-resolution reconstruction method and will not be further elaborated here.

[0070] In an exemplary embodiment, a multi-image super-resolution reconstruction device is provided, comprising: a model construction module for constructing a multi-image super-resolution reconstruction model; the multi-image super-resolution reconstruction model comprises: a feature extraction module, a feature alignment module and a feature fusion and reconstruction module connected in sequence; the feature extraction module is constructed based on an inter-frame staggered attention mechanism and an inter-window staggered attention mechanism with overlapping areas; the inter-frame staggered attention mechanism includes N+1 inter-frame attention mechanisms, wherein the input ends of the first N inter-frame attention mechanisms are used to input the image to be reconstructed, and the output ends of the first N inter-frame attention mechanisms are all connected to the input end of the N+1th inter-frame attention mechanism; the inter-window staggered attention mechanism with overlapping areas is to replace the Swin Transformer Block in the Swin Transformer with an ST Block, and the ST Block comprises a first convolution and a second convolution; the first input end of the second convolution is connected to the output end of the first convolution; the second input end of the second convolution is connected to the second input end of the first convolution, and is the first input end of the ST Block; the first input end of the first convolution is the second input end of the ST Block; the feature fusion and reconstruction module is constructed based on the VMamba model.

[0071] The reconstruction module is used to input multiple frames of images to be reconstructed into the multi-image super-resolution reconstruction model to obtain reconstructed images.

[0072] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal. The internal structure diagram thereof may be as follows: Figure 10 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store multi-image super-resolution reconstruction data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a multi-image super-resolution reconstruction method is implemented.

[0073] Those skilled in the art will understand that Figure 10The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present application and does not constitute a limitation on the computer device to which the solution of the present application is applied. A specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement. In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the above-mentioned method embodiments when executing the computer program.

[0074] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, which implements the above-mentioned method embodiments when executed by a processor.

[0075] In an exemplary embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the above method embodiments are implemented.

[0076] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0077] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0078] The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may include, but are not limited to, general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic units, data processing logic units based on quantum computing, and the like.

[0079] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0080] This document uses specific examples to illustrate the principles and implementation methods of this application. The description of the above examples is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.

Claims

1. A multi-image super-resolution reconstruction method, characterized in that: The multi-image super-resolution reconstruction method comprises: Construct a multi-image super-resolution reconstruction model; the multi-image super-resolution reconstruction model includes: a feature extraction module, a feature alignment module and a feature fusion and reconstruction module connected in sequence; the feature extraction module is constructed based on an inter-frame staggered attention mechanism and an inter-window staggered attention mechanism with overlapping areas; the inter-frame staggered attention mechanism includes N+1 inter-frame attention mechanisms, wherein the input ends of the first N inter-frame attention mechanisms are used to input the image to be reconstructed, and the output ends of the first N inter-frame attention mechanisms are all connected to the input end of the N+1 inter-frame attention mechanism; the inter-window staggered attention mechanism with overlapping areas is to replace the Swin Transformer Block in the Swin Transformer with an ST Block, and the ST Block includes a first convolution and a second convolution; the first input end of the second convolution is connected to the output end of the first convolution; the second input end of the second convolution is connected to the second input end of the first convolution, and is the first input end of the ST Block; the first input end of the first convolution is the second input end of the STBlock; the feature fusion and reconstruction module is constructed based on the VMamba model; Multiple frames of images to be reconstructed are input into the multi-image super-resolution reconstruction model to obtain reconstructed images.

2. The multi-image super-resolution reconstruction method according to claim 1, characterized in that The multi-image super-resolution reconstruction model further includes: The first upsampling layer, wherein the input end of the first upsampling layer is used to input multiple frames of images to be reconstructed, and the output end of the first upsampling layer is connected to the second input end of the feature fusion and reconstruction module.

3. The multi-image super-resolution reconstruction method according to claim 1, characterized in that The feature extraction module specifically includes: An optical flow prediction network and a composite interleaved attention subsystem; the input end of the optical flow prediction network and the input end of the composite interleaved attention subsystem are both used to input multiple frames of images to be reconstructed, and the output end of the optical flow prediction network and the output end of the composite interleaved attention subsystem are both connected to the input end of the feature alignment module; The composite interleaved attention subsystem includes: a plurality of composite interleaved attention modules connected in sequence; the composite interleaved attention module includes: a shallow neural network, a first inter-frame interleaved attention mechanism, an inter-window interleaved attention mechanism with overlapping areas, a first merging operation, a first residual network, and a second inter-frame interleaved attention mechanism; The input end of the shallow neural network is the input end of the composite interleaved attention module; the output end of the shallow neural network is respectively connected to the input end of the first inter-frame interleaved attention mechanism and the input end of the inter-window interleaved attention mechanism with overlapping areas, the output end of the shallow neural network, the output end of the first inter-frame interleaved attention mechanism and the output end of the inter-window interleaved attention mechanism with overlapping areas are all connected to the input end of the first merging operation, the output end of the first merging operation is connected to the input end of the first residual network, and the output end of the first residual network is connected to the input end of the second inter-frame interleaved attention mechanism; the output end of the second inter-frame interleaved attention mechanism is the output end of the composite interleaved attention module.

4. The multi-image super-resolution reconstruction method according to claim 3, characterized in that: The first residual network includes: a second merging operation and a first layer normalization unit and a multilayer perceptron connected in sequence; The input end of the first layer normalization unit is connected to the output end of the first merging operation, the input end of the second merging operation is connected to the output end of the multilayer perceptron and the output end of the first merging operation respectively, and the output end of the second merging operation is connected to the input end of the second inter-frame staggered attention mechanism.

5. The multi-image super-resolution reconstruction method according to claim 2, characterized in that: The feature fusion and reconstruction module specifically includes: a residual VMamba subsystem and a third merging operation; the input end of the residual VMamba subsystem is connected to the output end of the feature alignment module; the output end of the residual VMamba subsystem and the output end of the first upsampling layer are both connected to the input end of the third merging operation; The residual VMamba subsystem includes: multiple residual VMamba modules connected in sequence; the residual VMamba module includes: a second residual network, a third residual network and a second upsampling layer connected in sequence; the second residual network is constructed based on the VMamba model; the input end of the second residual network is the input end of the residual VMamba module, and the output end of the second upsampling layer is the output end of the residual VMamba module.

6. The multi-image super-resolution reconstruction method according to claim 5, characterized in that: The third residual network includes: a first skip connection operation, a fourth merging operation, and a second layer normalization unit, a first convolutional layer, and a channel attention network connected in sequence; the channel attention network has a compression mechanism; The input end of the second layer normalization unit and the input end of the first skip connection operation are both connected to the output end of the second residual network; the output end of the first skip connection operation and the output end of the channel attention network are both connected to the input end of the fourth merging operation, and the output end of the fourth merging operation is connected to the input end of the second upsampling layer.

7. A multi-image super-resolution reconstruction device, characterized in that: The multi-image super-resolution reconstruction device comprises: A model construction module is used to construct a multi-image super-resolution reconstruction model; the multi-image super-resolution reconstruction model includes: a feature extraction module, a feature alignment module and a feature fusion and reconstruction module connected in sequence; the feature extraction module is constructed based on an inter-frame staggered attention mechanism and an inter-window staggered attention mechanism with overlapping areas; the inter-frame staggered attention mechanism includes N+1 inter-frame attention mechanisms, wherein the input ends of the first N inter-frame attention mechanisms are used to input the image to be reconstructed, and the output ends of the first N inter-frame attention mechanisms are all connected to the input end of the N+1 inter-frame attention mechanism; the inter-window staggered attention mechanism with overlapping areas is to replace the Swin Transformer Block in the Swin Transformer with an ST Block, and the ST Block includes a first convolution and a second convolution; the first input end of the second convolution is connected to the output end of the first convolution; the second input end of the second convolution is connected to the second input end of the first convolution, and is the first input end of the ST Block; the first input end of the first convolution is the second input end of the ST Block; the feature fusion and reconstruction module is constructed based on the VMamba model; The reconstruction module is used to input multiple frames of images to be reconstructed into the multi-image super-resolution reconstruction model to obtain reconstructed images.

8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the multi-image super-resolution reconstruction method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the multi-image super-resolution reconstruction method according to any one of claims 1 to 6 is implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the multi-image super-resolution reconstruction method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Traffic flow prediction method based on time-space diagram attention and computing device

    CN116596151A

  • Image super-resolution method and device

    CN118799181A

  • Lightweight hybrid Transform-Mama network-TransMama for super-resolution

    CN119337937A

  • VMama-based remote sensing image change detection method and system

    CN120495758A

  • Operation acceleration processing method, operation accelerator use method, and operation accelerator

    WO2023123453A1