Lightweight image super-resolution method and storage medium
By using shallow feature extraction module, deep feature extraction module, large-core dynamic self-attention module and adaptive multi-head sparse self-attention module in the image super-resolution method, the shortcomings of the existing methods in local feature extraction and global modeling are solved, and higher quality image reconstruction is achieved.
Patent Information
- Application Number
- CN202510196242.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-10
AI Technical Summary
The existing Transformer-based image super-resolution method has shortcomings in local feature extraction and global modeling, resulting in poor detail recovery and poor global consistency.
A lightweight image super-resolution method is proposed, using shallow feature extraction module and deep feature extraction module to extract low-frequency information and deep features through single-layer convolution and residual hybrid learning, and combined with large-core dynamic self-attention module and adaptive multi-head sparse self-attention module to enhance the model's feature extraction and image reconstruction capabilities.
While maintaining computing efficiency, it significantly improves image reconstruction quality, improves local detail extraction and global context modeling capabilities, and improves the performance of super-resolution tasks.
Smart Images

Figure CN120125435A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and particularly relates to a lightweight image super-resolution method and a storage medium. Background Art
[0002] Image super-resolution technology is an important image processing method. Its core objective is to reconstruct high-resolution images from low-resolution images. This technology can not only significantly improve the clarity and detail performance of images, but also provide high-quality data support for subsequent image analysis and processing. Therefore, it has broad application value in multiple fields. Traditional super-resolution methods usually rely on interpolation techniques, and speculate on missing information by performing simple mathematical operations on image pixels. The calculation process is relatively simple and the speed is relatively fast. However, the quality of the images reconstructed by such methods is limited. It is often difficult to retain rich details in the images, and there are prone to blurred edges and lost texture information, which cannot meet the actual application scenarios that require high-precision images. With the rise of deep learning, single-image super-resolution methods based on convolutional neural networks (CNNs) have gradually become the mainstream. For example, SRCNN, FSRCNN, CARN, etc. Although these methods have achieved good results in local feature extraction, CNN methods are usually better at capturing local information and have relatively weak global context modeling capabilities. This makes their performance limited when dealing with some images with complex structures and long-range dependencies.
[0003] In recent years, image super-resolution methods based on Transformer have developed rapidly. Transformer uses the self-attention mechanism and can effectively capture long-range dependencies in images. Therefore, it has unique advantages in modeling global features. For example, Vision Transformer (ViT) realizes the modeling of global features with fewer parameters through the self-attention mechanism, showing better performance than traditional CNNs, especially in capturing the global context information of images. However, existing Transformer-based methods still face some challenges: First, the limitations of the window self-attention mechanism limit the ability to extract local effective features, resulting in poor detail recovery; Second, the modeling ability of discontinuous windows is weak, making it difficult for the model to fully capture the global information in the images, thus affecting the detail reconstruction and global consistency in the super-resolution task. Summary of the Invention
[0004] To solve the above technical problems, the present invention provides a lightweight image super-resolution method and a storage medium.
[0005] To solve the above technical problems, the present invention adopts the following technical solutions:
[0006] A lightweight image super-resolution method inputs a low-resolution image into a trained super-resolution model to reconstruct a high-resolution image. The training process of the super-resolution model includes:
[0007] Obtain a low-resolution image from the high-resolution image, and form a training set with the high-resolution image and the low-resolution image;
[0008] Construct a super-resolution model, which includes a shallow feature extraction module, a deep feature extraction module, and an image reconstruction module;
[0009] Input the low-resolution image I LR into the shallow feature extraction module, and use a single-layer convolution to extract low-frequency information from the low-resolution image to obtain the shallow feature F'; 0 ;
[0010] Input the shallow feature F' 0 into the deep feature extraction module for residual mixing learning to obtain the deep feature F fin ;
[0011] Input the deep feature F fin into the image reconstruction module for image reconstruction to obtain the high-resolution image I SR ;
[0012] For a given pair of high-resolution image I HR and low-resolution image pair I LR in the training set, input the low-resolution image I LR into the super-resolution model to obtain the reconstructed high-resolution image I SR , and train the super-resolution model through the high-resolution image I HR and the reconstructed high-resolution image I SR .
[0013] In one embodiment, the obtaining of the low-resolution image from the high-resolution image specifically includes:
[0014] Obtain the corresponding low-resolution image by scaling the high-resolution image by a scale factor; use bicubic interpolation to ensure the image quality during the scaling process, and then divide the paired high-resolution images and low-resolution images into a training set and a test set according to a set ratio.
[0015] In one embodiment, the inputting of the shallow feature F' 0 into the deep feature extraction module for residual mixing learning to obtain the deep feature F fin specifically includes:
[0016] The deep feature extraction module contains N residual hybrid learning groups. After passing through the residual hybrid learning groups, the shallow features obtain deep features at different levels:
[0017]
[0018] Among them, represents the operation of the i-th residual hybrid learning group, and F′ i+1 represents the deep feature output by the i-th residual hybrid learning group;
[0019] After the shallow features are input into the deep feature extraction module and learned through N residual hybrid learning groups, the deep feature F fin is obtained by applying the residual mechanism to the single-layer convolution:
[0020] F fin = Conv(F′ N + F′ 0 );
[0021] Conv represents the convolution operation.
[0022] In one embodiment, the residual hybrid learning group consists of M large-kernel dynamic Transformer blocks and a single-layer convolution applying the residual mechanism; there is a pair of large-kernel dynamic self-attention modules and an adaptive multi-head sparse self-attention module in each large-kernel dynamic Transformer block:
[0023] F′ ij = L AMHSSA (L LKDSA (F′ i ));
[0024] Among them, F i ′ j is the feature processed by the j-th large-kernel dynamic Transformer block in the i-th residual hybrid learning group, where i = 0, 1, 2... N and j = 0, 1, 2... M; L LKDSA represents the large-kernel dynamic self-attention module, and L AMHSSA represents the adaptive multi-head sparse self-attention module.
[0025] In one embodiment, the large-kernel dynamic self-attention module dynamically extracts features within a set range through a dilated filter, uses the extracted features as the weights of the dynamic filter to process the input features to achieve local feature aggregation; and then uses a convolutional feed-forward network to improve the feature representation after attention weighting.
[0026] In one embodiment, the large-core dynamic self-attention module dynamically extracts features within a set range through a dilated filter, and uses the extracted features as the weights of the dynamic filter to process the input features to achieve local feature aggregation; then uses a convolutional feed-forward network to improve the feature representation after attention weighting, specifically including:
[0027] Input feature Y in After normalization, it passes through a group of depthwise separable convolutions to obtain a preliminary focused feature Y f :
[0028] Y f = DSConv(Y in );
[0029] Where DSConv represents depthwise convolution Depth-wiseConv and point convolution PointConv;
[0030] The preliminary focused feature is input into a large-core grouped convolution with dilated attributes for processing, and the output result size is changed:
[0031] W = Reshape(Conv(DilationDW(Y f )));
[0032] Where DilationDW represents grouped convolution with dilation rate, and W is the convolution kernel parameter dynamically learned;
[0033] The dynamic convolution is applied to the input feature Y in a channel-sharing manner in To obtain the feature Y after self-attention out :
[0034] Y out = Conv(Dynamic W (Y in ));
[0035] Where Dynamic w Represents a dynamic convolution with convolution kernel parameter W. After applying the attention mechanism to the input feature, it is necessary to deeply explore the high-weight part, that is, input Y out Into the convolutional feed-forward network to enhance the model's ability to express features, and obtain the output result Y of the large-core dynamic self-attention module l :
[0036] Y l = CFF(Y out );
[0037] Where CFF(·) represents a feed-forward neural network in convolutional form.
[0038] In one of the embodiments, after the adaptive multi-head sparse self-attention module obtains the output result Y of the large-kernel dynamic self-attention module, it passes through normalization and single-layer convolution to fuse each channel to obtain the feature Y l : p Y
[0039] = Conv(LayerNorm(Y p )); l
[0040] where LayerNorm(·) represents the normalization operation, and then depth convolution is used to filter each channel, and it is divided into three self-attention matrices Q, K, and V along the channel dimension;
[0041] Q, K, V = chunk(DWConv(Y p ));
[0042] where chunk(·) represents the chunking operation along the channel dimension, and DWConv(·) represents the depthwise separable convolution operation that performs convolution independently on each channel;
[0043] Convert the self-attention matrix size and perform the sparse operation:
[0044]
[0045] where is the multi-head attention matrix obtained through size transformation, PReLU(·) represents the parametric rectified linear unit, is the sparse attention matrix obtained through processing, and α is a learnable parameter;
[0046] Multiply the sparse attention matrix with and perform single convolution processing to obtain the global attention expression· ag :
[0047]
[0048] where @ represents matrix multiplication;
[0049] Perform depth extraction on the global attention, that is, input Y ag into the convolutional feed-forward network for processing to obtain the global feature Y g :
[0050] Y g = CFF(Y ag );
[0051] CFF(·) represents the convolutional feed-forward neural network.
[0052] In one embodiment, the depth feature F fin is input into an image reconstruction module for image reconstruction to obtain a high-resolution image I SR , which specifically includes:
[0053] I SR = L shuffle (Conv(F fin + F′ 0 ))
[0054] where L shuffle represents sub-pixel convolution for upsampling operation, and Conv represents convolution operation.
[0055] In one embodiment, the super-resolution model is trained by the high-resolution image I HR and the reconstructed high-resolution image I SR , which specifically includes:
[0056] The super-resolution model is trained using the L1 loss function, and the loss function Loss = |I HR - I SR |.
[0057] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method in any one of the embodiments are implemented.
[0058] Compared with the prior art, the beneficial technical effects of the present invention are:
[0059] 1. Lightweight dynamic hybrid self-attention model: The present invention proposes a lightweight dynamic hybrid self-attention network model, which can efficiently complete the image super-resolution task with a small number of network parameters and computational overhead, and has higher feasibility and practicality.
[0060] 2. Large-kernel dynamic self-attention module: The present invention designs a large-kernel dynamic self-attention module, which realizes the application of convolutional local self-attention by using dilated large-kernel convolution and dynamic convolution with a channel weight sharing mechanism. This module not only effectively expands the receptive field and improves the feature extraction ability, but also significantly reduces the computational complexity, and has higher efficiency and applicability.
[0061] 3. Adaptive multi-head sparse self-attention module: The present invention proposes an adaptive multi-head sparse self-attention module, which effectively improves the sparsity of attention through multi-head adaptive learning while fully mining global features. Compared with the traditional attention mechanism, this module can not only reduce redundant calculations and improve the model efficiency, but also capture global context information more accurately, thereby enhancing the expression ability and generalization performance of the model. Brief Description of the Drawings
[0062] Figure 1 This is a schematic diagram of the super-resolution network model in the present invention.
[0063] Figure 2 This is a schematic diagram of the residual hybrid learning group of the present invention.
[0064] Figure 3 This is a schematic diagram of the large kernel dynamic self-attention module of the present invention.
[0065] Figure 4 This is a schematic diagram of the adaptive multi-head sparse self-attention module of the present invention.
[0066] Figure 5 This is a schematic diagram of the convolutional feed-forward network of the present invention. Detailed Embodiment
[0067] A preferred embodiment of the present invention will be described in detail below with reference to the accompanying drawings.
[0068] The present invention provides a lightweight image super-resolution method. A low-resolution image is input into a trained super-resolution model to reconstruct a high-resolution image. The training process of the super-resolution model includes the following steps:
[0069] S1. Obtain a low-resolution image from a high-resolution image, and form a training set with the high-resolution image and the low-resolution image;
[0070] S2. Construct a super-resolution model, which includes a shallow feature extraction module, a deep feature extraction module, and an image reconstruction module;
[0071] Input the low-resolution image I LR into the shallow feature extraction module, and use a single-layer convolution to extract low-frequency information from the low-resolution image to obtain a shallow feature F′ 0 ;
[0072] Input the shallow feature F′ 0 into the deep feature extraction module for residual hybrid learning to obtain a deep feature F fin ;
[0073] Input the deep feature F fin into the image reconstruction module for image reconstruction to obtain a high-resolution image I SR ;
[0074] For a given pair of high-resolution image I HR and low-resolution image pair I LR in the training set, input the low-resolution image I LR into the super-resolution model to obtain the reconstructed high-resolution image I SR, through the high - resolution image I HR and the reconstructed high - resolution image I SR train the super - resolution model.
[0075] The present invention effectively captures low - frequency basic information through the single - layer convolutional structure of the shallow feature extraction module, retains key structural features for subsequent processing, and through the residual mixing learning architecture of the deep feature extraction module, innovatively integrates the local detail enhancement and global context modeling capabilities. This technical solution successfully solves the key bottlenecks of existing Transformer methods in local feature extraction and global modeling through structural innovation and algorithm optimization, significantly improves the image reconstruction quality while maintaining computational efficiency, and provides a new technical implementation path for high - precision super - resolution applications.
[0076] In one embodiment, obtaining the low - resolution image according to the high - resolution image in step S1 specifically includes:
[0077] By scaling the high - resolution image by a scale factor to obtain the low - resolution image at the corresponding scale; using bicubic interpolation to ensure the image quality during the scaling process, and then dividing the paired high - resolution and low - resolution images into a training set and a test set according to a set ratio. Among them, the high - resolution image is denoted as I HR , the low - resolution image is denoted as I LR , and the high - resolution image obtained after model reconstruction is denoted as I SR .
[0078] Specifically, the shallow feature extraction module consists of a single - layer convolution.
[0079] The low - resolution image I LR passes through the single - layer convolution Conv 3x3 (·) to extract the shallow feature F′ 0 :
[0080] F′ 0 =Conv 3x3 (I LR ).
[0081] In one embodiment, inputting the shallow feature F′ 0 into the deep feature extraction module for residual mixing learning to obtain the deep feature F fin , specifically includes:
[0082] The deep feature extraction module contains N residual mixing learning groups, and the shallow feature obtains different - level deep features after passing through the residual mixing learning groups:
[0083]
[0084] Among them, represents the i-th residual hybrid learning group operation, F′ i+1 represents the depth feature output by the i-th residual hybrid learning group;
[0085] After the shallow feature is input into the depth feature extraction module and undergoes learning through N residual hybrid learning groups, it is combined with a single-layer convolution Conv 3x3 (·) applies the residual mechanism to obtain the depth feature F fin :
[0086] F fin = Conv 3x3 (F′ N + F′ 0 ).
[0087] In one embodiment, the residual hybrid learning group consists of M large-kernel dynamic Transformer blocks and a single convolution applying the residual mechanism; there is a pair of large-kernel dynamic self-attention modules and an adaptive multi-head sparse self-attention module in each large-kernel dynamic Transformer block:
[0088] F′ ij = L AMHSSA (L LKDSA (F′ i ));
[0089] Among them, F′ ij is the feature processed by the j-th large-kernel dynamic Transformer block in the i-th residual hybrid learning group, i = 0, 1, 2…N, j = 0, 1, 2…M; L LKDSA represents the large-kernel dynamic self-attention module, and L AMHSSA represents the adaptive multi-head sparse self-attention module.
[0090] In one embodiment, the large-kernel dynamic self-attention module dynamically extracts features within a set range through a dilated filter, uses the extracted features as the weights of the dynamic filter to process the input features to achieve local feature aggregation; and then uses a convolutional feed-forward network to improve the feature representation after attention weighting.
[0091] Specifically, the input feature Y in after normalization passes through a group of depthwise separable convolutions to obtain a preliminary focused feature Y f :
[0092] Y f = DSConv 3x3 (Y in );
[0093] Among them, DSConv represents Depth-wise Conv and Point Conv.
[0094] Input the preliminary focused features into the grouped convolution with a large kernel with dilation attributes for processing, and change the size of the output result:
[0095] W = Reshape(Conv 1x1 (DilationDW 9x9 (Y f )));
[0096] Among them, DilationDW 9x9 represents the grouped convolution with a kernel size of 9x9 with a dilation rate. W is the convolution kernel parameter obtained by dynamic learning.
[0097] Apply the dynamic convolution to the input feature Y in a channel-sharing manner in to obtain the feature Y after self-attention out :
[0098] Y out = Conv 1x1 (Dynamic W (Y in ));
[0099] Among them, Dynamic W represents the dynamic convolution with the convolution kernel parameter W. After applying the attention mechanism to the input feature, it is necessary to deeply explore the high-weight part, that is, input Y out into the convolutional feed-forward network to enhance the model's ability to express features, and obtain the output result Y of the large-kernel dynamic self-attention module l :
[0100] Y l = CFF(Y out );
[0101] Among them, CFF(·) represents the feed-forward neural network in convolutional form.
[0102] In one embodiment, after obtaining the output result of the large-kernel dynamic self-attention module, the adaptive multi-head sparse self-attention module fuses each channel through normalization and single-layer convolution to obtain the feature Y p :
[0103] Y p = Conv 1x1 (LayerNorm(Y l ));
[0104] Among them, LayerNorm(·) represents the normalization operation. Then, each channel is filtered through depthwise convolution and split into three self-attention matrices Q, K, and V along the channel dimension:
[0105] Q, K, V = chunk(DWConv 3x3 (Y p ));
[0106] Among them, chunk(·) represents the chunking operation along the channel dimension, and DWConv 3x3 (·) represents the efficient depthwise separable convolution operation that independently convolves each channel, and the kernel size is 3x3.
[0107] Then, the size of the self-attention matrix is transformed and a sparsity operation is performed:
[0108]
[0109] Among them, is the multi-head attention matrix obtained after size transformation, PReLU(·) represents the parametric rectified linear unit function, is the sparse attention matrix obtained after processing, and α is a learnable parameter.
[0110] Finally, the sparse attention matrix is multiplied by through matrix multiplication and then processed by a single convolution to obtain the global attention expression Y ag :
[0111]
[0112] Among them, @ represents matrix multiplication.
[0113] Similarly, deep extraction of the global attention is required, that is, Y ag is input into the convolutional feed-forward network for processing to obtain the global feature Y g :
[0114] Y g = CFF(Y ag ).
[0115] In one embodiment, inputting the depth feature F fin in step S2 into the image reconstruction module for image reconstruction to obtain the high-resolution image I SR , specifically includes:
[0116] I SR = L shuffle (Conv 3x3 (F fin + F′ 0 ));
[0117] Among them, L shuffle represents sub-pixel convolution for upsampling operation.
[0118] In one embodiment, in step S2, through the high-resolution image I HR and the reconstructed high-resolution image I SR train the super-resolution model, specifically including:
[0119] Use the L1 loss function to train the super-resolution model, and the loss function Loss = |I HR - I SR |.
[0120] After completing the training of the super-resolution model, input the low-resolution images in the test dataset into the trained super-resolution network model to obtain the corresponding high-resolution image I SR .
[0121] The low-resolution image and the high-resolution image of the present invention are a set of relative concepts, and the resolution of the high-resolution image is higher than that of the low-resolution image.
[0122] It should be understood that although the steps in the flowchart of the accompanying drawings of the specification are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limitation, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings of the specification may include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same moment, but can be executed at different moments, and the execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps or stages in other steps.
[0123] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory including instructions, and the above instructions can be executed by a processor to complete the above method. The storage medium can be a computer-readable storage medium. For example, the computer-readable storage medium can be ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0124] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above-described exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention, and any reference signs in the claims should not be regarded as limiting the claims involved.
[0125] In addition, it should be understood that although this specification is described in terms of embodiments, not every embodiment only contains an independent technical solution. This narrative manner of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A lightweight image super-resolution method, which inputs a low-resolution image into a trained super-resolution model to reconstruct a high-resolution image. The training process of the super-resolution model includes: Obtain low-resolution images based on high-resolution images, and combine the high-resolution images and low-resolution images into a training set; Construct a super-resolution model, which includes a shallow feature extraction module, a deep feature extraction module, and an image reconstruction module; The low-resolution image I LR Input into the shallow feature extraction module, use a single layer of convolution to extract low-frequency information from the low-resolution image to obtain the shallow feature F′0; The shallow feature F′0 is input into the deep feature extraction module for residual hybrid learning to obtain the deep feature F fin ; The deep feature F fin Input the image reconstruction module to reconstruct the image and obtain a high-resolution image I SR ; For a given pair of high-resolution images I in the training set HR Compare with low resolution image LR , the low-resolution image I LR Input into the super-resolution model to obtain the reconstructed high-resolution image I SR , through high-resolution image I HR and the reconstructed high-resolution image I SR Train the super-resolution model.
2. The lightweight image super-resolution method according to claim 1, characterized in that: The step of obtaining a low-resolution image according to the high-resolution image specifically includes: The high-resolution image is scaled by a proportional factor to obtain a low-resolution image of the corresponding scale. Bicubic interpolation is used to ensure the image quality during the scaling process, and then the pairs of high-resolution and low-resolution images are divided into training sets and test sets according to the set ratio.
3. The lightweight image super-resolution method according to claim 1, characterized in that: The shallow feature F′0 is input into the deep feature extraction module for residual hybrid learning to obtain the deep feature F fin , including: The deep feature extraction module contains N residual mixed learning groups. After the shallow features pass through the residual mixed learning groups, they obtain deep features of different levels: in, represents the i-th residual hybrid learning group operation, F′ i+1 represents the deep features output by the i-th residual hybrid learning group; The shallow feature input deep feature extraction module is learned by N residual mixed learning groups, and then the residual mechanism is applied to the single-layer convolution to obtain the deep feature F fin : F fin =Conv(F′ N +F′0); Conv represents the convolution operation.
4. The lightweight image super-resolution method according to claim 3, characterized in that: The residual hybrid learning group consists of M large-core dynamic Transformer blocks and a single convolutional residual mechanism; each large-core dynamic Transformer block has a pair of large-core dynamic self-attention modules and an adaptive multi-head sparse self-attention module: In ij =L AMHSSA (L LKDSA (F′ i )); Among them, F′ ij is the feature processed by the j-th large-core dynamic Transformer block in the i-th residual hybrid learning group, i = 0, 1, 2 ... N, j = 0, 1, 2 ... M; L LKDSa represents the large-core dynamic self-attention module, L AMHSSA Represents the adaptive multi-head sparse self-attention module.
5. The lightweight image super-resolution method according to claim 4, characterized in that: The large-core dynamic self-attention module dynamically extracts features within a set range through a hole filter, and uses the extracted features as the weights of the dynamic filter to process the input features to achieve local feature aggregation; it then uses a convolutional feedforward network to improve the feature representation after attention weighting.
6. The lightweight image super-resolution method according to claim 5, characterized in that: The large-core dynamic self-attention module dynamically extracts features within a set range through a hole filter, uses the extracted features as the weights of the dynamic filter to process the input features to achieve local feature aggregation; and then uses a convolutional feedforward network to improve the feature representation after attention weighting, specifically including: Input feature Y in After normalization, a set of depth-separable convolutions are performed to obtain the initial focus feature Y f : AND f =DSConv(Y in ); Among them, DSConv represents depth convolution Depth-wiseConv and point convolution PointConv; The preliminary focused features are input into a large kernel group convolution with a hole attribute for processing, and the output result size is changed: W=Reshape(Conv(DilationDW(Y f ))); Among them, DilationDW represents the group convolution with dilation rate, and W is the convolution kernel parameter obtained by dynamic learning; Apply dynamic convolution to the input feature Y in a channel-sharing manner in Get the feature Y after self-attention out : Y out =Conv(Dynamic W (Y in )); Among them, Dynamic W It represents a dynamic convolution with a convolution kernel parameter of W. After the attention mechanism is applied to the input features, it is necessary to conduct a deep exploration of the high-weight part, that is, Y out Input the convolutional feedforward network to enhance the model's ability to express features and obtain the output result Y of the large-core dynamic self-attention module l : Y l =CFF(Y out ); Among them, CFF(·) represents a feed-forward neural network in the form of convolution.
7. The lightweight image super-resolution method according to claim 4, characterized in that: The adaptive multi-head sparse self-attention module obtains the output result Y of the large core dynamic self-attention module l After normalization and single-layer convolution, each channel is fused to obtain feature Y p : AND p =Conv(LayerNorm(Y l )); Among them, LayerNorm(·) represents the normalization operation, and then each channel is filtered through deep convolution and divided into three self-attention matrices Q, K, and V along the channel dimension; Q、K、V=chunk(DWConv(Y p )); Among them, chunk(·) represents the block operation along the channel dimension, and DWConv(·) represents the depth-wise separable convolution operation that performs convolution on each channel independently; Convert the self-attention matrix size and implement sparse operations: in, is the multi-head attention matrix obtained after size transformation, PReLU(·) represents a linear rectification function with learnable parameters, is the processed sparse attention matrix, and α is a learnable parameter; The sparse attention matrix and Perform matrix multiplication and single convolution to obtain the global attention expression Y ag : Among them, @ represents matrix multiplication; Deeply extract the global attention, that is, Y ag The global feature Y is obtained by processing the input convolutional feedforward network g : Y g =CFE(Y ag ); CFF(·) represents a convolutional feedforward neural network.
8. The lightweight image super-resolution method according to claim 1, characterized in that: The deep feature F fin Input the image reconstruction module to reconstruct the image and obtain a high-resolution image I SR , including: I SR =L shuffle (Conv(F fin +F′0)) Among them, L shuffle It represents sub-pixel convolution for upsampling operation, and Conv represents convolution operation.
9. The lightweight image super-resolution method according to claim 1, characterized in that: The high resolution image I HR and the reconstructed high-resolution image I SR Training the super-resolution model includes: The super-resolution model is trained using the L1 loss function, with loss function Loss = |I HR -I SR |.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.
Citation Information
Patent Citations
Image super-resolution reconstruction model and method based on residual mixed attention network
CN115222601A
Method for constructing image super-resolution model based on neural network
CN115908139A
Light field video time-angle super-resolution network based on deep learning
CN116152070A
Image super-resolution reconstruction method and system based on convolution and Transform hybrid architecture
CN118780984A
Image super-resolution method and system
WO2020238558A1
Cited By
Self-adaptive image super-resolution reconstruction method based on compressed sensing driving
CN120471774A
Remote sensing image deep learning super-resolution reconstruction method
CN121073781A
Single image super-resolution reconstruction method and device, computer equipment and storage medium
CN122492458A