Image enhancement method based on mask self-attention feature separation and multi-scale fusion
Through the image enhancement method of mask self-attention feature separation and multi-scale fusion, the problems of high model complexity and large calculation amount in the prior art are solved, and low-complexity and efficient image quality enhancement effect are achieved.
Patent Information
- Application Number
- CN202510407197.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-04-02
Smart Images

Figure CN120259110A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to an image enhancement method based on masked self-attention feature separation and multi-scale fusion. Background Art
[0002] In modern society, images are widely used in many fields, such as digital photography, environmental monitoring, medical imaging, satellite remote sensing, etc. However, during the processes of image acquisition, transmission, and storage, images are often affected by various factors, resulting in a decline in image quality, such as image blurring, detail loss, noise interference, etc. Single Image Enhancement Technology (SIET) can generate high-quality (HQ) images from low-quality (LQ) images. This technology can be used to improve the above-mentioned problems and enhance the visual effect and usability of images. Therefore, SIET technology has important research significance in various fields. In recent years, deep learning methods have made significant progress in the field of image enhancement. For example, some methods based on Convolutional Neural Network (CNN) perform image enhancement by learning the feature representation of images. However, these methods have some limitations. Traditional CNNs perform poorly in dealing with long-range dependencies, and when extracting multi-scale features, they often require a large number of parameters, leading to an increase in computational complexity and model complexity. In addition, some methods introduce a large computational cost and memory occupancy while improving image quality, which limits their application in resource-constrained environments.
[0003] In view of the above problems of the prior art, how to design a lightweight, efficient, and effective method for enhancing image quality has become the focus of current research. Summary of the Invention
[0004] In view of the above defects of the prior art, the present invention provides an image enhancement method based on masked self-attention feature separation and multi-scale fusion to solve the technical problems of high model complexity and large computational amount.
[0005] To achieve the above and other related purposes, the present invention provides an image enhancement method based on masked self-attention feature separation and multi-scale fusion, including: obtaining a first image to be processed; processing the first image with a trained image enhancement network to obtain a high-quality second image, wherein the expression of the image enhancement network is as follows: X HQ =H RC (H SF (X LQ ) + H DF (H SF(X LQ ))) where X LQ is the first image, X HQ is the second image, X SF is the shallow feature extraction unit, H DF is the deep feature enhancement unit, H RC is the high-quality image reconstruction unit.
[0006] In an embodiment of the present invention, the deep feature enhancement unit includes a masked self-attention feature separation unit and a dual-window multi-scale fusion unit; the masked self-attention feature separation unit separates the shallow features extracted by the shallow feature extraction unit based on the masked self-attention mechanism to obtain an associated feature and a differential feature, wherein the associated feature is used to characterize the global dependence relationship between different regions in the image, and the differential feature is used to characterize local detail information; the dual-window multi-scale fusion unit is used to process the associated feature and the differential feature to obtain deep features.
[0007] In an embodiment of the present invention, the masked self-attention feature separation unit separates the shallow features according to the following steps: performing a normalization operation on the shallow features to obtain a first feature; respectively processing the first feature by using three linear layers to obtain a query matrix, a key matrix, and a value matrix; performing a depthwise separable convolution operation on the query matrix and the key matrix and then performing a reshaping operation to respectively obtain a first matrix and a second matrix; performing a permutation operation on the second matrix to obtain a third matrix; performing a matrix multiplication calculation on the first matrix and the third matrix and then performing a normalization operation to obtain a fourth matrix; processing the fourth matrix by using a trainable binarization layer to obtain a binary matrix; performing a matrix multiplication calculation on the value matrix and the binary matrix to obtain the associated feature; subtracting the associated feature from the first feature to obtain the differential feature.
[0008] In an embodiment of the present invention, the dual-window multi-scale fusion unit includes a dual-window self-attention module and a multi-scale convolution module; the dual-window self-attention module processes the associated feature to enhance the modeling ability of long-range dependence relationships in the image; the multi-scale convolution module enhances the differential feature to extract local detail information at different scales; the outputs of the dual-window self-attention module and the multi-scale convolution module are fused to obtain the deep features.
[0009] In an embodiment of the present invention, the expression of the dual-window self-attention module is as follows: X1 = H RWLAB (LN(X lin )) + X lin ; X2 = MLP(LN(X1)) + X1; X3 = H TWLAB(LN(X2)) + X2; X4 = MLP(LN(X3)) + X3; where, X lin is the associated feature, X4 is the output of the double-window self-attention module, LN is normalization, MLP is a multi-layer perceptron, H RWLAB is a rectangular window self-attention block, H TWLAB is a triangular window self-attention block.
[0010] In an embodiment of the present invention, the multi-scale convolution module is formed by connecting convolution layers of multiple scales in parallel.
[0011] In an embodiment of the present invention, the expression of the multi-scale convolution module is as follows: X5 = Conv 3×3 (X dif ) + Conv 5×5 (X dif ) + Conv 7×7 (X dif ) + Conv 9×9 (X dif ), where Conv i×i represents a convolution operation with a convolution kernel size of i×i.
[0012] In an embodiment of the present invention, when training the image enhancement network, the loss function is calculated according to the following formula: , where is the real image corresponding to the input data, is the prediction result after enhancing the input data by the image enhancement network. The superscript i is used to indicate the correspondence between the prediction result and the real image, and N is the total number of samples in each batch.
[0013] Advantages of the present invention: An image enhancement method based on masked self-attention feature separation and multi-scale fusion proposed by the present invention. This method builds an image enhancement network based on masked self-attention feature separation and double-window multi-scale feature fusion. Based on the masked self-attention mechanism, the features of the input image are separated into associated features and differential features, and then the associated features and differential features are processed and fused by double-window multi-scale features, combined with shallow features, to generate high-quality enhanced images. This method has low complexity and can effectively enhance the image quality. Description of the Drawings
[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0015] Figure 1 Flow chart of the image enhancement method provided by an embodiment of the present invention; Figure 2 Architecture diagram of the image enhancement network provided by an embodiment of the present invention; Figure 3 Architecture diagram of the masked self-attention feature separation unit provided by an embodiment of the present invention; Figure 4 Architecture diagram of the dual-window multi-scale fusion unit provided by an embodiment of the present invention; Figure 5 Architecture diagram of the dual-window self-attention module provided by an embodiment of the present invention; Figure 6 Architecture diagram of the multi-scale convolution module provided by an embodiment of the present invention. Detailed implementation manners
[0016] The following specific embodiments are used to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. Except for the specific methods, devices, and materials used in the embodiments, according to the knowledge of those skilled in the art in the technical field and the description of the present invention, any methods, devices, and materials similar or equivalent to those described in the embodiments of the present invention can also be used to implement the present invention.
[0017] It should be understood that the terms used in the embodiments of the present invention are for the purpose of describing specific implementation manners and are not intended to limit the protection scope of the present invention. Unless otherwise defined, all technical and scientific terms used in the present invention have the same meaning as commonly understood by those skilled in the technical field of the present invention.
[0018] In the following description, a large number of details are explored to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In some of these embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.
[0019] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of methods and computer program products that can be implemented according to various embodiments disclosed in the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions denoted in the blocks may occur in a different order than that denoted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0020] Please refer to Figure 1 , Figure 1 An image enhancement method based on mask self-attention feature separation and multi-scale fusion provided for an embodiment of the present invention includes the following two steps: obtaining a first image to be processed; using a trained image enhancement network to process the first image to obtain a high-quality second image. Among them, the image enhancement network is as Figure 2 shown, and its expression is as follows: X HQ =H RC (H SF (X LQ )+H DF (H SF (X LQ ))), where X LQ is the first image, that is, the image to be processed, X HQ is the second image, that is, the enhanced image output by the network model, X SF is the shallow feature extraction unit, H DF is the deep feature enhancement unit, and H RC is the high-quality image reconstruction unit.
[0021] The shallow feature extraction unit here is equivalent to the preprocessing before the deep feature enhancement unit. In a specific embodiment of the present invention, it can be composed of a single 3×3 convolution, for example, to extract shallow features , and H SF (X LQ ) in the expression is X s . The deep feature enhancement unit is the focus of the present invention. It enhances the shallow feature X s based on the mask self-attention feature separation and multi-scale fusion method to obtain the deep feature , and H in the expressionDF (H SF (X LQ )) That is, X d . Finally, the high-quality image reconstruction unit will perform fusion on the shallow feature X s and the deep feature X d to output the second image X HQ .
[0022] Please refer to Figure 2 . In a specific embodiment of the present invention, the deep feature enhancement unit includes a masked self-attention feature separation unit and a dual-window multi-scale fusion unit. The masked self-attention feature separation unit separates the shallow features extracted by the shallow feature extraction unit based on the masked self-attention mechanism to obtain the correlation feature and the differential feature. Among them, the correlation feature is used to represent the global dependence relationship between different regions in the image, and the differential feature is used to represent the local detail information; the dual-window multi-scale fusion unit is used to process the correlation feature and the differential feature to obtain the deep feature.
[0023] The masked self-attention feature separation unit is used to decompose the shallow features to obtain the correlation feature and the differential feature, and then use the dual-window multi-scale fusion unit to process these two features respectively. After processing, they are fused to obtain the deep feature.
[0024] Please refer to Figure 3 . In a specific embodiment of the present invention, the architecture diagram of the masked self-attention feature separation unit is as Figure 3 shown. As a key component of the image enhancement network, this unit plays an important role in feature extraction and feature separation operations, and can not only achieve efficient calculation but also improve the network's ability to represent image features. The masked self-attention feature separation unit separates the shallow features according to the following steps.
[0025] (1) Perform a normalization operation on the shallow feature X s to obtain the first feature X s '. The normalization operation here corresponds to Figure 3 Norm in, whose full name is Normalization. Before feature extraction, normalizing the input features is to stabilize the training process and accelerate convergence.
[0026] (2) Use three linear layers to process the first feature respectively to obtain the query matrix, the key matrix, and the value matrix. The linear layer is Figure 3 Linear in, specifically, three linear layers L Q , L K , L V will be used to process the first feature X s' is processed to obtain a query matrix Q, a key matrix K, and a value matrix V. Among them, the number of channels of the value matrix V is the same as that of the shallow feature X s remains consistent.
[0027] (3) After performing depthwise separable convolution operations on the query matrix and the key matrix and then performing a reshaping operation, a first matrix and a second matrix are obtained respectively. The depthwise separable convolution operation is the DWconv in Figure 3 , and the reshaping operation is the R in Figure 3 . After these two operations, the number of channels of the query matrix Q and the key matrix K will be compressed to C / r 2 , obtaining and , where Q' is the first matrix and K' is the second matrix. C is the number of channels of the shallow feature X s , r is a hyperparameter used to control the compression ratio of the number of channels, N is the batch size (i.e., the number of samples input into the model at one time), and S 2 represents the spatial resolution of the feature map (i.e., height × width).
[0028] (4) A permutation operation is performed on the second matrix to obtain a third matrix. The permutation operation is the P (Permutation) in Figure 3 . After performing the permutation operation on the second matrix K', a third matrix K'' is obtained to make it suitable for the subsequent self-attention operation.
[0029] (5) After performing matrix multiplication on the first matrix and the third matrix and then performing a normalization operation, a fourth matrix is obtained. The matrix multiplication is the × (Matrix Multiplication) in Figure 3 , and the normalization operation is the σ in Figure 3 . Here, normalization is performed using the Softmax function, which is different from the previous Norm. Norm scales the data to a specific range (usually [0,1] or [-1,1]), or standardizes the distribution of the data (mean is 0, variance is 1); Softmax converts a set of numerical values into a probability distribution such that their sum is 1.
[0030] (6) The fourth matrix is processed using a trainable binarization layer to obtain a binary matrix. The trainable binarization layer is the TBL (Trainable Binarization Layer) in Figure 3 . The obtained binary matrix can be denoted as A. This binarization process is trainable, and the model can optimize the parameters of the binarization layer through backpropagation, so as to better adapt to different input feature distributions. By learning an appropriate threshold, the values corresponding to the associated features are set to 1, and the values corresponding to the differential features are set to 0, thereby obtaining the binary matrix A.
[0031] (7) Perform matrix multiplication on the value matrix and the binary matrix to obtain the associated features. As mentioned above, the binary matrix A obtained through the processing of the trainable binarization layer, where the values corresponding to the associated features are set to 1 and the values corresponding to the differential features are set to 0. Therefore, after multiplying the value matrix V by the binary matrix A, these associated features can be retained, and the associated features are X lin (i.e., Figure 3 the linked features in
[0032] ), which is equal to V×A. (8) Subtract the associated features from the first feature to obtain the differential features. By making an identity mapping at the start of the unit and subtracting from the associated feature X lin at the output port, the differential features can be obtained, and the differential features are X dif (i.e., Figure 3 the difference features in s ), which is equal to X lin ’ - X S
[0033] Through the above steps (1) to (8), the input shallow features X lin can be partitioned to obtain the associated feature X dif and the differential feature X
[0034] Please refer to Figure 4 In a specific embodiment of the present invention, the dual-window multi-scale fusion unit includes a dual-window self-attention module and a multi-scale convolution module. Among them, the dual-window self-attention module processes the associated features to enhance the modeling ability of long-range dependencies in the image; the multi-scale convolution module enhances the differential features to extract local detail information at different scales; finally, the outputs of the dual-window self-attention module and the multi-scale convolution module are fused to obtain the deep features.
[0035] Please refer to Figure 5 In a specific embodiment of the present invention, the expression of the dual-window self-attention module is as follows: X1 = H RWLAB (LN(X lin )) + X lin X2 = MLP(LN(X1)) + X1; X3 = H TWLAB (LN(X2)) + X2; X4 = MLP(LN(X3)) + X3; Among them, X linFor associated features, X1 to X3 are intermediate features. The above four expressions can be written as a general formula. In the general formula, these intermediate features can be omitted, and X4 is the output of the dual-window self-attention module. In the above formula, LN is normalization, corresponding to Norm in Figure 5 ; MLP is a multi-layer perceptron; H RWLAB is a rectangular window self-attention block, corresponding to RWLAB in Figure 5 ; H TWLAB is a triangular window self-attention block, corresponding to TWLAB in Figure 5 . Due to the defect that the rectangular window local self-attention technology is vulnerable to distortion at the boundary, to address this problem, in the dual-window self-attention module, a rectangular window self-attention block and a triangular window self-attention block are added. The two work in series synchronously. First, the image is divided into rectangular windows, and then divided according to triangular windows. Such a connection method has the best enhancement effect on the image quality. Subsequently, the features of adjacent windows are overlapped and cross-fused, and cross-attention is constructed between these windows. Such a design can reduce boundary distortion, capture long-range and multi-scale features, and more effectively enhance various features of the image.
[0036] The rectangular window self-attention block and the triangular window self-attention block are two variants of the local window-based self-attention mechanism, mainly used to improve the computational efficiency and performance of the Transformer model. They reduce the computational complexity by restricting the attention range while retaining the ability to capture local information.
[0037] Please refer to Figure 4 , in a specific embodiment of the present invention, the multi-scale convolution module is composed of convolution layers of multiple scales in parallel. Its function is to perform further feature extraction on the differential feature X dif output by the masked self-attention feature separation unit to achieve a better image enhancement effect.
[0038] Please refer to Figure 5 , in a specific embodiment of the present invention, the expression of the multi-scale convolution module is as follows: X5 = Conv 3×3 (X dif ) + Conv 5×5 (X dif ) + Conv 7×7 (X dif ) + Conv 9×9 (X dif ), where Conv i×i represents a convolution operation with a convolution kernel size of i×i.
[0039] Finally, the deep feature X d extracted by the deep feature enhancement unit = H DF (X S) = X4 + X5. Then, after adding the deep feature X d and the shallow feature X s and processing through the high-quality image reconstruction unit, a high-quality second image can be obtained.
[0040] In a specific embodiment of the present invention, after obtaining the high-quality result, we use the L1 loss function to calculate and measure the error between the high-quality result X HQ and the real data X RHQ , and use the backpropagation method to optimize the network parameters, thus completing the training of the image enhancement network based on mask self-attention feature separation and double-window multi-scale fusion. Specifically, the task of training the network using the training samples includes the following steps: substituting the image X LQ in the LQ-HQ image pair into the image enhancement network to obtain the image X HQ ; calculating the loss loss according to the image X HQ , the image X RHQ in the LQ-HQ image pair, and the following loss function calculation formula: , wherein, is the real image corresponding to the input data, is the predicted result after the input data is enhanced by the image enhancement network, the superscript i is used to indicate the correspondence between the predicted result and the real image, and N is the total number of samples in each batch. The above steps can conveniently help us achieve the training of the image enhancement network parameters. The lightweight image enhancement network with double-window self-attention and feature separation obtained after training can directly be used to enhance the feature quality of the input first image X LQ to obtain its corresponding second image X HQ .
[0041] To further illustrate the effect of the method of the present invention, the effect of the present invention will be further described below in combination with experiments.
[0042] 1. Experimental conditions: The computer hardware environment for the experiments of the present invention is an Intel Core i9-10980XE CPU and a GTX3090 GPU, the software environment is the Centos 7.6 operating system, the compilation environment is PyCharm, and the deep learning framework is PyTorch; all subsequent training and testing are based on this platform. Table 1 partially shows the quantitative analysis of the method in the present invention based on the ×2, ×3, and ×4 scaling factors on the DIV2K and DF2K (DIV2K + Flickr2K) training datasets, comprehensively tests the model performance, and makes a performance comparison with other stream graphics processing models.
[0043] 2. Image Enhancement Evaluation Metrics: To evaluate the enhancement performance of the image feature enhancement method of the present invention, the present invention uses two evaluation metrics, Peak Signal to Noise Ratio (PSNR) and Structural Similarity (SSIM), to evaluate the enhancement results.
[0044] PSNR is an objective evaluation metric for images, and its expression is as follows: , where MAX represents the maximum value of the image point color, and MSE represents the mean square error of X HQ and its corresponding X RHQ . However, the PSNR value does not necessarily match the visual quality perceived by the human eye. To overcome this drawback, we also use SSIM to evaluate the feature enhancement results, and its expression is as follows: , , , , where l(x,y) represents the comparison of image brightness, c(x,y) represents the comparison of image contrast, s(x,y) represents the comparison of image structure, μ represents the mean, σ represents the standard deviation, and σ xy represents the covariance. c1, c2, and c3 are constants used to maintain stability and are calculated by the following formula: , where L is the maximum value of the pixel values in the image, k1 = 0.01, k2 = 0.03. In practical applications, we often use the simplified SSIM formula: , SSIM models distortion as a combination of three different factors: brightness, contrast, and structure, and can better reflect image quality than PSNR.
[0045] 3. Experimental Content and Result Analysis: The method in the present invention was compared with the currently common and most advanced methods. Here, EDSR, SAN, and HAN based on CNN, as well as advanced architectures such as IPT, SwinIR, Swin2SR, ACT, ART, EDT, and HAT based on Transformer, were selected to quantitatively compare the model of the present invention in terms of PSNR and SSIM. HAT was considered the best model before, and its results were better than any other SOTA model. However, due to the distortion-free and rich feature exploration ability of the image enhancement network in the present invention, it can produce superior performance to HAT for all five benchmark datasets and all scaling factors. These excellent results confirm the importance of using triangular window attention in the present invention. The performance metrics of all the compared models are from their respective papers, and the comparison results are shown in Tables 1 - 3. Among them, the best results are in bold, and the second-best are underlined.
[0046] Table 1: Comparison of Multiple Methods at ×2 Scale
[0047] Table 2: Comparison of Multiple Methods at ×3 Scale
[0048] Table 3: Comparison of Multiple Methods at ×4 Scale
[0049] It can be seen from the quantitative comparison results in Tables 1 - 3 that the method proposed in the present invention uses a more flexible method to process the information of different features in the image and achieves the optimal image enhancement results in all cases.
[0050] To show that our proposed model has a better trade-off between effectiveness and efficiency, we also qualitatively compared the task performance of each method when inferring the dataset at ×4 image scale, and the results are shown in Table 4.
[0051] Table 4: Quantitative Trade-off between the Effectiveness and Efficiency of the Model
[0052] Here, two currently better-performing Transformer-based methods are selected, and their PSNR, SSIM, Params (number of parameters), FLOPs (Floating Point Operations, number of floating-point operations per second), Memory (maximum video memory occupancy), and Latency (inference time) when inferring Set14 at a ×4 image scale are compared. From the results, compared with other methods, the present invention has a very significant advantage in terms of inference time.
[0053] In summary, the present invention proposes an image enhancement method based on masked self-attention feature separation and dual-window multi-scale fusion. A masked self-attention feature separation unit and a dual-window multi-scale fusion unit are built, and the two units form a deep feature enhancement unit in the image enhancement network. It not only solves the problems of large video memory occupancy and long calculation time of the graphics card, but also can achieve higher-quality feature enhancement tasks.
[0054] The above embodiments are only illustrative of the principles and effects of the present invention, and are not intended to limit the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes completed by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed by the present invention should still be covered by the claims of the present invention.
Claims
1. An image enhancement method based on mask self-attention feature separation and multi-scale fusion, characterized in that Including: Obtain a first image to be processed; Process the first image by using a trained image enhancement network to obtain a second image of high quality, wherein the expression of the image enhancement network is as follows: X HQ =H RC (H SF (X LQ )+H DF (H SF (X LQ ))), where X LQ is the first image, X HQ is the second image, X SF is the shallow feature extraction unit, H DF is the deep feature enhancement unit, H RC is the high-quality image reconstruction unit.
2. The image enhancement method based on mask self-attention feature separation and multi-scale fusion according to claim 1, wherein The deep feature enhancement unit includes a masked self-attention feature separation unit and a dual-window multi-scale fusion unit; The masked self-attention feature separation unit separates the shallow features extracted by the shallow feature extraction unit based on the masked self-attention mechanism to obtain an associated feature and a differential feature, wherein the associated feature is used to represent the global dependence relationship between different regions in the image, and the differential feature is used to represent local detail information; The dual-window multi-scale fusion unit is used to process the associated feature and the differential feature to obtain deep features.
3. The image enhancement method based on mask self-attention feature separation and multi-scale fusion according to claim 2, wherein, The masked self-attention feature separation unit separates the shallow features according to the following steps: Perform a normalization operation on the shallow features to obtain a first feature; Process the first feature by using three linear layers respectively to obtain a query matrix, a key matrix, and a value matrix; Perform a depthwise separable convolution operation on the query matrix and the key matrix and then perform a reshaping operation to obtain a first matrix and a second matrix respectively; Perform a permutation operation on the second matrix to obtain a third matrix; Perform a matrix multiplication calculation on the first matrix and the third matrix and then perform a normalization operation to obtain a fourth matrix; Process the fourth matrix by using a trainable binarization layer to obtain a binary matrix; Perform a matrix multiplication calculation on the value matrix and the binary matrix to obtain the associated feature; Subtract the associated feature from the first feature to obtain the differential feature.
4. The image enhancement method based on mask self-attention feature separation and multi-scale fusion according to claim 2, characterized in that The dual-window multi-scale fusion unit includes a dual-window self-attention module and a multi-scale convolution module; The dual-window self-attention module processes the associated feature to enhance the modeling ability of long-range dependence relationships in the image; The multi-scale convolution module enhances the differential feature to extract local detail information of different scales; Fuse the outputs of the dual-window self-attention module and the multi-scale convolution module to obtain the deep features.
5. The image enhancement method based on mask self-attention feature separation and multi-scale fusion according to claim 4, wherein The expression of the dual-window self-attention module is as follows: X1 = H RWLAB (LN(X lin )) + X lin ; X2 = MLP(LN(X1)) + X1; X3=H TWLAB (LN(X2))+X2; X4 = MLP(LN(X3)) + X3; Among them, X lin is the associated feature, X4 is the output of the double-window self-attention module, LN is normalization, MLP is a multi-layer perceptron, H RWLAB is a rectangular window self-attention block, H TWLAB is a triangular window self-attention block.
6. The image enhancement method based on mask self-attention feature separation and multi-scale fusion according to claim 4, wherein The multi-scale convolution module is composed of convolution layers of multiple scales connected in parallel.
7. The image enhancement method based on mask self-attention feature separation and multi-scale fusion according to claim 6, characterized in that, The expression of the multi-scale convolution module is as follows: X5 = Conv 3×3 (X dif ) + Conv 5×5 (X dif ) + Conv 7×7 (X dif ) + Conv 9×9 (X dif ) In the formula, Conv i×i represents a convolution operation with a convolution kernel size of i×i.
8. The image enhancement method based on mask self-attention feature separation and multi-scale fusion according to claim 1, characterized in that When training the image enhancement network, calculate the loss function according to the following formula: , In the formula, is the true image corresponding to the input data, is the prediction result after enhancing the input data through the image enhancement network. The superscript i is used to indicate the correspondence between the prediction result and the true image, and N is the total number of samples in each batch.
Citation Information
Patent Citations
Image super-resolution reconstruction method based on multi-scale pyramid network
CN111402128A
Image denoising method for enhancing gating Transform
CN117726540A
Space-time adaptive thermal infrared tracking method based on coordinate information
CN118429385A
Image super-resolution reconstruction method and device based on high-frequency feature enhancement
CN119228651A
Hydropower station floater detection method based on global relation modeling
CN119516474A