Wavelet-space dual attention image deraining method and system guided by prior knowledge

CN118014890BActive Publication Date: 2026-08-07SUN YAT SEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SUN YAT SEN UNIV
Filing Date
2024-01-17
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0005]尽管现有的方法在图像去雨方面取得了一定的成果,但仍面临一些挑战,例如,现有的网络结构仍有优化空间,去雨性能也有待提升,此外,大多数方法在合成数据集上表现良好,但在真实场景中的泛化能力较差

Benefits of technology

[0052]本发明提供了一种先验知识引导的小波-空间双注意力图像去雨方法及系统,所述方法通过获取原始降雨图像作为初始输入图像;利用基于先验知识引导的小波-空间双注意力去雨深度学习网络模型对初始输入图像进行去雨处理,得到无雨图像;其中,基于先验知识引导的小波-空间双注意力去雨深度学习网络模型包括输入映射模块、先验知识引导模块、基于小波-空间双注意力的层次化编码-解码网络和输出映射模块。与现有技术相比,该方法结合小波-空间双注意力和先验知识引导机制,使雨天条件下包含丰富结构信息的残差通道先验可以有效的引导网络的学习,有效提取降雨图像的空间域信息和频率域信息,提高了图像去雨精度,同时采用多尺度特征交互策略降低网络学习难度,使得网络能够适应各种复杂的降雨场景,稳定提升去雨效果的准确性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118014890B_ABST
    Figure CN118014890B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of image processing, and more particularly to a priori knowledge guided wavelet-space double attention image rain removal method and system, comprising: obtaining an original rainfall image as an initial input image; using a priori knowledge guided wavelet-space double attention rain removal deep learning network model to perform rain removal processing on the initial input image to obtain a rain-free image; the a priori knowledge guided wavelet-space double attention rain removal deep learning network model comprises an input mapping module, a priori knowledge guiding module, a wavelet-space double attention based hierarchical encoding-decoding network and an output mapping module. The rain removal deep learning network model used in the present application combines a priori knowledge guiding mechanism, wavelet-space double attention and multi-scale feature interaction strategy to enable the network model to better understand the structure and features of rainy day images, enhance feature extraction, improve image detail and texture quality, and reduce network learning difficulty.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to a wavelet-spatial dual attention image deraining method and system guided by prior knowledge. Background Technology

[0002] Rain is a very common weather phenomenon. Under rainy conditions, the quality of images taken by cameras will be severely degraded, which has a significant impact on fields that rely on visual images, such as security monitoring, autonomous driving, and aerospace. Therefore, how to effectively remove the effects of rain from images has become a very important research direction in the field of computer vision.

[0003] Existing image deraining methods are mainly divided into two categories: traditional image processing methods and deep learning-based methods. Traditional image deraining methods mainly rely on manually designed prior information based on rain patterns and statistical features of clear images, such as filter-based methods, image decomposition and reconstruction, motion blur estimation, and deconvolution. However, these methods have limited versatility and adaptability when facing complex rain scenes, and their applicability is poor when the application scenario changes.

[0004] With the advancement of computer technology, deep learning methods have gradually become mainstream. Compared with traditional image processing methods, deep learning methods can automatically learn the required features from a large amount of data, rather than simply relying on prior design. Their performance and generalization are significantly improved compared to traditional methods. Currently, the deep learning-based methods are mainly fully supervised rain removal methods represented by convolutional neural networks (CNNs). These methods learn the mapping from rainy images to rainless images by constructing deep learning networks and have verified their effectiveness on multiple datasets. In addition to convolutional neural networks (CNNs), Transformer networks have also played an important role in the field of image rain removal. Transformer-based methods have also been gradually applied to the field of image rain removal, further improving the performance of rain removal.

[0005] Although existing methods have achieved some success in image deraining, they still face some challenges. For example, there is still room for optimization in existing network structures, and the deraining performance needs to be improved. In addition, most methods perform well on synthetic datasets, but their generalization ability in real-world scenes is poor.

[0006] In rainy environments, rain appears in images as rain streaks, which significantly affect the high-frequency texture structure information of the image. Therefore, it is crucial to utilize frequency domain information for image deraining. However, most existing methods design and model networks in the spatial domain, neglecting frequency domain information. Although some methods consider modeling frequency domain features, their performance still has room for improvement due to the lack of reasonable construction methods. In addition, most methods lack effective prior knowledge guidance, leading to distortion of the recovered image texture and structural information and deviation from the real image. Therefore, how to better combine prior knowledge, optimize network structure, and improve deraining performance and generalization ability is an important research direction that urgently needs to be addressed. Summary of the Invention

[0007] The purpose of this invention is to provide a wavelet-spatial dual attention image deraining method and system guided by prior knowledge, so that the residual channel prior containing rich structural information under rainy conditions can effectively guide the network's learning and improve the accuracy of image deraining.

[0008] To address the above technical problems, this invention provides a wavelet-spatial dual attention image deraining method and system guided by prior knowledge.

[0009] In a first aspect, the present invention provides a wavelet-spatial dual-attention image deraining method guided by prior knowledge, the method comprising the following steps:

[0010] Obtain the original rainfall image as the initial input image;

[0011] The initial input image is processed to remove rain using a wavelet-spatial dual-attention rain removal deep learning network model guided by prior knowledge, resulting in a rain-free image.

[0012] The wavelet-spatial dual attention rain removal deep learning network model based on prior knowledge guidance includes an input mapping module, a prior knowledge guidance module, a hierarchical encoder-decoder network based on wavelet-spatial dual attention, and an output mapping module.

[0013] In a further implementation, the prior knowledge guidance module includes a prior information extraction module and a prior knowledge fusion module connected in series;

[0014] The prior information extraction module includes a residual channel conversion unit, an input convolutional layer, an SE-ResBlock stacked model, and an output convolutional layer connected in sequence; the SE-ResBlock stacked model includes several SE-ResBlock structures, and each SE-ResBlock structure includes two convolutional layers and an SE-Block.

[0015] The hierarchical encoder-decoder network includes multiple encoder-decoder stages and a multi-scale information interaction module connecting two adjacent encoder-decoder stages. Each encoder-decoder stage includes a Transformer module based on wavelet-spatial dual attention.

[0016] In a further implementation, the step of using a wavelet-spatial dual-attention deraining deep learning network model guided by prior knowledge to perform deraining processing on the initial input image to obtain a rain-free image includes:

[0017] The initial input image is mapped to the feature space through the input mapping module to obtain image mapping features. At the same time, the initial input image is input in parallel into the prior information extraction module for residual channel prior feature extraction.

[0018] The residual channel prior features and the image mapping features are input into the prior knowledge fusion module to enhance the background structure features, thereby obtaining prior fusion enhancement information;

[0019] Based on a multi-scale information interaction strategy, the prior fusion enhancement information is processed through the hierarchical encoder-decoder network to generate a rainless image.

[0020] In a further implementation, the step of inputting the initial input image in parallel into the prior information extraction module for residual channel prior feature extraction includes:

[0021] The initial input image is converted into a residual channel image by the residual channel conversion unit, and the residual channel image is mapped to the feature space by the input convolutional layer to obtain a residual channel feature matrix containing residual channel information.

[0022] The residual channel feature matrix is ​​input into the SE-ResBlock stacked model to learn prior background structure information, and then processed through the output convolutional layer to output the residual channel prior features.

[0023] In a further implementation, the step of inputting the residual channel prior features and the image mapping features into the prior knowledge fusion module for background structure feature enhancement to obtain prior fusion enhancement information includes:

[0024] The residual channel prior features are converted into query vectors through linear layers and depthwise separable convolutional layers, and auxiliary value vectors containing prior information about the background structure are generated.

[0025] The image mapping features are converted into key vectors and value vectors using linear layers and depthwise separable convolutional layers.

[0026] Calculate the dot product attention between the query vector and the key vector, and enhance the background structure prior information based on the dot product attention to obtain a prior attention map;

[0027] The prior attention map is element-wise multiplied with the auxiliary value vector and the value vector respectively to obtain the auxiliary value multiplication result and the value multiplication result. The auxiliary value multiplication result and the value multiplication result are concatenated to obtain the value concatenation feature. The value concatenation feature is input into the linear layer to obtain the multi-head attention output feature.

[0028] The multi-head attention output features are input into a local augmentation feedforward network for dimensional expansion and compression to obtain prior fusion augmentation information.

[0029] In a further implementation, the step of processing the prior fusion enhancement information through the hierarchical encoder-decoder network based on the multi-scale information interaction strategy to generate a rainless image includes:

[0030] The prior fusion enhancement information is used as the encoding input feature of the first encoding and decoding stage. In each encoding and decoding stage, the input encoding input feature is divided into wavelet domain and spatial domain by a Transformer module based on wavelet-space dual attention with a preset attention head segmentation ratio to obtain wavelet domain feature map and spatial domain feature map.

[0031] Based on the wavelet domain feature map and the spatial domain feature map, the wavelet domain multi-head attention output features and the spatial domain multi-head attention output features are calculated.

[0032] The wavelet domain attention output features and the spatial domain attention output features are connected in parallel, and wavelet-spatial dual attention output features at different scales are obtained through a local enhancement feedforward network. Between two adjacent encoding and decoding stages, the wavelet-spatial dual attention output features are obtained by convolution downsampling or transposed convolution upsampling to obtain multi-scale feature representations.

[0033] Based on a multi-scale information interaction strategy, multi-scale feature representations at different scales are bridged to obtain global multi-scale output features.

[0034] The output mapping module performs a residual connection between the global multi-scale output features and the initial input image to generate a rainless image.

[0035] In a further implementation, the step of calculating the wavelet domain multi-head attention output features and the spatial domain multi-head attention output features based on the wavelet domain feature map and the spatial domain feature map includes:

[0036] The wavelet domain feature map is linearly transformed into a wavelet domain linear feature map, and a wavelet domain query vector is generated.

[0037] The discrete wavelet transform is used to decompose the wavelet domain linear feature map into four wavelet sub-bands. The four wavelet sub-bands are then connected side by side along the channel dimension to obtain the wavelet sub-band splicing feature.

[0038] A convolutional layer is used to apply spatial locality to the wavelet subband splicing features to obtain local context features, and the local context features are linearly transformed into wavelet domain key vectors and wavelet domain value vectors according to the preset attention head segmentation ratio.

[0039] Multi-head attention is calculated for each group of wavelet domain key vectors, wavelet domain value vectors and wavelet domain query vectors to obtain the wavelet domain multi-head attention output features.

[0040] The spatial domain feature map is windowed into several spatial domain feature matrices, and three parallel linear layers are used to perform linear transformation on the spatial domain feature matrices to obtain spatial domain key vectors, spatial domain value vectors, and spatial domain query vectors.

[0041] Multi-head attention is performed on each set of spatial domain key vectors, spatial domain value vectors, and spatial domain query vectors to obtain the spatial domain multi-head attention output features.

[0042] In a further implementation, the step of bridging multi-scale feature representations at different scales based on a multi-scale information interaction strategy to obtain global multi-scale output features includes:

[0043] The output of the last encoding and decoding stage is used as the global feature, and a multi-scale information matrix is ​​obtained based on the multi-scale feature representation at different scales.

[0044] The global features and the multi-scale information matrix are sequentially input into the layer normalization, the linear layer and the activation function to obtain the corresponding standard global features and standard multi-scale information matrix.

[0045] A multi-axis cross-gating module is used to interactively operate the standard global features and the standard multi-scale information matrix to obtain global multi-scale output features, which include global semantically enhanced multi-scale output features and semantically optimized global output features.

[0046] Secondly, this invention provides a priori knowledge-guided wavelet-spatial dual-attention image deraining system, the system comprising:

[0047] The image acquisition module is used to acquire the original rainfall image as the initial input image;

[0048] The rain removal module is used to perform rain removal processing on the initial input image using a wavelet-spatial dual-attention rain removal deep learning network model guided by prior knowledge, so as to obtain a rain-free image.

[0049] The wavelet-spatial dual attention rain removal deep learning network model based on prior knowledge guidance includes an input mapping module, a prior knowledge guidance module, a hierarchical encoder-decoder network based on wavelet-spatial dual attention, and an output mapping module.

[0050] In addition, in a third aspect, the present invention also provides a computer device, including a processor and a memory, the processor being connected to the memory, the memory being used to store a computer program, and the processor being used to execute the computer program stored in the memory, so that the computer device performs the steps of implementing the above-described method.

[0051] Fourthly, the present invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described method.

[0052] This invention provides a wavelet-spatial dual-attention image deraining method and system guided by prior knowledge. The method acquires an original rainfall image as the initial input image; it then uses a wavelet-spatial dual-attention deraining deep learning network model guided by prior knowledge to process the initial input image for deraining, resulting in a rain-free image. The wavelet-spatial dual-attention deraining deep learning network model includes an input mapping module, a prior knowledge guidance module, a hierarchical encoder-decoder network based on wavelet-spatial dual attention, and an output mapping module. Compared with existing technologies, this method combines wavelet-spatial dual attention and prior knowledge guidance mechanisms, enabling the residual channel priors containing rich structural information under rainy conditions to effectively guide the network's learning. This effectively extracts spatial and frequency domain information from the rainfall image, improving the image deraining accuracy. Simultaneously, the use of a multi-scale feature interaction strategy reduces the learning difficulty of the network, allowing it to adapt to various complex rainfall scenarios and consistently improve the accuracy of the deraining effect. Attached Figure Description

[0053] Figure 1 This is a schematic diagram of the prior knowledge-guided wavelet-spatial dual attention image rain removal method provided in the embodiments of the present invention;

[0054] Figure 2 This is a schematic diagram of the overall structure of the wavelet-spatial dual attention rain removal deep learning network model provided in an embodiment of the present invention;

[0055] Figure 3 This is a schematic diagram of the prior information extraction module structure provided in an embodiment of the present invention;

[0056] Figure 4 This is a schematic diagram of the prior information fusion module structure provided in an embodiment of the present invention;

[0057] Figure 5 This is a schematic diagram of the Transformer module structure based on wavelet-spatial dual attention provided in an embodiment of the present invention;

[0058] Figure 6 This is a schematic diagram of the multi-scale information interaction module structure provided in an embodiment of the present invention;

[0059] Figure 7 This is a block diagram of the wavelet-spatial dual attention image deraining system guided by prior knowledge provided in an embodiment of the present invention;

[0060] Figure 8 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present invention. Detailed Implementation

[0061] The embodiments of the present invention are described in detail below with reference to the accompanying drawings. The embodiments are given for illustrative purposes only and should not be construed as limiting the present invention. The accompanying drawings are for reference and illustration only and do not constitute a limitation on the scope of patent protection of the present invention, because many changes can be made to the present invention without departing from the spirit and scope of the present invention.

[0062] refer to Figure 1 This invention provides a priori knowledge-guided wavelet-spatial dual-attention image deraining method, such as... Figure 1 As shown, the method includes the following steps:

[0063] S1. Obtain the original rainfall image as the initial input image.

[0064] S2. The initial input image is processed to remove rain using a wavelet-spatial dual-attention deraining deep learning network model guided by prior knowledge, resulting in a rain-free image.

[0065] The wavelet-spatial dual-attention rain removal deep learning network model proposed in this embodiment is an end-to-end fully supervised deep learning network. This network model can directly learn the mapping relationship between rain-free images and rain-free images from rainfall images. The input of the network model is a 3-channel RGB image of the rainfall conditions, and the output is the predicted rain-free image. During the training process, this embodiment uses a dataset with positive and negative samples to train the network model. The positive sample dataset consists of images with clear backgrounds, while the negative sample dataset consists of rainfall images with synthetic rain lines on the same background. After training with the dataset with positive and negative samples, the network model can be validated on the test set of the dataset and in real-world scenarios to ensure its generalization ability.

[0066] In this embodiment, the wavelet-spatial dual-attention deraining deep learning network model guided by prior knowledge includes an input layer, an input mapping module, a prior knowledge guidance module, a hierarchical encoder-decoder network based on wavelet-spatial dual attention, an output mapping module, and an output layer connected in sequence. The input mapping module is also connected in parallel with the prior knowledge guidance module at the output end of the input layer. The prior knowledge guidance module includes a prior information extraction module and a prior knowledge fusion module connected in series. The hierarchical encoder-decoder network includes multiple encoding / decoding stages and a multi-scale information interaction module connecting adjacent encoding / decoding stages. Each encoding / decoding stage includes a Transformer module based on wavelet-spatial dual attention. In this embodiment, the step of using the wavelet-spatial dual-attention deraining deep learning network model guided by prior knowledge to derain the initial input image and obtain a rain-free image includes:

[0067] The initial input image is mapped to the feature space through the input mapping module to obtain image mapping features;

[0068] The initial input image is input in parallel into the prior information extraction module. The prior information extraction module extracts residual channel prior features from the initial input image. The residual channel prior features and the image mapping features are then input into the prior knowledge fusion module for background structure feature enhancement to obtain prior fusion enhancement information.

[0069] Based on a multi-scale information interaction strategy, the prior fusion enhancement information is processed through the hierarchical encoder-decoder network to generate a rainless image.

[0070] Figure 2 This is a schematic diagram of the overall structure of the wavelet-spatial dual-attention deraining deep learning network model. To better understand the image deraining method proposed in this embodiment, this embodiment first... Figure 2 The overall network structure and flow are briefly summarized as follows: the input image is initially mapped to the feature space through the input mapping module to obtain the image mapping features X. input ∈R H×W×DThe input mapping module consists of a 3×3 standard convolution and an activation function. In this embodiment, the input image is a 3-channel RGB image. Simultaneously, the input image is input in parallel to the prior knowledge guidance module, which includes a prior information extraction module and a prior knowledge fusion module. The prior information extraction module first converts the input image into a residual channel image and extracts features from it. The information extracted from the prior information extraction module contains richer background structural details. Then, in this embodiment, the features mapped from the input image and the extracted residual channel prior features are jointly input into the prior information fusion module for processing. The role of the prior information fusion module is to enhance the structural features of the background and guide the subsequent extraction of rain line-related features. Next, the output of the prior knowledge guidance module is input into the subsequent hierarchical encoder-decoder network. Finally, the output of the hierarchical encoder-decoder network is mapped back to a 3×H×W size through an output mapping module, and the final output is obtained by connecting it to the residual of the original input.

[0071] In this embodiment, the prior information extraction module includes a residual channel transformation unit, an input convolutional layer, an SE-ResBlock stacked model, and an output convolutional layer connected in sequence. The SE-ResBlock stacked model includes several SE-ResBlock structures, each comprising two convolutional layers: Conv ReLU and SE-Block. The step in this embodiment, where the prior information extraction module extracts residual channel prior features from the initial input image, is as follows: Figure 3 As shown, the prior information extraction module first uses the residual channel conversion unit to convert the initial input image I∈R H×W×3 Convert to residual channel image I res ∈R H×W×1 The residual channel image is mapped to the feature space through the input convolutional layer to obtain a residual channel feature matrix containing residual channel information. The size of the input convolutional layer is 3×3. The conversion process of converting the initial input image into a residual channel image in this embodiment is as follows:

[0072]

[0073] In the formula, I c I d Used to represent the color channels of the initial input image.

[0074] Then, in this embodiment, the residual channel feature matrix is ​​input into the SE-ResBlock stacked model to learn the prior background structure information, and processed through the output convolutional layer to output the residual channel prior features. Specifically, the output convolutional layer is 3×3 in size. In this embodiment, n=3 SE-ResBlock structures are stacked to extract features. The SE-ResBlock structure can effectively extract information in spatial and channel dimensions, and achieve rain removal effect more accurately.

[0075] This embodiment will use the residual channel prior features. and the image mapping feature X input ∈R H ×W×D The step of inputting the residual channel prior features and the image mapping features into the prior knowledge fusion module for background structure feature enhancement to obtain prior fusion enhancement information includes:

[0076] The residual channel prior features are converted into query vectors through linear layers and depthwise separable convolutional layers, and auxiliary value vectors containing prior information about the background structure are generated.

[0077] The image mapping features are converted into key vectors and value vectors using linear layers and depthwise separable convolutional layers.

[0078] Calculate the dot product attention between the query vector and the key vector, and enhance the background structure prior information based on the dot product attention to obtain a prior attention map;

[0079] The prior attention map is element-wise multiplied with the auxiliary value vector and the value vector respectively to obtain the auxiliary value multiplication result and the value multiplication result. The auxiliary value multiplication result and the value multiplication result are concatenated to obtain the value concatenation feature. The value concatenation feature is input into the linear layer to obtain the multi-head attention output feature.

[0080] The multi-head attention output features are input into a local augmentation feedforward network for dimensional expansion and compression to obtain prior fusion augmentation information.

[0081] This embodiment is based on the Transformer architecture and uses Prior Fusion Attention (PFA) to replace the original Multi-Head Self-Attention in the prior information fusion module structure. Figure 4 This is a schematic diagram of the prior information fusion module structure, and its operation process is as follows:

[0082]

[0083] P = P′ + LeFF(LN(P′))

[0084] In the formula, P′ represents the output of the prior fusion attention PFA; LN represents layer normalization; and P represents the output of the local enhancement feedforward network LeFF, i.e., the prior fusion enhancement information.

[0085] The multi-head attention mechanism in the standard Transformer architecture creates query, key, and value vectors from the input information. It then calculates an attention score by performing a dot product attention calculation between the query and key to capture the relationships between the input's own features. In this embodiment, the original input is replaced with residual channel prior features containing richer background structure information as the query vector, and a key is generated from the original input to calculate the dot product attention between them, resulting in an attention map enhanced with background structure information. The prior fusion attention calculation process is as follows: Figure 4 As shown on the right, specifically, the residual channel prior features are first processed through a combination of linear layers and deep convolutional layers. Convert to query vector And an auxiliary value vector was generated. To store prior feature-related information, the auxiliary value vector contains richer prior information about the background structure. Simultaneously, this embodiment employs similar linear layers and deep convolutional layers to convert image mapping features into key vectors. Sum value vector It should be noted that this embodiment preferentially uses depthwise separable convolutional layers as depthwise convolutional layers, enabling the network to more effectively capture local features of the image; then, the query vector generated from the residual channel prior features is calculated. and the key vector generated from the image mapping features The dot product attention between the two vectors yields a prior attention map M that enhances the background structure information. The prior attention map M is then compared with the value vector. and auxiliary value vector Element-wise multiplication is performed to enhance background structural information and guide subsequent feature extraction, yielding auxiliary value multiplication results and value multiplication results. It should be noted that multi-head operation is used here. Finally, the two sets of multiplication results are concatenated and passed through a linear layer to obtain the final multi-head attention output features. The computation process is as follows:

[0086]

[0087]

[0088]

[0089]

[0090] In the formula, M jThis represents the prior attention map output by the j-th attention head; This represents the attention output feature of the j-th attention head; the subscript j is used to label the attention head; W s This represents a linear transformation matrix.

[0091] This embodiment sets up a Local-enhanced Feed-Forward Network (LeFF) after the prior fused attention PFA. It should be noted that the feed-forward network (FFN) is an important component of the Transformer. The feed-forward network in the standard Transformer contains two linear layers that expand and compress the dimensions of each token, and incorporates non-linear characteristics using activation functions. To enable the FFN to better utilize local context information, this embodiment adds a deep convolutional block to the FFN, forming the Local-enhanced Feed-Forward Network (LeFF). The multi-head attention output features are then input into the Local-enhanced Feed-Forward Network for dimensional expansion and compression to obtain the prior fused enhanced information. Specifically, the multi-head attention output features are input into the Local-enhanced Feed-Forward Network after layer normalization. The Local-enhanced Feed-Forward Network then processes the input layer-normalized multi-head attention output features Z... l ∈R H×W×C The extended multi-head attention output features are obtained after the first linear layer expansion. Where k is the expansion factor, the expanded multi-head attention output features are transformed into a two-dimensional input containing structural and positional information via tensor reshaping, and this two-dimensional input is fed into a 3*3 deep convolutional layer to capture contextual information to obtain Z′. l ∈R H×W×kC Then, after being flattened back and compressed through subsequent activation function layers and linear layers, the final output is obtained as the prior fusion enhancement information. The calculation process for LeFF is as follows:

[0092]

[0093]

[0094]

[0095] This embodiment uses a locally enhanced feedforward network to output feature maps with richer contextual information and can better capture local features of the input data. In addition, the locally enhanced feedforward network will adaptively weight and fuse the features to further improve the representation ability of the features.

[0096] Traditional Transformer-based image deraining methods primarily focus on spatial domain attention, neglecting to fully utilize frequency information. To address this issue, this embodiment employs a Wavelet-Spatial Dual Attention (WSDA) mechanism to replace the multi-head self-attention mechanism in the traditional Transformer, forming a Wavelet-Spatial Dual Attention-based Transformer module. This module leverages wavelet and spatial domain information to enhance feature extraction, improving the quality of recovered image details and textures. Furthermore, a Local-enhanced Feed-Forward Network (LeFF) is also used within this module. Therefore, the Wavelet-Spatial Dual Attention-based Transformer module comprises a layer normalization layer (LN), a Wavelet-Spatial Dual Attention module (WSDA), a layer normalization layer (LN), and a Local-enhanced Feed-Forward Network (LeFF) connected sequentially. The LeFF calculation process is similar to that described above. A schematic diagram of the Wavelet-Spatial Dual Attention-based Transformer module structure is shown below. Figure 5 As shown, the calculation process is as follows:

[0097] X′ l =X l-1 +WSDA(LN(X l-1 ))

[0098] X l =X′ l +LeFF(LN(X l-1 ))

[0099] In the formula, X l LN(·) represents the output of WSDA in the l-th Transformer module based on wavelet-spatial attention; LN(·) represents layer normalization; WSDA(·) represents wavelet-spatial dual attention; X l ′ represents the output of the locally enhanced feedforward network LeFF in the l-th wavelet-spatial attention-based Transformer module.

[0100] In this embodiment, the step of generating a rainless image by processing the prior fusion enhancement information through the hierarchical encoder-decoder network based on a multi-scale information interaction strategy includes:

[0101] The prior fusion enhancement information is used as the encoding input feature of the first encoding and decoding stage. In each encoding and decoding stage, the input encoding input feature is divided into wavelet domain and spatial domain by a Transformer module based on wavelet-space dual attention with a preset attention head segmentation ratio to obtain wavelet domain feature map and spatial domain feature map.

[0102] Based on the wavelet domain feature map and the spatial domain feature map, the wavelet domain multi-head attention output features and the spatial domain multi-head attention output features are calculated.

[0103] The wavelet domain attention output features and the spatial domain attention output features are connected in parallel, and wavelet-spatial dual attention output features at different scales are obtained through a local enhancement feedforward network. Between two adjacent encoding and decoding stages, the wavelet-spatial dual attention output features are obtained by convolution downsampling or transposed convolution upsampling to obtain multi-scale feature representations.

[0104] Based on a multi-scale information interaction strategy, multi-scale feature representations at different scales are bridged to obtain global multi-scale output features.

[0105] The output mapping module performs a residual connection between the global multi-scale output features and the initial input image to generate a rainless image.

[0106] In the standard multi-head attention mechanism of the Transformer module, the attention score is obtained by performing a dot product operation on all features in the spatial domain. However, this attention calculation method cannot fully extract detailed and texture information in image deraining tasks. Furthermore, calculating on all pixels in high-resolution images incurs significant overhead. Therefore, the key to optimizing network performance lies in designing a more efficient attention calculation mechanism. To effectively extract information from degraded images, this embodiment divides the input into wavelet domain and spatial domain for processing. The calculation process is as follows: Figure 5 As shown on the right, in a conventional multi-head attention mechanism, all attention heads typically share the key and value. However, in this embodiment, the attention heads are divided into two groups with an attention head split ratio of α, where α·N h The head is used for wavelet domain attention calculation, and the rest (1-α)·N h The head is used for spatial domain attention calculation, after obtaining the wavelet domain multi-head attention (WA) output features WA. α (X) and spatial domain multi-head attention (SA) output features SA 1―α After (X)], it is concatenated in parallel to obtain the final wavelet-spatial dual attention output feature, the mathematical expression of which is:

[0107] WSDA(X) = Concat[WA α (X),SA 1―α (X)]

[0108] In this embodiment, the calculation process for wavelet domain attention and spatial domain attention includes the following steps: calculating the wavelet domain multi-head attention output features and the spatial domain multi-head attention output features based on the wavelet domain feature map and the spatial domain feature map.

[0109] The wavelet domain feature map is linearly transformed into a wavelet domain linear feature map, and a wavelet domain query vector is generated.

[0110] The discrete wavelet transform is used to decompose the wavelet domain linear feature map into four wavelet sub-bands. The four wavelet sub-bands are then connected side by side along the channel dimension to obtain the wavelet sub-band splicing feature.

[0111] A convolutional layer is used to apply spatial locality to the wavelet subband splicing features to obtain local context features, and the local context features are linearly transformed into wavelet domain key vectors and wavelet domain value vectors according to the preset attention head segmentation ratio.

[0112] Multi-head attention is calculated for each group of wavelet domain key vectors, wavelet domain value vectors and wavelet domain query vectors to obtain the wavelet domain multi-head attention output features.

[0113] The spatial domain feature map is windowed into several spatial domain feature matrices, and three parallel linear layers are used to perform linear transformation on the spatial domain feature matrices to obtain spatial domain key vectors, spatial domain value vectors, and spatial domain query vectors.

[0114] Multi-head attention is performed on each set of spatial domain key vectors, spatial domain value vectors, and spatial domain query vectors to obtain the spatial domain multi-head attention output features.

[0115] To better understand the computational details of wavelet domain attention and spatial domain attention, this embodiment will describe in detail the computational details of the output features of wavelet domain multi-head attention and spatial domain multi-head attention respectively. For wavelet domain attention, this embodiment uses 2D wavelet domain feature maps... The embedding matrix is ​​linearly transformed into a wavelet domain linear feature map. Then, in this embodiment, the discrete wavelet transform (DWT) is used to transform the input wavelet domain linear feature map. Decomposed into four wavelet subbands X LL X LH X HL , The four wavelet subbands are then connected side-by-side along the channel dimension to obtain the wavelet subband splicing feature. The resolution of the wavelet subband after wavelet transform is downsampled by half compared to the original image. In this embodiment, the classic Haar wavelet is preferred for Discrete Wavelet Transform (DWT). Then, a 3*3 convolution is applied to apply spatial locality to the image. The local context feature X is obtained above. c And according to the preset attention head segmentation ratio, X c Linear transformation into wavelet domain key vector and wavelet domain value vector Simultaneously, the wavelet domain feature map X∈R H×W×D Perform a linear transformation to obtain the wavelet domain query vector Q. w ∈R H×W×(α·D) Similarly, the query based on wavelet multi-head attention learning is performed on the downsampled key / value corresponding to each head. To reduce the learning difficulty and increase the richness of information, this embodiment also performs downsampled feature X after the local context features. c Applying inverse wavelet transform, and based on wavelet reconstruction theory, the reconstructed feature map X is obtained. r The mathematical expression for the multi-head attention computation process in the wavelet domain, which preserves every detail of the original input, is as follows:

[0116]

[0117]

[0118]

[0119] In the formula, W represents the output feature of the m-th attention head in the wavelet domain; w This represents the wavelet domain linear transformation matrix.

[0120] In the Discrete Wavelet Transform (DWT) decomposition process, a low-pass filter is first applied along the row direction of the DWT. and high-pass filter Wavelet domain linear feature map Encoded as two wavelet subbands X L and X H Next, along the wavelet subband X L and X H The columns use the same low-pass filter f. L and low-pass filter f H We obtain all four wavelet subbands: Among them, X LL X is a low-frequency component that reflects the basic object structure at a coarse-grained level. LH X HL X HHThis refers to a high-frequency component that preserves the texture details of an object at a fine-grained level.

[0121] While wavelet attention mechanisms are helpful in capturing directional high-frequency information and global low-frequency information, they have limitations in extracting local detail features of images. Therefore, this embodiment sets up a spatial attention branch to supplement local detail information in space. In image restoration problems, the correlation between local pixels is very important. Therefore, in the specific calculation process of spatial domain attention, this embodiment performs multi-head attention calculation on locally windowed pixels, rather than performing global attention on all tokens as in standard multi-head attention. This embodiment adopts simple non-overlapping window segmentation. Compared with time-consuming window shifting or multi-scale window segmentation, this method is more hardware-friendly. The specific process is as follows:

[0122] In this embodiment, the input 2D spatial domain feature map is first windowed into... A size of H win ×W win The spatial domain feature matrix is ​​multiplied by D, and then passed through three independent linear layers to obtain the spatial domain query vector, spatial domain key vector, and spatial domain value vector. Similar to the attention calculation process in the wavelet domain, this embodiment performs intra-domain multi-head attention calculation on each set of spatial domain query vector, spatial domain key vector, and spatial domain value vector to obtain the spatial domain multi-head attention output features. The mathematical expression for the calculation process is as follows:

[0123] SA 1―α (X) = Concat(head0) s head1 s ,…,head (1―α)Nh W s

[0124] head n s =Attention s (Q n s ,K n s V n s )

[0125]

[0126] In the formula, head n s W represents the output feature of the nth attention head in spatial domain attention; s This represents the linear transformation matrix in the spatial domain.

[0127] After calculating the wavelet domain multi-head attention output features and the spatial domain multi-head attention output features, this embodiment concats the wavelet domain multi-head attention output features and the spatial domain multi-head attention output features together to obtain the final output wavelet-spatial dual attention output features.

[0128] While the U-Shaped architecture (i.e., encoding downsampling followed by decoding upsampling, with features of the same dimension connected via skip connections) commonly found in Transformer-based image restoration methods has advantages in handling multi-scale information, the skip connections in the U-shaped structure can lead to significant semantic differences between the encoder and decoder. Therefore, to address the common U-Shaped architecture problem in Transformer-based image restoration methods, this embodiment designs a multi-scale information interaction strategy to improve image restoration performance. The steps of bridging multi-scale feature representations at different scales at various stages based on the multi-scale information interaction strategy to obtain global multi-scale output features include:

[0129] The output of the last encoding and decoding stage is used as the global feature, and a multi-scale information matrix is ​​obtained based on the multi-scale feature representation at different scales.

[0130] The global features and the multi-scale information matrix are sequentially input into the layer normalization, the linear layer and the activation function to obtain the corresponding standard global features and standard multi-scale information matrix.

[0131] A multi-axis cross-gating module is used to interactively operate the standard global features and the standard multi-scale information matrix to obtain global multi-scale output features, which include global semantically enhanced multi-scale output features and semantically optimized global output features.

[0132] In this embodiment, the encoding and decoding processes of the hierarchical encoder-decoder network each include four stages, with each stage containing N... iThere are [1,2,3,4] wavelet-spatial dual attention Transformer modules. These modules are used to extract features at different scales. Between each stage, the feature information is obtained through convolution downsampling or transposed convolution upsampling to obtain a multi-scale representation of the features. In terms of network structures for processing hierarchical multi-scale information, the classic U-shaped structure achieves information interaction by directly connecting information of the same scale through skip connections. However, such direct skip connections will lead to insufficient information interaction at different scales in image restoration problems, and will also cause excessive semantic gaps due to skip connections. Therefore, this embodiment introduces a multi-scale information interaction strategy, which is implemented through a multi-scale information interaction module, effectively alleviating the above problems. Finally, the output of the hierarchical encoder-decoder network is mapped back to a size of 3×H×W through an output mapping module, and the final output is obtained through a residual connection with the original input.

[0133] For the output of the encoder in its four stages, the multi-scale information interaction strategy uses a multi-scale information interaction module to bridge information at different scales. The overall structure of the multi-scale information interaction module is as follows: Figure 6 As shown, the input to the multi-scale information interaction module in the i-th stage includes the global feature Y. i and the multi-scale information matrix X of each stage of the encoder i For global information Y i This embodiment uses the output of the last stage as the global feature and interacts with the features of all previous stages. For different stages, this embodiment adjusts the dimensions through upsampling. Note that, except for the last stage, the input of the CGB global information for the previous stages is the global information processed by CGB in the previous stage. For the multi-scale information matrix X... i In this embodiment, the output E of each encoder is first processed. i∈1,2,3 By upsampling / downsampling, the outputs of each stage are adjusted to have the same dimension. Then, the channel dimension is adjusted through parallel concatenation and linear layers to obtain a multi-scale information matrix X containing multi-scale information. i Its mathematical expression is as follows:

[0134] X i =Linear(concat[E i ,scale(E j ),scale(E k )])

[0135]

[0136] In the formula, scale(·) represents the upsampling / downsampling operation; G iE represents the input of the multi-scale information interaction module in the i-th stage; j E k This represents the output of the encoder at stages other than the i-th stage.

[0137] This embodiment will use the global feature Y i and the multi-scale information matrix X of each stage of the encoder i As a multi-scale information interaction module G i Input, such as Figure 6 As shown, Y i and X i Y is obtained after passing through LN, linear layer and activation function respectively. i ′ and X i Then, it enters the multi-axis cross gating module for interactive computation. After processing by the multi-axis gating module, the globally semantically enhanced multi-scale output features are obtained. And semantically optimized global output features with increased semantic richness The exchange of different semantic information was achieved, and its computation process can be summarized as follows:

[0138] X′ i =GELU(linear(LN(X) i )))

[0139] Y i =GELU(linear(LN(Y)) i )))

[0140]

[0141]

[0142] In the formula, represents element-wise dot product; g(·) represents multi-axis cross-gating module;

[0143] The computation process of the multi-axis cross-gating module is similar to that of MLP, but the difference lies in transforming the processed information dimension from channel dimension to spatial dimension between two linear layers. It uses features segmented into modular and rasterized spatial dimensions as input to achieve interaction between local and global spatial dimension features. Finally, and The final output Y is obtained through a linear layer and a residual connection to the input. i+1 and X i+1It should be noted that, in this embodiment, after obtaining the global multi-scale output features using the multi-scale information interaction module, the input of the i-th level decoder should include the upsampled output of the previous level decoder and the outputs of all multi-scale information interaction modules containing multi-scale information. Similar to the input of the multi-scale information interaction module, the multi-scale information is connected in parallel and the channel size is adjusted using linear operations.

[0144] The prior knowledge-guided wavelet-spatial dual-attention image deraining method proposed in this embodiment is a fully supervised approach. It obtains predicted values ​​by inputting negative samples into the network, then calculates the loss function between the predicted values ​​and positive samples. The network parameters are updated using backpropagation and gradient descent algorithms until the network converges. Regarding the loss function, this embodiment employs a composite loss function combining L1-loss, MSE-loss, and SSIM-loss to train the prior knowledge-guided wavelet-spatial dual-attention deraining deep learning network model. This composite loss function maximizes the performance of the deraining deep learning network model, improving its accuracy and generalization ability. The total loss function is defined as follows:

[0145] L total =L1+L MSE +L SSIM

[0146]

[0147] L MSE =||x pre ―x gt || 2

[0148] L SSIM =―SSIM(x pre ―x gt )

[0149] In the formula, L total L1 represents the total loss function; L1 represents the L1 loss function; L MSE L represents the mean squared error loss function; SSIM Denotes the structural similarity loss function; x pre Indicates the predicted value; x gt The true value is represented by SSIM, which stands for Structural Similarity Index. SSIM is an index used to measure the similarity between two images. It models the structural information of an image as a combination of three different factors: luminance, contrast, and structure. Specifically, SSIM uses the image mean as an estimate of luminance, the standard deviation as an estimate of contrast, and the covariance as an estimate of structural similarity. Its calculation formula is as follows:

[0150]

[0151]

[0152]

[0153] In the formula, x and y represent two images of the structural similarity to be evaluated; μ x μ y σ represents the mean of the two images. x σ y σ represents the variance of the two images. xy represents the covariance of image x and image y; c represents constant to avoid cases where the denominator is zero.

[0154] From the above formula, the calculation formula for SSIM can be obtained as follows:

[0155] SSIM(x,y)=[l(x,y) η ·c(x,y) β ·s(x,y) γ ]

[0156] If we make η = β = γ = 1, then we have:

[0157]

[0158] In the actual training process of the network, this embodiment uses AdamW as the optimizer, and the initial learning rate is set to 4×10. -4 Then, the learning rate is gradually reduced to 2×10 using a cosine annealing strategy. -5 In this embodiment, 400 rounds of iterative training were performed. After actual verification, the network model reached a convergent state after 400 iterations and can effectively complete the image deraining task.

[0159] This invention provides a wavelet-spatial dual-attention image deraining method guided by prior knowledge. The method acquires the original rainfall image as the initial input image and uses a wavelet-spatial dual-attention deraining deep learning network model guided by prior knowledge to process the initial input image for deraining, resulting in a rain-free image. The wavelet-spatial dual-attention deraining deep learning network model includes an input mapping module, a prior knowledge guidance module, a hierarchical encoder-decoder network based on wavelet-spatial dual attention, and an output mapping module. Compared with existing technologies, the wavelet-spatial dual-attention image deraining method proposed in this embodiment, combined with a prior knowledge guidance mechanism, not only effectively guides the network learning under rainy conditions using residual channel priors containing rich structural information, but also effectively extracts spatial and frequency domain information from the rainfall image. Furthermore, this embodiment solves the problems of insufficient information interaction at different scales and excessive semantic gaps caused by skip connections in multi-scale deep learning networks through a multi-scale feature interaction strategy, reducing the difficulty of network learning, improving the accuracy of image deraining, and making the deraining effect more stable and reliable. It has advantages such as efficient deraining, strong generalization ability, and saving computational resources.

[0160] It should be noted that the sequence number of each process does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0161] In one embodiment, such as Figure 7 As shown, this embodiment of the invention provides a wavelet-spatial dual-attention image deraining system guided by prior knowledge. The system includes:

[0162] Image acquisition module 101 is used to acquire the original rainfall image as the initial input image;

[0163] The rain removal processing module 102 is used to perform rain removal processing on the initial input image using a wavelet-spatial dual-attention rain removal deep learning network model guided by prior knowledge, to obtain a rain-free image; wherein, the wavelet-spatial dual-attention rain removal deep learning network model guided by prior knowledge includes an input mapping module, a prior knowledge guidance module, and a hierarchical encoder-decoder network based on wavelet-spatial dual attention.

[0164] Specific limitations regarding a priori knowledge-guided wavelet-spatial dual-attention image deraining system can be found in the above-described limitations regarding a priori knowledge-guided wavelet-spatial dual-attention image deraining method, and will not be repeated here. Those skilled in the art will recognize that the various modules and steps described in conjunction with the embodiments disclosed in this application can be implemented in hardware, software, or a combination of both. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0165] This invention provides a wavelet-spatial dual-attention image deraining system guided by prior knowledge. The system utilizes a wavelet-spatial dual-attention deraining deep learning network model guided by prior knowledge to process an initial input image for deraining, resulting in a rain-free image. The wavelet-spatial dual-attention deraining deep learning network model includes an input mapping module, a prior knowledge guidance module, and a hierarchical encoder-decoder network based on wavelet-spatial dual attention. Compared to existing technologies, this application combines a prior knowledge guidance mechanism, wavelet-spatial dual attention, and a multi-scale feature interaction strategy to construct the wavelet-spatial dual-attention deraining deep learning network model. Wavelet-spatial dual attention effectively extracts spatial and frequency domain information from the rain image, improving the deraining accuracy. The prior knowledge guidance mechanism ensures the recovered image retains more complete structural information, while the multi-scale feature interaction strategy reduces the learning difficulty of the network, making the deraining effect more stable and reliable, improving the overall image quality and visual effect, and adapting to various complex rainy weather scenarios.

[0166] Figure 8 This invention provides a computer device including a memory, a processor, and a transceiver, which are connected to each other via a bus. The memory is used to store a set of computer program instructions and data, and can transmit the stored data to the processor. The processor can execute the program instructions stored in the memory to perform the steps of the above method.

[0167] The memory may include volatile memory or non-volatile memory, or both; the processor may be a central processing unit, a microprocessor, an application-specific integrated circuit, a programmable logic device, or a combination thereof. By way of example, but not limitation, the programmable logic device described above may be a complex programmable logic device, a field-programmable gate array, a general-purpose array logic, or any combination thereof.

[0168] In addition, memory can be a physically independent unit or integrated with the processor.

[0169] Those skilled in the art will understand that Figure 8 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have the same component arrangement.

[0170] In one embodiment, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described method.

[0171] This invention provides a prior knowledge-guided wavelet-spatial dual-attention image deraining method and system. The prior knowledge-guided wavelet-spatial dual-attention image deraining method constructs a wavelet-spatial dual-attention deraining deep learning network model based on wavelet-spatial dual attention, a prior knowledge-guided mechanism, and a multi-scale feature interaction strategy. By extracting spatial and frequency domain information from rainfall images and combining this with prior knowledge-guided learning of residual channel prior features that enrich structural information, the method better understands the structure and features of rainy images, reduces the difficulty of network learning, makes deraining processing more efficient, and improves the performance and accuracy of the image deraining model.

[0172] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., SSD), etc.

[0173] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when the computer program is executed, it can include the processes of the embodiments of the above methods.

[0174] The embodiments described above are merely preferred embodiments of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various improvements and substitutions without departing from the technical principles of this invention, and these improvements and substitutions should also be considered within the scope of protection of this application. Therefore, the scope of protection of this patent application should be determined by the scope of the claims.

Claims

1. A wavelet-spatial dual-attention image rain removal method guided by prior knowledge, characterized in that, Includes the following steps: Obtain the original rainfall image as the initial input image; The initial input image is processed to remove rain using a wavelet-spatial dual-attention rain removal deep learning network model guided by prior knowledge, resulting in a rain-free image. The deep learning network model for rain removal based on prior knowledge-guided wavelet-spatial dual attention includes an input mapping module, a prior knowledge guidance module, a hierarchical encoder-decoder network based on wavelet-spatial dual attention, and an output mapping module. The prior knowledge guidance module includes a cascaded prior information extraction module and a prior knowledge fusion module. The prior information extraction module includes a residual channel transformation unit, an input convolutional layer, an SE-ResBlock stacked model, and an output convolutional layer connected in sequence. The SE-ResBlock stacked model includes several SE-ResBlock structures, each of which includes two convolutional layers and an SE-Block. The hierarchical encoder-decoder network includes multiple encoding and decoding stages and a multi-scale information interaction module connecting adjacent encoding and decoding stages. Each encoding and decoding stage includes a Transformer module based on wavelet-spatial dual attention. The steps of using a wavelet-spatial dual-attention deraining deep learning network model guided by prior knowledge to process the initial input image for deraining and obtain a rain-free image include: The initial input image is mapped to the feature space through the input mapping module to obtain image mapping features. At the same time, the initial input image is input in parallel into the prior information extraction module for residual channel prior feature extraction. The residual channel prior features and the image mapping features are input into the prior knowledge fusion module to enhance the background structure features, thereby obtaining prior fusion enhancement information; Based on a multi-scale information interaction strategy, the prior fusion enhancement information is processed through the hierarchical encoder-decoder network to generate a rainless image; The step of generating a rainless image by processing the prior fusion enhancement information through the hierarchical encoder-decoder network based on the multi-scale information interaction strategy includes: The prior fusion enhancement information is used as the encoding input feature of the first encoding and decoding stage. In each encoding and decoding stage, the input encoding input feature is divided into wavelet domain and spatial domain by a Transformer module based on wavelet-space dual attention with a preset attention head segmentation ratio to obtain wavelet domain feature map and spatial domain feature map. Based on the wavelet domain feature map and the spatial domain feature map, the wavelet domain multi-head attention output features and the spatial domain multi-head attention output features are calculated. The wavelet domain attention output features and the spatial domain attention output features are connected in parallel, and wavelet-spatial dual attention output features at different scales are obtained through a local enhancement feedforward network. Between two adjacent encoding and decoding stages, the wavelet-spatial dual attention output features are obtained by convolution downsampling or transposed convolution upsampling to obtain multi-scale feature representations. Based on a multi-scale information interaction strategy, multi-scale feature representations at different scales are bridged to obtain global multi-scale output features. The output mapping module performs a residual connection between the global multi-scale output features and the initial input image to generate a rainless image. The steps of calculating the wavelet domain multi-head attention output features and the spatial domain multi-head attention output features based on the wavelet domain feature map and the spatial domain feature map include: The wavelet domain feature map is linearly transformed into a wavelet domain linear feature map, and a wavelet domain query vector is generated. The discrete wavelet transform is used to decompose the wavelet domain linear feature map into four wavelet sub-bands. The four wavelet sub-bands are then connected side by side along the channel dimension to obtain the wavelet sub-band splicing feature. A convolutional layer is used to apply spatial locality to the wavelet subband splicing features to obtain local context features, and the local context features are linearly transformed into wavelet domain key vectors and wavelet domain value vectors according to the preset attention head segmentation ratio. Multi-head attention is calculated for each group of wavelet domain key vectors, wavelet domain value vectors and wavelet domain query vectors to obtain the wavelet domain multi-head attention output features. The spatial domain feature map is windowed into several spatial domain feature matrices, and three parallel linear layers are used to perform linear transformation on the spatial domain feature matrices to obtain spatial domain key vectors, spatial domain value vectors, and spatial domain query vectors. Multi-head attention is performed on each set of spatial domain key vectors, spatial domain value vectors, and spatial domain query vectors to obtain the spatial domain multi-head attention output features.

2. The wavelet-spatial dual-attention image deraining method guided by prior knowledge as described in claim 1, characterized in that, The step of inputting the initial input image in parallel into the prior information extraction module for residual channel prior feature extraction includes: The initial input image is converted into a residual channel image by the residual channel conversion unit, and the residual channel image is mapped to the feature space by the input convolutional layer to obtain a residual channel feature matrix containing residual channel information. The residual channel feature matrix is ​​input into the SE-ResBlock stacked model to learn prior background structure information, and then processed through the output convolutional layer to output the residual channel prior features.

3. The wavelet-spatial dual-attention image deraining method guided by prior knowledge as described in claim 1, characterized in that, The step of inputting the residual channel prior features and the image mapping features into the prior knowledge fusion module for background structure feature enhancement to obtain prior fusion enhancement information includes: The residual channel prior features are converted into query vectors through linear layers and depthwise separable convolutional layers, and auxiliary value vectors containing prior information about the background structure are generated. The image mapping features are converted into key vectors and value vectors using linear layers and depthwise separable convolutional layers. Calculate the dot product attention between the query vector and the key vector, and enhance the background structure prior information based on the dot product attention to obtain a prior attention map; The prior attention map is element-wise multiplied with the auxiliary value vector and the value vector respectively to obtain the auxiliary value multiplication result and the value multiplication result. The auxiliary value multiplication result and the value multiplication result are concatenated to obtain the value concatenation feature. The value concatenation feature is input into the linear layer to obtain the multi-head attention output feature. The multi-head attention output features are input into a local augmentation feedforward network for dimensional expansion and compression to obtain prior fusion augmentation information.

4. The wavelet-spatial dual-attention image deraining method guided by prior knowledge as described in claim 1, characterized in that, The steps for bridging multi-scale feature representations at different scales to obtain global multi-scale output features based on the multi-scale information interaction strategy include: The output of the last encoding and decoding stage is used as the global feature, and a multi-scale information matrix is ​​obtained based on the multi-scale feature representation at different scales. The global features and the multi-scale information matrix are sequentially input into the layer normalization, the linear layer and the activation function to obtain the corresponding standard global features and standard multi-scale information matrix. A multi-axis cross-gating module is used to interactively operate the standard global features and the standard multi-scale information matrix to obtain global multi-scale output features, which include global semantically enhanced multi-scale output features and semantically optimized global output features.

5. A wavelet-spatial dual-attention image deraining system guided by prior knowledge, characterized in that, The system employs the prior knowledge-guided wavelet-spatial dual attention image deraining method as described in any one of claims 1 to 4, wherein the system comprises: The image acquisition module is used to acquire the original rainfall image as the initial input image; The rain removal module is used to process the initial input image using a wavelet-spatial dual-attention rain removal deep learning network model guided by prior knowledge, so as to obtain a rain-free image. The wavelet-spatial dual attention rain removal deep learning network model based on prior knowledge guidance includes an input mapping module, a prior knowledge guidance module, a hierarchical encoder-decoder network based on wavelet-spatial dual attention, and an output mapping module.

6. A computer device, characterized in that: The device includes a processor and a memory, the processor being connected to the memory for storing computer programs, and the processor for executing the computer programs stored in the memory to cause the computer device to perform the method as described in any one of claims 1 to 4.