Multi-stage image denoising method based on cascaded double subnets

Through the multi-stage image denoising method of cascading Gemini networks, combined with the spatial detail subnet and the cross-scale fusion subnet, the shortcomings of the image denoising method in the prior art in terms of local details and global feature fusion are solved, and a higher quality image denoising effect is achieved.

CN120339108APending Publication Date: 2025-07-18TIANJIN UNIVERSITY OF FINANCE AND ECONOMICS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510428578.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing image denoising method based on a single network structure has insufficient in the balanced fusion of non-local context features and local spatial details, resulting in poor high-frequency detail extraction results, and the encoder-decoder architecture has a characterization performance imbalance between global semantic information modeling and local structure maintenance.

Method used

Using a multi-stage image denoising method of cascading Gemini networks, combining spatial detail subnets and cross-scale fusion subnets, local spatial features are extracted and non-local context features are fused, and model performance is optimized using edge-weighted loss function.

Benefits of technology

Improve image denoising quality, enhance spatial detail information extraction and multi-scale non-local context information fusion, and improve image denoising effect and visual effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339108A_ABST
    Figure CN120339108A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of image de-noising, and discloses a multi-stage image de-noising method based on a cascaded double subnet, and the method comprises the following specific steps: 1, constructing a multi-stage de-noising network of the cascaded double subnet, the de-noising network comprises n de-noising stages, n belongs to [2, 5], and each de-noising stage comprises a space detail subnetwork and a cross-scale fusion subnetwork; according to the method, a multi-stage framework is provided, an image denoising task is decomposed into a plurality of sub-tasks, a user can select and use several stages to realize denoising according to actual needs of the user, and the more the stages are, the better the denoising effect is; by utilizing the advantages of branch network addition, a single-scale network is improved, and space detail information extraction is enriched; a network is constructed by using a cross-scale fusion method, multi-scale non-local context information is fused, and the image denoising quality is improved; an edge weighted loss function is provided, and compared with common pixel-based loss, the overall visual effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image denoising, and specifically relates to a multi-stage image denoising method based on a cascaded dual subnet. Background Technique

[0002] Image denoising is a classic inverse problem in low-level computer vision, aiming to recover a clean image x from a noisy image y containing noise n, which can be expressed as x = y - n, where y is the noisy image, x is the potential clean image, and n is the noise. Since both the noise and the clean image are unknown, the denoising task is highly ill-posed. Removing noise and reasonably preserving the detailed information of the image is crucial for high-quality image denoising, and it can be used as a preprocessing step for subsequent high-level computer vision tasks.

[0003] Currently, the image denoising methods based on multi-stage networks have shown significant advantages in the balanced fusion of non-local context features and local spatial details. However, the models of traditional convolutional neural networks based on a single network structure have insufficient ability to extract high-frequency details, resulting in poor local spatial detail extraction effects; although the models based on the encoder-decoder architecture can capture large-scale context features through hierarchical feature abstraction, their inherent single-scale information transmission mechanism still has a problem of imbalance in representation efficiency between global semantic information modeling and local structure preservation, which mainly stems from the following two aspects: First, the step-by-step downsampling operation in the encoding and decoding processes is prone to irreversible loss of high-frequency details; second, it is difficult to achieve deep fusion of multi-granularity visual features after simply adding the skip connection mechanism for cross-level feature interaction. Summary of the Invention

[0004] The purpose of the present invention is to provide a multi-stage image denoising method based on a cascaded dual subnet to solve the problems proposed in the above background technique.

[0005] To achieve the above purpose, the present invention provides the following technical solution: A multi-stage image denoising method based on a cascaded dual subnet, and the specific steps are as follows:

[0006] Step 1:

[0007] Construct a multi-stage denoising network of a cascaded dual subnet. The denoising network includes: n denoising stages, n ∈ [2, 5], and each denoising stage includes: a spatial detail subnet and a cross-scale fusion subnet. The spatial detail subnet includes: an edge enhancement branch network based on U-Net and a backbone network of a cascaded DRCAB module. The edge enhancement branch also returns edge information for calculating the edge weighted loss function.

[0008] Step 2:

[0009] Extract shallow features before the start of each stage. In each stage, a convolutional layer is used to extract the shallow features F of the noisy image y, and the formula is as follows:

[0010] F = W * y

[0011] where W is the convolutional kernel that expands the number of feature maps, and * represents the convolution operation;

[0012] Step 3:

[0013] The shallow features of the first stage first pass through the spatial detail sub-network with an edge enhancement branch to obtain the local spatial features of the image, and then pass through a cross-scale fusion sub-network to fully fuse the non-local context features;

[0014] Step 4:

[0015] The features obtained in the first denoising stage are added to the shallow features obtained in the second stage, and then first pass through the spatial detail sub-network with an edge enhancement branch to obtain the local spatial features of the image, and then pass through a cross-scale fusion sub-network to fully fuse the non-local context features;

[0016] Step 5:

[0017] Repeat the above steps until the features of the nth denoising stage are obtained. The features output in the nth stage are processed by a convolutional layer to obtain the clear image x n .

[0018] As a preferred technical solution of the present invention, the spatial detail sub-network described in step one is divided into two parts: the backbone network of the cascaded DRCAB module and the edge enhancement branch based on U-Net. The backbone network composed of the cascaded DRCAB is used to extract potential local features and realize the recovery of local spatial information of the image; the edge enhancement branch based on U-Net is used to supplement edge texture features and enrich local spatial features.

[0019] As a preferred technical solution of the present invention, the specific method of the edge enhancement branch network based on U-Net described in step one is as follows: First, in the downsampling stage, it passes through three encoding blocks in sequence. Each encoding block includes, in sequence: a convolutional layer - BN layer - ELU activation function, a convolutional layer - BN layer - ELU activation function, a convolutional layer - BN layer - ReLU activation function. Each encoding block reduces the size of the feature map by increasing the stride and doubles the number of channels using the convolution operation. The formula is as follows:

[0020]

[0021] where y k-1 represents the initial feature map obtained by the (k - 1)th downsampling, Denote the feature map obtained in the k-th downsampling process, y k Denote the feature map obtained by the k-th downsampling, where W1, W2, W3 represent the convolutional kernel parameters, b1, b2, b3 represent the convolutional kernel bias terms, k ∈ {1, 2, 3}, BN represents the batch normalization layer, ELU represents the ELU activation function, δ represents the ReLU activation function, * represents the convolution operation, y k ∈R C×H×W ,y k-1 ∈R C×H×W Denote the dimension of the feature map, where C represents the number of channels of the feature map, H represents the height of the feature map, and W represents the width of the feature map. Simplify the above process to: y k =DS(y k-1 ), then the formula for the downsampling process is as follows:

[0022] y3=DS(DS(DS(F)))

[0023] where DS represents downsampling, the initial feature map F = x0, and the feature maps x3 are obtained after three consecutive downsamplings;

[0024] Then, through two convolution operations, only double and halve the number of channels respectively, and the formula is as follows:

[0025]

[0026] where, Denote the obtained feature map, W4, W5 represent the convolutional kernel parameters, b4, b5 represent the convolutional kernel bias terms, BN represents the batch normalization layer, ELU represents the ELU activation function, δ represents the ReLU activation function, * represents the convolution operation;

[0027] Finally, in the upsampling stage, three decoding blocks are also passed through. Each decoding block sequentially includes: transposed convolutional layer - BN layer - ReLU activation function, convolutional layer - BN layer - ELU activation function, convolutional layer - BN layer - ELU activation function. First, fuse the multi-scale features in the encoding stage, and then use the transposed convolution to enlarge the size of the feature map. The formula is as follows:

[0028]

[0029] where, Denote the feature map obtained after the k-th upsampling, Denote the k-th upsampling transposed convolution operation, W6, W7 represent the convolutional kernel parameters, b6, b7 represent the convolutional kernel bias terms, [] represents the concatenation operation, y 4-k Denote the feature map of x0 after the (4 - k)-th downsampling, Denote the feature map after the (4 - k)-th upsampling, is the feature map in the (4-k)-th downsampling process, where k ∈ {1, 2, 3}, BN represents the batch normalization layer, ELU represents the ELU activation function, δ represents the ReLU activation function. The above process is simplified as: Then the formula for the upsampling process is as follows:

[0030]

[0031] where ES represents upsampling, and is obtained by continuously performing upsampling three times and then passing through a convolutional layer to obtain the final output sp;

[0032] In summary, the above process is represented by the following formula:

[0033] sp = EEBNet(F)

[0034] where EEBNet() represents an edge enhancement branch network with a U-Net structure, sp represents the extracted edge features, and F represents the extracted shallow features;

[0035] The edge information is normalized before being fused into the backbone network to generate a mask mask, and its formula is:

[0036] mask = σ(w1 * sp) - w2 * σ(sp),

[0037] where σ represents the Sigmoid function, the function w1 represents the amplification factor, w2 represents the bias factor, * represents the convolution operation, and sp represents the edge information learned from EEBNet.

[0038] As a preferred technical solution of the present invention, the cascaded DRCAB module described in step one includes m DRCAB modules. Each DRCAB module sequentially includes three residual blocks, global average pooling, one residual block, sigmoid activation function, weighted operation, and long connection operation.

[0039] As a preferred technical solution of the present invention, the specific operation of each DRCAB module is as follows: First, use three consecutive residual blocks composed of a convolutional layer - ReLU activation function - convolutional layer and residual connection operation to extract the initial feature h of the input feature f, and the formula is as follows:

[0040] h1 = W2 * (δ(W1 * f + b1)) + b2

[0041] h2 = W4 * (δ(W3 * (f + h1) + b3)) + b4

[0042] h3 = W 6* (δ(W5 * (f + h1 + h2) + b5)) + b6

[0043] h = h3 + h2 + h1

[0044] Among them, W1, W2, W3, W4, W5, W6 represent the convolutional kernel parameters, b1, b2, b3, b4, b5, b6 represent the convolutional kernel bias terms, δ represents the ReLU activation function, * represents the convolution operation, h1, h2, h3 represent the feature maps after passing through each residual block. Taking the first stage as an example, f = F1;

[0045] Then, perform global average pooling operation on the initial feature h for each channel, and the formula is as follows:

[0046]

[0047] Among them, GAP represents the global average pooling operation, h c represents the feature of the c-th channel, h c ∈h, h c (i, j) represents the feature value at the coordinate (i, j) of the c-th channel, z c represents the feature statistic of the c-th channel, where c ∈ {0, 1…, C};

[0048] Then, concatenate the feature statistics of all channels, and apply the residual block and Sigmoid activation function to extract the channel attention weight s, as shown in the following formula:

[0049] z = [z1, z2, …, z c

[0050] s = σ(W8 * (δ(W7 * z + b7)) + b8)

[0051] Among them, [] represents the concatenation operation, σ represents the sigmoid activation function, W7, W8 represent the convolutional kernel parameters respectively, b7, b8 represent the convolutional kernel bias terms, δ represents the ReLU activation function, * represents the convolution operation;

[0052] Then, use the channel attention weight s to weight the feature z to obtain as the output of the DRCAB module, as shown in the following formula:

[0053]

[0054] Among them, · represents the element-wise multiplication operation;

[0055] Finally, add a long connection from the initial feature h to The formula is as follows:

[0056]

[0057] ​In summary, the processing procedure of each DRCAB module can be simplified as: f DRCAB = DRCAB(f).

[0058] As a preferred technical solution of the present invention, taking the first stage as an example for the specific operation of each DRCAB module, the input of the first DRCAB module is the shallow feature F1. After the above operations, the output of the first DRCAB module is obtained. The output of the first DRCAB module is used as the input of the second DRCAB module. The above operations are looped, and finally the output of the cascaded DRCAB modules is obtained, as shown in the following formula:

[0059] F d = DRCAB m (DRCAB m-1 (…DRCAB1(F)))

[0060] where F d represents the local spatial feature of the image extracted by the cascaded DRCAB blocks, DRCAB1() represents the first DRCAB block, DRCAB m represents the m-th DRCAB block, and F represents the feature map extracted at the initial contact;

[0061] Finally, the local spatial detail information obtained by the cascaded DRCAB blocks is added to the mask mask1 obtained by the edge enhancement branch to obtain the final local spatial detail information. Taking the first stage as an example, the formula is as follows:

[0062] F 11 = W * (F d + mask1)

[0063] where F 11 represents the spatial detail feature extracted by the spatial detail subnetwork, W represents the convolution kernel, mask1 represents the mask output by the edge enhancement branch, and F d represents the local spatial feature of the image extracted by the cascaded DRCAB blocks.

[0064] As a preferred technical solution of the present invention, the specific steps of the cross-scale fusion subnetwork described in step three and step four are as follows:

[0065] First, the input feature X passes through dilated convolutions with dilation rates of 1, 2, 3, and 4 in parallel to generate multi-scale feature maps. The formula is expressed as follows:

[0066] F k = W k*d X + b k

[0067] where *d represents the dilated convolution operation with dilation rate d, and W kis the weight of the convolutional kernel, b k is the bias term of the convolutional kernel, d ∈ {1, 2, 3, 4}, and the input feature X = F 11 + F1;

[0068] Fuse the four feature maps obtained by the four dilated convolutions in pairs. For the feature maps F i and F j (i ∈ {1, 2}, j ∈ {3, 4}), as shown in the following formula:

[0069] G i,j = W i,j * Concat(F i , F j ) + b i,j

[0070] Among them, G i,j represents the fused feature map, Concat() represents the concatenation in the channel dimension, W i,j is the convolutional kernel, b i,j is the bias term, and * represents the convolution operation;

[0071] Fuse the fused feature maps again to obtain the feature map H, as shown in the following formula:

[0072] H = W h * Concat({G i,j}) + b h

[0073] Among them, W h is the convolutional kernel parameter, b h is the bias term of the convolutional kernel, and Concat() represents the concatenation in the channel dimension;

[0074] Then, respectively perform the channel attention-weighted operation and the spatial attention-weighted operation to extract the channel attention weight s1 and the spatial attention weight s2, as shown in the following formula:

[0075] s1 = σ(W2 * fc{W1 * GAP(h) + b1) + b2)

[0076]

[0077] Among them, represents the feature map after channel attention and spatial attention, W1, W2, W3 represent convolutional kernel parameters, b1, b2, b3 represent convolutional kernel bias terms, * represents the convolution operation, GAP() is the global average pooling, σ is the sigmoid activation function, fc represents the fully connected layer, and · represents the element-wise multiplication operation;

[0078] Finally, the obtained feature map As the input, it passes through the ReLU activation function - convolutional layer - DRCAB module - convolutional layer in sequence to obtain the final output result, and its formula is as follows:

[0079]

[0080] Among them, l represents the final output feature map, * represents the convolution operation, DRCAB() represents the DRCAB module, W4 and W5 represent the convolutional kernel parameters, b4 and b5 represent the convolutional kernel bias terms, and δ represents the ReLU activation function;

[0081] To sum up, the above process is represented by the following formula:

[0082] F s = CSFNet(X)

[0083] Among them, CSFNet() represents the cross-scale fusion subnet, X represents the input feature map, and F s represents the output feature map of the cross-scale fusion subnet.

[0084] As a preferred technical solution of the present invention, the n denoising stages described in step five are arranged as the first-stage denoising, the second-stage denoising,..., the nth-stage denoising. At the end of the first-stage denoising, the second-stage denoising,..., the nth-stage denoising, a convolutional layer is used to process the features extracted in each stage, and then the long connection from the original image y is added to obtain the clear images x1, x2,..., x corresponding to each denoising stage n , and the clarity of the clear images x1, x2,..., x n gradually increases.

[0085] As a preferred technical solution of the present invention, the calculation steps of the edge weighted loss function described in step one are as follows:

[0086] First, an improved Laplacian operator is used to train the edge enhancement branch to obtain the edge gradient, and its formula is as follows:

[0087]

[0088] Among them, F Laplace is the defined Laplacian operator;

[0089] Then, L edge is used to represent the edge loss of each stage, and its formula is as follows:

[0090] L edge = ||sp i - F Laplace (x gt )|| 2, where \(i\in(1,2,\cdots,n)\)

[0091] Among them, \(sp\) i represents the edge information obtained by the edge enhancement branch;

[0092] Finally, we use the L2 loss function to measure the denoised image \(x\) at each stage i (\(i\in\{1,2,\cdots,n\}\)) and the ground truth \(x\) gt The difference between them is as follows:

[0093] \(L\) out \(=\|x\) i - x gt \| 2 , where \(i\in(1,2,3)\)

[0094] Among them, \(L\) out represents the L2 loss function at each stage;

[0095] In summary, the proposed edge-weighted loss function is as follows:

[0096]

[0097] Among them, the parameter \(\theta\) controls the loss term of the edge loss, and \(L\) edge represents the edge loss at each stage, and \(L\) out represents the L2 loss at each stage.

[0098] The beneficial effects of the present invention are as follows:

[0099] The present invention proposes a multi-stage framework, which decomposes the image denoising task into multiple subtasks. Users can also choose to use several stages to achieve denoising according to their actual needs. The more stages are used, the better the denoising effect. By taking advantage of the added advantages of the branch network, the single-scale network is improved, and the extraction of spatial detail information is enriched. The network is constructed by using the method of cross-scale fusion, which fuses multi-scale non-local context information and greatly improves the image denoising quality. An edge-weighted loss function is proposed, which improves the overall visual effect compared with the ordinary pixel-based loss. Brief Description of the Drawings

[0100] Figure 1 is the overall architecture diagram of the multi-stage denoising network of the cascade dual subnet of the present invention;

[0101] Figure 2 is the architecture diagram of the EEBNet based on the U-Net edge enhancement branch network of the present invention;

[0102] Figure 3 is the structure diagram of the DRCAB module of the present invention;

[0103] Figure 4 This is the architecture diagram of the CSFNet for the cross-scale fusion sub-network of the present invention. Detailed implementation manners

[0104] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0105] As Figures 1 to 4 shown, the embodiments of the present invention provide a multi-stage image denoising method based on a cascaded twin network, and the specific steps are as follows:

[0106] Step 1:

[0107] Construct a multi-stage denoising network of a cascaded twin network. The denoising network includes: n denoising stages, where n ∈ [2, 5]. Each denoising stage includes: a spatial detail sub-network and a cross-scale fusion sub-network. The spatial detail sub-network includes: an edge enhancement branch network based on U-Net and a backbone network of cascaded DRCAB modules. The edge enhancement branch also returns edge information for calculating the edge weighted loss function;

[0108] Step 2:

[0109] Extract shallow features before the start of each stage. In each stage, a convolutional layer is used to extract the shallow features F of the noisy image y, and the formula is as follows:

[0110] F = W * y

[0111] where W is the convolutional kernel that expands the number of feature maps, and * represents the convolution operation;

[0112] Step 3:

[0113] The shallow features of the first stage first pass through the spatial detail sub-network with an edge enhancement branch to obtain the local spatial features of the image, and then pass through a cross-scale fusion sub-network to fully fuse the non-local context features;

[0114] Step 4:

[0115] The features obtained in the first denoising stage are added to the shallow features obtained in the second stage, and then first pass through the spatial detail sub-network with an edge enhancement branch to obtain the local spatial features of the image, and then pass through a cross-scale fusion sub-network to fully fuse the non-local context features;

[0116] Step 5:

[0117] Repeat the above steps until the features of the nth denoising stage are obtained. The features output in the nth stage are processed by a convolutional layer to obtain the clear image x n 。

[0118] Before each stage, the DRCAB module is used to extract the shallow features of the noisy image; the shallow features of the first stage are first passed through the spatial detail sub-network with an edge enhancement branch to extract local detail information; the edge enhancement branch is a U-Net-based network, which obtains edge information through training, calculates the difference as the edge loss, and then fuses it into the backbone network cascading multiple DRCAB modules; then it passes through the cross-scale fusion sub-network to fuse the non-local context information to obtain the features of the first stage; the features of the first stage are concatenated with the shallow features of the second stage, and the concatenated features are then passed through the spatial detail sub-network and the cross-scale fusion sub-network to obtain the features of the second stage; repeat the above steps to obtain the features of the nth stage, and then obtain the clear image of each stage after convolutional layer processing.

[0119] Among them, the spatial detail sub-network in step one is divided into two parts: the backbone network of cascaded DRCAB modules and the U-Net-based edge enhancement branch. The backbone network composed of cascaded DRCABs is used to extract potential local features and realize the recovery of local spatial information of the image; the U-Net-based edge enhancement branch is used to supplement edge texture features and enrich local spatial features.

[0120] The DRCAB module has a single receptive field. The backbone network of the spatial detail sub-network composed of cascading multiple DRCAB modules will pay more attention to the local details of the image. The U-Net-based edge enhancement branch will add edge information to the backbone network, which will assist in restoring the edge texture features of the image and improve the extraction ability of local detail information.

[0121] Among them, the specific method of the U-Net-based edge enhancement branch network in step one is as follows: First, in the downsampling stage, it passes through three encoding blocks in sequence. Each encoding block includes: convolutional layer - BN layer - ELU activation function, convolutional layer - BN layer - ELU activation function, convolutional layer - BN layer - ReLU activation function. Each encoding block reduces the size of the feature map by increasing the stride and doubles the number of channels using convolutional operations. The formula is as follows:

[0122]

[0123] Among them, y k-1 represents the initial feature map obtained by the (k - 1)th downsampling, represents the feature map obtained during the kth downsampling process, y kDenote the feature map obtained by the $k$-th downsampling, $W_1$, $W_2$, $W_3$ denote the convolutional kernel parameters, $b_1$, $b_2$, $b_3$ denote the convolutional kernel bias terms, $k \in \{1, 2, 3\}$, $BN$ denotes the batch normalization layer, $ELU$ denotes the ELU activation function, $\delta$ denotes the ReLU activation function, $*$ denotes the convolution operation, $y$ k $\in \mathbb{R}$ C×H×W , $y$ k-1 $\in \mathbb{R}$ C×H×W denotes the dimension of the feature map, where $C$ denotes the number of channels of the feature map, $H$ denotes the height of the feature map, $W$ denotes the width of the feature map. Simplify the above process to: $y$ k $= DS(y$ k-1 ). Then the formula for the downsampling process is as follows:

[0124] $y_3 = DS(DS(DS(F)))$

[0125] where $DS$ denotes downsampling, the initial feature map $F = x_0$, and the feature maps $x_3$ are obtained after three consecutive downsamplings respectively;

[0126] Then, after two convolution operations, only double and halve the number of channels respectively. The formula is as follows:

[0127]

[0128] where, denotes the obtained feature map, $W_4$, $W_5$ denote the convolutional kernel parameters, $b_4$, $b_5$ denote the convolutional kernel bias terms, $BN$ denotes the batch normalization layer, $ELU$ denotes the ELU activation function, $\delta$ denotes the ReLU activation function, $*$ denotes the convolution operation;

[0129] Finally, in the upsampling stage, also go through three decoding blocks. Each decoding block includes in turn: transposed convolutional layer - BN layer - ReLU activation function, convolutional layer - BN layer - ELU activation function, convolutional layer - BN layer - ELU activation function. First, fuse the multi-scale features in the encoding stage, and then use the transposed convolution to enlarge the size of the feature map. The formula is as follows:

[0130]

[0131] where, denotes the feature map obtained after the $k$-th upsampling, denotes the $k$-th upsampling transposed convolutional operation, $W_6$, $W_7$ denote the convolutional kernel parameters, $b_6$, $b_7$ denote the convolutional kernel bias terms, $[]$ denotes the concatenation operation, $y$ 4-k denotes the feature map of $x_0$ after the $(4 - k)$-th downsampling, denotes the feature map after the $(4 - k)$-th upsampling, is the feature map in the (4 - k)-th downsampling process, where k ∈ {1, 2, 3}, BN represents the batch normalization layer, ELU represents the ELU activation function, δ represents the ReLU activation function. The above process is simplified as: Then the formula for the upsampling process is as follows:

[0132]

[0133] where ES represents upsampling, and is obtained by consecutive three upsamplings and then passes through a convolutional layer to obtain the final output sp;

[0134] In summary, the above process is represented by the following formula:

[0135] sp = EEBNet(F)

[0136] where EEBNet() represents an edge enhancement branch network with a U-Net structure, sp represents the extracted edge features, and F represents the extracted shallow features;

[0137] The edge information is normalized before being fused into the backbone network to generate a mask mask, and its formula is:

[0138] mask = σ(w1 * sp) - w2 * σ(sp),

[0139] where σ represents the Sigmoid function, the function w1 represents the amplification factor, w2 represents the bias factor, * represents the convolution operation, and sp represents the edge information learned from EEBNet.

[0140] The edge enhancement branch network based on U-Net is a deep learning model that combines the U-Net architecture and edge enhancement technology, mainly used for tasks such as cultivated land extraction from high-resolution remote sensing images. Its network structure includes the U-Net basic architecture, the edge enhancement branch, and the decoder and feature fusion.

[0141] Among them, the cascaded DRCAB module in step one contains m DRCAB modules, and each DRCAB module sequentially includes three residual blocks, global average pooling, one residual block, sigmoid activation function, weighted operation, and long connection operation.

[0142] The cascaded DRCAB module is mainly composed of a convolutional layer, a channel attention mechanism, and a residual connection, and can extract deeper image features, which helps to improve the quality of image reconstruction or processing.

[0143] Among them, the specific operation of each DRCAB module is as follows: First, use three consecutive residual blocks composed of a convolutional layer - ReLU activation function - convolutional layer and residual connection operations to extract the initial feature h of the input feature f. The formula is as follows:

[0144] h1 = W2 * (δ(W1 * f + b1)) + b2

[0145] h2 = W4 * (δ(W3 * (f + h1) + b3)) + b4

[0146] h3 = W 6* (δ(W5 * (f + h1 + h2) + b5)) + b6

[0147] h = h3 + h2 + h1

[0148] Among them, W1, W2, W3, W4, W5, W6 represent convolutional kernel parameters, b1, b2, b3, b4, b5, b6 represent convolutional kernel bias terms, δ represents the ReLU activation function, * represents the convolutional operation, h1, h2, h3 represent the feature maps after passing through each residual block respectively. Taking the first stage as an example, f = F1;

[0149] Then, perform global average pooling operation on the initial feature h for each channel. The formula is as follows:

[0150]

[0151] Among them, GAP represents the global average pooling operation, h c represents the feature of the c-th channel, h c ∈h, h c (i, j) represents the feature value at the coordinate (i, j) of the c-th channel, z c represents the feature statistic of the c-th channel, where c ∈ {0, 1…, C};

[0152] Then, concatenate the feature statistics of all channels, and apply a residual block and a Sigmoid activation function to extract the channel attention weight s, as shown in the following formula:

[0153] z = [z1, z2, …, z c

[0154] s = σ(W8 * (δ(W7 * z + b7)) + b8)

[0155] Among them, [] represents the concatenation operation, σ represents the sigmoid activation function, W7, W8 represent convolutional kernel parameters respectively, b7, b8 represent convolutional kernel bias terms, δ represents the ReLU activation function, * represents the convolutional operation;

[0156] ​Then, the feature z is weighted using the channel attention weight s to obtain which is the output of the DRCAB module, as shown in the following equation:

[0157]

[0158] where · represents the element-wise multiplication operation;

[0159] Finally, a long connection from the initial feature h to is added, and the formula is as follows:

[0160]

[0161] In summary, the processing procedure of each DRCAB module can be simplified as: f DRCAB = DRCAB(f).

[0162] Each DRCAB module extracts image features through a convolutional layer and weights the features through a channel attention mechanism. By weighting the features through the channel attention mechanism, it can enhance important features, suppress unimportant features, and improve the performance of the network. The residual connection helps train deeper networks and avoid problems such as gradient vanishing or explosion.

[0163] Among them, taking the first stage as an example for the specific operation of each DRCAB module, the input of the first DRCAB module is the shallow feature F1. After the above operations, the output of the first DRCAB module is obtained. The output of the first DRCAB module is used as the input of the second DRCAB module. The above operations are looped, and finally the output of the cascaded DRCAB module is obtained, as shown in the following equation:

[0164] F d = DRCAB m (DRCAB m-1 (… DRCAB1(F)))

[0165] where F d represents the local spatial features of the image extracted by the cascaded DRCAB blocks, DRCAB1() represents the first DRCAB block, and DRCAB m represents the m-th DRCAB block, and F represents the initially extracted feature map;

[0166] Finally, the local spatial detail information obtained by the cascaded DRCAB blocks is added to the mask mask1 obtained from the edge enhancement branch to obtain the final local spatial detail information. Taking the first stage as an example, the formula is as follows:

[0167] F 11 = W * (F d + mask1)

[0168] Among them, F 11 represents the spatial detail features extracted by the spatial detail sub-network, W represents the convolutional kernel, mask1 represents the mask output by the edge enhancement branch, and F d represents the local spatial features of the image extracted by the cascaded DRCAB blocks.

[0169] The mask mask is a matrix with the same size as the original image, used to identify specific regions in the image. In the edge enhancement branch, the mask is used to highlight the edge information in the image. The mask can be binary (only containing 0 and 1) or grayscale (containing multiple gray levels). In a binary mask, 1 usually represents the edge region and 0 represents the non-edge region.

[0170] Among them, the specific steps of the cross-scale fusion sub-network in Step 3 and Step 4 are as follows:

[0171] First, the input feature X passes through four dilated convolutions with dilation rates of 1, 2, 3, and 4 in parallel to generate multi-scale feature maps. The formula is as follows:

[0172] F k = W k*d X + b k

[0173] Among them, *d represents the dilated convolution operation with dilation rate d, W k is the weight of the convolutional kernel, b k is the bias term of the convolutional kernel, d ∈ {1, 2, 3, 4}, and the input feature X = F 11 + F1;

[0174] Fuse the four feature maps obtained from the four dilated convolutions pairwise. For the feature maps F i , F j (i ∈ {1, 2}, j ∈ {3, 4}, as shown in the following formula:

[0175] G i,j = W i,j *Concat(F i , F j ) + b i,j

[0176] Among them, G i,j represents the fused feature map, Concat() represents the concatenation in the channel dimension, W i,j is the convolutional kernel, b i,j is the bias term, and * represents the convolution operation;

[0177] Fuse the fused feature maps again to obtain the feature map H, as shown in the following formula:

[0178] H = W h *Concat({G i,j ) + b h

[0179] where W h is the convolution kernel parameter, b h is the convolution kernel bias term, and Concat() represents concatenation in the channel dimension;

[0180] Then, through channel attention - weighted operation and spatial attention - weighted operation respectively, channel attention weight s1 and spatial attention weight s2 are extracted, as shown in the following formula:

[0181] s1 = σ(W2 * fc(W1 * GAP(h) + b1) + b2)

[0182]

[0183] where represents the feature map after channel attention and spatial attention, W1, W2, W3 represent convolution kernel parameters, b1, b2, b3 represent convolution kernel bias terms, * represents convolution operation, GAP() is global average pooling, σ is the sigmoid activation function, fc represents the fully - connected layer, and · represents element - wise multiplication operation;

[0184] Finally, taking the obtained feature map as the input, it passes through the ReLU activation function - convolution layer - DRCAB module - convolution layer in sequence to obtain the final output result, and its formula is as follows:

[0185]

[0186] where l represents the final output feature map, * represents convolution operation, DRCAB() represents the DRCAB module, W4, W5 represent convolution kernel parameters, b4, b5 represent convolution kernel bias terms, and δ represents the ReLU activation function;

[0187] In summary, the above process is represented by the following formula:

[0188] F s = CSFNet(X)

[0189] where CSFNet() represents the cross - scale fusion subnet, X represents the input feature map, and F s represents the output feature map of the cross - scale fusion subnet.

[0190] The cross - scale fusion subnet is a network structure used to process multi - scale features in deep learning, aiming to effectively fuse features from different scales to improve the performance and robustness of the model.

[0191] Among them, the n denoising stages in step five are arranged as the first-stage denoising, the second-stage denoising, …, the nth-stage denoising. At the end of the first-stage denoising, the second-stage denoising, …, the nth-stage denoising, a convolutional layer is used to process the features extracted in each stage, and together with the long connection from the original image y, clear images x1, clear images x2, …, clear images x corresponding to each denoising stage are obtained. n , and the clarity of clear images x1, clear images x2, …, clear images x n gradually increases.

[0192] Through the method proposed by the present invention, the clarity of the clear image is gradually increased, thereby realizing multi-stage image denoising.

[0193] Among them, the calculation steps of the edge weighted loss function in step one are as follows:

[0194] First, an edge enhancement branch is trained with an improved Laplacian operator to obtain an edge gradient, and its formula is as follows:

[0195]

[0196] Among them, F Laplace is the defined Laplacian operator;

[0197] Then, L edge is used to represent the edge loss in each stage, and its formula is as follows:

[0198] L edge =||sp i -F Laplace (x gt )|| 2 , i ∈ (1, 2, …, n)

[0199] Among them, sp i represents the edge information obtained by the edge enhancement branch;

[0200] Finally, we use the L2 loss function to measure the difference between the denoised image x i (i ∈ {1, 2, …, n}) and the ground truth x gt , and the formula is as follows:

[0201] L out =||x i -x gt || 2 , i ∈ (1, 2, 3)

[0202] Among them, L out represents the L2 loss function in each stage;

[0203] In summary, the proposed edge-weighted loss function has the following formula:

[0204]

[0205] where the parameter θ controls the loss term of the edge loss, and L edge represents the edge loss at each stage, and L out represents the L2 loss at each stage.

[0206] By assigning higher weights to the edge regions, the edge-weighted loss function enables the model to pay more attention to edge information during training, thereby improving the accuracy of edge detection, the precision of image segmentation, and the robustness of object recognition. It plays an important role in computer vision tasks that require fine edge processing and is an effective technical means to enhance the model performance.

[0207] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device.

[0208] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A multi-stage image denoising method based on a cascaded dual subnet, characterized in that The specific steps are as follows: Step 1: Construct a multi-stage denoising network for a cascaded dual subnet. The denoising network includes: n denoising stages, where n ∈ [2, 5]. Each denoising stage includes: a spatial detail subnet and a cross-scale fusion subnet. The spatial detail subnet includes: an edge enhancement branch network based on U-Net and a backbone network with cascaded DRCAB modules. The edge enhancement branch also returns edge information for calculating the edge weighted loss function; Step 2: Extract shallow features before the start of each stage. In each stage, a convolutional layer is used to extract the shallow features F of the noisy image y, and the formula is as follows: F = W * y where W is the convolutional kernel that expands the number of feature maps, and * represents the convolutional operation; Step 3: The shallow features of the first stage first pass through the spatial detail subnet with an edge enhancement branch to obtain the local spatial features of the image, and then pass through a cross-scale fusion subnet to fully fuse the non-local context features; Step 4: The features obtained in the first denoising stage are added to the shallow features obtained in the second stage, and then first pass through the spatial detail subnet with an edge enhancement branch to obtain the local spatial features of the image, and then pass through a cross-scale fusion subnet to fully fuse the non-local context features; Step 5: Repeat the above steps until the features of the nth denoising stage are obtained. The features output by the nth stage are processed by a convolutional layer to obtain the clear image x n .

2. The multi-stage image denoising method based on a cascaded twin network according to claim 1, wherein: The spatial detail subnet described in Step 1 is divided into two parts: the backbone network with cascaded DRCAB modules and the edge enhancement branch based on U-Net. The backbone network composed of cascaded DRCABs is used to extract potential local features and restore the local spatial information of the image; the edge enhancement branch based on U-Net is used to supplement edge texture features and enrich the local spatial features.

3. A multi-stage image denoising method based on a cascaded dual subnet according to claim 1, characterized in that: The specific method of the edge enhancement branch network based on U-Net described in Step 1 is as follows: First, in the downsampling stage, it passes through three encoding blocks in sequence. Each encoding block includes, in sequence: a convolutional layer - BN layer - ELU activation function, a convolutional layer - BN layer - ELU activation function, a convolutional layer - BN layer - ReLU activation function. Each encoding block reduces the size of the feature map by increasing the stride and doubles the number of channels using convolutional operations. The formula is as follows: Among them, y k-1 represents the initial feature map obtained by the (k - 1)-th downsampling, represents the feature map obtained during the k-th downsampling, y k represents the feature map obtained by the k-th downsampling, W1, W2, and W3 represent the convolutional kernel parameters, b1, b2, and b3 represent the convolutional kernel bias terms, k ∈ {1, 2, 3}, BN represents the batch normalization layer, ELU represents the ELU activation function, δ represents the ReLU activation function, * represents the convolution operation, y k ∈ R C×H×W ,y k-1 ∈ R C×H×W represents the dimension of the feature map, where C represents the number of channels of the feature map, H represents the height of the feature map, and W represents the width of the feature map. Simplify the above process to: y k = DS(y k-1 ), then the formula for the downsampling process is as follows: y3 = DS(DS(DS(F))) where DS represents downsampling, the initial feature map F = x0, and the feature maps x3 are obtained after three consecutive downsamplings; Then, two convolutional operations are performed to only double the number of channels and reduce it by half respectively. The formula is as follows: Among them, represents the obtained feature map, W4 and W5 represent the convolutional kernel parameters, b4 and b5 represent the convolutional kernel bias terms, BN represents the batch normalization layer, ELU represents the ELU activation function, δ represents the ReLU activation function, and * represents the convolution operation; Finally, in the upsampling stage, it also passes through three decoding blocks in sequence. Each decoding block includes, in sequence: a transposed convolutional layer - BN layer - ReLU activation function, a convolutional layer - BN layer - ELU activation function, a convolutional layer - BN layer - ELU activation function. First, it fuses the multi-scale features in the encoding stage, and then uses transposed convolution to enlarge the size of the feature map. The formula is as follows: Among them, represents the feature map obtained after the k-th upsampling, represents the k-th transposed convolutional operation of upsampling, W6 and W7 represent the convolutional kernel parameters, b6 and b7 represent the convolutional kernel bias terms, [] represents the concatenation operation, and y 4-k represents the feature map of x0 after the (4 - k)-th downsampling, represents the feature map after the (4 - k)-th upsampling, is the feature map during the (4 - k)-th downsampling process, where k ∈ {1, 2, 3}, BN represents the batch normalization layer, ELU represents the ELU activation function, δ represents the ReLU activation function. Simplify the above process to: Then the formula for the upsampling process is as follows: Among them, ES represents upsampling, which is obtained by continuously performing upsampling three times and then passing through a convolutional layer to obtain the final output sp; In summary, the above process is represented by the following formula: sp = EEBNet(F) where EEBNet() represents an edge enhancement branch network with a U-Net structure, sp represents the extracted edge features, and F represents the extracted shallow features; The edge information is normalized before being fused into the backbone network to generate a mask mask, and its formula is: mask = σ(w1 * sp) - w2 * σ(sp), where σ represents the Sigmoid function, the function w1 represents the amplification factor, w2 represents the bias factor, * represents the convolution operation, and sp represents the edge information learned from EEBNet.

4. A multi-stage image denoising method based on a cascaded dual subnet according to claim 1, characterized in that: The cascaded DRCAB module described in step one contains m DRCAB modules. Each DRCAB module sequentially includes three residual blocks, global average pooling, one residual block, sigmoid activation function, weighted operation, and long connection operation.

5. A multi-stage image denoising method based on a cascaded dual subnet according to claim 4, characterized in that: The specific operation of each DRCAB module is as follows: First, use three consecutive residual blocks composed of a convolutional layer - ReLU activation function - convolutional layer and residual connection operation to extract the initial feature h of the input feature f. The formula is as follows: h1 = W2 * (δ(W1 * f + b1)) + b2 h2 = W4 * (δ(W3 * (f + h1) + b3)) + b4 h3 = W 6* (δ(W5 * (f + h1 + h2) + b5)) + b6 h = h3 + h2 + h1 where W1, W2, W3, W4, W5, W6 represent convolutional kernel parameters, b1, b2, b3, b4, b5, b6 represent convolutional kernel bias terms, δ represents the ReLU activation function, * represents the convolution operation, and h1, h2, h3 represent the feature maps after passing through each residual block. Taking the first stage as an example, f = F1; Then, perform global average pooling operation on the initial feature h for each channel. The formula is as follows: Among them, GAP represents the global average pooling operation, and h c represents the feature of the c-th channel, and h c ∈ h, h c (i, j) represents the eigenvalue of the feature with coordinates (i, j) in the c-th channel, and z c represents the feature statistic of the c-th channel, where c ∈ {0, 1 …, C}; Then, concatenate the feature statistical magnitudes of all channels, and apply a residual block and Sigmoid activation function to extract the channel attention weight s, as shown in the following formula: z = [z1, z2, …, z c ​ s = σ(W8 * (δ(W7 * z + b7)) + b8) where [] represents the concatenation operation, σ represents the sigmoid activation function, W7, W8 respectively represent convolutional kernel parameters, b7, b8 represent convolutional kernel bias terms, δ represents the ReLU activation function, * represents the convolution operation; Then, the feature z is weighted using the channel attention weight s to obtain which is the output of the DRCAB module, as shown in the following equation: where · represents the element-wise multiplication operation; Finally, add a long connection from the initial feature h to as follows: In summary, the processing procedure of each DRCAB module can be simplified to: f DRCAB = DRCAB(f).

6. The multi-stage image denoising method based on a cascaded dual subnet according to claim 5, wherein: Taking the first stage as an example for the specific operation of each DRCAB module, the input of the first DRCAB module is the shallow feature F1. After the above operations, the output of the first DRCAB module is obtained. The output of the first DRCAB module is used as the input of the second DRCAB module. Repeat the above operations, and finally obtain the output of the cascaded DRCAB module, as shown in the following formula: F d = DRCAB m (DRCAB m-1 (…DRCAB1(F))) Among them, F d represents the local spatial features of the image extracted by the cascaded DRCAB blocks. DRCAB1() represents the first DRCAB block, and DRCAB m represents the m-th DRCAB block, and F represents the feature map initially extracted by contact; Finally, add the local spatial detail information obtained from the cascaded DRCAB block to the mask mask1 obtained from the edge enhancement branch to obtain the final local spatial detail information. Taking the first stage as an example, its formula is as follows: F 11 = W * (F d + mask1) Among them, F 11 represents the spatial detail features extracted by the spatial detail sub-network, W represents the convolution kernel, mask1 represents the mask output by the edge enhancement branch, and F d represents the local spatial features of the image extracted by the cascaded DRCAB blocks.

7. A multi-stage image denoising method based on a cascaded dual subnet according to claim 1, characterized in that: The specific steps of the cross-scale fusion sub-network described in steps three and four are as follows: First, the input feature X passes through dilated convolutions with dilation rates of 1, 2, 3, and 4 in parallel to generate multi-scale feature maps. The formula is expressed as follows: F k = W k*d X + b k Among them, *d represents the dilated convolution operation with a dilation rate of d, and W k is the weight of the convolution kernel, and b k is the bias term of the convolution kernel. d ∈ {1, 2, 3, 4}, and the input feature X = F 11 + F1; Fuse the four feature maps obtained by four dilated convolutions pairwise, for the feature maps F of any two branches i 、F j (i ∈ {1, 2}, j ∈ {3, 4}), as shown in the following formula: G i,j = W i,j *Concat(F i , F j ) + b i,j Among them, G i,j represents the fused feature map, Concat() represents the concatenation in the channel dimension, W i,j is the convolution kernel of multiple consecutive convolutional layers, b i,j is the bias term, and * represents the convolution operation; The fused feature maps are fused again to obtain the feature map H, as shown in the following formula: H = W h *Concat({G i,j}) + b h Among them, W h is the convolutional kernel parameter, b h is the convolutional kernel bias term, and Concat() represents concatenation in the channel dimension; Then, through the channel attention-weighted operation and the spatial attention-weighted operation respectively, the channel attention weight s1 and the spatial attention weight s2 are extracted, as shown in the following formula: s1 = σ(W2 * fc(W1 * GAP(h) + b1) + b2) Among them, represents the feature map after channel attention and spatial attention, W1, W2, and W3 represent the convolutional kernel parameters, b1, b2, and b3 represent the convolutional kernel bias terms, * represents the convolution operation, GAP() is the global average pooling, σ is the sigmoid activation function, fc is the fully connected layer, and · represents the element-wise multiplication operation; Finally, the obtained feature map is used as the input, and passes through the ReLU activation function - convolutional layer - DRCAB module - convolutional layer in sequence to obtain the final output result. The formula is as follows: where l represents the final output feature map, * represents the convolution operation, DRCAB() represents the DRCAB module, W4 and W5 represent the convolution kernel parameters, b4 and b5 represent the convolution kernel bias terms, and δ represents the ReLU activation function; In summary, the above process is represented by the following formula: F s = CSFNet(X) Among them, CSFNet() represents the cross-scale fusion subnet, X represents the input feature map, and F s represents the output feature map of the cross-scale fusion subnet.

8. A multi-stage image denoising method based on a cascaded dual subnet according to claim 1, characterized in that: The n denoising stages described in step five are arranged as the first-stage denoising, the second-stage denoising, ……, the nth-stage denoising. At the end of each of the first-stage denoising, the second-stage denoising, ……, the nth-stage denoising, a convolutional layer is used to process the features extracted in each stage, and then the long connection from the original image y is added to obtain the clear images x1, x2, ……, x corresponding to each denoising stage n , and the clarity of the clear images x1, x2, ……, x n gradually improves.

9. A multi-stage image denoising method based on a cascaded dual subnet according to claim 1, characterized in that: The calculation steps of the edge-weighted loss function described in Step 1 are as follows: First, the edge gradient is obtained by training the edge enhancement branch with the improved Laplacian operator, and its formula is as follows: where, F Laplace is the defined Laplace operator; Then, use L edge to represent the edge loss for each stage, and its formula is as follows: L edge = ||sp i - F Laplace (x gt )|| 2 , i ∈ (1, 2,..., n) Among them, sp i represents the edge information obtained by the edge enhancement branch; Finally, we use the L2 loss function to measure the denoised image x at each stage i (i ∈ {1, 2, ..., n}) and the ground truth x gt The difference between them is given by the following formula: L out = ||x i - x gt || 2 , i ∈ (1, 2, 3) Among them, L out represents the L2 loss function for each stage; In summary, the proposed edge-weighted loss function has the following formula: Among them, the parameter θ controls the loss term of the edge loss, and L edge represents the edge loss at each stage, and L out represents the L2 loss at each stage.