Remote sensing image cloud processing and semantic segmentation method based on adaptive information fusion
By adopting an adaptive information fusion method in remote sensing image processing, combining the encoder and decoder of the convolutional layer and channel self-attention mechanism, multi-task joint training for cloud removal and semantic segmentation is realized, solving the limitations of independent task optimization in the existing technology, and significantly improving image quality and segmentation accuracy.
Patent Information
- Application Number
- CN202510217521.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art has limitations in cloud removal and surface semantic segmentation of remote sensing images, resulting in limited technical performance, lack of direct complementary information transmission between tasks, and serious error accumulation problems.
Adaptive information fusion method is adopted, and by building an encoder and decoder based on convolutional layer and channel self-attention mechanism, combined with an expert neural network with UNet structure, multi-task joint training for cloud removal and semantic segmentation is realized, and feature fusion and weight allocation are used for gated units.
It significantly improves the quality of cloud removal and the accuracy of semantic segmentation in remote sensing images, reduces redundant calculations, and overcomes the efficiency problems caused by multi-model separation operation.
Smart Images

Figure CN120147633A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of remote sensing image processing, specifically to image cloud and fog removal and surface semantic segmentation using a deep learning neural network model, and particularly to a method for remote sensing image cloud and fog processing and semantic segmentation with adaptive information fusion. Background Art
[0002] Currently, remote sensing images have important applications in the fields of geographic information science, urban planning, agricultural monitoring, environmental protection, etc. However, due to the existence of cloud and fog occlusion, remote sensing images face great challenges in practical applications.
[0003] (1) Cloud and fog removal technology: Traditional methods include image enhancement and spectral information compensation techniques. Image enhancement methods such as histogram equalization and Retinex algorithm improve the quality by adjusting the contrast and brightness of the image, but have limited effects on areas blocked by thick clouds and fog. Spectral information compensation uses the different penetration abilities of multiple spectral bands of multi-spectral images for compensation, but requires high-quality multi-spectral data support and has insufficient restoration ability for completely cloud and fog-blocked areas. In recent years, deep learning-based methods, such as generative adversarial networks (GAN), have made certain progress in complex scenarios, but there is still room for improvement in terms of detail restoration and computational efficiency.
[0004] (2) Surface semantic segmentation technology: Traditional semantic segmentation methods rely on manual feature design and classification algorithms, such as support vector machines and random forests, but it is difficult to capture complex surface information in high-dimensional data. Deep learning methods, such as fully convolutional networks (FCN) and U-Net, have significantly improved the segmentation performance through end-to-end training. However, when the quality of the input image is poor, the segmentation effect is often affected.
[0005] Existing multi-task joint strategies mainly focus on independent module optimization or the combination of image enhancement and segmentation. The main defects of these methods are as follows:
[0006] (1) Limited technical performance: Traditional strategies perform worse than deep learning models under complex surface conditions and rely on the quality of preprocessing.
[0007] (2) Lack of direct complementary information transfer between tasks: There is no cooperation mechanism between the cloud and fog removal and semantic segmentation tasks. The segmentation only depends on the result after cloud removal and cannot dynamically utilize intermediate features; the cloud removal task also cannot use the surface semantics generated by the segmentation to provide assistance for restoration.
[0008] (3) Error accumulation problem: The defects in cloud and fog removal are directly transmitted to the segmentation task, and the lack of features further deteriorates the segmentation result, forming an error accumulation chain. Summary of the Invention
[0009] To solve the problem that existing technologies have their own advantages in the independent tasks of cloud and fog removal and surface semantic segmentation, but due to the failure to jointly optimize, the improvement of their respective performances is limited, the present invention provides a method for remote sensing image cloud and fog processing and semantic segmentation with adaptive information fusion, which can effectively improve the quality of cloud and fog removal of remote sensing images and the accuracy of semantic segmentation, and significantly reduce redundant calculations. The method mainly includes:
[0010] S1: Obtain the original image obscured by clouds and fog, and preprocess the original image;
[0011] Build an encoder based on convolutional layers and channel self-attention mechanism, input the preprocessed image into the encoder, and generate image features input to the expert neural network;
[0012] S2: Build three expert neural networks based on UNet, namely a cloud and fog removal expert neural network, a shared expert neural network, and a semantic segmentation expert neural network. After splicing the features of the corresponding layers through skip connections, each expert neural network generates a feature map input to the gating unit;
[0013] S3: Build a gating unit based on the channel self-attention mechanism, design an independent gating unit for each task. The single-task expert neural network and the shared expert neural network are spliced and then obtain a one-dimensional global feature through global average pooling, and then use a 1D convolutional layer to fuse the information between channels, multiply with the original image data to obtain the output of the gating unit, and input the output result into the decoder; the single-task expert neural network is the cloud and fog removal expert neural network or the semantic segmentation expert neural network;
[0014] S4: Build a decoder based on convolutional layers and channel self-attention mechanism. The decoder includes a cloud and fog removal decoder and a semantic segmentation decoder. The cloud and fog removal decoder is used to obtain the image after cloud and fog removal, and the semantic segmentation decoder is used to obtain the ground semantic segmentation result image;
[0015] S5: According to the image after cloud and fog removal and the ground semantic segmentation result image, calculate the loss L cloud of the cloud and fog removal task and the loss L seg of the semantic segmentation task, add the two weighted, obtain the overall loss of the model, and reversely optimize the multi-task learning model of deep learning. The model includes an encoder, a cloud and fog removal expert neural network, a shared expert neural network, a semantic segmentation expert neural network, a gating unit, a cloud and fog removal decoder, and a semantic segmentation decoder, and finally obtain the optimized multi-task learning model of deep learning for cloud and fog removal and semantic segmentation processing of remote sensing images obscured by rain and fog.
[0016] A computer-readable storage medium stores a computer program, which when executed by a processor, implements the steps of the above method.
[0017] The beneficial effects brought by the technical solution provided by the present invention are as follows:
[0018] 1. Enhanced cloud and fog processing: The prior information of surface categories introduced by the segmentation module strengthens the model's perception of the semantic structure in occluded areas, significantly optimizing the performance of the cloud removal model in texture restoration of the target area. The multi-task joint training mode constructs multi-level feature mappings in complex scenarios, and its cloud and fog removal ability is particularly superior to the single-task benchmark model under heterogeneous terrain conditions.
[0019] 2. Improved segmentation performance: The features of the cloud and fog removal module enhance the segmentation module's ability to capture high-frequency details and boundary features, reducing the classification error caused by input image noise; through feature reconstruction, the segmentation network achieves refined differentiation and classification of surface objects, avoiding the performance bottleneck of traditional segmentation models relying on low-quality images.
[0020] 3. Unified framework reduces computational complexity: The shared network architecture guides the efficient coupling of multiple tasks in the parameter space, significantly reducing redundant calculations and compressing the occupancy of hardware resources. The single-frame multi-task collaborative inference mode reduces the overall computational complexity by an order of magnitude, overcoming the efficiency problems brought by the separate operation of multiple models. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The present invention will be further described below in conjunction with the drawings and embodiments. In the drawings:
[0022] Figure 1 is a flowchart of a method for remote sensing image cloud and fog processing and semantic segmentation with adaptive information fusion in an embodiment of the present invention;
[0023] Figure 2 is a schematic diagram of the overall model structure in an embodiment of the present invention;
[0024] Figure 3 is a schematic diagram of the shared encoder structure in an embodiment of the present invention;
[0025] Figure 4 is a schematic diagram of the expert neural network structure in an embodiment of the present invention;
[0026] Figure 5 is a schematic diagram of the gating unit structure in an embodiment of the present invention;
[0027] Figure 6 is a schematic diagram of the decoder structure in an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0028] In order to have a clearer understanding of the technical features, objectives, and effects of the present invention, the specific embodiments of the present invention will now be described in detail with reference to the drawings.
[0029] Example 1
[0030] Please refer to Figure 1 , Figure 1 which is a flowchart of a method for processing remote sensing image cloud and fog and semantic segmentation with adaptive information fusion in an embodiment of the present invention. This method uses the preprocessed remote sensing image data to provide to the model for dual-task training, saves the model after the training is completed, and then the model can generate a clear image without cloud and fog occlusion and the result of surface semantic segmentation according to the original remote sensing image occluded by cloud and fog. The specific steps are as follows:
[0031] S1: Obtain the original remote sensing image occluded by cloud and fog, and perform preprocessing on the original remote sensing image; the preprocessing includes a normalization operation;
[0032] Build an encoder based on convolutional layers and a channel self-attention mechanism. The encoder includes convolutional layer 1, convolutional layer 2, SE_Bolck (channel self-attention layer), convolutional layer 3, convolutional layer 4, and a splicing output layer arranged in this order. After each convolutional layer, it passes through the Relu activation function;
[0033] Input the preprocessed image into the encoder to generate image features input to the expert neural network;
[0034] S2: Build three expert neural networks based on UNet, namely a cloud and fog removal expert neural network, a shared expert neural network, and a semantic segmentation expert neural network. Gradually reduce the size of the image through 5 convolutional operations. Each layer includes operations such as 3*3 convolutional operations, ReLU activation, and 2*2 max pooling, and uses deconvolution to gradually restore it to a segmentation result of the same size as the input image. Concatenate the features of the corresponding layers through skip connections. Finally, each expert neural network generates a 16-channel feature map;
[0035] S3: Build a gating unit based on the channel self-attention mechanism. Design an independent gating unit for each task. After the single-task expert neural network and the shared expert neural network are concatenated, obtain a one-dimensional global feature through global average pooling (GAP), and then use a 1D convolutional layer (here, the convolutional kernel size is selected as 5) to fuse the information between channels, so as to ensure that the most appropriate cross-channel information can be captured. Finally, obtain the output of the gating unit after multiplying with the original data;
[0036] S4: Build a decoder based on the convolutional layer and channel self-attention mechanism. The decoder can be regarded as the inverse process of the encoder. Different from the encoder, different-sized convolutional kernels are used in the decoder. The decoder includes a 3×3 convolutional layer, a 5×5 convolutional layer, and a 7×7 convolutional layer arranged in parallel. After passing the results through these 3 convolutional layers and then through the Relu activation function, they are then calculated by the SE_Bolck (channel self-attention layer). After splicing the results of the three, finally, through a 1×1 convolutional operation, the image is reduced to 3 channels to obtain the final result;
[0037] S5: Calculate the loss L of the dehazing task cloud and the loss L of the semantic segmentation task seg , and add them together with weights to obtain the overall loss of the model, and perform reverse optimization on the multi-task learning model of deep learning as Figure 2 shown. The model includes an encoder, a dehazing expert neural network, a shared expert neural network, a semantic segmentation expert neural network, a gating unit, a dehazing decoder, and a semantic segmentation decoder.
[0038] The core of the multi-task learning model of deep learning is to use the encoder to extract the general features of the original image obscured by clouds and fully exploit the common information in the image. Subsequently, through three carefully designed expert neural networks, the task relationships are accurately modeled, and each expert neural network focuses on learning the features crucial for a specific task. Among them, the shared expert neural network serves as a key hub, deeply interacting with the other two tasks respectively to achieve joint optimization. Moreover, with the help of the gating unit, the model can automatically and intelligently adjust the weight distribution between the shared information and the specific information in the specific task according to the characteristics of the input image, ensuring that each task can obtain the most suitable information support. Finally, through the decoder of the specific task, the intermediate result is converted into the final result, that is, a clear image after removing clouds and an accurate semantic segmentation result are obtained. The specific implementation is as follows:
[0039] (1) Extract the general features of the cloud-obscured image through the encoder as Figure 3 shown. In this process, the channel self-attention mechanism will be used to highlight the key features. An encoder based on the convolutional layer and channel self-attention mechanism will be built. The structure of this encoder sequentially includes: convolutional layer 1, convolutional layer 2, SE_Block (channel self-attention layer), convolutional layer 3, and splicing output layer. After each convolutional operation, the Relu activation function will be used for activation processing, and finally, the input features required by the expert network will be generated. The following is the training process of this encoder:
[0040] Step a: Pass the remotely sensed image x obscured by clouds through a convolutional layer 1 with a convolutional kernel size of 3×3, a stride of 1, and 16 channels, and pass it through the Relu activation function:
[0041] O 1 = Relu(F 1.1 (x)) (1)
[0042] Step b: Pass O 1 through a convolutional layer 2 with a kernel size of 3×3, a stride of 1, and 32 channels, and pass it through the Relu activation function:
[0043] O 2 = Relu(F 1.2 (O 1 )) (2)
[0044] Step c: Pass O 2 through a channel self-attention layer, perform global average pooling on the feature map to generate a 1*1*32 vector z, and process the obtained vector z through two fully connected layers to obtain the desired channel weight value s. Different values represent the weight information of different channels, and different weights are assigned to the channels. Assign weights to O 2 generated in step b to obtain the weight map O 3 :
[0045]
[0046] s = σ(W 2 δ(W 1 z)) (4)
[0047] O 3 = sO 2 (5)
[0048] Step d: Pass O 3 through a convolutional layer 3 with a kernel size of 3×3, a stride of 1, and 3 channels, and pass it through the Relu activation function:
[0049] O 4 = Relu(F 1.3 (O 3 )) (6)
[0050] Step e: Concatenate O 4 with the original image x to obtain the encoder output feature x encode
[0051] x encode = O 4 + x (7)
[0052] where, F 1.n represents the convolution operation of convolutional layer n; for example, F 1.1 represents the convolution operation of convolutional layer 1; F 1.2Represents the convolution operation of the convolutional layer 2; O n Represents the output image of the nth layer; H and W represent the height and width of the image; u c Represents the global average pooling operation on the feature map; W 1 ,W 2 Represents the weight matrices in two fully connected layers; s represents the channel weight value, σ and δ represent activation functions; δ(W 1 z) represents the first fully connected layer, σ(W 2 δ(W 1 z)) represents the second fully connected layer. Formula (4) represents the self-attention operation on the globally pooled vector, and formula (5) represents the channel weight adjustment for O 2 ; The sum of formulas (3), (4), and (5) represents the entire channel self-attention layer.
[0053] (2) Use the UNet-based expert neural network as shown in Figure 4 to extract and model the cloud feature, semantic segmentation feature, and shared feature respectively, and optimize specific tasks. The specific implementation process is as follows:
[0054] Step f: Perform a convolutional layer with a convolution kernel size of 3×3, a stride of 1, and an output channel of 64 on the x encode feature of the encoder, and pass it through the Relu activation function, and then perform a max pooling operation to reduce the image size to half of the original image:
[0055] E 1 = Relu(F 2.1 (x encode )) (8)
[0056] E 1 = Max_Pooling 2.1 (E 1 )) (9)
[0057] Repeat step f continuously until E 2 , E 3 , E 4 , E 5 ;
[0058] Step g: Perform a transposed convolution operation on E 5 with a convolution kernel size of 2×2, a stride of 2, and an unchanged output channel to restore the image to the size of E 4 , and splice it with E 4 and then perform convolution to obtain P 4 , P 4 is then dimensionally reduced to obtain the same number of channels as E 4 :
[0059] P 4 = Relu(Tran_F 2.1 (E 5 )) + E 4 (10)
[0060] Repeat step g continuously until the original image size is restored to obtain the output features of the expert neural network. Each expert neural network performs the above steps simultaneously, and finally generates the output features E cloud , E share , E seg .
[0061] (3) Adaptive task interaction and information fusion mechanism, a two-way information interaction channel established between the cloud removal and semantic segmentation tasks. For example, the cloud removal task adaptively adjusts its own processing strategy based on the category information provided by semantic segmentation; the semantic segmentation task uses the images of cloud removal to improve the segmentation accuracy; and in the process of fusing shared information and task-specific information, the core logic of the adaptive weight allocation algorithm adopted by the gating unit. This unique design of task interaction and information fusion is an important protected content. Specifically:
[0062] Use Figure 5 The gating unit shown in the figure to achieve the fusion of shared features and task-specific features, and enhance the information interaction between tasks through the adaptive weight allocation algorithm. The specific implementation of the gating unit is similar to the channel self-attention layer in the encoder. The difference is that after the single-task expert neural network and the shared expert neural network are concatenated and the channel weight value s is obtained through global average pooling (GAP), a 1D convolutional layer (in this embodiment, the convolutional kernel size of the 1D convolutional layer is 5) is used to fuse the information between channels:
[0063] Step h: Perform channel self-attention information fusion on E cloud , E seg respectively with E share :
[0064] Gate cloud = Conv1D 3.1 (s)(E cloud + E share ) (11)
[0065] Gate seg = Conv1D 3.2 (s)(E cloud + E seg ) (12)
[0066] (4) The decoder structure is as Figure 6 shown, and multi-scale features are fused in the decoding stage to output clear images and high-precision segmentation results
[0067] Step i: Pass Gate cloud through a convolutional layer with a 3×3 convolutional kernel, padding of 1, and 16 channels, and pass through the Relu activation function:
[0068] D 1 = Relu(F 4.1 (Gate cloud )) (13)
[0069] Step j: Pass Gate cloud through a convolutional layer with a 5×5 convolutional kernel, padding of 2, and 16 channels, and pass through the Relu activation function:
[0070] D 2 = Relu(F 4.2 (Gate cloud )) (14)
[0071] Step k: Pass Gate cloud through a convolutional layer with a 3×3 convolutional kernel, padding of 3, and 16 channels, and pass through the Relu activation function:
[0072] D 3 = Relu(F 4.3 (Gate cloud )) (15)
[0073] Step l: Concatenate D 1 , D 2 , D 3 and pass through a convolutional layer with a 1×1 convolutional kernel, padding of 1, and 3 channels, and pass through the Relu activation function to obtain the final clear image X clear :
[0074] X clear = Relu(F 4.4 (D 1 + D 2 + D 3 )) (16)
[0075] The result X of semantic segmentation seg is obtained from Gate seg through steps i to k.
[0076] (5) Calculate the loss and optimize the model backward
[0077] Step m: Compare the results X clear , X seg obtained by the model with the true labels Y clear , Y seg, the calculation formula is as follows:
[0078]
[0079]
[0080] L = α·L cloud +β·L seg (19)
[0081] Wherein, H and W represent the height and width of the image, N represents the total number of semantic segmentation categories, α represents the weight parameter of the de - clouding task, β represents the weight parameter of the semantic segmentation task, and in this embodiment, α and β are set to 0.65 and 0.35 through multiple experimental tests.
[0082] The present invention can be applied to urban expansion monitoring: processing remote sensing images of a city at different times, removing cloud interference and performing semantic segmentation, which can clearly distinguish different regions such as urban built - up areas, undeveloped land, and green spaces; estimating the planting area of crops: in agricultural production, the remote sensing images after removing clouds and accurate semantic segmentation results can help the agricultural department accurately count the planting areas of different crops, and infrastructure planning, etc.
[0083] Embodiment 2
[0084] A computer - readable storage medium stores a computer program, which when executed by a processor, implements the steps of the above - mentioned method.
[0085] The above - mentioned are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A remote sensing image cloud processing and semantic segmentation method based on adaptive information fusion, characterized in that: include: S1: Obtain the original remote sensing image blocked by clouds and fog, and preprocess the original remote sensing image; Build an encoder based on convolutional layers and channel self-attention mechanism, input the preprocessed image into the encoder, and generate image features that are input into the expert neural network; S2: Build three expert neural networks based on UNet, namely, de-clouding expert neural network, sharing expert neural network, and semantic segmentation expert neural network. After splicing the features of the corresponding layers through skip connections, each expert neural network generates features that are input to the gating unit; S3: Build a gating unit based on the channel self-attention mechanism, design an independent gating unit for each task, and obtain a one-dimensional global feature by global average pooling after splicing the single-task expert neural network and the shared expert neural network. Then use the 1D convolution layer to fuse the information between channels, and obtain the output of the gating unit by multiplying it with the original image data. The output result is input into the decoder; the single-task expert neural network is a defogging expert neural network or a semantic segmentation expert neural network; S4: Build a decoder based on convolutional layers and channel self-attention mechanism. The decoder includes a declouding decoder and a semantic segmentation decoder. The declouding decoder is used to obtain the declouded image, and the semantic segmentation decoder is used to obtain the landmark semantic segmentation result image. S5: Calculate the loss L of the declouding task based on the declouded image and the landmark semantic segmentation result image. cloud and the loss L for the semantic segmentation task seg , the weighted addition of the two is performed to obtain the overall loss of the model, and the multi-task learning model of deep learning is reversely optimized. The model includes an encoder, a declouding expert neural network, a shared expert neural network, a semantic segmentation expert neural network, a gating unit, a declouding decoder, and a semantic segmentation decoder. Finally, the optimized multi-task learning model of deep learning is obtained, which is used for declouding and semantic segmentation of remote sensing images obscured by rain and fog.
2. The method for cloud and fog processing and semantic segmentation of remote sensing images with adaptive information fusion according to claim 1, characterized in that: In S1, the encoder includes convolution layer 1, convolution layer 2, channel self-attention layer, convolution layer 3, convolution layer 4, and concatenation output layer. Each convolution layer is followed by a Relu activation function.
3. The method for cloud and fog processing and semantic segmentation of remote sensing images with adaptive information fusion according to claim 2, characterized in that: The implementation process of the encoder is: Step a: The original remote sensing image x blocked by clouds passes through a convolution layer 1 with a convolution kernel size of 3×3, a step size of 1, and a channel of 16, and the weight map O1 is obtained through the Relu activation function: O1=Relu(F 1.1 (x)) (1) Step b: Pass O1 through a convolution layer 2 with a convolution kernel size of 3×3, a step size of 1, and a channel of 32, and obtain the weight map O2 through the Relu activation function: O2=Relu(F 1.2 (O1)) (2) Step c: Pass O2 through a channel self-attention layer, perform global average pooling on the feature map, generate a 1*1*32 vector z, and process the obtained vector z through two fully connected layers to obtain the channel weight value s. Different values represent the weight information of different channels, giving different weights to the channels; assign weights to O2 generated in step b to obtain the weight map O3: s=σ(W2δ(W1z)) (4) O3=sO2 (5) Step d: Pass O3 through a convolution layer 3 with a convolution kernel size of 3×3, a step size of 1, and a channel of 3, and obtain the weight map O4 through the Relu activation function: O4=Relu(F 1.3 (O3)) (6) Step e: Concatenate O4 with the original remote sensing image x to obtain the image feature x output by the encoder encode : x encode =O4+x (7) Among them, F 1.n represents the convolution operation of convolution layer n; O n represents the output image of the nth layer, n = 1, 2, 3, 4; H, W represent the height and width of the image; u c Indicates the global average pooling operation on the feature map; W1, W2 represent the weight matrices in the two fully connected layers; s represents the channel weight value, σ and δ represent the activation function.
4. The method for cloud and fog processing and semantic segmentation of remote sensing images with adaptive information fusion according to claim 1, characterized in that: In S2, each expert neural network gradually reduces the size of the image through 5 layers of convolution operations. Each layer contains a 3*3 convolution operation, ReLU activation, and a 2*2 maximum pooling operation, and uses deconvolution to gradually restore the image after the convolution operation to a segmentation result of the same size as the input image.
5. The method for remote sensing image cloud processing and semantic segmentation based on adaptive information fusion according to claim 1, characterized in that: The implementation process of the expert neural network is: Step f: transform feature x encode The image is input to a convolutional layer with a kernel size of 3×3, a stride of 1, and an output channel of 64, and then activated by the Relu function and then a maximum pooling to reduce the image size to half of the original remote sensing image: e1=Relu(F 2.1 (x encode )) (8) E1=Max_Pooling 2.1 (e1)) (9) Among them, F 2.1 Represents the convolution operation in the expert neural network; e1 represents the result obtained after the convolution operation and the Relu activation function; Max_Pooling 2.1 Represents the maximum pooling operation in an expert neural network; Repeat step f until E2, E3, E4, E5 are obtained; Step g: Deconvolve E5 with a kernel size of 2×2, a step size of 2, and output channels unchanged. The image is restored to the size of E4 and concatenated with E4 to obtain P4. P4 is then reduced in dimension to obtain the same number of channels as E4: P4=Relu(Tran_F 2.1 (E5))+E4 (10) Repeat step g until the original remote sensing image size is restored and the output features of the expert neural network are obtained. Each expert neural network performs the above steps at the same time, and finally generates the output features E of three expert neural networks. cloud ,E share ,E seg .
6. The method for cloud and fog processing and semantic segmentation of remote sensing images with adaptive information fusion according to claim 5, characterized in that: The implementation process of the gate control unit is: Step h: Add E cloud ,E seg Respectively with E share Perform channel self-attention for information fusion: Gate cloud =Conv1D 3.1 (s)(E cloud +E share ) (11) Gate seg =Conv1D 3.2 (S)(E cloud +E seg ) (12) Among them, Conv1D 3.1 represents the convolution operation of the 1D convolution layer in the gated unit, and s represents the channel weight value.
7. The method for cloud and fog processing and semantic segmentation of remote sensing images by adaptive information fusion according to claim 6, characterized in that: In S4, the implementation process of the decoder is: Step i: Gate cloud After a convolution layer with a convolution kernel size of 3×3, padding of 1, and channels of 16, and through the Relu activation function: D1=Relu(F 4.1 (Gate cloud )) (13) Step j: Gate cloud After a convolution layer with a convolution kernel size of 5×5, padding of 2, and channels of 16, and through the Relu activation function: D2=Relu(D 4.2 (Gate cloud )) (14) Step k: Gate cloud After a convolution layer with a convolution kernel size of 3×3, padding of 3, and channels of 16, and through the Relu activation function: D3=Relu(F 4.3 (Gate cloud )) (15) Step 1: Concatenate D1, D2, and D3 through a convolution layer with a convolution kernel size of 1×1, padding of 1, and channels of 3, and then pass the Relu activation function to obtain the final image X clear : X clear =Relu(F 4.4 (D1+D2+D3)) (16) Semantic segmentation result X seg By Gate set Obtained through steps i~k.
8. The method for cloud and fog processing and semantic segmentation of remote sensing images by adaptive information fusion according to claim 7, characterized in that: The formula for calculating the loss in S5 is: L=α·L cloud +β·L seg (19) Among them, Y clesr ,Y seg represents the true label; H and W represent the height and width of the image, N represents the total number of semantic segmentation categories, α represents the weight parameter of the declouding task, and β represents the weight parameter of the semantic segmentation task.
9. A computer-readable storage medium, characterized in that: A computer program is stored, and when the program is executed by a processor, the steps of the remote sensing image cloud processing and semantic segmentation method with adaptive information fusion as described in any one of claims 1 to 8 are implemented.
Citation Information
Cited By
Aviation riveting assembly error and omission automatic detection method based on multi-task learning
CN120913149A
Multi-modal remote sensing image classification method under any modal missing condition
CN122023948A