Feature correction and redundancy elimination-based semantic segmentation method for optical and SAR (Synthetic Aperture Radar) images

By adopting a cross-modal feature correction and de-redundant neural network model in semantic segmentation of optical images and SAR images, the problem of redundant information coverage during feature fusion is solved, and the segmentation accuracy is significantly improved.

CN119942546APending Publication Date: 2025-05-06NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510011406.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing optical image and SAR image fusion semantic segmentation networks are easily covered by a large amount of redundant information when feature fusion, resulting in the masking of specific advantageous features, affecting the accuracy of segmentation recognition.

Method used

A neural network model based on cross-modal feature correction and de-redundancy is adopted. By extracting components such as the backbone module, the cross-modal feature correction module, the interactive de-redundancy module and other components of the backbone module, the channel dimension and the spatial dimension, the disadvantages in each modal feature are corrected and redundant information is removed, so as to retain the advantageous features when the feature fusion is performed.

Benefits of technology

The accuracy of synergistic semantic segmentation of optical images and SAR images is improved by 1.5% and mIoU is improved by 2.1% compared with past algorithms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942546A_ABST
    Figure CN119942546A_ABST
Patent Text Reader

Abstract

The invention provides an optical and SAR image semantic segmentation method based on feature correction and redundancy elimination, and mainly solves the problem that a large amount of redundant information covers specific advantage features during feature fusion. On the basis of a cross-modal feature correction and redundancy elimination method, advantage features of one modal are utilized to correct disadvantage features of the other modal, defects contained in a single modal are removed, interference to subsequent feature extraction and fusion is avoided, then the features of the two modals are compared, redundant information contained in the two modals is removed, and therefore the accuracy of the method is improved. During feature fusion, a large amount of redundant information is prevented from covering specific advantage features, and the precision of collaborative semantic segmentation of the optical image and the SAR image is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of image processing, and mainly relates to an optical and SAR image semantic segmentation method based on cross-modal feature correction and redundancy removal. Background Art

[0002] The semantic segmentation task is a refined classification task. The classification task is to classify an image into a category, while the semantic segmentation task is to classify each pixel on the image into a category. Therefore, the semantic segmentation task is more difficult than the classification task and has a wider range of application scenarios, including autonomous driving, remote sensing image interpretation, medical image analysis and many other fields.

[0003] Existing multimodal remote sensing image semantic segmentation methods can be divided into data-level fusion, feature-level fusion, and decision-level fusion. Data-level fusion refers to fusion at the data level, fusing data from multiple modalities into a single data, and then using the single-modal semantic segmentation method for prediction; feature-level fusion refers to extracting features from each modality using a neural network, connecting multiple features into a single high-dimensional feature through a feature fusion method, forming a multimodal fusion feature, and then using the fusion feature to form a prediction result; decision-level fusion refers to extracting features from each modality using a neural network and performing semantic segmentation prediction separately, and then fusing multiple prediction results to form the final prediction result. At present, feature-level fusion is the most commonly used fusion method in the task of multimodal remote sensing image semantic segmentation. Compared with data-level fusion and decision-level fusion, feature-level fusion can better capture the complementary information of different modalities while maintaining the unique characteristics of each modality, which is more in line with the fact that humans express in a complementary and redundant manner.

[0004] With the development of remote sensing imaging technology, different types of sensor data can be more easily obtained. Optical images can provide high spatial resolution, rich spectral and texture information, but optical sensors are easily affected by weather; Synthetic Aperture Radar (SAR) sensors can work in any weather conditions and can penetrate certain ground objects, providing rich ground object geometric information. In complex scenes, semantic segmentation using only data obtained by a single sensor will inevitably encounter bottlenecks. Therefore, how to make full use of optical image and SAR image data for joint segmentation and study related information collaborative processing technology to improve the accuracy and reliability of segmentation is of great value.

[0005] Optical images and SAR images have their own advantages and disadvantages that are unique to their modalities. For example, there are coherent speckle noise and front contraction problems in SAR images, and there are cloud occlusion and illumination problems in optical images. These inherent problems of the modalities will have a negative impact on neural network feature extraction and fusion. At the same time, the two modalities have similarities, and more redundant information will be generated during feature extraction. If the redundant information in the features is not removed, a large number of redundant features will cover specific advantageous features during fusion, thereby affecting subsequent segmentation and recognition. Therefore, it is necessary to deal with the inherent problems and redundancy of the modality to improve the accuracy of semantic segmentation. Existing semantic segmentation networks for the fusion of optical images and SAR images often only focus on reducing the feature differences between the two modalities, but ignore the advantages and disadvantages of optical images and SAR images that are unique to their modalities. For example, there are coherent speckle noise and front contraction problems in SAR images, and there are cloud occlusion and illumination effects in optical images. These inherent problems of the modalities will have a negative impact on the feature extraction and fusion of neural networks. At the same time, the two modalities have similarities, and more redundant information will be generated during feature extraction. If the redundant information in the features is not removed, a large number of redundant features will cover specific advantageous features during fusion, thereby affecting subsequent segmentation and recognition.

[0006] In response to the above-mentioned problems, the present invention is based on a cross-modal feature correction and de-redundancy method, which uses the advantageous features of one modality to correct the inferior features of another modality, removes the defects contained in a single modality, and avoids interference with subsequent feature extraction and fusion. The features of the two modalities are then compared and the redundant information contained in the two modalities is removed. During feature fusion, a large amount of redundant information is avoided from covering specific advantageous features, thereby improving the accuracy of collaborative semantic segmentation of optical images and SAR images. Summary of the invention

[0007] In order to overcome the shortcomings of the prior art, to solve the problem that a large amount of redundant information covers specific advantageous features during feature fusion, and to improve the accuracy of collaborative semantic segmentation of optical images and SAR images, the present invention provides an optical and SAR image semantic segmentation method based on feature correction and de-redundancy.

[0008] A semantic segmentation method for optical and SAR images based on feature correction and redundancy removal, including a neural network model for feature correction and redundancy removal and a semantic segmentation method based on the neural network model for feature correction and redundancy removal;

[0009] The neural network model for feature correction and redundancy removal includes a self-feature extraction backbone module, a cross-modal feature correction module in the channel dimension, a cross-modal feature correction module in the spatial dimension, an interactive redundancy removal module, a feature decoder module and a loss function;

[0010] The self-feature extraction backbone module includes a self-feature extraction module for optical images and a self-extraction module for SAR images; the input of the self-feature extraction backbone module is the optical image and the SAR image obtained after preprocessing; the output of the self-feature extraction backbone module is the optical mode self-feature and the SAR mode self-feature;

[0011] The self-feature extraction module of the optical image includes three convolution blocks connected in series, namely and

[0012] The SAR image self-extraction module includes three convolution blocks connected in series, namely and

[0013] The cross-modal feature correction module of the channel dimension obtains the channel adaptive pooling features of the optical mode and the channel adaptive pooling features of the SAR mode by average pooling operation and maximum pooling operation respectively on the optical mode self-features and the SAR mode self-features, and splices the channel adaptive pooling features of the optical mode and the channel adaptive pooling features of the SAR mode in the channel dimension to obtain fused multi-scale pooling features; performs feature fusion and channel conversion on the fused multi-scale pooling features to obtain a fused channel correction vector; and then uses the channel correction vector to correct the self-features of the optical mode and the self-features of the SAR mode to obtain a feature map of the optical mode and a feature map of the SAR mode after the channel dimension correction;

[0014] The cross-modal feature correction module of the spatial dimension downsamples the feature map of the optical modality and the feature map of the SAR modality after the channel dimension correction in the channel dimension through average pooling operation and maximum pooling operation, and obtains the spatial multi-scale pooling features of the optical modality and the spatial multi-scale pooling features of the SAR modality respectively; splices the spatial multi-scale pooling features of the optical modality and the spatial multi-scale pooling features of the SAR modality along the channel dimension to obtain a fused spatial multi-scale pooling feature; adaptively selects and changes the channel of the fused spatial multi-scale pooling features to obtain a fused spatial correction weight map; performs channel segmentation on the fused spatial correction weight map to obtain an optical spatial correction weight map and a SAR spatial correction weight map; uses the optical spatial correction weight map and the SAR spatial correction weight map to correct the feature maps of the optical modality and the SAR modality after the channel dimension correction, and obtains a spatial dimension corrected optical modality feature map and a spatial dimension corrected SAR modality feature map;

[0015] The interactive redundancy removal module includes an optical modality feature map removal module, a SAR modality feature map removal module and a self-attention encoding network module;

[0016] The optical modality feature map removal module straightens the optical modality feature map after the spatial dimension correction to obtain optical modality interaction features and optical modality residual features respectively; the SAR modality feature map removal module straightens the feature map of the SAR modality after the spatial dimension correction to obtain SAR modality interaction features and SAR modality residual features respectively; the self-attention encoding network module modifies the linear transformation weight parameters from three channel dimensions to four channel dimensions to obtain four elements of key, self-query, mutual query and value;

[0017] Inputting the optical modality interaction feature into the self-attention encoding network module to obtain the key of the optical feature, the self-query of the optical feature, and the value of the mutual feature of the optical feature; inputting the SAR modality interaction feature into the self-attention encoding network module to obtain the key of the SAR feature, the self-query of the SAR feature, the mutual query of the SAR feature, and the value of the SAR feature;

[0018] According to the output of the self-attention encoding network module, the redundant attention matrix and self-attention matrix of the optical modality as well as the redundant attention matrix and self-attention matrix of the SAR modality are calculated using the softmax function respectively;

[0019] Subtract the optical modality self-attention matrix from the optical modality redundant attention matrix to obtain the optical modality attention matrix after removing redundancy; subtract the SAR modality self-attention matrix from the SAR modality redundant attention matrix to obtain the SAR modality attention matrix after removing redundancy;

[0020] The attention matrix of the optical modality after removing redundancy is multiplied by the value of the optical feature to obtain the optical modality feature after removing redundancy; the attention matrix of the SAR modality after removing redundancy is multiplied by the value of the SAR feature to obtain the SAR modality feature after removing redundancy;

[0021] The optical modal features after removing redundancy are added element by element to the optical modal residual features and normalized to obtain the optical modal normalized de-redundancy features;

[0022] The SAR modal features after removing redundancy are added element by element to the SAR modal residual features and normalized to obtain the SAR modal normalized de-redundancy features;

[0023] The normalized de-redundant features of the optical mode and the normalized de-redundant features of the SAR mode are processed using the sigmoid function respectively, and the processed results are subjected to channel changes respectively to obtain the features to be fused of the optical mode and the features to be fused of the SAR module; the features to be fused of the optical mode and the features to be fused of the SAR module are spliced ​​in the channel dimension to obtain the fused features;

[0024] The input of the feature decoder module is the fusion feature, and the output is the final classification result y;

[0025] The feature correction and de-redundancy neural network model is trained using the training set. When the loss function of the feature correction and de-redundancy neural network model is less than 0.3 or the training rounds reach 100 rounds, the feature correction and de-redundancy neural network model is the feature correction and de-redundancy neural network model with the best performance.

[0026] The steps of the semantic segmentation method of the redundant neural network model are as follows:

[0027] Step 1: Obtain optical and SAR image data sets and preprocess them, and divide the preprocessed optical image and SAR image data into a training set and a test set;

[0028] Step 2: Input the training set into the neural network model for feature correction and redundancy removal for training;

[0029] Step 3: Use the neural network model with the best performance for feature correction and redundancy removal to predict the test set data and obtain the segmentation result.

[0030] Furthermore, the preprocessing respectively normalizes the width, height and number of channels of the original optical image and the original SAR image, and classifies them; and divides the optical image and SAR image obtained after the preprocessing into a training set and a test set.

[0031] Furthermore, the loss function Loss of the neural network model for feature correction and redundancy removal is the cross entropy loss between the predicted value y and the category Y of the corresponding pixel point:

[0032] Where N = H × W, class is the number of categories.

[0033] Furthermore, the feature decoder module adopts two convolution blocks with shared parameters to decode and segment the fused features by interactively updating the parameters, and then upsamples the segmentation results to the original image size by interpolation to finally obtain the classification result y.

[0034] Furthermore, the interpolation method includes a bilinear interpolation method, a nearest neighbor algorithm and a bicubic interpolation algorithm.

[0035] The beneficial effects of the present invention are as follows: based on the cross-modal feature correction and interactive redundancy removal method, the present invention uses the advantageous features of one modality to correct the inferior features of another modality, removes the defects contained in a single modality, avoids interference with subsequent feature extraction and fusion, and then compares the features of the two modalities and removes the redundant information contained in the two modalities, avoids a large amount of redundant information covering specific advantageous features during feature fusion, and improves the accuracy of the collaborative semantic segmentation of optical images and SAR images. Compared with the previous collaborative semantic segmentation algorithm of optical images and SAR images, the algorithm proposed in the present invention has an accuracy increase of 1.5% and mIoU increase of 2.1% in the WHU-OPT-SAR dataset. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 It is an overall diagram of the present invention;

[0037] Figure 2 Schematic diagram of the cross-modal feature correction module in the channel dimension;

[0038] Figure 3 Schematic diagram of the cross-modal feature correction module in the spatial dimension;

[0039] Figure 4 Schematic diagram of the interactive redundancy removal module. DETAILED DESCRIPTION

[0040] The present invention is further described below in conjunction with the accompanying drawings and embodiments.

[0041] A method for semantic segmentation of optical and SAR images based on feature correction and de-redundancy includes a neural network model based on feature correction and de-redundancy, wherein the neural network model based on feature correction and de-redundancy includes a preprocessing module, a self-feature extraction backbone module, a cross-modal feature correction module in a channel dimension, a cross-modal feature correction module in a spatial dimension, an interactive de-redundancy module and a feature decoder module, as shown in the attached figure. Figure 1 To Attachment Figure 4 As shown;

[0042] The neural network model based on feature correction and redundancy removal includes the following processing steps:

[0043] Step 1: Preprocess the optical image and SAR image data sets;

[0044] Step 101: Obtain an original optical image F covering the same geographical area S OPT and the original SAR image F S SAR , R represents the pixel value of the original optical image and the original SAR image is a real number, H 1 ×W 1is the number of pixels of the original optical image, C 1 is the number of channels of the original optical image; H 1 is the height of the original optical image, W 1 The width of the original optical image, H 2 is the height of the original SAR image, W2 is the width of the original SAR image, C 2 is the number of channels of the original SAR image, H 2 ×W 2 is the number of pixels of the original SAR image;

[0045] Step 102: The original optical image and the original SAR image are respectively subjected to preprocessing operations of registration, cropping and labeling in sequence to obtain an optical image F with N pixels. OPT SAR image F SAR and label Y, Y∈R H×W×class ,Y is the category of the corresponding pixel, R H×W×class , where N = H × W, class is the number of segmentation categories, H is the height of the image, and W is the width of the image;

[0046] Step 103: Divide the optical image and SAR image data obtained in step 102 into a training set and a test set;

[0047] Step 2: construct a self-feature extraction backbone module to extract the self-features of optical image and SAR image data;

[0048] The structure of the self-feature extraction backbone module is a CNN structure;

[0049] Step 201: The self-feature extraction module of the optical image includes three convolution blocks: The self-extraction module of SAR images consists of three convolutional blocks: The dimensions of the output features of the self-feature extraction module of optical images and the self-extraction module of SAR images are both H′×W′×C′, where H′ is the height of the output feature, W′ is the width of the output feature, and C′ is the number of channels of the output feature;

[0050] Step 202: Input the optical image and SAR image data into the self-feature extraction module constructed in step 201 to obtain the optical mode self-feature and SAR modal self-characteristics in

[0051] Step 3: Construct a cross-modal feature correction module in the channel dimension to correct the optical mode self-features and SAR mode self-features in the channel dimension;

[0052] Step 301: Optical self-features are pooled by average pooling and maximum pooling. and SAR self-signature Downsampling in the spatial dimension;

[0053] Using four different pooling factors, the results of the average pooling operation and the maximum pooling operation are straightened to And splice along the spatial dimension to obtain channel multi-scale pooling features and in N i is the size of the pooling factor, i = 1, 2, 3, 4; the size of the pooling factor is N 1 ×N 1 、N 2 ×N 2 、N 3 ×N 3 、N 4 ×N 4 , L is an intermediate variable; the subscripts ap and mp represent the result of the average pooling operation and the result of the maximum pooling operation, respectively;

[0054] Process the channel multi-scale pooling features to obtain the channel adaptive pooling features of the optical modality and And channel adaptive pooling features of SAR modality and in The steps for processing the channel multi-scale pooling features are:

[0055]

[0056] Where σ is the sigmoid function, is the parameter with dimension L×1 to be optimized, is the parameter with dimension 1 to be optimized;

[0057] Step 302: Adaptively pool the obtained channel features and Splicing is performed in the channel dimension to obtain fused multi-scale pooling features

[0058] Perform feature fusion and channel conversion on the fused multi-scale pooling features to obtain the fused channel correction vector

[0059] Where σ is the sigmoid function, is the parameter to be optimized with a dimension of 4C′×2, is the parameter with dimension 2 to be optimized; Perform channel segmentation to obtain the channel correction vector of the optical mode and the channel correction vector of the SAR mode in

[0060] Step 303: Use the channel correction vector to correct the optical mode self-characteristic and the SAR mode self-characteristic. The correction steps are as follows:

[0061]

[0062] Finally, the characteristic map of the optical mode after channel dimension correction is obtained And the characteristic diagram of SAR mode in

[0063] Step 4: Construct a cross-modal feature correction module in the spatial dimension to correct the features of the optical image and the SAR image in the spatial dimension;

[0064] Step 401: Downsample the feature map of the optical mode and the feature map of the SAR mode after the channel dimension correction in the channel dimension through average pooling operation and maximum pooling operation, and use four pooling factors to splice the pooling results along the spatial dimension to obtain the spatial multi-scale pooling features of the optical mode respectively. and and spatial multi-scale pooling features of SAR modalities and in in The subscripts ap and mp represent the results of the average pooling operation and the maximum pooling operation respectively; the four different pooling factor sizes are N' 1 、N' 2 、N' 3 and N' 4 , the size of the pooling result is

[0065] Step 402: The obtained spatial multi-scale pooling features of the optical mode and the SAR mode and Splicing along the channel dimension to obtain the fusion space multi-scale pooling features

[0066] Adaptive selection and channel change of multi-scale pooling features in fusion space to obtain fusion space correction weight map

[0067]

[0068] Where σ is the sigmoid function, The parameter to be optimized is of dimension 4M×2×3×3. is the parameter with dimension 2 to be optimized;

[0069] Will Perform channel segmentation to obtain the optical spatial correction weight map and SAR spatial correction weight map in

[0070] Step 403: Correct the optical modality feature map and the SAR modality feature map after channel dimension correction using the optical spatial correction weight map and the SAR spatial correction weight map to obtain the optical modality feature map after spatial dimension correction. Characteristic map of SAR mode after spatial dimension correction

[0071]

[0072] in

[0073] Step 5: construct an interactive redundancy removal module to remove redundant information contained in the optical modal feature map and the SAR modal feature map, and fuse the features after removing the redundant information;

[0074] Step 501: The step of removing the optical modal feature map is:

[0075] Will Straighten to Where N = H′×W′;

[0076] Generating optical modal interaction features and the optical modal residual characteristics

[0077] in σ is the sigmoid function, is the parameter to be optimized with dimension C′×2C′, is the parameter of dimension 2C′ to be optimized;

[0078] Will The straightening

[0079] Generating optical modal interaction features and the optical modal residual characteristics

[0080]

[0081] in σ is the sigmoid function, is the parameter to be optimized with dimension C′×2C′, The parameter to be optimized is of dimension 2C′, N = H′×W′;

[0082] The steps to remove the SAR modal feature map are the same as those to remove the optical modal feature map, generating the SAR modal interaction feature map. and SAR modal residual characteristics

[0083] Step 502: Optical modal interaction features Interaction characteristics with SAR modalities Input the self-attention encoding network module respectively;

[0084] The self-attention encoding network module modifies the linear transformation parameters and outputs four elements: key, self-query, mutual query and value;

[0085] The linear transformation of the existing self-attention encoding network will linearly transform the input features to obtain three elements: key, query, and value. The self-attention encoding network used in this method modifies the linear transformation weight parameters from three channel dimensions to four channel dimensions to obtain four elements: key, self-query, mutual query, and value.

[0086] The self-attention encoding network module generates the key K of optical features OPT , Self-query of optical features Optical feature cross-check and the value of the optical characteristic V OPT :

[0087]

[0088] The key K of the self-attention encoding network module to generate SAR features SAR , Self-query of SAR features Mutual query of SAR features and the SAR characteristic value V SAR :

[0089]

[0090] Where K OPT ∈R N×C′ , V OPT ∈R N×C′ , K SAR ∈R N×C′ , V SAR∈R N×C′ ,σ is the sigmoid function, and is the parameter to be optimized with dimension C′×4C′, and is the parameter of dimension 4C′ to be optimized;

[0091] Computing the optical modality redundancy attention matrix and the optical modality self-attention matrix

[0092]

[0093]

[0094] in

[0095] Compute the redundant attention matrix for SAR modality And the self-attention matrix of SAR modality

[0096]

[0097] in, Softmax is a function, T is a matrix transpose operation, d k is the scaling factor;

[0098] Calculate the attention matrix after removing redundancy of the optical modality Attention matrix after removing redundancy with SAR modality

[0099]

[0100] in

[0101] Calculate the optical mode characteristics after removing redundancy And the SAR modal characteristics after removing redundancy

[0102]

[0103] in

[0104] The redundant features and residual features are added element by element and normalized. The optical modality is normalized to remove redundant features. and SAR modal normalization to remove redundant features for:

[0105]

[0106] Among them, LN is layer normalization, which completes the de-redundancy of feature redundant information;

[0107] Step 503: extracting features from the optical mode and SAR mode features after redundancy removal;

[0108] Intermediate variables and intermediate variables for:

[0109]

[0110] Where σ is the sigmoid function, W 3 OPT and is the parameter of dimension C′×C′ to be optimized, and is the parameter of dimension C′ to be optimized,

[0111] Will Perform channel changes to obtain the features to be fused in the optical modality Will Perform channel changes to obtain the features to be fused by the SAR module And concatenate them in the channel dimension to get the fusion feature X Fuse ∈R H′×W′×2C′ ;in

[0112] Step 6: Input the fused features into the feature decoder module and train to obtain a complete neural network model based on feature correction and redundancy removal;

[0113] Step 601: Fusion feature X Fuse In the input feature decoder module, two convolution blocks with shared parameters are used and By interactively updating parameters, the fusion feature X Fuse Decode and segment, then upsample the segmentation result to the original image size by bilinear interpolation, and finally obtain the classification result y, where y∈R H ×W×class , class is the number of segmentation categories; the segmentation result can be upsampled to the original image size using other interpolation methods such as the nearest neighbor algorithm and the bicubic interpolation algorithm;

[0114] Step 602: input the training set data into the neural network model based on feature correction and redundancy removal to obtain a fully trained neural network model based on feature correction and redundancy removal;

[0115] The loss function Loss of the neural network model based on feature correction and redundancy removal is the cross entropy loss between the predicted value y and the category Y of the corresponding pixel point:

[0116] Where N = H × W, class is the number of categories;

[0117] When the loss function of the neural network model for feature correction and de-redundancy is less than 0.3 or the number of training rounds reaches 100, the neural network model for feature correction and de-redundancy is the neural network model with the best performance; the training data is sent to the model for training in sequence, and one round of training is a complete traversal;

[0118] Step 7: Use a complete neural network model based on feature correction and redundancy removal to predict the test set data and obtain the segmentation results.

Claims

1. A method for semantic segmentation of optical and SAR images based on feature correction and redundancy removal, characterized by: A neural network model for feature correction and de-redundancy and a semantic segmentation method based on the neural network model for feature correction and de-redundancy; the neural network model for feature correction and de-redundancy includes a self-feature extraction backbone module, a cross-modal feature correction module in the channel dimension, a cross-modal feature correction module in the spatial dimension, an interactive de-redundancy module, a feature decoder module and a loss function; The self-feature extraction backbone module includes a self-feature extraction module for optical images and a self-extraction module for SAR images; the input of the self-feature extraction backbone module is an optical image and a SAR image; the output of the self-feature extraction backbone module is an optical modality self-feature and a SAR modality self-feature; the self-feature extraction module for optical images includes three convolution blocks connected in series in sequence; the self-extraction module for SAR images includes three convolution blocks connected in series in sequence; The cross-modal feature correction module of the channel dimension obtains the channel adaptive pooling features of the optical mode and the channel adaptive pooling features of the SAR mode by average pooling operation and maximum pooling operation respectively on the optical mode self-features and the SAR mode self-features, and splices the channel adaptive pooling features of the optical mode and the channel adaptive pooling features of the SAR mode in the channel dimension to obtain fused multi-scale pooling features; performs feature fusion and channel conversion on the fused multi-scale pooling features to obtain a fused channel correction vector; and then uses the channel correction vector to correct the self-features of the optical mode and the self-features of the SAR mode to obtain a feature map of the optical mode and a feature map of the SAR mode after the channel dimension correction; The cross-modal feature correction module in the spatial dimension downsamples the feature map of the optical modality and the feature map of the SAR modality corrected in the channel dimension in the channel dimension through average pooling operations and maximum pooling operations, and obtains the spatial multi-scale pooling features of the optical modality and the spatial multi-scale pooling features of the SAR modality respectively; The spatial multi-scale pooling features of the optical mode and the spatial multi-scale pooling features of the SAR mode are spliced ​​along the channel dimension to obtain the fused spatial multi-scale pooling features; Adaptively select and change the channels of the fused spatial multi-scale pooling features to obtain a fused spatial correction weight map; perform channel segmentation on the fused spatial correction weight map to obtain an optical spatial correction weight map and a SAR spatial correction weight map; use the optical spatial correction weight map and the SAR spatial correction weight map to correct the feature map of the optical modality and the feature map of the SAR modality after channel dimension correction, and obtain the feature map of the optical modality after spatial dimension correction and the feature map of the SAR modality after spatial dimension correction; The interactive redundancy removal module includes an optical modality feature map removal module, a SAR modality feature map removal module and a self-attention encoding network module; The optical modal feature map removal module straightens the optical modal feature map after the spatial dimension correction to obtain optical modal interaction features and optical modal residual features respectively; the SAR modal feature map removal module straightens the feature map of the SAR modality after the spatial dimension correction to obtain SAR modal interaction features and SAR modal residual features respectively; the self-attention encoding network module modifies the linear transformation weight parameter from three channel dimensions to four channel dimensions to obtain four elements of key, self-query, mutual query and value; the optical modal interaction feature is input into the self-attention encoding network module to obtain the key of the optical feature, the self-query of the optical feature, the mutual query of the optical feature and the value of the optical feature; the SAR modal interaction feature is input into the self-attention encoding network module to obtain the key of the SAR feature, the self-query of the SAR feature, the mutual query of the SAR feature and the value of the SAR feature; According to the output of the self-attention encoding network module, the redundant attention matrix and self-attention matrix of the optical modality as well as the redundant attention matrix and self-attention matrix of the SAR modality are calculated using the softmax function respectively; Subtracting the optical modality self-attention matrix from the optical modality redundant attention matrix to obtain the optical modality attention matrix after removing redundancy; Subtract the self-attention matrix of the SAR modality from the redundant attention matrix of the SAR modality to obtain the redundantly removed attention matrix of the SAR modality; Multiplying the redundantly removed attention matrix of the optical modality with the value of the optical feature to obtain the redundantly removed optical modality feature; The redundantly removed attention matrix of the SAR modality is multiplied by the value of the SAR feature to obtain the redundantly removed SAR modality feature; The optical modal features after removing redundancy are added element by element to the optical modal residual features and normalized to obtain the optical modal normalized de-redundancy features; The SAR modal features after removing redundancy are added element by element to the SAR modal residual features and normalized to obtain the SAR modal normalized de-redundancy features; The normalized de-redundant features of the optical mode and the normalized de-redundant features of the SAR mode are processed using the sigmoid function respectively, and the processed results are subjected to channel changes to obtain the features to be fused of the optical mode and the features to be fused of the SAR module; The features to be fused of the optical modality and the features to be fused of the SAR module are spliced ​​in the channel dimension to obtain fused features; The input of the feature decoder module is the fusion feature, and the output is the final classification result y; The feature correction and de-redundancy neural network model is trained using the training set. When the loss function of the feature correction and de-redundancy neural network model is less than 0.3 or the training rounds reach 100 rounds, the feature correction and de-redundancy neural network model is the feature correction and de-redundancy neural network model with the best performance.

2. The method for semantic segmentation of optical and SAR images based on feature correction and redundancy removal according to claim 1, characterized in that: The steps of the semantic segmentation method of the redundant neural network model are as follows: Step 1: Obtain optical and SAR image data sets and preprocess them, and divide the preprocessed optical image and SAR image data into a training set and a test set; Step 2: Input the training set into the neural network model for feature correction and redundancy removal for training; Step 3: Use the trained and optimally performing neural network model with feature correction and redundancy removal to predict the test set data and obtain the segmentation result.

3. The method for semantic segmentation of optical and SAR images based on feature correction and redundancy removal according to claim 2, characterized in that: The preprocessing normalizes the width, height and number of channels of the original optical image and the original SAR image respectively, and classifies them; The optical images and SAR images obtained after preprocessing are divided into training sets and test sets.

4. The method for semantic segmentation of optical and SAR images based on feature correction and redundancy removal according to claim 1 or claim 2, characterized in that: The loss function Loss of the neural network model for feature correction and redundancy removal is the cross entropy loss between the predicted value y and the category Y of the corresponding pixel point: Where N = H × W, class is the number of categories.

5. The method for semantic segmentation of optical and SAR images based on feature correction and redundancy removal according to claim 1 or claim 2, characterized in that: The feature decoder module adopts two convolution blocks with shared parameters to decode and segment the fused features by interactively updating the parameters, and then upsamples the segmentation results to the original image size by interpolation to finally obtain the classification result y.

6. The method for semantic segmentation of optical and SAR images based on feature correction and redundancy removal according to claim 5, characterized in that: The interpolation methods include bilinear interpolation, nearest neighbor algorithm and bicubic interpolation algorithm.

7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed, the optical and SAR image semantic segmentation method based on feature correction and redundancy removal according to any one of claims 1 to 6 is implemented.