Multi-modal remote sensing image semantic segmentation method and device based on asymmetric semantic alignment

By using a multimodal remote sensing image semantic segmentation method with asymmetric semantic alignment, the problem of poor design of optical and SAR image fusion methods is solved, and high-precision remote sensing image segmentation under different weather conditions is achieved.

CN121811033APending Publication Date: 2026-04-07PERCEPTION WORLD (BEIJING) INFORMATION TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-18
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing remote sensing image semantic segmentation methods do not fully utilize the complementarity between optical and SAR images, resulting in poor performance in fusion design, especially with reduced segmentation accuracy under adverse weather conditions.

Method used

A multimodal remote sensing image semantic segmentation method with asymmetric semantic alignment is adopted. By acquiring registered optical and SAR images of the same area, a set of land cover category mask images is generated. After preprocessing, feature encoders and decoders are used for encoding and decoding. The asymmetry between modes is fully considered to align the feature space and channels, thereby reducing noise interference.

Benefits of technology

It improves the accuracy and quality of semantic segmentation of remote sensing images and enhances the segmentation effect under different weather conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121811033A_ABST
    Figure CN121811033A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a multi-mode remote sensing image semantic segmentation method and device based on asymmetric semantic alignment. The method comprises the following steps: acquiring a registered optical image and a registered SAR image of the same area; generating a ground object category mask image set according to the registered optical image and the registered SAR image; pre-processing the registered optical image, the registered SAR image and the ground object category mask image set to obtain a pre-processed optical image, a pre-processed SAR image and a pre-processed ground object category mask image set; encoding the preprocessed optical image and the preprocessed SAR image to obtain a fusion feature image set; and inputting the fused feature image set into a decoder to obtain a segmented feature image. According to the embodiment of the invention, by fusing and aligning the complementary information between the two modals, the interference of noise information is reduced, and the precision and quality of the remote sensing image semantic segmentation image are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing, and more particularly to a method, server, medium, and related equipment for multimodal remote sensing image semantic segmentation with asymmetric semantic alignment. Background Technology

[0002] Currently, semantic segmentation of remote sensing images mainly relies on single-source satellite imagery, such as optical imagery, for semantic feature extraction. Optical satellite imagery has high resolution and is rich in spectral and textural information. However, due to the limitations of optical imagery imaging methods, it is susceptible to cloud cover and severe weather, which can obscure ground feature information and reduce the accuracy of semantic segmentation. SAR sensors can operate normally under various weather conditions and can penetrate certain types of ground cover. In such cases, optical and SAR imagery can complement each other. In recent years, to fully utilize the complementarity of optical and SAR imagery, most approaches focus on fusion interpretation based on the complementarity of the two modalities, mainly using three methods: data-level, feature-level, and decision-level. However, while current fusion methods retain consistent feature information between the two modalities, their design does not fully utilize the unique characteristics between the two modalities, resulting in poor semantic segmentation performance. Summary of the Invention

[0003] To address the aforementioned problems, embodiments of the present invention provide a multimodal remote sensing image semantic segmentation method, server, medium, and related equipment with asymmetric semantic alignment. A remote sensing image semantic segmentation method integrating two modalities is proposed.

[0004] One aspect of the present invention provides a multimodal remote sensing image semantic segmentation method with asymmetric semantic alignment, comprising:

[0005] Step 101: Acquire registered optical images and registered SAR images of the same area;

[0006] Step 102: Generate a set of land cover category mask images based on the above-registered optical images and the above-registered SAR images;

[0007] Step 103: Preprocess the above-mentioned registered optical image, the above-mentioned registered SAR image and the above-mentioned land cover category mask image set to obtain the preprocessed optical image, the preprocessed SAR image and the preprocessed land cover category mask image set.

[0008] Step 104: Encode the preprocessed optical image and the preprocessed SAR image to obtain a fused feature map set;

[0009] Step 105: Input the above fused feature map set into the decoder to obtain the segmentation feature map.

[0010] Optionally, the above-mentioned encoding of the preprocessed optical image and the preprocessed SAR image to obtain a fused feature map set includes: inputting the preprocessed optical image, the preprocessed SAR image, and the preprocessed land cover category mask image set into a feature encoder to obtain the fused feature map set. The feature encoder includes a predetermined number of feature extraction layers, all of which have a symmetrical structure and include a convolutional layer module, a feature attention module, and a semantic fusion module. The fused feature map set includes a predetermined number of fused feature maps, and the preprocessed land cover category mask image set serves as a label.

[0011] Optionally, the above-mentioned inputting the preprocessed optical image, the preprocessed SAR image, and the preprocessed land cover category mask image set into the feature encoder to obtain the fused feature map set includes: inputting the preprocessed optical image and the preprocessed SAR image into the first convolutional layer module respectively to obtain the initial feature map output by the optical branch and the initial feature map output by the SAR branch; inputting the initial feature map output by the optical branch and the initial feature map output by the SAR branch into the first feature attention module to obtain the first weighted optical feature map and the first weighted SAR feature map; and inputting the above-mentioned... The first weighted optical feature map and the first weighted SAR feature map are input into the first semantic fusion module to obtain the first-level fused feature map. The initial feature maps output from the optical branch and the SAR branch are input into the second convolution module to obtain the second optical feature map and the second SAR feature map. The second optical feature map and the second SAR feature map are input into the second feature attention module to obtain the second weighted optical feature map and the second weighted SAR feature map. The second weighted optical feature map and the second weighted SAR feature map are input into the second semantic fusion module to obtain the second-level fused feature map. The third optical feature map and the third SAR feature map are then fed into a third convolutional module to obtain a third optical feature map and a third SAR feature map. These three features are then fed into a third feature interest module to obtain a third weighted optical feature map and a third weighted SAR feature map. Finally, these features are fed into a third semantic fusion module to obtain a third-level fused feature map. The third optical feature map and the third SAR feature map are then fed into a fourth convolutional module to obtain... The fourth optical feature map and the fourth SAR feature map are input into the fourth feature attention module to obtain the fourth weighted optical feature map and the fourth weighted SAR feature map. The fourth weighted optical feature map and the fourth weighted SAR feature map are input into the fourth semantic fusion module to obtain the fourth-level fused feature map. The first-level fused feature map, the second-level fused feature map, the third-level fused feature map and the fourth-level fused feature map are saved, wherein the resolution of the fused feature map at each level decreases layer by layer, and the semantic information increases layer by layer.

[0012] Optionally, the first feature attention module, the second feature attention module, the third feature attention module, and the fourth feature attention module have the same structure and function, and are respectively connected to the first convolution module, the second convolution module, the third convolution module, and the fourth convolution module; the first feature attention module, the second feature attention module, the third feature attention module, and the fourth feature attention module perform the following steps on the initial feature map output by the optical branch and the initial feature map output by the SAR branch, the second optical feature map and the second SAR feature map, the third optical feature map and the third SAR feature map and the fourth optical feature map and the fourth SAR feature map, respectively: generating a difference feature map combining the SAR feature map and the optical feature map based on the received SAR feature map and the optical feature map; inputting the difference feature map to the pooling layer to generate initial weight values; inputting the initial weight values ​​to the channel perceptron to obtain global difference weights; and performing weight correction on the received SAR feature map and the optical feature map based on the global difference weights to obtain weighted optical feature maps and weighted SAR feature maps.

[0013] Optionally, the results and functions of the first, second, third, and fourth semantic fusion modules are all the same, and they are connected to the first, second, third, and fourth feature attention modules, respectively; the first, second, third, and fourth semantic fusion modules respectively process the first weighted optical feature map and the first weighted SAR feature map, the second weighted optical feature map and the second weighted SAR feature map, the third weighted optical feature map and the fourth feature attention module, respectively. The third weighted SAR feature map, the fourth weighted optical feature map, and the fourth weighted SAR feature map are processed using the following steps: feature alignment and calibration are performed on the received weighted optical feature map and weighted SAR feature map to obtain a pre-processed optical feature map and a pre-processed SAR feature map; spatial alignment and calibration are performed on the pre-processed optical feature map and the pre-processed SAR feature map to obtain a calibrated optical feature map and a calibrated SAR feature map; the calibrated optical feature map and the calibrated SAR feature map are added pixel-level to obtain a fused feature map.

[0014] Optionally, the above-mentioned inputting the fused feature map set into the decoder to obtain the segmentation feature map includes: upsampling the fourth-level fused feature map by a factor of 2 to obtain a 2-sampled fourth fused map; fusing the third-level fused feature map with the 2-sampled fourth fused map to obtain a decoded third-level fused feature map; upsampling the decoded third-level fused feature map by a factor of 2 to obtain a 2-sampled third-level fused map; fusing the second-level fused feature map with the 2-sampled third-level fused map to obtain a decoded second-level fused feature map; upsampling the decoded second-level fused feature map by a factor of 2 to obtain a 2-sampled second-level fused map; fusing the first-level fused feature map with the 2-sampled second-level fused map to obtain a finally decoded first-level fused feature map; inputting the finally decoded first-level fused feature map into a convolutional layer to obtain a segmentation score map; and classifying each pixel according to the segmentation score map to obtain the above-mentioned segmentation feature map.

[0015] Optionally, the aforementioned feature encoder and decoder constitute a semantic segmentation model for multimodal remote sensing images. During the training phase, the semantic segmentation model for multimodal remote sensing images compares the feature map output by the decoder, which has the same pixel size as the original image, with the ground truth using a multi-class cross-entropy loss function.

[0016] Another aspect of the present invention provides a server, characterized in that the server runs the aforementioned remote sensing image change detection method.

[0017] Another aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions, wherein the computer program instructions, when executed by a processor, cause the processor to perform any of the methods described above.

[0018] Three other aspects of the present invention also provide an electronic device, including: a processor; and a memory storing instructions executable by the processor; wherein the processor is configured to execute the instructions to implement any of the methods described above.

[0019] The multimodal remote sensing image semantic segmentation method and related equipment with asymmetric semantic alignment provided by the embodiments of the present invention fully consider the asymmetry between modalities when designing the network, and perform alignment in feature space and channels, so that it can adaptively calibrate and align complementary information between two modalities, reduce the interference of noise information, and improve the accuracy and quality of remote sensing image semantic segmentation.

[0020] Further aspects and scope of adaptation become apparent from the description provided herein. It should be understood that various aspects of this application may be implemented individually or in combination with one or more other aspects. It should also be understood that the description and specific embodiments herein are for illustrative purposes and are not intended to limit the scope of this application. Attached Figure Description

[0021] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:

[0022] Figure 1 This is an overall flowchart of a multimodal remote sensing image semantic segmentation method with asymmetric semantic alignment according to an embodiment of the present invention;

[0023] Figure 2 This is an overall structural diagram of the semantic segmentation model of multimodal remote sensing images in an embodiment of the present invention;

[0024] Figure 3 This is a structural diagram of the feature attention module in an embodiment of the present invention;

[0025] Figure 4 This is a structural diagram of the semantic fusion module in an embodiment of the present invention. Detailed Implementation

[0026] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0027] like Figure 1 As shown, the asymmetric semantic alignment multimodal remote sensing image semantic segmentation method in this embodiment of the invention includes:

[0028] Step 101: Acquire registered optical images and registered SAR images of the same area. The registration process may include geometric correction and spatial alignment. The optical images can be images that record the spectral information of ground features based on the reflection of sunlight. The SAR images can be images obtained through active microwave remote sensing, which transmits and receives electromagnetic waves and records information such as the dielectric constant, surface roughness, and geometric structure of ground features. In practice, registered optical images and registered SAR images can be acquired through methods such as calling from internal memory or external databases, or receiving raw data from sensors in real time and then performing the registration operation.

[0029] Step 102: Generate a set of land cover category mask images based on the registered optical image and the registered SAR image. In practice, corresponding mask images can be created based on the land cover information in the above-mentioned registered optical image and the above-mentioned registered SAR image. Taking typical land cover such as buildings, water bodies, and roads as examples, the pixel value corresponding to the mask image of the building area is 1, the pixel value corresponding to the mask image of the water body area is 2, and the pixel value corresponding to the mask image of the road area is 3.

[0030] Step 103: Preprocess the registered optical image, the registered SAR image, and the land cover category mask image set to obtain the preprocessed optical image, the preprocessed SAR image, and the preprocessed land cover category mask image set. The preprocessing may include random rotation, image brightness, and saturation adjustments.

[0031] Step 104: Encode the preprocessed optical image and the preprocessed SAR image to obtain the fused feature map set.

[0032] In some optional embodiments, the above-mentioned encoding of the preprocessed optical image and the preprocessed SAR image to obtain a fused feature map set includes inputting the preprocessed optical image, the preprocessed SAR image, and the preprocessed land cover category mask image set into a feature encoder to obtain the fused feature map set. The feature encoder includes a predetermined number of feature extraction layers, all of which are symmetrical structures and include a convolutional layer module, a feature attention module, and a semantic fusion module. The fused feature map set includes a predetermined number of fused feature maps. The predetermined number can be 4. The overall structure of the feature encoder can be as follows: Figure 2 The encoder section is shown in the diagram. The feature encoder described above can be a two-branch ResNet50 structure. The preprocessed land cover category mask image set is used as the label, which can be used as the ground truth label (which can be a binary image) for subsequent models in practice.

[0033] In some alternative embodiments, inputting the preprocessed optical image and the preprocessed SAR image into the feature encoder to obtain the fused feature map may include the following steps:

[0034] The first step involves inputting the preprocessed optical image and the preprocessed SAR image into the first convolutional layer module, respectively, to obtain the initial feature maps for the optical branch and the SAR branch. The first convolutional layer module may include a symmetric convolutional layer, a feature attention module, and a semantic fusion module. The initial feature map for the optical branch output may include basic visual features extracted from the optical image, such as edges, colors, textures, and simple shapes; it is a first-level abstraction of the original pixels, denoted as F. rgb The initial feature map output by the aforementioned SAR branch can include basic scattering feature maps extracted from SAR images, such as brightness (corresponding to backscattering intensity), speckle texture, and simple geometric contours, denoted as F. sar Each of the above feature maps has a size of C×H×W, representing the number of channels, height, and width of the feature map, respectively.

[0035] The second step involves inputting the initial feature maps from the optical branch and the SAR branch into the first feature attention module to obtain the first weighted optical feature map and the first weighted SAR feature map. The first feature attention module can be a neural network component based on a channel attention mechanism embedded in each level of the encoder. The first weighted optical feature map can be denoted as... The above first weighted SAR feature map

[0036] The third step involves inputting the first weighted optical feature map and the first weighted SAR feature map into the first semantic fusion module to obtain the first-level fused feature map. The first semantic fusion module can be a neural network component employing a dual-path attention mechanism, immediately following the feature attention module. The first-level fused feature map can be a high-resolution, low-semantic-information feature map that correlates features between the two modalities and retains the high-resolution, low-semantic-information feature map generated by the fusion of the first convolutional module in the dual-branch encoder, denoted as...

[0037] Fourth, the initial feature maps output from the optical branch and the SAR branch are input into the second convolutional module to obtain the second optical feature map and the second SAR feature map. The second convolutional module is connected after the first convolutional module, and its structure and composition can be the same as the first convolutional module. The second optical feature map can be a more detailed optical feature map, i.e., it combines more complex textures and patterns (such as window arrangements or roof structures), denoted as F2. rgb The second SAR feature map mentioned above can be a more detailed SAR feature map, reflecting a more macroscopic geometric structure (such as the strong reflection profile of large buildings, or uniform scattering in flat areas), denoted as F2. sar .

[0038] Fifth, the second optical feature map and the second SAR feature map are input into the second feature attention module to obtain the second weighted optical feature map and the second weighted SAR feature map. The second feature attention module is linked after the first feature attention module, and its structure and function can be the same as the first feature attention module. The second weighted optical feature map can be denoted as... The second weighted SAR feature map mentioned above can be denoted as:

[0039] Step 6: Input the second weighted optical feature map and the second weighted SAR feature map into the second semantic fusion module to obtain the second-level fused feature map. The second semantic fusion module can be connected after the first semantic fusion module, and its function and structure can be the same as the first semantic fusion module. The second-level fused feature map can be a feature map that correlates features between the two modalities and retains the medium-resolution, low-semantic-information feature map generated by the fusion of the second convolutional module in the dual-branch encoder, denoted as...

[0040] Step 7: Input the second weighted optical feature map and the second weighted SAR feature map into the third convolutional module to obtain the third optical feature map and the third SAR feature map, respectively. The third convolutional module is connected after the second convolutional module, and its structure and function can be the same as the second convolutional module. The third optical feature map can be a more advanced optical feature map, capable of identifying object components (such as the overall outline of a building or the shape of a vehicle), denoted as F3. rgb The aforementioned third SAR feature map can be a more advanced SAR feature map, capturing the scattering patterns of typical ground features (such as strong bright lines on metal roofs and scattering characteristics of rough surfaces), denoted as F3. sar .

[0041] Step 8: Input the second optical feature map and the second SAR feature map into the third feature attention module to obtain the third weighted optical feature map and the third weighted SAR feature map. The third feature attention module can be connected after the second feature attention module, and its structure and function can be the same as the second feature attention module. The third weighted optical feature map can be denoted as... The third weighted SAR feature map mentioned above can be denoted as:

[0042] Step nine involves inputting the third weighted optical feature map and the third weighted SAR feature map into the third semantic fusion module to obtain the third-level fused feature map. This third semantic fusion module can be connected after the second semantic fusion module, and its structure and function can be the same. The third-level fused feature map can be a feature map that correlates features between two modalities and retains the lower-resolution, moderately semantically informational feature map generated by the fusion of features from the third convolutional module in the dual-branch encoder; denoted as...

[0043] Step 10: Input the third optical feature map and the third SAR feature map into the fourth convolutional module to obtain the fourth optical feature map and the fourth SAR feature map, respectively. The fourth convolutional module can be a module connected after the third convolutional module, and its structure and function can be the same as the third convolutional module. The fourth optical feature map can be the final optical feature map, containing the strongest semantic information and capable of clearly distinguishing high-level concepts such as "building complex," "water area," and "road network," denoted as F4. rgb The aforementioned fourth SAR feature map, the final SAR feature map, includes top-level semantic information and understands the scene from the radar's perspective. For example, it can distinguish between open waters (specular reflection, dark) and urban areas (multiple reflections, bright) based on scattering mechanisms, and is denoted as F4. sar .

[0044] Step 11: Input the aforementioned third optical feature map and third SAR feature map into the fourth feature attention module to obtain the fourth weighted optical feature map and fourth weighted SAR feature map. The fourth feature attention module can be connected after the third feature attention module, and its structure and function can be the same as the third feature attention module. The aforementioned fourth weighted optical feature map can be denoted as... The above fourth weighted SAR feature map can be denoted as:

[0045] Step 12: Input the fourth weighted optical feature map and the fourth weighted SAR feature map into the fourth semantic fusion module to obtain the fourth-level fused feature map. The fourth semantic fusion module can be connected after the third semantic fusion module, and its structure and function can be the same as the third semantic fusion module. The fourth-level fused feature map can be a feature map that correlates the features between the two modalities and retains the low-resolution, high-semantic-information feature map generated by the fusion of the third convolutional module in the dual-branch encoder, denoted as...

[0046] Step 13: Save the first-level fusion feature map, the second-level fusion feature map, the third-level fusion feature map, and the fourth-level fusion feature map. The resolution of the fusion feature map at each level decreases layer by layer, while the semantic information increases layer by layer.

[0047] In some optional embodiments, the first feature attention module, the second feature attention module, the third feature attention module, and the fourth feature attention module all have the same structure and function, and are respectively connected to the first convolutional module, the second convolutional module, the third convolutional module, and the fourth convolutional module. The structure of each feature attention module can be as follows: Figure 3 As shown. The first feature attention module, the second feature attention module, the third feature attention module, and the fourth feature attention module can respectively perform the following steps on the initial feature map output by the optical branch, the initial feature map output by the SAR branch, the second optical feature map, the second SAR feature map, the third optical feature map, the third SAR feature map, and the fourth optical feature map and the fourth SAR feature map:

[0048] The first step is to generate a difference feature map combining the received SAR and optical feature maps. In practice, the generation formulas for each difference feature map can be as follows (taking the initial feature map output from the optical branch and the initial feature map output from the SAR branch as examples):

[0049]

[0050] in, It can be a difference feature map on optical images. This can be a difference feature map on a SAR image. The generation method for each set of difference feature maps is the same; for example, the formula for generating the feature difference map of the second optical feature map and the second SAR feature map mentioned above is as follows: in It can be a second optical feature difference map. It can be a feature difference map of the second SAR feature map.

[0051] The second step involves inputting the aforementioned difference feature maps into the pooling layer to generate initial weight values. In practice, global max pooling (GMP) can be performed on a set of (optical, SAR) difference feature maps to obtain the initial optical channel weights. and initial SAR channel weights

[0052] The third step involves inputting the initial weight values ​​into the channel perceptron to obtain the global difference weights. The channel perceptron can be a module composed of a preset number of convolutional layers and GELU activation functions; this preset number is not specifically limited. The global difference weights may include optical global difference weights. and SAR global difference weights The specific formula can be seen as follows:

[0053]

[0054]

[0055] Where F GMP This indicates global max pooling. F1 and F2 represent two different convolutional layers with output dimensions of... C can be the output dimension, F GELU This represents the GELU activation function.

[0056] The fourth step involves adjusting the weights of the received SAR and optical feature maps based on the aforementioned global difference weights, resulting in weighted optical and SAR feature maps. In practice, the weights can be normalized using the Sigmoid function and then multiplied channel-by-channel with the original feature map to complete the weight adjustment. The formula is shown below:

[0057]

[0058] and This is the feature map output after weight correction by the feature attention module. σ can be the sigmoid function. This represents pixel-level multiplication. Here, w can be 1, 2, 3, or 4, which respectively represent the pixel-level multiplication. and stated The and stated The and stated and the and stated

[0059] In some optional embodiments, the results and functions of the first semantic fusion module, the second semantic fusion module, the third semantic fusion module, and the fourth semantic fusion module are all the same, and they are respectively connected to the first feature attention module, the second feature attention module, the third feature attention module, and the fourth feature attention module. The structure of each semantic fusion module can be as follows: Figure 4As shown. The first semantic fusion module, the second semantic fusion module, the third semantic fusion module, and the fourth semantic fusion module respectively perform the following steps on the first weighted optical feature map and the first weighted SAR feature map, the second weighted optical feature map and the second weighted SAR feature map, the third weighted optical feature map and the third weighted SAR feature map, and the fourth weighted optical feature map and the fourth weighted SAR feature map:

[0060] The first step involves feature alignment and calibration of the received weighted optical feature map and weighted SAR feature map to obtain pre-processed optical feature map and pre-processed SAR feature map. In practice, the two feature maps F can be first... rgb ,F sar (Taking the initial feature maps output by the optical branch and the initial feature maps output by the SAR branch as examples, the combination of other stages is similar.) These are then concatenated, and convolutional layers are used to fuse the concatenated feature maps. Global average pooling is then used to obtain the feature vector s∈R. C×1×1 Global average pooling compresses the spatial dimension of the connected feature maps to generate channel-level statistics. The feature vector s is then input into two independent multilayer perceptrons (MLPs). After processing with the sigmoid function, each modality receives different fusion channel weights. Multiplying these weights by the original feature map yields the optical feature map obtained after the initial processing. And the SAR feature map after preliminary processing mentioned above The specific formula is shown below:

[0061]

[0062] Where cat represents connecting feature maps along the channel dimension, F GAP Represents global average pooling. This represents a convolutional layer with an input dimension of 2C and an output dimension of C. F MLP1 and F MLP2 These represent MLP, respectively. σ is the sigmoid function. This represents pixel-level multiplication. The attention module can learn the attention weights for different channels, thereby enabling adaptive selection and calibration of the original features from both branches.

[0063] The second step involves spatially aligning and calibrating the pre-processed optical and SAR feature maps to obtain calibrated optical and SAR feature maps. In practice, the feature maps can be first... and The features are connected together, and the convolutional processing of the connected feature maps is performed to fuse information, generating a feature map Z∈R. 1×H×W The feature map is then convolved with two different sets of weights to enhance the feature representation. A softmax function is then used to obtain weight scores for each modality dimension, and these scores are used to refine the feature map. and By performing weighted summaries, the calibrated optical feature map described above is obtained. And the SAR feature map that has been calibrated as described above The specific formula is as follows:

[0064]

[0065] Here, cat represents connecting feature maps along the channel dimension. F3 represents a convolutional layer with an input dimension of 2C and an output dimension of C. F4 can represent different convolutional layers, but the input and output dimensions are both C. This indicates pixel-level multiplication, and softmax is the softmax function.

[0066] The third step involves pixel-wise addition of the calibrated optical feature map and the calibrated SAR feature map to obtain a fused feature map. In practice, this can be done... and Perform pixel-level addition to generate a fused feature map. and The formula is as follows:

[0067]

[0068] Where i = 1, 2, 3, 4, and + represents pixel-level addition operation.

[0069] In some optional embodiments, the feature encoder and the decoder constitute a semantic segmentation model for multimodal remote sensing images, and the overall structure of the semantic segmentation model for multimodal remote sensing images can be as follows: Figure 2 As shown. The semantic segmentation model for the aforementioned multimodal remote sensing image compares the feature map output by the decoder (with the same pixel size as the original image) with the ground truth during the training phase, using a multi-class cross-entropy loss function for comparison. The formula for the multi-class cross-entropy loss function is shown below:

[0070]

[0071] Where M represents the number of categories, y ic The sign function is p, which takes the value 1 if the true class of sample i equals the predicted class c, and 0 otherwise. icLet be the predicted probability that observed sample i belongs to category c. Finally, when the value of the contrastive loss function no longer decreases, the model generated during the training phase is retained and used for semantic segmentation of multimodal remote sensing images.

[0072] Step 105: Input the fused feature map set into the decoder to obtain the segmentation feature map.

[0073] In some alternative embodiments, inputting the fused feature map set into the decoder to obtain the segmentation feature map may include the following steps:

[0074] The first step is to upsample the fourth-level fused feature map by a factor of two, resulting in a double-sampled fourth fused map. In practice, this double upsampling can be achieved using transposed convolution or bilinear interpolation.

[0075] The second step involves fusing the aforementioned third-level fused feature map with the aforementioned 2x sampled fourth-level fused feature map to obtain the decoded third-level fused feature map. This decoded third-level fused feature map can be a fused feature map with a higher resolution than the aforementioned fourth-level fused feature map. In practice, this fusion can be achieved through concatenation or element-wise addition.

[0076] The third step involves upsampling the decoded third-level fused feature map by a factor of two to obtain a double-sampled third-level fused map. In practice, the method for implementing this double upsampling can be the same as the first step described above.

[0077] The fourth step involves fusing the second-level fused feature map and the double-sampled third-level fused feature map to obtain the decoded second-level fused feature map. This decoded second-level fused feature map can have a higher resolution than the decoded third-level fused feature map, incorporating more and finer spatial details. In practice, the fusion process can be implemented in the same way as the second step.

[0078] The fifth step is to upsample the decoded second-level fused feature map by a factor of two to obtain the double-sampled second-level fused map. In practice, the double-upsampling method can be the same as the first step described above.

[0079] The sixth step involves fusing the first-level fused feature map and the second-level fused map (sampled twice) to obtain the final decoded first-level fused feature map. This decoded first-level fused feature map can have a higher resolution than the decoded second-level fused feature map, incorporating the richest low-level details such as edges and textures. In practice, the fusion process can be implemented in the same way as in the second step.

[0080] Step 7: Input the first-level fused feature map of the final decoded layer into the convolutional layer to obtain the segmentation score map. In practice, the c-dimensional feature vector of each pixel can be mapped to an m-dimensional vector, where c and m are constants without specific limitations. The m mentioned above can be the total number of semantic segmentation categories. For example, in a segmentation task, if the categories are "background (0)", "building (1)", "water area (2)", and "road (3)", then m = 4. Then a score map or probability map is obtained, which is the segmentation score map mentioned above.

[0081] Step 8: Classify each pixel based on the segmentation score map to obtain the segmentation feature map. In practice, you can first apply the Argmax function to the m-dimensional vector at each pixel position (h,w) along the channel dimension m of the segmentation score map of size [m,h,w]. (The Argmax function finds the index of the largest element in this m-dimensional vector, and this index (0,1,2, or3) directly corresponds to the predicted class label of the pixel). Then, after performing the Argmax operation on all h x w pixels, you can obtain a matrix of size [h,w]. (Each value in this matrix is ​​an integer representing the predicted class of the pixel. This [h,w] matrix can be the final semantic segmentation map. It can be visualized as a pseudo-color image, where different colors represent different land cover categories).

[0082] It is understood that, in addition to the above, this invention also includes some conventional structures and methods, which are well known and will not be elaborated upon further. However, this does not mean that these structures and methods are absent from this invention.

[0083] Those skilled in the art will recognize that, although numerous exemplary embodiments of the invention have been shown and described in detail herein, many other variations or modifications conforming to the principles of the invention can be directly determined or derived from the disclosure of the invention without departing from its spirit and scope. Therefore, the scope of the invention should be understood and construed as covering all such other variations or modifications.

[0084] In this invention, an asymmetric semantic alignment network is designed. When designing the network, the asymmetry between modes is fully considered, and alignment is performed on the feature space and channels, so that it can adaptively calibrate and align the complementary information between the two modes, reduce the interference of noise information, and improve the performance of semantic segmentation of remote sensing images.

[0085] Those skilled in the art will understand that various aspects of this application can be implemented as a system, method, or program product. Therefore, various aspects of this application can be specifically implemented in the following forms: a completely hardware implementation, a completely software implementation (including firmware, microcode, etc.), or a combination of hardware and software implementations, collectively referred to herein as a "circuit," "module," or "system."

[0086] In embodiments of this disclosure, a computer-readable storage medium is also provided, on which computer program instructions are stored. When executed, for example, by a processor, these computer program instructions can implement the steps of the asymmetric semantic alignment multimodal remote sensing image semantic segmentation method described in any of the above embodiments. In some possible implementations, various aspects of this application can also be implemented as a program product comprising program code. When the program product is run on a terminal device, the program code causes the terminal device to perform the steps described in the asymmetric semantic alignment multimodal remote sensing image semantic segmentation method according to various exemplary embodiments of this application.

[0087] The program product for implementing the above-described method according to embodiments of this application may employ a portable compact disc read-only memory (CD-ROM) and include program code, and may run on a terminal device, such as a personal computer. However, the program product of this application is not limited thereto. In this document, a readable storage medium may be any tangible medium that contains or stores a program that may be used by or in conjunction with an instruction execution system, apparatus, or device.

[0088] The above-described program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0089] The aforementioned computer-readable storage medium may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The readable storage medium may also be any readable medium other than a readable storage medium, capable of transmitting, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.

[0090] Program code for performing the operations of this application can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java and C++, and conventional procedural programming languages ​​such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0091] Those skilled in the art will understand that various aspects of this application can be implemented as a system, method, or program product. Therefore, various aspects of this application can be specifically implemented in the following forms: a completely hardware implementation, a completely software implementation (including firmware, microcode, etc.), or a combination of hardware and software implementations, collectively referred to herein as a "circuit," "module," or "system."

[0092] The above description is merely a preferred embodiment of this disclosure and is not intended to limit the scope of this disclosure. Various modifications and variations can be made to this disclosure by those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A multimodal remote sensing image semantic segmentation method with asymmetric semantic alignment, characterized in that, The method includes: Acquire registered optical images and registered SAR images of the same area; Based on the registered optical image and the registered SAR image, a set of ground feature category mask images is generated; The registered optical image, the registered SAR image, and the land cover category mask image set are preprocessed to obtain the preprocessed optical image, the preprocessed SAR image, and the preprocessed land cover category mask image set. The preprocessed optical image and the preprocessed SAR image are encoded to obtain a fused feature map set; The fused feature map set is input into the decoder to obtain the segmentation feature map.

2. The multimodal remote sensing image semantic segmentation method with asymmetric semantic alignment according to claim 1, wherein, The process of encoding the preprocessed optical image and the preprocessed SAR image to obtain a fused feature map set includes: The preprocessed optical image, the preprocessed SAR image, and the preprocessed land cover category mask image set are input into a feature encoder to obtain the fused feature map set. The feature encoder includes a predetermined number of feature extraction layers, all of which have a symmetrical structure and include a convolutional layer module, a feature attention module, and a semantic fusion module. The fused feature map set includes a predetermined number of fused feature maps, and the preprocessed land cover category mask image set consists of labels.

3. The method according to claim 2, wherein, The step of inputting the preprocessed optical image, the preprocessed SAR image, and the preprocessed land cover category mask image set into the feature encoder to obtain the fused feature map set includes: The preprocessed optical image and the preprocessed SAR image are respectively input into the first convolutional layer module to obtain the initial feature map of the optical branch output and the initial feature map of the SAR branch output; The initial feature map output by the optical branch and the initial feature map output by the SAR branch are input into the first feature interest module to obtain the first weighted optical feature map and the first weighted SAR feature map; The first weighted optical feature map and the first weighted SAR feature map are input into the first semantic fusion module to obtain the first-level fused feature map. The initial feature map output by the optical branch and the initial feature map output by the SAR branch are respectively input into the second convolution module to obtain the second optical feature map and the second SAR feature map. The second optical feature map and the second SAR feature map are input into the second feature interest module to obtain the second weighted optical feature map and the second weighted SAR feature map. The second weighted optical feature map and the second weighted SAR feature map are input into the second semantic fusion module to obtain the second-level fused feature map; The second weighted optical feature map and the second weighted SAR feature map are respectively input into the third convolution module to obtain the third optical feature map and the third SAR feature map; The third optical feature map and the third SAR feature map are input into the third feature attention module to obtain the third weighted optical feature map and the third weighted SAR feature map; The third weighted optical feature map and the third weighted SAR feature map are input into the third semantic fusion module to obtain the third-level fused feature map. The third optical feature map and the third SAR feature map are respectively input into the fourth convolution module to obtain the fourth optical feature map and the fourth SAR feature map; The fourth optical feature map and the fourth SAR feature map are input into the fourth feature interest module to obtain the fourth weighted optical feature map and the fourth weighted SAR feature map; The fourth weighted optical feature map and the fourth weighted SAR feature map are input into the fourth semantic fusion module to obtain the fourth-level fused feature map. The first-level fusion feature map, the second-level fusion feature map, the third-level fusion feature map, and the fourth-level fusion feature map are saved, wherein the resolution of the fusion feature map at each level decreases layer by layer, and the semantic information increases layer by layer.

4. The method according to claim 3, wherein, The first feature attention module, the second feature attention module, the third feature attention module, and the fourth feature attention module have the same structure and function, and are respectively connected to the first convolution module, the second convolution module, the third convolution module, and the fourth convolution module; the first feature attention module, the second feature attention module, the third feature attention module, and the fourth feature attention module respectively perform the following steps on the initial feature map output by the optical branch and the initial feature map output by the SAR branch, the second optical feature map and the second SAR feature map, the third optical feature map and the third SAR feature map, and the fourth optical feature map and the fourth SAR feature map: Based on the received SAR feature map and optical feature map, a difference feature map combining the SAR feature map and optical feature map is generated; The difference feature map is input into the pooling layer to generate initial weight values; The initial weight values ​​are input into the channel perceptron to obtain the global difference weights; The received SAR feature map and optical feature map are weighted and corrected according to the global difference weight to obtain the weighted optical feature map and the weighted SAR feature map.

5. The method according to claim 3, wherein, The first, second, third, and fourth semantic fusion modules have the same results and functions, and are respectively connected to the first, second, third, and fourth feature attention modules. The first, second, third, and fourth semantic fusion modules perform the following steps on the first weighted optical feature map and the first weighted SAR feature map, the second weighted optical feature map and the second weighted SAR feature map, the third weighted optical feature map and the third weighted SAR feature map, and the fourth weighted optical feature map and the fourth weighted SAR feature map: The received weighted optical feature map and weighted SAR feature map are aligned and calibrated to obtain the preliminary processed optical feature map and the preliminary processed SAR feature map. Spatial alignment and calibration are performed on the pre-processed optical feature map and the pre-processed SAR feature map to obtain calibrated optical feature map and calibrated SAR feature map; The calibrated optical feature map and the calibrated SAR feature map are added pixel-by-pixel to obtain a fused feature map.

6. The method according to claim 3, wherein, The step of inputting the fused feature map set into the decoder to obtain the segmentation feature map includes: The fourth-level fusion feature map is upsampled by a factor of two to obtain a fourth fusion map with a factor of two. The third-level fused feature map is fused with the fourth fused map that has been doubled in sampling to obtain the decoded third-level fused feature map; The decoded third-level fused feature map is upsampled by a factor of two to obtain a third-level fused map that has been upsampled by a factor of two. The second-level fused feature map and the double-sampled third-level fused map are fused to obtain the decoded second-level fused feature map; The decoded second-level fused feature map is upsampled by a factor of two to obtain a second-level fused map that has been upsampled by a factor of two. The first-level fused feature map and the double-sampled second-level fused map are fused to obtain the final decoded first-level fused feature map; The first-level fused feature map of the final decoded layer is input into the convolutional layer to obtain the segmentation score map; Each pixel is classified according to the segmentation score map to obtain the segmentation feature map.

7. The method according to claim 2, wherein, The feature encoder and the decoder constitute a semantic segmentation model for multimodal remote sensing images. During the training phase, the semantic segmentation model for multimodal remote sensing images compares the feature map output by the decoder, which has the same pixel size as the original image, with the ground truth using a multi-class cross-entropy loss function.

8. A server, characterized in that, The server runs the remote sensing image change detection method as described in claims 1-7.

9. A computer-readable storage medium having computer program instructions stored thereon, wherein, When executed by a processor, the computer program instructions cause the processor to perform the method according to any one of claims 1 to 7.

10. An electronic device, comprising: processor; as well as A memory that stores instructions executable by the processor; The processor is configured to execute the instructions to implement the method according to any one of claims 1 to 7.