Remote sensing image semantic change detection method based on graph incentive prompt joint learning

Through the method of graph-inspired prompt joint learning, the main-auxiliary network joint learning and multi-level prompt strategy are utilized to optimize prediction uncertainty, improve the accuracy and precision of semantic change detection in remote sensing images, and solve the shortcomings of changed area identification.

CN120635721APending Publication Date: 2025-09-12FUZHOU UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510958935.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing multi-task semantic change detection models fail to effectively reason and make decisions about changed areas in remote sensing images, affecting the accuracy of semantic change recognition.

Method used

A joint learning method of graph excitation and prompting is adopted. Through joint learning of main and auxiliary networks, a multi-level prompting strategy and a dual-phase graph excitation module are combined to optimize prediction uncertainty and construct a multi-task decoder to improve the recognition accuracy of changed areas.

Benefits of technology

It improves the accuracy and precision of semantic change detection, improves the reasoning and decision-making performance of changed areas, and provides important technical support for government management departments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635721A_ABST
    Figure CN120635721A_ABST
Patent Text Reader

Abstract

The invention relates to a remote sensing image semantic change detection method based on graph excitation prompt joint learning, and belongs to the field of dual-time-phase remote sensing image land utilization / land coverage semantic change detection research. According to the method, an image deep learning technology is combined to construct a multi-class earth surface semantic change detection model for remote sensing image land utilization / land coverage (LULC), and two groups of main-auxiliary twin neural networks with non-shared weights are adopted for double-temporal remote sensing images obtained in different periods to extract deep features of the double-temporal images respectively; reasoning output and joint optimization of an LULC change area and a semantic change type are realized in cooperation with a double-time-phase diagram excitation module, a multi-level feature prompt method and a multi-task decoder, and prediction probability distribution of the change area is improved in a prediction uncertainty optimization mode. According to the invention, on the basis of a conventional multi-task semantic change detection model, through optimizing the detection performance of the change area, the dual-temporal high remote sensing image semantic change detection with higher precision is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a remote sensing image semantic change detection method based on graph-inspired prompting joint learning. Background Art

[0002] Obtaining information on semantic changes in land use / land cover (LULC) between remote sensing images from different time periods is a crucial process for surface observation and analysis of geographic phenomena, prompting the exploration of semantic change detection methods. SCD is a specialized and valuable field of multi-class change detection, aiming to accurately identify and classify the semantic attributes, spatial location, and extent of transformed features across different image periods. This research explores the underlying spatiotemporal semantic transformations, providing powerful decision-making support for urban planning, environmental monitoring, and resource conservation.

[0003] Semantic change detection can simultaneously identify regions of change and discriminate the types of features within the region in the preceding and following imagery. Currently, semantic change detection models based on multi-task learning separate semantic change detection into semantic segmentation and binary change detection (i.e., region of change detection). These models use the corresponding LULC classification reference and the region of change reference from the two imagery phases to perform parallel training and optimization of the twin networks. However, during the model's inference process, the accuracy of semantic change recognition is often directly affected by the recognition performance of the region of change. The inability to effectively reason about and constrain the decision-making of the region of change to improve the accuracy of semantic change detection is a major challenge faced by existing multi-task methods.

[0004] In existing scientific research and production applications, multi-network joint learning has been successfully applied in different pattern recognition fields, especially in visual tasks. Its basic idea is to use different feature extraction tasks or different feature observation perspectives to compensate for each other and perform parallel reasoning on the entire model, thereby improving the performance of model learning. At present, joint learning has also been proven to be significantly effective in applications such as remote sensing semantic segmentation and target detection. However, its application in the field of change detection is relatively small. Based on this, the present invention applies joint learning to semantic change detection of LULC and constructs a remote sensing image semantic change detection method using graph-inspired joint learning. Summary of the Invention

[0005] The purpose of the present invention is to provide a remote sensing image semantic change detection method based on graph-inspired prompting joint learning. It flexibly utilizes the main-auxiliary network joint learning, multi-level prompting strategy, dual-phase graph-inspired modeling and prediction uncertainty optimization to overcome the constraints of traditional multi-task semantic change detection models that fail to effectively reason and make decisions on the changed areas, thereby improving the accuracy of semantic change recognition results.

[0006] To achieve the above-mentioned purpose, the technical solution of the present invention is: a remote sensing image semantic change detection method with graph excitation and prompting joint learning, combining image deep learning technology to construct a multi-category surface semantic change detection model for remote sensing image land use / land cover LULC, and using two sets of main-auxiliary twin neural networks with non-shared weights to extract the deep features of the dual-phase remote sensing images acquired in different periods, and coordinating with the dual-phase graph excitation module, multi-level feature prompting method and multi-task decoder to realize the inference output and joint optimization of LULC change areas and semantic change types, and improve the predicted probability distribution of the change area by optimizing the prediction uncertainty.

[0007] Furthermore, the method comprises the following steps:

[0008] Step S1, obtaining remote sensing images corresponding to two different periods of the same surface area, and preprocessing the images, including radiation correction, orthorectification, atmospheric correction, image fusion, image registration and image resampling operations;

[0009] Step S2: performing tile cropping and random data enhancement on the pre-processed images, and constructing a high-resolution remote sensing image semantic change detection sample set for multivariate patterns through manual screening and correction;

[0010] Step S3: Based on the semantic change detection sample set constructed in step S2, the corresponding bi-temporal image sample pairs in the sample set are input into the main network bi-temporal feature extractor and the auxiliary network bi-temporal feature extractor respectively for bi-temporal image high-level feature extraction;

[0011] Step S4: construct a bi-temporal image excitation module to perform same-dimensional coupling on the bi-temporal image high-level features obtained by the auxiliary network bi-temporal feature extractor in step S3; and use a multi-level feature prompting method to cascade the coupled features on the bi-temporal image high-level features obtained by the main network bi-temporal feature extractor to achieve prompt mapping of the auxiliary network to the main network.

[0012] Step S5: Fusing the high-level features of the bi-temporal images obtained by the main network bi-temporal feature extractor in step S3, and constructing a main network multi-task feature decoder based on the fused features: one task for detecting the changed area, and two tasks for identifying the corresponding ground feature types in the previous and next temporal images within the changed area;

[0013] Step S6: constructing an auxiliary network multi-task feature decoder for the dual-temporal image high-level features obtained by the auxiliary network dual-temporal feature extractor in step S3: using two dual-task feature decoders with shared parameter weights to detect non-changing areas of the dual-temporal image high-level features, and simultaneously fusing the dual-temporal image high-level features, and using a single-task feature decoder to independently detect changing areas of the fused features;

[0014] Step S7: Optimize the prediction uncertainty of multiple change area outputs of the main network and the auxiliary network in step S5 and step S6; perform normalized reverse transposition on the prediction probability of the non-change area obtained by the auxiliary network multi-task feature decoder, and perform accuracy calculation on the transposed probability of the auxiliary network in the non-change area prediction, the independent change area prediction probability obtained by the auxiliary network single-task feature decoder, and the independent change area prediction probability obtained by the main network single-task feature decoder respectively with the real sample labels, obtain their respective proportion weights, assign corresponding prediction probabilities respectively, and calculate the threshold distribution result of the total to obtain the final change area output result;

[0015] Step S8: perform spatial mask superposition on the changed area output result obtained in step S7 and the bi-temporal feature type recognition output result of the main network in step S5 to obtain the semantic change detection result corresponding to the final bi-temporal image;

[0016] Step S9: input the training data of the semantic change detection sample set into the model for training, and use the trained model to complete the recognition of the semantic change type.

[0017] The present invention also provides a computer-readable storage medium on which computer program instructions that can be executed by a processor are stored. When the processor executes the computer program instructions, the steps of any of the above methods can be implemented.

[0018] Compared with the existing technology, the present invention has the following beneficial effects: (1) Through the design of a model for joint learning of the main and auxiliary networks, the synchronous learning and optimization of non-change, binary change and semantic change are performed, which prompts the overall performance of the model in recognizing semantic changes; (2) By utilizing a multi-level prompting strategy, a dual-phase image excitation module and prediction uncertainty optimization, the reasoning and decision-making performance of the model in the change area is improved. By improving the accuracy of changing area identification, the accuracy of LULC semantic change detection in dual-phase remote sensing images is improved, providing important methodological and technical support for government management departments. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 Schematic diagram of a method flow in an embodiment of the present invention.

[0020] Figure 2Schematic diagram of the structure of a semantic change detection model for graph-inspired prompting joint learning according to an embodiment of the present invention.

[0021] Figure 3 Schematic diagram of the dual-phase diagram excitation module structure of an embodiment of the present invention.

[0022] Figure 4 This is a diagram showing the semantic change detection results of part of the Hi-UCD dataset according to an embodiment of the present invention. DETAILED DESCRIPTION

[0023] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings.

[0024] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present application belongs.

[0025] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.

[0026] The present invention provides a remote sensing image semantic change detection method based on graph-inspired prompting joint learning. A multi-category surface semantic change detection model for land use / land cover LULC in remote sensing images is constructed by combining image deep learning technology. Two sets of main-auxiliary twin neural networks with non-shared weights are used to extract deep features of dual-temporal remote sensing images acquired at different periods. The inference output and joint optimization of LULC change areas and semantic change types are achieved by combining a dual-temporal graph excitation module, a multi-level feature prompting method and a multi-task decoder. The predicted probability distribution of the change area is improved by optimizing the prediction uncertainty. Figure 1 As shown, the method comprises the following steps:

[0027] Step S1, obtaining remote sensing images corresponding to two different periods of the same surface area, and preprocessing the images, including radiation correction, orthorectification, atmospheric correction, image fusion, image registration and image resampling operations;

[0028] Step S2: performing tile cropping and random data enhancement on the pre-processed images, and constructing a high-resolution remote sensing image semantic change detection sample set for multivariate patterns through manual screening and correction;

[0029] Step S3: Based on the semantic change detection sample set constructed in step S2, the corresponding bi-temporal image sample pairs in the sample set are input into the main network bi-temporal feature extractor and the auxiliary network bi-temporal feature extractor respectively for bi-temporal image high-level feature extraction;

[0030] Step S4: construct a bi-temporal image excitation module to perform same-dimensional coupling on the bi-temporal image high-level features obtained by the auxiliary network bi-temporal feature extractor in step S3; and use a multi-level feature prompting method to cascade the coupled features on the bi-temporal image high-level features obtained by the main network bi-temporal feature extractor to achieve prompt mapping of the auxiliary network to the main network.

[0031] Step S5: Fusing the high-level features of the bi-temporal images obtained by the main network bi-temporal feature extractor in step S3, and constructing a main network multi-task feature decoder based on the fused features: one task for detecting the changed area, and two tasks for identifying the corresponding ground feature types in the previous and next temporal images within the changed area;

[0032] Step S6: constructing an auxiliary network multi-task feature decoder for the dual-temporal image high-level features obtained by the auxiliary network dual-temporal feature extractor in step S3: using two dual-task feature decoders with shared parameter weights to detect non-changing areas of the dual-temporal image high-level features, and simultaneously fusing the dual-temporal image high-level features, and using a single-task feature decoder to independently detect changing areas of the fused features;

[0033] Step S7: Optimize the prediction uncertainty of multiple change area outputs of the main network and the auxiliary network in step S5 and step S6; perform normalized reverse transposition on the prediction probability of the non-change area obtained by the auxiliary network multi-task feature decoder, and perform accuracy calculation on the transposed probability of the auxiliary network in the non-change area prediction, the independent change area prediction probability obtained by the auxiliary network single-task feature decoder, and the independent change area prediction probability obtained by the main network single-task feature decoder respectively with the real sample labels, obtain their respective proportion weights, assign corresponding prediction probabilities respectively, and calculate the threshold distribution result of the total to obtain the final change area output result;

[0034] Step S8: perform spatial mask superposition on the changed area output result obtained in step S7 and the bi-temporal feature type recognition output result of the main network in step S5 to obtain the semantic change detection result corresponding to the final bi-temporal image;

[0035] Step S9: Use the training data of the semantic change detection sample set to input the model for training; wherein, the AdamW optimizer, cross entropy loss function and binary cross entropy loss function are used to perform back propagation optimization of the model parameters; use the trained model to complete the identification of semantic change types.

[0036] In this embodiment, step S3 is specifically implemented as follows:

[0037] Two twin neural networks with the same structure but no shared parameter weights are constructed as the main network and auxiliary network. The feature encoders of the main network and the auxiliary network use ResNet34 for feature extraction, which specifically includes: (1) a single convolution layer with a convolution kernel size of 7×7 and a maximum pooling layer with a convolution kernel size of 3×3 are used to extract the first layer features of the initial image; (2) residual blocks are used as the basic module to extract the features of the second, third, fourth and fifth layers, and the residual block combination of each layer is 3, 4, 6 and 3 respectively; each residual block contains two convolution layers (CNL) with a convolution kernel size of 3×3, batch normalization (BN), activated linear unit (ReLU) and a skip connection layer of CNL and BN combination. In the process of inputting data into each feature layer, the spatial size is gradually reduced by 1 / 2, 1 / 4 and 1 / 8, and the channel dimension is gradually increased by 64, 128, 256 and 512.

[0038] In this embodiment, step S4 is specifically implemented as follows:

[0039] Step S41: In the dual-phase image excitation module, the dual-phase features and Graph similarity modeling mapping and fusion incentive mapping are performed simultaneously; in the graph similarity modeling mapping part, and Different fusion features are obtained by three metric calculation methods: spatial pixel superposition, spatial pixel dot product and spatial pixel difference, and the fusion features are compared with and Construct feature H by channel concatenation (Concat) A , graph structured modeling is performed through graph convolutional layer (GCL), and maximum pooling (Maxpool), multi-layer perceptron (MLP) and normalization (Sigmoid) are used to automatically assign weights to different fusion features. A , obtain graph similarity modeling feature G A :

[0040]

[0041] W A =Sigmoid(MLP(Maxpool(H A )))

[0042] G A =W A *GCL(H A )

[0043] Step S42: In the fusion excitation mapping part, the dual-phase feature and After the mapping of CNL, the channels are cascaded (Concat), and the compression-excitation SE module and CNL are used to obtain the fusion excitation mapping feature S A :

[0044]

[0045] Step S43: Model the graph similarity feature G A and fusion excitation mapping feature S A Perform channel cascade and obtain the output Y of the final dual-phase image excitation module through a single residual block A :

[0046]

[0047] In this embodiment, step S5 is specifically implemented as follows:

[0048] A multi-task decoder with a multi-branch upsampling layer is constructed as the main network. This decoder includes a branch for detecting regions with two temporal changes, a branch for recognizing semantic types of ground surfaces with pre-temporal changes, and a branch for recognizing semantic types of ground surfaces with post-temporal changes. Each of these three branches uses an independent decoder with the same structural structure: three deconvolutional layers (TCNL) with a 3×3 kernel size and three upsampling layers (Upsample) using bilinear interpolation.

[0049] In this embodiment, step S6 is specifically implemented as follows:

[0050] A multi-task decoder with a multi-branch upsampling layer is constructed as an auxiliary network. This decoder includes a branch for detecting regions of change between two temporal phases and two branches for detecting regions of non-change between the preceding and following phases. The branch for detecting regions of non-change between the preceding and following phases uses two identical decoders with shared weights to identify regions of non-change. The branch for detecting regions of change uses an independent decoder for identifying regions of change. Each decoder is composed of the same three TCNLs and three upsamples.

[0051] In this embodiment, step S7 is specifically implemented as follows:

[0052] Step S71: Output the decoder result of the double-branch non-changing area detection in the auxiliary network and Superimpose and calculate the corresponding change area prediction probability distribution by normalization transposition At the same time, the independent change region detection decoder output in the auxiliary network and the independent change region detection decoder output in the main network Calculate the predicted probability distribution separately and

[0053]

[0054]

[0055]

[0056] Step S72: Calculate the F1 score and mIoU accuracy of the three predicted probability distributions in step S71 and the actual change area labels to obtain their respective F1 evaluation scores. and And their respective mIoU evaluation scores and Sum the total scores based on the corresponding assessment scores and get the ratio of each score to the total score value and

[0057]

[0058]

[0059]

[0060] Step S73: The ratio of step S72 and As the weight given to each original probability, the final change area prediction result P is obtained by summing C :

[0061]

[0062] In this embodiment, step S8 is specifically implemented as follows:

[0063] Step S81: The change area prediction result P obtained in step S73 is C The decoder of the main network's pre-phase semantic change surface semantic type recognition obtains the pre-phase semantic change type recognition The predicted probability is optimized in the form of vector dot product masking to achieve the final front-phase semantic change detection results:

[0064]

[0065] Step S82: The change area prediction result P obtained in step S73 is C The post-phase semantic change type recognition obtained by the decoder of the post-phase semantic change surface semantic type recognition of the main network The predicted probability is used as a vector dot product mask to optimize the final temporal semantic change detection results:

[0066]

[0067] This example uses the preprocessed and cropped Hi-UCD semantic change detection dataset with a spatial resolution of 0.1m and red, green, and blue bands. This example uses 2,594 pairs of tile images with a size of 512×512 after tile cropping and corresponding semantic change labels for model training and validation.

[0068] like Figure 2 This is a structural diagram of the semantic change detection model for graph-excitation-prompt joint learning in this embodiment. The model consists of two twin neural networks with the same structure but unshared parameter weights. The two networks are defined as the main network and the auxiliary network respectively. The feature encoders of the main network and the auxiliary network use ResNet34 for feature extraction. Each feature coding layer of the encoder of the auxiliary network uses a dual-phase graph excitation module for feature coupling, and is mapped to the feature coding layer of the corresponding dimension of the main network in a multi-level prompting manner. The main network uses a multi-task decoder to identify dual-phase semantic change types and detect change areas, and the auxiliary network uses a multi-task decoder to identify dual-phase non-change areas and detect change areas. The final change area detection result is obtained by predictive uncertainty optimization, and the final semantic change detection result is obtained by mask optimization.

[0069] like Figure 3 This is a schematic diagram of the dual-phase image excitation module structure used in this embodiment. The module performs graph similarity modeling mapping and fusion excitation mapping on the dual-phase features. In the graph similarity modeling mapping part, different fusion features are obtained through spatial pixel superposition, spatial pixel dot product and spatial pixel difference, and the fusion features are cascaded with the dual-phase feature channels and the graph convolution layer for graph structured modeling. The weights of different fusion features are automatically assigned to obtain graph similarity modeling features; in the fusion excitation mapping part, the dual-phase features are cascaded through the channels after being mapped by the convolution layer, and the fusion excitation mapping features are obtained by using the compression-excitation module and the convolution layer. The graph similarity modeling features and the fusion excitation mapping features are cascaded through the channels, and the final dual-phase image excitation module output is obtained through a single residual block.

[0070] like Figure 4Figure 1 shows some experimental results from this embodiment on a processed Hi-UCD semantic change detection dataset. As can be seen, the semantic change detection predictions obtained using the method of the present invention are highly consistent with the true labels. The accuracy and refinement of the change shapes in detecting the change regions between the previous and next image phases are both high, with fewer missed and false detections. Furthermore, the identification of object types within the change regions between the previous and next image phases is relatively accurate. These experimental results demonstrate the effectiveness of this method for semantic change detection in remote sensing imagery.

[0071] Compared with existing methods, the present invention improves the model's overall performance in identifying semantic changes through the design of a joint learning model of primary and secondary networks, synchronously learning and optimizing non-changes, binary changes, and semantic changes. It also utilizes a multi-level prompting strategy, a dual-temporal image excitation module, and prediction uncertainty optimization to improve the model's reasoning and decision-making performance in changed areas. By improving the accuracy of changed area identification, the accuracy of LULC semantic change detection in dual-temporal remote sensing images is enhanced, providing important methodological and technical support for government management departments.

[0072] The present invention also provides a computer-readable storage medium on which computer program instructions that can be executed by a processor are stored. When the processor executes the computer program instructions, the steps of any of the above methods can be implemented.

[0073] The above are preferred embodiments of the present invention. Any changes made according to the technical solution of the present invention, as long as the resulting functions and effects do not exceed the scope of the technical solution of the present invention, shall fall within the scope of protection of the present invention.

Claims

1. A remote sensing image semantic change detection method based on graph-inspired and -prompted joint learning, characterized by: Combining image deep learning technology, a multi-category surface semantic change detection model for land use / land cover LULC in remote sensing images is constructed. Two sets of main-auxiliary twin neural networks with non-shared weights are used to extract the deep features of the dual-temporal remote sensing images acquired at different periods. The dual-temporal image excitation module, multi-level feature prompting method and multi-task decoder are used to realize the inference output and joint optimization of LULC change areas and semantic change types, and the predicted probability distribution of the change area is improved by optimizing the prediction uncertainty.

2. The remote sensing image semantic change detection method based on graph-inspired and -prompted joint learning according to claim 1 is characterized in that: The steps include: Step S1, obtaining remote sensing images corresponding to two different periods of the same surface area, and preprocessing the images, including radiation correction, orthorectification, atmospheric correction, image fusion, image registration and image resampling operations; Step S2: performing tile cropping and random data enhancement on the pre-processed images, and constructing a high-resolution remote sensing image semantic change detection sample set for multivariate patterns through manual screening and correction; Step S3: Based on the semantic change detection sample set constructed in step S2, the corresponding bi-temporal image sample pairs in the sample set are input into the main network bi-temporal feature extractor and the auxiliary network bi-temporal feature extractor respectively for bi-temporal image high-level feature extraction; Step S4: construct a bi-temporal image excitation module to perform same-dimensional coupling on the bi-temporal image high-level features obtained by the auxiliary network bi-temporal feature extractor in step S3; and use a multi-level feature prompting method to cascade the coupled features on the bi-temporal image high-level features obtained by the main network bi-temporal feature extractor to achieve prompt mapping of the auxiliary network to the main network. Step S5: Fusing the high-level features of the bi-temporal images obtained by the main network bi-temporal feature extractor in step S3, and constructing a main network multi-task feature decoder based on the fused features: one task for detecting the changed area, and two tasks for identifying the corresponding ground feature types in the previous and next temporal images within the changed area; Step S6: constructing an auxiliary network multi-task feature decoder for the dual-temporal image high-level features obtained by the auxiliary network dual-temporal feature extractor in step S3: using two dual-task feature decoders with shared parameter weights to detect non-changing areas of the dual-temporal image high-level features, and simultaneously fusing the dual-temporal image high-level features, and using a single-task feature decoder to independently detect changing areas of the fused features; Step S7: Optimize the prediction uncertainty of multiple change area outputs of the main network and the auxiliary network in step S5 and step S6; perform normalized reverse transposition on the prediction probability of the non-change area obtained by the auxiliary network multi-task feature decoder, and perform accuracy calculation on the transposed probability of the auxiliary network in the non-change area prediction, the independent change area prediction probability obtained by the auxiliary network single-task feature decoder, and the independent change area prediction probability obtained by the main network single-task feature decoder respectively with the real sample labels, obtain their respective proportion weights, assign corresponding prediction probabilities respectively, and calculate the threshold distribution result of the total to obtain the final change area output result; Step S8: perform spatial mask superposition on the changed area output result obtained in step S7 and the bi-temporal feature type recognition output result of the main network in step S5 to obtain the semantic change detection result corresponding to the final bi-temporal image; Step S9: input the training data of the semantic change detection sample set into the model for training, and use the trained model to complete the recognition of the semantic change type.

3. The remote sensing image semantic change detection method based on graph-inspired and -prompted joint learning according to claim 2 is characterized in that: Step S3 is specifically implemented as follows: Step S31: construct two twin neural networks with the same structure but no shared parameter weights as the main network and the auxiliary network, wherein the feature encoders of the main network and the auxiliary network use ResNet34 for feature extraction; in the initial input part, a single convolution layer with a convolution kernel size of 7×7 and a maximum pooling layer with a convolution kernel size of 3×3 are used to extract the first layer features of the initial image; Step S32: Use the residual block as the basic module to perform feature extraction on the second, third, fourth and fifth layers, and the number of residual blocks in each layer is 3, 4, 6 and 3 respectively; wherein each residual block contains two 3×3 convolutional layers (CNL) with a convolution kernel size of , batch normalization (BN), an activated linear unit (ReLU) and a jump connection layer of a combination of CNL and BN. In the process of inputting the data into each feature layer, the spatial size is gradually reduced by 1 / 2, 1 / 4 and 1 / 8, and the channel dimension is gradually increased by 64, 128, 256 and 512.

4. The remote sensing image semantic change detection method based on graph-inspired and -prompted joint learning according to claim 2 is characterized in that: Step S4 is specifically implemented as follows: Step S41: In the dual-phase image excitation module, the dual-phase features and Graph similarity modeling mapping and fusion incentive mapping are performed simultaneously; in the graph similarity modeling mapping part, and Different fusion features are obtained by three metric calculation methods: spatial pixel superposition, spatial pixel dot product and spatial pixel difference, and the fusion features are compared with and Construct feature H by channel concatenation A , graph structured modeling is performed through the graph convolution layer GCL, and the maximum pooling Maxpool, multi-layer perceptron MLP and normalization processing Sigmoid are used to automatically assign weights to different fusion features W A , obtain graph similarity modeling feature G A : W A =Sigmoid(MLP(Maxpool(H A ))) G A =W A *GCL(H A ) Step S42: In the fusion excitation mapping part, the dual-phase feature and After the mapping of CNL, the channels are cascaded and Concat is performed respectively, and the fusion excitation mapping feature S is obtained by using the compression-excitation SE module and CNL. A : Step S43: Model the graph similarity feature G A and fusion excitation mapping feature S A Perform channel cascade and obtain the output Y of the final dual-phase image excitation module through a single residual block A :

5. The remote sensing image semantic change detection method based on graph-inspired and -prompted joint learning according to claim 2 is characterized in that: Step S5 is specifically implemented as follows: Step S51: construct a multi-branch independent upsampling decoding layer as the main network multi-task feature decoder; select three TCNLs with a convolution kernel size of 3×3 and three upsampling layers Upsample using bilinear interpolation as decoder branches for detecting change regions between dual-temporal images to interpret the change regions; Step S52: Select three convolution kernels with a size of 3×3TCNL and three Upsamples as decoder branches for identifying the semantic type of the previous temporal phase change surface to interpret the semantic type of the previous temporal phase change surface; Step S53: Select three convolution kernels with a size of 3×3TCNL and three Upsamples as decoder branches for post-phase-varying surface semantic type recognition to interpret the post-phase-varying surface semantic type.

6. The remote sensing image semantic change detection method based on graph-inspired and -prompted joint learning according to claim 5 is characterized in that: Step S6 is specifically implemented as follows: Step S61: construct a multi-branch independent upsampling decoding layer as an auxiliary network multi-task feature decoder; select three convolution kernels of size 3×3 TCNL and three Upsample as decoder branches for detecting the previous phase unchanged region to interpret the previous phase unchanged region; Step S62: Select three convolution kernels with a size of 3×3TCNL and three Upsamples as decoder branches for detecting the post-phase non-changing region to interpret the post-phase non-changing region; Step S63 : Select three convolution kernels of size 3×3 TCNL and three Upsamples as decoder branches for detecting the changed regions between the two-temporal images to interpret the changed regions between the two-temporal images.

7. The remote sensing image semantic change detection method based on graph-inspired and -prompted joint learning according to claim 6 is characterized in that: Step S7 is specifically implemented as follows: Step S71: Output the decoder result of the double-branch non-changing area detection in the auxiliary network and Superimpose and calculate the corresponding change area prediction probability distribution by normalization transposition At the same time, the independent change region detection decoder output in the auxiliary network and the independent change region detection decoder output in the main network Calculate the predicted probability distribution separately and Step S72: Calculate the F1 score and mIoU accuracy of the three predicted probability distributions in step S71 and the actual change area labels to obtain their respective F1 evaluation scores. and And their respective mIoU evaluation scores and Sum the total scores based on the corresponding assessment scores and get the ratio of each score to the total score value and Step S73: The ratio of step S72 and As the weight given to each original probability, the final change area prediction result P is obtained by summing C :

8. The remote sensing image semantic change detection method based on graph-inspired and -prompted joint learning according to claim 7 is characterized in that: Step S8 is specifically implemented as follows: Step S81: The change area prediction result P obtained in step S73 is C The decoder of the main network's pre-phase semantic change surface semantic type recognition obtains the pre-phase semantic change type recognition The predicted probability is optimized in the form of vector dot product masking to achieve the final front-phase semantic change detection results: Step S82: The change area prediction result P obtained in step S73 is C The post-phase semantic change type recognition obtained by the decoder of the post-phase semantic change surface semantic type recognition of the main network The predicted probability is used as a vector dot product mask to optimize the final temporal semantic change detection results:

9. The remote sensing image semantic change detection method based on graph-inspired and -prompted joint learning according to claim 2 is characterized in that: In step S9, the AdamW optimizer, the cross entropy loss function and the binary cross entropy loss function are used to perform back propagation optimization of the model parameters.

10. A computer-readable storage medium storing computer program instructions that can be executed by a processor, wherein when the processor executes the computer program instructions, the steps of the method according to any one of claims 1 to 9 can be implemented.

Citation Information

Cited By

  • Low-altitude remote sensing image-oriented ground feature fine-grained attribute extraction method

    CN121661531A