Dual-temporal remote sensing image semantic change detection method and system based on multiple networks

Through the multi-network-based dual-time phase remote sensing image semantic change detection method, combined with sliding window, joint semantic change detection network and dual-branch semantic segmentation network, the problem that traditional methods are difficult to capture abstract features and handle complex tasks is solved, and efficient and accurate change detection results are achieved.

CN120014442APending Publication Date: 2025-05-16POWERCHINA FUJIAN ELECTRIC POWER SURVEY & DESIGN INST CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510051413.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Traditional dual-time phase remote sensing image change detection methods are difficult to accurately capture abstract features in high-dimensional space, and have limited ability to handle complex tasks, making it difficult to deal with diverse areas to be detected.

Method used

The semantic change detection method of dual-time phase remote sensing image based on multi-network is adopted, and the data set is generated using sliding windows and random rotations, combined with the joint semantic change detection network and the dual-branch semantic segmentation network, and the final change detection result is generated through the cross-attention fusion module and the two-stage decision fusion strategy, and the network is optimized using auxiliary loss functions.

Benefits of technology

It significantly improves the diversity of data samples and the generalization ability of the model, thereby improving the comprehensiveness of feature extraction and fusion and the reliability of the detection results, ensuring the stability and accuracy of the detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014442A_ABST
    Figure CN120014442A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-network-based dual-time-phase remote sensing image semantic change detection method and system. The method comprises the following steps: acquiring a dual-time-phase remote sensing image and preprocessing the dual-time-phase remote sensing image; performing image cutting by using a sliding window, and performing random rotation to obtain a dual-time-phase remote sensing image data set; a multi-network joint learning strategy is used, a joint semantic change detection network is used as a main network, a double-branch semantic segmentation network is used as an auxiliary network, and a double-temporal change enhancement network is constructed; wherein the trans-attention fusion module is used for carrying out feature image fusion, a two-stage decision fusion strategy is used, output of a main network is restrained in a pixel-level mask form, and a final dual-temporal remote sensing image change detection result is generated; performing optimization training by using an auxiliary loss function to obtain a trained dual-time-phase change enhanced network; and obtaining a to-be-detected dual-time-phase remote sensing image, and inputting the trained dual-time-phase change enhancement network to obtain a detection result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and mainly to a method and system for detecting semantic changes in dual-temporal remote sensing images based on multiple networks. Background Art

[0002] Detecting the type of surface changes between different high-resolution remote sensing images is a key task in earth observation applications, such as identifying surface changes in disaster losses, land use or non-agricultural information of cultivated land, etc. As an extension of Binary Change Detection (BCD), Semantic Change Detection (SCD) is essentially a multi-element semantic segmentation task, that is, determining what content has changed based on the changed area.

[0003] However, when dealing with complex and ever-changing image data, traditional methods have exposed many shortcomings that are difficult to overcome. Traditional methods have always been limited by their own relatively rigid algorithmic logic and processing framework, making it difficult to accurately capture abstract features hidden in high-dimensional space. At the same time, as time goes on, images will be affected by time factors such as seasonal changes in light and dynamic evolution of objects, further increasing the complexity of the data. In such a complex situation, traditional methods are often unable to effectively sort out key abstract feature clues in high-dimensional space, resulting in extremely unstable performance in actual application scenarios. Faced with images from different sensors, it is difficult to adapt to different imaging styles and data deviations, and it is difficult to meet high-precision requirements.

[0004] For example, a Chinese invention patent with publication number "CN118212534A" discloses a "method and system for change detection of dual-phase remote sensing images", which specifically discloses "obtaining a dual-phase remote sensing image of the area to be detected; the dual-phase remote sensing image includes two remote sensing images of different phases; performing image correction on the dual-phase remote sensing image to generate a dual-phase remote sensing corrected image; calculating similar features of the two remote sensing corrected images of different phases, and differential features of the depth features of the two remote sensing corrected images of different phases, and performing feature splicing on the similar features and the differential features to form a fusion feature; performing binary classification on each pixel point corresponding to the fusion feature to obtain the change detection result of the area to be detected", but this method only relies on calculating similar features and differential features to splice the fusion features, and it is difficult to obtain high-dimensional and complex abstract features, which will miss key information and limit the accuracy of the change detection results; in addition, this method has limited ability to handle complex tasks and is difficult to cope with diverse situations in the area to be detected. Summary of the invention

[0005] In order to solve the above problems existing in the prior art, the present application provides a method and system for semantic change detection of dual-temporal remote sensing images based on multiple networks.

[0006] The technical solution of this application is as follows:

[0007] On the one hand, the present invention proposes a method for detecting semantic changes in dual-temporal remote sensing images based on multiple networks, the method comprising:

[0008] Acquire dual-temporal remote sensing images, and preprocess the dual-temporal remote sensing images; crop the preprocessed dual-temporal remote sensing images into preset pixel specifications using a sliding window, and randomly rotate the cropped dual-temporal remote sensing images to obtain a dual-temporal remote sensing image data set;

[0009] A dual-temporal change enhancement network is constructed by using a multi-network joint learning strategy, with a joint semantic change detection network as the main network and a dual-branch semantic segmentation network as the auxiliary network; wherein, a cross-attention fusion module is used to fuse the feature images extracted from the main network and the auxiliary network, and a two-stage decision fusion strategy is used to constrain the output of the main network in the form of a pixel-level mask to generate the final dual-temporal remote sensing image change detection result; the dual-temporal change enhancement network is optimized using an auxiliary loss function, and the dual-temporal remote sensing image data set is input into the dual-temporal change enhancement network for iterative training to obtain a trained dual-temporal change enhancement network;

[0010] A dual-temporal remote sensing image to be detected is obtained, and the dual-temporal remote sensing image to be detected is input into a trained dual-temporal change enhancement network to obtain a corresponding detection result of the dual-temporal remote sensing image to be detected.

[0011] Preferably, the preprocessing includes radiation calibration, atmospheric correction, orthorectification, image fusion, image vectorization and image format conversion, wherein:

[0012] The image fusion specifically comprises fusing the multispectral and panchromatic images using a fusion algorithm;

[0013] The image vectorization specifically involves vectorizing the boundaries of land types in the remote sensing image, outlining different land types in the image, including water bodies, forest land, cultivated land and buildings, and assigning values ​​to pixels of different land types;

[0014] The image format conversion is specifically to convert the format using a vector to raster algorithm.

[0015] Preferably, the joint semantic change detection network includes a semantic change detection encoder, a semantic change attention module and a semantic change detection decoder. The dual-temporal remote sensing image dataset is input into the joint semantic change detection network, which is expressed as:

[0016]

[0017] In the formula, The characteristic image of semantic change representing the previous tense; Indicates the semantic change feature image of the later phase; I T1 Indicates the previous phase image; I T2 Indicates post-phase image; Siam SCD () represents the semantic change detection encoder with twin feature extraction; Represents the concatenation result of the bi-temporal features in the channel dimension; Concat represents the concatenation function of the bi-temporal features in the channel cascade; Classifier1 represents the first semantic change classification decoding layer; Classifier2 represents the second semantic change classification decoding layer; Classifier3 represents the third semantic change classification decoding layer; Represents the output feature map of semantic changes of the previous phase image; Output feature map representing semantic changes of later-phase images; Represents the output feature map of joint semantic change detection.

[0018] Preferably, the dual-branch semantic segmentation network includes a semantic segmentation encoder, a semantic segmentation attention module and a semantic segmentation decoder. The dual-temporal remote sensing image dataset is input into the dual-branch semantic segmentation network, which is expressed as:

[0019]

[0020] In the formula, A semantically segmented feature image representing the front phase; Represents the semantic segmentation feature image of the posterior phase; Siam SS () represents a dual-branch semantic segmentation encoder with twin feature extraction; Represents the semantic output feature map of the previous phase image; Represents the semantic output feature map of the later phase image; Classifier4 represents the first semantic segmentation classification decoding layer.

[0021] Preferably, the cross-attention fusion module is used to fuse the feature images extracted from the main network and the auxiliary network, which can be expressed as:

[0022]

[0023] A F =A1×A2;

[0024] O=v1×v2+A F ×(v1×v2);

[0025] Wherein, A1 represents the feature map of the extracted semantic segmentation encoder; Softmax represents the activation function; q1 represents the spatial domain vector of the preset semantic segmentation encoder; k1 represents the spatial domain vector of the preset semantic segmentation encoder; T represents the transposition operation; A2 represents the feature map of the extracted semantic change detection encoder; q2 represents the spatial domain vector of the preset semantic change detection encoder; k2 represents the spatial domain vector of the preset semantic change detection encoder; A F Represents the fused feature map; × represents the staggered dot product; O represents the output map of the cross-attention fusion module; v1 represents the spatial domain vector of the preset cross-attention fusion module; v2 represents the spatial domain vector of the preset cross-attention fusion module.

[0026] Preferably, a two-stage decision fusion strategy is used to constrain the output of the main network in the form of a pixel-level mask to generate the final dual-temporal remote sensing image change detection result, specifically:

[0027] In the first stage, the binary change detection map of the dual-branch semantic segmentation is obtained by binary superposition of the semantic output feature map of the previous phase image and the semantic output feature map of the subsequent phase image at the spatial vector level; the dot product operation of the pixel-level mask is performed on the binary change detection map and the output map of the cross-attention fusion module to calculate the intersection map;

[0028] In the second stage, a dot product operation is performed on the intersection map and the output feature map of the joint semantic change detection to generate the final dual-temporal remote sensing image change detection result.

[0029] Preferably, the dual-phase change enhancement network is optimized using an auxiliary loss function, which is expressed as:

[0030]

[0031]

[0032] L joint =L BCD +L SCD +L SS +L Aux ;

[0033] In the formula, represents the pre-temporal loss for joint semantic change detection; represents the previous phase probability distribution of the reference label preset for the joint semantic change detection of the i-th category; represents the previous phase probability distribution of the joint semantic change detection prediction result of the i-th category; represents the posterior temporal loss for joint semantic change detection; represents the posterior temporal probability distribution of the reference label preset for joint semantic change detection of the i-th category; represents the posterior phase probability distribution of the joint semantic change detection prediction result of the i-th category; L BCD represents the loss of binary change detection map; Representing the probability distribution of the preset reference labels of the prediction results of joint semantic change detection; represents the probability distribution of the prediction results of joint semantic change detection; L SCD represents the total loss of joint semantic change detection; Represents the pre-temporal loss of a two-branch semantic segmentation network; Represents the previous time-phase probability distribution of the reference label preset by the dual-branch semantic segmentation network of the i-th category; Represents the previous phase probability distribution of the prediction result of the dual-branch semantic segmentation network of the i-th category; Represents the posterior temporal loss of a two-branch semantic segmentation network; Represents the posterior temporal probability distribution of the reference label preset by the dual-branch semantic segmentation network of the i-th category; represents the posterior phase probability distribution of the prediction result of the dual-branch semantic segmentation network of the i-th category; L SS represents the total loss of the two-branch semantic segmentation network; represents the cosine similarity for joint semantic change detection; Represents the cosine similarity of a two-branch semantic segmentation network; represents the cosine similarity of the binary change detection graph; L Aux represents the sum of cosine similarity; L joint Represents the final auxiliary loss function; N represents the total number of categories of the dual-temporal remote sensing image dataset; i represents the index value of the i-th category.

[0034] On the other hand, the present invention also proposes a dual-temporal remote sensing image semantic change detection system based on multiple networks, the system comprising a data acquisition module, a network construction module, an image detection module and a result output module, wherein:

[0035] The data acquisition module is used to acquire dual-phase remote sensing images and preprocess the dual-phase remote sensing images; use a sliding window to crop the preprocessed dual-phase remote sensing images into a preset pixel specification, and randomly rotate the cropped dual-phase remote sensing images to obtain a dual-phase remote sensing image data set; transmit the dual-phase remote sensing image data set to the network construction module;

[0036] The network construction module is used to use a multi-network joint learning strategy, with a joint semantic change detection network as the main network and a dual-branch semantic segmentation network as the auxiliary network to construct a dual-phase change enhancement network; wherein, a cross-attention fusion module is used to fuse the feature images extracted from the main network and the auxiliary network, and a two-stage decision fusion strategy is used to constrain the output of the main network in the form of a pixel-level mask to generate the final dual-phase remote sensing image change detection result; the dual-phase change enhancement network is optimized using an auxiliary loss function, and the dual-phase remote sensing image data set is input into the dual-phase change enhancement network for iterative training to obtain a trained dual-phase change enhancement network;

[0037] The image detection module is used to obtain a dual-phase remote sensing image to be detected, input the dual-phase remote sensing image to be detected into the trained dual-phase change enhancement network, and obtain a corresponding detection result of the dual-phase remote sensing image to be detected;

[0038] The result output module is used to display the corresponding detection results of the dual-phase remote sensing image to be detected.

[0039] On the other hand, the present invention also proposes an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the multi-network-based dual-temporal remote sensing image semantic change detection method as described in any embodiment of the present invention is implemented.

[0040] On the other hand, the present invention further proposes a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the multi-network-based dual-temporal remote sensing image semantic change detection method as described in any embodiment of the present invention.

[0041] Compared with the prior art, the present invention has the following beneficial effects:

[0042] 1) The present invention provides a method and system for detecting semantic changes in dual-temporal remote sensing images based on a multi-network. The preprocessed dual-temporal remote sensing images are cropped into preset pixel specifications using a sliding window, and the cropped images are randomly rotated to obtain a dual-temporal remote sensing image dataset, thereby enhancing the diversity of data samples, greatly improving data richness, and improving the generalization ability of the model.

[0043] 2) The present invention provides a dual-temporal remote sensing image semantic change detection method and system based on multiple networks. The method utilizes a multi-network joint learning strategy, takes a joint semantic change detection network as the main network, and a dual-branch semantic segmentation network as the auxiliary network to construct a dual-temporal change enhancement network, thereby enhancing the comprehensiveness of feature extraction and fusion. Compared with a single network architecture, the method improves the ability to mine deep semantic information from images. At the same time, the feature images extracted from the main network and the auxiliary network are fused with the help of a cross-attention fusion module, thereby further improving the quality and efficiency of feature fusion and the reliability of the final detection result.

[0044] 3) The present invention provides a dual-temporal remote sensing image semantic change detection method and system based on multiple networks, optimizes the dual-temporal change enhancement network by using an auxiliary loss function, and adopts a two-stage decision fusion strategy. As well as the introduction of an auxiliary loss function, the scientificity and precision of network training are improved, the rationality of parameter adjustment during iterative training of the model is enhanced, the overall performance of the network is significantly improved, and the stability and accuracy of the detection results are guaranteed. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 is a method flow chart of an embodiment of the present invention;

[0046] Figure 2 Schematic diagram of the structure of a dual-phase change enhancement network according to an embodiment of the present invention. DETAILED DESCRIPTION

[0047] The specific implementation modes of the present invention are described below to facilitate those skilled in the art to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific implementation modes. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the attached claims, these changes are obvious, and all inventions and creations utilizing the concept of the present invention are protected.

[0048] The present invention provides the following technical solution: a method and system for detecting semantic changes in dual-temporal remote sensing images based on multiple networks.

[0049] Example 1

[0050] See Figure 1 This embodiment provides a method for detecting semantic changes in dual-temporal remote sensing images based on multiple networks, and the specific steps include:

[0051] S1, obtaining a dual-temporal remote sensing image, and preprocessing the dual-temporal remote sensing image; using a sliding window to crop the preprocessed dual-temporal remote sensing image into a preset pixel specification, and randomly rotating the cropped dual-temporal remote sensing image to obtain a dual-temporal remote sensing image data set;

[0052] S11, the preprocessing includes radiation calibration, atmospheric correction, orthorectification, image fusion, image vectorization and image format conversion;

[0053] The image fusion specifically comprises fusing the multispectral and panchromatic images using a fusion algorithm;

[0054] The image vectorization specifically involves vectorizing the boundaries of land types in the remote sensing image, outlining different land types in the image, including water bodies, forest land, cultivated land and buildings, and assigning values ​​to pixels of different land types;

[0055] The image format conversion specifically includes format conversion using a vector-to-raster algorithm;

[0056] S2. Please refer to Figure 2 ,Using the multi-network joint learning strategy, the joint semantic change detection network is used as the main network, and the dual-branch semantic segmentation network is used as the auxiliary network,to construct a dual temporal change enhancement network;

[0057] S21, the joint semantic change detection network includes a semantic change detection encoder, a semantic change attention module and a semantic change detection decoder, and the dual-temporal remote sensing image dataset is input into the joint semantic change detection network, which is expressed as:

[0058]

[0059] In the formula, The characteristic image of semantic change representing the previous tense; Indicates the semantic change feature image of the later phase; I T1 Indicates the previous phase image; I T2 Indicates post-phase image; Siam SCD () represents the semantic change detection encoder with twin feature extraction; Represents the concatenation result of the bi-temporal features in the channel dimension; Concat represents the concatenation function of the bi-temporal features in the channel cascade; Classifier1 represents the first semantic change classification decoding layer; Classifier2 represents the second semantic change classification decoding layer; Classifier3 represents the third semantic change classification decoding layer; Represents the output feature map of semantic changes of the previous phase image; Output feature map representing semantic changes of later-phase images; Represents the output feature map of joint semantic change detection;

[0060] S22, the dual-branch semantic segmentation network includes a semantic segmentation encoder, a semantic segmentation attention module and a semantic segmentation decoder, and the dual-temporal remote sensing image dataset is input into the dual-branch semantic segmentation network, which is expressed as:

[0061]

[0062] In the formula, A semantically segmented feature image representing the front phase; Represents the semantic segmentation feature image of the posterior phase; Siam SS () represents a dual-branch semantic segmentation encoder with twin feature extraction; Represents the semantic output feature map of the previous phase image; Represents the semantic output feature map of the later phase image; Classifier4 represents the first semantic segmentation classification decoding layer;

[0063] S3, where the cross-attention fusion module is used to fuse the feature images extracted from the main network and the auxiliary network, which can be expressed as:

[0064]

[0065] A F =A1×A2;

[0066] O=v1×v2+A F ×(v1×v2);

[0067] Wherein, A1 represents the feature map of the extracted semantic segmentation encoder; Softmax represents the activation function; q1 represents the spatial domain vector of the preset semantic segmentation encoder; k1 represents the spatial domain vector of the preset semantic segmentation encoder; T represents the transposition operation; A2 represents the feature map of the extracted semantic change detection encoder; q2 represents the spatial domain vector of the preset semantic change detection encoder; k2 represents the spatial domain vector of the preset semantic change detection encoder; A F Represents the fused feature map; × represents the staggered dot product; O represents the output map of the cross-attention fusion module; v1 represents the spatial domain vector of the preset cross-attention fusion module; v2 represents the spatial domain vector of the preset cross-attention fusion module;

[0068] S4, using a two-stage decision fusion strategy to constrain the output of the main network in the form of pixel-level masks to generate the final dual-temporal remote sensing image change detection results;

[0069] In the first stage, the binary change detection map of the dual-branch semantic segmentation is obtained by binary superposition of the semantic output feature map of the previous phase image and the semantic output feature map of the subsequent phase image at the spatial vector level; the dot product operation of the pixel-level mask is performed on the binary change detection map and the output map of the cross-attention fusion module to calculate the intersection map;

[0070] In the second stage, the intersection map is subjected to a dot product operation with the output feature map of the joint semantic change detection to generate the final dual-temporal remote sensing image change detection result;

[0071] S5. Optimizing the dual-phase change enhancement network using an auxiliary loss function, which is expressed as:

[0072]

[0073] L joint =L BCD +L SCD +L SS +L Aux ;

[0074] In the formula, represents the pre-temporal loss for joint semantic change detection; represents the previous phase probability distribution of the reference label preset for the joint semantic change detection of the i-th category; represents the previous phase probability distribution of the joint semantic change detection prediction result of the i-th category; represents the posterior temporal loss for joint semantic change detection; represents the posterior temporal probability distribution of the reference label preset for joint semantic change detection of the i-th category; represents the posterior phase probability distribution of the joint semantic change detection prediction result of the i-th category; L BCD represents the loss of binary change detection map; Representing the probability distribution of the preset reference labels of the prediction results of joint semantic change detection; represents the probability distribution of the prediction results of joint semantic change detection; L SCD represents the total loss of joint semantic change detection; Represents the pre-temporal loss of a two-branch semantic segmentation network; Represents the previous time-phase probability distribution of the reference label preset by the dual-branch semantic segmentation network of the i-th category; Represents the previous phase probability distribution of the prediction result of the dual-branch semantic segmentation network of the i-th category; Represents the posterior temporal loss of a two-branch semantic segmentation network; Represents the posterior temporal probability distribution of the reference label preset by the dual-branch semantic segmentation network of the i-th category; represents the posterior phase probability distribution of the prediction result of the dual-branch semantic segmentation network of the i-th category; L SS represents the total loss of the two-branch semantic segmentation network; represents the cosine similarity for joint semantic change detection; Represents the cosine similarity of a two-branch semantic segmentation network; represents the cosine similarity of the binary change detection graph; L Aux represents the sum of cosine similarity; L joint represents the final auxiliary loss function; N represents the total number of categories in the dual-temporal remote sensing image dataset; i represents the index value of the i-th category;

[0075] S6, inputting the dual-phase remote sensing image data set into the dual-phase change enhancement network for iterative training to obtain a trained dual-phase change enhancement network;

[0076] S7, obtaining a dual-temporal remote sensing image to be detected, inputting the dual-temporal remote sensing image to be detected into the trained dual-temporal change enhancement network, and obtaining a corresponding detection result of the dual-temporal remote sensing image to be detected.

[0077] Example 2

[0078] This embodiment provides a dual-temporal remote sensing image semantic change detection system based on multiple networks, the system includes a data acquisition module, a network construction module, an image detection module and a result output module, wherein:

[0079] The data acquisition module is used to acquire dual-phase remote sensing images and preprocess the dual-phase remote sensing images; use a sliding window to crop the preprocessed dual-phase remote sensing images into a preset pixel specification, and randomly rotate the cropped dual-phase remote sensing images to obtain a dual-phase remote sensing image data set; transmit the dual-phase remote sensing image data set to the network construction module;

[0080] The network construction module is used to use a multi-network joint learning strategy, with a joint semantic change detection network as the main network and a dual-branch semantic segmentation network as the auxiliary network to construct a dual-phase change enhancement network; wherein, a cross-attention fusion module is used to fuse the feature images extracted by the main network and the auxiliary network, and a two-stage decision fusion strategy is used to generate the final dual-phase remote sensing image change detection result; the dual-phase change enhancement network is optimized using an auxiliary loss function, and the dual-phase remote sensing image data set is input into the dual-phase change enhancement network for iterative training to obtain a trained dual-phase change enhancement network;

[0081] The image detection module is used to obtain a dual-phase remote sensing image to be detected, input the dual-phase remote sensing image to be detected into the trained dual-phase change enhancement network, and obtain a corresponding detection result of the dual-phase remote sensing image to be detected;

[0082] The result output module is used to display the corresponding detection results of the dual-phase remote sensing image to be detected.

[0083] Example 3

[0084] This embodiment provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the multi-network-based dual-temporal remote sensing image semantic change detection method as described in any embodiment of the present invention is implemented.

[0085] Example 4

[0086] This embodiment provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the method for detecting semantic changes in dual-temporal remote sensing images based on multiple networks as described in any embodiment of the present invention is implemented.

[0087] It is worth noting that the system, electronic device and computer-readable storage medium described in the present invention are based on the same principles as the method described in Example 1, and will not be described in detail here.

[0088] The above descriptions are merely embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. A multi-network based method for semantic change detection of dual-temporal remote sensing images, characterized in that: The method comprises: Acquire dual-temporal remote sensing images, and preprocess the dual-temporal remote sensing images; crop the preprocessed dual-temporal remote sensing images into preset pixel specifications using a sliding window, and randomly rotate the cropped dual-temporal remote sensing images to obtain a dual-temporal remote sensing image data set; A dual-temporal change enhancement network is constructed by using a multi-network joint learning strategy, with a joint semantic change detection network as the main network and a dual-branch semantic segmentation network as the auxiliary network; wherein, a cross-attention fusion module is used to fuse the feature images extracted from the main network and the auxiliary network, and a two-stage decision fusion strategy is used to constrain the output of the main network in the form of a pixel-level mask to generate the final dual-temporal remote sensing image change detection result; the dual-temporal change enhancement network is optimized using an auxiliary loss function, and the dual-temporal remote sensing image data set is input into the dual-temporal change enhancement network for iterative training to obtain a trained dual-temporal change enhancement network; A dual-temporal remote sensing image to be detected is obtained, and the dual-temporal remote sensing image to be detected is input into a trained dual-temporal change enhancement network to obtain a corresponding detection result of the dual-temporal remote sensing image to be detected.

2. The method for detecting semantic changes in dual-temporal remote sensing images based on multiple networks according to claim 1, characterized in that: The preprocessing includes radiometric calibration, atmospheric correction, orthorectification, image fusion, image vectorization and image format conversion, wherein: The image fusion specifically comprises fusing the multispectral and panchromatic images using a fusion algorithm; The image vectorization specifically involves vectorizing the boundaries of land types in the remote sensing image, outlining different land types in the image, including water bodies, forest land, cultivated land and buildings, and assigning values ​​to pixels of different land types; The image format conversion is specifically to convert the format using a vector to raster algorithm.

3. The method for detecting semantic changes in dual-temporal remote sensing images based on multiple networks according to claim 1, characterized in that: The joint semantic change detection network includes a semantic change detection encoder, a semantic change attention module and a semantic change detection decoder. The dual-temporal remote sensing image dataset is input into the joint semantic change detection network, which is expressed as: In the formula, The characteristic image of semantic change representing the previous tense; The characteristic image representing the semantic change of the later phase; I T1 Indicates the previous phase image; I T2 Indicates post-phase image; Siam SCD () represents the semantic change detection encoder with twin feature extraction; Represents the concatenation result of the bi-temporal features in the channel dimension; Concat represents the concatenation function of the bi-temporal features in the channel cascade; Classifier1 represents the first semantic change classification decoding layer; Classifier2 represents the second semantic change classification decoding layer; Classifier3 represents the third semantic change classification decoding layer; Represents the output feature map of semantic changes of the previous phase image; Output feature map representing semantic changes of later-phase images; Represents the output feature map of joint semantic change detection.

4. The method for detecting semantic changes in dual-temporal remote sensing images based on multiple networks according to claim 1, characterized in that: The dual-branch semantic segmentation network includes a semantic segmentation encoder, a semantic segmentation attention module and a semantic segmentation decoder. The dual-temporal remote sensing image dataset is input into the dual-branch semantic segmentation network, which is expressed as: In the formula, A semantically segmented feature image representing the front phase; Represents the semantic segmentation feature image of the posterior phase; Siam SS () represents a dual-branch semantic segmentation encoder with twin feature extraction; Indicates the semantic segmentation result of the previous phase image; Represents the semantic segmentation result of the later phase image; Classifier4 represents the first semantic segmentation classification decoding layer.

5. The method for detecting semantic changes in dual-temporal remote sensing images based on multiple networks according to claim 4, characterized in that: The cross-attention fusion module is used to fuse the feature images extracted from the main network and the auxiliary network, which can be expressed as: A F =A1×A2; O=v1×v2+A F ×(v1×v2); Wherein, A1 represents the feature map of the extracted semantic segmentation encoder; Softmax represents the activation function; q1 represents the spatial domain vector of the preset semantic segmentation encoder; k1 represents the spatial domain vector of the preset semantic segmentation encoder; T represents the transposition operation; A2 represents the feature map of the extracted semantic change detection encoder; q2 represents the spatial domain vector of the preset semantic change detection encoder; k2 represents the spatial domain vector of the preset semantic change detection encoder; A F Represents the fused feature map; × represents the staggered dot product; O represents the output map of the cross-attention fusion module; v1 represents the spatial domain vector of the preset cross-attention fusion module; v2 represents the spatial domain vector of the preset cross-attention fusion module.

6. The method for detecting semantic changes in dual-temporal remote sensing images based on multiple networks according to claim 5, characterized in that: Using a two-stage decision fusion strategy, the output of the main network is constrained in the form of a pixel-level mask to generate the final dual-temporal remote sensing image change detection result, which is expressed as: In the first stage, the binary change detection map of the dual-branch semantic segmentation is obtained by binary superposition of the semantic output feature map of the previous phase image and the semantic output feature map of the subsequent phase image at the spatial vector level; the dot product operation of the pixel-level mask is performed on the binary change detection map and the output map of the cross-attention fusion module to calculate the intersection map; In the second stage, a dot product operation is performed on the intersection map and the output feature map of the joint semantic change detection to generate the final dual-temporal remote sensing image change detection result.

7. The method for detecting semantic changes in dual-temporal remote sensing images based on multiple networks according to claim 1, characterized in that: The dual-phase change enhancement network is optimized using an auxiliary loss function, which is expressed as: L joint =L BCD +L SCD +L SS +L Aux ; In the formula, represents the pre-temporal loss for joint semantic change detection; represents the previous phase probability distribution of the reference label preset for the joint semantic change detection of the i-th category; represents the previous phase probability distribution of the joint semantic change detection prediction result of the i-th category; represents the posterior temporal loss for joint semantic change detection; represents the posterior temporal probability distribution of the reference label preset for joint semantic change detection of the i-th category; represents the posterior phase probability distribution of the joint semantic change detection prediction result of the i-th category; L BCD represents the loss of binary change detection map; Representing the probability distribution of the preset reference labels of the prediction results of joint semantic change detection; represents the probability distribution of the prediction results of joint semantic change detection; L SCD represents the total loss of joint semantic change detection; Represents the pre-temporal loss of a two-branch semantic segmentation network; Represents the previous time-phase probability distribution of the reference label preset by the dual-branch semantic segmentation network of the i-th category; Represents the previous phase probability distribution of the prediction result of the dual-branch semantic segmentation network of the i-th category; Represents the posterior temporal loss of a two-branch semantic segmentation network; Represents the posterior temporal probability distribution of the reference label preset by the dual-branch semantic segmentation network of the i-th category; represents the posterior phase probability distribution of the prediction result of the dual-branch semantic segmentation network of the i-th category; L SS represents the total loss of the two-branch semantic segmentation network; represents the cosine similarity for joint semantic change detection; Represents the cosine similarity of a two-branch semantic segmentation network; represents the cosine similarity of the binary change detection graph; L Aux represents the sum of cosine similarity; L joint Represents the final auxiliary loss function; N represents the total number of categories of the dual-temporal remote sensing image dataset; i represents the index value of the i-th category.

8. A multi-network based dual-temporal remote sensing image semantic change detection system, characterized in that: The system includes a data acquisition module, a network construction module, an image detection module and a result output module, wherein: The data acquisition module is used to acquire dual-phase remote sensing images and preprocess the dual-phase remote sensing images; use a sliding window to crop the preprocessed dual-phase remote sensing images into a preset pixel specification, and randomly rotate the cropped dual-phase remote sensing images to obtain a dual-phase remote sensing image data set; transmit the dual-phase remote sensing image data set to the network construction module; The network construction module is used to use a multi-network joint learning strategy, with a joint semantic change detection network as the main network and a dual-branch semantic segmentation network as the auxiliary network to construct a dual-phase change enhancement network; wherein, a cross-attention fusion module is used to fuse the feature images extracted from the main network and the auxiliary network, and a two-stage decision fusion strategy is used to constrain the output of the main network in the form of a pixel-level mask to generate the final dual-phase remote sensing image change detection result; the dual-phase change enhancement network is optimized using an auxiliary loss function, and the dual-phase remote sensing image data set is input into the dual-phase change enhancement network for iterative training to obtain a trained dual-phase change enhancement network; The image detection module is used to obtain a dual-phase remote sensing image to be detected, input the dual-phase remote sensing image to be detected into the trained dual-phase change enhancement network, and obtain a corresponding detection result of the dual-phase remote sensing image to be detected; The result output module is used to display the corresponding detection results of the dual-phase remote sensing image to be detected.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the multi-network-based dual-temporal remote sensing image semantic change detection method as described in any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method for detecting semantic changes in dual-temporal remote sensing images based on multiple networks as described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Method and system for detecting change of dual-time-phase remote sensing image

    CN118212534A