Remote sensing image semantic change detection method and related equipment

By using the SCGNet network model, combined with the ResNet34 module, semantic branch, and change branch, the problems of insufficient feature fusion and semantic consistency in semantic change detection of remote sensing images are solved, achieving accurate detection of semantic change regions in remote sensing images and improving the accuracy and effectiveness of detection.

CN121074673APending Publication Date: 2025-12-05XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511266836.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-05
Publication Date
2025-12-05

AI Technical Summary

Technical Problem

Existing remote sensing image semantic change detection technologies suffer from insufficient dynamic adaptability due to dual-temporal feature fusion, inadequate semantic consistency structure layer modeling, and insufficient integration of multi-scale analysis with semantic information, resulting in inadequate detection accuracy and effectiveness.

Method used

The SCGNet network model is adopted, combined with the ResNet34 module, semantic branch and change branch. The SCE module extracts invariant region information, the SGF module highlights semantically different regions, and the MSCA module performs multi-scale analysis to achieve the synergistic effect of semantic information and change features and accurate detection.

Benefits of technology

It improves the accuracy and effectiveness of semantic change detection in remote sensing images, can accurately identify semantically changed regions, enhances the ability to model unchanged regions and the sensitivity to key changed regions, and improves the ability to identify change features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121074673A_ABST
    Figure CN121074673A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of remote sensing image change detection, in particular to a remote sensing image semantic change detection method and related equipment, and the core is as follows: firstly, obtaining dual-tense remote sensing image data, and inputting the dual-tense remote sensing image data into a trained semantic change detection model for processing; the model comprises a ResNet34 module, a semantic branch and a change branch. The ResNet34 module is responsible for extracting multi-scale features of the dual-tense image, generating feature maps of different scales and synchronously inputting a semantic branch and a change branch; the semantic branch is in close communication connection with the change branch, and enhanced semantic features are generated by extracting semantic information of the feature map and combining unchanged information of the change branch; and the change branch fuses the double-tense features and performs change region activation processing in combination with semantic information to obtain improved change features, and finally, a semantic change graph is output through the improved change features and semantic features. According to the method, precise detection of the semantic change of the dual-tense remote sensing image is realized through a branch cooperation mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of remote sensing image change detection, and particularly relates to a remote sensing image semantic change detection method and related equipment. BACKGROUND

[0002] In the field of remote sensing technology, semantic change detection is a key technology that mainly identifies semantic changes of ground objects, such as changes from farmland to building and changes in vegetation coverage, by analyzing double-time remote sensing images (remote sensing images of the same area taken at different times). This kind of technology is widely used in scenarios such as environmental monitoring, urban planning, and disaster assessment. For example, by comparing remote sensing images at different times, changes such as deforestation and urban expansion can be detected in a timely manner, providing data support for decision-making.

[0003] However, in practical applications, semantic change detection faces many challenges. First, double-time features are difficult to efficiently fuse, and image features at different times need to be effectively combined to accurately capture changes. Second, there is a large intra-class variance (large differences in the same class of ground objects in the image) and a small inter-class difference (similar performance of different classes of ground objects in the image) in remote sensing images, increasing the difficulty of recognition. In addition, the scale and shape of the change area are diverse, ranging from small new houses to large-scale land use conversion, and the shape is often irregular.

[0004] To address these issues, existing technologies have formed some solutions. In terms of double-time feature fusion, feature concatenation (combining two time features by dimension) or absolute difference (calculating the absolute difference of corresponding elements of two time features) is commonly used. To strengthen semantic consistency, some methods introduce specific loss functions (functions used to measure the difference between predicted results and true results when training the model) to guide the network to improve semantic expression ability. For the scale and shape of the change area, the mainstream approach is to use multi-scale analysis strategy to adapt to diverse change scales by processing features at different resolutions.

[0005] However, existing solutions still have obvious shortcomings. Feature concatenation or absolute difference is a static operation that cannot dynamically adjust the fusion strategy according to image content, making it difficult to extract truly key change features. Relying on loss functions to strengthen semantic consistency only stays at the supervision level and fails to explicitly model semantic consistency from the network structure, limiting its effectiveness in complex scenarios. Although multi-scale analysis strategies can handle scale issues, they often fail to fully incorporate semantic information, making it difficult to accurately capture the details and overall shape of the change area when detecting small changes or large-scale complex changes. These problems collectively constrain the performance improvement of semantic change detection networks. SUMMARY

[0006] The technical problems to be solved by the present application are to provide a remote sensing image semantic change detection method and related equipment to solve the technical problems of the dynamic adaptability of the dual-time feature fusion, the structure layer modeling of the semantic consistency, and the combination of the multi-scale analysis and the semantic information in the existing semantic detection method.

[0007] The purpose of the present application is achieved by the following technical solutions: In a first aspect, the present application provides a remote sensing image semantic change detection method, comprising: Obtaining dual-time remote sensing image data, inputting the dual-time remote sensing image data into a trained semantic change detection model to obtain a semantic change map; The semantic change detection model adopts an SCGNet network model; the SCGNet network model sequentially includes a ResNet34 module, a semantic branch, and a change branch; The ResNet34 module is used to learn multi-scale information in the dual-time remote sensing image data to obtain feature maps of different scales, and the feature maps are input into the semantic branch and the change branch at the same time; The semantic branch and the change branch are closely connected and are used to extract semantic information in the feature maps, and based on the semantic information and unchanging information in the change branch, reinforced semantic features are obtained; The change branch is used to fuse dual-time features to obtain change features, and based on the change features, semantic change region activation is performed to obtain improved change features, and based on the improved change features and the semantic features, a semantic change map is obtained.

[0008] As a further improvement of the present application, the semantic branch includes an SCE module and a semantic segmentation classifier; The semantic segmentation classifier is used to segment the semantic information in the dual-time feature map to obtain a semantic segmentation map; The SCE module is used to extract an unchanged region map in the input dual-time feature map to obtain a control matrix; based on the control matrix, semantic information is obtained, and the semantic information is added to the dual-time feature map based on a residual error to obtain semantic features.

[0009] As a further improvement of the present application, the SCE module execution step includes:

[0010] In the formula, , and respectively represent the classifier function, convolution, and convolution, is a change feature, is an invariant region graph, is a control function corresponding to the first i of the bitemporal feature, is semantic information corresponding to the first feature, is semantic information corresponding to the second feature, is the first of the bitemporal feature graph, is the second of the bitemporal feature graph, is the first i of the bitemporal feature graph corresponding to the semantic feature.

[0011] As a further improvement of the present application, the SGF module is used to perform the following steps: The input bitemporal feature graph is processed semantically by a semantic correlation attention mechanism unit to form a single-channel semantic change prompt graph; According to the generated single-channel semantic change prompt graph, mark the potential change region, Connect the bitemporal feature graph in the channel dimension to obtain a feature graph, multiply it with the semantic change prompt graph to filter redundancy, and generate an initial fusion feature through a convolution layer; Element-wise absolute difference is performed on the initial fusion feature and the input bitemporal feature graph to obtain a difference feature graph; Connect the difference feature graph and the initial fusion feature in the channel dimension, and then pass through a convolution layer to obtain a single-scale fusion feature; Integrate each scale feature in the multi-scale fusion feature to obtain a change feature.

[0012] As a further improvement of the present application, the semantic correlation attention mechanism unit includes a shared weight convolution layer, a cosine similarity calculation submodule, and a normalization mask generation submodule; The bitemporal feature graph is input into the shared weight convolution layer, and the bitemporal feature is projected into a unified feature space based on the shared weight convolution layer; The cosine similarity calculation submodule is used to calculate the preliminary similar region of the feature space; The normalization mask generation submodule is used to map the similarity region graph into a single-channel semantic change prompt graph.

[0013] As a further improvement of the present application, the change branch includes an MSCA module; the MSCA module includes a change region activation part and a multi-scale feature aggregation part; The change region activation part performs element-wise absolute difference operation on the bitemporal semantic feature to generate a semantic difference feature, and after multi-scale average pooling operation is performed on the semantic difference feature and the change feature, multi-scale semantic difference features and change features are obtained; The single-scale semantic difference feature is mapped to a single-channel mask through a convolution layer with a sigmoid activation function; The single-scale change feature is element-wise multiplied with the single-channel mask to activate a semantic change region, to obtain a single-scale semantic enhanced change feature; The single-scale semantic enhanced change features are fused in multiple scales through a multi-scale feature aggregation part, to further refine the change region, to obtain an improved change feature.

[0014] As a further improvement of the application, the change branch further comprises a change classifier, which is used to classify the improved change feature to obtain a change map, and the change map is combined with the semantic segmentation map to obtain a semantic change map.

[0015] In a second aspect, the application provides a remote sensing image semantic change detection system, comprising: a data acquisition module, configured to acquire remote sensing image data; a semantic detection module, in communication connection with the data acquisition module, comprising a semantic change detection model configured to perform semantic processing on the remote sensing image data; The semantic change detection model adopts an SCGNet network model; the SCGNet network model comprises a ResNet34 module, a semantic branch and a change branch in sequence; The ResNet34 module is configured to learn multi-scale information in the remote sensing image data, to obtain feature maps of different scales, and the feature maps are simultaneously input into the semantic branch and the change branch; The semantic branch and the change branch are in communication connection, configured to extract semantic information in the feature maps, and further obtain semantic features according to the semantic information and change features in the change branch; The change branch is configured to fuse the double-time feature maps to obtain change features, and further perform semantic change region activation based on the change features to obtain improved change features, and obtain a semantic change map based on the improved change features and the semantic features.

[0016] In a third aspect, the application provides a computer readable storage medium storing one or more programs, the one or more programs including instructions which, when executed by a computing device, cause the computing device to perform the remote sensing image semantic change detection method described above.

[0017] In a fourth aspect, the application provides a computing device, comprising: one or more processors, a memory and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs comprise instructions for performing the steps in the remote sensing image semantic change detection method described above.

[0018] The beneficial effects of the present application are that the embodiments of the present application provide a remote sensing image semantic change detection method, which, by means of the multi-scale feature learning ability of the ResNet34 module, combines the cooperative communication mechanism of the semantic branch and the change branch, makes the semantic information and the change features interact with each other, and optimizes the change features through the semantic change region activation of the change branch, so as to realize the accurate detection of the semantic change region in the double-time remote sensing image, and the present application effectively learns the multi-scale information of the double-time remote sensing image data through the ResNet34 module, lays a good foundation for the extraction of subsequent semantic information and change features, and by means of the communication connection and respective function realization of the semantic branch and the change branch, the semantic information and the change features can be fully combined, and the improved change features obtained by the change branch through the semantic change region activation can more accurately obtain the semantic change map, and thus the accuracy and effectiveness of the double-time remote sensing image semantic change detection are improved.

[0019] Further, the semantic branch of the present application proposes an SCE module, which improves the modeling ability of the network for unchanged regions from the structure based on the guiding mechanism of invariant regions, and enhances the alignment degree of double-time semantic features.

[0020] Further, the SGF module is introduced in the present application, which can actively highlight the regions with significant semantic differences, thereby improving the sensitivity to key change regions and significantly improving the discrimination ability of the model.

[0021] Further, the MSCA module is designed in the present application, which organically fuses multi-scale analysis and semantic enhancement information, can activate fine-grained change regions, and can also improve the boundary recognition ability of large-scale regions. BRIEF DESCRIPTION OF DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present application, and those skilled in the art can also obtain other drawings according to these drawings without creative labor.

[0023] Figure 1 is the SCGNet framework diagram in the embodiments of the present application; Figure 2 is the SGF structure schematic diagram in the embodiments of the present application; Figure 3 is the SCE structure schematic diagram in the embodiments of the present application; Figure 4 is the MSCA structure schematic diagram in the embodiments of the present application; Figure 5Fig. 1 is a structural schematic diagram of a computer device in an embodiment of the present application. DETAILED DESCRIPTION

[0024] In order to make the purpose and technical solutions of the present application clearer and more convenient to understand, the present application will be further described in detail below in combination with the drawings and embodiments. The specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0025] English abbreviation explanation: SCGNet (Semantic Correlation Guided Network): semantic correlation guided network, a deep learning network architecture combined with semantic correlation information; SGF (Semantic Guided Fusion): semantic guided fusion; SCE (Semantic Consistency Enhancement): semantic consistency enhancement; MSCA (Multi-Scale Change Activation): multi-scale change activation; SRA (Semantic Relevance Attention): semantic relevance attention.

[0026] The technical solutions of the present application will be described clearly and completely below in combination with the drawings and specific embodiments. The described embodiments are only a part of the embodiments of the present application, but not all the embodiments of the present application.

[0027] Embodiment 1 The present embodiment provides a remote sensing image semantic change detection method, mainly comprising: acquiring double-time remote sensing image data, inputting the double-time remote sensing image data into a trained semantic change detection model, and obtaining a semantic change map.

[0028] The semantic change detection model adopts an SCGNet network model. The SCGNet network model sequentially comprises a ResNet34 module, a semantic branch, and a change branch. The ResNet34 module is used to learn multi-scale information in the double-time remote sensing image data to obtain double-time feature maps of different scales, and the double-time feature maps are simultaneously input into the semantic branch and the change branch. The semantic branch and the change branch are communicatively connected, and are used to extract semantic information in the feature maps, and further obtain semantic features according to the semantic information and change features in the change branch. The change branch is used to fuse the double-time feature maps to obtain change features, and then perform semantic change region activation based on the change features to obtain improved change features, and finally obtain a semantic change map based on the improved change features and the semantic features. The semantic change map is obtained by multiplying a full-image semantic segmentation map and a change map.

[0029] In this embodiment, by obtaining double-time remote sensing image data and inputting the trained SCGNet network model, the ResNet34 module is used to learn multi-scale information therein to obtain double-time feature maps of different scales. After the feature maps are simultaneously input into the semantic branch and the change branch, the semantic branch and the change branch are communicatively connected to extract semantic information and obtain semantic features in combination with change features of the change branch. The change branch first fuses the double-time feature maps to obtain change features, and then performs semantic change region activation to obtain improved change features. Finally, a semantic change map is obtained based on the improved change features and the semantic features. The technical principle is to use the multi-scale feature learning capability of the ResNet34 module, and combine the cooperative communication mechanism of the semantic branch and the change branch, so that the semantic information and the change features interact with each other, and the change features are optimized through semantic change region activation of the change branch, thereby realizing accurate detection of the semantic change region in the double-time remote sensing image. Therefore, the method of this embodiment can effectively fuse the multi-scale features, semantic information, and change features of the double-time remote sensing image, and improve the accuracy and reliability of the semantic change map.

[0030] It is worth noting that the double-time remote sensing image data in this embodiment refers to remote sensing images obtained at different time points, for example Figure 1 As shown in FIG. 1, remote sensing images at T1 and T2 are obtained respectively. The dimensions of the double-time remote sensing image data are consistent, and both are data-cleaning images.

[0031] Further, the semantic branch in this embodiment comprises an SCE module and a semantic segmentation classifier. The semantic segmentation classifier is used to segment the semantic information in the double-time feature maps to obtain a semantic segmentation map. The SCE module is used to extract an invariant region map in the input double-time feature maps to obtain a control matrix. The semantic information is obtained based on the control matrix, and the semantic information is added to the double-time feature maps based on the residual to obtain semantic features. The SCE module execution steps include:

[0032] wherein, , and denote a classifier function, convolution and convolution, a varying feature, an invariant region map, a control function corresponding to the i-th feature in the bi-temporal feature, a semantic information corresponding to the 1st feature, a semantic information corresponding to the 2nd feature, a 1st in the bi-temporal feature map, a 2nd in the bi-temporal feature map, a semantic feature corresponding to the i-th in the bi-temporal feature map.

[0033] By setting the SCE module, the modeling capability of the network for the unchanged region is structurally improved based on the invariant region guided mechanism, and the alignment degree of the bi-temporal semantic feature is enhanced.

[0034] The changing branch includes an SGF module, and the SGF module is configured to perform the following steps: performing semantic processing on the input bi-temporal feature map through a semantic correlation attention mechanism unit to form a single-channel semantic change prompt map; connecting the bi-temporal feature map in the channel dimension according to the generated single-channel semantic change prompt map to mark the potential changing region, and multiplying the feature map after the connection with the semantic change prompt map to filter the redundancy, and then generating an initial fusion feature through a convolution layer; performing element-wise absolute difference on the initial fusion feature and the input bi-temporal feature map to obtain a difference feature map; connecting the difference feature map and the initial fusion feature in the channel dimension, and then obtaining a single-scale fusion feature through a convolution layer; and integrating each scale feature in the multi-scale fusion feature to obtain a varying feature.

[0035] The semantic correlation attention mechanism unit includes a shared weight convolution layer, a cosine similarity calculation submodule, and a normalized mask generation submodule. The bi-temporal feature map is input into the shared weight convolution layer, and the bi-temporal feature is projected into a unified feature space based on the shared weight convolution layer; the cosine similarity calculation submodule is used to calculate the preliminary similar region of the feature space; and the normalized mask generation submodule is used to map the similarity region map into a single-channel semantic change prompt map.

[0036] The SGF module is introduced to highlight the regions with significant semantic differences in the bi-temporal image. The module guides the network to focus on the key regions that have actually changed, and realizes more efficient and more discriminative bi-temporal feature fusion.

[0037] Further, the change branch includes an MSCA module; a change region activation part and a multi-scale feature aggregation part of the MSCA module. The semantic difference features are generated by performing element-wise absolute difference operation on the double-time semantic features through the change region activation part, and after multi-scale average pooling operation is performed on the semantic difference features and the change features, multi-scale semantic difference features and change features are obtained. The single-scale semantic difference features are mapped to a single-channel mask through a convolution layer with a sigmoid activation function; the single-scale change features are multiplied with the single-channel mask element by element to activate the semantic change region, and single-scale semantic enhanced change features are obtained. After multi-scale fusion is performed on the semantic enhanced change features of each single scale through the multi-scale feature aggregation part, the change region is further refined, and improved change features are obtained. The semantic features enhanced by the SCE module are integrated with multi-scale analysis through the MSCA module, which effectively improves the adaptability of the network to the scale and shape diversity of the change region, and realizes accurate positioning and refinement of the change region.

[0038] Further, the change branch further includes a change classifier, which is used to classify the change region to obtain a change map. After the change map is obtained, the semantic segmentation map output by the semantic branch is fused to obtain a semantic change map. The embodiment realizes sufficient information interaction between the semantic branch and the change branch, the semantic information can guide change detection, and the change branch can also guide semantic expression in reverse, thereby improving the accuracy of the final semantic change map.

[0039] Embodiment 2 The embodiment based on embodiment 1 provides a specific implementation of a remote sensing image semantic change detection method, and more detailed description of the semantic change detection model is as follows.

[0040] The semantic correlation guided network SCGNet for remote sensing image semantic change detection proposed in embodiment 1 has a main framework as shown in Figure 1 The architecture mainly includes ResNet34, a semantic branch and a change branch. The semantic branch includes an SCE module and a semantic segmentation classifier, and the change branch includes an SGF module, an MSCA module and a change classifier. The SGF highlights the regions with significant semantic difference between the double-time images, enhances the perception of the key change region, and realizes more efficient and more adaptive feature fusion. The SCE uses invariant information as a filter to strengthen the representation of the invariant region, thereby improving the overall semantic consistency. The MSCA combines the SCE enhanced semantic information with multi-scale analysis, effectively activates and refines the change region, and further improves the detection accuracy.

[0041] When a pair of remote sensing images When inputting into the SCGNet, ResNet34 first learns multi-scale information from it. Based on the special structure of ResNet34, the embodiment can obtain four pairs of feature maps corresponding to different scales according to the dual-temporal remote sensing images. The embodiment selects the features of the shallowest layer (rich in texture information) and the deepest layer (rich in semantic information) for semantic change detection, and they are denoted as and respectively, where , . It is worth noting that the embodiment adopts an improved ResNet34, that is, the down-sampling is removed after the second residual stage. Then, the obtained multi-scale features are simultaneously input into the semantic branch and the change branch. In order to ensure the quality of the semantic change map, the modules corresponding to the semantic branch and the change branch jointly learn and influence each other. Specifically, the SGF module deeply excavates the change clues hidden in the multi-scale features to generate change features . The SCE module aims to discover the semantics in the remote sensing images according to the invariable clues. To this end, the multi-scale features and the change features are jointly used to realize the semantic features corresponding to the dual-temporal remote sensing images. Next, in order to further activate the change area and take into account the semantics of the remote sensing images, the embodiment applies the MSCA module to the change features and the semantic features to obtain improved change features . Finally, by means of the semantic segmentation and the change classifier, the improved change features , the semantic features and are used to obtain the change map and the semantic segmentation map respectively. After a simple multiplication operation, the final semantic change map can be generated.

[0042] The SGF module aims to explore the change clues in the dual-temporal remote sensing images. Unlike the conventional cascading or absolute difference operation, the SGF module focuses on capturing the complex change patterns in the remote sensing images. To this end, the embodiment develops a semantic relevance attention (SRA) mechanism, which is committed to highlighting the areas with significant semantic differences in the dual-temporal features. Then, the SGF module is constructed based on the SRA.

[0043] The SRA contains three key components: a shared weight convolution layer, a cosine similarity calculation module, and a normalized mask generation module. Specifically, given a pair of feature maps , the SRA first applies a shared weight convolution to project the feature maps into a unified feature space. This step helps to realize semantic alignment and reduce the number of channels. Next, the cosine similarity calculation module is used to calculate the preliminary similarity map between the two features. The value range of the preliminary similarity map is between the two, which is not conducive to the network focusing on the area with a large semantic difference. Therefore, the normalized mask generation module is adopted to map the preliminary similarity map to a single-channel semantic change hint map , the value range of which is , indicating the degree of attention that should be given to each area. The above SRA process can be represented as:

[0044] wherein represents the cosine distance calculation. On the one hand, SRA can weaken / highlight the false / true change area, thereby improving the positioning accuracy of the change area by SCGNet. On the other hand, since SRA gives a higher weight to the area with a significant semantic difference, it helps the model focus more on the true change area, further improving the distinguishing ability of the change feature.

[0045] With the help of SRA, SGF is constructed, the structure of which is shown in Figure 2 . For the input feature map , firstly, an is generated by using SRA to indicate the potential change area. At the same time, the input features and are connected in the channel dimension to obtain the first feature map . The first feature map retains the rich information of the dual-phase remote sensing image, but also contains a considerable amount of redundancy. In order to filter out the redundant clues, the potential change area is multiplied by the first feature map , and the result is subjected to a convolution layer to generate the initial fusion feature . This operation suppresses the interference of irrelevant areas and helps SGF adaptively fuse the features of the change area. Subsequently, the element-wise absolute difference operation is performed on and to obtain the difference feature map . The difference feature map , although contains useful explicit change information, is not complete and clear enough due to the lack of context and semantic support. In contrast, the initial fusion feature has stronger semantic integrity and area consistency. Therefore, the element-wise absolute difference operation is performed on and in the channel dimension to make full use of their complementarity and further improve the expression ability of the change area. Finally, a convolution layer is adopted to obtain the fusion feature . The above process is represented as:

[0046] wherein denotes the absolute difference function, denotes convolution, denotes channel concatenation. After processing the single-scale features, the single-scale features are fused to obtain the change feature :

[0047] wherein, is the change feature, which is obtained based on the convolution of .

[0048] The embodiment effectively fuses the double-time features by the SGF and obtains high-quality change features .

[0049] The semantic change detection task also focuses on the semantic change of the change region, and the semantics of these regions are essentially different, which conflicts with the promotion of semantic consistency. However, there are still a large number of unchanged regions, and the double-time semantics of the unchanged regions remain consistent. If the semantic consistency expression of these unchanged regions can be effectively enhanced, the overall feature expression ability can be further improved, so that the semantic features of the same category in the double-time image are more consistent. Therefore, the embodiment proposes an SCE module, the structure of which is shown in Figure 3 . It mainly consists of a classifier, two convolution layers and two depth modulation convolution layers, and they are connected through residual connection. When the SCE module receives input data (the input data is fused by the up-sampled and , wherein denotes the time phase) and the change feature , first, the change feature is processed through a simple classifier (consisting of one convolution layer and a sigmoid function) to generate a preliminary change prediction map, from which the embodiment further derives the unchanged region map , wherein . Next, the unchanged region map is multiplied by the input feature and to extract the unchanged region in the double-time remote sensing image. Combined with two convolutions and sigmoid activation functions, two control matrices and are obtained. The two control matrices are used to filter and , thereby modeling the cross-time semantic consistency to obtain semantic information and By this invariant information guiding mechanism, the SCE module can utilize the knowledge of invariant regions to make the mutual transmission and alignment of semantics of the bi-temporal features, thus effectively enhancing the expression of semantic consistency. The semantic information is added to the corresponding input features in a residual manner, further enhancing the expression of unchanged regions. Finally, the two enhanced features are modulated by convolution to obtain robust semantic features and . The process is represented as:

[0050] wherein, , and represent the classifier function, convolution and convolution respectively, is the changed feature, is the invariant region map, is the control function corresponding to the i-th feature in the bi-temporal feature, is the semantic information corresponding to the 1st feature, is the semantic information corresponding to the 2nd feature, is the 1st in the bi-temporal feature map, is the 2nd in the bi-temporal feature map, is the semantic feature corresponding to the i-th in the bi-temporal feature map.

[0051] After obtaining the enhanced semantic features and , the semantic features are used to enhance the changed feature . Considering that the changed region usually has multiple scales and complex shapes, in addition to the semantic information, multi-scale clues should also be noted. Therefore, the present embodiment also proposes an MSCA module to enhance the changed feature by simultaneously utilizing semantic information and multi-scale context exploration. The MSCA module flow chart is shown in Figure 4 , which consists of a changed region activation part and a multi-scale feature aggregation part.

[0052] Element-wise absolute difference operation is performed on the semantic features and to generate a semantic difference feature map . In order to enhance the changed feature by utilizing semantic clues, the semantic difference feature map is first mapped to a single-channel mask by a convolution layer with a sigmoid activation function The semantic change region is highlighted. Then, a single-channel mask is obtained by performing an element-wise product with the change feature , thereby combining the semantic change and the binary change to achieve semantic enhancement of the change feature . This process can be understood as change region activation.

[0053] To explore multi-scale context knowledge, the above process is extended to a multi-scale version in this embodiment. First, multi-scale versions of the change feature and the semantic difference feature map are obtained by using an average pooling operation, for convenience, denoted as , where denotes the level of the feature map. The size of these feature maps is , where and . Then, semantic change region activation is performed at different scales. Taking the first scale as an example, semantic change region activation can be expressed as:

[0054]

[0055] where denotes a sigmoid function. When is obtained, they gradually fuse from deep to shallow, which can be expressed as:

[0056] where denotes a bilinear interpolation function. Among them, bilinear interpolation aims to adjust the size of the feature map, while and convolution focuses on analyzing and fusing different feature maps. By gradually refining and fusing change clues at different levels, MSCA can effectively enhance the model's perception of the change region, and ultimately achieve accurate characterization of the region features.

[0057] Embodiment 3 The embodiment provides a remote sensing image semantic change detection system, which is used to implement the remote sensing image semantic change detection method in embodiments 1 and 2. The system specifically comprises: a data acquisition module for acquiring remote sensing image data.

[0058] a semantic detection module in communication connection with the data acquisition module, comprising a semantic change detection model for performing semantic processing on the remote sensing image data.

[0059] The semantic change detection model adopts an SCGNet network model; the SCGNet network model sequentially comprises a ResNet34 module, a semantic branch and a change branch.

[0060] The ResNet34 module is used to learn multi-scale information in remote sensing image data, and obtain feature maps of different scales, which are input into the semantic branch and the change branch at the same time. The semantic branch and the change branch are in communication connection, and are used to extract semantic information in the feature maps, and obtain semantic features according to the semantic information and change features in the change branch. The change branch is used to fuse the double-time feature maps to obtain change features, and then perform semantic change region activation based on the change features to obtain improved change features, and obtain a semantic change map based on the improved change features and the semantic features.

[0061] Embodiment 4 In another embodiment of the present application, a computer readable storage medium is provided as a storage component in a terminal device, and its function is to store programs and data. It should be noted that the computer readable storage medium herein not only covers the built-in storage component of the terminal device, but also includes the expansion storage component supported by the device. Its essence is a tangible medium that can contain or store programs, which can be called or cooperated with an instruction execution system, device or instrument. The storage medium provides a storage area for the operating system of the terminal, and also stores one or more instructions suitable for the processor to load and run, which can constitute one or more computer programs containing program codes.

[0062] Specifically, examples (non-exclusive list) of the computer readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory, an optical fiber, a portable optical disk read-only memory, an optical storage device, a magnetic storage device, or any reasonable combination of the above types.

[0063] The storage medium can also include a data signal propagating as a baseband part or as part of a carrier wave, which carries readable program codes. Such a propagating data signal can take various forms, including but not limited to electromagnetic signals, optical signals or any reasonable combination of the two. In addition, the computer readable storage medium can also refer to other readable media other than traditional readable storage media, which can send, propagate or transmit programs for use or cooperation of instruction execution systems, devices or instruments. The program code on the storage medium can be transmitted through any suitable medium, including but not limited to wireless, wired, optical cable, etc. Transmission methods, or any reasonable combination thereof.

[0064] Program code implementing the operations for the present application can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computing device, partly on the user's device, as a stand-alone software package, partly on the user's device and partly on a remote computing device or entirely on the remote computing device or server. When the program code is executed on the remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network ("LAN"), a wide area network ("WAN"), the Internet, or the like.

[0065] The processor is capable of loading and running one or more instructions stored in the computer readable storage medium to implement the corresponding steps of the remote sensing image semantic change detection method described in Embodiment 1.

[0066] Embodiment 5 Figure 5 A schematic diagram of a computer device provided by an embodiment of the present application.

[0067] Please refer to Figure 5 The terminal device is a computer device, and the computer device 60 of this embodiment includes a processor 61, a memory 62, and a computer program 63 stored in the memory 62 and executable on the processor 61. The computer program 63, when executed by the processor 61, implements the remote sensing image semantic change detection method in the embodiment. To avoid repetition, details are not repeated here. Alternatively, the computer program 63, when executed by the processor 61, implements the functions of the models / units in the computing system in the remote sensing image semantic change detection processing method of the embodiment. To avoid repetition, details are not repeated here.

[0068] The computer device 60 can be a desktop computer, a notebook computer, a palm computer, a cloud server, and the like. The computer device 60 can include, but is not limited to, a processor 61 and a memory 62. Those skilled in the art can understand that Figure 5 The computer device 60 is merely an example and does not constitute a limitation on the computer device 60, and can include more or fewer components than shown, or combine certain components, or different components, for example, the computer device can also include an input / output device, a network access device, a bus, and the like.

[0069] The processor 61 can be a central processing unit (CPU), and can also be other general-purpose processors, central processing units, graphics processing units, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic components, quantum computing-based data processing logic components, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.

[0070] The memory 62 can be an internal storage unit of the computer device 60, such as a hard disk or a memory of the computer device 60. The memory 62 can also be an external storage device of the computer device 60, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc.

[0071] Further, the memory 62 can include both an internal storage unit and an external storage device of the computer device 60. The memory 62 is used to store computer programs and other programs and data required by the computer device. The memory 62 can also be used to temporarily store data that has been output or will be output.

[0072] Any reference to a memory, database or other medium herein includes at least one of volatile and non-volatile memories, and can include database(s) or other data storage. Non-volatile memories can include read only memories, floppy diskettes, compact disc read only memories, optical disks, flash memories, programmable read only memories, erasable programmable read only memories, electrically erasable programmable read only memories, magnetic tapes, and the like. Volatile memories can include random access memories, static random access memories, dynamic random access memories, and the like. By way of illustration, a RAM can be a static random access memory (SRAM), a dynamic random access memory (DRAM), or the like. The memory can be used for storing data or executable instructions as described herein.

[0073] The database(s) can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a blockchain, without limitation. The processor(s) can be a general processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic, a data processing logic based on quantum computing, without limitation.

Claims

1. A method for detecting semantic change in remote sensing images, characterized in that, The method comprises the following steps: acquiring double-time remote sensing image data, inputting the double-time remote sensing image data into a trained semantic change detection model to obtain a semantic change map; the semantic change detection model adopts an SCGNet network model; the SCGNet network model sequentially comprises a ResNet34 module, a semantic branch and a change branch; the ResNet34 module is used for learning multi-scale information in the double-time remote sensing image data to obtain double-time feature maps of different scales, and the double-time feature maps are input into the semantic branch and the change branch at the same time; the semantic branch and the change branch are communicatively connected, and are used for extracting semantic information in the feature maps, and further obtaining semantic features according to the semantic information and change features in the change branch; the change branch is used for fusing the double-time feature maps to obtain change features, performing semantic change region activation based on the change features to obtain improved change features, and obtaining the semantic change map based on the improved change features and the semantic features. 2.The method of claim 1, wherein, the semantic branch comprises an SCE module and a semantic segmentation classifier; the semantic segmentation classifier is used for segmenting the semantic information in the double-time feature maps to obtain a semantic segmentation map; the SCE module is used for extracting an invariant region map in the input double-time feature map to obtain a control matrix, obtaining semantic information based on the control matrix, and adding the semantic information to the double-time feature map according to a residual error to obtain semantic features. 3.The method of claim 2, wherein, the SCE module performs the following steps: wherein, , and denote a classifier function, convolution and convolution, a varying feature, an invariant region map, a control function corresponding to the i-th feature in the bi-temporal feature map, semantic information corresponding to the 1st feature, semantic information corresponding to the 2nd feature, the 1st in the bi-temporal feature map, the 2nd in the bi-temporal feature map, semantic feature corresponding to the i-th in the bi-temporal feature map. 4.The method of claim 1, wherein, the change branch comprises an SGF module, and the SGF module is used for performing the following steps: forming a single-channel semantic change prompt map by performing semantic processing on the input double-time feature map through a semantic correlation attention mechanism unit; marking a potential change region according to the generated single-channel semantic change prompt map, connecting the double-time feature maps in a channel dimension to obtain a feature map, multiplying the feature map by the semantic change prompt map to filter redundancy, and generating an initial fusion feature through a convolution layer; performing element-wise absolute difference on the initial fusion feature and the input double-time feature map to obtain a difference feature map; connecting the difference feature map and the initial fusion feature in the channel dimension, and then obtaining a single-scale fusion feature through a convolution layer; integrating each scale feature in the multi-scale fusion feature to obtain a change feature. 5.The method of claim 4, wherein, the semantic correlation attention mechanism unit comprises a shared weight convolution layer, a cosine similarity calculation submodule and a normalized mask generation submodule; inputting the double-time feature map into the shared weight convolution layer, and projecting the double-time feature map into a unified feature space based on the shared weight convolution layer; calculating a preliminary similar region of the feature space by using the cosine similarity calculation submodule; mapping the similarity region map into a single-channel semantic change prompt map by using the normalized mask generation submodule. 6.The method of claim 4, wherein, the change branch comprises an MSCA module; the MSCA module comprises a change region activation part and a multi-scale feature aggregation part; performing element-wise absolute difference operation on the double-time semantic features through the change region activation part to generate semantic difference features, and performing multi-scale average pooling operation on the semantic difference features and the change features to obtain multi-scale semantic difference features and change features; The single-scale semantic difference feature is mapped to a single-channel mask through a convolution layer with a sigmoid activation function; The single-scale change feature is element-wise multiplied with the single-channel mask to activate the semantic change region, to obtain a single-scale semantic enhanced change feature; The single-scale semantic enhanced change features are fused in multiple scales through a multi-scale feature aggregation part, to further refine the change region, to obtain an improved change feature.

7. The method of claim 6, wherein, The change branch further includes a change classifier, which classifies the improved change feature to obtain a change map, and the change map is combined with a semantic segmentation map to obtain a semantic change map.

8. A remote sensing image semantic change detection system, characterized in that, The method comprises the following steps: The data acquisition module is configured to acquire remote sensing image data; The semantic detection module is in communication connection with the data acquisition module, and comprises a semantic change detection model configured to perform semantic processing on the remote sensing image data; The semantic change detection model adopts an SCGNet network model; the SCGNet network model comprises a ResNet34 module, a semantic branch and a change branch in sequence; The ResNet34 module is configured to learn multi-scale information in the remote sensing image data, to obtain feature maps of different scales, and the feature maps are input into the semantic branch and the change branch at the same time; The semantic branch and the change branch are in communication connection, and are configured to extract semantic information in the feature maps, and further obtain semantic features according to the semantic information and change features in the change branch; The change branch is configured to fuse the double-time feature maps to obtain change features, and further perform semantic change region activation based on the change features to obtain improved change features, and obtain a semantic change map based on the improved change features and the semantic features.

9. A computer-readable storage medium storing one or more programs, the one or more programs comprising instructions that when executed by a computer cause the computer to perform a method of any of claims 1-8. The one or more programs include instructions that, when executed by a computing device, cause the computing device to perform the remote sensing image semantic change detection method of any one of claims 1 to 7.

10. A computing device, comprising: The method comprises the following steps: One or more processors, memories and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs comprise steps for executing the remote sensing image semantic change detection method of any one of claims 1 to 7.