Lightweight semantic enhancement and change integration remote sensing image semantic change detection method

By adopting a lightweight multi-task network architecture and deep separable convolution in remote sensing image semantic change detection, the interaction ability between semantic information and changing information is improved, and the problems of high computational cost and insufficient interaction in the existing technology are solved, and efficient and accurate semantic change detection is achieved.

CN120071176AActive Publication Date: 2025-05-30WUHAN UNIV

Patent Information

Application Number
CN202510123288.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-26
Publication Date
2025-05-30
Estimated Expiration
2045-01-26

AI Technical Summary

Technical Problem

The existing semantic change detection methods in the field of intelligent interpretation of remote sensing images have problems such as insufficient model lightweighting, high computational cost, slow inference speed, and insufficient interaction between semantic information and change information, which affect detection efficiency and accuracy.

Method used

The lightweight multi-task weight sharing encoder and multi-task decoder are adopted, combining jump connection and depth separation convolution, and designing semantic enhancement fusion modules and change integration conversion modules to improve the interaction ability between semantic information and change information.

Benefits of technology

It realizes efficient and accurate semantic change detection, reduces the number of parameters and calculations of the model, improves the inference speed, adapts to different remote sensing image data types, and is efficient and scalable.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071176A_ABST
    Figure CN120071176A_ABST
Patent Text Reader

Abstract

The invention provides a lightweight semantic enhancement and change integration remote sensing image semantic change detection method. The method comprises the following steps: acquiring a dual-temporal remote sensing image and constructing a sample library; constructing a multi-task semantic change detection network model, and performing network model training optimization based on the sample library; an encoder part of the multi-task semantic change detection network model adopts a lightweight multi-task weight sharing encoder, and supports semantic segmentation and binary change detection tasks at the same time; a decoder part of the multi-task semantic change detection network model comprises a semantic segmentation decoder corresponding to a first time phase, a semantic segmentation decoder corresponding to a second time phase, and a binary change detection decoder; an image of a first time phase and an image of a second time phase input into the encoder part are respectively processed to generate feature maps with different resolutions, and the feature maps are transmitted to the decoder part through jump connection; and inputting a dual-temporal remote sensing image to the trained multi-task semantic change detection network model for semantic change detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent interpretation of remote sensing images, and particularly to a lightweight semantic enhancement and change integration remote sensing image semantic change detection technical solution. Background Art

[0002] Changes on the Earth's surface are everywhere and present diverse forms at different scales. Currently, remote sensing change detection has been widely applied in fields such as urban planning, resource and environmental monitoring, and disaster emergency response. Remote sensing change detection methods are mainly divided into Binary Change Detection (BCD) and Semantic Change Detection (SCD). BCD is usually used to detect a single target, such as buildings, floods, or farmland, and only needs to identify the changed areas. Although it is widely applied, it is difficult to meet the monitoring requirements for multiple land cover types in practical applications. In contrast, SCD can simultaneously identify the changed areas and their corresponding Land Use and Land Cover (LULC) types and provide "from-to" transition information. Although the research is relatively less, it shows higher efficiency and value in practical applications. Therefore, it is urgent to study more mature technologies for popularization and application.

[0003] With the development of computer vision technology, the application of convolutional neural networks in remote sensing image analysis has gradually become popular. However, traditional deep learning models often have a large number of parameters and high computational complexity, and are not suitable for rapid processing in resource-constrained environments. The introduction of lightweight models such as MobileViT has significantly reduced the computational cost while maintaining high performance. This innovation in model architecture makes efficient and accurate change detection possible in practical applications and promotes the development of remote sensing image intelligent interpretation technology.

[0004] Currently, deep learning-based SCD methods usually adopt a three-branch structure, integrating the Semantic Segmentation (SS) and Binary Change Detection (BCD) tasks into a multi-task framework. The decoding stage usually includes two SS branches and one BCD branch, and the detection accuracy is improved through the synergistic effect between the branches. Although existing research has optimized the temporal correlation and task correlation, the existing methods still face two main challenges: one is the insufficient lightweight of the model, resulting in high computational cost and slow inference speed, which limits the efficiency and scalability of SCD in practical applications; the other is the insufficiency in dealing with the interaction between semantic information and change information, resulting in inconsistent output results of each branch, directly affecting the semantic change detection results. Summary of the Invention

[0005] In view of the deficiencies of the semantic change detection methods in the field of intelligent interpretation of remote sensing images, the present invention combines deep learning and high-resolution remote sensing images to provide a lightweight semantic enhancement and change integration remote sensing image semantic change detection method and device, which can effectively meet the requirements of many practical applications such as natural resource monitoring.

[0006] The technical solution provided by the present invention is a lightweight semantic enhancement and change integration remote sensing image semantic change detection method, including:

[0007] Collecting dual-temporal remote sensing images and constructing a sample library;

[0008] Constructing a multi-task semantic change detection network model and optimizing the network model training based on the sample library; the encoder part of the multi-task semantic change detection network model adopts a lightweight multi-task weight sharing encoder, and the lightweight multi-task weight sharing encoder supports both semantic segmentation and binary change detection tasks, including an encoder branch corresponding to the first temporal phase and an encoder branch corresponding to the second temporal phase; the decoder part of the multi-task semantic change detection network model includes a semantic segmentation decoder corresponding to the first temporal phase and a semantic segmentation decoder corresponding to the second temporal phase, and a binary change detection decoder; the images of the first temporal phase and the second image input to the encoder part are respectively processed to generate feature maps of different resolutions, and are transmitted to the decoder part through skip connections to achieve multi-level feature integration, and finally a binary change map and two semantic segmentation maps are generated, and a semantic change map is extracted by combining mask operations;

[0009] Inputting the dual-temporal remote sensing images into the trained multi-task semantic change detection network model for semantic change detection.

[0010] Moreover, in the lightweight multi-task weight sharing encoder, the encoder branch corresponding to the first temporal phase and the encoder branch corresponding to the second temporal phase have the same structure and share weights, and each includes two stages. In the first stage, a lightweight feature extraction network is used to perform feature encoding on the input image to extract deep semantic information, and then three groups of features of different scales are output. In the second stage, these three groups of features are respectively input into three receptive field attention convolutional blocks to extract receptive field spatial features, and three groups of optimized features are obtained.

[0011] Moreover, each semantic segmentation decoder includes three spatio-temporal semantic enhancement and fusion modules connected in series and a depthwise separable residual convolutional block in sequence; the binary change detection decoder includes three multi-type temporal change integration and conversion modules connected in series and a depthwise separable residual convolutional block in sequence.

[0012] Moreover, a pair of dual-temporal RGB three-band optical remote sensing images T 1 and T 2After passing through the corresponding encoder branches of the lightweight multi-task weight-sharing encoder respectively, the outputs are features with different scales corresponding to T 1 features with corresponding different scales and features with different scales corresponding to T 2 features with corresponding different scales and and are transmitted to the decoder part through skip connections.

[0013] Moreover, the spatio-temporal semantic enhancement and fusion module receives feature inputs from two time phases, uses skip connections and depthwise separable convolutions to fuse change information and generate richer feature representations; in this module, change information is captured by performing absolute value difference operations on the features, and combined with the directional attention mechanism, significantly enhancing the model's semantic expression ability for the changed regions.

[0014] Moreover, the multi-type temporal change integration and transformation module generates differential features by performing various forms of operations on the features at different times, and preserves the context information of the features; further extracts spatial and channel features, generates enhanced spatial and channel weights, and performs feature integration, and enhances the change feature representation in the depth features through the blending transformation mechanism, and outputs a more accurate binary change result.

[0015] Moreover, the depthwise separable residual convolution block is a residual block constructed based on the depthwise separable convolution DSConv.

[0016] On the other hand, the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the lightweight semantic enhancement and change integration remote sensing image semantic change detection method as described above.

[0017] On the other hand, the present invention provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the lightweight semantic enhancement and change integration remote sensing image semantic change detection method as described above.

[0018] On the other hand, the present invention provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the lightweight semantic enhancement and change integration remote sensing image semantic change detection method as described above.

[0019] In summary, the present invention proposes a method and device for semantic change detection in remote sensing images that combines lightweight semantic enhancement and change integration. Based on the design concept of lightweight skeletons and skip connections, a network model with a new decoding mechanism is constructed to better address the interaction problem between semantic information and change information in the SCD task. While maintaining high detection performance, the model has a low number of parameters, low computational complexity, and fast inference speed, and can better meet the requirements in practical applications such as urban planning, resource and environmental monitoring, and disaster emergency response. In addition, the detection model of the present invention is efficient and scalable, adapts to different types of remote sensing image data, and can be fine-tuned according to requirements to adapt to other change detection tasks. This flexibility makes it widely applicable in fields such as emergency response, agricultural monitoring, and infrastructure management. In short, the present invention provides an efficient, accurate, and scalable solution, significantly improving the efficiency of remote sensing image analysis and applications.

[0020] Specifically, the beneficial effects of the technical solution provided by the present invention are as follows:

[0021] (1) High efficiency and low computational cost: By adopting a lightweight network structure and skip connection design, the number of parameters and computational complexity of the model are significantly reduced, and the inference speed is greatly improved. This efficient design not only reduces the consumption of hardware resources but also enables the model to achieve real-time processing in resource-constrained environments and better meet the requirements in practical applications.

[0022] (2) Accurate semantic change detection ability: The present patent innovatively designs a semantic enhancement fusion module and a change integration transformation module, effectively improving the interaction ability between semantic information and change information. By enhancing the intra-class similarity and inter-class separability of the changed regions, the model achieves higher semantic segmentation accuracy and change detection performance in complex scenarios, providing more reliable data support for fields such as urban planning and environmental monitoring.

[0023] (3) Wide applicability and flexible adjustment ability: The model has high adaptability and scalability. It is not only applicable to various types of remote sensing image data but also can be fine-tuned according to different application requirements to meet the change detection tasks in multiple fields such as urban planning, resource and environmental monitoring, and disaster emergency response. Its flexibility makes it have great application potential and value in actual production and emergency scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 It is a structural diagram of the network model according to an embodiment of the present invention.

[0025] Figure 2 It is a structural diagram of the depthwise separable residual convolution block according to an embodiment of the present invention.

[0026] Figure 3Schematic diagram of the spatio-temporal semantic enhancement and fusion module according to an embodiment of the present invention.

[0027] Figure 4 Schematic diagram of the multi-type temporal change integration and conversion module according to an embodiment of the present invention. Detailed implementation manners

[0028] The technical solution of the present invention will be specifically described below in conjunction with the accompanying drawings and embodiments.

[0029] Compared with traditional binary change detection methods, semantic change detection can not only identify change regions, but also detect the land use and land cover types of these regions simultaneously, providing "from-to" transition information, thus showing higher efficiency and application value in actual production. In response to the urgent needs in the fields of urban planning, resource and environmental monitoring, and disaster emergency response, the present invention designs an efficient and accurate remote sensing semantic change detection scheme.

[0030] An embodiment of the present invention proposes a lightweight semantic enhancement and change integration remote sensing image semantic change detection method, including:

[0031] 1) Collect dual-temporal remote sensing images and construct a sample library;

[0032] In specific implementation, two-phase high-resolution remote sensing image data of the same area can be obtained, preprocessed and labeled to generate samples required for subsequent steps.

[0033] 2) Construct a multi-task semantic change detection network model and optimize the network model training based on the sample library; the encoder part of the multi-task semantic change detection network model uses a lightweight multi-task weight sharing encoder, and the lightweight multi-task weight sharing encoder supports semantic segmentation and binary change detection tasks simultaneously, including an encoder branch corresponding to the first phase and an encoder branch corresponding to the second phase; the decoder part of the multi-task semantic change detection network model includes a semantic segmentation decoder corresponding to the first phase and a semantic segmentation decoder corresponding to the second phase, and a binary change detection decoder; the images of the first phase and the second image input to the encoder part are respectively processed to generate feature maps of different resolutions, and are transmitted to the decoder part through skip connections to achieve multi-level feature integration, and finally generate a binary change map and two semantic segmentation maps, and combine masking operations to extract a semantic change map.

[0034] In an embodiment of the present invention, a semantic change detection network model based on MobileViTv3, a multi-task architecture, and a multi-task decoding module is constructed, including a semantic enhancement fusion and change integration conversion module, and a multi-task decoding module is designed to more effectively extract semantic change features in the network model; more preferably, a multi-class cross-entropy loss function, a binary cross-entropy loss function, and a semantic consistency loss function can be used to optimize the multi-task network model from multiple perspectives, making it more focused on the change region and semantic categories, thereby improving the detection effect.

[0035] 3) Input the dual-temporal remote sensing images into the trained multi-task semantic change detection network model for semantic change detection.

[0036] Using the trained network model to perform semantic change detection on the dual-temporal remote sensing images to be processed can generate high-precision detection results.

[0037] A lightweight semantic enhancement and change integration remote sensing image semantic change detection method provided by an embodiment of the present invention further preferably adopts an implementation manner including the following process:

[0038] I. Collect dual-temporal remote sensing images and construct a sample library

[0039] The present invention further proposes to improve the change detection ability of dual-temporal high-resolution remote sensing images through a carefully designed data processing flow. First, when collecting images, preferentially select remote sensing images containing a large number of change regions to enrich the change samples of the model. Subsequently, preprocess the images, including size standardization, cropping invalid regions, denoising, etc., to ensure the stability of the image quality. In the annotation stage, use professional tools to finely annotate the change regions and their land cover types to provide diverse label data for the model. Then, expand the sample diversity through data augmentation techniques such as affine transformation and hue adjustment to improve the robustness and generalization ability of the model. Finally, divide the dataset into a training set, a validation set, and a test set to provide scientific data support for the learning, parameter tuning, and performance evaluation of the model, ensure its stable performance on different data, and thus improve the overall detection effect.

[0040] As a preferred embodiment thereof, the specific implementation of the sample library construction of the embodiment includes the following sub-steps:

[0041] 1) First, for the collection of dual-temporal high-resolution remote sensing images, the present invention emphasizes selecting those remote sensing images containing a large number of change regions during the collection process to ensure that the model can obtain rich change samples in subsequent training.

[0042] 2) After obtaining these images, strict preprocessing work is recommended. These preprocessing steps include adjusting all images to a unified size format, cropping the images to remove invalid areas, and eliminating noise interference in the images through denoising processing. These steps can improve the quality of the image data and ensure the accuracy and consistency of subsequent analysis.

[0043] 3) In the data annotation stage, it is preferably recommended to use professional annotation tools to finely annotate the changed areas in each remote sensing image. First, outline the changed areas in the multi-temporal images and mark the positions with obvious differences in different time phases.

[0044] Then, for these changed areas, detail the land cover types of each time phase, such as building, water body, road, vegetation and other categories, so as to provide diverse label data for the model. This fine annotation process can not only improve the model's recognition ability of different category changes, but also provide detailed data basis for subsequent analysis.

[0045] 4) After completing the annotation, data augmentation processing of the image data is a key step. By applying data augmentation techniques such as affine transformation (including scaling, rotation and translation), tone adjustment (such as brightness and contrast changes), random cropping and flipping, the diversity of sample data can be significantly expanded. These augmentation methods effectively increase the robustness of the model in the face of different environments, lighting conditions and perspective changes, and help reduce the overfitting phenomenon that occurs during the training process of the model, thus improving the overall performance and generalization ability of the model.

[0046] 5) Finally, in order to enable the dataset to fully play its role during the training process, divide the dataset into a training set, a validation set and a test set according to the actual training requirements. In the subsequent network model training and optimization stage, the training set is used for the learning and parameter optimization of the model, the validation set is used for parameter tuning and effect evaluation during the model training process, and the test set is used to evaluate the final performance of the model after the training is completed. This division method can ensure the stable performance of the model on different data and effectively avoid overfitting and underfitting problems, providing solid data support for the training and evaluation of the network model.

[0047] II. Construct a multi-task semantic change detection network model based on MobileViTv3

[0048] The present invention designs a brand-new network structure, which includes two encoders with shared weights and three decoders with different functions, and makes full use of lightweight design, skip connections, and innovative decoding strategies. At the encoder end, MobileViTv3 (the third-generation mobile-friendly vision transformer) and Receptive-FieldAttention convolution (RFAConv) are preferably adopted to construct a lightweight shared encoder applicable to the dual tasks of semantic segmentation (SS) and binary change detection (BCD). While ensuring sensitivity to detailed features, the computational load and model complexity are reduced. In the decoder part, with three spatio-temporal semantic enhancement fusion modules (TSEF) and a residual block based on depthwise separable convolution as the core, two semantic segmentation decoders with the same function are constructed to output two semantic segmentation maps, and the consistency of homogeneous features and the distinguishability of heterogeneous features are improved; and with three multi-type temporal change integration and conversion modules (Temporal-spatial semantic enhancement fusion module, TSEF) and a residual block based on depthwise separable convolution as the core, a binary change detection decoder is constructed. In this decoding branch, the output of BCD is used as auxiliary information for the semantic segmentation decoder, so as to ensure the capture of key regions and improve the accuracy of change information processing. Generally speaking, the model receives a pair of images T1 and T2, generates feature maps with different resolutions after processing, and transmits them to the decoder through skip connections to achieve multi-level feature integration. Finally, the model generates accurate binary change maps and two semantic segmentation maps, and combines mask operations to extract high-precision semantic change maps, demonstrating efficient change detection capabilities under limited resources.

[0049] For the specific implementation of MobileViTv3, please refer to the relevant literature, which will not be elaborated in this invention: Wadekar S N, Chaurasia A. Mobilevitv3: Mobile-friendly vision transformer with simple and effective fusion of local, global and input features[J]. arXiv preprint arXiv:2209.15159, 2022.

[0050] The specific implementation of RFAConv can be found in the relevant literature and will not be elaborated in this invention: Zhang X, Liu C, Yang D, et al. RFAConv: Innovating spatial attention and standard convolutional operation [J]. arXiv preprint arXiv:2304.03198, 2023.

[0051] As a preferred embodiment among them, the network model structure of the embodiment of the present invention is as Figure 1 shown and is specifically described as follows:

[0052] At the encoding end, the following processing is included:

[0053] A1) First, a pair of dual-temporal RGB three-band optical remote sensing images is input into the lightweight multi-task weight-sharing encoder and The superscript 3 indicates that the input image has 3 bands (which can also be replaced with other multi-band images), and H and W respectively represent the height and width of the input image. is the spatial structure. T 1 and T 2 will simultaneously perform feature encoding along two branches with shared parameters of the lightweight multi-task weight-sharing encoder.

[0054] A2) The lightweight multi-task weight-sharing encoder includes two stages. The first stage is the lightweight feature extraction network MobileViTv3. Taking T 1 as an example, MobileViTv3 will perform feature encoding on the input image, extract rich depth semantic information, and then output three groups of features with different scales These three groups of features have different sizes. M 1 contains relatively rich spatial detail information, and M 2 and M 3 contain relatively rich semantic information.

[0055] Next, the three groups of features M 1 , M 2 and M 3 will be passed to the second stage of the encoder. In the second stage, these three groups of features are respectively input into three receptive field attention convolutional blocks RFAConv to extract receptive field spatial features, obtain more global and local information, and obtain three groups of optimized features

[0056] At the same time, T 2 will go through the same process as T 1The same processing is performed to obtain three sets of optimized features.

[0057] A3)T 1 and T 2 After passing through the lightweight multi-task weight-sharing encoder, six sets of features will be output: and The numbers in the upper right corner represent different time phases, and the numbers in the lower right corner represent which RFAConv block outputs the feature. These features will subsequently be combined with skip connections and input to different nodes of different decoders to play a role.

[0058] In the decoding stage, the embodiment has three decoders: two identical semantic segmentation decoders and one binary change detection decoder. Each semantic segmentation decoder consists of three TSEF modules and one depthwise separable residual convolution block, and these four modules are connected in series in sequence. The binary change detection decoder consists of three MTCT modules and one depthwise separable residual convolution block, and they are also connected in series in sequence.

[0059] As a preferred embodiment, the structure of the depthwise separable residual convolution block Residualblock further provided by the embodiment of the present invention is as Figure 2 shown, which is a residual block constructed based on depthwise separable convolution (DSConv):

[0060] Assume that the input feature of the depthwise separable residual convolution block is The output feature is C DS is the corresponding number of feature channels, H DS and W DS respectively represent the height and width of the corresponding feature. The input feature K will first pass through depthwise separable convolution 1, batch normalization layer 1, activation function layer 1, depthwise separable convolution 2, and batch normalization layer 2 in sequence, and then add the feature obtained through the above five steps to the input feature K and transfer it to activation function layer 2 to obtain the final output feature L.

[0061] The DSConv modules used in Depthwise Separable Convolution 1 and Depthwise Separable Convolution 2 are prior art. For the specific implementation, please refer to the relevant literature, and the present invention will not elaborate: Chollet F. Xception: Deep learning with depthwise separable convolutions[C] / / Proceedings of the IEEE conference on computer vision and pattern recognition. 2017:1251-1258.

[0062] As a preferred embodiment thereof, the processing procedure of the semantic segmentation decoder is as follows:

[0063] B11) Still taking the semantic segmentation decoder corresponding to T 1 as an example, the features and are first passed to the first TSEF module of the semantic segmentation decoder for spatio-temporal semantic enhancement and upsampling to obtain optimized features;

[0064] B12) Then, the optimized features obtained in B11) will be combined with and transmitted through the skip connection and jointly input into the second TESF module for spatio-temporal semantic enhancement and upsampling to obtain enhanced features with further increased resolution;

[0065] B13) Next, the enhanced features obtained in B12) will be combined with and transmitted through the skip connection and jointly input into the third TESF module for spatio-temporal semantic enhancement and upsampling again to obtain optimized and enhanced features;

[0066] B14) Finally, the optimized and enhanced features obtained in step B14) will be independently input into the depthwise separable residual convolution block for further decoding extraction and upsampling to obtain the semantic segmentation result N represents the number of ground object categories in the semantic segmentation task.

[0067] Meanwhile, the semantic segmentation decoder corresponding to T 2 will also perform the same operations (steps B11, B12, B13, and B14) to obtain the semantic segmentation result N represents the number of ground object categories in the semantic segmentation task. Specifically as follows:

[0068] B21) Still taking the semantic segmentation decoder corresponding to T 2 as an example, the features and First, it is passed to the first TSEF module of the semantic segmentation decoder for spatio-temporal semantic enhancement and upsampling to obtain optimized features;

[0069] B22) Then, the optimized features obtained in B21) will be combined with and transmitted through the skip connection and jointly input into the second TESF module for spatio-temporal semantic enhancement and upsampling to obtain enhanced features with further improved resolution;

[0070] B23) Next, the enhanced features obtained in B22) will be combined with and transmitted through the skip connection and jointly input into the third TESF module for spatio-temporal semantic enhancement and upsampling again to obtain optimized and enhanced features;

[0071] B24) Finally, the optimized and enhanced features obtained in step B24) will be independently input into the depthwise separable residual convolution block for further decoding extraction and upsampling to obtain the semantic segmentation result N represents the number of ground object categories in the semantic segmentation task.

[0072] As a preferred embodiment, the processing process of the binary change detection decoder is as follows:

[0073] C1) For the binary change detection decoder, the features and are first passed to the first MTCT module of this decoder for temporal integration conversion and upsampling to obtain preliminary change features;

[0074] C2) Then, the preliminary change features obtained in step C1) will be combined with and transmitted through the skip connection and jointly input into the second MTCT module for temporal integration conversion and upsampling to obtain optimized features with further improved resolution;

[0075] C3) Next, the optimized features obtained in step C2) will be combined with and transmitted through the skip connection and jointly input into the third MTCT module for temporal integration conversion and upsampling again to obtain optimized and enhanced change features;

[0076] C4) Finally, the optimized and enhanced change features obtained in step C3) will be independently input into the depthwise separable residual convolution block for further change decoding and upsampling to obtain the binary change detection result Y S1 、Y S2 and Y C The height and width of are related to T1 and T 2 is consistent, that is, it has returned to the original input size.

[0077] Finally, the two semantic segmentation results Y S1 and Y S2 are respectively multiplied by the binary change detection result Y C by masking, aiming to only retain the semantic segmentation results of the changed areas, so as to obtain the final two high-precision semantic change detection result maps, denoted as semantic change maps Y 1 、Y 2 , and Their sizes are already consistent with that of the input image.

[0078] Under the multi-decoder architecture, the network model of this embodiment can still provide efficient semantic change detection capabilities under the limited hardware conditions through multi-level integration and expression of feature information. And the network model of this embodiment is lightweight, has low computational cost, slow inference speed, and is easy to apply in actual needs; in addition, the semantic change detection results output by this network model have high consistency, and the provided semantic change information is more reliable.

[0079] The following focuses on the spatio-temporal semantic enhancement fusion module and the multi-type temporal phase change integration conversion module proposed and constructed by the present invention. These two modules cooperate with each other, effectively improving the change representation ability of deep features, and providing support for generating more accurate land cover change detection results. This innovative design provides a reliable technical basis for remote sensing image analysis.

[0080] The present invention proposes a spatio-temporal semantic enhancement fusion module (Temporal-spatial semantic enhancement fusion module, TSEF), aiming to improve the performance of the remote sensing image change detection model. The core function of the TSEF module is to receive feature inputs from two time phases, use skip connections and depthwise separable convolutions to fuse change information and generate richer feature representations. The module captures change information through absolute value difference operations on features, combined with a directional attention mechanism, thereby significantly enhancing the model's semantic expression ability for changed areas. This innovative design provides a reliable technical basis for remote sensing image analysis.

[0081] The Temporal-spatial Semantic Enhancement Fusion module (TSEF) is a key component in the decoder of Semantic Segmentation (SS). To enhance the interaction between semantic information and change information, the present invention innovatively introduces change information into the TSEF. This innovation not only significantly improves the accuracy of the SS task but also enables the network to pay more attention to the semantic features of the changed regions. In addition, the construction of the TSEF not only improves the correlation between the two tasks but also enhances the temporal correlation between the bi-temporal features. This means that the semantic change detection network model can better understand the semantic change process with the help of change information.

[0082] As a preferred embodiment thereof, the detailed structure of the proposed TSEF in the embodiment is as Figure 3 shown, taking Figure 1 the second TSEF module corresponding to the first-phase semantic segmentation decoder in

[0083] Step T1, the TSEF module receives the two-phase features optimized by the receptive field attention convolution block RFAConv obtained from the encoding end as inputs, namely the first phase and the second phase and passes them jointly through the skip connection and Figure 1 the output feature of the first TSEF module in the first-phase semantic segmentation decoder in

[0084] For the convenience of general representation, in this embodiment, the first input feature is represented by the second input feature is represented by the output feature of the previous TSEF module is represented by Q, and the specific process of the TSEF is elaborated. C 1 and C 2 are the number of channels of the corresponding features, H 1 and W 1 are the height and width of the corresponding features. It can be understood that in the semantic segmentation decoder of the second phase, Figure 3 the RF 1 in 2 represents the corresponding feature transmitted in the second phase, and RF

[0085] 2 In step T2, in the processing flow of the TSEF module, first, the first input feature RF 1And the output feature X of the upper - layer TESF module T are concatenated in the channel dimension and cross - channel fusion is performed through 1×1 convolution.

[0086] Step T3, based on the result obtained in step T2, subsequently, it is optimized using a 3×3 depth - wise separable convolution (Depthwise Separable Convolution, DSConv) to obtain a new feature representation as C 3 represents the number of channels of the output feature of the TESF module. This operation helps to extract a richer feature representation. This process can be expressed as:

[0087]

[0088] where ∪ represents concatenation in the channel dimension, represents 1×1 convolution, represents 3×3 depth - wise separable convolution.

[0089] The process of the third TSEF module in the semantic segmentation decoder is the same as that of the second TSEF module, as shown in Equation (1). However, the first TSEF module does not have the feature X from the previous module T , so there are only two inputs in the first TESF module, which are RF 1 and RF 2 , and at the same time, the channel concatenation and 1×1 convolution in step T2 are not performed. Instead, it is directly input into the 3×3 depth - wise separable convolution block of this step to obtain the first output feature Z 1 , and at this time, the number of channels C 3 , C 1 and C 2 are the same. So in the first TSEF module, this step is expressed as shown in Equation (2):

[0090]

[0091] Step T4, then, the second input feature RF 2 is introduced and an absolute - value difference operation is performed with the first output feature Z 1 to capture the change information.

[0092] Step T5, after processing the result obtained in step T4 through the Softmax function, a change weight is generated and applied to the first output feature Z 1 to obtain an enhanced semantic feature.

[0093] Step T6: The enhanced semantic features obtained in step T5 then enter the first depthwise separable residual convolution block (DS_ResBlock1) constructed based on DSConv to extract deeper second output features. Implement lightweight feature learning.

[0094] The calculation processes of steps T4 - T6 can be expressed as follows:

[0095]

[0096] where |.| represents the absolute difference operation, ρ() represents the Softmax function, and × represents feature multiplication. represents the first depthwise separable residual convolution block of TSEF.

[0097] Step T7: To enhance the spatio-temporal semantic expression ability, the TSEF module performs two-dimensional adaptive average pooling on the feature Z 2 to generate different types of channel weights by encoding in the horizontal and vertical directions. The pooling results in the horizontal and vertical directions respectively go through convolution - activation - convolution operations to form a weight-sharing twin structure, accurately capturing the correlations between channels and enhancing spatial information.

[0098] Then, the directional attention weights are obtained through the Sigmoid function and corresponding to the horizontal and vertical directions respectively.

[0099] These processes can be expressed as:

[0100]

[0101] where represents the activation operation, σ() represents the Sigmoid function, and represent the first and second weight-sharing 1×1 convolution structures respectively, and P H and P W represent the horizontal pooling and vertical pooling operations respectively.

[0102] Step T8: Finally, Z 2 is multiplied by the attention weights A H and A W simultaneously, then further optimized through the second depthwise separable residual convolution block DS_ResBlock2, and then added to Z 1 . Finally, the output feature is generated through transposed convolutional upsampling which can be expressed as:

[0103]

[0104] Among them represents a 2×2 transposed convolution for upsampling represents the second depthwise separable residual convolution block of TSEF, where × represents feature multiplication and + represents feature addition

[0105] The above content shows the working process of the TSEF module designed based on the correlation between semantic information and change information in the embodiment. In this module, the embodiment innovatively integrates the feature change information from two time phases in the SS branch and combines spatio-temporal semantic enhancement in the channel and spatial dimensions. This method significantly improves the quality of the model's feature representation of the change region, enabling it to more effectively distinguish the change region from the non-change region and providing a more consistent semantic understanding of the same type of change region. This design of the TSEF module enables the TSEF module to effectively fuse change information in the semantic branch, improve the feature expression ability of the model in the change region, and thus more accurately identify the types of land cover changes

[0106] The present invention proposes a Multi-type Temporal Change Integration Transformation module (MTCT). The MTCT module generates differential features by performing various forms of operations (such as subtraction, addition, and multiplication) on the features at different times and preserves the context information of the features. The module further extracts spatial and channel features, generates enhanced spatial and channel weights, and performs feature integration, aiming to integrate the results of various temporal difference representation methods and enhance the change feature representation in the deep features through an efficient fusion transformation mechanism, and output a more accurate binary change result

[0107] As a preferred embodiment thereof, the detailed structure of the MTCT proposed in the embodiment is as shown in Figure 4 Shown, taking Figure 1 the second MTCT module of the corresponding BCD decoder in as an example, the specific steps of the feature in this structure are as follows

[0108] Step M1, the MTCT module receives, as inputs, the two temporal features optimized by the receptive field attention convolution block RFAConv obtained at the encoding end, namely the first time phase and the second time phase and passes them through the skip connection and the output feature Figure 1 of the first MTCT module of the BCD decoder in For the convenience of general representation, in this embodiment, is represented by represents is represented by represents is represented by Denote as V, elaborate on the specific process of MTCT, C 1 and C 2 is the number of channels corresponding to the feature, H 1 and W 1 are the height and width of the corresponding feature. Here, RF 1 and RF 2 are the same as those in step T1.

[0109] Step M2, in the MTCT module, different ways of representing phase differences have different functions. RF 1 and RF 2 generate difference features through three ways: feature subtraction, addition, and multiplication, and keep the same number of channels.

[0110] Step M3, concatenate X M , RF 1 and RF 2 on the channel dimension and then pass them to a 1×1 convolutional layer to achieve cross-channel information fusion, and then process them through a 3×3 depthwise separable convolution (DSConv) layer to enhance the feature expression ability and obtain the output feature can be expressed as:

[0111]

[0112] where ∪ represents feature concatenation on the channel dimension, represents 1×1 convolution, represents 3×3 depthwise separable convolution. This can retain more context information and variation details.

[0113] The process of the third MTCT module in the BCD decoder is the same as that of the second MTCT module, as shown in Equation (6). However, the first MTCT module does not have the feature X M from the previous module, so there are only two inputs in the first MTCT module, which are RF 1 and RF 2 , so only RF 1 and RF 2 are concatenated, and the corresponding Equation (6) will become Equation (7):

[0114]

[0115] Step M4, the features generated in step M2 are further processed through two branches (spatial branch and channel branch).

[0116] In the spatial branch, each feature representation will be subjected to max-pooling and average-pooling in the spatial dimension, and the results will be averaged to enhance the generation effect of the attention weights. Figure 4 It is simplified as (AvgS + MaxS) / 2 in []. These operations can also be expressed as:

[0117]

[0118] where |.| represents taking the absolute value, - represents feature subtraction, P MaxS and P AvgS represent the max-pooling and average-pooling operations on the spatial dimension respectively, × represents feature multiplication, and + represents feature addition.

[0119] The obtained pooled features P S1 、P S2 、P S3 are concatenated and integrated through a 7×7 convolutional layer to generate diverse features, and then processed by the Sigmoid function to obtain the final spatial weights which can be expressed as:

[0120]

[0121] where σ represents the Sigmoid function, represents the 7×7 convolution, and ∪ represents the feature concatenation in the channel dimension.

[0122] The obtained spatial weights A S are multiplied by the concatenated features F C obtained in step M3 and then enter the first depthwise separable residual convolutional block DS_ResBlock1 of MTCT to obtain the enhanced features of the spatial branch

[0123]

[0124] where represents the first depthwise separable residual convolutional block of MTCT, and × represents feature multiplication.

[0125] In the channel branch, each feature representation obtained in step M2 extracts channel features through adaptive max-pooling and average-pooling, and the results are averaged to enhance the ability to represent the significance of each channel. Figure 4 It is simplified as (AvgC + MaxC) / 2 in []. These operations can also be expressed as:

[0126]

[0127]

[0128] where P MaxCand P AvgC respectively represent the adaptive max pooling and average pooling operations on the channel dimension, where |.| represents taking the absolute value, - represents feature subtraction, × represents feature multiplication, and + represents feature addition.

[0129] The obtained channel features P C1 、P C2 、P C3 are added and then passed through the Sigmoid function to obtain the channel weights

[0130] A C =σ(P C1 +P C2 +P C3 ) (16)

[0131] The channel weights A C are multiplied by the feature F C , and then enter the second depthwise separable residual convolution block DS_ResBlock2 of MTCT to obtain the enhanced features of the channel branch

[0132]

[0133] where represents the second depthwise separable residual convolution block of MTCT, and × represents feature multiplication.

[0134] Step M5, the enhanced features E S and E C generated by the spatial branch and the channel branch are added to perform feature integration and information supplementation, and then optimized feature extraction is performed through the third depthwise separable residual convolution block DS_ResBlock3 of MTCT, and a residual connection is made with the cascaded feature F C .

[0135] Step M6, finally, the result obtained in Step M5 is processed by transposed convolution to improve the spatial resolution and feature refinement, and the final transformed decoded feature

[0136]

[0137] where represents the third depthwise separable residual convolution block of MTCT, + represents feature addition, represents a 2×2 transposed convolution. The output feature Y M of this MTCT module has the same size as the output feature Y T in Step T8.

[0138] The above content provides the MTCT processing flow proposed to enhance the ability to identify change decoding branches. First, through the blending of various types of temporal changes, multiple differential feature representations are generated in the spatial and channel dimensions. Then, these features are integrated into representative features using cascading and addition methods, and are input into the corresponding attention weight generation component for blending transformation, finally outputting features optimized from different perspectives. Therefore, the MTCT module can effectively enhance the change representation ability in deep features, providing more accurate change region information for generating the final SCD result.

[0139] III. Optimizing the network model training using a multi-task loss function

[0140] In a deep learning network, the backpropagation of the loss function is used to optimize the network weights. However, as the complexity of the network structure increases, the problem of gradient disappearance may occur during parameter update, resulting in unstable optimization and reducing the change detection effect. To solve this problem, the present invention proposes a method that combines three loss functions to jointly supervise and optimize network training, specifically for semantic segmentation, binary change detection, and semantic consistency. First, for the semantic segmentation task, a multi-class cross-entropy loss function is used to guide network training. Second, the binary cross-entropy loss function is used to supervise the change detection branch for the binary change detection task. In addition, a semantic consistency loss function is introduced to ensure consistent semantic predictions for the two temporal branches in the non-change region. These loss functions are combined to form a total loss function, which systematically optimizes the performance of the network model by balancing the influence of different loss functions, thereby enhancing its robustness and efficiency in complex change detection tasks. This method enhances the performance of the model from multiple perspectives and improves the accuracy of change detection.

[0141] As a preferred embodiment thereof, the multi-task loss function of the embodiment is specifically constructed and implemented as follows:

[0142] For the semantic segmentation task (SS task), the embodiment uses a multi-class cross-entropy loss function to guide the training of the network, and its definition is as follows:

[0143]

[0144] where N represents the total number of pixels, C represents the number of classes, is the true probability that pixel i belongs to class C, and is the probability that pixel i is predicted to be class C, t represents the temporal phase, represents the multi-class cross-entropy loss at temporal phase t.

[0145] In this semantic change detection framework, since there are two temporal branches, and each branch corresponds to a semantic segmentation task, the total semantic segmentation loss L segis defined as the sum of two phase losses:

[0146]

[0147] where and represent the losses of the two phase semantic segmentation branches respectively.

[0148] For the binary change detection task, the embodiment adopts the binary cross-entropy loss function L bcd to supervise the change detection branch, and its definition is:

[0149]

[0150] where N is the total number of pixels, and y i indicates whether pixel i has changed (1 for changed, 0 for unchanged), and p i is the probability predicted by the model that this pixel has changed.

[0151] To further improve the stability and accuracy of the model, the embodiment uses a semantic consistency loss function (Semantic Consistency Loss, SCLoss), denoted as L sc . The purpose of this loss function is to reward the cases where the two phase branches have consistent semantic predictions in the non-changing regions, so as to ensure that the semantic information of images at different times remains consistent. Its definition is as follows:

[0152]

[0153] where x 1 and x 2 refer to the pixel vectors of the two semantic segmentation results respectively, and y i indicates whether pixel i has changed (1 for changed, 0 for unchanged).

[0154] To balance the influence of each loss function during the training process, the embodiment combines these loss functions into a total loss function L total :

[0155]

[0156] By combining the semantic segmentation loss, the binary change detection loss and the semantic consistency loss, this total loss function systematically optimizes the overall performance of the network model from multiple task objectives, making the model more robust and efficient when dealing with complex change detection tasks.

[0157] IV. Inputting dual-temporal remote sensing images for semantic change detection

[0158] The semantic change detection model obtained through training and its optimal weight parameters ensure the performance stability and detection accuracy in subsequent applications. In actual operation, users only need to input the dual-temporal remote sensing image data to be detected into the pre-trained model, and high-precision semantic change detection maps can be automatically generated. These maps visually display the areas where changes have occurred in the remote sensing images, accurately locating the change positions and their category information. The generated change detection results can not only be saved but also be used for downstream analysis tasks such as urban planning, land use management, ecological monitoring, and natural resource management. For example, in urban planning, the change detection maps help identify urban expansion, changes in construction land, and reduction of green spaces, thus formulating scientific and reasonable development plans. In natural resource monitoring, the detection results effectively track dynamic changes such as deforestation and wetland changes, supporting environmental protection and resource management.

[0159] In the previous steps, the semantic change detection model obtained through training and the corresponding optimal weight parameters can ensure the performance stability and detection accuracy of the model during subsequent use. In practical applications, when semantic change detection is required, users only need to input the dual-temporal remote sensing image data to be detected into the pre-trained network model, and high-precision semantic change detection maps can be automatically generated. These change detection maps can visually display the areas where changes have occurred in the remote sensing images, accurately locating the change positions and category information.

[0160] In addition, the generated semantic change detection results can not only be directly saved but also be used in downstream analysis tasks such as urban planning, land use management, ecological environment monitoring, and natural resource management. Specifically, in urban planning, the semantic change detection maps can help planners identify issues such as urban expansion areas, changes in construction land, and reduction of green spaces, thus formulating more scientific and reasonable development plans. In the field of natural resource monitoring, the detection results can effectively track the dynamic changes of natural resources such as deforestation, wetland changes, and water body distribution, supporting environmental protection decision-making and the optimization of resource management strategies.

[0161] On this basis, the detection model of the present invention also has high efficiency and scalability. It can not only adapt to different types of remote sensing image data but also fine-tune the network model according to requirements to adapt to other change detection tasks. This flexibility makes the model widely applicable in various application scenarios, such as emergency disaster response, agricultural monitoring, and infrastructure management. Generally speaking, the present invention provides an efficient, accurate, and scalable solution in the field of semantic change detection, which can significantly improve the efficiency of remote sensing image analysis and application.

[0162] In specific implementation, the method proposed by the technical solution of the present invention can be automatically run by those skilled in the art using computer software technology. The system device for implementing the method, such as a computer-readable storage medium storing the corresponding computer program of the technical solution of the present invention and a computer device including the operation of the corresponding computer program, should also be within the protection scope of the present invention.

[0163] Next, a lightweight semantic enhancement and change integration remote sensing image semantic change detection device provided by the present invention will be described. The lightweight semantic enhancement and change integration remote sensing image semantic change detection device described below can be correspondingly referred to the lightweight semantic enhancement and change integration remote sensing image semantic change detection method described above.

[0164] In another embodiment, the present invention provides an electronic device, which may include: a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus. The processor can call the logical instructions in the memory to execute the lightweight semantic enhancement and change integration remote sensing image semantic change detection method.

[0165] In addition, when the logical instructions in the above-mentioned memory are implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the essence of the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0166] In another embodiment, the present invention further provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the lightweight semantic enhancement and change integration remote sensing image semantic change detection method provided by the above-mentioned methods.

[0167] In another embodiment, the present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements a lightweight semantic enhancement and change integration remote sensing image semantic change detection method provided by the above-mentioned various methods.

[0168] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0169] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0170] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or equivalently replace some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A lightweight semantic enhancement and change integration remote sensing image semantic change detection method, comprising: Collect dual-temporal remote sensing images and build a sample library; Build a multi-task semantic change detection network model and optimize the network model training based on the sample library; The encoder part of the multi-task semantic change detection network model adopts a lightweight multi-task weight sharing encoder, which supports semantic segmentation and binary change detection tasks at the same time, including an encoder branch corresponding to the first phase and an encoder branch corresponding to the second phase; the decoder part of the multi-task semantic change detection network model includes a semantic segmentation decoder corresponding to the first phase and a semantic segmentation decoder corresponding to the second phase, and a binary change detection decoder; The first phase image and the second phase image of the input encoder are processed to generate feature maps of different resolutions, and then transmitted to the decoder through skip connections to achieve multi-level feature integration, and finally generate a binary change map and two semantic segmentation maps, and extract the semantic change map in combination with mask operation; Input the dual-temporal remote sensing image into the trained multi-task semantic change detection network model for semantic change detection.

2. The method for detecting semantic changes in remote sensing images by integrating lightweight semantic enhancement and changes according to claim 1 is characterized by: In the lightweight multi-task weight-sharing encoder, the encoder branch corresponding to the first phase and the encoder branch corresponding to the second phase have the same structure and share weights, and each includes two stages. In the first stage, a lightweight feature extraction network is used to perform feature encoding on the input image, extract deep semantic information, and then output three groups of features of different scales. In the second stage, the three groups of features are respectively input into three receptive field attention convolution blocks to extract receptive field spatial features to obtain three groups of optimized features.

3. The method for detecting semantic changes in remote sensing images by integrating lightweight semantic enhancement and changes according to claim 2 is characterized by: Each semantic segmentation decoder consists of three spatiotemporal semantic enhancement fusion modules connected in series and a deep separable residual convolution block; the binary change detection decoder consists of three multi-type temporal change integrated transformation modules connected in series and a deep separable residual convolution block.

4. The method for detecting semantic changes in remote sensing images by integrating lightweight semantic enhancement and changes according to claim 3 is characterized by: A pair of dual-phase RGB three-band optical remote sensing images T1 and T2 are respectively passed through the corresponding encoder branches of the lightweight multi-task weight sharing encoder, and the output features of different scales corresponding to T1 are and features of different scales corresponding to T2 and And passed to the decoder part through skip connection.

5. The method for detecting semantic changes in remote sensing images by integrating lightweight semantic enhancement and changes according to claim 4 is characterized by: The spatiotemporal semantic enhancement fusion module receives feature inputs from two time phases, utilizes skip connections and depthwise separable convolutions to fuse change information and generate richer feature representations; this module captures change information by performing absolute value difference operations on features, and combines the directional attention mechanism to significantly enhance the model's ability to express semantics for change areas.

6. The method for detecting semantic changes in remote sensing images by integrating lightweight semantic enhancement and changes according to claim 4, characterized in that: The multi-type time-phase change integrated conversion module generates difference features by performing various forms of operations on features at different times and maintains the contextual information of the features; further extracts spatial and channel features, generates enhanced spatial and channel weights, and integrates features, and enhances the change feature representation in deep features through a fusion transformation mechanism, outputting more accurate binary change results.

7. The method for detecting semantic changes in remote sensing images by integrating lightweight semantic enhancement and changes according to claim 4, characterized in that: The depthwise separable residual convolution block is a residual block constructed based on the depthwise separable convolution DSConv.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the lightweight semantic enhancement and change integration remote sensing image semantic change detection method as described in any one of claims 1 to 7 is implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for detecting semantic changes in remote sensing images with lightweight semantic enhancement and change integration is implemented as claimed in any one of claims 1 to 7.

10. A computer program product, comprising a computer program, characterized in that: When the computer program is executed by a processor, the method for detecting semantic changes in remote sensing images with lightweight semantic enhancement and change integration is implemented as claimed in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Consistency loss guided building semi-supervised change detection method and device

    CN116343033A

  • Remote sensing image semantic change detection method based on multi-task twin network guided by enhanced change information

    CN118196622A

  • Network model for dual-temporal remote sensing image semantic change detection

    CN118397480A

Cited By

  • Double-flow remote sensing image change detection method fused with Mmba enhancement

    CN120298906A

  • Three-dimensional point cloud segmentation method and system for assembled integral steel structure based on BIM (Building Information Modeling)

    CN120931808A

  • Image segmentation method and device, electronic equipment and storage medium

    CN121074386A

  • Remote sensing semantic change detection method based on local detail continuity keeping

    CN121121528A

  • A method and system for detecting semantic changes in remote sensing images

    CN122574684A