A lake area semantic change detection method based on remote sensing data

By using multi-source remote sensing data fusion and water level-driven feature alignment technology, the accuracy and robustness issues of change detection in the regulation and storage lake area have been solved, and efficient and accurate semantic interpretation of land cover changes in the lake area has been achieved, which is applicable to practical operations such as ecological monitoring and flood control and disaster reduction.

CN122200345APending Publication Date: 2026-06-12CENTRAL SOUTH UNIVERSITY OF FORESTRY AND TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CENTRAL SOUTH UNIVERSITY OF FORESTRY AND TECHNOLOGY
Filing Date
2026-03-13
Publication Date
2026-06-12

AI Technical Summary

Technical Problem

Existing technologies struggle to achieve accurate, robust, and semantically interpretable change detection in dynamic scenarios within lake regulation areas. They cannot effectively distinguish between natural hydrological changes and changes caused by human activities. Furthermore, multi-source data fusion methods fail to fully consider the cloudy and rainy climate of the lake area and the noise characteristics of SAR data, resulting in insufficient detection accuracy.

Method used

A multi-source remote sensing data fusion method is adopted, combining optical remote sensing images and SAR images. Through cross-temporal feature alignment and multi-scale feature extraction driven by water level information, semantic association modeling is carried out by combining attention mechanism to suppress noise and output the location information and semantic category of the changed area.

Benefits of technology

It improves the accuracy and robustness of change detection, can simultaneously capture features of large-scale natural features and small-scale artificial features, enhances the interpretability, adaptability and efficiency of the results, and is suitable for large-scale monitoring and edge device deployment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122200345A_ABST
    Figure CN122200345A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of remote sensing information processing and artificial intelligence, and discloses a lake area semantic change detection method based on remote sensing data; the method aims at the technical difficulties of spatial displacement of ground objects in the regulation and storage lake area caused by violent water level fluctuation, insufficient utilization of multi-source data, and missed detection of small-scale ground objects, acquires double-time optical and SAR images, carries out pretreatment and feature fusion, and constructs multi-channel input features; a multi-scale feature extraction network with overlapping blocks is adopted to capture different scale ground object information; water level information is introduced to drive cross-time feature space correction to inhibit pseudo changes caused by water level changes; noise self-adaptive suppression and feature compensation are carried out on the multi-modal features; finally, cross-time semantic correlation modeling and decoding are carried out to synchronously output the change area and the semantic categories before and after the change area. The application effectively improves the accuracy, robustness and interpretability of change detection under the dynamic lake area environment, and is suitable for ecological monitoring, flood control and disaster reduction and other practical businesses.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of remote sensing information processing and artificial intelligence technology, specifically to a method for detecting semantic changes in lake areas based on remote sensing data. Background Technology

[0002] Remote sensing change detection technology is an important tool for land cover monitoring, ecological environment assessment, and natural resource supervision. With the increasing abundance of remote sensing data sources, methods combining multi-temporal imagery for change detection have been widely applied. Traditional methods are mainly divided into two categories: one is binary change detection, which can only identify whether changes have occurred on the land surface; the other is semantic segmentation based on single-temporal imagery, which can only provide information on land cover categories at a specific moment.

[0003] However, in applications involving large-scale regulating lakes, existing technologies suffer from significant seasonal fluctuations in lake water levels, leading to large-scale, regular shifts in the spatial locations of features such as water boundaries, mudflats, and wetlands. Current change detection methods typically employ fixed feature matching or location alignment strategies, which are ill-suited to this dynamic spatial change. This can easily result in misinterpreting simple location shifts as genuine semantic changes in land features, producing numerous false change results. Furthermore, to overcome the limitations of single data sources, fusing multi-source remote sensing data, such as optical and synthetic aperture radar (SAR), has become a trend. However, existing multi-source data fusion methods often employ simple channel stitching or fixed weight strategies, failing to adequately consider the lack of optical data in cloudy and rainy weather conditions in lake areas, as well as the inherent speckle noise characteristics of SAR data. This results in insufficient quality and robustness of the fused features, affecting the accuracy of subsequent change detection. In addition, lake areas have complex land feature types, including large areas of water and wetlands, as well as numerous small-scale ditches, embankments, and scattered aquaculture ponds—key artificial features. Existing semantic change detection models based on deep learning often struggle to capture both broad contextual information and preserve detailed features of small features, resulting in a high rate of missed detections of important features at small scales.

[0004] Therefore, existing technologies are insufficient to achieve accurate, robust, and semantically interpretable change detection in dynamic scenarios such as reservoirs, and cannot effectively distinguish between natural hydrological changes and changes caused by human activities, thus limiting their application in practical operations such as ecological monitoring and flood control and disaster reduction. To address this, a semantic change detection method for lake areas based on remote sensing data is proposed. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention provides a method for detecting semantic changes in lake areas based on remote sensing data, thereby resolving the problems mentioned in the background.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a method for detecting semantic changes in lake areas based on remote sensing data, comprising the following steps: S1. Data Acquisition and Preprocessing: Acquire multi-source remote sensing image data of the target lake area in the first and second time phases. The multi-source remote sensing image data includes at least optical remote sensing images and synthetic aperture radar (SAR) images. Perform preprocessing on the multi-source remote sensing image data. The preprocessing includes radiometric correction, geometric correction, and spatial registration to obtain a spatially aligned multi-source remote sensing image dataset. S2. Multi-source feature construction: Based on the preprocessed optical remote sensing image, calculate spectral index features including at least water index and vegetation index, and combine the spectral index features with the original optical image band data to form first feature data; based on the preprocessed SAR image, extract backscattering features to form second feature data; fuse the first feature data and the second feature data in the channel dimension to construct multi-channel input features; S3. Multi-scale feature extraction: The multi-channel input features are divided into blocks using a convolution operation with overlapping windows to extract multi-scale feature maps containing different spatial scales; S4. Water level-driven cross-temporal feature alignment: Obtain lake water level information corresponding to the first and second time phases; determine spatial location offset parameters based on the water level information, and use the offset parameters to spatially correct the multi-scale feature maps of the first and second time phases to achieve cross-temporal feature alignment; S5. Multimodal feature optimization and noise suppression: The dual-temporal multi-scale feature map aligned in step S4 is subjected to noise suppression processing. The noise suppression processing includes: for areas in the optical features affected by cloud and fog shadows, feature compensation is performed using SAR features; for speckle noise in the SAR features, adaptive filtering is performed based on the spatial structure characteristics of ground features. S6. Cross-temporal semantic reasoning and change discrimination: Based on the optimized dual-temporal features, cross-temporal semantic association modeling is performed to generate change features; the change features are decoded, and the location information of the changed area, the semantic category of the land cover before the change, and the semantic category of the land cover after the change are output synchronously; S7. Output the semantic change detection results.

[0007] According to claim 1, a method for detecting semantic changes in lake areas based on remote sensing data is characterized in that, in step S1, the multi-source remote sensing image data further includes auxiliary data that is in the same temporal phase as the optical remote sensing image and the SAR image, and the auxiliary data includes at least one of hydrological data, ecological survey data or land use vector data.

[0008] Preferably, in step S2, the spectral index features further include a building index or a bare soil index for characterizing artificial features.

[0009] Preferably, step S3 specifically involves processing the multi-channel input features through a multi-layer feature extraction network, wherein the multi-layer feature extraction network includes convolutional layers with overlapping sliding windows and outputs multi-level feature maps with resolutions ranging from high to low, thereby forming the multi-scale feature maps.

[0010] Preferably, in step S4, determining the spatial offset parameter based on the water level information specifically includes: based on the correspondence model between historical water level data and the spatial offset of land feature boundaries, calculating the estimated spatial offset as the offset parameter according to the water level difference between the current time phases.

[0011] Preferably, in step S5, the adaptive filtering of speckle noise in SAR features specifically involves: using large window filtering for water areas and small window filtering for linear features or edge areas.

[0012] Preferably, in step S6, the cross-temporal semantic association modeling is implemented using a semantic reasoning module based on an attention mechanism. The semantic reasoning module integrates a window self-attention component for capturing long-range dependencies, a channel attention component for enhancing the channel weights of the spectral index, and a spatial attention component for focusing on change-sensitive regions.

[0013] Preferably, in step S6, the decoding of the change features is implemented using a decoder network with a jump-connection structure, so as to fuse deep semantic features with shallow detail features and output a high-resolution semantic map and change map.

[0014] Preferably, the method further includes a post-processing step for the preliminary discrimination results, wherein the post-processing includes logical verification and correction of the location information and / or semantic category of the changed area based on prior knowledge of land cover categories or spatial context relationships.

[0015] Compared with the prior art, the present invention has the following beneficial effects: 1. This invention acquires and utilizes lake water level information to drive spatial correction of cross-temporal features, effectively overcoming the spatial offset of ground features caused by drastic seasonal fluctuations in lake water levels, reducing the probability of misjudging such positional changes as semantic changes, and improving the accuracy of change detection in dynamic hydrological environments.

[0016] 2. This invention employs a multi-source feature construction method that integrates optical images, spectral indices, and SAR backscattering features, combined with targeted noise suppression processing. This enables the use of SAR data to compensate for the loss of optical information in cloudy weather, suppress speckle noise, and improve data utilization and robustness of change detection.

[0017] 3. By employing a multi-scale feature extraction network with overlapping windows and combining it with semantic association modeling based on attention mechanisms, this invention can simultaneously capture the detailed features of small-scale artificial features and the contextual information of large-scale natural features, thereby improving the detection capability of key small-scale features such as ditches and embankments.

[0018] 4. This invention adopts an end-to-end discrimination and decoding architecture that can simultaneously output the changed location and the semantic category before and after the change, directly providing semantic-level change results including the relationship of land cover type transfer, enhancing the interpretability of the results and better supporting ecological assessment and regulatory decisions.

[0019] 5. By combining specific training strategies and network structure design, this invention improves the efficiency and adaptability of the method while ensuring detection accuracy, making it easier to transform into practical engineering application scenarios such as large-scale monitoring and edge device deployment.

[0020] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures pointed out in the description, claims and drawings. Attached Figure Description

[0021] Figure 1 This is a flowchart of the lake area semantic change detection method based on remote sensing data according to the present invention; Figure 2 This is a flowchart of the data preprocessing and multi-source feature construction process of the present invention; Figure 3 This is a schematic diagram of the multi-scale feature extraction and water level-driven alignment of the present invention; Figure 4 This is a structural diagram of the multimodal feature optimization and semantic reasoning module of the present invention; Figure 5 This is a flowchart of the output and post-processing of the results of this invention. Detailed Implementation

[0022] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0023] Please see Figures 1-5 This invention provides a method for detecting semantic changes in lake areas based on remote sensing data.

[0024] S1. Data Acquisition and Refined Preprocessing This step forms the basis for all subsequent analyses, and its goal is to obtain high-quality, spatiotemporally consistent multi-source data. A large regulating lake area with typical seasonal water level fluctuations is selected as the monitoring target.

[0025] 1. Data Collection: Optical remote sensing data: acquiring the target lake area in the first time phase Second phase Imagery was used. Priority was given to Level-2 products from the domestically produced Gaofen-2 (GF-2) satellite, which has a panchromatic spatial resolution of 0.8 meters and a multispectral resolution of 3.2 meters. These images have undergone radiometric calibration, atmospheric correction, and geometric fine correction, directly providing surface reflectance data. Simultaneously, Level-2A products from Sentinel-2, with a spatial resolution of 10-20 meters and a short revisit period, were used to supplement temporal information.

[0026] SAR remote sensing data: Sentinel-1 wide-swath (IW) ground range detection (GRD) products corresponding to optical imagery were acquired, including VV and VH dual-polarization channels, with a spatial resolution of approximately 10 meters. This data was used to penetrate cloud cover and compensate for data loss in optical imagery under adverse weather conditions.

[0027] Supplementary data: obtained from relevant hydrological departments and Average water level observations in the lake area during the time period and Simultaneously, we acquired wetland ecological survey reports for the area (including the distribution of major vegetation types and the range of migratory bird habitats), historical and current land use / land cover vector maps (with particular attention to the ownership or marked boundaries of legal and illegal aquaculture ponds, dikes, and other artificial structures), and digital elevation models (DEMs) as prior knowledge and validation data. When selecting data, we controlled the cloud cover ratio of optical imagery to be below 20% and ensured that the interval between bi-temporal imagery was within 3-6 months to focus on seasonal changes and the impact of human activities.

[0028] 2. Data preprocessing: Basic radiometric and geometric corrections: For GF-2 data, its Level-2 product has been confirmed to meet application requirements. For Sentinel-2 data, its Level-2A product has undergone atmospheric correction using the Sen2Cor processor, outputting surface reflectance. ,in Represents the band index. For Sentinel-1 GRD data, the following processing chain is performed using ESA SNAP software: thermal noise removal → radiometric calibration (output backscattering coefficients). (Unit: dB) → Terrain correction (simulation using a 30-meter resolution DEM to eliminate geometric and radiometric distortions caused by terrain). Orthorectify all imagery to a common geodetic coordinate system (e.g., WGS84 UTM).

[0029] High-precision spatial registration: with High-resolution GF-2 imagery was used as the geometric reference. Subpixel-level registration was achieved using a strategy of "feature point matching + model optimization". a. Feature extraction: The Scale Invariant Feature Transform (SIFT) algorithm is used to extract features from both the reference image and the image to be registered. Optical imaging, and Key points and their descriptors are extracted from SAR images.

[0030] b. Feature matching and refinement: Initial matching is performed using the nearest neighbor distance ratio (NNDR), followed by iterative estimation of the optimal affine transformation model using the Random Sample Consensus (RANSAC) algorithm. And remove outliers (mismatched points).

[0031] c. Image resampling: using the estimated transformation matrix Bilinear interpolation is used to resample the images to be registered, ensuring strict alignment with the reference image. After registration, the root mean square error (RMSE) of the spatial positional error of corresponding pixels among all images must be less than 0.5 pixels.

[0032] Cloud and Shadow Detection and Processing: For optical images, binary masks for clouds and cloud shadows are generated using Sentinel-2's Scene Classification Layer (SCL) or the s2cloudless algorithm. For pixel regions marked as clouds / shadows, no padding is performed. As a preferred implementation, a feature completion method based on Generative Adversarial Networks (GANs) is employed. This method trains a completion model using texture features of similar land features (such as wetlands) in cloudless phase images of the same region to restore the image features of cloud-shadowed areas. Furthermore, in subsequent multi-source feature fusion and optimization steps, features from SAR data will be used to collaboratively supplement and verify the information for this region.

[0033] Adaptive SAR speckle noise filtering: The registered SAR image SS is filtered. An improved Lee filter is used, and the filtering window is adaptively selected based on the land cover type. , in, This is the filtered value. It is a filter window The mean of the inner pixel values, This is the original value of the center pixel. Weight From the variance within the window and the overall mean Decide, , This is the equivalent number of views. The key adaptive strategy lies in the window. Size: In homogeneous areas (such as open water bodies, determined by the VV / VH ratio threshold), a larger window (such as 9×99×9) is used to fully suppress speckle; in edge and detail areas (such as dikes and ditches, initially determined by edge detection), a smaller window (such as 3×33×3) is used to protect structural information.

[0034] Study area cropping and resampling: Based on the lake basin boundary vector, a uniform region is cropped from all processed images. All data layers are then resampled to the same spatial resolution. (For example, uniformly set to 1 meter), resulting in dimensions of Standardized data (height × width): Optical reflectivity cube SAR backscattering data ,in , and These are the optical and SAR band numbers, respectively.

[0035] S2. Multi-source feature construction This step aims to fuse spectral, index, and radar information to construct a robust feature input.

[0036] Calculation of spectral indices: based on preprocessed optical data Calculate spectral indices sensitive to land features in the lake area and generate index feature maps. This mainly includes: Normalized Difference Water Index (NDWI): .

[0037] Normalized Difference Vegetation Index (NDVI): .

[0038] Normalized Difference Building Index (NDBI): This index is used to enhance information on man-made structures and bare land.

[0039] Feature channel fusion: combining the original optical bands Spectral index characteristics and filtered SAR data Concatenate the data along the channel dimension to form the final multi-channel input feature tensor. , Among them, the total number of channels This yields the dual-temporal input characteristics. and .

[0040] S3. Multi-scale feature extraction based on overlapping blocks This step uses a carefully designed encoder network to extract multi-level feature representations from the fused features.

[0041] Overlapping block embedding: The first layer of the encoder uses a stride of... The kernel size is The convolutional layer satisfies In this embodiment, the following is adopted: The convolutions allow for inter-convolution windows to be convolved, which results in a certain distance between adjacent convolution windows. The overlap of pixels is minimal. The direct effect of this design is to ensure that the local contextual information of small-scale features, such as ditches and embankments only a few pixels wide, can be completely captured by the same or adjacent receptive fields, avoiding the problem that small feature details might be fragmented in the early stages of encoding, which could be caused by traditional non-overlapping patching (such as Patch Embedding in Vision Transformer). This layer will input... Mapped to the initial deep feature map .

[0042] Constructing a multi-scale feature pyramid: The encoder core is a hierarchical convolutional neural network (e.g., a modification of the EfficientNet or ResNet architecture). The network contains multiple stages, each consisting of several convolutional layers, normalization layers, and activation functions, with downsampling between stages (e.g., using convolutions with a stride of 2). The network ultimately outputs feature maps at four different spatial scales, forming a feature pyramid: , in, The space size is approximately It retains high-resolution detail information; The space size is approximately It possesses a broad receptive field and contains rich semantic and global contextual information. This step provides a complete feature representation from local details to global semantics for subsequent processing, enabling the model to focus on both subtle changes at the dike edge and to understand the large-scale correlation between lake expansion and wetland inundation.

[0043] S4. Water level-driven cross-temporal feature alignment This step is the core of solving the problem of pseudo-changes in the lake area, using water level information to correct the spatial positional offset of ground features.

[0044] Constructing a water level-offset relationship model: Before model training or application, an empirical model of water level changes and horizontal displacement of typical land cover types (mainly land-water boundaries) needs to be established. By analyzing long-term water level data (WLWL) and synchronized remote sensing imagery, a function or lookup table can be fitted. For the current input, the water level difference is calculated. By querying the model, the estimated global spatial offset vector at the original image resolution is obtained. ,in and These are the offset pixels in the column direction (east) and the row direction (north), respectively.

[0045] Feature map space transformation correction: Image-level offset Adapted to feature maps of different scales Above. For the l-th layer feature map, its spatial scaling factor relative to the original image is... (For example, The offset on the feature map of that layer is... Using the Spatial Transformer Network (STN) module or direct bilinear interpolation, for... The feature map of the time phase is translated to achieve the same result as... Alignment of phase feature maps: , in, Represents a spatial transformation function. This is the corrected feature map. This step compensates for the overall positional movement of natural features (water bodies, mudflats, wetlands) caused by the rise and fall of water levels at the feature level, so that the features of "same object in different position" are re-aligned in space, which greatly reduces the probability of misjudging this positional shift as a semantic change in the future.

[0046] S5. Multimodal Feature Optimization and Noise Suppression This step, based on feature alignment, further cleans up the features and suppresses residual noise.

[0047] Preliminary calculation and optimization of feature differences: Design a feature optimization module to suppress noise and enhance the real variation signal; as an efficient implementation, this module can be designed as a dual-branch structure to process two paths simultaneously: 1. Differential Enhancement Path: Calculate the Euclidean distance between the aligned bi-temporal features to obtain the primary differential features. .

[0048] 2. Consistency-weighted path: Calculate the cosine similarity of bi-temporal features. And it is mapped to a consistency weight graph through a small neural network or a sigmoid function. In stable regions where no change has occurred, the feature similarity is high and the weight is close to 1; in noisy or pseudo-change regions, the weight is low.

[0049] Adaptive noise suppression integration: Integrating prior information from step S1 (such as cloud mask) The features are then integrated. For regions corresponding to cloud shadows in the optical feature channels, aligned SAR features are used for weighted supplementation. Finally, the optimized difference features are fused using a gating mechanism. ,in This represents an adaptive filtering operation that incorporates land cover type (water body / edge), echoing the SAR filtering principle in S1. This module can automatically attenuate non-semantic differential responses caused by cloud shadows, residual speckles, and seasonal phenological changes (such as vegetation yellowing), while preserving and potentially enhancing signals of real land cover transformations (such as farmland becoming aquaculture ponds).

[0050] S6. Cross-temporal semantic reasoning and change detection This step is the core of information fusion and decision-making, responsible for inferring semantic changes from optimized features.

[0051] Multi-scale attention fusion and semantic reasoning: Constructing a semantic reasoning module. This module receives optimized differential features across all scales. and the original alignment features of the two temporal phases Internally, it performs the following operations: 1. Feature reshaping: Adjusting different features through upsampling or convolution. Features are aligned to the same intermediate size and then spliced ​​together.

[0052] 2. Triple attention mechanism: Channel Attention (CAM-L): Calculates the channel weight vector This allows the model to automatically focus on relevant index channels such as NDWI, NDVI, and NDBI, as well as SAR channels that are sensitive to texture. The formula can be expressed as... ,in is the feature tensor after feature reshaping and concatenation in step 2, and GAP is global average pooling.

[0053] Spatial Attention (SAM-L): Generating Spatial Weight Maps This focuses on areas sensitive to change, such as boundaries and abrupt texture changes. It can be obtained by convolving the feature map with a small convolution kernel followed by a sigmoid function.

[0054] Long-range contextual attention (CTM-L): Employs window self-attention (such as WindowAttention in Swing Transformer) or non-local network mechanisms to compute the relationships between feature points within a larger window (such as a region on a 512x512 feature map) to capture long-range dependencies.

[0055] One specific implementation of this module, which integrates the aforementioned channel, spatial, and long-range context attention components, can be called SR-Block-L.

[0056] 3. Feature Fusion: After normalizing the three attention weights mentioned above, the reorganized features are weighted and fused to generate a unified feature representation rich in multi-scale context and semantic information of change. .

[0057] Dual-task joint decoding output: Employs a decoder with a dense skip connection structure (such as UNet++). The decoder uses high-level semantic features. To guide, and through a jump connection structure, the early stages of the encoder (such as...) The high-resolution details are gradually incorporated and upsampled to the original image size. .

[0058] The decoder ends with two output heads connected in parallel: Dual-temporal semantic segmentation head: Each head consists of a 1×1 convolutional layer followed by a softmax layer, outputting pixel-level class probability maps respectively. , where NN is the total number of land cover categories.

[0059] Change detection head: Also a 1×1 convolutional layer followed by a sigmoid function, outputting a change probability map. .

[0060] The jump-connection structure combines the high-level semantic information of deep features with the fine spatial information of shallow features, ensuring that the final output semantic map and change map have both accurate category judgment and clear target boundaries, especially for small-scale features.

[0061] 3. Model Training Strategy To enable the aforementioned network model (including the encoder, feature optimization module, semantic reasoning module, and decoder) to detect semantic changes, it needs to be trained using labeled data. Considering the scarcity and noise complexity of labeled data in the lake area, a two-stage semi-supervised training scheme is adopted, which effectively combines the advantages of unlabeled data and a limited amount of labeled data.

[0062] Phase 1: Unsupervised pre-training (noise suppression and difference perception learning) The goal of this stage is to use a large number of unlabeled bi-temporal image pairs to enable the model to initially learn to distinguish between noise and real changes without the need for manual annotation.

[0063] 1. Training data: collection For example, 3000 pairs of dual-temporal images covering the lake area ( These data do not require any manually labeled semantic or change tags.

[0064] 2. Training setup: Unfreeze and train only the encoder (the multi-scale feature extraction network described in step S3) and the feature optimization module (the module described in step S5), and freeze the semantic reasoning module and the decoder.

[0065] 3. Loss Function: A noise suppression and difference enhancement loss function is designed. This loss function guides the model to learn two tasks: Image reconstruction: requires the model to be able to reconstruct a clear feature representation from features that may be contaminated by noise.

[0066] Difference amplification: The model is encouraged to amplify the differences between two temporal features in regions of real change (pseudo-labels can be generated through a simple pixel difference thresholding method), while suppressing differences in regions of no change or noise.

[0067] One feasible approach is to design a loss function centered on feature reconstruction and region consistency constraints, guiding the model to automatically filter out noise rather than relying on unreliable pseudo-labels. The loss function can be designed as follows: : , in, As input features, The output features are those processed by the encoder and feature optimization module. This is the reconstruction loss (e.g., mean squared error). This is a diagram showing the differences in characteristics between two time phases. For example, a pseudo-change mask generated by unsupervised methods (such as pixel-level comparison combined with morphological filtering) can be created by applying bi-temporal input features. and The result is obtained by summing the absolute differences of each channel, and then removing noise through threshold segmentation (such as the OTSU algorithm) and morphological opening and closing operations; For difference alignment loss (such as binary cross-entropy); and To balance the weighting coefficients of the two tasks, a higher penalty weight is applied to known optical cloud shadow regions and SAR speckle noise regions in this loss, forcing the model to actively learn to suppress these noises.

[0068] 4. Training Results: Through this stage of training, the encoder and feature optimization module can initially learn the stability characteristics of land features in the lake area (such as the consistency of water spectral density and the continuity of wetland texture), and have basic noise suppression and coarse localization capabilities for changing areas.

[0069] Phase 2: Semi-supervised joint fine-tuning (semantic and change co-learning) Based on the pre-trained model in the first stage, this stage introduces a small amount of labeled data for end-to-end fine-tuning, enabling the model to have accurate semantic understanding and change discrimination capabilities.

[0070] 1. Training Data: Construct a dedicated dataset for detecting semantic changes in lake areas, containing... There are 4700 pairs of samples. Of these, approximately 70% have complete pixel-level annotations (including semantic label maps of preceding and following time phases). and change label map Approximately 30% of the samples have only weak image-level labels (e.g., only knowing whether there are major categories of changes such as "dike construction" or "wetland degradation" in the image).

[0071] 2. Training setup: Unfreeze all network parameters, including the encoder, feature optimization module, semantic reasoning module, and decoder, and perform joint training.

[0072] 3. Loss Function: A multi-task joint loss function, L_joint, is adopted to simultaneously optimize both semantic segmentation and change detection tasks, and a consistency constraint is introduced: , The loss is semantic segmentation loss, such as cross-entropy loss. To improve the ability to identify key features in the lake area, when calculating this loss, the prediction errors of categories such as "wetland" and "embankment" are given higher weights (e.g., 1.5 times). For change detection loss, a combination of Dice loss and binary cross-entropy loss can be used; This refers to cross-temporal semantic consistency loss. Its core idea is: for a pixel region labeled as "unchanged," the semantic prediction probability distribution between its two preceding and succeeding temporal phases... and Consistency should be maximized, and this inconsistency can be measured and penalized using Kullback-Leibler Divergence. , , These are the weighting coefficients for each type of loss.

[0073] 4. Training strategy: Data augmentation: Strong data augmentation is applied during training, including random rotation, flipping, cropping, color jittering, and mixup, to improve the robustness of the model to different imaging conditions and ground features.

[0074] Iterative pseudo-labeling: For weakly labeled samples, the current model is used for forward inference to select pixels with a prediction confidence higher than a threshold θ (e.g., 0.9). The predicted semantics and change category of these pixels are then used as "pseudo-labels" and mixed with manually labeled strong-label data for the next round of training. This process is iterative, gradually expanding the effective training data, which is particularly beneficial for improving the model's ability to identify rare change types such as "dike demolition" and "newborn aquaculture in small ponds."

[0075] Lake area characteristic constraints: Prior knowledge of the lake area can be further incorporated into the loss function. For example, the weight of the "water body" category in the semantic segmentation loss can be dynamically adjusted based on the water level data corresponding to the input image (the weight increases when the water level is high); or the weight of the semantic consistency loss can be enhanced within the vector range of ecological protection areas (such as core wetlands) to make the model's predictions in these areas more stable.

[0076] 5. Training Results: Through this stage of training, the model can deeply integrate multi-scale features, accurately align cross-temporal semantics, and ultimately achieve high-precision semantic change detection that simultaneously outputs the location and type of change.

[0077] S7. Output and Post-processing of Change Results This step transforms the probability graph output by the network into the final application product.

[0078] Outcome-based decision-making and integration: For the probability diagram of change Set threshold (e.g., 0.5) to obtain a binary transformation mask. .

[0079] for The pixels marked as "changed" are respectively from and The category with the highest probability is used as its semantic label in the preceding and following time phases. and .

[0080] For "unchanged" pixels, it is generally assumed that they are consistent before and after, i.e. .

[0081] Post-processing optimization: Verification based on the spatiotemporal logic of ground feature changes.

[0082] Spatial context verification: For example, an isolated pixel identified as "water" that is surrounded by a large area of ​​"farmland" may be a classification error and can be corrected based on the dominant category of its neighborhood.

[0083] Change logic verification: Utilize a prior knowledge base to eliminate impossible or extremely low-probability change types. For example, it is extremely rare for "water body" to directly change into "building" in a short period of time; such changes require thorough review or correction.

[0084] Edge smoothing: Morphological opening and closing operations or conditional random fields (CRF) are used to smooth the boundaries of the changing regions, removing small holes and burrs.

[0085] Generate structured products: The system automatically outputs: Thematic raster maps: including the final variation mask map, Phase classification diagram Phase classification diagram.

[0086] Transformed patch vector file: The polygon boundaries of the transformed area are extracted using a raster-to-vector conversion algorithm, and each patch is associated with an attribute table.

[0087] Change Detection Report: The attribute table includes fields such as patch ID, area, perimeter, previous time phase category, subsequent time phase category, change type, and change confidence level. It can also calculate the change area and transition matrix by land type.

[0088] Through the detailed explanation of the above seven steps, this embodiment fully and clearly demonstrates the complete technical implementation of the present invention's method from raw data to the final semantic change product. Each step is interconnected and progressively advanced, comprehensively utilizing multiple technical means such as multi-source data processing, deep learning modeling, and physical prior fusion, effectively solving the key challenges in semantic change detection in lake areas.

Claims

1. A method for detecting semantic changes in lake areas based on remote sensing data, characterized in that, Includes the following steps: S1. Data Acquisition and Preprocessing: Acquire multi-source remote sensing image data of the target lake area in the first and second time phases. The multi-source remote sensing image data includes at least optical remote sensing images and synthetic aperture radar (SAR) images. Perform preprocessing on the multi-source remote sensing image data. The preprocessing includes radiometric correction, geometric correction, and spatial registration to obtain a spatially aligned multi-source remote sensing image dataset. S2. Multi-source feature construction: Based on the preprocessed optical remote sensing image, calculate spectral index features including at least water index and vegetation index, and combine the spectral index features with the original optical image band data to form first feature data; based on the preprocessed SAR image, extract backscattering features to form second feature data; fuse the first feature data and the second feature data in the channel dimension to construct multi-channel input features; S3. Multi-scale feature extraction: The multi-channel input features are divided into blocks using a convolution operation with overlapping windows to extract multi-scale feature maps containing different spatial scales; S4. Water level-driven cross-temporal feature alignment: Obtain lake water level information corresponding to the first and second time phases; determine spatial location offset parameters based on the water level information, and use the offset parameters to spatially correct the multi-scale feature maps of the first and second time phases to achieve cross-temporal feature alignment; S5. Multimodal feature optimization and noise suppression: The dual-temporal multi-scale feature map aligned in step S4 is subjected to noise suppression processing. The noise suppression processing includes: for areas in the optical features affected by cloud and fog shadows, feature compensation is performed using SAR features; for speckle noise in the SAR features, adaptive filtering is performed based on the spatial structure characteristics of ground features. S6. Cross-temporal semantic reasoning and change discrimination: Based on the optimized dual-temporal features, cross-temporal semantic association modeling is performed to generate change features; the change features are decoded, and the location information of the changed area, the semantic category of the land cover before the change, and the semantic category of the land cover after the change are output synchronously; S7. Output the semantic change detection results.

2. The method for detecting semantic changes in lake areas based on remote sensing data according to claim 1, characterized in that, In step S1, the multi-source remote sensing image data also includes auxiliary data that is in the same time phase as the optical remote sensing image and SAR image. The auxiliary data includes at least one of hydrological data, ecological survey data, or land use vector data.

3. The method for detecting semantic changes in lake areas based on remote sensing data according to claim 1, characterized in that, In step S2, the spectral index features also include building indices or bare soil indices used to characterize artificial features.

4. The method for detecting semantic changes in lake areas based on remote sensing data according to claim 1, characterized in that, Step S3 specifically involves processing the multi-channel input features through a multi-layer feature extraction network. The multi-layer feature extraction network includes convolutional layers with overlapping sliding windows and outputs multi-level feature maps with resolutions ranging from high to low to form the multi-scale feature maps.

5. The method for detecting semantic changes in lake areas based on remote sensing data according to claim 1, characterized in that, In step S4, determining the spatial offset parameter based on water level information specifically includes: based on the correspondence model between historical water level data and the spatial offset of land feature boundaries, calculating the estimated spatial offset as the offset parameter according to the water level difference between the current time phases.

6. The method for detecting semantic changes in lake areas based on remote sensing data according to claim 1, characterized in that, In step S5, the adaptive filtering of speckle noise in SAR features specifically involves: using large window filtering for water areas and small window filtering for linear features or edge areas.

7. The method for detecting semantic changes in lake areas based on remote sensing data according to claim 1, characterized in that, In step S6, the cross-temporal semantic association modeling is implemented using a semantic reasoning module based on an attention mechanism. The semantic reasoning module integrates a window self-attention component for capturing long-range dependencies, a channel attention component for enhancing the channel weights of the spectral index, and a spatial attention component for focusing on change-sensitive regions.

8. The method for detecting semantic changes in lake areas based on remote sensing data according to claim 1, characterized in that, In step S6, the decoding of the change features is implemented using a decoder network with a jump connection structure to fuse deep semantic features with shallow detail features and output a high-resolution semantic map and change map.

9. A method for detecting semantic changes in lake areas based on remote sensing data according to claim 1, characterized in that, The method also includes a post-processing step for the preliminary discrimination results, which includes logical verification and correction of the location information and / or semantic category of the changed area based on prior knowledge of land cover categories or spatial context relationships.