Cultivated land semantic change detection method and system based on boundary guidance and dynamic memory network, medium and electronic equipment
By combining a lightweight convolutional network and a DINOv2-small network with a dynamic memory network, the fusion of local spatial features and global semantic features and the updating of differential features in the farmland semantic change detection method are realized. This solves the problem of insufficient accuracy and reliability in farmland change detection in existing technologies and improves the sensitivity and stability of detection.
Patent Information
- Application Number
- CN202511772621.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-28
- Publication Date
- 2026-02-03
AI Technical Summary
Existing methods for detecting semantic changes in cultivated land are difficult to effectively integrate low-level spatial features with high-level semantic features, and are prone to producing spurious changes under seasonal differences, illumination variations, and noise interference, resulting in insufficient detection accuracy and reliability.
A lightweight convolutional network is used to extract local spatial features, combined with a DINOv2-small network for global semantic feature extraction, and cross-layer fusion is achieved through a self-attention mechanism. The difference features and boundary enhancement of dual-temporal images are utilized, and a dynamic memory network is combined for feature update. Finally, the detection results are output through a classifier.
It enhances the comprehensive perception capability of farmland change detection, strengthens the sensitivity and robustness to subtle changes, significantly improves detection accuracy and stability, and effectively suppresses spurious change interference.
Smart Images

Figure CN121459201A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of image recognition technology, specifically relating to a method, system, medium, and electronic device for detecting semantic changes in cultivated land based on boundary guidance and dynamic memory networks. Background Technology
[0002] Arable land, as a vital natural resource for human survival and development, plays an irreplaceable role in food security, economic construction, and ecological environmental protection. Changes in arable land use not only affect agricultural productivity but also have a profound impact on regional ecological balance and socio-economic development. Therefore, continuous and accurate monitoring of arable land dynamics is of great significance for formulating scientific and rational land use policies and ensuring the sustainable use of arable land resources.
[0003] With the rapid development of remote sensing technology, the ability to acquire multi-temporal remote sensing images has been greatly improved, achieving significant progress in both spatial and spectral resolution. Change detection methods based on remote sensing images have gradually become important tools in fields such as land resource management, agricultural surveys, and ecological environment monitoring. The core task of change detection is to identify changes in land cover types by analyzing remote sensing images acquired at different times in the same area. Existing change detection methods can be broadly classified into two categories: one is binary change detection methods, which aim to identify areas of change and those that have not, meeting the needs of some applications that only require focusing on the scope of change; the other is semantic change detection methods, which can not only locate the location of changes but also identify specific types of land cover changes, such as changes from cultivated land to water areas or from forest land to construction land. These methods have higher application value in refined land use management and resource planning decisions.
[0004] While the introduction of deep learning methods has propelled the development of semantic change detection, several unresolved issues remain in practical applications. First, existing models often struggle to effectively integrate low-level spatial features with high-level semantic features simultaneously, leading to insufficient recognition of subtle changes in land cover. Second, commonly used differential feature modeling methods, such as simple feature subtraction or concatenation, are insufficient to fully represent the complex temporal dynamics in multi-temporal remote sensing images. Finally, due to the influence of seasonal differences, lighting conditions, and noise interference on remote sensing image acquisition, most existing methods often produce spurious changes, making it difficult to identify genuine changes in land cover. These problems are particularly prominent in farmland dynamic monitoring scenarios, as farmland changes are often subtle and seasonal, increasing the difficulty of change detection.
[0005] Therefore, there is an urgent need for a method for detecting semantic changes in cultivated land that can take into account both local and global information, possess robust difference modeling capabilities, and effectively suppress spurious change interference, so as to improve the detection accuracy and reliability of land cover changes. Summary of the Invention
[0006] To address the problems of existing technologies that struggle to simultaneously capture both low-level spatial details and high-level semantic features, overly simplistic difference modeling methods, and susceptibility to false changes when faced with seasonal variations, lighting conditions, and image noise, this invention provides a method, system, medium, and electronic device for detecting semantic changes in cultivated land based on boundary guidance and dynamic memory. This approach effectively improves the accuracy and robustness of change detection.
[0007] The method of this invention is as follows: First, a lightweight convolutional network is used to extract local spatial features from the input image. Simultaneously, multi-level global semantic features are extracted from the input image based on DINOv2-small, and cross-layer fusion is achieved through a self-attention mechanism. Then, the local spatial features and global semantic features are fused to obtain fused features that take into account both local spatial and global semantic aspects. Second, the fused features from the dual-temporal images are used for differential feature extraction and boundary enhancement to generate enhanced differential features. Then, a dynamic memory network is used to update these differential features, obtaining updated features. Finally, the updated features are input into a classifier to obtain the farmland semantic change detection results.
[0008] The specific technical solution adopted in this invention is as follows:
[0009] A method for detecting semantic changes in cultivated land based on boundary guidance and dynamic memory networks, the method comprising:
[0010] Collecting dual-temporal remote sensing images of farmland and and the image and Local spatial features and global semantic features are extracted separately, and then the extracted local spatial features and global semantic features are fused to obtain the image. and Corresponding fusion features and ;
[0011] Based on fusion features and Perform differential feature extraction to obtain initial differential features. , and then from Extract and process boundary information, and then process the boundary information. and Weighted fusion yields enhanced differential features. ;
[0012] Using dynamic memory networks to enhance differential features Perform an update to obtain the updated features. ;
[0013] Based on updated features The semantic change detection results of cultivated land are obtained by using a classifier.
[0014] Preferably, the result is the same as and Corresponding fusion features and The process includes:
[0015] Lightweight convolutional networks are used to extract images separately. and The local spatial features are used to obtain the image. Local spatial features ,image Local spatial features ;
[0016] Images were extracted using the DINOv2-small backbone network. and global semantic features To obtain the image global semantic features ,image global semantic features ;
[0017] Images are fused separately and Local spatial features and global semantic features are used to obtain images and Corresponding fusion features and .
[0018] Preferably, the enhanced differential features are obtained. The process includes:
[0019] Fusion features and Perform differential feature extraction to obtain initial differential features. ;
[0020] From initial differential characteristics Extract and process the boundary information to obtain the processed boundary information. ;
[0021] Based on spatial attention and channel attention mechanisms, the initial differential features are... With processed boundary information Weighted fusion is performed to obtain enhanced differential features. .
[0022] Preferably, the updated features are obtained. The process includes:
[0023] Constructing a semantic memory bank for cultivated land ;
[0024] right Perform the operation to generate a query vector. ;
[0025] Calculate query vector With memory bank Middle memory item The similarity is then updated. hour The corresponding weights.
[0026] Weights and memory Memory items in The updated term is obtained by performing a weighted summation, and then the updated term is compared with the enhanced difference features. Adding them together yields the updated features. .
[0027] Preferably, the query vector is calculated. With memory bank Middle memory item The similarity method is as follows:
[0028]
[0029] in, For temperature coefficient, For specific query vectors, For memory items.
[0030] The second objective of this invention is to provide a farmland semantic change detection system based on boundary guidance and dynamic memory networks, the system comprising:
[0031] The local-global feature extraction module is used for image processing. and Local spatial features and global semantic features are extracted separately, and then the extracted local spatial features and global semantic features are fused to obtain the image. and Corresponding fusion features and ;
[0032] The differential feature extraction module is based on fused features. and Perform differential feature extraction to obtain initial differential features. , and then from Extract and process boundary information, and then process the boundary information. and Weighted fusion yields enhanced differential features. ;
[0033] The update module utilizes a dynamic memory network to process the enhanced differential features. Perform an update to obtain the updated features. ;
[0034] Output module, based on updated features The classifier is used to obtain the semantic change detection results of cultivated land;
[0035] The system is based on a method for semantic change of cultivated land that is guided by boundaries and uses a dynamic memory network.
[0036] A third objective of this invention is to provide a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute a method for semantic change of cultivated land based on boundary guidance and dynamic memory networks.
[0037] A fourth objective of this invention is to provide an electronic device comprising:
[0038] At least one processor, and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform a method for semantic change of cultivated land based on boundary guidance and dynamic memory networks.
[0039] The beneficial effects of this invention are: This invention discloses a method, system, medium, and electronic device for detecting semantic changes in cultivated land based on boundary guidance and dynamic memory networks. Compared with the prior art, the improvement of this invention lies in:
[0040] (1) This invention combines a lightweight convolutional network with DINOv2-small to achieve complementarity between local spatial details and global semantic information in dual-temporal remote sensing images, making feature representation more comprehensive and effectively improving the comprehensive perception capability of the method of this invention in detecting changes in cultivated land.
[0041] (2) The present invention introduces a boundary guidance mechanism and combines differential operators with spatial-channel joint attention, which can highlight the significant change areas of dual-temporal remote sensing images, suppress irrelevant information, make the difference modeling more refined, and improve the sensitivity of the method of the present invention to the detection of subtle changes in cultivated land.
[0042] (3) The present invention adopts a dynamic memory network, with the category prototype as the core for feature query and update, and gradually optimizes the feature distribution, thereby effectively distinguishing between real changes and pseudo changes in dual-temporal remote sensing images, making the method of the present invention more robust in detecting semantic changes in cultivated land, and significantly enhancing the stability of the method of the present invention in complex environments. Attached Figure Description
[0043] Figure 1 This is a flowchart of the present invention;
[0044] Figure 2 This is a structural diagram of the method of the present invention. Detailed Implementation
[0045] To enable those skilled in the art to better understand the technical solutions of the present invention, the technical solutions of the present invention will be further described below in conjunction with the accompanying drawings and embodiments. The following embodiments are used to illustrate the present invention, but should not be used to limit the scope of the present invention.
[0046] Example 1:
[0047] See attached document Figure 1-2 The method for detecting semantic changes in cultivated land based on boundary guidance and dynamic memory networks, as shown, includes:
[0048] S1. Collect dual-temporal remote sensing images of cultivated land. and and the image and Local spatial features and global semantic features are extracted separately, and then the extracted local spatial features and global semantic features are fused to obtain the result. and Corresponding fusion features and .
[0049] S101. Use a lightweight convolutional network to extract images respectively. and The local spatial features are used to obtain the image. Local spatial features ,image Local spatial features , among which, image and These are remote sensing images of cultivated land acquired at different times, and , and These represent the height and width of the image, respectively.
[0050] Lightweight convolutional networks for image extraction and Local spatial features The expression is:
[0051] (1)
[0052] in, For convolution operations, It is a non-linear activation function.
[0053] The extraction process of the lightweight convolutional network is as follows: The lightweight convolutional network consists of two parts: the first part is a regular 3×3 convolution, used to extract basic local features; the second part is a 3×3 dilated convolution with a dilation rate, used to expand the receptive field with low computational cost. The outputs of the two convolutions are concatenated along the channel dimension, and then feature fusion and compression are achieved through 1×1 convolutions. Furthermore, to meet the feature processing requirements of deep networks, a convolution with a stride of 4 is added to the output of the lightweight convolutional network to achieve downsampling. This structure, while maintaining lightweight design, can simultaneously acquire local details and large-scale contextual features, thus providing efficient features for remote sensing image change detection.
[0054] S102. Images are extracted using the DINOv2-small backbone network. and global semantic features To obtain the image global semantic features ,image global semantic features .
[0055] Specifically, the image is extracted using multiple intermediate layers of the DINOv2-small backbone network. and Multi-level features Where L represents the number of network layers; and a self-attention mechanism is used to merge features from multiple levels to generate an image. global semantic features and images global semantic features .
[0056] Extracting global semantic features The expression is:
[0057] (2)
[0058] in, For self-attention mechanism, This is a channel-dimensional splicing operation.
[0059] In this embodiment, the DINOv2-small network is used as the image. and The backbone network for feature extraction. The DINOv2-small network is a visual Transformer model based on self-supervised learning. Its core idea is to obtain discriminative general visual features by performing feature alignment training on unlabeled images through a self-distillation mechanism.
[0060] Specifically, DINOv2-small is based on the Vision Transformer-small architecture, including an image block embedding layer and multiple Transformer encoder modules. Input image and Then, first image and The images are divided into fixed-size blocks (e.g.) The image is transformed into a sequence of feature vectors through a linear mapping; subsequently, these feature vector sequences are sequentially processed by a 12-layer Transformer encoder for feature extraction, yielding the resulting images. and Multi-level features of bi-temporal images. Each Transformer encoder layer contains a multi-head self-attention and feedforward neural network structure to capture global dependencies and contextual information.
[0061] In this embodiment, to take into account multi-level feature information, the outputs of layers 2, 5, 8, and 11 of the 12-layer Transformer encoder are selected as multi-level features with multi-scale characteristics; among them, the shallow layer (layer 2 and layer 5) features can preserve image information. and Deeper features (layers 8 and 11) provide more spatial structure and edge texture information, thus preserving more image data. and Stronger semantic information.
[0062] Next, the multi-level features of the dual-temporal images are fused to obtain the final global semantic features. and Specifically, with Taking the acquisition of the image as an example, firstly... Features from different levels are unified to the same feature dimension through linear projection to ensure consistency of multi-scale features in both spatial and channel dimensions. Subsequently, the features from each level are flattened, serialized, and concatenated to form a feature sequence containing multi-scale information. Based on this, a multi-head self-attention mechanism is introduced to globally model this feature sequence, adaptively adjusting the weights of features at different scales by calculating the correlation between levels, thereby achieving cross-level information interaction and feature supplementation. The attention-enhanced feature sequence is then normalized and remapped back to the original spatial structure; finally, the fused features are obtained by averaging across the scale dimensions. This process effectively integrates the image... The semantic and detailed information of features at different levels are combined to generate more discriminative multi-scale fusion features.
[0063] S103, Merge images separately and Local spatial features and global semantic features , obtain with image and Corresponding fusion features and .
[0064] Specifically, the images are... Local spatial features With global semantic features and images Local spatial features With global semantic features Perform scale alignment and linear transformation to achieve image... of and The fusion yields fusion characteristics. ,image of and The fusion yields fusion characteristics. ,
[0065] and The fusion expression is:
[0066] (3)
[0067] in, For upsampling operation, In response to The convolution projection matrix, In response to The convolution projection matrix, For fusion operation.
[0068] The specific process of fusion is as follows: first, the features are... Perform upsampling and 1×1 convolutional projection, making it compatible with the convolutional features. Maintaining consistency in spatial resolution and number of channels, the convolutional features are projected using a 1×1 convolution to align the channels. Then, the two types of features are concatenated along the channel dimension and input into a channel attention unit. This unit generates attention weights based on global average pooling, which weight the concatenated features to highlight key information and suppress redundant information. Finally, the output convolutional unit (containing 1×1 convolution, batch normalization, and a non-linear activation function) performs feature compression and enhancement, resulting in a fused feature that simultaneously possesses local spatial details and global semantic information. .
[0069] S2, Based on Fusion Features and Perform differential feature extraction to obtain initial differential features. Then from the initial difference characteristics Extract and process boundary information, and then process the boundary information. and Weighted fusion yields enhanced differential features. The specific process is as follows:
[0070] S201, Regarding fusion features and Perform differential feature extraction to obtain initial differential features. The specific process is as follows:
[0071] Fusion features and Perform pixel-level difference operations to highlight the degree of change between the two; simultaneously, fuse the features. and Channel-level concatenation is performed to preserve complete contextual information. Then, the difference results and concatenated features are input into a convolutional unit, processed by batch normalization and a non-linear activation function to generate initial differential features. .
[0072] Based on fusion features and Initial differential feature extraction is performed to obtain the initial differential features. The calculation expression is:
[0073] (4)
[0074] in, This represents pixel-level difference operations. This represents the convolution mapping operation. It is a non-linear activation function.
[0075] S202, From the initial difference characteristics Extract and process the boundary information to obtain the processed boundary information. The specific process includes:
[0076] After obtaining the initial differential characteristics Based on this, the present invention introduces a boundary guidance mechanism, utilizing the Sobel operator to start from the initial difference features Boundary information is extracted from the boundary data, and then refined through convolution and normalization operations to ensure the accuracy and robustness of boundary detection. The final result is the feature data of the boundary information, i.e., the processed boundary information. .
[0077] Obtain the processed boundary information The expression is:
[0078] (5)
[0079] in, This indicates a normalization operation. This represents the convolution operation. This is the Sobel operator.
[0080] S203. Based on spatial attention and channel attention mechanisms, the initial differential features are... With processed boundary information Weighted fusion is performed to obtain enhanced differential features. .
[0081] Initial differential features With processed boundary information The expression for weighted fusion is:
[0082] (6)
[0083] in, This represents the convolution operation. This indicates the concatenation of channel dimensions. Indicates channel attention. This represents spatial attention.
[0084] In formula (6), spatial attention is applied to the initial differential features through a multi-head self-attention mechanism. Modeling is performed to enhance the spatial dependencies of changing regions; channel attention learns initial differential features through adaptive pooling and nonlinear mapping. The importance distribution of different channels is used to selectively enhance features, i.e., through weighting. Finally, the weighted differential features are compared with the processed boundary information. The pieces are spliced together to achieve fusion, and then... Convolution and batch normalization are used to obtain enhanced differential features. .
[0085] Through the above steps, this embodiment can effectively capture the changing areas of the image while using boundary information to suppress noise interference, thereby highlighting the real changing areas and improving detection accuracy.
[0086] S3. Utilizing dynamic memory networks to enhance differential features Perform an update to obtain the updated features. .
[0087] The dynamic memory network employs a dynamic memory optimization mechanism to enhance the differential features obtained in step S2. The update is performed using a dynamic memory optimization mechanism based on category prototypes to enhance the robustness of the method to pseudo-changes in cultivated land.
[0088] S301. Constructing a semantic memory database for cultivated land ;
[0089] Specifically, the semantic memory bank of cultivated land Multiple memory items Each memory item is used to represent the semantic features of different farmland changes.
[0090] S302, to Process the data to generate a query vector. The query vector Contains multiple specific .
[0091] Specifically, for Channel normalization is performed, and a query vector is generated through convolution mapping. ;
[0092] Generate query vectors The expression is:
[0093] (7)
[0094] in, B represents the batch size. for The dimension of a vector when it is flattened in space. For query dimensions, for Normalization results This indicates a linear projection operation.
[0095] S303, Calculate the query vector With memory bank Middle memory item The similarity is then updated. hour Corresponding weights .
[0096] Specifically, for a given query vector and a certain category of memory items The expression for calculating the similarity between the two is:
[0097] (8)
[0098] in, This is a temperature coefficient used to adjust the smoothness of the feature similarity distribution. In embodiments of the present invention, Setting it to 0.1 is relatively small and can effectively enhance the model's sensitivity to highly similar samples.
[0099] The update is obtained by using the similarity between the query vector and the memory item. hour Corresponding weight The expression is:
[0100] (9)
[0101] in, This involves normalizing the similarity scores to transform the weights between samples into a probability distribution. This operation ensures that all weight values are non-negative and that their sum is 1, thus achieving [the desired consistency]. In update A relative measure of contribution over time.
[0102] S304, Weight With memory bank Middle memory item The updated term is obtained by performing a weighted summation, and then the updated term is compared with the enhanced difference features. Adding them together yields the updated features. The entire process is expressed as:
[0103] (10)
[0104] Through the above steps, dynamic memory networks can fuse memory retrieval results with differential features. Updated features are then obtained. This feature, while preserving real change information, effectively suppresses spurious change interference, thereby enhancing the change features and iteratively optimizing the memory, thus improving the accuracy and robustness of the overall change detection model.
[0105] S4, Update Features As input to the classifier, the output yields the results of semantic change detection of cultivated land.
[0106] The classifier can be a conventional classifier, such as a lightweight and stable convolution classifier.
[0107] Example 2:
[0108] This implementation also presents a farmland semantic change detection model based on boundary guidance and dynamic memory networks, which includes:
[0109] The local spatial feature extraction submodule is used to extract dual-temporal remote sensing images. and The local spatial features are obtained, and this process is achieved through step S101;
[0110] The global semantic feature extraction submodule is used to extract dual-temporal remote sensing images. and The global semantic features are obtained, and this process is achieved through step S102;
[0111] Local and global feature fusion submodules are used to fuse images separately. and Local spatial features and global semantic features are used to obtain images and Corresponding fusion features and This process is achieved through step S103;
[0112] The difference information extraction and boundary enhancement submodule is based on fusion features. and Perform differential feature extraction to obtain initial differential features. , and then from Extract and process boundary information, and then process the boundary information. and Weighted fusion yields enhanced differential features. This process is achieved through step S2;
[0113] The dynamic memory network submodule is used to process the enhanced differential features. Perform an update to obtain the updated features. This process is achieved through step S3;
[0114] The classifier submodule is based on updated features. The output yields the results of semantic change detection of cultivated land.
[0115] Before use, this model requires supervised training. During training, labeled data serves as the supervision signal, and the parameters of each module are continuously updated by optimizing the loss function. Simultaneously, the dynamic memory network submodule iterates synchronously during model training, gradually adapting the category prototype memory items to the real data distribution. After multiple rounds of training, the overall capability of the model is steadily improved, enabling it to more accurately distinguish between real and spurious changes, thus achieving reliable farmland change detection.
[0116] To ensure the memory bank in the dynamic memory network submodule Memory items of various categories To effectively reflect the distribution of features across different categories, the dynamic memory network submodule employs a dynamic update mechanism combining Short-Term Memory (STM) and Long-Term Memory (LTM) in its memory bank. STM temporarily stores feature information from recent samples to quickly respond to changes in feature distribution; LTM stores relatively stable category features to maintain the overall consistency and representativeness of the model. When an item in the STM becomes stable after multiple iterations, the system incorporates it into the LTM, achieving a transfer and fusion from short-term to long-term memory.
[0117] Record the dataset Involving Similar to changes in cultivated land, for sample pairs composed of dual-temporal remote sensing images. and The corresponding pixel-level semantic change detection label for cultivated land is For each category , Using tags Select the set of valid pixels for the corresponding category And for each short-term memory item Before selecting based on similarity Update the high-confidence samples:
[0118] conduct The updated expression is:
[0119] (11)
[0120] in, For the updated memory item, The momentum coefficient, This indicates a normalization operation. These are similarity-based weighted coefficients. For category corresponding The high-confidence sample set in Set to 0.01, Set it to 0.8.
[0121] When category samples When there is insufficient memory, a skip update mechanism is used to delay the memory entries. Update. When the memory item When the cumulative number of updates reaches a threshold, the gating mechanism is triggered, and reliable short-term memory items are written to it. At the same time, reset the corresponding It is a zero vector, and the update count is reset to 0.
[0122] Through the above process, the dynamic memory network submodule can fuse memory retrieval results with differential features. Enhanced features were subsequently obtained. This feature, while preserving real change information, effectively suppresses spurious change interference, thereby enhancing the change features and iteratively optimizing the memory, thus improving the accuracy and robustness of the overall change detection model.
[0123] During training, the memory bank expression for the dynamic memory network submodule is:
[0124]
[0125] in, This is a weighting factor used to control the relative contributions of long-term memory and short-term memory in the fusion process. The value is set to 0.7, which means that long-term memory has a higher weight, thus maintaining feature consistency while taking into account the sensitivity to new changes.
[0126] This model is implemented based on the PyTorch framework. During the model training phase, the cross-entropy loss function is used as a supervision signal. This loss function measures the difference between the predicted distribution and the true distribution by comparing the predicted probability of each pixel output by the model with the corresponding true class label. During training, the model's learnable parameters (including convolutional kernel weights, attention weights, etc.) are iteratively updated through the backpropagation algorithm, continuously optimizing feature extraction and classification capabilities, thereby gradually reducing the overall prediction error and improving the accuracy of change area detection. After training and evaluation on multiple remote sensing image datasets, this model demonstrates excellent performance in the task of detecting semantic changes in cultivated land, accurately identifying the true change areas of cultivated land, and exhibiting strong robustness to interference factors.
[0127] Example 3:
[0128] To verify the effectiveness and applicability of the method in Example 1, this example selects the representative large-scale cultivated land semantic change detection dataset JL1 for experiments. The JL1 dataset, constructed by Changguang Satellite Technology Co., Ltd. based on Jilin-1 satellite images, contains approximately 8000 pairs of high-resolution dual-temporal remote sensing images with a spatial resolution better than 0.75 meters, covering various typical cultivated land change types. This dataset is large in scale and finely labeled, fully supporting the training and evaluation of the model. This dataset has a large-scale and rich sample of cultivated land changes, comprehensively reflecting the actual dynamic situation of cultivated land. This example conducts experiments based on this dataset, which can fully verify the effectiveness and generalization ability of the proposed method.
[0129] To comprehensively evaluate the performance of the method of this invention, this embodiment employs four commonly used evaluation metrics: Overall Accuracy (OA), Mean Intersection over Union (mIoU), Mean F1 Score (mF1), and the F1 metric for semantic change detection. Among them, OA is used to measure the overall classification accuracy and can reflect the model's overall ability to distinguish all categories; mIoU emphasizes the spatial overlap between the predicted results and the true labels and is the core indicator of semantic segmentation tasks; mF1 balances precision and recall across categories and can better reflect the model's comprehensive performance across different semantic change categories. The evaluation specifically targets the detection performance in changed areas, unaffected by the "unchanged" category, and can more objectively reflect the model's ability to identify real changes. All the above indicators show a trend of "higher values indicate better performance." Combining these four indicators allows for comparison of the overall accuracy, spatial localization ability, and sensitivity of change detection of different methods, thereby comprehensively and objectively verifying the effectiveness and advantages of this invention.
[0130] To further verify the advantages of the method of the present invention, we selected a total of 5 comparison algorithms: BiSRNet, SCanNet, CdSC, BT-SCD, and EGMS-Net. By applying the above existing methods and the method of the present invention to the JL1 dataset, the experimental results are shown in Table 1.
[0131] Table 1. Experimental results of the method of this invention and existing methods on the JL1 dataset.
[0132] method OA (%) mIoU(%) mF1(%) Fscd(%) BiSRNet 69.57 31.73 44.69 38.01 SCanNet 68.59 25.39 36.67 34.08 CdSC 75.27 36.59 50.58 51.79 BT-SCD 71.34 30.08 42.34 42.61 EGMS-Net 69.67 34.53 48.53 40.65 Method of the present invention 76.78 39.83 53.53 55.26
[0133] As clearly shown in Table 1, the method of this invention outperforms existing methods in all four evaluation metrics, indicating that the method of this invention has significant advantages in capturing farmland change information, suppressing spurious changes, and improving classification accuracy. On the JL1 dataset, the method of this invention achieves an mF1 value of 53.53% and an mIoU of 39.83%. The accuracy is 55.26%, and the OA is 76.78%. Compared with existing methods, the method of this invention shows significant improvements in all indicators. For example, compared with EGMS-Net, the method of this invention improves in OA, mIoU, mF1, and mIoU. The four indicators were 7.11, 5.3, 5, and 14.61 percentage points higher, respectively, indicating that the method of the present invention is more accurate and reliable in identifying areas of actual change in cultivated land.
[0134] In summary, the experimental results of the method of this invention on the JL1 dataset show that it can effectively capture semantic changes in cultivated land, improve detection accuracy and robustness, and has significant performance advantages compared with existing methods.
[0135] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.
Claims
1. A method for detecting semantic changes in cultivated land based on boundary guidance and dynamic memory networks, characterized in that, The method includes: Collecting dual-temporal remote sensing images of farmland and and the image and Local spatial features and global semantic features are extracted separately, and then the extracted local spatial features and global semantic features are fused to obtain the image. and Corresponding fusion features and ; Based on fusion features and Perform differential feature extraction to obtain initial differential features. , and then from Extract and process boundary information, and then process the boundary information. and Weighted fusion yields enhanced differential features. ; Using dynamic memory networks to enhance differential features Perform an update to obtain the updated features. ; Based on updated features The semantic change detection results of cultivated land are obtained by using a classifier.
2. The method according to claim 1, characterized in that, Get with image and Corresponding fusion features and The process includes: Lightweight convolutional networks are used to extract images separately. and Local spatial features To obtain the image Local spatial features ,image Local spatial features ; Images were extracted using the DINOv2-small backbone network. and global semantic features To obtain the image global semantic features ,image global semantic features ; Images are fused separately and Local spatial features and global semantic features , obtain with image and Corresponding fusion features and .
3. The method according to claim 2, characterized in that, Enhanced differential features The process includes: Fusion features and Perform differential feature extraction to obtain initial differential features. ; From the characteristics of differences Extract and process the boundary information to obtain the processed boundary information. ; Based on spatial attention and channel attention mechanisms, differential features are... With processed boundary information Weighted fusion is performed to obtain enhanced differential features. .
4. The method according to claim 3, characterized in that, Get updated features The process includes: Constructing a semantic memory bank for cultivated land ; right Perform the operation to generate a query vector. ; Calculate query vector With memory bank Middle memory item The similarity is then updated. hour The corresponding weights; Weights and memory Memory items in The updated term is obtained by performing a weighted summation, and then the updated term is compared with the enhanced difference features. Adding them together yields the updated features. .
5. The method according to claim 4, characterized in that, Query vector With memory bank Middle memory item The similarity calculation method is as follows: ; in, For temperature coefficient, For specific query vectors, For memory items.
6. A system for detecting semantic changes in cultivated land based on boundary guidance and dynamic memory networks, the system comprising: The local-global feature extraction module is used to acquire dual-temporal remote sensing images of cultivated land. and and the image and Local spatial features and global semantic features are extracted separately, and then the extracted local spatial features and global semantic features are fused to obtain the image. and Corresponding fusion features and ; The differential feature extraction module is based on fused features. and Perform differential feature extraction to obtain initial differential features. , and then from Extract and process boundary information, and then process the boundary information. and Weighted fusion yields enhanced differential features. ; The update module utilizes a dynamic memory network to process the enhanced differential features. Perform an update to obtain the updated features. ; Output module, based on updated features The classifier is used to obtain the semantic change detection results of cultivated land; The system is implemented based on the method described in any one of claims 1-5.
7. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-5.
8. An electronic device, comprising: At least one processor, and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-5.