Cultivated land change detection method based on double-time-phase remote sensing image and semantic knowledge enhancement

By combining dual-temporal remote sensing imagery with semantic knowledge enhancement, and utilizing masked attention mechanism and spatiotemporal self-attention module, the problem of inaccurate detection in complex scenarios of farmland change detection is solved, achieving efficient and accurate farmland change detection.

CN120976776APending Publication Date: 2025-11-18SOUTH CHINA NORMAL UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510890445.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing technologies are difficult to effectively adapt to complex scenarios in farmland change detection, especially in terms of high detail of multi-temporal remote sensing data, spectral characteristic analysis and data registration, resulting in inaccurate detection results.

Method used

By combining dual-temporal remote sensing imagery and semantic knowledge enhancement methods, and through mask generation, weighted mask self-attention mechanism, multi-scale spatial feature extraction, temporal feature extraction, and spatiotemporal self-attention module, accurate detection of farmland changes is achieved.

Benefits of technology

It significantly improves the focus and accuracy of farmland change detection, reduces the false detection rate, and generates farmland change detection results with clear boundaries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976776A_ABST
    Figure CN120976776A_ABST
Patent Text Reader

Abstract

The invention discloses a cultivated land change detection method based on a double-time-phase remote sensing image and semantic knowledge enhancement. The method comprises the following steps: acquiring the double-time-phase remote sensing image and simultaneous thematic data; generating a mask of the dual-time-phase remote sensing image, and calculating a weighted dual-time-phase feature map; extracting a dual-time-phase spatial feature pyramid through down-sampling, up-sampling, element-by-element addition from top to bottom and transverse connection; and dynamically weighting the difference region through channel attention, carrying out attention calculation, and obtaining a cultivated land change detection result through a full connection layer. More concentrated cultivated land features are extracted under the constraint of a weighted mask self-attention mechanism, timing sequence features and spatial features of cultivated land changes are captured through a time and space attention mechanism, the timing sequence features and the spatial features are fused with the cultivated land features, a cultivated land change detection result with a clear boundary is generated, and the accuracy of cultivated land change detection is improved. And the data advantages of various data are fully utilized to remarkably improve the extraction accuracy and fineness, so that efficient and accurate cultivated land change detection is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the technical fields of remote sensing science and computer vision, specifically relating to a method for detecting farmland changes based on dual-temporal remote sensing images and semantic knowledge enhancement. Background Technology

[0002] With the rapid development of remote sensing Earth observation technology, acquiring multi-temporal and multispectral remote sensing data using advanced sensors mounted on satellites has become more convenient and efficient. The acquisition of multi-temporal remote sensing data provides high-resolution geographic information and the ability to monitor dynamic environmental changes. This technology reveals the temporal patterns of change in land features by comparing and analyzing surface observation data from different time points. Multi-temporal remote sensing data has significant application value in change detection, environmental monitoring, and resource management, and its high timeliness and comprehensive observation advantages provide a reliable data foundation for scientific research and decision support.

[0003] Developing methods for detecting farmland changes can effectively and promptly monitor these changes. Timely monitoring of farmland changes provides data support for formulating scientific and rational farmland protection and utilization policies, thereby ensuring food security. Deep learning, through multi-layered neural network structures, extracts and represents complex features, automatically learning important patterns and features from data. It possesses powerful non-linear representation capabilities and can handle large-scale data; it performs excellently in tasks such as image recognition, natural language processing, and speech recognition; it eliminates the need for manual feature design, reducing human intervention and bias; and with the support of big data and high-performance computing hardware, deep learning models can continuously improve their accuracy and generalization ability. The results of farmland change detection obtained using dual-temporal remote sensing and deep learning technologies represent a cutting-edge research area in remote sensing science and computer vision. However, facing challenges in high detail, spectral characteristic analysis, and data registration of multi-temporal remote sensing data, farmland change detection also faces numerous technical difficulties.

[0004] Over the past decade, numerous new methods have emerged in the field of remote sensing image change detection, including statistical machine learning-based and object-oriented extraction methods. However, these methods often rely excessively on low / mid-level hand-designed features when dealing with surface feature extraction, making them ill-suited for complex scenarios. In recent years, deep convolutional neural networks (CNNs) have been widely applied in remote sensing image processing, demonstrating exceptional performance, particularly in change detection. The introduction of Residual Neural Networks (ResNet) has provided new possibilities for the application of deep learning models in change detection. By introducing residual connections, ResNet effectively alleviates the performance bottleneck caused by gradient vanishing or degradation problems in traditional deep networks, enabling the network to maintain high learning capacity at greater depths. This structure makes ResNet better suited to capturing complex spatiotemporal features in change detection tasks, such as subtle surface changes or dynamic evolution of local areas. Further integration of attention mechanisms into change detection models significantly improves the accuracy and robustness of feature extraction. Attention mechanisms adaptively enhance the feature representation of changed areas while suppressing background noise and interference from irrelevant regions. This method loses information when learning detailed semantic features of images, resulting in inaccurate farmland change detection results during the upsampling process. Summary of the Invention

[0005] The main objective of this invention is to overcome the shortcomings and deficiencies of existing technologies and provide a method for detecting farmland changes based on dual-temporal remote sensing images and semantic knowledge enhancement. This method focuses on solving the problem of accurate mapping in farmland change detection methods. By combining dual-temporal remote sensing data with contemporaneous thematic data, deep learning technology is used to achieve automated feature extraction from remote sensing data. Furthermore, an improved mask attention semantic knowledge fusion model is used to accurately detect farmland changes.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: One aspect of the present invention provides a method for detecting farmland change based on dual-temporal remote sensing imagery and semantic knowledge enhancement, comprising the following steps: Acquire bi-temporal remote sensing images, as well as thematic data from the same period as the bi-temporal remote sensing images; thematic data includes permanent basic farmland data, annual cultivated land datasets, and point-of-interest data; The center point and boundary point of permanent basic farmland data and the center point of annual cultivated land data are extracted as positive cultivated land sample points, and the non-cultivated land interest point data in the interest point data are extracted as negative cultivated land sample points. A mask for dual-temporal remote sensing imagery is generated using positive and negative sample points of cultivated land. The weighted mask self-attention mechanism is used to calculate the weights of the mask of the dual-temporal remote sensing image to obtain the weighted dual-temporal feature map; The weighted dual-temporal feature map and dual-temporal remote sensing image are input into the multi-scale spatial feature extraction module based on FPN. The dual-temporal spatial feature pyramid is obtained by downsampling, upsampling and element-wise addition from top to bottom and horizontal connection. By utilizing the temporal feature extraction module that takes into account both temporal phases, the difference regions of the spatial feature pyramid of both temporal phases are dynamically weighted by channel attention to obtain the weighted feature map of both temporal phases; The spatiotemporal self-attention module is used to perform attention calculation on the bi-temporal weighted feature map to obtain the spatiotemporal feature map; The spatiotemporal feature map is input into the fully connected layer to obtain the farmland change detection results.

[0007] As a preferred technical solution, the step of extracting the center points and boundary points of permanent basic farmland data and the center points of annual cultivated land datasets as positive cultivated land sample points, and extracting non-cultivated land interest point data from the interest point data as negative cultivated land sample points, specifically involves: Based on the scope of permanent basic farmland data acquisition, a center point and the points in the four directions furthest from the center point, as well as the center point in the annual cultivated land dataset, are determined as positive sample points for cultivated land. Specifically, this includes: 1) Perform geometric correction on the dual-temporal remote sensing images to correct geographical location deviations and align the dual-temporal remote sensing images with the actual geographic coordinates; 2) Obtain the set of boundary coordinate points of the dual-temporal remote sensing image ( x 1, y 1),( x 2, y 2),…,( x n , y n ); 3) Obtain the center point using the centroid formula. C x , C y The formula for calculating ) is: , ; in, A The area of ​​the polygon: ; 4) With the center point ( C x , C y Using the center point as a reference, define four directions, and find the point farthest from the center point in each direction. As a boundary point; The center points extracted from the non-arable land interest point data were used as negative sample points for arable land.

[0008] As a preferred technical solution, positive and negative sample points of cultivated land are input into the SAM model to generate a mask for dual-temporal remote sensing images. The SAM model includes a cue encoder, an image encoder, and a mask decoder; The prompt encoder is used to encode the input positive and negative farmland sample points to generate feature representations related to the target area. The image encoder employs a pre-trained ViT architecture to capture contextual relationships in dual-temporal remote sensing images; The mask decoder combines the features output by the cue encoder and the image encoder, and generates the final segmentation mask using an autoregressive or parallel generation method.

[0009] As a preferred technical solution, the step of using a weighted mask self-attention mechanism to calculate the weights of the mask of the dual-temporal remote sensing image to obtain a weighted dual-temporal feature map is as follows: A set of mask weight matrices is generated using the mask, where the value of each pixel position represents whether the position belongs to a cultivated land area and its importance, as shown in the following formula: W mask(x,y) = Sigmoid ( M ( x , y )* α ); in, W mask(x,y) It is the first x , y The mask weight matrix of the position, M ( x , y ) represents the pixel value of the mask image; 1 represents farmland, and 0 represents the background. α It is a scalar parameter; The spatiotemporal features of the dual-temporal remote sensing images are weighted using a mask weight matrix, as shown in the following equation: ; in, F ( i , j (This is a feature map extracted from dual-temporal remote sensing images.) W mask(i,j) It is the mask weight matrix generated by mask attention. F’ It is a weighted dual-phase feature map.

[0010] As a preferred technical solution, a multi-level scaling factor with multiplied coefficients is introduced to dynamically adjust the scalar parameters in the mask weight matrix. α .

[0011] As a preferred technical solution, the multi-scale spatial feature extraction module based on FPN performs three downsampling convolutions on the input to obtain four layers of features; after each layer of features is bilinearly upsampled, it is added element-wise with the features of the previous layer, and then the upsampling artifacts are eliminated by 3×3 convolution. At the same time, by horizontally connecting the features of each layer, a dual temporal spatial feature pyramid with different resolutions is formed from top to bottom; wherein, parallel dilated convolutional layers with different dilation rates are embedded in downsampling×1 and downsampling×2.

[0012] As a preferred technical solution, the time feature extraction module considering both temporal phases includes a splicing layer, a global average pooling layer, and a weight generation layer, specifically: The splicing layer is used to integrate the dual-temporal spatial feature pyramid. After being concatenated along the channel dimension, the data is input into the global average pooling layer. The global average pooling layer is used to generate channel description vectors. And input it into the weight generation layer; The weight generation layer comprises two fully connected layers and a nonlinear activation function, as shown below: W c =σ( W 2∙ δ ( W 1∙ z )); in, , These are the parameters for the fully connected layer. δ σ is the ReLU function, and σ is the Sigmoid function. W c For the generated channel weights; Finally, the channel weights are multiplied channel by channel by the dual-temporal spatial feature pyramid to obtain the dual-temporal weighted feature map.

[0013] As a preferred technical solution, the input of the spatiotemporal self-attention module is a dual-temporal weighted feature map with dimensions H×W×C, corresponding to the spatial height, width, and number of channels of the dual-temporal remote sensing image, respectively. It includes feature projection, multi-head attention calculation, and fusion output, specifically: 1) Feature projection: Mapping the input to a linear transformation. Q , K and V The dimension of the projection matrix is C × d k ,in dk It is the dimension of each attention head; 2) Multi-head attention calculation: The global correlation between features is calculated using a multi-head self-attention mechanism, as shown in the following formula: ; in, Attention ( Q , K , V The attention weights are the obtained values. 3) Fusion output: The original features are preserved through residual connections, and the attention output is linearly transformed in channel dimension to match the dimension of the input. Finally, the spatiotemporal feature map is output through a convolutional layer and a ReLU activation function. The spatiotemporal self-attention module controls the input resolution through downsampling.

[0014] As a preferred technical solution, the spatiotemporal feature map is subjected to three or more layers of convolution to extract high-level features, and a two-dimensional change probability map is output through a fully connected layer to represent the probability that each pixel belongs to the change region. After the two-dimensional change probability map is generated, a change decision is made for each pixel based on the probability value to determine whether it belongs to the change region. After setting a threshold, pixels with probability values ​​greater than the threshold in the change probability map are classified as change regions, and pixels with probability values ​​less than the threshold are classified as non-change regions, thus obtaining the farmland change detection result.

[0015] Another aspect of the present invention provides a farmland change detection system based on dual-temporal remote sensing images and semantic knowledge enhancement, applied to the above-mentioned farmland change detection method based on dual-temporal remote sensing images and semantic knowledge enhancement, including a data acquisition module, a mask generation module, a weighted mask self-attention mechanism, a multi-scale spatial feature extraction module based on FPN, a temporal feature extraction module considering dual temporal phases, a spatiotemporal self-attention module, and an output module. The data acquisition module is used to acquire dual-temporal remote sensing images, as well as thematic data from the same period as the dual-temporal remote sensing images; thematic data includes permanent basic farmland data, annual cultivated land datasets, and point-of-interest data; The mask generation module is used to extract the center point and boundary point of permanent basic farmland data and the center point of annual cultivated land dataset as cultivated land positive sample points, and to extract non-cultivated land interest point data from the interest point data as cultivated land negative sample points; and to generate a mask for dual-temporal remote sensing image using cultivated land positive sample points and cultivated land negative sample points. The weighted mask self-attention mechanism is used to calculate the weights of the mask of the dual-temporal remote sensing image to obtain a weighted dual-temporal feature map. The multi-scale spatial feature extraction module based on FPN takes a weighted dual-temporal feature map and dual-temporal remote sensing image as input, and extracts the dual-temporal spatial feature pyramid through downsampling, upsampling-elemental summation and lateral connection. The time feature extraction module that takes into account both temporal phases is used to dynamically weight the difference regions of the spatial feature pyramid of both temporal phases through channel attention to obtain a weighted feature map of both temporal phases. The spatiotemporal self-attention module is used to perform attention calculations on the bi-temporal weighted feature map to obtain the spatiotemporal feature map; The output module is used to input the spatiotemporal feature map into the fully connected layer to obtain the farmland change detection results.

[0016] Compared with the prior art, the present invention has the following advantages and beneficial effects: (1) This application combines multi-source thematic data and effectively supplements the knowledge that is difficult to obtain in dual-temporal images through thematic data mining, significantly improving the focus of farmland change detection, thereby achieving farmland change detection with lower false detection rate and stronger robustness.

[0017] (2) In the face of the challenges of high detail, spectral characteristic analysis and data registration of multi-phase remote sensing data, this application utilizes an improved semantic knowledge fusion model based on mask attention to extract more concentrated farmland features under the constraint of weighted mask self-attention mechanism. Finally, the temporal and spatial features of farmland change are captured by temporal and spatial attention mechanism and fused with farmland features to generate farmland change detection results with clear boundaries. The data advantages of various types of data are fully utilized to significantly improve the accuracy and detail of extraction, thereby achieving efficient and accurate farmland change detection. Attached Figure Description

[0018] Figure 1 This is a hardware structure block diagram of a mobile terminal for a farmland change detection method based on dual-temporal remote sensing images and semantic knowledge enhancement according to an embodiment of the present invention. Figure 2 This is a flowchart illustrating the method for detecting farmland changes based on dual-temporal remote sensing images and semantic knowledge enhancement according to an embodiment of the present invention. Figure 3 This is a schematic diagram illustrating the process of automatically extracting high-confidence masks for cultivated land areas using the SAM model, according to an embodiment of this application. Figure 4 This is a schematic diagram of the architecture of an improved semantic knowledge fusion model based on mask attention, according to an embodiment of this application.

[0019] The above figures include the following reference numerals: 102. Processor; 104. Memory; 106. Transmission device; 108. Input / output device. Detailed Implementation

[0020] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of the present application, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative effort are within the scope of protection of the present application.

[0021] It should be noted that the terms "first," "second," etc., used in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be used interchangeably where appropriate for the embodiments of this application described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0022] Example: As introduced in the background section, existing technologies face challenges in terms of high detail, spectral characteristic analysis, and data registration of multi-temporal remote sensing data. Farmland change detection also faces many technical difficulties. To solve the above problems, this embodiment provides a farmland change detection method based on dual-temporal remote sensing imagery and semantic knowledge enhancement.

[0023] The methods and embodiments provided in this application can be executed on a mobile terminal, computer terminal, or similar computing device. Taking running on a mobile terminal as an example, Figure 1 This is a hardware structure block diagram of a mobile terminal for a farmland change detection method based on dual-temporal remote sensing and deep learning, according to an embodiment of the present invention. Figure 1 As shown, a mobile terminal may include one or more ( Figure 1 Only one is shown in the diagram. A processor 102 (which may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.) and a memory 104 for storing data are also shown. The mobile terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the mobile terminal described above. For example, the mobile terminal may also include components that are more... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.

[0024] The memory 104 can be used to store computer programs, such as application software programs and modules, like the computer program corresponding to the dual-temporal remote sensing and deep learning method for detecting farmland changes in this embodiment of the invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, thereby implementing the above-described method. The memory 104 may include high-speed random access memory and non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the mobile terminal via a network. Examples of the aforementioned networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. The transmission device 106 is used to receive or send data via a network. Specific examples of the aforementioned networks may include wireless networks provided by the mobile terminal's communication provider. In one instance, the transmission device 106 includes a Network Interface Controller (NIC), which can be connected to other network devices via a base station to communicate with the Internet. In one example, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0025] This embodiment provides a method for detecting farmland changes based on dual-temporal remote sensing images and semantic knowledge enhancement, which runs on a mobile terminal, computer terminal, or similar computing device. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Also, although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than that shown here.

[0026] Figure 2 This is a flowchart illustrating a method for detecting farmland changes based on dual-temporal remote sensing imagery and semantic knowledge enhancement, according to an embodiment of this application. Figure 2 As shown, the method includes the following steps: Step S1: Acquire high-resolution dual-temporal remote sensing images and thematic data from the same period as the remote sensing images. The thematic data includes permanent basic farmland data, annual cultivated land datasets, and points of interest data.

[0027] In one or more preferred embodiments, the high-resolution dual-temporal remote sensing imagery is acquired via remote sensing satellites; the thematic data contemporaneous with the remote sensing imagery can be acquired through various means, among which the permanent basic farmland database is usually released by authoritative institutions such as the Ministry of Agriculture; the annual arable land dataset adopts the China Annual Arable Land Dataset, which is open data generated based on nationwide arable land coverage analysis; Points of Interest (POI) data can be obtained using the location and classification information of commercial platforms such as Baidu Maps and Gaode Maps, and the points of interest for agricultural and non-agricultural land within the target area can be obtained through API or manual filtering, and stored and managed in a standard GIS format.

[0028] Step S2: Extract the center point and boundary point of the permanent basic farmland data, as well as the center point and boundary point of the annual cultivated land data as positive sample points of cultivated land, and extract the non-cultivated land interest point data from the interest point data as negative sample points of cultivated land.

[0029] In one or more preferred embodiments, step S2 specifically includes the following steps: S21. Based on the scope of permanent basic farmland data acquisition, determine the center point and the points in the four directions furthest from the center point, as well as the center point in the annual cultivated land dataset, as positive sample points for cultivated land. Specifically, this includes: 1) Perform geometric correction on the dual-temporal remote sensing images to correct geographical location deviations and align the dual-temporal remote sensing images with the actual geographic coordinates; 2) Obtain the set of boundary coordinate points of the dual-temporal remote sensing image ( x 1, y 1),( x 2, y 2),…,( x n , y n ); 3) Obtain the center point using the centroid formula. C x , C y The formula for calculating ) is: , ; in, A The area of ​​the polygon: ; 4) With the center point ( C x , C y Using the center point as a reference, define four directions, and find the point farthest from the center point in each direction. As a boundary point; S22. Use the center points extracted from the non-arable land interest point data as negative sample points for arable land.

[0030] Step S3: Generate a mask for the dual-temporal remote sensing image using positive and negative farmland sample points.

[0031] In one or more preferred embodiments, the mask can be generated using methods such as SAM, SegGPT, Grounded-SAM, Diffusion-based, SAM2, and RS-SAM.

[0032] Taking the automatic extraction of high-confidence masks for cultivated land areas using the SAM (Segment Anything Model) model as an example, such as... Figure 3 As shown, it specifically includes: Positive and negative sample points of cultivated land are input into the SAM model as cue information. By utilizing the spectral and spatial features of the image and the positive and negative sample features in the cue information, the cultivated land area is accurately segmented to generate a mask for dual-temporal remote sensing image.

[0033] The SAM model includes a prompt encoder, an image encoder, and a mask decoder. The prompt encoder serves as an input module, encoding prompt information to generate feature representations related to the target region, which can be used to flexibly define regions of interest. The image encoder adopts a large-scale pre-trained ViT (Vision Transformer) architecture, focusing on the extraction of high-dimensional features and the ability to adapt to cross-resolution images. It captures complex contextual relationships in dual-temporal remote sensing images through a multi-layer ViT architecture. The mask decoder combines the features output by the cue encoder and the image encoder, and uses an autoregressive or parallel generation method to generate the final high-confidence segmentation mask, clearly defining the extent of the cultivated land area.

[0034] Step S4: Calculate the weights of the mask of the dual-temporal remote sensing image using a weighted mask self-attention mechanism to obtain a weighted dual-temporal feature map and extract more concentrated farmland features.

[0035] In one or more preferred embodiments, step S4 specifically includes the following steps: S41. Generate a set of mask weight matrices using the mask, where the value of each pixel position represents whether the position belongs to a cultivated land area and its importance, as shown in the following formula: W mask(x,y) = Sigmoid ( M( x , y )* α ); in, W mask(x,y) It is the first x , y The mask weight matrix of the position, M ( x , y ) represents the pixel value of the mask image; 1 represents farmland, and 0 represents the background. α It is a learnable scalar parameter; S42. Utilize the mask weight matrix to weight the spatiotemporal features of dual-temporal remote sensing images to highlight the changing characteristics of cultivated land areas and suppress noise in background areas, as shown in the following formula: ; in, F ( i , j (This is a feature map extracted from dual-temporal remote sensing images.) W mask(i,j) It is the mask weight matrix generated by mask attention. F’ It is a weighted dual-temporal feature map; through the above mechanism, the mask attention module realizes the deep integration of farmland semantic knowledge and spatiotemporal features.

[0036] Furthermore, under the weighted mask self-attention mechanism, by introducing a multi-level scaling factor with learnable coefficient multiplication, the scalar parameters in the mask weight matrix are dynamically adjusted. α Higher weights are assigned to high-confidence regions, while lower weights are assigned to low-confidence regions to mitigate their impact. Specifically: By introducing a learnable scaling factor, the mask weight matrix is ​​dynamically adjusted. W mask scalar parameters in α This allows it to adaptively magnify high-confidence areas and suppress low-confidence areas. Multi-level scaling factors are designed to cover different spatial scales, with a global scaling factor capturing the overall trend of cultivated land distribution and local scaling factors targeting details within patches. It is linked to the temporal characteristics of differences between dual-temporal images, automatically triggering [a feature] when significant differences are detected (e.g., cultivated land converted to construction land). α Increasing the parameter amplifies the effect. W mask Suppress weights for stable regions (such as unchanged cultivated land). Ultimately, the goal is to constrain the spatial distribution of features using supervisory signals and extract features from highly concentrated cultivated land.

[0037] Step S5: Input the weighted dual-temporal feature map and dual-temporal remote sensing image into the multi-scale spatial feature extraction module based on FPN. The dual-temporal spatial feature pyramid is obtained by downsampling, upsampling and element-wise addition of upper and lower layers and lateral connection.

[0038] In one or more preferred embodiments, the multi-scale spatial feature extraction module based on FPN performs three downsampling convolutions on the input to obtain four layers of features; after each layer of features is bilinearly upsampled, it is added element-wise with the features of the previous layer, and then the upsampling artifacts are eliminated by 3×3 convolution. At the same time, by horizontally connecting the features of each layer, a dual temporal spatial feature pyramid with different resolutions (1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the original image, respectively) is formed from top to bottom. Furthermore, to enhance the detection capability of small-scale farmland boundaries, parallel dilated convolutional layers with different dilation rates are embedded in downsampling×1 and downsampling×2. Dilated convolution captures multi-scale contextual information without increasing the number of parameters by expanding the receptive field.

[0039] Step S6: Using the temporal feature extraction module that takes into account both temporal phases, the difference regions of the spatial feature pyramid of both temporal phases are dynamically weighted by channel attention to obtain the weighted feature map of both temporal phases. In one or more preferred embodiments, the temporal feature extraction module considering both temporal phases includes a splicing layer, a global average pooling layer, and a weight generation layer, specifically: The splicing layer is used to integrate the dual-temporal spatial feature pyramid. After being concatenated along the channel dimension, the data is input into the global average pooling layer. The global average pooling layer is used to generate channel description vectors. And input it into the weight generation layer; The weight generation layer consists of two fully connected layers (with a middle dimension of C / 4) and a non-linear activation function, as shown below: W c =σ( W 2∙ δ ( W 1∙ z )); in, , For learnable fully connected layer parameters, δ σ is the ReLU function, and σ is the Sigmoid function. W c For the generated channel weights; Finally, the channel weights are multiplied channel by channel in the dual-temporal spatial feature pyramid to highlight the channels with changes and suppress noise, resulting in a dual-temporal weighted feature map.

[0040] Step S7: Use the spatiotemporal self-attention module to perform attention calculation on the dual-temporal weighted feature map to obtain the spatiotemporal feature map.

[0041] In one or more preferred embodiments, the spatiotemporal self-attention module outputs features that fuse global spatiotemporal relationships, directly introducing global change features into the feature interaction stage. This significantly improves the ability to perceive changing regions, and is particularly suitable for change detection tasks that require attention to differences between two time phases, as detailed below: The input to the spatiotemporal self-attention module is a dual-temporal weighted feature map with dimensions H×W×C, corresponding to the spatial height, width, and number of channels of the dual-temporal remote sensing image, respectively. It includes feature projection, multi-head attention calculation, and fusion output, specifically: 1) Feature projection: Mapping the input to a linear transformation. Q , K and V The dimension of the projection matrix is C × d k ,in d k It is the dimension of each attention head; 2) Multi-head attention calculation: The global correlation between features is calculated using a multi-head self-attention mechanism, as shown in the following formula: ; in, Attention ( Q , K , V The attention weights obtained are used to capture the feature interaction information of dual-temporal remote sensing images in the spatiotemporal dimension; 3) Fusion output: The original features are preserved through residual connections, and the attention output is linearly transformed in channel dimension to match the dimension of the input; to enhance the local feature representation capability, a spatiotemporal feature map (consistent with the input, H×W×C) is output through a convolutional layer (3×3 kernel size, the number of channels remains unchanged) and the ReLU activation function. Furthermore, the input feature resolution of the spatiotemporal self-attention module is controlled within the range of H / 4×W / 4 by downsampling, thereby improving inference efficiency while ensuring performance.

[0042] Furthermore, the parameter configuration of the number of channels and the number of attention heads was optimized: 8 attention heads were selected, with each head having a dimension of 64.

[0043] Step S8: Input the spatiotemporal feature map into the fully connected layer to obtain the farmland change detection results.

[0044] In one or more preferred embodiments, step S8 specifically comprises: The spatiotemporal feature map is processed through three or more convolutional layers to extract high-level features, and a two-dimensional change probability map is output through a fully connected layer, representing the probability that each pixel belongs to a change region. After the two-dimensional change probability map is generated, a change decision needs to be made for each pixel based on the probability value to determine whether it belongs to a change region. After setting a threshold, pixels with probability values ​​greater than the threshold are classified as change regions, and pixels with probability values ​​less than the threshold are classified as non-change regions, thus obtaining the farmland change detection result.

[0045] like Figure 4 As shown, steps S4-S8 together constitute the core network of this embodiment, which is called an improved semantic knowledge fusion model based on mask attention.

[0046] In another embodiment of this application, a farmland change detection system based on dual-temporal remote sensing images and semantic knowledge enhancement is provided. The system includes a data acquisition module, a mask generation module, a weighted mask self-attention mechanism, a multi-scale spatial feature extraction module based on FPN, a temporal feature extraction module that takes into account dual temporal phases, a spatiotemporal self-attention module, and an output module. The data acquisition module is used to acquire dual-temporal remote sensing images, as well as thematic data from the same period as the dual-temporal remote sensing images; thematic data includes permanent basic farmland data, annual cultivated land datasets, and point-of-interest data; The mask generation module is used to extract the center point and boundary point of permanent basic farmland data and the center point of annual cultivated land dataset as cultivated land positive sample points, and to extract non-cultivated land interest point data from the interest point data as cultivated land negative sample points; and to generate a mask for dual-temporal remote sensing image using cultivated land positive sample points and cultivated land negative sample points. The weighted mask self-attention mechanism is used to calculate the weights of the mask of the dual-temporal remote sensing image to obtain a weighted dual-temporal feature map. The multi-scale spatial feature extraction module based on FPN takes a weighted dual-temporal feature map and dual-temporal remote sensing image as input, and extracts the dual-temporal spatial feature pyramid by downsampling, upsampling and element-wise addition from top to bottom and horizontal connection. The time feature extraction module that takes into account both temporal phases is used to dynamically weight the difference regions of the spatial feature pyramid of both temporal phases through channel attention to obtain a weighted feature map of both temporal phases. The spatiotemporal self-attention module is used to perform attention calculations on the bi-temporal weighted feature map to obtain the spatiotemporal feature map; The output module is used to input the spatiotemporal feature map into the fully connected layer to obtain the farmland change detection results.

[0047] It should be noted that the system provided in the above embodiments is only an example of the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure can be divided into different functional modules to complete all or part of the functions described above. The system is applied to the farmland change detection method based on dual-temporal remote sensing images and semantic knowledge enhancement in the above embodiments.

[0048] In another embodiment of this application, a storage medium is also provided, storing a program that, when executed by a processor, implements the farmland change detection method based on dual-temporal remote sensing imagery and semantic knowledge enhancement described in the above embodiments, specifically as follows: S1. Acquire dual-temporal remote sensing images, as well as thematic data from the same period as the dual-temporal remote sensing images; thematic data includes permanent basic farmland data, annual cultivated land datasets, and point-of-interest data; S2. Extract the center point and boundary point of permanent basic farmland data and the center point of annual cultivated land data as positive cultivated land sample points, and extract the non-cultivated land interest point data from the interest point data as negative cultivated land sample points. S3. Generate a mask for dual-temporal remote sensing images using positive and negative sample points of cultivated land. S4. The weighted mask self-attention mechanism is used to calculate the weight of the mask of the dual-temporal remote sensing image to obtain the weighted dual-temporal feature map. S5. Input the weighted dual-temporal feature map and dual-temporal remote sensing image into the multi-scale spatial feature extraction module based on FPN. The dual-temporal spatial feature pyramid is obtained by downsampling, upsampling and element-wise addition from top to bottom and horizontal connection. S6. Using the temporal feature extraction module that takes into account both temporal phases, the difference regions of the spatial feature pyramid of both temporal phases are dynamically weighted by channel attention to obtain the weighted feature map of both temporal phases. S7. Use the spatiotemporal self-attention module to perform attention calculation on the dual-temporal weighted feature map to obtain the spatiotemporal feature map; S8. Input the spatiotemporal feature map into the fully connected layer to obtain the farmland change detection results.

[0049] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0050] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any changes, modifications, substitutions, combinations, or simplifications made without departing from the spirit and principle of the present invention shall be considered equivalent substitutions and shall be included within the protection scope of the present invention.

Claims

1. A method for detecting farmland change based on dual-temporal remote sensing imagery and semantic knowledge enhancement, characterized in that, Includes the following steps: Acquire bi-temporal remote sensing images, as well as thematic data from the same period as the bi-temporal remote sensing images; thematic data includes permanent basic farmland data, annual cultivated land datasets, and point-of-interest data; A mask for dual-temporal remote sensing imagery is generated using positive and negative sample points of cultivated land. The weighted mask self-attention mechanism is used to calculate the weights of the mask of the dual-temporal remote sensing image to obtain the weighted dual-temporal feature map; The weighted dual-temporal feature map and dual-temporal remote sensing image are input into the multi-scale spatial feature extraction module based on FPN. The dual-temporal spatial feature pyramid is obtained by downsampling, upsampling and element-wise addition from top to bottom and horizontal connection. By utilizing the temporal feature extraction module that takes into account both temporal phases, the difference regions of the spatial feature pyramid of both temporal phases are dynamically weighted by channel attention to obtain the weighted feature map of both temporal phases; The spatiotemporal self-attention module is used to perform attention calculation on the bi-temporal weighted feature map to obtain the spatiotemporal feature map; The spatiotemporal feature map is input into the fully connected layer to obtain the farmland change detection results.

2. The method for detecting farmland change based on dual-temporal remote sensing imagery and semantic knowledge enhancement according to claim 1, characterized in that, The process involves extracting the center points and boundary points of permanent basic farmland data, as well as the center points of the annual cultivated land dataset, as positive cultivated land sample points, and extracting non-cultivated land interest point data from the interest point data as negative cultivated land sample points. Specifically: Based on the scope of permanent basic farmland data acquisition, a center point and the points in the four directions furthest from the center point, as well as the center point in the annual cultivated land dataset, are determined as positive sample points for cultivated land. Specifically, this includes: 1) Perform geometric correction on the dual-temporal remote sensing images to correct geographical location deviations and align the dual-temporal remote sensing images with the actual geographic coordinates; 2) Obtain the set of boundary coordinate points of the dual-temporal remote sensing image ( x 1, y 1),( x 2, y 2),…,( x n , y n ); 3) Obtain the center point using the centroid formula. C x , C y The formula for calculating ) is: ; ; in, A The area of ​​the polygon: ; 4) With the center point ( C x , C y Using the center point as a reference, define four directions, and find the point farthest from the center point in each direction. As a boundary point; The center points extracted from the non-arable land interest point data were used as negative sample points for arable land.

3. The method for detecting farmland change based on dual-temporal remote sensing imagery and semantic knowledge enhancement according to claim 1, characterized in that, Input the positive and negative sample points of cultivated land into the SAM model to generate a mask for dual-temporal remote sensing images. The SAM model includes a cue encoder, an image encoder, and a mask decoder; The prompt encoder is used to encode the input positive and negative farmland sample points to generate feature representations related to the target area. The image encoder employs a pre-trained ViT architecture to capture contextual relationships in dual-temporal remote sensing images; The mask decoder combines the features output by the cue encoder and the image encoder, and generates the final segmentation mask using an autoregressive or parallel generation method.

4. The method for detecting farmland change based on dual-temporal remote sensing imagery and semantic knowledge enhancement according to claim 1, characterized in that, The weighted mask self-attention mechanism is used to calculate the weights of the mask of the dual-temporal remote sensing image to obtain a weighted dual-temporal feature map, specifically as follows: A set of mask weight matrices is generated using the mask, where the value of each pixel position represents whether the position belongs to a cultivated land area and its importance, as shown in the following formula: W mask(x,y) = Sigmoid ( M ( x , y )* α ); in, W mask(x,y) It is the first x , y The mask weight matrix of the position, M ( x , y ) represents the pixel value of the mask image; 1 represents farmland, and 0 represents the background. α It is a scalar parameter; The spatiotemporal features of the dual-temporal remote sensing images are weighted using a mask weight matrix, as shown in the following equation: ; in, F ( i , j (This is a feature map extracted from dual-temporal remote sensing images.) W mask(i,j) It is the mask weight matrix generated by mask attention. F’ It is a weighted dual-phase feature map.

5. The method for detecting farmland change based on dual-temporal remote sensing imagery and semantic knowledge enhancement according to claim 4, characterized in that, By introducing a multi-level scaling factor for coefficient multiplication, the scalar parameters in the mask weight matrix are dynamically adjusted. α .

6. The method for detecting farmland change based on dual-temporal remote sensing imagery and semantic knowledge enhancement according to claim 1, characterized in that, The multi-scale spatial feature extraction module based on FPN performs three downsampling convolutions on the input to obtain four layers of features. After bilinear upsampling of each layer of features, it is added element-wise with the features of the previous layer, and then the upsampling artifacts are eliminated by 3×3 convolution. At the same time, by connecting the features of each layer laterally, a dual temporal spatial feature pyramid with different resolutions is formed from top to bottom. Parallel dilated convolutional layers with different dilation rates are embedded in downsampling×1 and downsampling×2.

7. The method for detecting farmland change based on dual-temporal remote sensing imagery and semantic knowledge enhancement according to claim 1, characterized in that, The temporal feature extraction module considering both temporal phases includes a splicing layer, a global average pooling layer, and a weight generation layer, specifically: The splicing layer is used to integrate the dual-temporal spatial feature pyramid. After being concatenated along the channel dimension, the data is input into the global average pooling layer. The global average pooling layer is used to generate channel description vectors. And input it into the weight generation layer; The weight generation layer comprises two fully connected layers and a nonlinear activation function, as shown below: W c =σ( W 2∙ δ ( W 1∙ z )); in, , These are the parameters for the fully connected layer. δ σ is the ReLU function, and σ is the Sigmoid function. W c For the generated channel weights; Finally, the channel weights are multiplied channel by channel by the dual-temporal spatial feature pyramid to obtain the dual-temporal weighted feature map.

8. The method for detecting farmland change based on dual-temporal remote sensing imagery and semantic knowledge enhancement according to claim 1, characterized in that, The input to the spatiotemporal self-attention module is a dual-temporal weighted feature map with dimensions H×W×C, corresponding to the spatial height, width, and number of channels of the dual-temporal remote sensing image, respectively. It includes feature projection, multi-head attention calculation, and fusion output, specifically: 1) Feature projection: Mapping the input to a linear transformation. Q , K and V The dimension of the projection matrix is C × d k ,in d k It is the dimension of each attention head; 2) Multi-head attention calculation: The global correlation between features is calculated using a multi-head self-attention mechanism, as shown in the following formula: ; in, Attention ( Q , K , V The attention weights are the obtained values. 3) Fusion output: The original features are preserved through residual connections, and the attention output is linearly transformed in channel dimension to match the dimension of the input. Finally, the spatiotemporal feature map is output through a convolutional layer and a ReLU activation function. The spatiotemporal self-attention module controls the input resolution through downsampling.

9. The method for detecting farmland change based on dual-temporal remote sensing imagery and semantic knowledge enhancement according to claim 1, characterized in that, The spatiotemporal feature map is processed through three or more convolutional layers to extract high-level features, and a two-dimensional change probability map is output through a fully connected layer, representing the probability that each pixel belongs to a change region. After the two-dimensional change probability map is generated, a change decision is made for each pixel based on the probability value to determine whether it belongs to a change region. After setting a threshold, pixels with probability values ​​greater than the threshold are classified as change regions, and pixels with probability values ​​less than the threshold are classified as non-change regions, thus obtaining the farmland change detection result.

10. A farmland change detection system based on dual-temporal remote sensing imagery and semantic knowledge enhancement, characterized in that: The method for detecting farmland changes based on dual-temporal remote sensing images and semantic knowledge enhancement, as applied in any one of claims 1-9, includes a data acquisition module, a mask generation module, a weighted mask self-attention mechanism, a multi-scale spatial feature extraction module based on FPN, a temporal feature extraction module that takes into account dual temporal phases, a spatiotemporal self-attention module, and an output module. The data acquisition module is used to acquire dual-temporal remote sensing images, as well as thematic data from the same period as the dual-temporal remote sensing images; thematic data includes permanent basic farmland data, annual cultivated land datasets, and point-of-interest data; The mask generation module is used to extract the center point and boundary point of permanent basic farmland data and the center point of annual cultivated land dataset as cultivated land positive sample points, and to extract non-cultivated land interest point data from the interest point data as cultivated land negative sample points; and to generate a mask for dual-temporal remote sensing image using cultivated land positive sample points and cultivated land negative sample points. The weighted mask self-attention mechanism is used to calculate the weights of the mask of the dual-temporal remote sensing image to obtain a weighted dual-temporal feature map. The multi-scale spatial feature extraction module based on FPN takes a weighted dual-temporal feature map and dual-temporal remote sensing image as input, and extracts the dual-temporal spatial feature pyramid through downsampling, upsampling-elemental summation and lateral connection. The time feature extraction module that takes into account both temporal phases is used to dynamically weight the difference regions of the spatial feature pyramid of both temporal phases through channel attention to obtain a weighted feature map of both temporal phases. The spatiotemporal self-attention module is used to perform attention calculations on the bi-temporal weighted feature map to obtain the spatiotemporal feature map; The output module is used to input the spatiotemporal feature map into the fully connected layer to obtain the farmland change detection results.

Citation Information

Cited By

  • Image change detection method and device, electronic equipment and storage medium

    CN121074038A