A method for detecting farmland non-agriculturalization by combining attention mechanism and Mamba
By combining attention mechanism and Mamba's non-agricultural detection method for cultivated land, the pseudo-change and insufficient feature extraction capabilities in the existing technology are solved, and more accurate and clear change detection results are achieved.
Patent Information
- Application Number
- CN202411178608.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-27
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-08-27
AI Technical Summary
The prior art has problems of pseudo-change and insufficient feature extraction capabilities in the non-agriculturalization detection of cultivated land, resulting in inaccurate detection results and rough edges of the change areas.
The non-agricultural detection method of arable land that combines attention mechanism and Mamba is adopted. The features of different levels and scales are retained through the feature extraction module, the deep semantic information and spatial detail information are combined, and the dual-time phase characteristics are fused through the cross attention mechanism, and the semantic segmentation is guided by using a binarized change mask.
Effectively capture global dependencies, improve the detection ability of small targets, clearly define the change areas, and improve the accuracy and accuracy of change detection.
Smart Images

Figure CN119169478B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of change detection technology, and in particular to a method for detecting the non-agriculturalization of cultivated land by combining an attention mechanism and Mamba. Background Art
[0002] Non-agriculturalization of cultivated land refers to the use of cultivated land for production and management activities other than agricultural production. It is a process of converting land originally used for agricultural production into non-agricultural use. With the continuous acceleration of urbanization, the red line of cultivated land is facing the threat of non-agriculturalization. In many areas, the non-agriculturalization of cultivated land has become a common phenomenon, which not only affects agricultural production, but also may have a negative impact on the ecological environment and biodiversity. At present, there are problems such as illegal occupation of cultivated land for afforestation, excessive construction of green corridors, occupation of permanent basic farmland, illegal digging of lakes for landscaping, expansion of nature reserves, illegal non-agricultural construction, and illegal and irregular land use. The government has issued a series of policy documents aimed at resolutely stopping the non-agriculturalization of cultivated land.
[0003] The non-agriculturalization of cultivated land not only affects the effective utilization of cultivated land resources, but also poses a threat to food production security. Therefore, how to effectively identify and stop the non-agriculturalization of cultivated land is an urgent problem to be solved. At present, with the continuous advancement of earth observation technology, the acquisition of remote sensing data has shown the characteristics of real-time, economical and efficient. A large amount of remote sensing data of different phases, scales, spectra and resolutions has been accumulated, which provides unprecedented data support for the monitoring of non-agriculturalization of cultivated land.
[0004] Change detection is to quantitatively analyze and determine the characteristics and processes of surface changes from remote sensing data of different periods. In recent years, with the continuous development of deep learning and the excellent feature extraction ability of deep learning in the field of image processing, deep learning has been widely used in the field of remote sensing image change detection. At present, remote sensing image change detection methods based on deep learning can be roughly divided into three categories from the perspective of network structure. One method adopts a single encoder-decoder structure. This method fuses two images of different phases and then inputs the fused image into a single-branch encoder-decoder network. This method uses the encoder to extract the deep features of the fused image and reconstructs the changed area through the decoder. The changes in the surface are identified by analyzing the difference between the features extracted by the encoder and the original image. Another method adopts a multi-encoder and single decoder structure. Each image of different phases is processed by an independent encoder, and its respective feature representation is extracted respectively. Then, the features are fused and sent to a shared decoder. The decoder combines the feature information from different time points to identify and map the changed area. The third type of method adopts a dual encoder-decoder structure. Each phase of the image is equipped with an independent encoder-decoder, which extracts the features of the image and fuses and analyzes them through the decoder. The difference between the features of two time points is quantified through a distance measurement mechanism, such as Euclidean distance or cosine similarity.
[0005] Although deep learning-based methods have achieved good results, there are still some challenges. When using a weight-sharing twin network to extract features from images of two phases separately, the interactive information and spatiotemporal dependencies between different phases are ignored, resulting in pseudo-changes in the detection results due to changes in imaging conditions. Traditional convolutional neural networks are affected by their limited receptive fields when extracting features, which limits their ability to capture global contextual information and cannot fully extract the features of dual-phase remote sensing images. In addition, continuous downsampling during feature extraction will lead to the loss of accurate spatial information, and the inability to effectively locate small targets and capture detailed features, resulting in problems such as missed detection and rough edges of changed areas. Therefore, a remote sensing image change detection method that can accurately locate the changed area and have clear edges of the changed area is crucial for the detection of cultivated land non-agriculturalization.
[0006] Although deep learning-based methods have made some progress, most existing research methods still have the following problems when detecting farmland non-agriculturalization:
[0007] First, when using a weight-sharing twin network to extract features from images of two phases separately, the interactive information and spatiotemporal dependencies between different phases are ignored, resulting in pseudo-changes in the detection results due to changes in imaging conditions.
[0008] Second, traditional convolutional neural networks are affected by their limited receptive field when extracting features, which limits their ability to capture global contextual information and cannot fully extract the features of dual-phase remote sensing images. In addition, continuous downsampling operations during the feature extraction process will lead to the loss of accurate spatial information, and cannot effectively locate small targets and capture detailed features, resulting in problems such as missed detections and rough edges of changed areas. Summary of the invention
[0009] The present invention provides a method for detecting non-agriculturalization of cultivated land by combining an attention mechanism and Mamba to solve the problems raised in the above background technology.
[0010] A method for detecting farmland non-agriculturalization by combining attention mechanism and Mamba, comprising the following steps:
[0011] Step 1: Obtain non-agricultural detection data set;
[0012] Step 2: Perform data enhancement on the non-agricultural detection data set in step 1;
[0013] Step 3: Build a network model for farmland non-agriculturalization detection based on attention mechanism and Mamba;
[0014] Step 4: input the enhanced data of step 2 into the farmland non-agriculturalization detection network model of step 3, train and verify the farmland non-agriculturalization detection network model; obtain the trained farmland non-agriculturalization detection network model;
[0015] Step 5: Using the test set in the non-agricultural detection data set, the pre-processed data is input into the trained farmland non-agricultural detection network model to obtain the farmland non-agricultural detection results.
[0016] Beneficial effects of the present invention:
[0017] 1. A method for detecting the non-agriculturalization of cultivated land of the present invention retains features of different scales at different levels through a feature extraction module, integrates deep semantic information and spatial detail information, and thoroughly scans information in all directions through selective scanning, which can effectively capture the global dependencies in remote sensing images, not only taking into account targets of different sizes, but also greatly improving the problem of missed detection of small targets. In addition, by adding a SAM stacking attention module after the second, third, and fourth groups of SSM blocks to obtain more discriminative features, the problem of blurred edges in the changed areas of remote sensing images is greatly improved, further improving the accuracy of change detection.
[0018] 2. A method for detecting non-agriculturalization of cultivated land in the present invention fuses the features of two time points through a cross-attention mechanism, thereby better capturing the spatiotemporal dependency of dual-phase images and improving model performance.
[0019] 3. A method for detecting the non-agriculturalization of cultivated land of the present invention uses a binary change mask to guide two semantic segmentation maps of the dual phase as output, which more clearly defines the change area, helps to more accurately identify the non-agriculturalization area of cultivated land, reduces the misjudgment of stable areas, and only analyzes the change area, effectively reducing the amount of calculation. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 It is a flow chart of a method for detecting farmland non-agriculturalization by combining an attention mechanism and Mamba according to the present invention;
[0021] Figure 2 It is a schematic diagram of a farmland non-agriculturalization detection network model based on attention mechanism and Mamba in the method of the present invention;
[0022] Figure 3 This is the schematic diagram of the multi-scale feature extraction module in the farmland non-agriculturalization detection network model based on the attention mechanism and Mamba;
[0023] Figure 4 The schematic diagram of the structure of the SSM block in the farmland non-agriculturalization detection network model based on the attention mechanism and Mamba;
[0024] Figure 5 The schematic diagram of stacking attention modules in the farmland non-agriculturalization detection network model based on attention mechanism and Mamba;
[0025] Figure 6 The schematic diagram of the cross-attention module in the farmland non-agriculturalization detection network model based on the attention mechanism and Mamba;
[0026] Figure 7 It is a schematic diagram of the change detection results of the present invention on the JL-1 data set; DETAILED DESCRIPTION
[0027] Combination Figures 1 to 7 This embodiment describes a method for detecting farmland non-agriculturalization by combining the attention mechanism and Mamba. Figure 1 As shown, the following steps are included:
[0028] Step 1: Obtain a non-agricultural detection dataset, which includes several land use type images. The images are dual-temporal remote sensing images, where the land use types are cultivated land, bare soil, forest / grassland, buildings, and roads. The specific process is as follows:
[0029] Step 1.1: Download the open source dataset JL-1 from the Internet. This dataset is a competition dataset for farmland types. These images come from the Jilin-1 remote sensing satellite and include 6,000 images. Each image is 256×256 pixels in size, composed of RGB channels, and has a spatial resolution of 0.75m / pixel. The JL-1 dataset provides 9 categories of "from to" labels, indicating the type of change between two time points. These categories include cultivated land-road, cultivated land-forest / grassland, cultivated land-building, cultivated land-bare soil, road-cultivated land, forest / grassland-cultivated land, building-cultivated land, bare soil-cultivated land, and unchanged areas. The JL-1 original dataset label categories are converted into dual-temporal land use type labels, namely cultivated land, roads, forest / grassland, buildings, and bare soil. The five land use types are labeled separately in the dual-temporal remote sensing images;
[0030] Step 1.2: Divide the images in the dataset in step 1.1 into training set, test set and validation set in a ratio of 6:3:1; the final dataset includes 4096 training image pairs, 2048 test image pairs and 682 validation image pairs;
[0031] Step 2: Read the non-agricultural detection dataset obtained in step 1 and perform data enhancement on it; specifically:
[0032] The training set data is processed by data enhancement operations such as horizontal flipping, vertical flipping, Gaussian noise and color transformation; during the network training phase, the image data of the training set and the validation set are read, where only the training set data is enhanced, and the validation set is not involved in the training, so the validation set is not enhanced. Both the training set and the validation set data need to be normalized and converted into tensor form;
[0033] Step 3: Construct a network model for farmland non-agriculturalization detection based on attention mechanism and Mamba, such as Figure 2 As shown; the farmland non-agriculturalization detection network model includes: an SSM structure multi-scale feature extraction module, a binary change detection module, a cross attention module, a classifier module and an upsampling module;
[0034] like Figure 3 As shown, in this embodiment, the SSM structure multi-scale feature extraction module includes an image block embedding operation block (Patches Embedding), four layers of SSM blocks and three stacked attention modules (SAM);
[0035] The SSM structure multi-scale feature extraction module uses a Mamba network based on a state space model (State Space Models, SSM) as a feature extractor, and uses two network structures with the same weights to receive the input of dual-phase remote sensing images; Patches Embedding is used to flatten the original two-dimensional image into a one-dimensional vector for the input image, and then four layers of SSM blocks are used to extract information in the image to obtain features of four different scales. The number of SSM blocks in each layer is 2, 2, 9, and 2, respectively, and the sizes of the output features are 1 / 2, 1 / 4, 1 / 8, and 1 / 16 of the input image, respectively, and their channels are 64, 128, 256, and 512, respectively.
[0036] like Figure 4 As shown, the SSM block first normalizes the input layer (LayerNormalization) and then divides the input into two branches. In the first branch, the input passes through a linear layer (LinearLayer); in the second branch, the input is processed by a linear layer (Linear Layer), a depth-wise separable convolution layer (Depth-wise Conv Layer), and then further extracts features through a two-dimensional selective scanning layer (2D Selective Scan Layer) to achieve a global receptive field at the cost of linear complexity. Subsequently, after passing through a linear layer (Linear Layer), the output of the first branch is used to perform element-by-element multiplication to merge the two paths. Finally, a linear layer (Linear Layer) is used to mix the features, and this result is residually connected with the input as the output result. This implementation uses Relu as the activation function by default.
[0037] In this embodiment, the two-dimensional selective scanning layer can implement a scan expansion operation, an S6 block operation, and a scan fusion operation; the scan expansion operation expands the input image into a sequence along four different directions (upper left to lower right, lower left to upper right, lower right to upper left, and upper right to lower left). These sequences are then feature extracted by the S6 block, which is based on a state space model SSM, a discretization for creating loops and convolutional representations, and a selection mechanism introduced on top of S4 composed of HIPPO for processing remote dependencies, and adjusts the parameters of SSM according to the input. This enables the model to distinguish and retain relevant information while filtering out irrelevant information. The input sequence is processed through a series of linear transformations and discretization processes, and each element in the one-dimensional array interacts with any previously scanned sample through a compressed hidden state, effectively reducing the quadratic complexity to linear, thereby ensuring that information in all directions is thoroughly scanned, thereby capturing different features. Subsequently, the scan fusion operation adds and merges the sequences from the four directions, restoring the output image to the same size as the input.
[0038] In this embodiment, the two-dimensional selective scanning layer uses a state space model, which is based on the principles of control theory and is defined by a set of linear ordinary differential equations:
[0039] h′(t)=Ah(t)+Bx(t)
[0040] y(t)=Ch(t)+Dx(t)
[0041] In the formula, input x(t)∈R L , output y(t)∈R, A∈C N×N , B∈R N×L , C∈R N , D∈R L is the system matrix, h(t) represents the hidden state vector at time t. The state transition matrix A of the model controls the evolution of the state vector h(t), while the input matrix B, output matrix C and feedthrough matrix D represent the relationship between input x(t), state h(t) and output y(t), respectively.
[0042] To achieve computational tractability and consistency with the data sampling rate, these continuous equations must be discretized and defined as follows:
[0043] Φ=e AΔT
[0044] Γ=(e AΔT -I)A -1 B
[0045] h k =Φh k-1 +Γx k
[0046] y k =Ch k +Dx k
[0047] where h k is the hidden state at discrete time step k, x k is the input signal, y k is the output state, Φ is the state transition matrix of the time step ΔT. Γ represents the state transition matrix, I represents the identity matrix, and when the input remains constant over each interval ΔT, Γ is derived as (e AΔT -I)A -1 B.
[0048] In this embodiment, three SAM blocks are added to the SSM structure multi-scale feature extraction module, and one SAM is applied to the output features of each layer of SSM blocks after the first layer of SSM blocks to obtain the output feature map of the layer. After obtaining the output feature maps of each layer, these feature maps are adjusted to half the scale of the original input images T1 and T2, three feature maps of the same size are spliced in the channel direction, and the number of channels is adjusted through a 1×1 convolution layer, so as to obtain a more discriminative feature map containing multi-scale information.
[0049] like Figure 5 As shown, in this implementation, each SAM block includes a position attention module (PAM) to capture spatial semantic information and a proxy attention module (PRAM) to implement effective global information; the specific process includes:
[0050] Set the input feature map F i , a tensor of the form C×H×W is fed into PAM, where C is the channel dimension, H is the image height, W is the image width, and F i is the feature map output by the SSM block of the corresponding layer; then the query Q is generated respectively through three convolutional layers i1 , key K i1 Sum value V i1 ;in:
[0051]
[0052] Where W is the corresponding query Q i1 , key K i1 Sum value V i1 The weight of query Q i1 With key K i1 Perform dot product operation to calculate the similarity score, normalize the obtained similarity score through the softmax function to obtain the attention weight (Weight1), and use matrix multiplication to convert the value V i1 Multiply it by Weight1 to get the attention weighted feature and input it into the PAM layer (PAMLayer); then add the attention weighted feature to the input feature map F i Add them element by element and balance them by parameter β, which is a trainable parameter used to balance the input feature map F i And the attention weighted features, the final output feature map is expressed as follows:
[0053]
[0054] At the same time, the input feature map F i Sent into PRAM, two convolutional layers are used to generate query Q i2 and key Ki2 ,in, And the query Q i2 and key K i2 Perform dot product operation with the proxy vector to calculate the proxy attention score. The proxy vector is set to 128, 256, and 512 in feature extraction. The obtained proxy attention score is normalized by the softmax function to obtain the proxy attention probability (Weight2). i2 Set as the input feature map F i , that is: V i2 =F i ; Set the value V i2 Multiply it with Weight2 and input it into the PRAM layer. The feature map is obtained by summing it up, and finally the output feature map is converted into the form of C×H×W. It can be expressed as:
[0055]
[0056] The feature map output by the PAM and the feature map output by the PRAM are added element by element to obtain the output of the SAM block; the output of the SAM block is then added to the input feature map F i Perform residual connection and obtain the output feature map F by element-by-element addition o , expressed as:
[0057] F o =F i +A(Q i1 ,K i1 ,V i1 )+A(Q i2 K i2 ,V i2 )
[0058] The output feature maps of the three layers are adjusted to half the scale of the original input images T1 and T2, and spliced in the channel direction to obtain the feature map of the phase; that is, the SSM structure multi-scale feature extraction module corresponds to the output T1 phase feature map and T2 phase feature map.
[0059] The obtained T1 phase feature map and T2 phase feature map are input into the binary change detection module. First, the two inputs are spliced in the channel dimension through the splicing operation, and then the feature map is downsampled through the residual module. The specific process includes:
[0060] The input form is set as a tensor feature map of C×H×W. First, the feature dimension is changed to 2C×H×W through the concatenation operation on the channel dimension, and then input into the residual block. There are four residual layers, each of which includes two 3×3 convolutions, two BN layers and two ReLU activation functions; before inputting into the second ReLU activation function, the feature is fused by adding the input feature element by element, and finally the size of the output feature map F becomes C×H×W.
[0061] like Figure 6 As shown, the T1 phase feature map and the T2 phase feature map obtained by the multi-scale feature extraction module are simultaneously input into the cross attention module, and the input feature maps are set to F1 and F2, which generate queries q1 and q2, keys k1 and k2, and values v1 and v2. Cross calculations are performed using q1 and k2 and q2 and k1, respectively, and then Softmax normalization is performed to obtain normalized attention weights Energy1 and Energy2, respectively.
[0062] Where v1 = W v1 F1,q1=W q1 F1, k1 = W k1 F1, k2 = W k2 F2,q2=W q2 F2, v2 = W v2 F2;
[0063] Use the value v2 and the attention weight Energy1 to calculate the weighted features of F1 and input them into the attention layer AttentionLayer1. Then use the value v1 and the attention weight Energy2 to calculate the weighted features of F2 and input them into the attention layer Attention Layer2. Use the learnable parameters γ and μ to scale the weighted features respectively, and then perform element-by-element addition operations with the input feature maps F1 and F2 respectively, and finally output feature maps out1 and out2, which can be expressed as follows:
[0064]
[0065] The feature map after the cross attention module and the binary change detection module is input into the classifier module, which has three classifiers, one binary classifier and two category classifiers. Specifically:
[0066] The feature map F after the binarization change detection module is input into the binarization classifier, which includes two 1×1 convolutions, a BN layer, a ReLU activation function and a Dropout layer, and outputs a binarization mask.
[0067] The dual-phase feature maps out1 and out2 after the cross-attention module are input into the category classifier respectively. The number of channels of the output out1 and out2 feature maps is mapped from 128 to the number of categories in the non-agricultural classification task through the classifier, which is set to 5.
[0068] The upsampling module uses bilinear interpolation to expand the size of the binary mask and the two feature maps to the same size as the input images of the T1 phase and the T2 phase, and the format is a tensor of C×H×W. C is the channel dimension (the number of non-agricultural task categories); in this embodiment, the tensor format of the binary mask is 1×256×256, and the format of the two feature maps is 5×256×256.
[0069] Step 4: Input the data read in step 2 into the farmland non-agriculturalization detection network model constructed in step 3, and train and verify it. The specific process is as follows:
[0070] In this implementation, a combined loss function is used to calculate the error between the predicted value and the true label:
[0071] L=L dice +L bce +L sca +L ce
[0072] Among them, L dice Dice Loss, L dice ) is a commonly used loss function in image segmentation tasks. It is a variant based on the intersection-over-union ratio (IoU) and is used to measure the similarity between the predicted segmented area and the true segmented area. It can be expressed as follows:
[0073]
[0074] L bce Binary Cross-Entropy Loss, L bce ) is a commonly used loss function in binary classification problems. It is used to measure the difference between the probability distribution predicted by the model and the probability distribution of the true label. It can be expressed as follows:
[0075]
[0076] L sca is the semantic change alignment loss (Semantic ChangeAlignment, L sca) is a loss function used to measure the logical correlation between predicted bi-temporal semantic segmentation images. The consistency between the predicted semantic segmentation image and the true change image is measured by calculating the cosine similarity between the change region and the unchanged region in the change image. It can be expressed as:
[0077]
[0078] L ce Categorical Cross-Entropy Loss, L ce ) is a loss function used to handle multi-classification problems. It measures the difference between the probability distribution predicted by the entire model and the probability distribution of the true label. In multi-classification problems, multiple categories each have a corresponding label. It can be expressed as:
[0079]
[0080] In the above formula, y i and Represent the label value and predicted value of pixel i respectively, N is the total number of pixels, y and Represent the true value change map and the predicted binary change map, respectively. and is the predicted dual-time semantic segmentation map, α is the set balance factor, p is the index of each category, y p is the probability of the true category corresponding to category p, is the probability that the predicted category is classified as p.
[0081] In this embodiment, during the training process, the AdamW optimizer is used to update and optimize the network parameters; the farmland non-agriculturalization detection network model adopts four evaluation indicators, namely F1 score (F1-score), mean intersection over union (mIoU), separation kappa coefficient (SeK) and overall accuracy (OA), by calculating the loss function between the change detection result and the true label, using the error back propagation algorithm, the parameters in the model are continuously iterated and optimized until the model converges or the training reaches the iteration number, and the training ends;
[0082] During the validation process, the model is validated using the validation set. The model's predictions are compared with the known outputs of the validation dataset to evaluate the model's performance. This process includes forward propagation, calculating validation loss, and evaluating metrics.
[0083] Step 5: Read the test set in step 1.2 and input it into the trained change detection model to obtain the final cultivated land non-agriculturalization detection result map, as shown in Figure 7As shown, red represents cultivated land, gray represents roads, dark green represents forests and grasslands, light green represents buildings, and blue represents bare soil. The left and right image pairs represent remote sensing images of the same area in two time phases, T1 and T2, respectively. The label pair represents the ground truth, and the prediction result pair represents the model's prediction result of the ground truth. In the figure, (a) shows that the model successfully detected the change from cultivated land to solar panel buildings, and the right area of (b) shows the scene of the change from cultivated land to bridge roads. In (c), the change from cultivated land to bare soil in the entire area was successfully detected. (d) shows the transition from forest to bare soil. The prediction method described in this embodiment is used for prediction, and the results show that the model fully extracts the features of the dual-phase remote sensing images, so that there is a small difference between the semantic segmentation map of the changed area and the label pair. And the changed area is clearly defined, small targets are effectively located, and the problem of rough edges in the changed area is improved.
[0084] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0085] The above-mentioned embodiments only express several implementation methods of the present invention, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.
Claims
1. A method for detecting farmland non-agriculturalization by combining attention mechanism and Mamba, characterized by: The method is implemented by the following steps: Step 1: Obtain non-agricultural detection data set; Step 2: Perform data enhancement on the non-agricultural detection data set in step 1; Step 3: Build a network model for farmland non-agriculturalization detection based on attention mechanism and Mamba; The farmland non-agriculturalization detection network model includes: a multi-scale feature extraction module, a binary change detection module, a cross attention module, a classifier module and an upsampling module; The multi-scale feature extraction module uses a Mamba network based on a state space model SSM as a feature extractor. The extractor uses two network structures with the same weight to receive the input of dual-phase remote sensing images and perform feature extraction, and outputs a T1 phase feature map and a T2 phase feature map; The cross attention module performs cross calculation on the T1 phase feature map and the T2 phase feature map to output dual-phase features, namely, outputting feature map out1 and feature map out2; The binary change detection module splices the T1 phase feature map and the T2 phase feature map in the channel dimension, performs a downsampling operation on the spliced feature map, and outputs a feature map F; The classifier module maps the number of channels of the feature map out1 and the feature map out2 from 128 to the number of categories in the non-agricultural classification task, and outputs the tensor form of the two feature maps; At the same time, the classifier performs a convolution operation on the feature map F and outputs a tensor form of a binary mask; The upsampling module expands the size of the binary mask and the two feature maps to the same size as the T1 phase input image and the T2 phase input image through bilinear interpolation operation, and the tensor forms of the two feature maps are all C×H×W tensors; where C is the channel dimension, H is the image height, and W is the image width; and finally outputs the binary mask, the T1 phase semantic segmentation map, and the T2 phase semantic segmentation map; The feature extractor includes an image block embedding operation block, four layers of SSM blocks and three SAM blocks; the image block embedding operation block is used to flatten the original two-dimensional image into a one-dimensional vector, and the four layers of SSM blocks are used to extract information in the image to obtain features of four different scales; the output features of each layer of SSM blocks after the first layer of SSM blocks use the corresponding SAM blocks, and the features output by the SSM blocks are adjusted to half the size of the original input image through the SAM blocks, and then three feature maps of the same size are spliced in the channel direction, and the number of channels is adjusted through a 1×1 convolution layer to obtain feature maps of multi-scale information; namely: T1 phase feature map and T2 phase feature map; Step 4: input the enhanced data in step 2 into the farmland non-agriculturalization detection network model in step 3, train and verify the farmland non-agriculturalization detection network model; obtain the trained farmland non-agriculturalization detection network model; Step 5: Using the test set in the non-agricultural detection data set, the pre-processed data is input into the trained farmland non-agricultural detection network model to obtain the farmland non-agricultural detection results.
2. According to claim 1, a method for detecting farmland non-agriculturalization by combining attention mechanism and Mamba, characterized in that: In step 2, the data enhancement operation is: processing the training set data using data enhancement operations such as horizontal flipping, vertical flipping, Gaussian noise, and color transformation; In the network training stage, the image data of the training set and the verification set in the non-agricultural detection data set described in step 1 are read, data enhancement operations are performed on the training set data, and both the training set and the verification set data are normalized and converted into tensor form.
3. According to claim 1, a method for detecting farmland non-agriculturalization by combining attention mechanism and Mamba, characterized in that: Each SAM block includes a position attention module PAM and a proxy attention module PRAM; The position attention module PAM is used to capture spatial semantic information; The proxy attention module PRAM is used to achieve effective global information.
4. According to claim 3, a method for detecting farmland non-agriculturalization by combining attention mechanism and Mamba is characterized in that: The specific process of the position attention module PAM outputting the feature map is as follows: The input feature map is input into PAM, and then the query Q is generated through three convolutional layers respectively. i1 , key K i1 Sum value V i1 , for query Q i1 With key K i1 Perform dot product operation to calculate the similarity score, normalize the obtained similarity score through the softmax function to obtain the attention weight, and convert the value V into i1 Multiply it by the attention weight to obtain the attention weighted feature. The PAM layer combines the attention weighted feature with the feature map F i Add each element one by one and balance it by parameter β to finally output the feature map; At the same time, the input feature map is sent to PRAM, and the query Q is generated through the linear layer. i2 and key K i2 , and query Q i2 and key K i2 Perform a dot product operation with the proxy vector to calculate the proxy attention score, and normalize the proxy attention score through the softmax function to obtain the proxy attention probability; i2 Set as input feature, and set the value V i2 After multiplying with the agent attention probability, it is input into the PRAM layer and the weighted feature map is obtained by summing; The feature map output by PAM and the feature map output by PRAM are added element by element as the output of SAM block; the output of SAM block is residually connected with the input feature map, and the output feature map F is obtained by element-by-element addition. o .
5. The method for detecting farmland non-agriculturalization by combining attention mechanism and Mamba according to claim 1, characterized in that: The implementation process of the binary change detection module is as follows: The T1 phase feature map and the T2 phase feature map output by the SSM structure multi-scale feature extraction module are input into the binary change detection module, and the T1 phase feature map and the T2 phase feature map are spliced in the channel dimension through a splicing operation, and then the feature map is downsampled through the residual module, and finally the feature map F is output.
6. The method for detecting farmland non-agriculturalization by combining attention mechanism and Mamba according to claim 1, characterized in that: The implementation process of the cross attention module is: Assume that the two input feature maps are F1 and F2, and output feature maps out1 and out2 through the following formula; Where, queries q1, q2 are generated by F1 and F2, keys k1, k2, values v1, v2, and γ, μ are learnable parameters.
7. The method for detecting farmland non-agriculturalization by combining attention mechanism and Mamba according to claim 1, characterized in that: The classifier module has three classifiers, one binary classifier and two category classifiers; The feature map F that has passed through the binarization change detection module is input into the binarization classifier, and after binarization, a binary mask is output; The dual-phase features after the cross-attention module are correspondingly output to two category classifiers, and the number of channels of the output feature maps out1 and out2 are mapped from 128 to the number of categories in the non-agricultural classification task through the two category classifiers.
Citation Information
Patent Citations
Remote sensing image change detection method based on twinborn multi-scale cross attention
CN117975267A