Model training method and device based on frequency domain feature interaction and pyramid mixed attention, cultivated land non-agrochemical detection method and device, electronic equipment and storage medium

Through the model training method of frequency domain feature interaction and pyramid mixed attention, the missed and mis-checking problems of the farmland change detection model in complex scenarios are solved, and the accurate extraction and recognition of farmland semantic information is achieved.

CN120298889APending Publication Date: 2025-07-11XINJIANG UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510357895.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing farmland change detection model is difficult to accurately extract farmland semantic information in complex scenarios, resulting in missed and missed detection, and the style differences between bi-time phase remote sensing images affect the detection effect.

Method used

A model training method based on frequency domain feature interaction and pyramid hybrid attention is adopted. Through multi-scale feature extraction fusion, Transformer module, pyramid hybrid attention module and decision-level fusion classification module, the deep learning model is optimized, style differences are reduced and semantic information extraction accuracy is improved.

Benefits of technology

Accurately identify the non-agricultural changes in arable land in complex scenarios, reduce missed inspections and missed inspections, and improve the accuracy and robustness of detection of arable land change.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298889A_ABST
    Figure CN120298889A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of remote sensing image detection, in particular to a model training method based on frequency domain feature interaction and pyramid mixed attention, a cultivated land non-agrochemical detection method and device, electronic equipment and a storage medium. The training set is used to train a preset deep learning model, a trained deep learning model is obtained, and the deep learning model comprises a multi-scale feature extraction fusion module, a Transform module, a pyramid mixed attention module and a decision-level fusion classification module; according to the method, characteristics related to cultivated land captured on different scales can be fused, semantic information of the cultivated land is accurately positioned in a complex scene, non-agrochemical changes of the cultivated land in the complex scene are accurately identified, and the problems that the semantic information of the cultivated land in the complex scene is not accurate enough and the cultivated land change is not accurate enough in the existing cultivated land change detection method are solved. And missing detection and false detection are easily caused.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of remote sensing image detection, and is a model training method, a cultivated land non-agriculturalization detection method, a device, an electronic device and a storage medium based on frequency domain feature interaction and pyramid hybrid attention. Background Technique

[0002] Accurately grasping the distribution and changes of cultivated land within a region is not only a need for technological development, but also a necessity for carrying out macro management of agricultural development. With the rapid development of remote sensing technology, in terms of cultivated land protection, remote sensing images play an important role in monitoring and evaluating land use changes.

[0003] Currently, the traditional method for extracting cultivated land from remote sensing images in each regulatory department mainly uses manual visual interpretation. However, with the gradual increase in data volume, especially when facing high-frequency and large-area cultivated land area comparisons, the traditional manual visual interpretation method is difficult to support the explosive growth of workload. Moreover, with the development of technologies such as artificial intelligence and big data, deep learning with the ability to learn massive data and abstract features has been applied in remote sensing image change detection.

[0004] However, the distribution of cultivated land is intertwined and mixed with the distribution of other ground objects such as buildings and forests in terms of spatial distribution. In such complex scenarios, existing cultivated land change detection models cannot effectively extract cultivated land-related features, easily causing missed detections and misclassifications. And due to the inability to maintain the consistency of conditions such as light and weather when obtaining dual-temporal remote sensing images, there is a problem of large style differences between dual-temporal remote sensing images. Existing cultivated land change detection models cannot interact with the features of dual-temporal images, easily leading to missed detections and misclassifications. Summary of the Invention

[0005] The present invention provides a model training method, a cultivated land non-agriculturalization detection method, a device, an electronic device and a storage medium based on frequency domain feature interaction and pyramid hybrid attention, which overcomes the above-mentioned deficiencies of the prior art and can effectively solve the problems existing in the existing cultivated land change detection methods, such as inaccurate semantic information of cultivated land in complex scenarios, easily leading to missed detections and false detections.

[0006] One of the technical solutions of the present invention is achieved by the following measures: A model training method based on frequency domain feature interaction and pyramid hybrid attention, including:

[0007] Obtain a sample set and divide it into a training set, a validation set and a test set according to a ratio, where the sample set includes a number of samples, and each sample includes historical dual-temporal remote sensing images of cultivated land and corresponding cultivated land non-agriculturalization detection result identification information;

[0008] Train a preset deep learning model using a training set to obtain a trained deep learning model, where the deep learning model includes a multi-scale feature extraction and fusion module, a Transformer module, a pyramid hybrid attention module, and a decision-level fusion classification module;

[0009] Use a validation set and a test set to validate and test the trained deep learning model, optimize the deep learning model, and obtain a cultivated land non-agriculturalization detection model.

[0010] The following is a further optimization or / and improvement of the above technical solution of the invention:

[0011] The above-mentioned training of a preset deep learning model using a training set to obtain a trained deep learning model includes:

[0012] Use the multi-scale feature extraction and fusion module to perform multi-scale feature extraction and fusion on the samples in the training set to obtain feature maps of different scales, where the process of multi-scale feature extraction and fusion includes multi-scale feature extraction, channel dimensionality reduction, frequency-domain feature interaction, and multi-scale frequency-domain feature fusion;

[0013] Use the Transformer module to perform self-attention encoding, decoding, and reconstruction on the feature maps of different scales to extract global features;

[0014] Use the pyramid hybrid attention module to fuse the features output by the Transformer module according to different scales to obtain feature maps of each scale after fusion;

[0015] Use the decision-level fusion classification module to perform cultivated land non-agriculturalization detection on the feature maps of each scale after fusion respectively, and fuse the results of each cultivated land non-agriculturalization detection to obtain the final cultivated land non-agriculturalization detection result;

[0016] Optimize and balance the loss between the cultivated land non-agriculturalization detection results of each scale in the training and the cultivated land non-agriculturalization detection result after fusion through an adaptive loss function;

[0017] Loop the above steps, output the training weights that meet the training conditions, and obtain the trained deep learning model.

[0018] The above multi-scale feature extraction and fusion module includes a dual-branch convolutional neural network, a channel dimensionality reduction module, a frequency-domain feature interaction module, and a multi-scale frequency-domain feature fusion module. The use of the multi-scale feature extraction and fusion module to perform multi-scale feature extraction and fusion on the samples in the training set to obtain feature maps of different scales includes

[0019] Use the dual-branch convolutional neural network to extract multi-scale features of the dual-temporal remote sensing images;

[0020] Use a channel dimensionality reduction module to perform channel dimensionality reduction on features of different scales using different convolutions;

[0021] Use a frequency domain feature interaction module to splice features of the same scale but different time phases, and use the fast Fourier transform to transform them into the frequency domain to obtain multi-scale frequency domain features;

[0022] Use a multi-scale frequency domain feature fusion module to splice multi-scale frequency domain features based on a channel fusion branch and a spatial fusion branch respectively, and transform the splicing results of the channel fusion branch and the spatial fusion branch into the frequency domain for feature fusion to obtain feature maps of different scales.

[0023] The above-mentioned obtaining of the sample set and dividing it into a training set, a validation set, and a test set according to a ratio includes:

[0024] Obtain a number of samples, and adjust the sizes of all samples according to the set pixels, where each sample includes the historical double-time-phase remote sensing images of cultivated land and the corresponding identification information of cultivated land non-agriculturalization detection results;

[0025] Perform enhancement processing on the samples to obtain a sample set, where the enhancement methods include random rotation, horizontal flipping, and vertical flipping;

[0026] Divide the sample set into a training set, a validation set, and a test set according to a ratio.

[0027] The second technical solution of the present invention is achieved by the following measures: A method for detecting cultivated land non-agriculturalization from remote sensing images, including:

[0028] Obtain the double-time-phase remote sensing image of the cultivated land to be detected and perform preprocessing on it;

[0029] Input the preprocessed double-time-phase remote sensing image of the cultivated land to be detected into the cultivated land non-agriculturalization detection model to obtain the cultivated land non-agriculturalization detection result, where the cultivated land non-agriculturalization detection model is trained by a model training method based on frequency domain feature interaction and pyramid hybrid attention.

[0030] The third technical solution of the present invention is achieved by the following measures: A model training device based on frequency domain feature interaction and pyramid hybrid attention, including:

[0031] A sample acquisition module, which acquires a sample set and divides it into a training set, a validation set, and a test set according to a ratio, where the sample set includes a number of samples, and each sample includes the historical double-time-phase remote sensing images of cultivated land and the corresponding identification information of cultivated land non-agriculturalization detection results;

[0032] The model training module uses the training set to train a preset deep learning model to obtain a trained deep learning model. The deep learning model includes a multi-scale feature extraction and fusion module, a Transformer module, a pyramid hybrid attention module, and a decision-level fusion classification module;

[0033] The model optimization module uses the validation set and the test set to validate and test the trained deep learning model, optimize the deep learning model, and obtain a cultivated land non-agriculturalization detection model.

[0034] The following is a further optimization or / and improvement of the above-mentioned inventive technical solution:

[0035] The above-mentioned model training module includes:

[0036] The multi-scale feature extraction and fusion module performs multi-scale feature extraction and fusion on the samples in the training set to obtain feature maps of different scales. The process of multi-scale feature extraction and fusion includes multi-scale feature extraction, channel dimensionality reduction, frequency-domain feature interaction, and multi-scale frequency-domain feature fusion;

[0037] The Transformer module performs self-attention encoding, decoding, and reconstruction on the feature maps of different scales to extract global features;

[0038] The pyramid hybrid attention module fuses the features output by the Transformer module according to different scales to obtain the fused feature maps of each scale;

[0039] The decision-level fusion classification module respectively performs cultivated land non-agriculturalization detection on the fused feature maps of each scale, and fuses the results of each cultivated land non-agriculturalization detection to obtain the final cultivated land non-agriculturalization detection result;

[0040] The evaluation module optimizes and balances the loss between the cultivated land non-agriculturalization detection results of each scale and the fused cultivated land non-agriculturalization detection result in the training through an adaptive loss function;

[0041] The output module outputs the training weights that meet the training conditions to obtain the trained deep learning model.

[0042] The fourth aspect of the technical solution of the present invention is achieved by the following measures: A remote sensing image cultivated land non-agriculturalization detection device includes:

[0043] The image acquisition module acquires the bi-temporal remote sensing image of the cultivated land to be detected and preprocesses it;

[0044] The detection module inputs the two - temporal remote sensing images of the cultivated land to be detected after pre - processing into the cultivated land non - agriculturalization detection model, and obtains the cultivated land non - agriculturalization detection result, where the cultivated land non - agriculturalization detection model is trained by a model training method based on frequency - domain feature interaction and pyramid hybrid attention.

[0045] The fifth technical solution of the present invention is realized by the following measures: An electronic device, characterized by comprising a processor and a memory, wherein a computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the steps in the model training method based on frequency - domain feature interaction and pyramid hybrid attention.

[0046] The sixth technical solution of the present invention is realized by the following measures: A storage medium, characterized in that a computer - readable computer program is stored on the storage medium, and the computer program is set to execute the steps in the model training method based on frequency - domain feature interaction and pyramid hybrid attention when running.

[0047] The present invention combines a multi - scale feature extraction and fusion module, a Transformer module, and a pyramid hybrid attention module, and solves the problems existing in the existing cultivated land change detection methods, that is, the semantic information of cultivated land in complex scenarios is not accurate enough, which is prone to missed detection and false detection. Specifically as follows:

[0048] Adopt a convolutional neural network structure with multi - branch parallelism to extract multi - scale features from the two - temporal remote sensing images. While each branch is independent of each other, they interact with each other, making full use of the advantages of different - scale features to facilitate the localization of the semantic features of cultivated land in complex backgrounds.

[0049] Adopt a multi - scale channel dimensionality reduction module to perform channel dimensionality reduction on different scales using different convolutional operations, and at the same time adopt an operation similar to channel attention to highlight important information in the channel dimension.

[0050] Adopt a frequency - domain feature interaction strategy to transform the features in the spatial domain into the frequency domain, and then interact with the features of the two - temporal phases in a way similar to spatial attention and channel attention. Adopting this strategy can minimize the style differences between the two - temporal remote sensing images to achieve the purpose of unified style, thereby suppressing non - semantic changes and missed detections caused by inconsistent styles.

[0051] A multi-scale frequency domain feature fusion module is adopted, and the downsampling operation is replaced by wavelet transform. Wavelet transform can better preserve detailed information. Moreover, this module is provided with a spatial branch and a channel branch. The channels most closely related to the variation are obtained on the channel branch. On the spatial branch, the present invention uses wavelet convolution to expand the receptive field as much as possible. Wavelet convolutions with different receptive fields are used for different scales, so as to adapt to features of different scales. Then, the features of the spatial branch and the channel branch are fused in the frequency domain, which can fuse the features related to cultivated land captured at different scales, facilitating the positioning of the semantic information of cultivated land in the reproduction scenario, and thus identifying the changes of cultivated land in micro regions and complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Appendix Figure 1 FIG. is a schematic diagram of an implementation environment provided by the present invention.

[0053] Appendix Figure 2 FIG. is a schematic diagram of the flow of the method for training the cultivated land non-agriculturalization detection model provided by the present invention.

[0054] Appendix Figure 3 FIG. is a schematic diagram of the flow of the method for obtaining the sample set provided by the present invention.

[0055] Appendix Figure 4 FIG. is a schematic diagram of the flow of the method for training the deep learning model provided by the present invention.

[0056] Appendix Figure 5 FIG. is a schematic diagram of the structure of a deep learning model that can extract features from three scales provided by the present invention.

[0057] Appendix Figure 6 FIG. is a schematic diagram of the structure of the channel dimensionality reduction module provided by the present invention.

[0058] Appendix Figure 7 FIG. is a schematic diagram of the structure of the frequency domain feature interaction module provided by the present invention.

[0059] Appendix Figure 8 FIG. is a schematic diagram of the structure of the multi-scale frequency domain feature fusion module provided by the present invention.

[0060] Appendix Figure 9 FIG. is a schematic diagram of the structure of the channel attention provided by the present invention.

[0061] Appendix Figure 10 FIG. is a schematic diagram of the structure of the spatial attention provided by the present invention.

[0062] Appendix Figure 11 FIG. is a schematic diagram of the structure of the pyramid hybrid attention module provided by the present invention.

[0063] Appendix Figure 12 FIG. is a schematic diagram of the structure of the decision-level fusion classification module provided by the present invention.

[0064] Appendix Figure 13 It is a schematic structural diagram of the classifier provided by the present invention.

[0065] Appendix Figure 14 It is a schematic diagram of the verification result provided by the present invention.

[0066] Appendix Figure 15 It is a schematic flow diagram of the method for detecting non-agricultural conversion of cultivated land in remote sensing images provided by the present invention.

[0067] Appendix Figure 16 It is a schematic structural diagram of the device for training the non-agricultural conversion detection model of cultivated land provided by the present invention.

[0068] Appendix Figure 17 It is a schematic structural diagram of the device for detecting non-agricultural conversion of cultivated land in remote sensing images provided by the present invention. Detailed implementation manners

[0069] The present invention is not limited by the following embodiments, and the specific implementation manners can be determined according to the technical solutions of the present invention and the actual situation.

[0070] The embodiments of the present invention provide a model training method, a method for detecting non-agricultural conversion of cultivated land, a device, an electronic device and a storage medium based on frequency domain feature interaction and pyramid hybrid attention. The preset deep learning model is trained by using a training set to obtain a trained deep learning model. The trained deep learning model is verified and tested by using a validation set and a test set to optimize the deep learning model and obtain a non-agricultural conversion detection model of cultivated land. The dual-temporal remote sensing images of the cultivated land to be detected are detected by using the non-agricultural conversion detection model of cultivated land to obtain the non-agricultural conversion detection result of the cultivated land.

[0071] Among them, the method provided by the embodiments of the present invention may involve artificial intelligence (AI) technology and can be implemented based on artificial intelligence technology. For example, in a deep learning manner, a corresponding model is trained by using samples.

[0072] Among them, machine learning (ML) is a multi-disciplinary cross-discipline that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specializes in studying how a computer simulates or implements human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve its own performance. Machine learning is the core of artificial intelligence and the fundamental way to make a computer intelligent, and its applications cover all fields of artificial intelligence.

[0073] Among them, deep learning (DL) specifically refers to machine learning based on deep neural network models and methods. It is developed on the basis of algorithm models such as statistical machine learning and artificial neural networks, combined with the development of contemporary big data and high computing power. The most important technical feature of deep learning is the ability to automatically extract features.

[0074] The above-mentioned machine learning and deep learning generally include technologies such as neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.

[0075] As shown in the Figure 1 accompanying figure, it shows a schematic diagram of the implementation environment provided by an embodiment of the present invention. The implementation environment may include: a training device and a using device.

[0076] Both the training device and the using device are computer devices; optionally, the computer device is a terminal device, such as electronic devices like mobile phones, tablets, PCs (Personal Computers), etc.; or, the computer device is a server, which can be a single server, or a server cluster composed of multiple servers, or a cloud computing service center. The embodiments of the present invention do not make any limitations in this regard.

[0077] The training device refers to a computer device with the ability to train and learn deep learning models. Optionally, the training device has the ability to obtain deep learning models and trains and learns them according to application requirements. For example, the training device obtains a deep learning model from other devices through a network, and then trains it with training samples according to application requirements, so that the deep learning model has the ability to obtain a cultivated land non-agriculturalization detection model; optionally, the training device has the ability to construct neural networks, and it can construct a deep learning model by itself according to application requirements, and then train and learn it. For example, in order to obtain the cultivated land non-agriculturalization detection result based on dual-temporal remote sensing images, the training device constructs a deep learning model by itself and then trains and learns it with samples according to application requirements.

[0078] The using device refers to a computer device with the need to use deep learning models. Optionally, the using device obtains a cultivated land non-agriculturalization detection model from other devices through a network according to application requirements. For example, the using device has the need for cultivated land non-agriculturalization detection, and it can obtain a cultivated land non-agriculturalization detection model that has completed training and learning from other devices through a network and use this cultivated land non-agriculturalization detection model for cultivated land non-agriculturalization detection.

[0079] Based on this, the technical solutions of the present invention will be introduced and illustrated below with several examples.

[0080] Embodiment 1: As shown in the Figure 2As shown, the embodiment of the present invention discloses a model training method based on frequency domain feature interaction and pyramid hybrid attention, including:

[0081] Step S110, obtaining a sample set, and dividing it into a training set, a validation set, and a test set in proportion, wherein the sample set includes a plurality of samples, each of which includes a historical dual-temporal remote sensing image of cultivated land and corresponding identification information of cultivated land non-agriculturalization detection results;

[0082] Step S120, training a preset deep learning model using the training set to obtain a trained deep learning model, wherein the deep learning model includes a multi-scale feature extraction and fusion module, a Transformer module, a pyramid hybrid attention module, and a decision-level fusion classification module;

[0083] Step S130, using the verification set and the test set to verify and test the trained deep learning model, optimize the deep learning model, and obtain a cultivated land non-agriculturalization detection model.

[0084] In this embodiment, each sample in step S110 includes a historical dual-temporal remote sensing image of cultivated land and corresponding identification information of cultivated land non-agriculturalization detection results, wherein the identification information of cultivated land non-agriculturalization detection results is an image that identifies the size and position of the non-agriculturalized cultivated land in the historical dual-temporal remote sensing image, and the image can be a feature map or a conversion result of the feature map.

[0085] In this embodiment, the deep learning model preset in step S120 includes a multi-scale feature extraction and fusion module, a Transformer module, a pyramid hybrid attention module, and a decision-level fusion classification module; the multi-scale feature extraction and fusion module can extract multi-scale features from dual-phase remote sensing images, and interactively fuse the features of each scale to ensure that the features of each scale in the same phase contain both detailed local information and a part of global information, which is conducive to locating the semantic information of cultivated land in complex scenes, thereby identifying cultivated land changes in small areas and complex scenes; the Transformer module can globally model and mine global information to make up for the disadvantage that the multi-scale feature extraction and fusion module can only use local information; the pyramid hybrid attention module uses different layers of pyramid hybrid attention for features of different scales to highlight features related to changes in channel dimensions and spatial dimensions, and suppress features unrelated to changes, so that cultivated land changes can be accurately identified under the conditions of complex backgrounds and large differences in target scales; the decision-level fusion classification module outputs cultivated land non-agriculturalization detection results for features of each scale, and fuses each cultivated land non-agriculturalization detection result in combination with the decision weights of each scale to obtain the final cultivated land change detection result.

[0086] In this embodiment, step S130 specifically includes:

[0087] Validate the trained deep learning model using the validation set, and optimize and update the model parameters by combining the validation results and the optimizer; the optimizer here can be the Adam optimizer, which can optimize and update the model parameters to accelerate the convergence process and improve the optimization efficiency. Use the cosine annealing learning rate scheduler to dynamically adjust the learning rate during training to avoid skipping the optimal solution or falling into a local optimum when approaching the optimal solution.

[0088] Test the deep learning model that has passed the validation using the test set, and output a qualified cultivated land change detection model.

[0089] The embodiment of the present invention provides a model training method based on frequency domain feature interaction and pyramid hybrid attention. By combining the multi-scale feature extraction and fusion module, the Transformer module and the pyramid hybrid attention module, multi-scale features are extracted from the dual-temporal remote sensing images, and each scale of features is processed for interactive fusion, global information mining, highlighting features related to changes, and suppressing features unrelated to changes. As a result, the model can not only fuse the features related to cultivated land captured at different scales, accurately locate the semantic information of cultivated land in complex scenarios, and accurately identify the non-agriculturalization changes of cultivated land in complex scenarios, but also solve the problems of inaccurate semantic information of cultivated land in complex scenarios and easy omission and misdetection existing in the existing cultivated land change detection methods.

[0090] Example 2: As shown in the appendix Figure 3 The embodiment of the present invention discloses a model training method based on frequency domain feature interaction and pyramid hybrid attention, which is a further optimization of the above embodiment. The sample set is obtained and divided into a training set, a validation set and a test set according to a certain proportion, including:

[0091] Step S210: Obtain a number of samples, and adjust the size of all samples according to the set pixels, where each sample includes the historical dual-temporal remote sensing images of cultivated land and the corresponding identification information of cultivated land non-agriculturalization detection results; the set pixels can be set as needed, and can be but not limited to 256*256 pixels or 512*512 pixels.

[0092] Step S220: Perform enhancement processing on the samples to obtain a sample set, where the enhancement methods include random rotation, horizontal flipping, and vertical flipping;

[0093] In this step, data augmentation is performed using a random number p. Specifically, random rotation (which can be set to 0.5 <= p < 0.75) is used to increase the diversity of the dataset and improve the robustness of the model to image rotation changes; horizontal flipping (which can be set to p < 0.25) is used to increase the diversity of the dataset and improve the robustness of the model to horizontal changes in the image; vertical flipping (which can be set to 0.25 <= p < 0.5) is used to increase the diversity of the dataset and improve the robustness of the model to vertical changes in the image.

[0094] Step S230: Divide the sample set into a training set, a validation set, and a test set according to a ratio.

[0095] The ratio used for dividing the sample set in this step can be set as needed.

[0096] Example 3: As shown in the appendix Figure 4 、 5 This embodiment of the present invention discloses a model training method based on frequency domain feature interaction and pyramid hybrid attention, which is a further optimization of the above embodiment. Training a preset deep learning model using the training set to obtain a trained deep learning model includes:

[0097] Step S310: Use a multi-scale feature extraction and fusion module to perform multi-scale feature extraction and fusion on the samples in the training set to obtain feature maps of different scales. The process of multi-scale feature extraction and fusion includes multi-scale feature extraction, channel dimensionality reduction, frequency domain feature interaction, and multi-scale frequency domain feature fusion.

[0098] In this step, the multi-scale feature extraction and fusion module includes a dual-branch convolutional neural network, a channel dimensionality reduction module, a frequency domain feature interaction module, and a multi-scale frequency domain feature fusion module, specifically as follows:

[0099] (1) Use a dual-branch convolutional neural network to extract multi-scale features of dual-temporal remote sensing images, that is, extract multi-scale features of dual-temporal remote sensing images through two convolutional neural networks respectively. In this embodiment, taking the extraction of features at three scales as an example, this process specifically includes:

[0100] As shown in the appendix Figure 5 Input the dual-temporal remote sensing image into a dual-branch convolutional neural network (Resnet) to extract multi-scale features of the two temporal phases. Resnet is composed of multiple residual blocks Block, including a non-dimensionality reduction residual block Block1, a dimensionality reduction residual block Block2, and a dilated dimensionality reduction residual block Block3;

[0101] F1(x) = Relu(BN(Conv1(x)))

[0102] F2(x) = Relu(BN(Conv3(F1(x))))

[0103]

[0104] F(x) = BN(Conv1(F2(x)))

[0105]

[0106] Block1(x) = Relu(x + F(X))

[0107] Block2(x) = Relu(Conv1(x) + F s2 (X))

[0108] Block3(x) = Relu(Conv1(x) + F d (X))

[0109] Among them, Conv1 is a 1*1 convolution operation; Conv3 is a 3*3 convolution operation; Conv3_2 is a 3*3 convolution operation with a stride of 2; Conv3d is a 3*3 dilated convolution operation with a dilation rate of 2, BN is layer normalization, and Relu is an activation function.

[0110] The backbone network of the dual-branch convolutional neural network includes four stages. The first stage consists of non-dimension-reducing residual blocks (Block1). In the second and third stages, the first residual block is a dimension-reducing residual block (Block2), and the rest are non-dimension-reducing residual blocks (Block1). In the fourth stage, the first residual block is a dilated dimension-reducing residual block (Block3), and the rest are non-dimension-reducing residual blocks (Block1). The number of residual blocks in each stage is (3, 4, 6, 3) respectively. The steps are as follows:

[0111] x = MaxPool3_2(Relu(BN(Conv7_2(x))))

[0112] x1 = Block1(x) i 1 <= <i <= 2

[0113] x2 = Block2(x1)

[0114] x2 = Block1(x2) i 1 <= <i <= 3

[0115] x3 = Block2(x2)

[0116] x3 = Block1(x3) i 1 <= <i <= 5

[0117] x4 = Block3(x3)

[0118] x4 = Block1(x4) i 1 <= i <= 2

[0119] where i is the index value of the current layer; Block(x) i is the i-th residual block; MaxPool3_2 is a max pooling with a size of 3*3 and a stride of 2; Conv7_2 is a convolution with a size of 7*7 and a stride of 2; The backbone network can obtain a deep convolutional neural network through the stacking of multiple residual blocks, which helps to extract richer image features.

[0120] (2) Use the channel dimensionality reduction module to perform channel dimensionality reduction on features of different scales using different convolutions;

[0121] Since the features extracted by the backbone network have many channels, which is not conducive to subsequent processing, this embodiment constructs a multi-scale channel dimensionality reduction module that uses different convolutions for different scales. The structure of the channel dimensionality reduction module is as shown in the appendix Figure 6 shown, and the specific channel dimensionality reduction steps include:

[0122] x1 = Conv7(x1)

[0123] w1 = Sigmoid(GConv(Relu(BN(x1))))

[0124] x1 = x1 * w1

[0125] x2 = Conv5(x2)

[0126] w2 = Sigmoid(GConv(Relu(BN(x2))))

[0127] x2 = x2 * w2

[0128] x4 = Conv3(x4)

[0129] w4 = Sigmoid(GConv(Relu(BN(x4))))

[0130] x4 = x4 * w4

[0131] where s1, s2, and s4 are the features extracted in the 1st, 2nd, and 4th stages of the backbone network respectively, with the scales from large to small, Conv7 is a 7*7 convolution; Conv5 is a 5*5 convolution; Conv3 is a 3*3 convolution; BN is batch normalization; Relu is an activation function; GConv is grouped convolution; Sigmoid is an activation function.

[0132] It should be noted that the number of the above channel dimensionality reduction modules corresponds to the number of scales.

[0133] (3) Use the frequency-domain feature interaction module to splice features of the same scale at different time phases, and use the fast Fourier transform to transform them into the frequency domain to obtain multi-scale frequency-domain features;

[0134] The above frequency-domain feature interaction module can reduce the style differences between dual-time-phase remote sensing images, achieve deep feature interaction between dual-time-phase remote sensing images, and suppress missed detections and misdetections caused by the style differences of dual-time-phase remote sensing images. In this embodiment, the frequency-domain feature interaction module can be as shown in the appendix Figure 7 and the specific steps are as follows:

[0135] (a) Splice the features of the same scale at different time phases of the input, perform a two-dimensional fast Fourier transform on the spliced tensor, and then adjust the shape of the resulting tensor to meet the input requirements of the average pooling layer.

[0136] x = torch.cat([x1, x2], dim = 1)

[0137] x fft = fft(x)

[0138] x fft = torch.view_as_real(x fft )

[0139] x fft = adjust(x fft )

[0140] where x1 and x2 are the feature maps of the same scale at different time phases of the input; fft is the fast Fourier transform; torch.view_as_real is to convert a complex tensor into a new tensor, where the last dimension represents the real and imaginary parts of the complex number; adjust is to adjust the shape of the tensor, and x fft is the result after final adjustment.

[0141] (b) Process the features transformed into the frequency domain, and then obtain a weight score. The specific steps are as follows.

[0142]

[0143] where Conv1 is a 1*1 convolution, Relu is an activation function, Sigmoid is an activation function, and avgpool is adaptive average pooling, that is, Adaptive AvgPooling in the appendix Figure 7 .

[0144] (c) Respectively transform the input features into the frequency domain through the fast Fourier transform, then multiply them with the obtained weights respectively, and finally return the processed features. The calculation steps are as follows:

[0145] x1 = fft(x1)

[0146] x2 = fft(x2)

[0147] x1 = x1 * weight

[0148] x2 = x2 * weight

[0149] x1 = ifft(x1)

[0150] x2 = ifft(x2)

[0151] Where, ifft is the inverse Fourier transform, which is the Inv FFT in the appendix Figure 7 in the appendix

[0152] (4) The multi-scale frequency domain feature fusion module splices the multi-scale frequency domain features based on the channel fusion branch and the spatial fusion branch respectively, and transforms the splicing results of the channel fusion branch and the spatial fusion branch into the frequency domain for feature fusion to obtain feature maps of different scales.

[0153] The above multi-scale frequency domain feature fusion module performs feature fusion between different scales at the same time phase, ensuring that the features of each scale at the same time phase contain both detailed local change information and global image information. In this embodiment, the structure of the multi-scale frequency domain feature fusion module can be as shown in the appendix Figure 8 as follows

[0154] (a) First, determine the scale to be fused, and then sample the features of other scales to the same size as the scale to be fused through upsampling or downsampling. To better retain the details of the image, in this embodiment, downsampling is performed through wavelet transform. Taking the second scale as the scale to be fused, the calculation steps are as follows

[0155] x l ,x h = DWTForward(x1)

[0156] x hh ,x lh ,x hl = x h

[0157] x1 = torch.cat([x hh ,x lh ,x hl , dim = 1)

[0158] x1 = Relu(BN(Conv3(x)))

[0159] x4 = Upsample2(BN(Conv1(x4)))

[0160] Among them, x1 and x4 are feature maps of the first scale and the third scale respectively; x lh is the horizontal high frequency; x hl is the vertical high frequency; x hh is the diagonal high frequency; x l is the low frequency component; DWTForward is the wavelet transform; Conv3 is a 3*3 convolution; Conv1 is a 1*1 convolution; BN is batch normalization; Relu is the activation function; Upsample2 is upsampling with a scale factor of 2.

[0161] (b) Input into the fusion module in ascending order of the size of the original scale. The fusion module has two branches, namely the channel fusion branch and the spatial fusion branch.

[0162] The calculation steps of the channel fusion branch are as follows:

[0163] x fuse = torch.cat([x 1, x2, x4], dim = 1)

[0164] x fuse = Conv1(x fuse )

[0165] out c = x fuse * channel_attention(x fuse )

[0166] Among them, x2 is the feature map of the second scale; channel_attention is the channel attention, that is, the CAM in Figure 8 ; Conv1 is a 1*1 convolution.

[0167] The channel attention is as shown in Figure 9 and the steps are as follows:

[0168] avg = Conv1(avgpool(x fuse ))

[0169] max = Conv1(maxpool(x fuse ))

[0170] weight c = Sigmoid(max + avg)

[0171] Among them, avgpool is the adaptive average pooling; maxpool is the adaptive max pooling.

[0172] The calculation steps of the spatial fusion branch are as follows:

[0173] x2 = WTConv2(x1)

[0174] x3 = WTConv3(x2)

[0175] x4 = WTConv4(x3)

[0176] x fuse_s = torch.cat([x 1, x2, x4], dim = 1)

[0177] x fuse_s = Conv1(x fuse_s )

[0178] out s = x fuse_s * spatial_attention(x fuse_s )

[0179] The spatial attention is as shown in the appendix Figure 10 and the steps are as follows:

[0180]

[0181] weight s = torch.cat([mean, max], dim = 1)

[0182] weight s = Conv1(weight s )

[0183] weight s = Sigmoid(weight s )

[0184] Among them, WTConv2 is the wavelet convolution of two wavelet transforms; WTConv3 is the wavelet convolution of three wavelet transforms; WTConv4 is the wavelet convolution of four wavelet transforms; Conv1 is a 1*1 convolution; spatial_attention is the spatial attention, that is, SAM in the appendix Figure 8

[0185] ​Among them, WTConv is a wavelet convolution, which is a convolutional block based on wavelet transform. Different from the Fourier transform, wavelet transform contains information in both the spatial domain and the frequency domain. First, Haar WT is selected as the basis. Each wavelet transform is divided into four components: low frequency, horizontal high frequency, vertical high frequency, and diagonal high frequency. Subsequently, cascading operations will be performed. The low-frequency component among the four obtained components will be wavelet-transformed again to obtain four lower-level components (the hierarchical structure of cascaded wavelet decomposition will generate new LL, LH, HL, and HH components with each decomposition, but these new components only come from the LL part of the previous decomposition). During the inverse transform, first, convolution operations (depth convolution) will be performed on them. Then, the low-frequency component is added to the result of the inverse wavelet transform of the lower level, and then the four components at this level are subjected to the inverse wavelet transform to restore the original size. The steps are as follows:

[0186] x = IWT(Conv(w, WT(x)))

[0187] Among them, x is the input tensor, and W is a weight tensor of k×k with the number of channels being 4 times that of x. Each frequency component (i.e., the four frequency components obtained by wavelet decomposition) is convolved using a small convolution kernel (k×k) respectively. Here, depth convolution is used, that is, convolution is performed one by one in the channel dimension.

[0188] This operation not only separates the convolution between frequency components but also allows smaller kernels to operate in a larger area of the original input. After wavelet transform, there will be a larger receptive field during convolution operation. The visualization of wavelet transform (WT) is shown in the appendix Figure 8 as follows, and the calculation steps are as follows:

[0189]

[0190] Among them, is the low-frequency component generated after the i-th wavelet transform; is the high-frequency component generated after the i-th wavelet transform, which can be further decomposed into horizontal high frequency vertical high frequency diagonal high frequency three high-frequency components.

[0191] The inverse wavelet transform formula is as follows:

[0192] IWT(x + y) = IWT(x) + IWT(y)

[0193]

[0194] Using the above formula to obtain the sum of convolutions at different levels, where z (i) is the result of the i-th inverse wavelet transform, z (i +1)is the result of the (i + 1)-th inverse wavelet transform.

[0195] WTConv: The overall calculation process is as follows:

[0196]

[0197] X 0 (LL) = X

[0198]

[0199] for i = 1, …, l do

[0200]

[0201] end for

[0202] Z (l+1) = 0

[0203] for i = 1, …, l do

[0204]

[0205] end for

[0206]

[0207] return Z (0)

[0208] Input a feature map, then perform wavelet transform to obtain four components, and then perform wavelet transform on the low-frequency component until the last layer. Then for each layer, first transform the four components through depth convolution, then add the low-frequency component to the result of the inverse wavelet transform of the previous level, and then perform inverse wavelet transform on the four components of this level to obtain the result of this layer, and this result is passed back to the previous layer until the first layer. The result of the first layer is added to the input feature, and finally the final result is obtained.

[0209] (c) Concatenate the results of the spatial branch and the channel branch, then transform them into the frequency domain for feature fusion, and return the fused features. The specific steps are as follows:

[0210] x = torch.cat([out c , out c , dim = 1)

[0211] x fft = rfft(x)

[0212] x fft = torch.view_as_real(x fft )

[0213] x fft = adjust(x fft )

[0214] x fft = Relu(BN(Conv1(x fft )))

[0215] x fft = adjust1(x fft )

[0216] x fft = torch.view_as_complex(x fft )

[0217] out = irfft(x fft )

[0218] out = BN(Conv1(out))

[0219] Among them, rfft is the Fourier transform; irfft is the inverse Fourier transform; adjust and adjust1 are tensor format adjustments; Conv1 is a 1*1 convolution; BN is batch normalization; Relu is an activation function; torch.view_as_real is to convert a complex tensor into a new tensor with the same data but regarded as a real tensor; torch.view_as_complex is to convert a real tensor (or data regarded as a real tensor) into a complex tensor.

[0220] In this step of the present embodiment, a convolutional neural network structure with multi-branch parallelism is adopted to extract multi-scale features from the dual-temporal remote sensing images. While being independent of each other, the branches interact with each other, making full use of the advantages of different-scale features to facilitate the localization of the semantic features of cultivated land in complex backgrounds; a multi-scale channel dimensionality reduction module is used to perform different convolutional operations for different scales to reduce the dimensionality of channels, and at the same time, an operation similar to channel attention is adopted to highlight important information in the channel dimension; a frequency-domain feature interaction strategy is adopted to transform the features in the spatial domain into the frequency domain, and then the features of the dual-temporal are interacted in a manner similar to spatial attention and channel attention. Adopting this strategy can minimize the style differences between the dual-temporal remote sensing images, achieving the purpose of unified style, thereby suppressing non-semantic changes and missed detections caused by inconsistent styles; a multi-scale frequency-domain feature fusion module is used to replace the downsampling operation with wavelet transform. Wavelet transform can better retain detailed information, and this module is provided with a spatial branch and a channel branch. The channels most closely related to the changes are obtained on the channel branch. On the spatial branch, the present invention uses wavelet convolution to expand the receptive field as much as possible, and wavelet convolutions with different receptive fields are used for different scales to adapt to features of different scales, and then the features of the spatial branch and the channel branch are fused in the frequency domain, which can fuse the features related to cultivated land captured at different scales, facilitating the localization of the semantic information of cultivated land in the reproduction scene, and thus identifying the changes of cultivated land in micro regions and complex scenes.

[0221] Step S320: Use the Transformer module to perform self-attention encoding, decoding, and reconstruction on the feature maps of different scales to extract global features.

[0222] The above-mentioned Transformer module performs global modeling to make up for the shortcoming that the convolutional neural network can only utilize local information. The specific steps are as follows:

[0223] (a) First, use the encoder of the transformer to encode the features of different scales using self-attention. Specifically, through the multi-head self-attention mechanism, dynamically focus on the features at different levels of the image, strengthen the focus on the changing areas, and improve the detection accuracy and robustness of the model in complex scenarios. It should be noted that one encoder is used for each scale.

[0224] The self-attention mechanism generates attention scores by calculating the dot products of the query, key, and value matrices, and normalizes these scores through the softmax function to finally generate attention weights. The calculation formula is as follows:

[0225] q, k, v = linear(norm(token)), linear(norm(token)), linear(norm(token))

[0226]

[0227] out = rearrange(x, 'bhn d->b n(hd)')

[0228] out = linear(out)

[0229] FFN(out) = (max(0, w1 * out + b1) * w2 + b2) + out

[0230] result = norm(out)

[0231] Among them, token is the token after the change of the feature map at a certain scale; d k is the number of columns of k; b is the batch; h is the number of attention heads; n is the token length; d is the token dimension.

[0232] (b) In the decoder part, in this embodiment, cross-attention can be used to effectively decode and reconstruct features at different scales. Specifically, the decoder is also based on the transformer architecture, but its main purpose is to decode the multi-scale features encoded by the encoder to enhance the model's understanding of global features. It should be noted that one decoder is used for each scale, and the main working process of the decoder is as follows:

[0233] Cross-attention is used to establish a dependency relationship between two different feature sets. In the multi-temporal remote sensing image change detection task, cross-attention generates a normalized attention weight by calculating the attention score between the decoder feature (query) and the encoder feature (key) and (vaule), and then uses these weights to perform weighted summation on the encoder feature to generate a new feature representation. The calculation formula is as follows:

[0234] q = linear(norm(token))

[0235] k, v = linear(norm(token1)), linear(norm(token1))

[0236]

[0237] out = rearrange(x, 'bhn d->b n(hd)')

[0238] out = linear(out)

[0239] FFN(out) = (max(0, w1 * out + b1) * w2 + b2) + out

[0240] result = norm(out)

[0241] Among them, token is the token of a feature map at a certain scale before encoding; token1 is the token of a feature map at a certain scale after encoding; token1 is the token after the change of a feature map at a certain scale; d k is the number of columns of k; b is the batch size; h is the number of attention heads; n is the token length; d is the dimension of the token.

[0242] Step S330, use the pyramid hybrid attention module to fuse the features output by the Transformer module at different scales to obtain the feature maps of each scale after fusion.

[0243] In this embodiment, the structure of the pyramid hybrid attention module can be as shown in the appendix Figure 11 as follows, and the specific steps are as follows:

[0244] (1) Features of the same scale do not share a single pyramid hybrid attention module (corresponding to PMA in the appendix Figure 5 ), and the number of layers of the pyramid hybrid attention module used for features of different scales is different. x1, x2, and x4 represent features of three different scales, from large to small. Based on the scale of x4, the pyramid hybrid attention module it uses is one layer, the pyramid hybrid attention module of x2 is two layers, and the pyramid hybrid attention module of x1 is three layers. The specific calculation steps of the three-layer pyramid hybrid attention module used by x1 are as follows:

[0245] The first layer:

[0246]

[0247] x1, x2 = x.chunk(2, dim = 2)

[0248]

[0249]

[0250] The second layer:

[0251]

[0252]

[0253] The third layer:

[0254] out1 = torch.cat(out1, dim = 2)

[0255] out2 = torch.cat(out2, dim=2)

[0256] x = torch.cat([x1, x2], dim=2)

[0257] out1 = out1 * channe_attention(x)

[0258] out2 = out2 * channe_attention(x)

[0259] out1 = out1 * spatial_attention(x)

[0260] out2 = out2 * spatial_attention(x)

[0261] Among them, is the feature of the first scale in the first phase; is the feature of the first scale in the second phase; chunk is to split the tensor into multiple smaller tensors along the specified dimension; channe_attention is channel attention; spatial_attention is spatial attention.

[0262] (b) The pyramid hybrid attention used by x2 has only two layers, and the specific steps are as follows:

[0263] The first layer:

[0264]

[0265] x1, x2 = x.chunk(2, dim=2)

[0266]

[0267] The second layer:

[0268] out1 = torch.cat(out1, dim=2)

[0269] out2 = torch.cat(out2, dim=2)

[0270] x = torch.cat([x1, x2], dim=2)

[0271] out1 = out1 * channe_attention(x)

[0272] out2 = out2 * channe_attention(x)

[0273] out1 = out1 * spatial_attention(x)

[0274] out2 = out2 * spatial_attention(x)

[0275] (c) The pyramid hybrid attention used by x4 has only one layer, and the specific steps are as follows:

[0276]

[0277] In the above steps of this embodiment, different numbers of layers of pyramid hybrid attention are used for features of different scales, which can not only process global features but also focus on local features. The hybrid attention can highlight the features related to changes and suppress the features unrelated to changes in both the channel dimension and the spatial dimension, so as to accurately identify cultivated land changes in the case of complex backgrounds and large differences in target scales.

[0278] Step S340: Use the decision-level fusion classification module to perform cultivated land non-agriculturalization detection on the fused feature maps of each scale respectively, and fuse the detection results of each cultivated land non-agriculturalization to obtain the final cultivated land non-agriculturalization detection result.

[0279] In this embodiment, the decision-level fusion classification module adopts a weighted voting method to dynamically adjust the decision weights of the output results of each scale, ensuring that the final classification result takes into account both shallow semantic features and deep semantic features. The structure of the decision-level fusion classification module is as shown in the appendix Figure 12 shown, where the structure of each classifier is as shown in the appendix Figure 13 shown, and the specific steps are as follows:

[0280] (1) First, input the feature maps of each scale into the corresponding classifier respectively to obtain three cultivated land non-agriculturalization detection results at different scales. The specific steps are as follows:

[0281]

[0282] x1 = interpolate(x1, size = imgsize)

[0283] x2 = interpolate(x2, size = imgsize)

[0284] x a = interpolate(x4, size = imgsize)

[0285] result1 = Classifier1(x1)

[0286] result2 = Classifier2(x2)

[0287] result4 = Classifier4(x4)

[0288] The specific process of the classifier Classifier is as follows:

[0289] result = Conv3(Relu(BN(Conv3(x1))))

[0290] interpolate: Upsampling function, Conv3: 3*3 convolution operation, BN: Batch normalization, Relu: Activation function, Classifier1, Classifier2, Classifier4: Three classifiers, one for each scale.

[0291] (2)Fuse the obtained cultivated land non-agriculturalization detection results at three different scales to obtain the final cultivated land non-agriculturalization detection result. The specific steps are as follows:

[0292] result = torch.cat([result1, result2, result4], dim = 1)

[0293] result = Conv1(result)

[0294] Among them, Conv1: 1*1 convolution.

[0295] Step S350, optimize and balance the loss between the cultivated land non-agriculturalization detection results at each scale and the fused cultivated land non-agriculturalization detection result in the training through an adaptive loss function. Thereby ensuring the accuracy and consistency of the detection results. The specific steps are as follows:

[0296] Use adaptive parameters to adjust the ratio between the four losses. The calculation process is as follows: loss can be replaced with other loss functions

[0297]

[0298] loss = cross_entropy()

[0299] L = a * loss(result1) + b * loss(result2) + c * loss(result4) +

[0300] d * loss(result)

[0301] Among them, c is the number of classes; y i is the true label of the i-th class (0 or 1, usually using one-hot encoding); p iis the probability of the i-th class predicted by the model; a, b, c, and d are adaptive parameters; cross_entyopy() is the cross-entropy function; result1, result3, result4, and result are the transformed result maps after the first scale, the second scale, the third scale, and the final fusion, respectively.

[0302] In the above steps of this embodiment, a new loss function is used for the results of different scales, and the loss weights between different scales and different classes are adjusted through adaptive parameters. In this way, the features of different scales can be fully utilized, and the semantic features of interest in complex backgrounds can be focused on, alleviating the problems of false detection and missed detection caused by complex backgrounds.

[0303] Step S360, loop the above steps, output the training weights that meet the training conditions, and obtain the trained deep learning model. The training conditions include the number of iterations, stable loss value, etc.

[0304] Embodiment 4: As shown in the appendix Figure 14 As shown, 6 groups of historical dual-temporal remote sensing images are obtained. The first row is the remote sensing images of phase one, the second row is the remote sensing images of phase two, the third row is the identification information of the cultivated land non-agriculturalization detection results of each group of historical dual-temporal remote sensing images, and the fourth row is the cultivated land non-agriculturalization detection results detected by using the cultivated land non-agriculturalization detection model trained by the present invention. Among them, black represents the unchanged area, white represents the changed area, green represents the false detection area, and red represents the missed detection area. It is found by comparison that the missed detection and false reduction areas are within the allowable range, and the training model of the present invention is applicable to the detection of cultivated land non-agriculturalization.

[0305] Embodiment 5: As shown in the appendix Figure 15 As shown, an embodiment of the present invention discloses a method for detecting cultivated land non-agriculturalization in remote sensing images, including:

[0306] Step S410, obtain the dual-temporal remote sensing images of the cultivated land to be detected and preprocess them;

[0307] Step S420, input the preprocessed dual-temporal remote sensing images of the cultivated land to be detected into the cultivated land non-agriculturalization detection model to obtain the cultivated land non-agriculturalization detection results, where the cultivated land non-agriculturalization detection model is trained by the model training method based on frequency domain feature interaction and pyramid hybrid attention described in the above embodiment.

[0308] Embodiment 6: As shown in the appendix Figure 16 As shown, an embodiment of the present invention discloses a model training device based on frequency domain feature interaction and pyramid hybrid attention, including:

[0309] A sample acquisition module that acquires a sample set and divides it into a training set, a validation set, and a test set according to a ratio. The sample set includes a number of samples, and each sample includes historical dual-temporal remote sensing images of cultivated land and corresponding cultivated land non-agriculturalization detection result identification information.

[0310] A model training module that trains a preset deep learning model using the training set to obtain a trained deep learning model. The deep learning model includes a multi-scale feature extraction and fusion module, a Transformer module, a pyramid hybrid attention module, and a decision-level fusion classification module.

[0311] A model optimization module that validates and tests the trained deep learning model using the validation set and the test set to optimize the deep learning model and obtain a cultivated land non-agriculturalization detection model.

[0312] Example 7: The embodiment of the present invention discloses a model training device based on frequency domain feature interaction and pyramid hybrid attention, which is a further optimization of the above embodiment. The model training module includes:

[0313] A multi-scale feature extraction and fusion module that performs multi-scale feature extraction and fusion on the samples in the training set to obtain feature maps of different scales. The process of multi-scale feature extraction and fusion includes multi-scale feature extraction, channel dimensionality reduction, frequency domain feature interaction, and multi-scale frequency domain feature fusion.

[0314] A Transformer module that performs self-attention encoding, decoding, and reconstruction on the feature maps of different scales to extract global features.

[0315] A pyramid hybrid attention module that fuses the features output by the Transformer module according to different scales to obtain feature maps of each scale after fusion.

[0316] A decision-level fusion classification module that performs cultivated land non-agriculturalization detection on the feature maps of each scale after fusion respectively, and fuses the results of each cultivated land non-agriculturalization detection to obtain the final cultivated land non-agriculturalization detection result.

[0317] An evaluation module that optimizes and balances the loss between the cultivated land non-agriculturalization detection results of each scale in the training and the cultivated land non-agriculturalization detection result after fusion through an adaptive loss function.

[0318] An output module that outputs the training weights that meet the training conditions to obtain the trained deep learning model.

[0319] Example 8: As shown in the appendix Figure 17 The embodiment of the present invention discloses a remote sensing image cultivated land non-agriculturalization detection device, including:

[0320] An image acquisition module acquires dual-temporal remote sensing images of the cultivated land to be detected and preprocesses them.

[0321] A detection module inputs the preprocessed dual-temporal remote sensing images of the cultivated land to be detected into a cultivated land non-agriculturalization detection model to obtain a cultivated land non-agriculturalization detection result, where the cultivated land non-agriculturalization detection model is trained by the model training method based on frequency domain feature interaction and pyramid hybrid attention described in the above embodiments.

[0322] Embodiment 9: An embodiment of the present invention discloses a storage medium, on which a computer program readable by a computer is stored, and the computer program is set to execute the model training method based on frequency domain feature interaction and pyramid hybrid attention when running.

[0323] The above storage medium may include, but is not limited to: various media such as USB flash drives, read-only memories, mobile hard disks, magnetic disks, or optical discs that can store computer programs.

[0324] Embodiment 10: An embodiment of the present invention discloses an electronic device, including a processor and a memory. A computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the model training method based on frequency domain feature interaction and pyramid hybrid attention.

[0325] The above processor may be a central processing unit CPU, a general-purpose processor, a digital signal processor DSP, an ASIC, an FPGA, or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in combination with the disclosure of the present application. It can also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of DSP and microprocessors, etc. The memory may include, but is not limited to: various media such as USB flash drives, read-only memories, mobile hard disks, magnetic disks, or optical discs that can store computer programs.

[0326] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes. The solutions in the embodiments of the present application can be implemented in various computer languages. For example, object-oriented programming languages such as Java and interpreted scripting languages such as JavaScript.

[0327] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, as well as the combination of flows and / or blocks in the flowchart and / or block diagram. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.

[0328] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.

[0329] The above content is only the specific implementation manner of the present application, which has strong adaptability and implementation effects. However, the protection scope of the present application is not limited thereto. Any person skilled in the art in the technical field disclosed by the present application can easily think of changes or substitutions, which should be covered by the protection scope of the present application. Therefore, equivalent changes made according to the claims of the present application still fall within the scope covered by the present application.

Claims

1. A model training method based on frequency domain feature interaction and pyramid hybrid attention, characterized in that Including: Obtain a sample set and divide it into a training set, a validation set, and a test set according to a ratio, where the sample set includes a number of samples, and each sample includes historical dual-temporal remote sensing images of cultivated land and corresponding identification information on the detection results of cultivated land non-agriculturalization; Use the training set to train a preset deep learning model to obtain a trained deep learning model, where the deep learning model includes a multi-scale feature extraction and fusion module, a Transformer module, a pyramid hybrid attention module, and a decision-level fusion classification module; Use the validation set and the test set to validate and test the trained deep learning model, and optimize the deep learning model to obtain a cultivated land non-agriculturalization detection model.

2. The model training method based on frequency domain feature interaction and pyramid hybrid attention according to claim 1, characterized in that The step of using the training set to train a preset deep learning model to obtain a trained deep learning model includes: Use the multi-scale feature extraction and fusion module to perform multi-scale feature extraction and fusion on the samples in the training set to obtain feature maps of different scales, where the process of multi-scale feature extraction and fusion includes multi-scale feature extraction, channel dimensionality reduction, frequency-domain feature interaction, and multi-scale frequency-domain feature fusion; Use the Transformer module to perform self-attention encoding, decoding, and reconstruction on the feature maps of different scales to extract global features; Use the pyramid hybrid attention module to fuse the features output by the Transformer module according to different scales to obtain feature maps of each scale after fusion; Use the decision-level fusion classification module to perform cultivated land non-agriculturalization detection on the feature maps of each scale after fusion respectively, and fuse the detection results of each cultivated land non-agriculturalization to obtain the final detection result of cultivated land non-agriculturalization; Optimize and balance the loss between the detection results of cultivated land non-agriculturalization at each scale and the detection result of cultivated land non-agriculturalization after fusion in the training through an adaptive loss function; Loop the above steps, output the training weights that meet the training conditions, and obtain a trained deep learning model.

3. The model training method based on frequency domain feature interaction and pyramid hybrid attention according to claim 2, characterized in that The multi-scale feature extraction and fusion module includes a dual-branch convolutional neural network, a channel dimensionality reduction module, a frequency-domain feature interaction module, and a multi-scale frequency-domain feature fusion module. The step of using the multi-scale feature extraction and fusion module to perform multi-scale feature extraction and fusion on the samples in the training set to obtain feature maps of different scales includes Use the dual-branch convolutional neural network to extract multi-scale features of the dual-temporal remote sensing images; Use the channel dimensionality reduction module to perform channel dimensionality reduction on the features of different scales using different convolutions; Use the frequency-domain feature interaction module to splice the features of the same scale in different temporal phases and transform them to the frequency domain using the fast Fourier transform to obtain multi-scale frequency-domain features; Use the multi-scale frequency-domain feature fusion module to splice the multi-scale frequency-domain features based on the channel fusion branch and the spatial fusion branch respectively, and transform the splicing results of the channel fusion branch and the spatial fusion branch to the frequency domain for feature fusion to obtain feature maps of different scales.

4. The model training method based on frequency-domain feature interaction and pyramid hybrid attention according to any one of claims 1 to 3, characterized in that The step of obtaining a sample set and dividing it into a training set, a validation set, and a test set according to a ratio includes: Obtain a number of samples, and adjust the sizes of all samples according to the set pixels, where each sample includes the historical dual-temporal remote sensing images of cultivated land and the corresponding identification information of cultivated land non-agriculturalization detection results; Perform enhancement processing on the samples to obtain a sample set, where the enhancement methods include random rotation, horizontal flipping, and vertical flipping; Divide the sample set into a training set, a validation set, and a test set according to a ratio.

5. A method for detecting the non-agriculturalization of cultivated land in remote sensing images, characterized in that, Include: Obtain the dual-temporal remote sensing images of the cultivated land to be detected, and perform preprocessing on them; Input the preprocessed dual-temporal remote sensing images of the cultivated land to be detected into the cultivated land non-agriculturalization detection model to obtain the cultivated land non-agriculturalization detection results, where the cultivated land non-agriculturalization detection model is trained by the model training method based on frequency domain feature interaction and pyramid hybrid attention described in any one of claims 1 to 4.

6. A model training device based on frequency domain feature interaction and pyramid hybrid attention that applies the method described in any one of claims 1 to 4, characterized in that, Include: A sample acquisition module that obtains a sample set and divides it into a training set, a validation set, and a test set according to a ratio, where the sample set includes a number of samples, and each sample includes the historical dual-temporal remote sensing images of cultivated land and the corresponding identification information of cultivated land non-agriculturalization detection results; A model training module that uses the training set to train a preset deep learning model to obtain a trained deep learning model, where the deep learning model includes a multi-scale feature extraction and fusion module, a Transformer module, a pyramid hybrid attention module, and a decision-level fusion classification module; A model optimization module that uses the validation set and the test set to verify and test the trained deep learning model, optimize the deep learning model, and obtain the cultivated land non-agriculturalization detection model.

7. The model training device based on frequency domain feature interaction and pyramid hybrid attention according to claim 6, wherein The model training module includes: A multi-scale feature extraction and fusion module that performs multi-scale feature extraction and fusion on the samples in the training set to obtain feature maps of different scales, where the process of multi-scale feature extraction and fusion includes multi-scale feature extraction, channel dimensionality reduction, frequency domain feature interaction, and multi-scale frequency domain feature fusion; A Transformer module that performs self-attention encoding, decoding, and reconstruction on the feature maps of different scales to extract global features; A pyramid hybrid attention module that fuses the features output by the Transformer module according to different scales to obtain the fused feature maps of each scale; A decision-level fusion classification module that respectively performs cultivated land non-agriculturalization detection on the fused feature maps of each scale, and fuses the cultivated land non-agriculturalization detection results to obtain the final cultivated land non-agriculturalization detection results; An evaluation module that optimizes and balances the loss between the cultivated land non-agriculturalization detection results of each scale and the fused cultivated land non-agriculturalization detection results in the training through an adaptive loss function; An output module that outputs the training weights that meet the training conditions to obtain the trained deep learning model.

8. A remote sensing image cultivated land non-agriculturalization detection device applying the method according to claim 6, characterized in that, Include: An image acquisition module that obtains the dual-temporal remote sensing images of the cultivated land to be detected and performs preprocessing on them; The detection module inputs the two - time - phase remote - sensing images of the cultivated land to be detected after pre - processing into the cultivated - land non - agriculturalization detection model, and obtains the cultivated - land non - agriculturalization detection result, where the cultivated - land non - agriculturalization detection model is trained by the model training method based on frequency - domain feature interaction and pyramid hybrid attention as described in any one of claims 1 to 4.

9. An electronic device, characterized in that, It includes a processor and a memory, and a computer program is stored in the memory. The computer program is loaded and executed by the processor to implement the steps in the method as described in any one of claims 1 to 4.

10. A storage medium, characterized in that, A computer program readable by a computer is stored on the storage medium, and the computer program is set to execute the steps in the method as described in any one of claims 1 to 4 when running.

Citation Information

Cited By

  • Remote sensing image road extraction method based on direction wavelet convolution

    CN120894700A