A multi-scale cultivated land change detection system based on frequency domain conversion

By adopting a multi-scale cultivated land change detection system based on frequency domain conversion in high-resolution remote sensing image change detection, the problems of false change detection, limited multi-scale feature extraction effect and difficulty in long-distance dependency capture are solved, and a more efficient, accurate and robust change detection effect is achieved.

CN119832439BActive Publication Date: 2025-06-27CHINA AGRI UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510317032.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-06-27
Estimated Expiration
2045-03-18

AI Technical Summary

Technical Problem

The existing high-resolution remote sensing image change detection technology has problems such as high false detection rate of pseudo-change errors, limited multi-scale feature extraction effect, and difficulty in capturing long-distance dependencies, resulting in limited detection accuracy and robustness in complex scenarios.

Method used

A multi-scale farmland change detection system based on frequency domain transformation is proposed. The two-time phase images are converted from the spatial domain to the frequency domain through fast Fourier transform (FFT) and replaced the low-frequency components to unify the style of two different phase images. Multi-core parallel convolution module and cross-scale feature aggregator (CSFA) are used for multi-scale feature extraction and enhancement, capturing global context information and fine-grained features.

Benefits of technology

It effectively reduces pseudo-change interference, avoids background noise caused by large convolution kernels and information loss caused by hollow convolution, improves the detection ability of scale-sensitive targets, and significantly improves the accuracy and robustness of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119832439B_ABST
    Figure CN119832439B_ABST
Patent Text Reader

Abstract

The present disclosure provides a multi-scale cultivated land change detection system based on frequency domain conversion, belonging to the technical fields of remote sensing image processing and computer vision. First, feature extraction is performed on the two-temporal images to obtain multi-scale first feature maps; these first feature maps are subjected to multi-core parallel convolution processing via small convolution kernels of different sizes, and then stitched and fused into second feature maps; the second feature maps are subjected to feature enhancement to generate extracted feature maps; the two-way extracted feature maps corresponding to the two-temporal images are merged and classified to obtain a detection map of cultivated land changes. At the same time, the two-temporal images used for training are subjected to low-frequency feature replacement in the frequency domain, effectively eliminating the visual feature differences caused by different imaging conditions and reducing the misdetection of pseudo-changes. The present invention reduces the interference of pseudo-changes through low-frequency component replacement. At the same time, the multi-core parallel convolution processing of small convolution kernels can avoid the background noise caused by large convolution kernels and the information loss caused by dilated convolution, while improving the detection ability for scale-sensitive targets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of remote sensing image processing and computer vision, and particularly relates to a multi-scale cultivated land change detection system based on frequency domain conversion. Background Art

[0002] The existing research on high-resolution remote sensing image change detection mainly has the following deficiencies:

[0003] Firstly, in high-resolution remote sensing image change detection, the domain gap problem has not been effectively solved. The existing technologies are highly sensitive to imaging condition differences, such as illumination changes, viewing angle changes, and background complexity, resulting in a relatively high false detection rate of pseudo-changes. Traditional preprocessing and normalization methods, such as histogram matching, radiometric correction, and feature normalization, are insufficient in adaptability and stability in complex scenarios and are difficult to effectively eliminate the negative impacts of imaging condition changes. In addition, deep learning models still have significant limitations in domain generalization ability. They usually rely on a large amount of labeled data for training, but their adaptability under different imaging conditions is limited, leading to a decline in performance in practical applications. Although generative adversarial networks and transfer learning methods have been explored in domain adaptation, their effects are still not ideal and are difficult to be widely applied to variable imaging environments.

[0004] Secondly, in terms of multi-scale feature extraction, the large convolution kernels relied on by existing technologies are prone to introducing a large amount of irrelevant background information, interfering with the effective discrimination of target regions, and often resulting in the blurring of edge information and the loss of details. This phenomenon is particularly obvious in high-resolution remote sensing images because high-resolution data itself contains rich detail information, and any blurring or information loss will directly affect the accuracy of the detection results. At the same time, although dilated convolution can effectively expand the receptive field and capture context information in a larger range, due to the sparsity of its sampling method and the damage to the structure, it is easy to miss key local details, thereby affecting the accuracy and robustness of the detection. At the same time, single-scale feature extraction methods are difficult to comprehensively capture context information at different levels and cannot effectively balance the local and global relationships, resulting in limited performance of the model when dealing with complex scenarios. Therefore, the existing multi-scale feature extraction schemes face challenges of high computational cost, complex design, and limited effects in practical applications, and there is an urgent need for more efficient, accurate, and scale-adaptive feature extraction methods to improve the overall performance of high-resolution remote sensing image change detection.

[0005] In addition, the existing technologies have significant deficiencies in capturing long-range dependencies in high-resolution remote sensing image changes. Local feature extraction methods, such as traditional convolutional neural networks (CNNs), are difficult to effectively integrate global context information, especially when dealing with large-scale changes. This is because the characteristics of local operations limit the model's ability to capture long-distance pixel relationships, resulting in difficulty in ensuring global consistency in complex scenarios. Current attention mechanisms (such as self-attention mechanisms) mainly focus on enhancing local features by weighting information within the neighborhood to improve feature representation, but they are still insufficient in capturing cross-region long-range dependencies, and their high computational complexity and memory consumption make it difficult to be efficiently applied in large-scale high-resolution images. Although traditional Siamese networks have a certain feature fusion ability and perform feature comparison through a dual-branch structure, due to their relatively simple feature fusion method and lack of in-depth integration design, they cannot fully utilize multi-scale features and long-range dependencies, restricting their application effects and robustness in high-resolution remote sensing image change detection tasks. Therefore, there is an urgent need to develop a new network architecture that can simultaneously capture global context information and fine-grained features and has efficient computing capabilities to improve the overall performance and application value of high-resolution remote sensing image change detection.

[0006] In summary, the existing technologies have the following deficiencies in high-resolution remote sensing image change detection:

[0007] 1. Traditional methods are difficult to effectively handle visual feature differences caused by differences in imaging conditions (such as illumination, shadows, backgrounds, etc.) between two-phase images, and false change misdetection is likely to occur.

[0008] 2. Existing multi-scale feature extraction methods such as dilated convolution and large convolution kernels are prone to losing key information or introducing background noise.

[0009] 3. Changing objects exhibit complex and diverse morphological characteristics, which are reflected in the different sizes and irregular shapes of changing ground objects. This makes it more challenging to simultaneously identify changing objects of different scales and accurately depict their change boundaries.

[0010] These deficiencies limit the accuracy and robustness of the model in complex scenarios, making it difficult to be widely applied to diverse practical tasks, and increasing the costs and difficulties of training and application. Summary of the Invention

[0011] In view of this, the present invention proposes a multi-scale cultivated land change detection system based on frequency domain conversion, aiming to reduce false change interference, avoid background noise caused by large convolution kernels and information loss caused by dilated convolution, and at the same time improve the detection ability for scale-sensitive targets.

[0012] To solve the above technical problems, the present invention is implemented as follows.

[0013] A multi-scale cultivated land change detection system based on frequency domain conversion, which includes a feature extraction model, a dual-channel merging unit and a training unit;

[0014] The feature extraction model processes each path of the input dual-temporal image to obtain two paths of extracted feature maps and ;

[0015] The feature extraction model includes a feature extraction module, a multi-core parallel convolution module, and a feature enhancement module; the feature extraction module is used to extract features from the two paths of images in the dual-temporal image to obtain a multi-scale first feature map;

[0016] The multi-core parallel convolution module contains multiple parallel convolution branches. Each parallel convolution branch uses m convolution kernels of different sizes, and the convolution kernels are less than or equal to 13×13; the value range of m is 3-7; the first feature map enters each parallel convolution branch for processing, and then is stitched and fused to obtain a second feature map ;

[0017] The feature enhancement module is used to enhance the features of the second feature map to generate an extracted feature map ;

[0018] The dual-channel merging unit merges and classifies the two paths of extracted feature maps and to obtain a detection map P of cultivated land change;

[0019] The training unit includes an FFT module and a multi-scale training module;

[0020] The FFT module is used to convert the dual-temporal image from the spatial domain to the frequency domain when training the feature extraction model, unify the low-frequency regions in the dual-temporal image by replacement, and then convert it back to the spatial domain to generate a dual-temporal image with the low-frequency components replaced, as the dual-temporal image used for training;

[0021] When training the feature extraction model, the multi-scale training module processes the multi-scale first feature map from the feature extraction module through the multi-core parallel convolution module, the feature enhancement module and the dual-channel merging unit respectively to generate a multi-scale detection map P j , j represents the j-th scale; the loss is jointly calculated using the detection maps P of all scales j to optimize the feature extraction model; during inference detection, only the detection map P corresponding to the first feature map of the deepest layer is used as the inference result.

[0022] Preferably, when the FFT module replaces the low-frequency region of the dual-temporal image, for a series of dual-temporal images, it randomly determines to copy the low-frequency region of the first image in the dual-temporal image to the second image, or copy the low-frequency region of the second image to the first image.

[0023] Preferably, the feature extraction module uses a ResNet-18 backbone network with the fully connected layer removed.

[0024] Preferably, the feature extraction module includes a first convolutional layer Conv1 and four residual layers Res1, Res2, Res3, and Res4 connected in sequence; during training, three-scale first feature maps are extracted from the residual layers Res1, Res2, and Res4 and respectively input into the multi-core parallel convolution module for processing to obtain three-scale second feature maps; during inference and prediction, only the deepest first feature map output by the residual layer Res4 is input into the multi-core parallel convolution module for processing.

[0025] Preferably, in the feature extraction module, the first convolutional layer Conv1 uses a 7×7 convolutional kernel, and the residual layers Res2 and Res3 perform downsampling with a stride of 2.

[0026] Preferably, in the multi-core parallel convolution module, there are 5 parallel convolution branches, and the convolutional kernels are 3×3, 5×5, 7×7, 9×9, and 11×11 respectively; the 5 processing results output by the parallel convolution branches are further fused through feature connection and convolution with a 1×1 convolutional kernel to obtain the second feature map.

[0027] Preferably, the feature enhancement module enhances the feature representation of the changed region through position attention, spatial attention, and cross-attention mechanisms.

[0028] Preferably, the feature enhancement module uses a cross-scale feature aggregator, which consists of a position attention PA module, a spatial attention SA module, and the encoder and decoder of the Transformer module;

[0029] The PA module enhances the key positions of the second feature map in the in-channel dimension and outputs a third feature map ; the SA module enhances the key regions of the third feature map in the multi-channel dimension to generate a feature map T; the encoder encodes the feature map T to generate context-containing tokens and transmits them to the decoder; the decoder uses a cross-attention mechanism, using the third feature map generated by the PA module as the query, and the tokens As key and value respectively, generate enhanced features, that is, extract feature maps .

[0030] Preferably, the PA module comprises a channel number uniform convolution layer, a feature adjustment module, a multi-channel point-by-point convolution layer, a normalization module and a weighting module connected in sequence;

[0031] The number of channels is unified in the convolution layer, which is used to convert the second feature map of the input into Unify to a fixed number of channels M, generate the first intermediate feature map X0 of M channels and output it to the feature adjustment module;

[0032] The feature adjustment module includes a batch normalization layer, a linear rectification activation function and a Dropout module, processes the first intermediate feature map X0, and generates a second intermediate feature map X1 of M channels;

[0033] The multi-channel point-by-point convolution layer includes M 1×1 convolution kernels, which perform convolution on corresponding channels of the second intermediate feature map X1 without sharing the convolution kernels, and generate the third intermediate feature X2;

[0034] The normalization module uses a Sigmoid activation function to normalize the value of each position of the third intermediate feature X2 to the range of [0,1], and generates a position attention weight map, in which each value represents the importance of the corresponding position in the first intermediate feature map X0;

[0035] The weighting module is used to perform element-by-element multiplication of the first intermediate feature map X0 and the position attention weight map to achieve position enhancement and obtain the third feature map .

[0036] Preferably, the two-way merging unit comprises a connection module and a classification module; the connection module extracts feature maps from the two paths and Splicing; The classification module uses a convolutional layer to compare the spliced ​​feature maps to obtain a detection map of cultivated land changes.

[0037] Beneficial effects:

[0038] (1) Adaptive Feature Extraction and Timeliness Enhancement: Traditional change detection methods rely on technicians to manually select features. The algorithm effect is limited by experience and ability, and a large number of repeated experiments are required to determine parameters, making it difficult to ensure high robustness and timeliness. The system of the present invention effectively combines the powerful feature extraction and parameter optimization capabilities of deep convolutional neural networks, and can adaptively extract features by using iterative operations and gradient updates according to the data characteristics. In the application stage, the trained weight model can be loaded to directly perform feature extraction without repeated experiments and manual tuning, significantly improving timeliness and application value. Compared with traditional methods, the system of the present invention reduces the dependence on manual intervention and avoids feature selection biases caused by insufficient experience of technicians. In addition, automated feature extraction and model optimization greatly shorten the development cycle and improve the speed and efficiency of model deployment.

[0039] (2) Multi-scale Feature Fusion Optimization: Existing methods are prone to introducing background noise or losing key information when capturing multi-scale features, and it is difficult to balance fine-grained features and global modeling. The system of the present invention adopts a multi-core Inception module (MKI) to capture diverse feature information from local details to global structures through convolutional kernels of different sizes, avoiding background noise caused by large convolutional kernels and information loss caused by dilated convolutions. At the same time, the MKI module integrates multi-scale information, maintaining the integrity and accuracy of features, and significantly enhancing the richness and accuracy of feature representation. Traditional methods usually use single-scale convolutional kernels, making it difficult to comprehensively capture multi-scale features in images and easily leading to background noise interference and loss of important information. The system of the present invention, through the design of parallel convolution branches, not only enhances the feature expression ability but also improves the understanding of complex scenes.

[0040] (3) Improvement of Adaptability to Complex Scenes: Existing methods are difficult to handle visual feature changes caused by differences in imaging conditions such as illumination, perspective, and background, and are prone to false change misdetection. The system of the present invention introduces the fast Fourier transform (FFT) to transform the two-temporal images from the spatial domain to the frequency domain and replace the low-frequency components, unifying the styles of the two different-temporal images. This method effectively eliminates the visual feature differences caused by different imaging conditions and reduces false change misdetection. Randomly applying FFT enhances the generalization ability of the model, enabling it to adapt to more diverse data distributions and thus perform more stably in practical applications. Traditional methods are prone to false change misdetection when processing images under different imaging conditions due to visual feature differences. The system of the present invention effectively eliminates these differences through frequency domain processing, enabling the model to maintain consistent detection performance under different imaging conditions and significantly improving robustness and reliability.

[0041] (4) Long-distance Dependence Capturing: In the preferred embodiment, considering that existing methods are difficult to effectively capture long-distance dependencies in images, especially performing poorly when dealing with large-scale changes; the feature enhancement module in the system of the present invention combines position attention (PA) and spatial attention (SA), enhances the feature representation within the key regions, and focuses on the key regions of the entire image. By using the Transformer encoder (TE) and decoder (TD), global context information is captured to achieve an effective understanding of long-distance dependencies. This design significantly improves the model's understanding and parsing capabilities for complex scenes. Traditional methods are often limited to local feature extraction and are difficult to capture long-distance dependencies in images, which limits their performance in large-scale change detection. The system of the present invention can more effectively capture and utilize global context information by introducing a multi-scale attention mechanism and a Transformer architecture, thereby significantly improving the detection accuracy and reliability. Description of the Drawings

[0042] Figure 1 It is a schematic diagram of the multi-scale cultivated land change detection system based on frequency domain conversion of the present invention.

[0043] Figure 2 It is a block diagram of the composition of the multi-scale cultivated land change detection system based on frequency domain conversion in the preferred embodiment of the present invention.

[0044] Figure 3 It is a schematic diagram of the FFT module.

[0045] Figure 4 It is a schematic diagram of the multi-scale CNN and multi-core Inception module.

[0046] Figure 5 It is a schematic diagram of the skip connection of the residual layer.

[0047] Figure 6 It is a schematic diagram of the cross-scale feature aggregator (CSFA).

[0048] Figure 7 It is a schematic diagram of the position attention (PA) module.

[0049] Figure 8 It is a schematic diagram of the Transformer encoder and decoder.

[0050] Figure 9Visualization of experimental results on the CLCD dataset; white represents true positive examples, black represents true negative examples, purple represents false positive examples, and green represents false negative examples. (a) Phase 1; (b) Phase 2; (c) Labels; (d) FC-EF; (e) FC-Siam-conc; (f) FC-Siam-diff; (g) DTCDSCN; (h) SNUNet; (i) BIT; (j) MSCANet; (k) SARASNet; (l) HATNet; (m) The proposed solution of the present invention.

[0051] Figure 10 Visualization of experimental results on the WHU-CD and LEVIR-CD datasets; white represents true positive examples, black represents true negative examples, purple represents false positives, and green represents false negatives. (a) Phase 1; (b) Phase 2; (c) Labels; (d) FC-EF; (e) FC-Siam-conc; (f) FC-Siam-diff; (g) DTCDSCN; (h) SNUNet; (i) BIT; (j) MSCANet; (k) SARASNet; (l) HATNet; (m) The proposed solution of the present invention.

[0052] Figure 11 Visualization of the FFT effect; (a) The left side shows the image of time stage 1, the middle shows the image of time stage 2, and the right side shows the processed image of time stage 1; (b) is the frequency domain image of the red band corresponding to (a).

[0053] Figure 12 Visualization of CSFA. (a) Dual-temporal image. (b) Multi-scale feature maps F1 and F2 extracted by MKI; (c) Attention maps P1 and P2 generated by PA; (d) Attention maps S1 and S2 generated by SA; (e) Fine feature maps T1 and T2 generated by Transformer. (f) Change map. Detailed implementation manners

[0054] The present invention provides a multi-scale cultivated land change detection system based on frequency domain conversion, and its core improvement lies in:

[0055] (1) During training, the present invention maps the dual-temporal images from the spatial domain to the frequency domain through the Fast Fourier Transform (FFT) and performs frequency domain feature interaction to optimize the feature representation, thereby effectively alleviating the inevitable visual feature differences in the change detection task and reducing the interference of pseudo-changes.

[0056] (2) The present invention constructs a feature extraction module and a multi-core parallel convolution module, and uses an improved ResNet-18 and a parallel small convolution kernel architecture in the Inception style to perform the extraction and progressive fusion of multi-scale features, promoting the information coupling between intra-layer features, and avoiding the background noise caused by large convolution kernels and the information loss caused by dilated convolutions.

[0057] (3) In a preferred embodiment, a Cross-scale feature aggregator (CSFA) is designed to integrate multi-scale features through attention mechanisms and Transformer structures at different scales and enhance context understanding, ensuring effective global modeling while capturing fine-grained features.

[0058] The present invention will be described in detail below with reference to the accompanying drawings and by way of examples.

[0059] Figure 1 The schematic diagram of the multi-scale cultivated land change detection system based on frequency domain conversion according to the present invention is shown as Figure 1 shown. The system includes a feature extraction model, a dual-path merging unit, and a training unit.

[0060] The feature extraction model processes each path of the input dual-temporal image to obtain two paths of extracted feature maps and ; The dual-path merging unit merges and classifies the two paths of extracted feature maps and to obtain the detection map P of cultivated land change.

[0061] The feature extraction model includes a feature extraction module, a multi-core parallel convolution module, and a feature enhancement module. Among them, the feature extraction module is used to extract features from the two paths of images in the dual-temporal image to obtain multi-scale first feature maps , where j represents the jth scale, and i = 1, 2.

[0062] The multi-core parallel convolution module contains multiple parallel convolution branches. Each parallel convolution branch uses m convolution kernels of different sizes, and the convolution kernels are less than or equal to 13×13; the value range of m is 3-7; the multi-scale first feature maps enter each parallel convolution branch for processing, and then are spliced and fused to obtain the second feature map .

[0063] The feature enhancement module is used to enhance the features of the second feature map to generate the extracted feature map .

[0064] The training unit includes an FFT module and a multi-scale training module; among them,

[0065] An FFT module, which is used to convert a dual-temporal image from the spatial domain to the frequency domain when training a feature extraction model, unify the low-frequency regions in the dual-temporal image through replacement, and then convert it back to the spatial domain to generate a dual-temporal image with the low-frequency components replaced, which is used as the dual-temporal image for training;

[0066] When training a feature extraction model, the multi-scale training module extracts multi-scale first feature maps from multiple residual layers of the feature extraction module , and the first feature maps of each scale are respectively processed by a multi-core parallel convolution module, a feature enhancement module and a dual-path merging unit to generate multi-scale detection maps P j ; The loss is jointly calculated using the detection maps P of all scales j to optimize the feature extraction model. During inference detection, only the detection map P corresponding to the first feature map of the deepest layer is used as the inference result.

[0067] In a preferred embodiment, the above-mentioned feature extraction module is implemented by a modified ResNet-18; the multi-core parallel convolution module is implemented by an Inception-style parallel small convolution kernel architecture; the feature extraction module and the multi-core parallel convolution module jointly achieve multi-scale feature extraction and progressive fusion, alleviate the problem of gradient disappearance, expand the receptive field while avoiding the introduction of background noise and information loss. These features together improve the accuracy and robustness of change detection, especially showing excellent performance when processing complex high-resolution remote sensing images. The experimental results verify the effectiveness and superiority of the module, providing a solid foundation for future research and applications.

[0068] The feature enhancement module uses a Cross-scale Feature Aggregator (CSFA) to fuse multi-scale features extracted in the early stage and adaptively assign different attention weights according to the scale differences of target changes, so as to effectively capture subtle spatial changes and context dependencies. CSFA can be applied in different computer vision tasks, such as image classification, object detection, semantic segmentation, etc.

[0069] The present invention further provides a preferred implementation of the above multi-scale cultivated land change detection system, which specifically provides an architecture - FIMANet for fine-grained feature extraction and global modeling applicable to different imaging conditions. By combining the Fast Fourier Transform (FFT), multi-scale feature extraction module (CNN), multi-kernel Inception module (MKI), and cross-scale feature aggregator (CSFA), this architecture effectively solves the problem of false change detection caused by differences in illumination, perspective, and background in dual-temporal images. This method converts the image from the spatial domain to the frequency domain through FFT, replaces the low-frequency components of the target image, and reduces the style differences caused by imaging conditions; the MKI module gradually extracts and refines multi-scale features to avoid background noise and information loss; the CSFA module captures global and local dependencies through the attention mechanism and Transformer to improve the detection accuracy. By mining multi-scale features in high-resolution remote sensing image change detection,

[0070] Figure 2 shows the functional block diagram of this preferred embodiment. The following will describe the present invention in detail in conjunction with Figure 2 the present invention.

[0071] As Figure 2 shown, the multi-scale cultivated land change detection system based on frequency domain conversion includes an FFT module, a multi-scale CNN module (Multi-Scale CNN), a multi-kernel Inception module (MKI, Multiple Kernel Inception), a cross-scale feature aggregator (CSFA, Cross-scale feature aggregator), and an output module (Output in the figure). The system should also include a multi-scale training module not shown in the figure. The system composed of these modules is called FIMANet (Frequency-Based Detailed Multi-Scale Approach for Remote Sensing Change Detection) in the present invention.

[0072] Among them, the FFT module and the multi-scale training module form a training unit, the multi-scale CNN module realizes the function of the feature extraction module, the multi-kernel Inception module serves as a multi-kernel parallel convolution module, the CSFA realizes the function of the feature enhancement module, and the output module realizes the function of the dual-channel merging unit. The multi-scale CNN module, the multi-kernel Inception module, and the CSFA form a feature extraction model, and the parameters and weights of each grid layer involved need to be trained, including but not limited to the convolutional layer, batch normalization layer, behavior of the Dropout layer, weights of the Sigmoid activation function, ResNet-18 backbone network, etc. mentioned below.

[0073] The training unit is responsible for model training. The FFT module is only used during training. During actual inference, the dual-phase images are directly input into the multi-scale CNN module. During training, the multi-scale CNN module extracts multi-scale feature maps. After each scale of feature map is processed by the multi-core Inception module, CSFA, and the output module, multi-scale detection maps are obtained. Then, using the detection maps P of all scales j calculate the loss jointly with the ground truth to optimize the feature extraction model. During actual inference, only the detection map P corresponding to the first feature map of the deepest layer is used as the inference result to obtain the predicted detection map P.

[0074] The following is a detailed description of each of the above modules.

[0075] (1) FFT Module

[0076] To reduce the domain difference between dual-phase images caused by changes in imaging conditions (such as illumination, shadow, and background differences), frequency-domain feature interaction is adopted to reduce the domain gap and promote information coupling of intra-layer representations. This FFT module transforms the dual-phase images from the spatial domain to the frequency domain; among them, regions with drastic intensity changes correspond to high-frequency components (mapping the edges and contours of the image), while relatively smooth regions correspond to low-frequency components (providing a comprehensive mapping of the overall image intensity). By replacing the frequency-domain components of the dual-phase in the frequency domain and then applying the inverse FFT to transform back to the spatial domain, the visual features of the two images can be effectively unified.

[0077] As Figure 3 shown, the specific implementation process of the FFT module is as follows:

[0078] Input image preprocessing: Input the dual-phase images. The dual-phase images are divided into two paths, denoted as and respectively, ensuring that they have been registered to ensure pixel alignment at the same geographical location.

[0079] Forward Fourier transform: Apply 2D FFT (two-dimensional fast Fourier transform) to the dual-phase images and respectively to generate the corresponding spectrograms F1 and F2. The spectrogram represents the distribution of the image in the frequency domain, where the low-frequency components correspond to the overall brightness and smooth regions of the image, and the high-frequency components correspond to the edges and changes of the image. Taking as an example, the conversion formula for each channel i of the RGB three-channel image is as follows:

[0080]

[0081] where and respectively represent the height and width of the image, is the coordinate in the spatial domain of the image, is the coordinate in the frequency domain of the image, represents the image in the corresponding spectrogram frequency value at the position, represents the image in pixel at the position.

[0082] Low-frequency component replacement: In the frequency domain, swap the low-frequency components of the images and Here, the swap means copying the low-frequency component of one image to the other image. Among them, the low-frequency component is determined by a set threshold, and the threshold size is determined according to the experimental results by enumerating on three public datasets. Let Δ represent the size of the frequency domain sampling points in the low-frequency region, and its value is determined by the enumeration experiment. In one embodiment, the size of Δ is determined to be 10, as shown in Table 2. Taking the example of replacing the low-frequency component in the spectrogram F1 with the corresponding low-frequency component in the spectrogram F2, the formula is expressed as:

[0083]

[0084] In the formula, is the result after the low-frequency component in the spectrogram F1 is replaced. represents the data in the spectrogram F2.

[0085] In practice, the low-frequency component in the spectrogram F2 can also be replaced with the corresponding low-frequency component in the spectrogram F1.

[0086] After replacing the low-frequency component, the spectrogram F1 changes to the spectrogram F1′, and the spectrogram F2 changes to the spectrogram F2′.

[0087] Inverse Fourier transform: Apply the inverse FFT to the spectrogram F1′, convert it back to the spatial domain, and generate a new image . Ensure that and are more visually consistent, reducing the impact of imaging condition differences on the model. The formula is expressed as:

[0088]

[0089] and constitute a new dual-temporal image, to be input into the multi-scale CNN module.

[0090] If the low-frequency component in the spectrogram F2 is replaced with the corresponding low-frequency component in the spectrogram F1, then apply the inverse FFT to the spectrogram F2′, convert it back to the spatial domain, and generate a new image 。 and constitute a new dual - temporal image, which is to be input into the multi - scale CNN module.

[0091] It should be noted that the object replaced by the low - frequency component of the FFT module is random. It is randomly determined to copy the low - frequency region of the first - path image in the dual - temporal image to the second - path image, or copy the low - frequency region of the second - path image to the first - path image. Formula 2 is used to determine which pixels to replace. By randomly applying the FFT transformation to the dual - temporal image, the generalization ability and adaptability of the model are increased, while no replacement is required in actual detection.

[0092] This FFT module is used during training and no replacement is performed during actual detection.

[0093] (2)Multi - scale CNN

[0094] The multi - scale CNN adopts an improved ResNet - 18. ResNet - 18 is a classic convolutional neural network with good feature extraction ability and relatively shallow network depth, which is suitable for the change detection task of high - resolution remote sensing images.

[0095] The multi - scale CNN is improved based on the ResNet - 18 structure. The improvement contents include: (1) Removing the final fully - connected layer: Focusing on multi - scale feature extraction, avoiding overfitting to specific classification tasks, aiming to capture multi - scale features between various resblocks. (2) Retaining the residual connection: Alleviating the problem of gradient disappearance through skip connections to ensure that deep networks can also be effectively trained.

[0096] As Figure 4 shown in the left - hand half of the figure, in this embodiment, the multi - scale CNN includes a first convolutional layer Conv1 and four residual layers Res1, Res2, Res3, and Res4 connected in sequence. During training, three - scale first - feature maps are extracted from the residual layers Res1, Res2, and Res4, denoted here as 、 、 , with sizes of 128×128×64, 64×64×128, 32×32×512, where 64, 128, and 512 are the number of channels. The three feature maps are respectively input into the subsequent multi - core Inception module for processing to obtain three - scale second - feature maps Second - feature maps , j = 1, 2, 3. During inference and prediction, these three second - feature maps still enter the multi - core Inception module to undergo subsequent various processes, but only the deepest - layer feature map is used to obtain the detection map P 3 as the final prediction result.

[0097] Among them, the first convolutional layer Conv1 directly extracts basic low-level features, such as edges and textures, from the input image using a 7x7 convolutional kernel. The parameters of the Conv1 layer are a stride of 2, a padding of 3, and the number of output channels is 64.

[0098] The residual block contains four residual layers Res1 to Res4. Each layer is responsible for gradually extracting and refining the image features, transitioning from low-level details to high-level semantic features. The residual connection alleviates the vanishing gradient phenomenon and effectively reduces the size of the feature map through a downsampling operation with a stride of 2, effectively halving the size of the feature map and expanding the receptive field. Inside each residual block, there are two paths. One is the ordinary convolutional path, and the other is the identity mapping path. The two are added together and then passed to the next layer. See Figure 5 . Figure 5 shows the structure of a residual layer. The input x is skip-connected to the output. At the same time, the input x passes through two weight layers to generate an intermediate product F(x), which is then added to the input x and output.

[0099] (3) Multi-Kernel Inception Module (MKI)

[0100] To further extract multi-scale features and avoid problems caused by using large convolutional kernels or dilated convolutions, FIMANet uses multiple parallel convolutional kernels to extract features of different scales, which can reduce the introduction of background noise and maintain the integrity of key information.

[0101] As Figure 4 shown in the right half of the figure, in this preferred embodiment, in the multi-kernel parallel convolution module, there are 5 parallel convolution branches, and the convolutional kernels are 3×3, 5×5, 7×7, 9×9, and 11×11 respectively; the 5 processing results output by the parallel convolution branches are further fused through feature connection and convolution with a 1×1 convolutional kernel to obtain the second feature map Second Feature Map . In the present invention, F is used to represent the feature map, the subscript i represents which path belongs to the dual-temporal image, the superscript j represents the jth scale, and the number in the parentheses represents the number of times of processing.

[0102] The output feature map corresponding to the convolutional kernel in each convolution branch can be expressed as follows:

[0103] (5)

[0104] Assume the input feature Figure X has a size of C. Here represents the convolution operation, its convolutional kernel size is , and the number of output channels is Multiple such operations can be concatenated along the channel dimension. Then, the output feature maps of all convolutional operations are concatenated along the channel dimension to produce a new feature map, that is:

[0105] (6)

[0106] Finally, the concatenated feature map is fused through a 1×1×C convolutional layer, and the resulting fused feature map ∈ . This fused feature map is the second feature map . The convolutional fusion is expressed as:

[0107] (7)

[0108] In the formula, is the weight of the 1×1 convolutional kernel on channel c.

[0109] As a channel integration technology, the 1×1 convolution combines features from different receptive field sizes, which enables the MKI module to include comprehensive context information while maintaining the integrity of local texture features.

[0110] (4) Cross-scale Feature Aggregator CSFA

[0111] CSFA enhances the feature representation of the changing regions through position attention, spatial attention, and cross-attention mechanisms.

[0112] As Figure 2 in the middle and Figure 6 shown, CSFA consists of a position attention (PA) module, a spatial attention (SA) module, and the encoder and decoder of the Transformer module. Its function is to perform adaptive attention assignment on multi-scale features and capture long-range dependencies using the Transformer.

[0113] The PA module enhances the key positions of the second feature map in the channel dimension, and outputs the third feature map . The SA module enhances the key regions of the third feature map in the multi-channel dimension to generate the feature map T. The encoder encodes the feature map T to generate context-rich tokens , which are passed to the decoder. The decoder adopts the cross-attention mechanism, using the third feature map generated by the PA module as the query, and the tokens from the encoderGenerate enhanced features, i.e., extract feature maps, using them as the key and value respectively 。

[0114] The PA module selectively enhances the key regions in the feature map, enabling the model to focus more on these significant regions in subsequent processing stages and improving the accuracy of feature extraction. Different from the position attention module that only focuses on the significant features within specific channels, the SA module aims to identify the key regions in the entire image globally and generates more semantically expressive semantic markers by assigning adaptive weights to these important regions. This mechanism allows the model to model in a broader spatial context, thereby capturing global feature relationships. And a Transformer encoder (TransformerEncoder, TE) is introduced to capture long-range dependencies. It can be seen that CSFA enhances features by combining the processing of three dimensions: position in the channel, between channels, and context. This mechanism effectively combines information from different feature sources and significantly enhances the modeling ability of global context and local details.

[0115] Figure 7 The structure of the PA module in the preferred embodiment of the present invention is shown. The feature of this structure lies in the design of the multi-channel pointwise convolutional layer: each output channel has its own independent convolutional kernel instead of all channels sharing a convolutional kernel. Each convolutional kernel only performs convolution with the corresponding input channel, avoiding the mixing between channels. The size and number of channels of the convolution result are still the number of output channels.

[0116] As shown in Figure 7, the PA module includes a channel number unified convolutional layer, a feature adjustment module, a multi-channel pointwise convolutional layer, a normalization module, and a weighting module connected in sequence.

[0117] The channel number unified convolutional layer is used to unify the input second feature map (also simply referred to as X in the figure) to a fixed number of channels M by performing a convolution operation, generating a first intermediate feature map X0 with M channels and outputting it to the feature adjustment module. In this embodiment, the channel number unified convolutional layer is of convolution operation, unifying the multi-scale feature maps (with channel numbers of 64, 128, and 512 respectively) extracted in the previous stage to a fixed number of channels M = 32. Such processing not only ensures the unity of the number of channels but also retains the expressive ability of multi-scale features.

[0118] The feature adjustment module includes a batch normalization layer, a linear rectification activation function, and a Dropout module, which processes the first intermediate feature map X0 to generate the second intermediate feature map X1 of M channels, further enhancing the stability and nonlinear expression ability of the features. Batch normalization, ReLU activation, and Dropout do not directly affect the generation of the weight matrix, but are used to adjust the representation of the feature map, prevent overfitting, and improve the nonlinear expression ability of the model.

[0119] The multi-channel point-by-point convolution layer is the core. This module includes M 1×1 convolution kernels, which perform convolution on the corresponding channels of the second intermediate feature map X1 without sharing the convolution kernels to generate the third intermediate feature X2. The 1×1 convolution kernel uses a 1×1 depth convolution, which is used to model each channel separately and mine more fine-grained features. The setting of this point-by-point convolution layer is that the convolution kernel size is 1 and the number of output channels is grouped as M, which means that channel-by-channel convolution is performed. In this embodiment, each output channel has its own independent convolution kernel, a total of 32, instead of all channels sharing the convolution kernel. Each convolution kernel is convolved only with the corresponding input channel to avoid mixing between channels. The size of the result and the number of channels after convolution are still the number of output channels.

[0120] The normalization module uses the Sigmoid activation function ( ) Normalize the value of each position of the third intermediate feature X2 to the range of [0, 1] and generate a position attention weight map, in which each value represents the importance of the corresponding position in the first intermediate feature map X0.

[0121] A weighting module is used to multiply the first intermediate feature map X0 and the position attention weight map element by element (⊙) to achieve position enhancement and obtain the third feature map , this feature map is a position weighted feature map, also simplified as Y in the figure.

[0122] This process of the PA module not only retains the original spatial information of the input features, but also enhances the expression ability of key areas, significantly improving the model's perception of multi-scale targets and its ability to capture local details.

[0123] Next, the feature map Y output by the PA module is passed to the SA module for further processing. Unlike the PA module, which only focuses on the salient features in a specific channel, the SA module aims to identify key areas in the entire image globally and generate more semantically expressive semantic tags by assigning adaptive weights to these important areas. This mechanism allows the model to model in a wider spatial context, thereby capturing global feature relationships.

[0124] Specifically, the process starts with steps similar to those in natural language processing, through a learnable kernel Perform pointwise convolution on the input feature map Y. The purpose of this operation is to divide the feature map into L semantic groups to achieve the grouping and representation of semantic features. This step can not only retain local features but also enhance the information separation and aggregation between different semantic groups through the learned weights. Subsequently, the model applies the Softmax operation on the spatial dimension of each semantic group to calculate the corresponding spatial attention weights. These weights represent the importance of each spatial position and can effectively capture the global spatial dependence. Finally, the calculated weights are multiplied element-wise with the input feature map Y to generate the final set of tokens , which integrates the global information of the spatial context and the details of local significant features. The entire process can be formally expressed as:

[0125] (8)

[0126] where (•) represents pointwise convolution, (•) denotes the Softmax operation along the channel dimension.

[0127] Finally, a Transformer Encoder (TE) is introduced to capture long-range dependencies. First, a set of trainable parameters are added element-wise to the tokens to complete position embedding, generating context-rich tokens . These tokens are then passed to the Transformer Decoder (TD).

[0128] In the decoder, a cross-attention mechanism is adopted, where the feature map generated by the PA module serves as the query (Q), while the outputs from the Transformer decoder serve as the key ( ) and value (V), respectively. The corresponding formula is as follows:

[0129] (9)

[0130] where , and represent the weight matrices of the linear layers used to project Q, K, and V into a specific representation space, respectively. This mechanism effectively combines information from different feature sources and significantly enhances the ability to model global context and local details.

[0131] See Figure 8 Figure 8 , the improved Siamese network structure of the preferred embodiment of the present invention is used for change map generation. When generating a change map by traditional methods, it is difficult to balance local details and overall structure, resulting in inaccurate detection results. The present invention innovatively splices and fuses features: the features processed by the CSFA module are decoded by the Transformer decoder (TD), and then the decoded features are spliced together and input into the prediction head (Prediction Head) to generate the final change map. At the same time, multi-scale feature integration is performed: through the improved Siamese network structure, effective global modeling is achieved while capturing fine-grained features. The generated change map shows the detected change areas, which can be directly used for analysis or auxiliary decision-making, providing more accurate and reliable change detection results.

[0132] (5) Output module (dual-path merging unit)

[0133] This output module includes a connection module and a classification module. Among them, the connection module concatenates the two extracted feature maps and ; the classification module uses a convolutional layer to compare the concatenated feature maps to obtain the detection map P of cultivated land change.

[0134] During training, the feature maps of multiple scales generated by the multi-scale CNN are processed by the MKI, CSFA modules and the output module to form detection maps P of multiple scales j . In actual detection and inference, only the detection map corresponding to the feature map output by the last residual layer of the multi-scale CNN is used as the detection result.

[0135] (6) Training unit

[0136] In the training unit, when the multi-scale training module jointly calculates the loss using all scales of detection maps P j and the ground truth, the binary cross-entropy (BCE) loss is used as the objective function. The BCE loss is one of the most widely used loss functions in binary classification and image segmentation because it can assign different weights to each pixel and evaluate the similarity of dual-temporal images based on spatial information. The definition of the BCE loss is as follows:

[0137] (10)

[0138] where N represents the total number of pixels in the image, represents the ground truth label of the i-th pixel in the ground truth map, taking values of 0 or 1. is the probability that the i-th pixel in the detection map P is predicted as class 1, usually the predicted value calculated by the sigmoid activation function.

[0139] , and represent the output detection maps at different scales respectively. Let G denote the ground truth. The total loss function of FIMANet uses the cross-entropy ( ) loss formula as follows:

[0140] (11)

[0141] Next, the effectiveness of the present invention is verified.

[0142] 1. Dataset Selection

[0143] To verify the effectiveness of the present invention, three challenging change detection datasets are selected: CLCD, WHU-CD, and LEVIR-CD. The CLCD dataset focuses on farmland changes and contains 600 pairs of bi-temporal images with a size of 512x512 pixels and a resolution of 0.5 - 2 meters. The dataset is divided into training set, validation set, and test set, which is suitable for evaluating change detection algorithms in the agricultural field. WHU-CD and LEVIR-CD focus on building change detection and contain 1 pair of bi-temporal images with a size of 32,507×15,354 pixels and 637 pairs of bi-temporal images with a size of 1024×1024 pixels respectively, with spatial resolutions of 0.2 meters and 0.5 meters. The data of both are cropped to generate small blocks of 256×256 pixels and randomly divided into training, validation, and test sets, which are suitable for evaluating the algorithm performance.

[0144] 2. Experimental Environment Setup

[0145] The experimental environment is based on PyTorch 2.2.2 and CUDA 11.8, and the deep learning model is trained on an NVIDIA RTX 4090 GPU. The Adam optimizer is adopted with an initial learning rate of 0.0001, beta parameters of (0.9, 0.999), and a cosine annealing scheduler is used to dynamically adjust the learning rate with the minimum value set to 1e-5. The training lasts for 100 epochs, and validation is performed after each epoch. The model with the highest intersection over union (IoU) on the validation set is selected as the final model.

[0146] 3. Dataset Training

[0147] The datasets (CLCD, WHU-CD, LEVIR-CD) are cropped, normalized, and data-augmented during the preprocessing stage to improve the generalization ability of the model. The preprocessed datasets are divided into training set, validation set, and test set. During the training process, the model is optimized using batch gradient descent, and a learning rate scheduler and early stopping method are used to prevent overfitting. At the same time, the performance stability of the model on different datasets is evaluated through cross-validation.

[0148] 4. Experimental Method Selection

[0149] To verify the effectiveness of the present invention, a variety of advanced change detection methods were selected for comparison, including: FC-EF (based on the U-Net architecture, extracting joint features through image-level fusion); FC-Siam-conc (based on the Siamese network, processing dual-temporal images and extracting features); FC-Siam-diff (calculating pixel-by-pixel differences to capture changes); DTCDSCN (combining channel and spatial attention mechanisms to enhance features); SNUNet (combining the Siamese network and Nested-Unet for multi-scale feature extraction); BIT (combining CNN and transformer, using self-attention to capture context information); MSCANet (integrating a multi-scale context aggregator and a multi-branch prediction head); SARASNet (introducing relationship awareness and scale sensitivity to solve the problem of scene changes); HATNet (combining self-attention and cross-attention mechanisms for building change detection).

[0150] 5. Result Comparison

[0151] In the result comparison stage, a variety of evaluation metrics were used to quantify the performance of different methods, including Precision, Recall, F1 Score, Intersection over Union (IoU), and overall accuracy (OA), etc. The experimental results show that the proposed solution of the present invention has achieved significantly better performance than the baseline methods on three datasets.

[0152] (1) Quantitative Experimental Comparison

[0153] Table 1 summarizes the quantitative results of each model on the CLCD, WHU-CD, and LEVIR-CD test sets. FIMANet is better than all comparison models in all five metrics.

[0154] Table 1

[0155]

[0156] In the above table, the best scores are shown in bold, and all results are presented as percentages (%).

[0157] (3) Qualitative Result Comparison

[0158] Figure 9 and Figure 10The map of the prediction results of the model on different datasets is shown, with different color codings used to present the detection results: True Positives (TP) are represented by white, True Negatives (TN) by black, False Positives (FP) by purple, and False Negatives (FN) by green. Through this color mapping method, the accuracy and error types of the model prediction are visually highlighted. The fewer green and purple areas in the figure indicate that the algorithm performs well in recall and precision, and can more effectively identify real change areas and reduce false detections. The experimental results prove that the proposed FIMANet achieves SOTA (state-of-the-art) performance on all three datasets.

[0159] (4)FFT Intermediate Frequency Domain Sampling Point Parameter Setting and Visualization

[0160] To evaluate the influence of the frequency domain sampling size ∆ on the formula, it is verified by gradually increasing in 11 steps with a step size of 5 sampling points). Table 2 shows that when the frequency domain sampling points increase from 0 to 10, the F1 score and IoU on the CLCD, WHU-CD, and LEVIR-CD datasets gradually increase, indicating that FFTFDFI effectively enhances the information coupling of the dual-temporal images and reduces the interference of pseudo-changes. However, as the interaction area further expands, the F1 score and IoU decrease, because excessive replacement leads to the loss of source image content and the generation of artifacts. Therefore, setting ∆ to 10 is the most appropriate.

[0161] To further analyze the influence, Figure 11 (a) in shows the image after FFT processing. To further analyze the influence, Figure 11 (a) in shows the image after FFT processing, visually closer to more. Figure 11 (b) in shows the frequency domain maps of the red bands of nine images, with the relative mean energy spectral density (RMESD) marked in the upper left corner of the figure. After logarithmic transformation and normalization, the RMESD values of (10548, 9349, 10413) are closer to (9956, 8928, 10108), while different from the RMESD values of (11295, 12547, 10463). In the frequency domain, the information of the image is decomposed into sine and cosine waves of different frequencies, enabling the observation of different-scale features or patterns in the image here, such as textures, periodic structures, etc. Figure 11 (b) in shows that the frequency domain image shows the distribution of the original image on different frequency components. The centrally symmetric pattern represents zero frequency, and the points far from the center correspond to high-frequency information. Frequency domain analysis helps to reveal the spatial variation characteristics of the image and is crucial for image processing tasks such as filtering and feature extraction.

[0162] Table 2

[0163]

[0164] In the above table, the best scores are shown in bold, and all results are expressed as percentages (%).

[0165] (5)Visualization of the CSFA Module

[0166] To more intuitively understand the CSFA module, an example is selected from the CLCD dataset here, and multiple representative dual-temporal feature maps at each stage are visualized. As Figure 12 shown, the multi-scale feature maps extracted by the MKI module and . Figure 12 (b) of shows clear boundary contours and a highly compact internal structure, but the model shows relatively low global attention. Subsequently, the PA module assigns per-pixel attention weights to and to generate the feature maps and ( Figure 12 (c) of). In these images, obvious contour features of key attention areas can be observed, such as individual farmland plots and backfill areas. Then, the feature maps and ( Figure 12 (d) of) obtained through the SA module further enhance the expression of intra-regional variation features, enabling the model to more effectively aggregate meaningful spatial features. Finally, the and ( Figure 12 (e) of) obtained through the Transformer module show a fine understanding of high-level semantic relationships. For example, in , all farmland areas are closely clustered together, while in , the boundary between the farmland and the backfill is clearly distinguished. The final detection result map also verifies the effectiveness of the CSFA module in accurately assigning attention to the dual-temporal feature maps.

[0167] The above specific embodiments only describe the design principle of the present invention. The shapes and names of the components in this description can be different and are not limited. Therefore, those skilled in the art of the present invention can modify or equivalently replace the technical solutions recorded in the foregoing embodiments; and these modifications and replacements do not depart from the spirit and technical solutions of the present invention, and shall all fall within the protection scope of the present invention.

Claims

1. A multi-scale farmland change detection system based on frequency domain conversion, the system comprising a feature extraction model, a dual-path merging unit and a training unit; characterized in that: The feature extraction model processes each channel of the input dual-phase image to obtain two channels of extracted feature maps. and ; The feature extraction model includes a feature extraction module, a multi-core parallel convolution module, and a feature enhancement module; A feature extraction module is used to extract features from two images in a dual-phase image to obtain a multi-scale first feature map; the feature extraction module uses a ResNet-18 backbone network with the fully connected layer removed, including a first convolutional layer Conv1 and four residual layers Res1, Res2, Res3 and Res4 connected in sequence; during inference prediction, only the deepest first feature map output by the residual layer Res4 is used to input into the multi-core parallel convolution module for processing; The multi-core parallel convolution module includes multiple parallel convolution branches. The parallel convolution branches use m convolution kernels of different sizes, and the convolution kernel is less than or equal to 13×13; the value range of m is 3-7; the first feature map enters each parallel convolution branch for processing, and the m processing results output by the parallel convolution branch are further fused through feature connection and convolution of 1×1 convolution kernel to obtain the second feature map ; Feature enhancement module, used to enhance the second feature map Perform feature enhancement and generate extracted feature maps ; The two-way merging unit extracts feature maps from two paths and Merge and classify to obtain the detection map P of cultivated land changes; The training unit includes an FFT module and a multi-scale training module; The FFT module is used to convert the dual-phase image from the spatial domain to the frequency domain when training the feature extraction model, unify the low-frequency area in the dual-phase image through the replacement operation, and then convert it back to the spatial domain to generate a dual-phase image after the low-frequency component is replaced as the dual-phase image used for training; the replacement operation is: the dual-phase image is divided into two paths, respectively recorded as and , in the frequency domain, swapping images and The exchange refers to copying the low-frequency component of one image to the other image; When training the feature extraction model, the multi-scale training module extracts the first feature maps of three scales from the residual layers Res1, Res2 and Res4 of the feature extraction module; the multi-scale first feature maps from the feature extraction module are processed by the multi-core parallel convolution module, the feature enhancement module and the dual-path merging unit to generate a multi-scale detection map P j , j represents the jth scale; using the detection map P of all scales j The loss is calculated jointly to optimize the feature extraction model. During inference detection, only the detection map P corresponding to the first feature map of the deepest layer is used as the inference result.

2. The multi-scale farmland change detection system based on frequency domain conversion according to claim 1, characterized in that: When the FFT module replaces the low frequency area of ​​the dual-phase image, for a series of dual-phase images, it randomly determines to copy the low frequency area of ​​the first image in the dual-phase image to the second image, or to copy the low frequency area of ​​the second image to the first image.

3. The multi-scale farmland change detection system based on frequency domain conversion according to claim 1, characterized in that: In the feature extraction module, the first convolution layer Conv1 uses a 7×7 convolution kernel, and the residual layers Res2 and Res3 are downsampled with a step size of 2.

4. The multi-scale farmland change detection system based on frequency domain conversion according to claim 1, characterized in that: In the multi-core parallel convolution module, there are 5 parallel convolution branches, and the convolution kernels are 3×3, 5×5, 7×7, 9×9 and 11×11 respectively; the 5 processing results output by the parallel convolution branches are further fused through feature connection and convolution of the 1×1 convolution kernel to obtain a second feature map.

5. The multi-scale farmland change detection system based on frequency domain conversion according to claim 1, characterized in that: The feature enhancement module enhances the feature representation of the changing area through position attention, spatial attention and cross attention mechanisms.

6. The multi-scale farmland change detection system based on frequency domain conversion according to claim 5, characterized in that: The feature enhancement module adopts a cross-scale feature aggregator; The cross-scale feature aggregator consists of a position attention PA module, a spatial attention SA module, and the encoder and decoder of the Transformer module; The dimension of the PA module within the channel affects the second feature map Enhance the key positions and output the third feature map ; The SA module performs the third feature map in the multi-channel dimension Enhance the key areas and generate feature map T; the encoder encodes the feature map T to generate a tag with context , passed to the decoder; The decoder uses a cross-attention mechanism to generate the third feature map of the PA module. As query, the tokens from the encoder As key and value respectively, generate enhanced features, that is, extract feature maps .

7. The multi-scale farmland change detection system based on frequency domain conversion according to claim 6, characterized in that: The PA module includes a convolution layer with a uniform number of channels, a feature adjustment module, a multi-channel point-by-point convolution layer, a normalization module and a weighting module connected in sequence; The number of channels is unified in the convolution layer, which is used to convert the second feature map of the input into Unify to a fixed number of channels M, generate the first intermediate feature map X0 of M channels and output it to the feature adjustment module; The feature adjustment module includes a batch normalization layer, a linear rectification activation function and a Dropout module, processes the first intermediate feature map X0, and generates a second intermediate feature map X1 of M channels; The multi-channel point-by-point convolution layer includes M 1×1 convolution kernels, which perform convolution on corresponding channels of the second intermediate feature map X1 without sharing the convolution kernels, and generate the third intermediate feature X2; The normalization module uses a Sigmoid activation function to normalize the value of each position of the third intermediate feature X2 to the range of [0,1] to generate a position attention weight map; each value in the position attention weight map represents the importance of the corresponding position in the first intermediate feature map X0; The weighting module is used to perform element-by-element multiplication of the first intermediate feature map X0 and the position attention weight map to achieve position enhancement and obtain the third feature map .

8. The multi-scale farmland change detection system based on frequency domain conversion according to claim 1, characterized in that: The dual-path merging unit includes a connection module and a classification module; the connection module extracts feature maps from the two paths and Splicing; The classification module uses a convolutional layer to compare the concatenated feature maps to obtain a detection map of cultivated land changes.

Citation Information

Patent Citations

  • Remote sensing image change detection method based on time-space interaction Transform model

    CN117095287A

  • Visual time sequence feature network implementation method based on multi-module feature fusion

    CN119206847A

  • Remote sensing image change detection network for multi-scale feature extraction and information mining

    CN119380210A