A multi-scale feature fusion cultivated land change detection method and system

By employing a multi-scale feature fusion method for farmland change detection, utilizing Siamese networks and multi-head attention computation, and combining enhancements to change and invariant features, the method solves the problems of unclear detection results and missed or false detections in farmland change detection, achieving high-precision and high-reliability farmland change detection.

CN116704367BActive Publication Date: 2025-12-09WUHAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310655413.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-02
Publication Date
2025-12-09
Estimated Expiration
2043-06-02

AI Technical Summary

Technical Problem

Existing methods for detecting changes in cultivated land are prone to problems such as unclear detection results, missed detections, and false detections when detecting minor changes.

Method used

A multi-scale, multi-level feature fusion method for detecting farmland change is adopted by using a multi-level sliding window. By using the multi-scale feature fusion method for detecting farmland change, a Siamese network is used for location encoding, and multi-head attention computation is introduced to obtain the dual-scale feature fusion farmland change detection results.

Benefits of technology

It improves the accuracy and reliability of farmland change detection and reduces the probability of missed and false detections of small changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116704367B_ABST
    Figure CN116704367B_ABST
Patent Text Reader

Abstract

The application provides a kind of multi-scale feature fusion cultivated land change detection method and system, belong to remote sensing technical field, including: to double time phase optical remote sensing image is carried out multi-scale feature extraction, obtains double time phase image multi-scale multi-level feature;Based on the change feature and invariant feature enhancement method of priori guide, extract double time phase image multi-level difference feature of double time phase image multi-scale multi-level feature;To double time phase image multi-level difference feature is carried out multi-level feature fusion and constructs cultivated land change feature map, obtains cultivated land change feature map change detection result using preset deep supervision strategy.The application aims at the problem that multiple small change targets exist in target cultivated land, leading to boundary blur and edge ambiguity in change detection, uses multi-level sliding window to obtain multi-scale multi-level features, and uses feature enhancement for both change features and invariant features, improving the saliency of change areas, thereby greatly reducing the probability of missed detection and false detection of small change targets.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of remote sensing, and in particular to a multi-scale feature fusion cultivated land change detection method and system. BACKGROUND

[0002] Dual temporal remote sensing image change detection technology aims to extract the change region in the dual temporal remote sensing image that people are interested in. It has been widely applied in various fields such as land use, urbanization process, natural disaster investigation, and has achieved remarkable results. Influenced by climate change and human activities, the phenomenon of cultivated land occupation is becoming more and more serious, which has a negative impact on the local agricultural ecosystem and global food security. Therefore, how to quickly and accurately extract the change region in the dual temporal remote sensing image within the cultivated land has become a hot spot and a difficult problem in the field of remote sensing.

[0003] The existing change detection methods can be roughly divided into two categories: traditional change detection methods and deep learning change detection methods. The traditional change detection method mainly obtains the change region by judging the spectral difference between single pixels or uses the spatial context information of the image for object-level change detection. However, this method either only focuses on the information within a single pixel, resulting in a certain degree of salt and pepper noise and fragmentation in the detection result, or is highly dependent on the object segmentation accuracy and is not sufficient to deal with complex ground conditions in multi-temporal high-resolution images. The deep learning-based change detection method can simultaneously extract features from dual temporal images and automatically identify the change region, realizing end-to-end remote sensing image change detection and outputting a change map. Thanks to the powerful information mining and feature representation capabilities of deep learning, the deep learning-based change detection method has the characteristics of fast speed, high accuracy, and high robustness. However, with the continuous improvement of the performance of remote sensing satellites, the width of remote sensing images is increasing, and the ground objects within the image are also becoming more refined. When detecting small changes within the cultivated land, the existing methods often result in unclear detection results, missed detection, and false detection.

[0004] Therefore, in view of the deficiencies of the above detection methods, a cultivated land change detection method with high precision and high reliability needs to be proposed. SUMMARY

[0005] The present application provides a multi-scale feature fusion cultivated land change detection method and system to solve the problem that the existing cultivated land change detection methods are not fine enough for detecting small changes, and are prone to have unclear detection results, missed detection, and false detection.

[0006] In a first aspect, the present application provides a multi-scale feature fusion cultivated land change detection method, comprising:

[0007] Collecting double-phase optical remote sensing images of the farmland to be detected, performing multi-scale feature extraction on the double-phase optical remote sensing images to obtain double-phase image multi-scale multi-level features;

[0008] Based on the prior guided change feature and invariant feature enhancement method, the double-phase image multi-level difference features of the double-phase image multi-scale multi-level features are extracted;

[0009] The double-phase image multi-level difference features are fused to construct a farmland change feature map, and a preset deep supervision strategy is used to obtain a change detection result of the farmland change feature map.

[0010] According to the multi-scale feature fusion farmland change detection method provided by the application, double-phase optical remote sensing images of the farmland to be detected are collected, multi-scale feature extraction is performed on the double-phase optical remote sensing images to obtain double-phase image multi-scale multi-level features, including:

[0011] The double-phase optical remote sensing images are position encoded to obtain the relative position information of each part of the double-phase optical remote sensing images;

[0012] Based on the encoding mode and size information of the double-phase optical remote sensing images, different sizes and different numbers of hierarchical windows corresponding to the double-phase optical remote sensing images are determined;

[0013] The encoding mode and size information of the double-phase optical remote sensing images are input into a twin network, and multi-head attention calculation is performed using the different sizes and different numbers of hierarchical windows to obtain double-phase image multi-scale features;

[0014] The double-phase image multi-scale features are fused by using a multilayer perceptron to obtain the double-phase image multi-scale multi-level features.

[0015] According to the multi-scale feature fusion farmland change detection method provided by the application, the double-phase optical remote sensing images are position encoded to obtain the relative position information of each part of the double-phase optical remote sensing images, including:

[0016] The double-phase optical remote sensing images are cut into a plurality of non-overlapping patch blocks, and the plurality of non-overlapping patch blocks are used as tokens;

[0017] The token dimension is obtained, and the token belongs to the position and a plurality of odd position elements and a plurality of even position elements in the token dimension are determined;

[0018] Based on the sine function of the token belonging to the position and the plurality of even position elements, and the cosine function of the token belonging to the position and the plurality of odd position elements, the relative position information of the plurality of non-overlapping patch blocks is obtained.

[0019] According to the multi-scale feature fusion cultivated land change detection method provided by the application, the different sizes and different number of hierarchical windows corresponding to the double-phase optical remote sensing images are determined based on the encoding mode and size information of the double-phase optical remote sensing images, and the multi-scale feature fusion cultivated land change detection method comprises the following steps:

[0020] The hierarchical and input size of the double-phase optical remote sensing images are obtained by using the original size information and network structure of the double-phase optical remote sensing images.

[0021] It is determined that the global structured features of the double-phase optical remote sensing images are extracted by using a preset large window, the local fine features of the double-phase optical remote sensing images are extracted by using a preset small window, and the number of hierarchical windows is determined according to the feature map size of the double-phase optical remote sensing images.

[0022] It is determined that the number of hierarchical window types of the preset shallow features is greater than the number of hierarchical window types of the preset deep features.

[0023] According to the multi-scale feature fusion cultivated land change detection method provided by the application, the encoding mode and size information of the double-phase optical remote sensing images are input into a twin network, multi-head attention calculation is performed by using the different sizes and different number of hierarchical windows, and double-phase image multi-scale features are obtained.

[0024] The double-phase optical remote sensing images are subjected to dimension-preserving linear transformation respectively, and matrix Q, matrix K and matrix V are obtained.

[0025] The matrix Q is multiplied by the transpose matrix K, divided by the square root of the element dimension of the matrix Q, and then normalized by using a softmax function to obtain an attention matrix Attn.

[0026] The matrix V is multiplied by the attention matrix Attn to obtain a correlation calculation matrix, the correlation calculation matrices corresponding to multiple self-attention are spliced, and the spliced matrix is converted by using a conversion matrix to obtain an output feature matrix.

[0027] It is determined that a sliding window is used in the different sizes and different number of hierarchical windows to obtain multiple output feature matrices.

[0028] The multiple output feature matrices are aggregated and connected to obtain the double-phase image multi-scale features.

[0029] According to the multi-scale feature fusion cultivated land change detection method provided by the application, based on the prior guided change feature and invariant feature enhancement method, the double-phase image multi-level difference features of the double-phase image multi-scale multi-level features are extracted, and the multi-scale feature fusion cultivated land change detection method comprises the following steps:

[0030] Calculate the cosine similarity between the double-time image multi-scale multi-level features to obtain double-time feature prior information;

[0031] Based on the double-time feature prior information, the change features and the invariant features in the double-time image multi-scale multi-level features are enhanced to obtain enhanced double-time features.

[0032] The enhanced double-time features are fused by using a multi-layer perception machine to extract multi-level difference features of the double-time images.

[0033] According to the multi-scale feature fusion cultivated land change detection method provided by the application, based on the double-time feature prior information, the change features and the invariant features in the double-time image multi-scale multi-level features are enhanced to obtain enhanced double-time features, including:

[0034] The first-time image multi-scale multi-level features and the second-time image multi-scale multi-level features of the double-time image multi-scale multi-level features are determined;

[0035] The absolute value of the difference between the first-time image multi-scale multi-level features and the second-time image multi-scale multi-level features is calculated to obtain a feature difference, and the sum of the first-time image multi-scale multi-level features and the second-time image multi-scale multi-level features is calculated to obtain a feature sum;

[0036] The first change region enhanced features and the first invariant region enhanced features of the double-time feature prior information and the first-time image multi-scale multi-level features, and the second change region enhanced features and the second invariant region enhanced features of the double-time feature prior information and the second-time image multi-scale multi-level features are calculated respectively;

[0037] The first change region enhanced features, the first invariant region enhanced features, the second change region enhanced features, the second invariant region enhanced features, the feature difference and the feature sum are sequentially aggregated, channel attention calculated and fused by using a multi-layer perception machine to obtain the enhanced double-time features.

[0038] According to the multi-scale feature fusion cultivated land change detection method provided by the application, the multi-level difference features of the double-time image are fused to construct a cultivated land change feature map, and a preset deep supervision strategy is used to obtain a change detection result of the cultivated land change feature map, including:

[0039] The multi-level difference features of the double-time image are decoded and restored level by level to obtain a plurality of cultivated land change feature maps, wherein the size of each cultivated land change feature map gradually decreases with the deepening of the network;

[0040] The cross-entropy loss function is used to determine pixel-by-pixel classification of the plurality of cultivated land change feature maps by the classifier to obtain the change detection result.

[0041] In a second aspect, the present application further provides a cultivated land change detection system based on multi-scale feature fusion, comprising:

[0042] A multi-scale feature extraction module is configured to collect double-time optical remote sensing images of cultivated land to be detected, and perform multi-scale feature extraction on the double-time optical remote sensing images to obtain double-time multi-scale multi-level features.

[0043] A difference feature extraction module is configured to extract double-time multi-level difference features of the double-time multi-scale multi-level features based on a priori guided change feature and invariant feature enhancement method.

[0044] A fusion detection module is configured to perform multi-level feature fusion on the double-time multi-level difference features to construct a cultivated land change feature map, and obtain a change detection result of the cultivated land change feature map by using a preset deep supervision strategy.

[0045] In a third aspect, the present application further provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned multi-scale feature fusion cultivated land change detection method when executing the program.

[0046] The multi-scale feature fusion cultivated land change detection method and system provided by the present application can solve the problem of blurred boundaries and unclear edges in change detection caused by multiple small change targets in the target cultivated land. The multi-level sliding window is used to obtain multi-scale multi-level features, and the features are enhanced for both change features and invariant features, which improves the saliency of the change region, thereby greatly reducing the probability of missed detection and false detection of small change targets. BRIEF DESCRIPTION OF DRAWINGS

[0047] In order to more clearly illustrate the technical solutions in the present application or prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor.

[0048] Figure 1 is one of the flowcharts of the multi-scale feature fusion cultivated land change detection method provided by the present application;

[0049] Figure 2 is the second flowchart of the multi-scale feature fusion cultivated land change detection method provided by the present application;

[0050] Figure 3 is a change detection network diagram provided by the present application;

[0051] Figure 4 is a structural schematic diagram of a multi-scale feature fusion cultivated land change detection system provided by the present application;

[0052] Figure 5 is a structural schematic diagram of an electronic device provided by the present application. DETAILED DESCRIPTION

[0053] In order to make the objects, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below in combination with the drawings in the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0054] Figure 1 is one of the flowcharts of the multi-scale feature fusion cultivated land change detection method provided by the embodiments of the present application, as shown in Figure 1 , comprising:

[0055] Step 100: collecting double-time optical remote sensing images of cultivated land to be detected, performing multi-scale feature extraction on the double-time optical remote sensing images, and obtaining double-time image multi-scale multi-level features;

[0056] Step 200: based on the change feature and invariant feature enhancement method guided by priori, extracting double-time image multi-level difference features of the double-time image multi-scale multi-level features;

[0057] Step 300: performing multi-level feature fusion on the double-time image multi-level difference features to construct a cultivated land change feature map, and obtaining a change detection result of the cultivated land change feature map by using a preset deep supervision strategy.

[0058] The embodiments of the present application propose a multi-scale feature fusion change detection method based on Transformer using hierarchical windows for cultivated land change detection scenarios, as shown in Figure 2 , comprising the following steps:

[0059] First, a feature extraction encoder is constructed, and the obtained double-time optical remote sensing images are input into the feature extraction encoder after position encoding for multi-scale feature extraction, to obtain corresponding double-time image multi-scale multi-level features, Figure 2The pre-phase image in the method is the remote sensing image before the change, and the post-phase image is the remote sensing image after the change, which are respectively input into a feature extraction encoder, and then weight sharing is performed by a twin network, and then a difference feature extraction module is constructed, multi-scale multi-level features of the double-phase images are input into the difference feature extraction module, multi-level difference features of the double-phase images are extracted by a change-invariant feature enhancement method guided by priori, and finally, a feature restoration decoder is constructed, multi-level feature fusion is performed on the multi-level difference features of the double-phase images processed by the feature extraction encoder and the difference feature extraction module, the cultivated land change feature map is gradually reconstructed, and deep supervision strategies are adopted to obtain change detection results at different depths.

[0060] It should be noted that the network framework adopted by the embodiment of the present application is generally in a U-shaped structure, and a classical encoder-decoder structure in a convolutional neural network is adopted, as shown in Figure 3 The network framework mainly includes three parts: a novel shared weight twin encoder for extracting multi-scale features of double-phase images, a double-phase difference feature extraction module based on change-invariant feature enhancement, and a decoder based on a deep supervision multi-level feature fusion strategy.

[0061] Firstly, like VIT, Swin and other pure Transformer networks, the double-phase RGB images 1 and 2 (the size is HxWx3, H represents the image height, W represents the image width, and 3 is the three channels of RGB) of the same region are non-overlappingly cropped into several patch blocks, which are aggregated and connected together to form a token sequence (H / 4xW / 4x48), and then input into the twin encoder for multi-scale multi-level feature extraction, forming a hierarchical output similar to the encoder of the convolutional neural network. Secondly, each level of double-phase multi-scale features output by the twin encoder is input into the difference feature extraction module for change-invariant region feature enhancement, and then aggregated into double-phase difference features. Finally, the multi-level double-phase fusion features are input into each level of the decoder for multi-level feature fusion, and gradually restored to the original size of the image through multiple decoding units. Finally, the different depths of the predicted change maps obtained by each deep supervision branch are aggregated together to form the final change map.

[0062] The present application obtains richer boundary detail features by fusing multi-scale information perceived by the hierarchical window, and then obtains a change result with clearer edges and better separation; secondly, the change and non-change regions in the double-phase images are respectively enhanced in features, the saliency of the change region is improved, the situation that small change targets in the cultivated land range are missed or misdetected is effectively reduced, and finally the high-precision and high-reliability detection of small change targets in the cultivated land range is realized.

[0063] On the basis of the above embodiment, the step 100 comprises:

[0064] positionally encode the dual-time optical remote sensing image to obtain relative position information of each part of the dual-time optical remote sensing image;

[0065] determine different sizes and different numbers of hierarchical windows corresponding to the dual-time optical remote sensing image based on the encoding mode and size information of the dual-time optical remote sensing image;

[0066] input the encoding mode and size information of the dual-time optical remote sensing image into a twin network, perform multi-head attention calculation using the different sizes and different numbers of hierarchical windows, and obtain dual-time image multi-scale features;

[0067] fuse the dual-time image multi-scale features using a multi-layer perception machine to obtain dual-time image multi-scale multi-level features.

[0068] The positional encoding of the dual-time optical remote sensing image to obtain the relative position information of each part includes:

[0069] cut the dual-time optical remote sensing image into a plurality of non-overlapping patch blocks, and use the plurality of non-overlapping patch blocks as tokens;

[0070] obtain token dimensions, determine the position to which the token belongs and a plurality of odd position elements and a plurality of even position elements in the token dimensions;

[0071] obtain the relative position information of the plurality of non-overlapping patch blocks based on the sine function of the position to which the token belongs and the plurality of even position elements, and the cosine function of the position to which the token belongs and the plurality of odd position elements, respectively.

[0072] The determination of different sizes and different numbers of hierarchical windows corresponding to the dual-time optical remote sensing image based on the encoding mode and size information of the dual-time optical remote sensing image includes:

[0073] obtain the level and input size of the dual-time optical remote sensing image using the original size information and network structure information of the dual-time optical remote sensing image;

[0074] determine that a preset large window is used to extract global structured features of the dual-time optical remote sensing image, a preset small window is used to extract local fine features of the dual-time optical remote sensing image, and the number of hierarchical windows is determined by the feature map size of the dual-time optical remote sensing image;

[0075] determine that the number of hierarchical window types of a preset shallow feature is greater than the number of hierarchical window types of a preset deep feature.

[0076] The encoding mode and size information of the dual-time optical remote sensing image are input into a twin network, multi-head attention calculation is performed by using different sizes and different numbers of hierarchical windows, and dual-time image multi-scale features are obtained, including:

[0077] The dual-time optical remote sensing image is subjected to a dimension-preserving linear transformation, respectively, to obtain a matrix Q, a matrix K, and a matrix V.

[0078] The matrix Q is multiplied by the transpose matrix K, divided by the square root of the dimension of the matrix K, and then normalized by using a softmax function to obtain an attention matrix Attn.

[0079] The matrix V is multiplied by the attention matrix Attn to obtain a correlation calculation matrix, the correlation calculation matrices corresponding to multiple self-attention are spliced, and the spliced matrix is converted by a conversion matrix to obtain an output feature matrix.

[0080] A sliding window is used in the different sizes and different numbers of hierarchical windows to obtain a plurality of output feature matrices.

[0081] The plurality of output feature matrices are aggregated and connected to obtain the dual-time image multi-scale features.

[0082] Specifically, first, the obtained dual-time optical remote sensing image is subjected to position encoding for extracting position information, and the position encoding method is specifically as follows:

[0083]

[0084]

[0085]

[0086] wherein, represents the i-th element in the position vector, k represents the odd-even position distinction of the i elements, t represents the position to which the token belongs, d model represents the dimension of the token.

[0087] Second, for a given position-encoded dual-time optical remote sensing image input, different scales and numbers of windows are divided according to different levels and input sizes to obtain the most suitable hierarchical window.

[0088] For a given input, the encoding unit sets hierarchical windows to capture multi-scale feature information of the image according to different levels and input sizes. The reason is that a larger window is suitable for extracting global structured features, and a smaller window is suitable for extracting local fine features, and the scale number of the hierarchical window is related to the size of the input feature map.

[0089] Meanwhile, when extracting multi-scale feature information, shallow features will use more window types than deep features, because shallow features have higher resolution, which represent the structure information and texture information of the object. More window types will make the extracted information more comprehensive. While deep features have smaller resolution and represent semantic information of the object, using as many window types as shallow features will cause information redundancy and computational burden. Therefore, the network divides different scales of windows for each stage, and the window type decreases with the deepening of the network.

[0090] Further, the double-time-phase optical remote sensing image after position encoding is input into the twin encoder, multi-head attention calculation is performed using hierarchical windows, and double-time-phase image multi-scale features are obtained.

[0091] Among them, the calculation of multi-head self-attention can be represented as:

[0092] Q=W q X

[0093] K=W k X

[0094] V=W v X

[0095]

[0096] Y=AttnxV

[0097] First, the position information X of the input double-time-phase optical remote sensing image is subjected to three dimension-preserving linear transformations to obtain Q (Query), K (Key), V (Value), W q , W k and W v represent the coefficients of X corresponding to Q, K and V respectively. Then, the matrix Q is multiplied by the matrix K to obtain the attention matrix Attn, and is used to normalize the attention matrix, d represents the element dimension of Q, the variance of the matrix element is corrected to 1, and the row vector of the matrix is normalized by using softmax to obtain the weight of each V, dim represents the dimension. Finally, Attn and V are multiplied to obtain the result Y after correlation calculation.

[0098]

[0099] In order to capture various correlations between vectors, multiple Q vectors and K vectors need to be introduced, and multiple calculation results Y 1 …Y H(H represents the number of heads) are spliced together, and a conversion matrix W is used to convert the spliced Y matrix to the original size to obtain the final output feature matrix O.

[0100] Multi-head self-attention is a commonly used module in various Transformer methods (VIT, DeiT, etc.). It extracts feature information by performing multiple global self-attention calculations on the input, but also brings a huge computational burden. Its computational complexity is the square of the input size. To alleviate the burden of huge computation, Swin proposed window multi-head self-attention, which performs self-attention calculation within each window by dividing the window in advance, greatly reducing the computation. At the same time, it improves the model performance by calculating the relative position encoding index.

[0101] Because the previous window division is not overlapping, only calculating self-attention within the window will cause a gap in information between multiple windows. To increase information exchange between adjacent windows, Swin adopts a sliding window multi-head self-attention module, which realizes information exchange with other windows by introducing a sliding window operation, and also ingeniously adopts an attention mask method, allowing window multi-head self-attention and sliding window multi-head self-attention to have equivalent calculation results under the same window number.

[0102] Using hierarchical window calculation to obtain multi-scale features of dual-time images;

[0103] F multl = Concat(swin1(F input ),…,Swin l (F input ))

[0104] F input ,F multl represent the input of the encoder and the calculated multi-scale features, respectively, Swin represents a unit of the encoder, l represents the serial number of the hierarchical window, and Concat represents aggregation connection.

[0105] Finally, the multi-scale features of the dual-time images are fused using a multi-layer perceptron to obtain the fused multi-scale features of the dual-time images;

[0106] Feature fusion through a multi-layer perceptron can be represented as:

[0107] F output = MLP{F multl}

[0108] MLP represents a multi-layer perceptron, and F output represents the output of the dual-time image multi-scale multi-level features.

[0109] On the basis of the above-mentioned embodiments, step 200 comprises:

[0110] calculating the cosine similarity between the multi-scale multi-level features of the dual-time images to obtain dual-time feature prior information;

[0111] based on the dual-time feature prior information, performing feature enhancement on the changing features and invariant features in the multi-scale multi-level features of the dual-time images to obtain enhanced dual-time features;

[0112] using a multi-layer perception to fuse the enhanced dual-time features to extract multi-level difference features of the dual-time images.

[0113] In the method, based on the dual-time feature prior information, the feature enhancement on the changing features and invariant features in the multi-scale multi-level features of the dual-time images to obtain enhanced dual-time features comprises:

[0114] determining the multi-scale multi-level features of the first-time images and the multi-scale multi-level features of the second-time images of the dual-time images;

[0115] calculating the absolute value of the difference between the multi-scale multi-level features of the first-time images and the multi-scale multi-level features of the second-time images to obtain feature difference, and calculating the sum of the multi-scale multi-level features of the first-time images and the multi-scale multi-level features of the second-time images to obtain feature sum;

[0116] respectively calculating the first changing region enhanced features and the first invariant region enhanced features of the dual-time feature prior information and the multi-scale multi-level features of the first-time images, and the second changing region enhanced features and the second invariant region enhanced features of the dual-time feature prior information and the multi-scale multi-level features of the second-time images;

[0117] sequentially performing aggregation connection, channel attention calculation and fusion using a multi-layer perception on the first changing region enhanced features, the first invariant region enhanced features, the second changing region enhanced features, the second invariant region enhanced features, the feature difference and the feature sum to obtain the enhanced dual-time features.

[0118] Specifically, for the input multi-scale multi-level features of the dual-time images, the dual-time feature prior information is obtained by calculating the cosine similarity between the multi-scale multi-level features of the dual-time images, and the calculation method is as follows:

[0119]

[0120] wherein, F1 and F2 respectively represent the multi-scale multi-level features of the first-time images and the multi-scale multi-level features of the second-time images in the multi-scale multi-level features of the dual-time images, and respectively represent any feature of the pre-phase image and any feature of the post-phase image, i represents any element, and n represents all elements.

[0121] Next, guided by prior information, the change-invariant region in the dual-phase feature is enhanced to obtain the enhanced dual-phase feature;

[0122] The change feature is indeed crucial in change detection, but for the binary classification task, enhancing the invariant region feature and enhancing the change region feature are equivalent, so the invariant feature should be on an equal footing with the change feature and should essentially give the same attention to the invariant feature and the change feature. Considering the feature that the cosine similarity can reveal the similarity between the dual-phase feature maps, the network innovatively proposes a change-invariant feature enhancement method, and the model can obtain more accurate change regions by comparing the enhanced feature maps. First, the calculated cosine similarity of the dual-phase feature maps is taken as a prior condition, which is multiplied by the corresponding elements of the input feature maps to obtain the activation feature maps of the corresponding feature maps. The activation feature maps retain the original feature structure while increasing the difference between the change region and the invariant region in the single feature map. The characteristics of the activation feature maps are that the feature values of the invariant region are positive, the feature values of the change region are negative, and the greater the change degree, the smaller the feature value (including negative increase). Therefore, by adding the activation feature maps to the original features to enhance the features of the invariant region, and subtracting the activation feature maps from the original features to enhance the features of the change region, the feature maps of the dual-phase feature maps after change-invariant feature enhancement are obtained.

[0123] wherein the feature enhancement method and the difference feature extraction method can be represented as:

[0124] F Absub = |F1-F2|; F Add = F1+F2

[0125] F CE = F x |Cossim-1|; F IE = F x (Cossim+1)

[0126]

[0127] wherein F Absub represents the feature difference, F Add represents the feature sum, F1, F2 represent the pre-phase image multi-scale multi-level feature and the post-phase image multi-scale multi-level feature in the dual-phase image multi-scale multi-level feature, F D represents the dual-phase image difference feature, Cossim represents the cosine similarity calculation, F CE represents the change region feature after enhancement, and FIE Concat represents aggregation connection, CA represents channel attention calculation, and MLP represents a multi-layer perception, respectively represent the features of the enhanced pre-phase change region features, the enhanced pre-phase invariant region features, the enhanced post-phase change region features, and the enhanced post-phase invariant region features.

[0128] Finally, the enhanced dual-phase features are fused by using a multi-layer perception to extract dual-phase image multi-level difference features.

[0129] On the basis of the above embodiment, step 300 comprises:

[0130] The dual-phase image multi-level difference features are decoded and recovered level by level to obtain a plurality of cultivated land change feature maps, wherein the size of each cultivated land change feature map gradually decreases with the deepening of the network;

[0131] A cross-entropy loss function is used to determine that the plurality of cultivated land change feature maps are classified pixel by pixel by a classifier to obtain the change detection result.

[0132] Specifically, the dual-phase image multi-level difference features are input into a feature recovery decoder, the size of the cultivated land feature map is gradually recovered step by step through a plurality of decoding stages, and finally a cultivated land feature map with the same size as the input image is obtained.

[0133] Finally, the plurality of change feature maps generated by the deep supervision strategy are classified pixel by pixel by a classifier to obtain the final change detection result.

[0134] Here, cross-entropy is used as a loss function:

[0135]

[0136] wherein n represents the number of categories contained in the training data, m=2 is set in cultivated land remote sensing image change detection, containing two categories of change and non-change; (i,j) represents the position of the pixel on the image, the value range of i is 0 to W-1, and the value range of j is 0 to H-1; y represents the real ground surface change result corresponding to the input image, y (i,j) represents the real category of the pixel at the position (i,j), and the value range of y (i,j) is 0 to m-1; x represents the prediction score of the network for the input image, x (i,j) [ ] represents the score of the network model predicting that the category of the pixel at the position (i,j) is s, and the value range of s is 0 to m-1.

[0137] The multi-scale feature fusion cultivated land change detection system provided by the present application is described below, and the multi-scale feature fusion cultivated land change detection system described below can be correspondingly referred to the multi-scale feature fusion cultivated land change detection method described above.

[0138] Figure 4 is a structural schematic diagram of the multi-scale feature fusion cultivated land change detection system provided by the embodiment of the present application, as Figure 4 shown, comprising: a multi-scale feature extraction module 41, a difference feature extraction module 42 and a fusion detection module 43, wherein:

[0139] The multi-scale feature extraction module 41 is used for collecting double-time optical remote sensing images of cultivated land to be detected, performing multi-scale feature extraction on the double-time optical remote sensing images, and obtaining double-time image multi-scale multi-level features; the difference feature extraction module 42 is used for extracting double-time image multi-level difference features of the double-time image multi-scale multi-level features based on a priori guided change feature and invariant feature enhancement method; the fusion detection module 43 is used for performing multi-level feature fusion on the double-time image multi-level difference features to construct a cultivated land change feature map, and obtaining a change detection result of the cultivated land change feature map by using a preset deep supervision strategy.

[0140] Figure 5 An example of an entity structure schematic diagram of an electronic device is shown in Figure 5 , which can include a processor 510, a communications interface 520, a memory 530 and a communications bus 540, wherein the processor 510, the communications interface 520 and the memory 530 complete mutual communication through the communications bus 540. The processor 510 can invoke the logic instructions in the memory 530 to execute the multi-scale feature fusion cultivated land change detection method, which includes: collecting double-time optical remote sensing images of cultivated land to be detected, performing multi-scale feature extraction on the double-time optical remote sensing images, and obtaining double-time image multi-scale multi-level features; based on a priori guided change feature and invariant feature enhancement method, extracting double-time image multi-level difference features of the double-time image multi-scale multi-level features; performing multi-level feature fusion on the double-time image multi-level difference features to construct a cultivated land change feature map, and obtaining a change detection result of the cultivated land change feature map by using a preset deep supervision strategy.

[0141] In addition, the logic instructions in the memory 530 described above can be implemented in the form of software functional units and sold or used as independent products, and can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the parts that contribute to the prior art or parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0142] In another aspect, the present application also provides a computer program product, which comprises a computer program, the computer program can be stored on a non-transitory computer readable storage medium, and the computer program can be executed by a processor to enable a computer to execute the multi-scale feature fusion cultivated land change detection method provided by the above-mentioned methods. The method comprises the following steps: collecting double-time optical remote sensing images of cultivated land to be detected, performing multi-scale feature extraction on the double-time optical remote sensing images to obtain double-time image multi-scale multi-level features; based on the change feature and invariant feature enhancement method guided by priori, extracting double-time image multi-level difference features of the double-time image multi-scale multi-level features; performing multi-level feature fusion on the double-time image multi-level difference features to construct a cultivated land change feature map, and obtaining a change detection result of the cultivated land change feature map by using a preset deep supervision strategy.

[0143] In another aspect, the present application also provides a computer program product, which comprises a computer program, the computer program can be stored on a non-transitory computer readable storage medium, and the computer program can be executed by a processor to enable a computer to execute the multi-scale feature fusion cultivated land change detection method provided by the above-mentioned methods. The method comprises the following steps: collecting double-time optical remote sensing images of cultivated land to be detected, performing multi-scale feature extraction on the double-time optical remote sensing images to obtain double-time image multi-scale multi-level features; based on the change feature and invariant feature enhancement method guided by priori, extracting double-time image multi-level difference features of the double-time image multi-scale multi-level features; performing multi-level feature fusion on the double-time image multi-level difference features to construct a cultivated land change feature map, and obtaining a change detection result of the cultivated land change feature map by using a preset deep supervision strategy.

[0144] The device embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed to multiple network units. Part or all of the modules can be selected to achieve the purposes of the embodiments according to actual needs. Those skilled in the art can understand and implement without creative labor.

[0145] Through the description of the above embodiments, those skilled in the art can clearly understand that the embodiments can be realized by means of software and the necessary general hardware platform, and of course can also be realized by hardware. Based on such understanding, the above technical solutions can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.

[0146] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A multi-scale feature fusion cultivated land change detection method, characterized in that, The method comprises the following steps: Collecting double-phase optical remote sensing images of the farmland to be detected, and performing multi-scale feature extraction on the double-phase optical remote sensing images to obtain double-phase image multi-scale multi-level features; Based on the prior guided change feature and invariant feature enhancement method, the double-phase image multi-level difference features of the double-phase image multi-scale multi-level features are extracted, including: Calculate the cosine similarity between the double-phase image multi-scale multi-level features to obtain double-phase feature prior information; Based on the double-phase feature prior information, the change features and invariant features in the double-phase image multi-scale multi-level features are enhanced to obtain enhanced double-phase features, including: Determine the first-phase image multi-scale multi-level features and the second-phase image multi-scale multi-level features of the double-phase image multi-scale multi-level features; Calculate the absolute value of the difference between the first-phase image multi-scale multi-level features and the second-phase image multi-scale multi-level features to obtain the feature difference, and the sum of the first-phase image multi-scale multi-level features and the second-phase image multi-scale multi-level features to obtain the feature sum; Calculate the first change region enhancement feature and the first invariant region enhancement feature of the double-phase feature prior information and the first-phase image multi-scale multi-level features, and the second change region enhancement feature and the second invariant region enhancement feature of the double-phase feature prior information and the second-phase image multi-scale multi-level features, respectively; The first change region enhancement feature, the first invariant region enhancement feature, the second change region enhancement feature, the second invariant region enhancement feature, the feature difference and the feature sum are sequentially aggregated, channel attention calculated and fused by using a multi-layer perception machine to obtain the enhanced double-phase features; The enhanced double-phase features are fused by using a multi-layer perception machine to extract the double-phase image multi-level difference features; Multi-level feature fusion is performed on the double-phase image multi-level difference features to construct a farmland change feature map, and a preset deep supervision strategy is used to obtain a change detection result of the farmland change feature map, including: The double-phase image multi-level difference features are decoded and restored level by level to obtain a plurality of farmland change feature maps, wherein the size of each farmland change feature map gradually decreases as the network deepens; A cross-entropy loss function is used to determine that the plurality of farmland change feature maps are classified pixel by pixel by a classifier to obtain the change detection result.

2. The multi-scale feature fusion cultivated land change detection method according to claim 1, characterized in that, Collecting double-phase optical remote sensing images of the farmland to be detected, and performing multi-scale feature extraction on the double-phase optical remote sensing images to obtain double-phase image multi-scale multi-level features, including: Position encoding is performed on the double-phase optical remote sensing images to obtain relative position information of each part of the double-phase optical remote sensing images; Based on the encoding mode and size information of the double-phase optical remote sensing images, different sizes and different numbers of hierarchical windows corresponding to the double-phase optical remote sensing images are determined; The encoding mode and size information of the double-phase optical remote sensing images are input into a twin network, and multi-head attention calculation is performed by using the different sizes and different numbers of hierarchical windows to obtain double-phase image multi-scale features; The multi-scale features of the dual-time-phase images are fused by using a multi-layer perception to obtain multi-scale multi-level features of the dual-time-phase images.

3. The multi-scale feature fusion cultivated land change detection method according to claim 2, characterized in that, The dual-time-phase optical remote sensing images are positionally encoded to obtain relative position information of each part of the dual-time-phase optical remote sensing images, including: The dual-time-phase optical remote sensing images are cropped into a plurality of non-overlapping patch blocks, and the plurality of non-overlapping patch blocks are taken as tokens. The token dimension is obtained, and a plurality of odd position elements and a plurality of even position elements in the token dimension are determined according to the position to which the token belongs. The relative position information of the plurality of non-overlapping patch blocks is obtained based on a sine function of the position to which the token belongs and the plurality of even position elements and a cosine function of the position to which the token belongs and the plurality of odd position elements, respectively.

4. The multi-scale feature fusion cropland change detection method according to claim 2, characterized in that, Based on the encoding mode and size information of the dual-time-phase optical remote sensing images, different sizes and different numbers of hierarchical windows corresponding to the dual-time-phase optical remote sensing images are determined, including: The level and input size of the dual-time-phase optical remote sensing images are obtained by using the original size information and network structure of the dual-time-phase optical remote sensing images. It is determined that a preset large window is used to extract the global structured features of the dual-time-phase optical remote sensing images, and a preset small window is used to extract the local fine features of the dual-time-phase optical remote sensing images, and the number of hierarchical windows is determined according to the feature map size of the dual-time-phase optical remote sensing images. It is determined that the number of hierarchical window types of the preset shallow features is greater than the number of hierarchical window types of the preset deep features.

5. The multi-scale feature fusion cropland change detection method according to claim 2, characterized in that, The encoding mode and size information of the dual-time-phase optical remote sensing images are input into a twin network, and multi-head attention calculation is performed by using the different sizes and different numbers of hierarchical windows to obtain multi-scale features of the dual-time-phase images, including: The dual-time-phase optical remote sensing images are subjected to dimension-preserving linear transformation respectively to obtain a matrix Q, a matrix K and a matrix V. The matrix Q is multiplied by the transpose matrix K and divided by the square root of the element dimension of the matrix Q, and then normalized by using a softmax function to obtain an attention matrix Attn. The matrix V is multiplied by the attention matrix Attn to obtain a correlation calculation matrix, a plurality of correlation calculation matrices corresponding to self-attention are spliced, and the spliced matrix is converted by a conversion matrix to obtain an output feature matrix. A sliding window is used in the different sizes and different numbers of hierarchical windows to obtain a plurality of output feature matrices. The plurality of output feature matrices are aggregated and connected to obtain the multi-scale features of the dual-time-phase images.

6. A multi-scale feature fusion cultivated land change detection system based on the multi-scale feature fusion cultivated land change detection method of any one of claims 1 to 5, characterized in that, It includes: A multi-scale feature extraction module is configured to acquire dual-time-phase optical remote sensing images of the cultivated land to be detected, and extract multi-scale features of the dual-time-phase optical remote sensing images to obtain multi-scale multi-level features of the dual-time-phase images. A difference feature extraction module is configured to extract multi-level difference features of the dual-time-phase images based on a priori guided change feature and invariant feature enhancement method. A fusion detection module is configured to perform multi-level feature fusion on the multi-level difference features of the dual-time-phase images to construct a cultivated land change feature map, and obtain a change detection result of the cultivated land change feature map by using a preset deep supervision strategy.

7. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor implements the multi-scale feature fusion cultivated land change detection method according to any one of claims 1-5 when executing the program.