A method and device for detecting land change types based on attention mechanism
The integration of an attention-based neural network with U-Net architecture for SAR and optical data processing addresses inefficiencies in land cover change detection, providing rapid and accurate results with high adaptability and precision.
Patent Information
- Application Number
- CN202211542600.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-02
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-12-02
AI Technical Summary
In the prior art, the land change type detection method is highly complex and has low applicability, making it difficult to adapt to efficient land management mode, especially in cloudy and rainy areas, which is difficult to obtain optical remote sensing images, which affects the change detection effect.
A neural network based on attention mechanism is adopted, combined with bi-time phase SAR data and optical data for preprocessing and training, and a U-Net network and attention module are used to fuse feature maps to achieve fast and accurate land change type detection.
It realizes fast and accurate land change type detection, with high applicability and no complex equations required, and is suitable for large-area change detection.
Smart Images

Figure CN116152669B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of land change type detection, and particularly relates to a method and device for land change type detection based on an attention mechanism. Background Art
[0002] Land change type detection usually involves quantitatively comparing, analyzing, and determining the surface change characteristics of remote sensing images of the same area at different times through change detection algorithms, that is, the study of the change from one type of ground object to another. With the rapid growth of the world's population, humans are currently facing serious problems such as severe forest degradation, desertification, soil erosion, land yield reduction, and biodiversity loss. The necessity of land cover change detection has become increasingly prominent. Remote sensing technology has the characteristics of being real-time, fast, wide coverage, and periodic. However, it is difficult to obtain high-quality images of optical remote sensing images in cloudy and rainy areas, which has a great impact on change detection. Synthetic Aperture Radar (SAR) data that is not affected by cloud and rain weather has become an important alternative data.
[0003] The current method for land change detection using SAR data is to draw change vectors and supplement them with on-site investigations. This method is interfered by human subjective factors and has low efficiency, which is not conducive to obtaining large-area change detection results and is difficult to adapt to an efficient land management mode. Some threshold segmentation, classification, and other algorithms have various defects when performing land change detection. For example, only initial features are used for calculation, high-dimensional and abstract features are not expressed, and features need to be manually defined and calculated. Iteration is performed through complex target equations, and then the change type is judged. The applicability of its algorithm is poor.
[0004] Therefore, how to provide a land change type detection method that is fast, accurate, and has high applicability is a technical problem to be solved by those skilled in the art. Summary of the Invention
[0005] The purpose of the present invention is to solve the technical problems of high complexity and low applicability in land change type detection in the prior art.
[0006] To achieve the above technical purpose, on the one hand, the present invention provides a method for land change type detection based on an attention mechanism, the method comprising:
[0007] Obtain training dual-temporal SAR data and training dual-temporal optical data of a training area, and the training dual-temporal SAR data and the training dual-temporal optical data are in one-to-one correspondence;
[0008] Perform the first preprocessing on the training dual-temporal SAR data to obtain the first training data, and perform the second preprocessing on the training dual-temporal optical data to obtain the second training data;
[0009] Input the first training data and the second training data into a preset neural network for training, where the preset neural network is a neural network established based on an attention module and a U-Net network;
[0010] Obtain the dual-temporal SAR data and the dual-temporal optical data of the area to be detected, perform the first preprocessing on the dual-temporal SAR data to obtain the first data, and perform the third preprocessing on the dual-temporal optical data to obtain the second data;
[0011] Input the first data and the second data into the trained preset neural network to obtain the land change type recognition and detection result, where the training area and the area to be detected are areas with the same topographic and geomorphic features.
[0012] Further, the land change types specifically include unchanged type, cultivated land to orchard type, cultivated land to road type, bare land to construction land type, and other change types.
[0013] Further, the preset neural network is specifically: insert the attention module after the convolution and before the pooling of each layer of the encoder in the U-Net network, and the U-Net network after the insertion is the preset neural network. The preset neural network includes an encoder and a decoder. The encoder is used for feature extraction. The decoder includes two types of upsampling, deconvolution and bilinear interpolation, and fuses the feature map passing through the attention module and the upsampled feature map through skip connections.
[0014] Further, both the dual-temporal SAR data and the training dual-temporal SAR data are dual-temporal fully polarized SAR data. The first preprocessing specifically includes:
[0015] Extract the SAR feature image according to the polarization decomposition method of the input data, where the input data is specifically the training dual-temporal SAR data or the dual-temporal SAR data;
[0016] Perform logarithmic ratio on the SAR feature image to obtain the feature difference image;
[0017] Standardize the feature difference image;
[0018] Perform cropping and resampling on the standardized feature difference image in sequence.
[0019] Further, the second preprocessing specifically includes:
[0020] Radiometric calibration, atmospheric correction, and registration are sequentially performed on the training dual-temporal optical data to obtain first-level training data;
[0021] Corresponding vector label data is produced according to the first-level training data;
[0022] Attributes are assigned to each land change type in the vector label data, and the attributes include unchanged attribute, cultivated land changed to orchard attribute, cultivated land changed to road attribute, bare land changed to construction land attribute, and other change attributes;
[0023] The vector label data after attribute assignment is rasterized to obtain a corresponding single-channel label image;
[0024] The single-channel label image is converted into a multi-channel RGB label image;
[0025] The RGB label image is sequentially cropped and resampled.
[0026] Further, the third preprocessing specifically includes:
[0027] Radiometric calibration, atmospheric correction, and registration are sequentially performed on the dual-temporal optical data to obtain first-level data;
[0028] The first-level data is sequentially cropped and resampled.
[0029] Further, inputting the first training data and the second training data into a preset neural network for training specifically includes:
[0030] Data augmentation processing is performed on the first training data and the second training data to obtain corresponding first augmented training data and second augmented training data;
[0031] The first augmented training data and the second augmented training data are allocated according to the ratio of 8:1:1 among training, testing, and validation;
[0032] The allocated first augmented training data and second augmented training data are input into the preset neural network for training.
[0033] On the other hand, the present invention also provides a land change type detection device based on an attention mechanism, and the device includes:
[0034] A first acquisition module, configured to acquire training dual-temporal SAR data and training dual-temporal optical data of a training area, and the training dual-temporal SAR data and the training dual-temporal optical data are in one-to-one correspondence;
[0035] A preprocessing module, configured to perform first preprocessing on the training dual-temporal SAR data to obtain first training data, and perform second preprocessing on the training dual-temporal optical data to obtain second training data;
[0036] A training module, configured to input the first training data and the second training data into a preset neural network for training, wherein the preset neural network is a neural network established based on an attention module and a U-Net network;
[0037] A second acquisition module, configured to acquire dual-temporal SAR data and dual-temporal optical data of a region to be detected, perform the first preprocessing on the dual-temporal SAR data to obtain first data, and perform third preprocessing on the dual-temporal optical data to obtain second data;
[0038] An identification and detection module, configured to input the first data and the second data into the trained preset neural network to obtain a land change type identification and detection result, wherein the training region and the region to be detected are regions with the same topographic and geomorphic features.
[0039] Further, the preset neural network is specifically: inserting the attention module respectively after the convolution and before the pooling of each layer of the encoder in the U-Net network, and the U-Net network after the insertion is the preset neural network. The preset neural network includes an encoder and a decoder. The encoder is used for feature extraction. The decoder includes two types of upsampling, namely transposed convolution and bilinear interpolation. The feature map passing through the attention module and the feature map passing through upsampling are fused through skip connections.
[0040] Compared with the prior art, the present invention has the following beneficial effects:
[0041] The present invention obtains training dual-temporal SAR data and training dual-temporal optical data of a training area, and the training dual-temporal SAR data and the training dual-temporal optical data are in one-to-one correspondence; performs first preprocessing on the training dual-temporal SAR data to obtain first training data, and performs second preprocessing on the training dual-temporal optical data to obtain second training data; inputs the first training data and the second training data into a preset neural network for training, where the preset neural network is a neural network established based on an attention module and a U-Net network; obtains dual-temporal SAR data and dual-temporal optical data of a detection area to be detected, performs the first preprocessing on the dual-temporal SAR data to obtain first data, and performs third preprocessing on the dual-temporal optical data to obtain second data; inputs the first data and the second data into the trained preset neural network, so as to obtain a land change type recognition and detection result, where the training area and the detection area to be detected are areas with the same topographic and geomorphic features, realizing fast and accurate detection of land change types, and without complex equations, and having high applicability. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present specification or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments described in this specification. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0043] Figure 1 The figure shows a schematic flowchart of a method for detecting land change types based on an attention mechanism provided by an embodiment of the present invention;
[0044] Figure 2 The figure shows a schematic structural diagram of a device for detecting land change types based on an attention mechanism provided by an embodiment of the present invention;
[0045] Figure 3 The figure shows a schematic structural diagram of a preset neural network provided by an embodiment of the present invention;
[0046] Figure 4 The figure shows a hardware structure block diagram of a server for detecting land change types based on an attention mechanism provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] To enable those of ordinary skill in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0048] As Figure 1 shown is a schematic flow chart of the land change type detection method based on the attention mechanism provided in the embodiments of this specification. Although this specification provides the method operation steps or device structures shown in the following embodiments or drawings, based on routine or without creative efforts, more or fewer operation steps or module units may be included in the method or device. In steps or structures where there is no necessary causal relationship logically, the execution order of these steps or the module structure of the device is not limited to the execution order or module structure shown in the embodiments or drawings of this specification. When the described method or module structure is applied to actual devices, servers or terminal products, it can be executed sequentially or in parallel according to the method or module structure shown in the embodiments or drawings (for example, in an environment of parallel processors or multi-threaded processing, even including an implementation environment of distributed processing and server clusters).
[0049] The land change type detection method based on the attention mechanism provided in the embodiments of this specification can be applied to terminal devices such as clients and servers. As Figure 1 shown, the method specifically includes the following steps:
[0050] Step S101: Obtain the training dual-temporal SAR data and the training dual-temporal optical data of the training area, and the training dual-temporal SAR data and the training dual-temporal optical data are in one-to-one correspondence;
[0051] Specifically, both the dual-temporal SAR data and the training dual-temporal SAR data are dual-temporal full-polarization SAR data, which more completely records the echo scattering information of the area to be detected. The classification result of the dual-temporal full-polarization SAR data can not only provide auxiliary information for further analysis or interpretation such as target detection and edge extraction, but also be used as the final result. This is to better identify and detect the land change type of the area to be detected. The dual-temporal optical data is GF-2 data, and the high-resolution data of the dual-temporal can clearly reflect the change type and provide data support for label making.
[0052] Step S101: Perform a first preprocessing on the training dual-temporal SAR data to obtain the first training data, and perform a second preprocessing on the training dual-temporal optical data to obtain the second training data.
[0053] Specifically, the area to be identified and detected is usually large, and the amount of data is also large. Conventional methods are cumbersome and complex. Therefore, a part of the training area is extracted first, and the data of the training area is used to train a preset neural network, so that the trained preset neural network can identify and detect the land change types in the area to be identified and detected. The land change types specifically include unchanged type, cultivated land to orchard type, cultivated land to road type, bare land to construction land type, and other change types.
[0054] In the embodiment of the present application, the first preprocessing specifically includes:
[0055] Extract the SAR feature image according to the polarization decomposition method of the input data. The input data is specifically training dual-temporal SAR data or dual-temporal SAR data;
[0056] Perform logarithmic ratio on the SAR feature image to obtain a feature difference image;
[0057] Standardize the feature difference image;
[0058] Successively crop and resample the standardized feature difference image.
[0059] Specifically, after obtaining the dual-temporal fully polarized SAR data of the study area, it is necessary to extract the required SAR feature image according to different polarization decomposition methods. The polarization methods include backscattering, Pauli polarization decomposition, and Freeman-Durden polarization decomposition. For the extraction of the backscattering feature image, it is necessary to perform radiometric calibration, multi-look processing, terrain correction, speckle filtering, and decibel value conversion operations on the SAR image in sequence. For the extraction of the Pauli polarization decomposition and Freeman-Durden polarization decomposition feature images, it is necessary to perform radiometric calibration, multi-look processing, terrain correction, polarization decomposition, and polarization filtering operations on the SAR image in sequence.
[0060] Since the land change type detection is carried out on remote sensing images of the same area at different times, after extracting the SAR feature image, it is necessary to use the logarithmic ratio method to process the SAR feature images of the front and back times to obtain the feature difference image.
[0061] Because the pixel values on the feature difference image have no dimension, in order to avoid interference with model training caused by chaotic input data, it is necessary to standardize the feature difference image, that is, standardize all pixels to 0-255. In the specific application process, according to the model, the difference feature image can be cropped into an image of 512×512 and resampled to 9.5 meters.
[0062] In the embodiments of the present application, the second preprocessing specifically includes:
[0063] Performing radiometric calibration, atmospheric correction, and registration on the training dual-temporal optical data in sequence to obtain first-level training data;
[0064] Producing corresponding vector label data according to the first-level training data;
[0065] Assigning attributes to each land change type in the vector label data, where the attributes include unchanged attribute, cultivated land to orchard attribute, cultivated land to road attribute, bare land to construction land attribute, and other change attributes;
[0066] Rasterizing the vector label data with assigned attributes to obtain a corresponding single-channel label image;
[0067] Converting the single-channel label image into a multi-channel RGB label image;
[0068] Performing cropping and resampling on the RGB label image in sequence.
[0069] Specifically, in deep learning, label data is used to identify the original data and add one or more meaningful information labels to provide context, so that the deep learning model can learn from it. For example, the label can indicate which type of ground object the SAR image contains and the specific change of the ground object type.
[0070] In order to enable the model to identify the change types, the land change types in the study area are classified, and each class is pixel-labeled with different numerical values. At the same time, in order to more intuitively observe different land change types, the single-channel raster image with pixel values is converted into a multi-channel RGB image, and different land change types can be labeled with different colors.
[0071] In order to correspond the label data set and the image data set, finally, the label map also needs to be cropped into an image of 512×512 and resampled to 9.5 meters, and the second training data can be obtained in this step.
[0072] In the embodiments of the present application, the finally identified land change types specifically include unchanged type, cultivated land to orchard type, cultivated land to road type, cultivated land to construction land type, and other change types.
[0073] Step S102: Input the first training data and the second training data into a preset neural network for training, where the preset neural network is a neural network established based on an attention module and a U-Net network.
[0074] In the embodiments of the present application, the inputting of the first training data and the second training data into a preset neural network for training specifically includes:
[0075] Performing data augmentation processing on the first training data and the second training data to obtain corresponding first augmented training data and second augmented training data;
[0076] Allocating the first augmented training data and the second augmented training data according to the ratio of 8:1:1 among training, testing, and validation respectively;
[0077] Inputting the allocated first augmented training data and second augmented training data into the preset neural network for training.
[0078] Specifically, during training, the number of training data input into the network each time can be set to 2, and the optimal solution is searched iteratively according to the gradient descent algorithm. The number of iterations is set to 400, that is, the iteration ends after 400 times. In order to search for the target optimal solution, a loss function is introduced for evaluation. What the loss function expresses is the difference between the test value and the true value of the model. The present invention uses the cross-entropy loss function. The cross-entropy loss function can continue training when the gradient is very small, and at the same time can accelerate the convergence of the preset neural network. In order to improve the training speed of the preset neural network, the training parameters are continuously updated through an optimizer. In this application, the Adam optimizer is used, and its learning rate is set to 0.01.
[0079] In the embodiments of the present application, the preset neural network is specifically: inserting the attention module after the convolution and before the pooling in each layer of the encoder of the U-Net network. The U-Net network after the insertion is the preset neural network. The preset neural network includes an encoder and a decoder. The encoder is used for feature extraction. The decoder includes two types of upsampling, transposed convolution and bilinear interpolation. The feature map passing through the attention module and the feature map passing through upsampling are fused through skip connections.
[0080] Specifically, the U-Net network is a decoder-encoder structure. The encoding part is used for feature extraction, and the decoding part is used for feature fusion. At the same time, the skip connections of U-Net combine the deep high-level features from the decoder with the shallow low-level features from the encoder. The attention module is the Coordinate Attention module. Coordinate Attention embeds position information into channel attention, enabling the mobile network to obtain information in a larger area while avoiding introducing a large overhead. Specifically, global pooling operations are performed on the input image in the x and y spatial directions respectively to generate two separate feature images along the vertical and horizontal directions. These two feature maps aggregate the features in the two spatial directions, allowing for capturing long-range dependencies along one spatial direction while retaining precise position information along the other spatial direction. Then, both feature maps are applied to the input image through multiplication to emphasize the attention area. Coordinate Attention can obtain the area that needs to be enhanced by multiplying the position weights, channel weights, and the input image without changing the size of the input image.
[0081] The present invention inserts the Coordinate Attention module, that is, the attention module, into the encoding part of each layer of U-Net. Specifically, it is before pooling and after the convolutional operation of each layer, while retaining the decoding and skip connections of U-Net. The U-Net network after completing these operations is the preset neural network, and the structure of the preset neural network can be as Figure 3 shown.
[0082] In the left encoding part of the preset neural network, starting from the input, each layer first performs 2 times of 3×3 convolution on the picture, that is, the training data or the first data. The convolution operation can extract the rich features contained in the feature image. After the convolution operation, the image size remains unchanged, but the number of channels changes. The more channels there are, the more image feature information can be extracted. It is not difficult to see that the number of channels in the 1st - 4th layers doubles, and the number of channels in the last layer remains unchanged. For example, the number of channels in each layer is 64, 128, 256, 512, 512 respectively. In order to enable the model to learn key features, obtain more detailed information, and at the same time accelerate the model calculation, the present invention inserts the CoordinateAttention attention module after each layer of convolution. The input and output sizes of the feature map passing through the attention module are the same, so that the convolution in the lower layer can learn features with spatial and channel attention. After convolution and the attention module, each layer also needs to perform 1 time of 2×2 max - pooling. The pooling will reduce the image size by 2 times. This model performs pooling 4 times in total. For example, if the initial picture is 512×512, then after 4 times of pooling, it will become four feature images with different sizes: 256×256, 128×128, 64×64, 32×32. It can be seen that for the image passing through the encoding part, both the number of channels and the image size change from shallow to deep, which is beneficial to extracting image features at multiple scales.
[0083] In the right decoding part of the preset neural network, from bottom to top, perform 3×3 de - convolution and bilinear interpolation on the 32×32 feature map output by the last layer of encoding. Both de - convolution and bilinear interpolation can restore the size of the feature map to obtain a 64×64 feature map. This feature map is fused with the 64×64 feature map with attention in the right encoding part of the same layer through skip connection to merge features, and then perform 2 times of 3×3 convolution on the connected feature map to extract features and restore the number of channels again. Similarly, in order to pass to the upper layer, after each layer of convolution, 3×3 de - convolution and bilinear interpolation are required to obtain a 128×128 feature map, and then skip connection and convolution operations are performed again. After a total of four de - convolution and bilinear interpolation operations, a result image with a size of 512×512, which is the same as the input image size, can be obtained.
[0084] During the encoding and decoding processes, in order to prevent the preset neural network from falling into over - fitting and to avoid the phenomenon of gradient disappearance that may occur during the process of backpropagation to update parameter weights, a BN layer and a ReLU activation function are added between each network layer.
[0085] Common evaluation metrics for deep learning network models include Accuracy, Precision, Recall, F1, IoU, etc. The present invention uses Accuracy, Precision, and Recall to quantitatively evaluate the change type detection results of the model. The formula expressions are as follows:
[0086] Accuracy = (TP+TN) / (TP+TN+FP+FN)
[0087] Precision = TP / (TP + FP)
[0088] Recall = TP / (TP + FN)
[0089] Among them, TP means the true category is positive and the network identifies it as positive; TN means the true category is positive and the network identifies it as negative; FP means the true category is negative and the network identifies it as positive; FN means the true category is positive and the network identifies it as negative;
[0090] After model training, finally, the overall accuracy, precision, and recall rate of the network model of the present invention all reached 97%. The precision rates of the 5 change types are 98.75%, 82.84%, 70.63%, 73.20%, and 81.30% respectively, and the recall rates are 98.07%, 87.49%, 72.82%, 84.03%, and 86.42% respectively.
[0091] Step S103: Obtain dual-temporal SAR data and dual-temporal optical data of the area to be detected, perform the first preprocessing on the dual-temporal SAR data to obtain first data, and perform third preprocessing on the dual-temporal optical data to obtain second data.
[0092] In the embodiment of the present application, the third preprocessing specifically includes:
[0093] Performing radiometric calibration, atmospheric correction, and registration on the dual-temporal optical data in sequence to obtain first-level data;
[0094] Performing cropping and resampling on the first-level data in sequence.
[0095] Step S104: Input the first data and the second data into the trained preset neural network to obtain the land change type recognition and detection result, where the training area and the area to be detected are areas with the same topographic and geomorphic features.
[0096] Specifically, through the above-trained preset neural network, the land change type recognition and detection of the area to be detected can be effectively performed. It should be noted that the training area and the area to be detected are areas with the same topographic and geomorphic features, such as both basin areas or both plain areas, etc. In the scheme expansion, for the land change type recognition and detection of different topographic and geomorphic features, a training area with the same topographic and geomorphic features can be used. After training is completed, the recognition and detection can be carried out.
[0097] Based on the above-mentioned land change type method based on the attention mechanism, one or more embodiments of this specification also provide a platform and a terminal for land change detection. The platform or terminal may include devices, software, modules, plugins, servers, clients, etc. that use the method described in the embodiments of this specification, and devices that combine necessary implementation hardware. Based on the same innovative concept, the systems in one or more embodiments provided by the embodiments of this specification are as described in the following embodiments. Since the implementation schemes for the systems to solve problems are similar to the methods, the implementation of the specific systems in the embodiments of this specification may refer to the implementation of the foregoing methods, and the repeated parts will not be elaborated. The term "unit" or "module" used hereinafter may be a combination of software and / or hardware that can achieve a predetermined function. Although the systems described in the following embodiments are preferably implemented in software, hardware and software combined implementations are also possible and contemplated.
[0098] Specifically, Figure 2 is a schematic diagram of the module structure of an embodiment of the land change type detection device based on the attention mechanism provided in this specification. As Figure 2 shown, the land change type detection device based on the attention mechanism provided in this specification includes:
[0099] A first acquisition module 201, configured to acquire training dual-temporal SAR data and training dual-temporal optical data of a training area, and the training dual-temporal SAR data and the training dual-temporal optical data are in one-to-one correspondence;
[0100] A preprocessing module 202, configured to perform first preprocessing on the training dual-temporal SAR data to obtain first training data, and perform second preprocessing on the training dual-temporal optical data to obtain second training data;
[0101] A training module 203, configured to input the first training data and the second training data into a preset neural network for training, where the preset neural network is a neural network established based on an attention module and a U-Net network;
[0102] A second acquisition module 204, configured to acquire dual-temporal SAR data and dual-temporal optical data of a region to be detected, perform the first preprocessing on the dual-temporal SAR data to obtain first data, and perform third preprocessing on the dual-temporal optical data to obtain second data;
[0103] An identification and detection module 205, configured to input the first data and the second data into the trained preset neural network to obtain a land change type identification and detection result, where the training area and the region to be detected are regions with the same topographic and geomorphic features.
[0104] It should be noted that the above system may also include other implementation manners according to the description of the corresponding method embodiments. The specific implementation manners may refer to the description of the corresponding method embodiments above, and will not be elaborated here one by one.
[0105] An embodiment of the present application further provides an electronic device, including:
[0106] A processor;
[0107] A memory for storing executable instructions of the processor;
[0108] The processor is configured to execute the method provided in the above embodiment.
[0109] For the electronic device provided in the embodiment of the present application, by storing the executable instructions of the processor in the memory, when the processor executes the executable instructions, it can first obtain the training dual-temporal SAR data and training dual-temporal optical data of the training area, and the training dual-temporal SAR data and training dual-temporal optical data are in one-to-one correspondence; perform a first preprocessing on the training dual-temporal SAR data to obtain first training data, and perform a second preprocessing on the training dual-temporal optical data to obtain second training data; input the first training data and the second training data into a preset neural network for training, where the preset neural network is a neural network established based on an attention module and a U-Net network; obtain the dual-temporal SAR data and dual-temporal optical data of the area to be detected, and perform the first preprocessing on the dual-temporal SAR data to obtain first data, and perform a third preprocessing on the dual-temporal optical data to obtain second data; input the first data and the second data into the trained preset neural network, so as to obtain the land change type recognition and detection result, where the training area and the area to be detected are areas with the same topographic and geomorphic features, realizing fast and accurate detection of land change types, and without complex equations, and having high applicability.
[0110] The method embodiments provided in the embodiments of this specification can be executed on a mobile terminal, a computer terminal, a server, or a similar computing device. Taking running on a server as an example, Figure 4 is a hardware structure block diagram of a land change type detection server based on an attention mechanism in an embodiment of this specification. The computer terminal may be the land change type detection server or the land change type detection device based on the attention mechanism in the above embodiment. It may include one or more (only one is shown in the figure) processors 100 (the processor 100 may include, but is not limited to, a processing device such as a microprocessor mcu or a field programmable gate array fpga), a non-volatile memory 200 for storing data, and a transmission module 300 for communication functions.
[0111] The non-volatile memory 200 can be used to store software programs and modules of application software, such as the program instructions / modules corresponding to the data security access method in the embodiments of this specification. The processor 100 executes various functional applications and resource data updates by running the software programs and modules stored in the non-volatile memory 200. The non-volatile memory 200 can include a high-speed random access memory, and can also include non-volatile memories, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the non-volatile memory 200 can further include memories remotely set relative to the processor 100, and these remote memories can be connected to the computer terminal through a network. Examples of the above network include but are not limited to the Internet, intranet, local area network, mobile communication network, and their combinations.
[0112] The transmission module 300 is used to receive or send data via a network. Specific examples of the above network can include the wireless network provided by the communication provider of the computer terminal. In one instance, the transmission module 300 includes a network interface controller (nic), which can be connected to other network devices through a base station so as to communicate with the Internet. In one instance, the transmission module 300 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0113] The above specific embodiments of this specification have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be executed in a different order from that in the embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0114] The method or device described in the above embodiments provided in this specification can implement the business logic through a computer program and record it on a storage medium, and the storage medium can be read and executed by a computer to achieve the effects of the solutions described in the embodiments of this specification, such as:
[0115] Obtain the training dual-temporal SAR data and training dual-temporal optical data of the training area, and the training dual-temporal SAR data and training dual-temporal optical data are in one-to-one correspondence;
[0116] Perform first preprocessing on the training dual-temporal SAR data to obtain first training data, and perform second preprocessing on the training dual-temporal optical data to obtain second training data;
[0117] Input the first training data and the second training data into a preset neural network for training, where the preset neural network is a neural network established based on an attention module and a U-Net network;
[0118] Obtain dual-temporal SAR data and dual-temporal optical data of the area to be detected, perform the first preprocessing on the dual-temporal SAR data to obtain first data, and perform third preprocessing on the dual-temporal optical data to obtain second data;
[0119] Input the first data and the second data into the trained preset neural network to obtain a land change type recognition and detection result, where the training area and the area to be detected are areas with the same topographic and geomorphic features.
[0120] The storage medium may include a physical device for storing information, usually by digitizing the information and then storing it in a medium using electrical, magnetic, or optical means. The storage medium may include: devices for storing information using electrical energy, such as various memories, such as RAM, ROM, etc.; devices for storing information using magnetic energy, such as hard disks, floppy disks, magnetic tapes, magnetic core memories, bubble memories, USB flash drives; devices for storing information using optical means, such as CDs or DVDs. Of course, there are also other ways of readable storage media, such as quantum memories, graphene memories, etc.
[0121] The embodiments of this specification are not limited to those that must conform to industry communication standards, standard computer resource data update and data storage rules, or the situations described in one or more embodiments of this specification. Some industry standards or implementation schemes slightly modified on the basis of the implementation described by a custom method or embodiment can also achieve the same, equivalent, or similar, or predictable implementation effects after deformation as the above embodiments. The embodiments obtained by applying these modified or deformed data acquisition, storage, judgment, processing methods, etc. still fall within the scope of the optional implementation schemes of the embodiments of this specification.
[0122] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to implement the same function by logically programming method steps so that the controller is in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be regarded as a hardware component, and the devices included therein for implementing various functions can also be regarded as structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or structures within the hardware component.
[0123] The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or plugins can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0124] These computer program instructions can also be loaded onto a computer or other programmable resource data update device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 a process or multiple processes and / or blocks Figure 1 steps for implementing the functions specified in one block or multiple blocks.
[0125] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the description of the method embodiments. In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this specification. In this specification, the schematic expression of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0126] Those of ordinary skill in the art will realize that the embodiments described herein are for helping the reader understand the principles of the present invention, and it should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations that do not deviate from the essence of the present invention based on the technical revelations disclosed in the present invention, and these deformations and combinations are still within the protection scope of the present invention.
Claims
1. A method for detecting land change types based on an attention mechanism, characterized in that The method includes: Obtaining training dual-temporal SAR data and training dual-temporal optical data of a training area, where the training dual-temporal SAR data and the training dual-temporal optical data are in one-to-one correspondence; Performing first preprocessing on the training dual-temporal SAR data to obtain first training data, and performing second preprocessing on the training dual-temporal optical data to obtain second training data; Both the dual-temporal SAR data and the training dual-temporal SAR data are dual-temporal full-polarization SAR data. The first preprocessing specifically includes: Extracting SAR feature images according to the polarization decomposition method of the input data, where the input data is specifically the training dual-temporal SAR data or the dual-temporal SAR data; Taking the logarithm ratio of the SAR feature images to obtain feature difference images; Normalizing the feature difference images; Sequentially cropping and resampling the normalized feature difference images; The second preprocessing specifically includes: Sequentially performing radiometric calibration, atmospheric correction, and registration on the training dual-temporal optical data to obtain first-level training data; Making corresponding vector label data according to the first-level training data; Assigning attributes to each land change type in the vector label data, where the attributes include unchanged attribute, cultivated land to orchard attribute, cultivated land to road attribute, bare land to construction land attribute, and other change attributes; Rasterizing the vector label data with attributes to obtain a corresponding single-channel label image; Converting the single-channel label image into a multi-channel RGB label image; Sequentially cropping and resampling the RGB label image; Inputting the first training data and the second training data into a preset neural network for training, where the preset neural network is a neural network established based on an attention module and a U-Net network; Obtaining dual-temporal SAR data and dual-temporal optical data of a detection area to be detected, performing the first preprocessing on the dual-temporal SAR data to obtain first data, and performing third preprocessing on the dual-temporal optical data to obtain second data; The third preprocessing specifically includes: Sequentially performing radiometric calibration, atmospheric correction, and registration on the dual-temporal optical data to obtain first-level data; Sequentially cropping and resampling the first-level data; Inputting the first data and the second data into the trained preset neural network to obtain a land change type recognition and detection result, where the training area and the detection area to be detected are areas with the same topographic and geomorphic features.
2. The method for detecting land change types based on the attention mechanism according to claim 1, wherein The land change types specifically include unchanged type, cultivated land to orchard type, cultivated land to road type, bare land to construction land type, and other change types.
3. The method for detecting land change types based on the attention mechanism according to claim 1, wherein The preset neural network is specifically: inserting the attention module respectively after the convolution and before the pooling of each layer of the encoder in the U-Net network. After the insertion, the U-Net network is the preset neural network. The preset neural network includes an encoder and a decoder. The encoder is used for feature extraction. The decoder includes two types of upsampling, deconvolution and bilinear interpolation, and fuses the feature maps passing through the attention module and the feature maps passing through the upsampling through skip connections.
4. The method for detecting land change types based on the attention mechanism according to claim 1, characterized in that, Inputting the first training data and the second training data into a preset neural network for training specifically includes: Performing data augmentation processing on the first training data and the second training data to obtain corresponding first augmented training data and second augmented training data; Allocating the first augmented training data and the second augmented training data according to the ratio of 8:1:1 among training, testing, and validation respectively; Inputting the allocated first augmented training data and second augmented training data into the preset neural network for training.
5. An apparatus for detecting land change types based on an attention mechanism, which is used to execute the method for detecting land change types based on an attention mechanism according to any one of claims 1-4, and is characterized in that, The device includes: A first acquisition module, configured to acquire training dual-temporal SAR data and training dual-temporal optical data of a training area, and the training dual-temporal SAR data and the training dual-temporal optical data are in one-to-one correspondence; A preprocessing module, configured to perform first preprocessing on the training dual-temporal SAR data to obtain first training data, and perform second preprocessing on the training dual-temporal optical data to obtain second training data; A training module, configured to input the first training data and the second training data into a preset neural network for training, wherein the preset neural network is a neural network established based on an attention module and a U-Net network; A second acquisition module, configured to acquire dual-temporal SAR data and dual-temporal optical data of a region to be detected, perform the first preprocessing on the dual-temporal SAR data to obtain first data, and perform third preprocessing on the dual-temporal optical data to obtain second data; An identification and detection module, configured to input the first data and the second data into the trained preset neural network to obtain a land change type identification and detection result, wherein the training area and the region to be detected are regions with the same topographic and geomorphic features.
6. The land change type detection device based on the attention mechanism according to claim 5, characterized in that, The preset neural network is specifically: inserting the attention module respectively after the convolution and before the pooling of each layer of the encoder in the U-Net network, and the U-Net network after the insertion is the preset neural network. The preset neural network includes an encoder and a decoder. The encoder is used for feature extraction. The decoder includes two types of upsampling, namely transposed convolution and bilinear interpolation. The feature maps passing through the attention module and the feature maps passing through upsampling are fused through skip connections.
Citation Information
Patent Citations
Method and device for extracting cultivated land parcels based on SE-U-Net + + model
CN114419430A
Method and system for image registration and change detection
US20110222781A1