Crop extraction method and system for coupling time sequence and space attention mechanism
By introducing timing and spatial attention mechanisms into the U-Net deep learning network, the problem of insufficient resource consumption and interpretation in crop remote sensing image extraction is solved, and efficient and accurate crop feature extraction and segmentation is achieved.
Patent Information
- Application Number
- CN202510661470.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-06-20
AI Technical Summary
The prior art requires a lot of hardware and human resources to be consumed in the extraction process of crop remote sensing images, and the deep learning model is insufficient to interpret the deep semantic information of crop remote sensing images.
The U-Net deep learning network model that uses a coupled timing and spatial attention mechanism is used to highlight the characteristics of the target crop by introducing channel attention modules and spatial attention modules, adaptively weighted feature channels and spatial locations.
It improves the efficiency and accuracy of feature extraction, reduces misjudgment, enhances the robustness and generalization ability of the model, and realizes high-precision remote sensing image segmentation of crops.
Smart Images

Figure CN120182836A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of agricultural remote sensing image extraction, and relates to a deep learning extraction method for crops, in particular to a crop extraction method and system that couples temporal and spatial attention mechanisms. Background Art
[0002] At present, the population in China is increasing continuously, and the demand for food is also growing. Timely and accurately obtaining the spatial distribution information of crops in China is of great significance for crop growth monitoring and yield estimation. The earliest acquisition of crop spatial distribution information required on-site inspections and records by management personnel. This process is not only cumbersome and complex, but also easily affected by human factors, resulting in inaccurate crop spatial distribution information. Remote sensing technology has the advantages of wide range, strong timeliness, and low cost, providing a new means for quickly and accurately obtaining crop spatial distribution information.
[0003] In the early stage, the extraction of crops using remote sensing technology was still based on some traditional methods. That is, first, the candidate regions were extracted, then features were designed manually, and finally, the classifier was used to output the category and coordinates of the target. In the process of extracting candidate regions, a large number of sliding windows often need to be set manually, which greatly consumes hardware and human resources. In addition, feature extraction mainly relies on visual information such as the color, contour, and texture of crops, which has great subjectivity. Traditional methods are also difficult to effectively extract crop features and lack robustness when dealing with crop remote sensing images of complex ground objects.
[0004] Entering the 21st century, deep learning has made breakthrough developments. Deep learning can obtain useful information by continuously stacking extremely deep networks and has very strong data mining and analysis capabilities. Therefore, it has attracted much attention from researchers in recent years. Deep learning can not only effectively extract the deep semantic information in images, but also has advantages over traditional algorithms in terms of robustness and application scope, so its development is rapid. However, due to the characteristics of crop remote sensing images themselves, existing deep learning detection methods cannot achieve good performance in crop extraction tasks. Current deep learning models pay more attention to the features of crops themselves and have less grasp of global and context information, resulting in insufficient interpretability of deep semantic information in crop remote sensing images by the models.
[0005] In view of the problems existing in current deep learning models, the present invention proposes a deep learning extraction method for crops that couples temporal and spatial feature attention mechanisms. This model can effectively solve the shortcomings of traditional methods that require a large amount of hardware and human resources and the limitation of existing deep learning models in the insufficient interpretability of deep semantic information in crop remote sensing images. Summary of the Invention
[0006] To solve the above problems, the present invention provides a crop extraction method and system that couples temporal and spatial attention mechanisms, which can effectively extract multi-scale features of target crops at different depths in images, resulting in high-precision recognition results, easy implementation, high efficiency, and strong practicability.
[0007] The technical solution adopted by the present invention is as follows:
[0008] A crop extraction method that couples temporal and spatial attention mechanisms, comprising the following steps:
[0009] Obtain a remote sensing image dataset of target crops and input it into a trained object detection model to output a crop coverage map of the target area; the object detection model is a deep learning network model of U-Net that introduces temporal and spatial attention mechanisms.
[0010] Furthermore, the training process of the object detection model includes:
[0011] Obtain a sample set of remote sensing images of target crops and perform preprocessing;
[0012] Construct a deep learning network model based on U-Net and introduce temporal and spatial attention mechanisms to obtain an object detection model;
[0013] Use the preprocessed sample set of remote sensing images of target crops to train the object detection model.
[0014] Furthermore, the preprocessing of the crop remote sensing image dataset specifically includes: classifying the crop remote sensing images, generating labels, and dividing them into a training set, a test set, and a validation set.
[0015] Furthermore, the deep learning network model based on U-Net includes a downsampling module, an upsampling module, and an output module; the downsampling module includes several sequentially connected downsampling units, each downsampling unit extracts deep features of the image through a convolutional layer, and then reduces the resolution through a pooling layer to output a feature map; the upsampling module includes several sequentially connected upsampling units, each upsampling unit performs upsampling on the feature map, splices it with the feature map of the same resolution in the corresponding downsampling stage, and then refines and adjusts the fused features through convolutional operations and the ReLu activation function.
[0016] Furthermore, the introduction of the temporal and spatial attention mechanisms specifically includes: connecting an attention module after the first downsampling unit and before the last upsampling unit, and the attention module includes a channel attention module and a spatial attention module.
[0017] Furthermore, the data processing process of the channel attention module includes:
[0018] Perform global average pooling and global max pooling on the input feature maps respectively to obtain the feature descriptors after average pooling and max pooling;
[0019] Process the feature descriptors after average pooling and max pooling respectively using a multi-layer perceptron, add the results of the average pooling and max pooling processed by the multi-layer perceptron, and then map the values to between 0 and 1 through a Sigmoid activation function to obtain the channel attention weights;
[0020] Multiply the attention weights with the input feature maps channel by channel to weight-adjust the channel features.
[0021] Furthermore, the data processing process of the spatial attention module includes:
[0022] The spatial attention module takes the feature map output by the channel attention module as input, performs max pooling and average pooling on the feature map respectively to obtain two single-channel feature maps;
[0023] Concatenate the two single-channel feature maps along the channel dimension to obtain a new feature map with 2 channels;
[0024] Process the concatenated feature map through a convolutional layer, and then map the values to between 0 and 1 through a Sigmoid activation function to obtain the spatial attention weight map;
[0025] Multiply the spatial attention weight map with the input feature map element by element to highlight the spatial region where the target crop is located in the feature map.
[0026] Furthermore, the output module converts the multiple channels of the feature map into the number of channels corresponding to the set segmentation categories through a convolutional layer, thereby converting the feature map into the final segmentation map, that is, the crop coverage map of the target area.
[0027] A crop deep learning extraction system coupling temporal and spatial attention mechanisms, comprising:
[0028] Data acquisition module: Acquire the crop remote sensing image dataset and perform preprocessing;
[0029] Model construction module: Construct a deep learning network model based on U-Net and introduce temporal and spatial attention mechanisms to obtain the target detection model;
[0030] Model training module: Train the target detection model using the preprocessed crop remote sensing image dataset.
[0031] A computer device, the computer device includes:
[0032] One or more processors;
[0033] A memory for storing one or more programs;
[0034] When the one or more programs are executed by the one or more processors, the one or more processors implement the crop extraction method that couples the timing and spatial attention mechanisms as described above.
[0035] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0036] (1) Feature extraction optimization. In terms of channel feature screening, the channel attention mechanism of CBAM can adaptively weight the feature channels extracted by U-Net, highlight the unique spectral feature channels of target crops in multi-spectral remote sensing images, suppress interference information channels, and make it focus on the features related to target crops. In terms of spatial position focusing, the spatial attention mechanism can enhance U-Net's perception of the spatial position of target crops, strengthen the edge features of the planting area, outline the contour, and reduce confusion with adjacent ground objects.
[0037] (2) Improved segmentation accuracy. In terms of small target detection, CBAM can guide the network to focus on the features of small targets such as the seedling stage, and improve the detection and segmentation ability of small-area target crop regions through weighting, which has obvious effects on small-area planting areas or experimental fields. In the case of complex backgrounds, CBAM can distinguish the feature differences between target crops and similar ground objects, reduce misjudgments, and improve the segmentation accuracy.
[0038] (3) Enhanced model robustness. In terms of noise resistance, CBAM enables U-Net to robustly extract features under noise, and the channel and spatial attention modules respectively reduce the impact of noise on features and boundaries. Secondly, the present invention can adjust the attention weights according to the changes in the planting patterns and growth stages of target crops, enhancing the robustness and generalization ability of the model. Description of the Drawings
[0039] Figure 1 is the structural diagram of the deep learning crop extraction method that couples the timing and spatial feature attention mechanisms in the embodiments of the present invention.
[0040] Figure 2 is the comparison diagram of the extraction results of the U-Net model combined with CBAM and the random forest model in the embodiments of the present invention.
[0041] Figure 3 is the statistical result diagram of the extraction errors of the U-Net model combined with CBAM and the random forest model in the embodiments of the present invention. Detailed Embodiments
[0042] To make the technical solutions of the present invention clearer, the following further describes the specific embodiments of the present invention with reference to the drawings.
[0043] In view of the demand for the extraction accuracy of target crops in existing models, the present invention proposes a deep learning extraction method for winter wheat that couples temporal and spatial feature attention mechanisms. First, based on the remote sensing images of target crops as the basic dataset, the images are classified and labeled to generate training sets, validation sets, and test sets. Then, a U-Net network is built and a temporal and spatial feature attention mechanism module is introduced. Through global average pooling, global max pooling, and MLP, important channel features are highlighted and unimportant channel features are suppressed. Then, concatenation is performed on the channel dimension and converted into a spatial attention map through a convolutional layer, and element-wise multiplication is performed with the original feature map to highlight the position information of the target crops in space. Finally, end-to-end testing is performed on the model, the model is adjusted, and the construction of the crop extraction model is completed.
[0044] The target in the input data of the U-Net model contains less information itself. Therefore, the network needs to use operations such as convolution and pooling to expand the receptive field and fully learn the information around the target to further mine deep features. General surrounding information is local, while the attention mechanism can help the network establish dependencies under long-distance constraints. The attention module (Convolutional Block Attention Module, CBAM) includes a channel attention module (Channel Attention Module, CAM) and a spatial attention module (Spatial Attention Module, SAM). The introduction of the attention mechanism allows the network to focus on relevant parts when processing input data, enabling the network to automatically learn and selectively focus on important information in the input data.
[0045] The specific steps of a crop extraction method that couples temporal and spatial attention mechanisms are as follows:
[0046] Obtain a remote sensing image dataset of target crops and input it into a trained target detection model to output a crop coverage map of the target area; the target detection model is a deep learning network model of U-Net that introduces temporal and spatial attention mechanisms.
[0047] The training process of the target detection model includes:
[0048] (1) Obtain a remote sensing image dataset of target crops and perform preprocessing.
[0049] The preprocessing specifically includes: after classifying the remote sensing images of target crops and labeling the target crops, generating corresponding labels, and then dividing the image training set, validation set, and test set according to a certain ratio, and trying to ensure the consistency of the number of target categories in each sub-dataset.
[0050] (2)The Sigmoid function is used as the activation function, and the cross-entropy loss function is used as the loss function to construct a deep learning network model based on U-Net. A temporal and spatial attention mechanism is introduced to obtain an object detection model. The input of the object detection model is crop remote sensing image data, and the output is the crop coverage map of the target area.
[0051] The deep learning network model based on U-Net includes a downsampling module, an upsampling module, and an output module. The downsampling module, that is, the encoding module, is responsible for gradually extracting the deep features of the image in the whole network. At the same time, the resolution of the feature map is reduced through pooling operations to increase the receptive field. The upsampling module, that is, the decoding module, is used to gradually restore the deep features extracted during the downsampling process to a resolution similar to that of the original input image, and fuse features at different levels to finally generate a high-quality segmentation image.
[0052] The downsampling module includes several sequentially connected downsampling units. Each downsampling unit extracts the deep features of the image through a convolutional layer, and then reduces the resolution through a pooling layer to output a feature map. The upsampling module includes several sequentially connected upsampling units. After each upsampling unit upsamples the feature map, it is concatenated with the feature map of the same resolution in the corresponding downsampling stage, and then the fused features are refined and adjusted through convolutional operations and the ReLu activation function.
[0053] The introduction of the temporal and spatial attention mechanism specifically includes: an attention module is connected after the first downsampling unit and before the last upsampling unit. The attention module includes a channel attention module and a spatial attention module. After adding the attention module, first the channel attention module plays a role. The channel attention module analyzes different channels of the input feature map to determine which channels contain features more critical for the segmentation task, and then assigns higher weights to these important channels. Then comes the spatial attention module. Based on the feature map adjusted by the channel attention, it calculates the attention in the spatial dimension to determine which spatial regions in the image contain more critical target object information and need to be focused on by the network.
[0054] The data processing process of the channel attention module includes:
[0055] Global average pooling and global max pooling are performed on each channel of the input feature map to compress the feature information of each channel into a scalar, obtaining the feature descriptors after average pooling and max pooling;
[0056] The feature descriptors after average pooling and max pooling are respectively processed using a multi-layer perceptron (MLP), which usually contains a shared hidden layer. The results of average pooling and max pooling processed by the multi-layer perceptron are added together, and then the values are mapped between 0 and 1 through a Sigmoid activation function to obtain the attention weights for each channel;
[0057] The attention weights are multiplied with the input feature map channel by channel, thereby weighted adjustment of the channel features is performed to highlight important channel features and suppress unimportant channel features.
[0058] The specific calculation formula of the channel attention module is as follows:
[0059] Among them, F represents the original input feature, represents the output feature after passing through the channel attention module, MLP represents being processed by the multi-layer perceptron, AvgPool represents being processed by average pooling, MacPool represents being processed by max pooling, represents the feature map in the channel attention module after average pooling, represents the feature map in the channel attention module after max pooling, represents the Sigmoid activation function, represents the weight of the fully connected layer.
[0060] Through the channel attention mechanism of CBAM in the present invention, the feature channels extracted by U-Net are adaptively weighted, highlighting the unique spectral feature channels of the target crop in the multi-spectral remote sensing image, suppressing the interference information channels, and making it focus on the features related to the target crop. It can adjust the attention weights according to the planting pattern and growth stage changes of the target crop, enhancing the robustness and generalization ability of the model.
[0061] The data processing process of the spatial attention module includes:
[0062] For the input feature map adjusted by the channel attention module, max pooling and average pooling are respectively performed on the feature map to obtain two feature maps of H×W×1. The feature map is a three-dimensional tensor with dimensions of H×W×C, where H represents the spatial resolution of the feature map in the vertical direction (y-axis), that is, the number of rows of the feature map, W represents the spatial resolution of the feature map in the horizontal direction (x-axis), that is, the number of columns of the feature map, and C represents the number of channels. Each channel corresponds to the activation response of a specific feature. In the case of a single channel (C = 1), the feature map degenerates into a two-dimensional matrix, called a single-channel feature map;
[0063] The two obtained feature maps of H×W×1 are concatenated along the channel dimension to obtain a new feature map of H×W×2. The new feature map has the same size as the original feature map in the spatial dimension, but the number of channels is 2. One channel stores the maximum value and the other stores the average value;
[0064] The concatenated feature map is processed through a convolutional layer, and then the values are mapped between 0 and 1 through a Sigmoid activation function to obtain a spatial attention weight map;
[0065] The spatial attention weight map is multiplied element-wise with the input feature map adjusted by the channel attention module, so as to highlight the spatial region where the target crop is located in the feature map.
[0066] The specific calculation formula of the spatial attention module is as follows:
[0067] In the formula, F represents the original input feature, represents the output feature after passing through the spatial attention module, MLP represents the processing through a multi-layer perceptron, AvgPool represents the processing through average pooling, and MacPool represents the processing through max pooling, represents the feature map in the spatial attention module after average pooling, represents the feature map in the spatial attention module after max pooling, represents the Sigmoid activation function, represents the 7×7 convolutional calculation.
[0068] The present invention enhances the U-Net's perception of the spatial position of the target crop through the spatial attention mechanism of CBAM, strengthens the edge features of the planting area, outlines the contour, and reduces the confusion with adjacent ground objects.
[0069] The output module converts the multiple channels of the feature map into the number of channels corresponding to the set segmentation categories through a convolutional layer, so as to convert the feature map into the crop coverage map of the final target area.
[0070] (3) Use the preprocessed remote sensing image dataset of the target crop to perform end-to-end training on the target detection model and test the model.
[0071] Next, taking the extraction of winter wheat as an example, the method of the present invention will be further described in detail. As Figure 1 described, the specific steps of the method in this embodiment are as follows:
[0072] Step 1: Select a large-scale remote sensing image dataset of winter wheat.
[0073] This embodiment uses a dataset publicly available on the Internet as a winter wheat remote sensing image dataset. The image data in the dataset is divided into a training set and a test set in a ratio of 4:1 for subsequent network training and testing.
[0074] Step 2: Build a deep learning network target detection model.
[0075] U-Net itself is a classic network structure for image segmentation. It has the process of encoding (downsampling) and decoding (upsampling). It generates the final segmented image by continuously extracting and fusing features. The CBAM attention mechanism consists of two parts: channel attention and spatial attention. Its purpose is to enable the network to adaptively focus on the more important feature information in the image, whether in the feature channel dimension or the spatial position dimension. Integrating CBAM into U-Net means inserting the CBAM module at the appropriate position of U-Net so that it can play a role in feature extraction and fusion, optimize feature representation, and thus improve the effect of image segmentation.
[0076] Step 2 includes:
[0077] Step 2.1: Channel attention module. The channel attention module mainly analyzes different channels of the input feature map, determines which channels contain features that are more critical to the segmentation task, and then assigns higher weights to these important channels.
[0078] Step 2.1.1: For the input feature map, first perform global average pooling and global maximum pooling operations respectively. Global average pooling is to find the average value of the entire feature map in the spatial dimension (that is, the width and height dimensions), compressing the feature information of each channel into a value; global maximum pooling is to take the maximum value of each channel in the spatial dimension, and also convert the feature information of each channel into a value. Through these two pooling methods, the characteristics of each channel can be summarized from different angles.
[0079] Step 2.1.2: After obtaining the two pooled feature descriptors, pass them through a structure containing a multi-layer perceptron (MLP). This MLP generally has a hidden layer, which is used to reduce and increase the dimension of the features, so that the network can learn the complex correlation between channels and mine more valuable channel feature information.
[0080] Step 2.1.3, add the results of average pooling and maximum pooling after MLP processing, and then use a Sigmoid activation function to map the obtained value to between 0 and 1. This value is the attention weight of the corresponding channel.
[0081] Step 2.1.4. Finally, multiply these attention weights with the original input feature map channel by channel, thus completing the weighted adjustment of channel features, highlighting important channel features and suppressing relatively unimportant channel features.
[0082] The specific calculation formula is as follows:
[0083]
[0084] In the formula, represents the Sigmoid activation function, represents the weight of the fully connected layer.
[0085] Step 2.2. Spatial attention module. The spatial attention module focuses on the importance of the feature map in spatial positions, that is, determining which spatial regions in the image contain more crucial information about the target object (winter wheat) and need to be focused on by the network.
[0086] Step 2.2.1. First, operate on the input feature map in the channel dimension, perform max pooling and average pooling on the feature map at each spatial position respectively, and obtain two single-channel feature maps with the same size as the original feature map in the spatial dimension.
[0087] Step 2.2.2. Then, concatenate these two single-channel feature maps along the channel dimension to obtain a new feature map with 2 channels. This new feature map combines the maximum feature and average feature information in spatial positions (one channel stores the maximum value and the other stores the average value).
[0088] Step 2.2.3. Then, process the concatenated feature map through a convolutional layer. This convolutional layer can learn the associations between spatial positions and how to highlight important spatial regions, and then pass through the Sigmoid activation function to obtain the spatial attention weight map, whose values are also between 0 and 1.
[0089] Step 2.2.4. Finally, multiply this spatial attention weight map with the original input feature map element by element, so that the spatial regions where the target object (winter wheat) is located in the feature map are highlighted, helping the network better locate and segment the target.
[0090] The specific calculation formula is as follows:
[0091]
[0092] In the formula, represents the Sigmoid activation function, represents the 7×7 convolutional calculation.
[0093] Step 2.3: U-Net network architecture with CBAM attention mechanism.
[0094] Step 2.3.1: Downsampling (encoding) module. The downsampling module is responsible for gradually extracting the deep features of the image throughout the network. At the same time, it reduces the resolution of the feature map through pooling operations to increase the receptive field. After adding CBAM, there are new changes in the downsampling stage.
[0095] Step 2.3.1.1: The first downsampling stage. After passing through a downsampling unit, that is, two convolutional layers and a pooling layer, the CBAM module follows immediately. The CBAM module in this stage will process the downsampled feature map according to the previous calculation methods of channel attention and spatial attention, adjust the channel feature weights of the feature map, and highlight the spatial region where the target object (winter wheat) is located, so that the subsequent network can focus on more critical feature information for deep feature extraction.
[0096] Step 2.3.1.2: Subsequent downsampling stages. The structure of the subsequent downsampling stages is similar to the first stage. However, as the stage progresses, the number of output channels of the convolutional layers will gradually increase (such as becoming 128, 256, 512 channels in sequence), in order to be able to extract increasingly complex and abstract image features.
[0097] Step 2.3.2: Upsampling (decoding) module. The main task of the upsampling module is to gradually restore the deep features extracted during the downsampling process to a resolution size similar to the original input image, and fuse features at different levels, finally generating a high-quality segmentation image. After adding CBAM, the process of the upsampling stage is as follows.
[0098] Step 2.3.2.1: The first upsampling stage. Concatenate the upsampled feature map with the feature map of the same resolution in the corresponding downsampling stage, and then refine and adjust the fused features through convolutional operations and ReLU activation functions.
[0099] Step 2.3.2.2: Subsequent upsampling stages. The subsequent upsampling stages also have a similar operation process. However, as the stage progresses, the number of output channels of the convolutional layers will gradually decrease (such as gradually changing back to 64 channels from 512 channels, etc.) to gradually restore to the final expected output size and number of channels. At the same time, in each stage, it also has to go through steps such as transposed convolution, feature concatenation, and convolutional processing. Finally, before the last downsampling unit, a CBAM module is added. The CBAM module in this stage will perform channel and spatial attention adjustments on the feature map after the above processing, highlighting the feature parts that are more important for the segmentation task, so that the network can better focus on the features of the target object during the subsequent upsampling and feature fusion processes.
[0100] Step 2.3.3, Output Module. After a series of downsampling, upsampling, feature fusion, and attention adjustment, finally, a 1×1 convolutional layer is used to convert the feature map into the final segmentation map. The role of this 1×1 convolutional layer is to convert the multiple channels of the feature map into the number of channels corresponding to the set segmentation categories according to the number of segmentation categories. Each channel represents a category, and the output value can be understood as the probability value of each pixel belonging to different categories. In this way, the conversion from the feature map to the segmentation result is completed.
[0101] It can be seen that the winter wheat extraction model combining the CBAM attention mechanism and the U-Net network can effectively extract multi-scale features at different depths in winter wheat remote sensing images, further refine the segmentation effect, and improve the extraction accuracy.
[0102] Step 3, Cascade network for end-to-end testing.
[0103] Use the training set in Step 1 to train the object detection model, and then input the winter wheat remote sensing images in the test set into the trained object detection model to evaluate its detection results. Comparative experiments are carried out using the traditional classification method for winter wheat remote sensing image processing and the winter wheat extraction model of the U-Net network combined with CBAM constructed in the present invention to detect various targets in the dataset. Compared with the traditional method, the detection effect is significantly improved.
[0104] To further evaluate the extraction effect of the U-Net model combined with CBAM, in this embodiment, the original images from April to May 2022 are used as a reference. In the areas where there are significant differences in the extraction results between the U-Net model combined with CBAM and the random forest model, six rectangular areas (a - f) with a size of 0.1°×0.1° are selected to magnify and compare the extraction results of the two models. The specific comparison details are as Figure 2 shown.
[0105] In the comparison results of a - f, the U-Net model shows a better classification effect than the random forest model, which is related to the spectral differences in the image dataset input to the model. To ensure the integrity of the images in the entire study area, there are differences in the acquisition dates of the study images, resulting in differences in the spectral characteristics of winter wheat reflected in the images. Therefore, the random forest model is prone to being interfered by error factors during the classification process and misses some classifications. In contrast, the U-Net model combined with CBAM can overcome the errors caused by the differences in image dates and successfully identify most of the winter wheat land classes in the study area.
[0106] Next, accuracy verification is carried out. Statistical data is used to verify the accuracy of the extraction results based on the U-Net model combined with CBAM.
[0107] The extraction results of the U-Net model combined with CBAM in the study area in 2022 were calculated from two aspects: the macro scale and the meso scale. Taking the statistical yearbooks of each province as reference data, the absolute error (AE), absolute percentage error (APE), and correlation coefficient (R 2 ) were used to evaluate the extraction results, and a comparison was made with the statistical situation of the extraction results of the random forest model. The specific results are shown in Table 1 and Figure 3 .
[0108] From the macro scale (Table 1), the average APE of the extraction results of the U-Net model was 7.85%, lower than that of the random forest (9.97%). Higher extraction accuracies were achieved in Region A, Region B, and Region D, with APEs of only 1.90%, 4.54%, and 5.84% respectively. Although the APE in Region C was too large, showing a poor extraction effect, overall, the U-Net model combined with CBAM had better performance in extraction.
[0109] Table 1 Validation accuracy based on statistical data (macro scale)
[0110]
[0111] To more objectively reflect the extraction accuracy of the U-Net model combined with CBAM, in this embodiment, from the meso scale, three indicators, namely the absolute error (AE), absolute percentage error (APE), and correlation coefficient (R 2 ), were used for evaluation, and a comparison was made with the evaluation results of the random forest model, as Figure 3 shown.
[0112] The statistical results show that the extraction accuracy of the U-Net model combined with CBAM is higher than that of the random forest model. From the extraction results of the U-Net model, the AE of all mesoscale regions is less than 100 thousand hectares. In the statistical results of APE, the distribution law of the number of mesoscale regions in different gradient intervals is similar to that of the random forest model. The APE of most mesoscale regions is below 30%. Among them, the number of mesoscale regions with APE in the intervals of 0-10% and 10-20% in the extraction results of the U-Net model combined with CBAM is larger, and the number of mesoscale regions with APE greater than 50% also exceeds 10, which is related to the winter wheat planting area in the mesoscale region. By comparing the experimental data, it is found that the winter wheat planting area of these mesoscale regions with APE greater than 50% is less than 80 thousand hectares, and most mesoscale regions are even below 50 thousand hectares. Therefore, even a small extraction error will cause the APE of these cities to show an abnormally high value numerically. The difference is that in the extraction results of the random forest model, some mesoscale regions with APE exceeding 50% have a large winter wheat planting area. In the correlation analysis, the U-Net model combined with CBAM also shows excellent classification performance, and its correlation coefficient reaches 0.96, which is significantly higher than that of the random forest model (0.92).
[0113] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0114] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0115] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to work in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more processes and / or blocks Figure 1 in the process Figure 1 or processes and / or boxes
[0116] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operational steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more processes and / or blocks Figure 1 in the process Figure 1 or processes and / or boxes
[0117] The foregoing is only a preferred embodiment of the present invention. Although the present invention has been disclosed above in preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make many possible changes and modifications to the technical solution of the present invention, or modify it into equivalent embodiments with equivalent changes, without departing from the scope of the technical solution of the present invention. Therefore, any simple modification, equivalent change and modification made to the above embodiments according to the technical essence of the present invention without departing from the content of the technical solution of the present invention are still within the scope of the protection of the technical solution of the present invention.
Claims
1. A method for crop extraction that couples temporal and spatial attention mechanisms, characterized in that, It includes the following steps: Obtain a remote sensing image dataset of target crops and input it into a trained target detection model to output a crop coverage map of the target area; the target detection model is a deep learning network model of U-Net introducing a temporal and spatial attention mechanism.
2. The method for crop extraction that couples temporal and spatial attention mechanisms according to claim 1, characterized in that, The training process of the target detection model includes: Obtain a sample set of remote sensing images of target crops and perform preprocessing; Construct a deep learning network model based on U-Net and introduce a temporal and spatial attention mechanism to obtain a target detection model; Use the preprocessed sample set of remote sensing images of target crops to train the target detection model.
3. The method for crop extraction that couples temporal and spatial attention mechanisms according to claim 2, characterized in that, Perform preprocessing on the remote sensing image dataset of crops. The specific steps are: classify the remote sensing images of crops, generate labels, and divide them into a training set, a test set, and a validation set.
4. The method for crop extraction that couples temporal and spatial attention mechanisms according to claim 1, characterized in that, The deep learning network model based on U-Net includes a downsampling module, an upsampling module, and an output module; the downsampling module includes several sequentially connected downsampling units. Each downsampling unit extracts deep features of the image through a convolutional layer, and then reduces the resolution through a pooling layer to output a feature map; the upsampling module includes several sequentially connected upsampling units. After each upsampling unit performs upsampling on the feature map, it is concatenated with the feature map of the same resolution in the corresponding downsampling stage, and then the fused feature is refined and adjusted through convolutional operations and the ReLu activation function.
5. The method for crop extraction that couples temporal and spatial attention mechanisms according to claim 4, characterized in that, The introduction of the temporal and spatial attention mechanism specifically includes: connecting an attention module after the first downsampling unit and before the last upsampling unit. The attention module includes a channel attention module and a spatial attention module.
6. The method for crop extraction that couples temporal and spatial attention mechanisms according to claim 5, characterized in that, The data processing process of the channel attention module includes: Perform global average pooling and global max pooling on the input feature map respectively to obtain average-pooled and max-pooled feature descriptors; Process the average-pooled and max-pooled feature descriptors respectively using a multi-layer perceptron, add the results of the average-pooled and max-pooled processed by the multi-layer perceptron, and then map the values to between 0 and 1 through the Sigmoid activation function to obtain channel attention weights; Multiply the attention weights with the input feature map channel by channel to weightedly adjust the channel features.
7. The method for crop extraction that couples temporal and spatial attention mechanisms according to claim 5, characterized in that, The data processing process of the spatial attention module includes: The spatial attention module takes the feature map output by the channel attention module as input, performs max pooling and average pooling on the feature map respectively to obtain two single-channel feature maps; Concatenate the two single-channel feature maps along the channel dimension to obtain a new feature map with 2 channels; Process the concatenated feature map through a convolutional layer, and then map the values to between 0 and 1 through the Sigmoid activation function to obtain a spatial attention weight map; Multiply the spatial attention weight map with the input feature map element by element to highlight the spatial region where the target crops are located in the feature map.
8. The method for crop extraction that couples temporal and spatial attention mechanisms according to claim 4, characterized in that, The output module converts multiple channels of the feature map into the number of channels corresponding to the set segmentation categories through a convolutional layer, thereby converting the feature map into a crop coverage map of the target area.
9. A crop extraction system that couples temporal and spatial attention mechanisms, characterized in that, It includes: Data acquisition module: Acquire the crop remote sensing image dataset and perform preprocessing; Model construction module: Construct a deep learning network model based on U-Net and introduce a temporal and spatial attention mechanism to obtain the target detection model; Model training module: Use the preprocessed crop remote sensing image dataset to train the target detection model.
10. A computer device, characterized in that, The computer device includes: One or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the crop extraction method with a coupled temporal and spatial attention mechanism as described in any one of claims 1-8.
Citation Information
Patent Citations
Crop drought detection method based on remote sensing image
CN115760866A
Cited By
Device and method for quickly processing remote sensing data of unmanned aerial vehicle
CN120358004A