An edge-guided remote sensing method for extracting fine spatial distribution of winter wheat
The edge-guided extraction network EGENet extracts the field shape in high-resolution remote sensing images, which solves the problem of inaccurate field shape extraction in the prior art, especially in the edge area, achieving higher segmentation accuracy and accuracy of field boundaries.
Patent Information
- Application Number
- CN202510191652.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-21
AI Technical Summary
In high-resolution remote sensing images, it is difficult for the prior art to accurately extract the shape of the field, especially in edge areas, resulting in limited segmentation accuracy.
A method for fine spatial distribution remote sensing extraction of winter wheat with edge-guided edge-guided extraction network EGENet is proposed. By establishing an edge-guided extraction network EGENet, the training sample data set is used to train EGENet, and the spatial feature information of remote sensing images, multi-scale semantic feature information and edge information of winter wheat field blocks are obtained, and fusion and prediction are carried out to retain more comprehensive feature information of the field blocks.
It improves the segmentation accuracy at the edge, ensures the accuracy of field boundaries, and significantly improves the accuracy of field shape extraction.
Smart Images

Figure CN119762986B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of agricultural remote sensing monitoring, and in particular to an edge-guided remote sensing extraction method for fine spatial distribution of winter wheat. Background Art
[0002] Winter wheat is one of the main grain crops in my country. The accurate extraction of its field shape is of great value to agricultural insurance and agricultural finance. Complete and accurate field boundary information can not only improve the accuracy of insurance claims, but also provide a scientific basis for agricultural loan assessment and risk control. However, although the pixel-by-pixel segmentation method has achieved remarkable results in accuracy, there is still a certain gap between the extracted field shape and the actual contour. Therefore, how to accurately extract the field shape in high-resolution remote sensing images, especially for the edge area, has become a very concerned issue in current remote sensing image processing.
[0003] Traditional remote sensing image segmentation methods mainly rely on manually designed feature extractors to obtain information such as spectral values, grayscale values, and texture features of images for segmentation. For example, by constructing spectral indices such as the Normalized Difference Vegetation Index (NDVI) and setting thresholds for classification. Combined with the image threshold method, the K-means clustering algorithm is used to segment remote sensing images. These methods have achieved good results on low- and medium-resolution remote sensing images. However, compared with low- and medium-resolution images, high-resolution remote sensing images usually have complex scenes and rich details, resulting in information overload and uneven distribution of target categories. It is difficult for the model to accurately capture and distinguish the features of different categories, resulting in limited segmentation accuracy.
[0004] With the development of deep learning, researchers have continuously proposed new methods for semantic segmentation of remote sensing images. Convolutional neural networks have become the most widely used algorithm in deep learning due to their powerful feature extraction capabilities. However, when segmenting using only convolutional neural networks, there is often a large difference between the feature values of edge pixels and the feature values of internal pixels in the same object, which affects the segmentation accuracy at the edge. In order to improve the segmentation accuracy at the edge, many scholars have introduced edge information into the feature extraction process. For example, edge information is used to guide the network, and edge information is used to guide the aggregation of context information in semantic segmentation. The existing problems are solved through dual-branch encoders and feature aggregation decoders, which provides new ideas for improving segmentation accuracy. Although these methods have improved the segmentation effect at the edge to a certain extent, there are still differences between the shape of the field extraction and the actual shape, especially in the edge area. Summary of the invention
[0005] In order to solve the above technical problems, this application proposes the following technical solutions:
[0006] In a first aspect, an embodiment of the present application provides an edge-guided remote sensing extraction method for fine spatial distribution of winter wheat, comprising:
[0007] Establish an edge-guided extraction network EGENet and train EGENet using the constructed training sample dataset, and adjust EGENet parameters through the validation sample set to obtain the optimal EGENet;
[0008] The acquired remote sensing images are input into EGENet to obtain the spatial feature information, multi-scale semantic feature information and edge information of winter wheat fields.
[0009] Then, the spatial feature information, multi-scale semantic feature information and edge information of the winter wheat field are fused and predicted to achieve information fusion and retain more comprehensive feature information of the winter wheat field.
[0010] In a possible implementation, the EGENet includes: an encoder, a decoder and a classifier, the remote sensing image input, different feature information extracted by the encoder is input into the decoder for fusion and prediction, and the fused features are restored to the same size as the original input image through bilinear interpolation;
[0011] The encoder comprises a first branch consisting of a plurality of semantic feature extraction modules and a second branch consisting of a plurality of edge feature extraction modules EEM, wherein the second branch is connected to an edge perception module HA-EAM based on an attention mechanism;
[0012] The decoder includes a feature fusion module FFM and an upsampling module, and different feature information extracted by the encoder passes through the FFM and the upsampling module in sequence;
[0013] The classifier includes a dynamic weighted classification module EWDC, which receives the output image from the upsampling module to implement classification.
[0014] In a possible implementation, the step of inputting the acquired remote sensing image into EGENet to obtain spatial feature information, multi-scale semantic feature information, and edge information of a winter wheat field of the remote sensing image includes:
[0015] Inputting the remote sensing image into the first branch and the second branch simultaneously to respectively obtain multi-scale semantic feature information and spatial feature information of the remote sensing image;
[0016] The spatial feature information obtained by the second branch is input into the HA-EAM to obtain the edge information of the winter wheat field.
[0017] In a possible implementation, the second branch obtains the spatial feature information of the remote sensing image, including: based on the Sobel operator in the EEM detecting the image features in the horizontal and vertical directions, adding the gradient calculation in two diagonal directions, detecting the edge features of the image in four directions, fusing the gradient features in each direction through the convolution layer, and compensating for the omission of edge features through the residual connection.
[0018] In a possible implementation, the spatial feature information obtained by the second branch is input into the HA-EAM to obtain the edge information of the winter wheat field, including:
[0019] HA-EAM uses the spatial feature image extracted by EEM , =2,3,4, make predictions;
[0020] The spatial feature image is upsampled to make the resolution consistent with the input image, where: , , Two, four, and eight times upsampling are required respectively;
[0021] Use mixed attention to further optimize edge features from both spatial and channel dimensions;
[0022] Finally, the 1×1 convolution kernel is used to splice and output the final edge prediction image.
[0023] In a possible implementation, the spatial feature information, the multi-scale semantic feature information and the edge information of the winter wheat field are fused and predicted to achieve information fusion, and more comprehensive feature information of the winter wheat field is retained, including:
[0024] FFM obtains an edge weight map through the edge prediction image;
[0025] The semantic feature map and the spatial detail feature map are weighted respectively, and the weighted feature maps are cascaded;
[0026] Then through The intermediate results generated by the convolutional layer enter the residual structure combined with the channel attention mechanism. When learning the correlation between input feature channels, it emphasizes important feature channels in the form of weights;
[0027] When performing feature classification, EWDC introduces edge density calculation to measure the richness of edge information in the image and dynamically adjusts the weighting coefficient of the classifier, thereby realizing an adaptive weighting mechanism.
[0028] In a possible implementation, when performing feature classification, EWDC introduces edge density calculation to measure the richness of edge information in the image and dynamically adjusts the weighting coefficient of the classifier, thereby realizing an adaptive weighting mechanism, including:
[0029] Dynamic weighted classification method based on edge density, adding weight term And the marginal weighted classification probability adjustment term EW, the formula is: , Define the formula: ,in, It is the logits output by the model, indicating that a certain pixel belongs to a category The score, represents the index score, Sum the index scores for all categories;
[0030] The edge density is measured by calculating the ratio of edge pixels to total pixels in the image. The higher the edge density, the richer the edge details in the image. The definition formula is: , where: is the edge density, is the sum of the edge strength values in the region, is the width of the image, is the height of the image, is the edge graph;
[0031] To enhance the classification of edge areas, The edge strength To weight the classification probability, more attention can be paid to edge information when classifying edge areas. Combined with logits, the edge information is integrated into the calculation of classification probability. The calculation formula is: .
[0032] In a possible implementation, the weighting coefficient is calculated based on the edge density of the image, and the calculation formula is:
[0033]
[0034] in, is a preset maximum value. It ensures that when the edge density is low, the classifier will still retain a certain Softmax classification weight to avoid over-reliance on edge information. By limiting the maximum value, It always changes within a reasonable range to ensure the balance of the classification process and can be adaptively adjusted according to the changes in edge information in the image.
[0035] In one possible implementation, a weighted loss function combining edge information is introduced in the training of EGENet, and its calculation formula is:
[0036]
[0037] in: , is the weight, is the binary cross entropy loss function, is the marginal loss function; The calculation formula is:
[0038]
[0039] in: is the total number of pixels, is the predicted image segmentation result, yes The pixels, is the actual label value, yes The pixels;
[0040] The calculation formula is:
[0041]
[0042] in: is the total number of pixels, is the predicted image edge result, No. pixels, is the true edge label value, yes The pixels.
[0043] In the embodiment of the present application, a convolutional neural network model is constructed based on the encoder-decoder structure. An edge feature extraction module and an edge perception module are introduced in the encoder to obtain edge information. In the subsequent feature classification process, in order to effectively utilize edge information, the model constructs a dynamic weighted classification method based on edge density. This method dynamically adjusts the output weight of Softmax by combining the edge weight map, so that the classification decision of the edge area is more dependent on accurate edge information, thereby achieving an improvement in classification accuracy at the edge, while ensuring the consistency of the classification result with the actual boundary of the object. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1A schematic diagram of a process of edge-guided remote sensing extraction of fine spatial distribution of winter wheat provided in an embodiment of the present application;
[0045] Figure 2 The EGENet network structure diagram provided for the embodiment of the present application;
[0046] Figure 3 A schematic diagram of the EGENet network training process provided in an embodiment of the present application;
[0047] Figure 4 A schematic diagram of an EEM provided in an embodiment of the present application;
[0048] Figure 5 Schematic diagram of HA-EAM provided in the embodiment of the present application;
[0049] Figure 6 Schematic diagram of the FFM module provided in the embodiment of the present application;
[0050] Figure 7 Schematic diagram of the EDWC module provided in the embodiment of the present application;
[0051] Figure 8 A schematic diagram of the comparison results of different models provided in the embodiments of the present application;
[0052] Fig. 9 Schematic diagram of field boundary extraction results of different models provided in the embodiments of the present application;
[0053] Fig.10 A schematic diagram of the comparison results of ablation experiments provided in the embodiments of the present application;
[0054] Fig.11 A schematic diagram comparing the test results of the two models provided in the embodiments of the present application. DETAILED DESCRIPTION
[0055] The present solution is described below in conjunction with the accompanying drawings and specific implementation methods.
[0056] See also Figure 1 The edge-guided remote sensing extraction method for fine spatial distribution of winter wheat in this embodiment includes:
[0057] S101, establish an edge-guided extraction network EGENet and train EGENet using the constructed training sample data set, and adjust EGENet parameters through the verification sample set to obtain the optimal EGENet.
[0058] The EGENet established in this application consists of multiple paths and multiple basic modules, which can be divided into six parts: input, encoder, decoder, classifier, loss function and output. The basic block diagram is as follows Figure 2As shown, the dotted line frame on the lower side is the encoder structure, and the dotted line frame on the upper side is the decoder structure.
[0059] The encoder consists of a first branch consisting of multiple semantic feature extraction modules and a second branch consisting of multiple edge feature extraction modules EEM. The semantic feature extraction module in the first branch is a residual network Res, and the second branch is connected to an edge perception module HA-EAM based on an attention mechanism. The decoder includes a feature fusion module FFM and an upsampling module. The different feature information extracted by the encoder passes through the FFM and the upsampling module in turn; the classifier includes a dynamic weighted classification module EWDC, which receives the output image from the upsampling module to implement classification.
[0060] In order to realize the training of EGENet, the Gaofen-2 PMS data is selected as the data source. The Gaofen-2 (GF-2) satellite was successfully launched on August 19, 2014, carrying two 1-meter resolution panchromatic cameras and a 4-meter resolution multispectral camera. It has the characteristics of high spatial resolution, wide coverage, long design life, high positioning accuracy, and fast attitude maneuverability. The parameters of the Gaofen-2 PMS data are shown in Table 1.
[0061] Table 1 GF-2 PMS band information
[0062]
[0063] The data comes from the China Resources Satellite Application Center (http: / / www.cresda.com). According to the comparison between ground survey data and remote sensing images, the crops grown in the same period as winter wheat in the GF-2 images from March to April are mainly garlic, and it can be clearly seen that winter wheat is dark green and garlic is light green, and the textures of the two are also significantly different. Therefore, a total of 42 scenes of GF-2 PMS data from March to April 2024 were selected as the basic data for extracting the spatial distribution information of winter wheat.
[0064] After stitching, the 42 images can completely cover the study area. The open source image processing toolkit is used for remote sensing image preprocessing, including radiometric calibration, atmospheric correction, orthorectification, image fusion, etc. The 42 GF-2 PMS data obtained after preprocessing are 4 bands with a spatial resolution of 1m.
[0065] By preprocessing the collected 42 high-resolution PMS images, image data suitable for experimental analysis were obtained. The land cover types in Tai'an City mainly include buildings, bare soil, waters, roads, wheat fields and other crops. Due to the high resolution and obvious feature differences between different types of land objects, winter wheat plots were identified by manual visual recognition combined with ground survey points, and vector data were marked by manual delineation to generate a labeled map containing two categories (winter wheat and non-winter wheat). The pixel value of the winter wheat category is 1, and the pixel value of the non-winter wheat category is 0. Then the original image and the corresponding annotated image were cropped to 512×512 pixel size, and finally 1000 groups of image blocks were obtained. Based on this, the training sample data set, the validation sample data set and the test sample data set were formed. Among them, the training sample is used to train the model, the validation sample is used to select the optimal parameters of the model, and the test sample is used to finally evaluate the model performance.
[0066] The loss function plays an important role in the training of the entire network. The binary cross entropy loss function is a commonly used loss function that can solve the problem of being unable to train due to too small a gradient. In addition, in the network proposed in the application, edge information needs to be used to guide the segmentation process, so it is necessary to add an edge loss function to the loss function. Therefore, the application proposes to use the root mean square error to calculate the error between the real edge and the predicted edge. Based on the above analysis, a weighted loss function combining edge information is proposed, and its calculation formula is:
[0067]
[0068] in, , is the weight, is the binary cross entropy loss function, is the marginal loss function. and The calculation formulas are:
[0069]
[0070] in, is the total number of pixels, is the predicted image segmentation result, yes The pixels, is the actual label value, yes The pixels.
[0071]
[0072] in, is the total number of pixels, is the predicted image edge result, No. pixels, is the true edge label value, yes The pixels.
[0073] By adopting the binary cross entropy loss function and the edge loss function as a joint loss function, the classification accuracy in the edge prediction module can be improved, making the predicted edge closer to the true value.
[0074] like Figure 3 As shown, in this embodiment, test samples, training samples and verification samples are first prepared. The model is trained using the constructed training sample data set. During the training process, the model parameters are initialized, samples are input, forward propagation is performed, the output probability of each category is calculated, and then the error is calculated and the model parameters are updated. This process is repeated until the number of learning times reaches the threshold, thereby obtaining the optimal model. The model parameters are then adjusted through the verification sample set to further optimize the model. After the model is determined, field extraction is performed, and then the spatial distribution extraction of winter wheat in Tai'an City is performed, followed by accuracy evaluation, and the model performance analysis is performed based on the results of the accuracy evaluation.
[0075] The model uses loss as the loss function, uses Adam optimizer as the optimization algorithm, and the input image size is 512 512, batch size is 4, number of iterations is 300, and learning rate is set to 0.05. The model is trained using the stochastic gradient descent algorithm. After each training, an experience file is generated for testing. After the training, the minimum validation loss value is selected as the optimal parameter to test the model performance.
[0076] S102, inputting the acquired remote sensing image into EGENet to obtain spatial feature information, multi-scale semantic feature information and edge information of the winter wheat field.
[0077] After EGENet training is completed, the remote sensing image is simultaneously input into the first branch and the second branch to obtain multi-scale semantic feature information and spatial feature information of the remote sensing image respectively. The spatial feature information obtained by the second branch is input into the HA-EAM to obtain the edge information of the winter wheat field.
[0078] In this embodiment, the second branch is composed of 4 EEMs, and the EEM is composed of an improved Canny algorithm. The traditional Canny algorithm uses the Sobel operator as the gradient operator, and uses the first-order finite difference in the horizontal and vertical directions to calculate the gradient of the filtered image, as shown in the formula ,in, is the gradient calculation in the horizontal direction, is the gradient calculation in the vertical direction. Therefore, this embodiment retains the Sobel operator to detect image features in the horizontal and vertical directions, and adds two diagonal directions (i.e. and ) is calculated as follows: As shown, is the gradient calculation in the first diagonal direction, The gradient calculation in the second diagonal direction is used to detect the edge features of the image in four directions. The gradient features in each direction are fused through the convolution layer, and the residual connection is used to compensate for the omission of edge features. The calculation process is as follows: Figure 4 shown.
[0079] In order to effectively predict edge images, this embodiment designs an edge perception module HA-EAM based on the attention mechanism. This embodiment adopts a hybrid attention mechanism, which can gradually compress the input feature image from the two dimensions of channel and space, and use the extracted feature weights to weight the initial feature image to achieve recalibration of the feature image. The specific process is as follows Figure 5 As shown. HA-EAM uses the feature image extracted by EEM { , =2,3,4} for prediction. First, the feature image is upsampled to make the resolution consistent with the input image. Specifically, , , The upsampling process is as follows: Figure 5 As shown in the black dashed box, transposed convolution and batch normalization are first performed, and the hybrid attention composed of the channel attention module and the spatial attention module is used to further optimize the edge features from the two dimensions of channel and space, and then the ReLU activation function is used. After upsampling, it is spliced after 1×1 convolution, and the final edge prediction image is output after the normalization function. .
[0080] S103, then fusing and predicting the spatial feature information, the multi-scale semantic feature information and the edge information of the winter wheat field, so as to retain more comprehensive feature information of the winter wheat field.
[0081] The decoder in this embodiment is composed of an FFM and an upsampling module, which is responsible for fusing and predicting the spatial feature information, edge information, and semantic features extracted by the encoder in the image, and restoring the fused features to the same size as the original input image through bilinear interpolation. Res shallow features contain rich spatial information while deep features contain more abstract semantic information, and the generated features also contain more spatial detail information. Directly fusing multiple features without considering their inherent differences may seriously affect the performance of the model. Therefore, this paper designs a feature fusion module FFM based on the attention mechanism and residual structure.
[0082] The schematic diagram of FFM is as follows Figure 6 As shown. FFM receives , spatial feature branch Two input features of the same size and edge weights First, through the edge prediction graph Get the edge weight graph The calculation process is as follows:
[0083]
[0084] Semantic feature maps and spatial detail feature maps Weighted, and feature concatenation of the weighted feature maps. Convolutional layers and The convolution layer produces intermediate results. The intermediate results enter the residual structure combined with the channel attention mechanism. When learning the correlation between input feature channels, it emphasizes important feature channels in the form of weights, and finally obtains the final result through feature splicing.
[0085] When performing feature classification, EWDC introduces edge density calculation to measure the richness of edge information in the image and dynamically adjusts the weighting coefficient of the classifier, thereby realizing an adaptive weighting mechanism.
[0086] The EDWC in this embodiment adds a weight term based on the traditional normalized exponential function Softmax. And the marginal weighted classification probability adjustment term EW, as shown in the formula:
[0087]
[0088] The definition is as follows:
[0089]
[0090] in, It is the logits output by the model, indicating that a certain pixel belongs to a category The score, represents the index score, The index scores are summed for all categories.
[0091] This design targets the problem that the normalized exponential function does not fully consider the relationship between each pixel and the edge when classifying the output feature map pixel by pixel, resulting in the problem of misclassification of edge areas. When performing feature classification, edge density calculation is introduced to measure the richness of edge information in the image and dynamically adjust the weighting coefficient of the classifier. , thereby realizing an adaptive weighting mechanism. This method can better integrate edge information during the classification process, significantly improve the classification accuracy of edge areas, and avoid the overfitting problem that may occur when edge information is insufficient.
[0092] Edge density refers to the number of edge pixels per unit area in an image. Edge density can be measured by calculating the ratio of edge pixels to total pixels in an image. The higher the edge density, the richer the edge details in the image. Its definition is as follows:
[0093]
[0094] In the formula is the edge density, is the sum of the edge strength values in the region, is the width of the image, is the height of the image, is the edge graph.
[0095] To enhance the classification of edge areas, The edge strength To weight the classification probability, so that more emphasis can be placed on edge information when classifying edge areas. The calculation formula is as follows:
[0096]
[0097] By increasing the edge strength Combined with logits, the edge information is incorporated into the calculation of classification probability.
[0098] coefficient The calculation method depends on the edge density of the image, and the calculation process is as follows:
[0099]
[0100] is a preset maximum value, determined based on experiments, which ensures that even when the edge density is low, the classifier will still retain a certain Softmax classification weight to avoid over-reliance on edge information. Always change within a reasonable range to ensure the balance of the classification process. Through this dynamic adjustment method, It can be adaptively adjusted according to the changes in edge information in the image, so that the classifier can process different image features more flexibly and improve classification accuracy, especially in edge areas.
[0101] The EDWC classifier structure is as follows Figure 7 As shown, Figure 7 (a) is the edge feature map of the input image. Figure 7 (b) is the feature map output by the last layer of the neural network. Figure 7 (c) is the final result after the classifier. In this classifier, the input of feature map and edge map is the key to the classification process. Specifically, the classifier not only classifies based on the internal features of the pixel, but also combines the edge density information of the edge map to adjust the final classification decision. By calculating the edge density of the edge map, the classifier can adaptively adjust the weighting coefficient , thereby dynamically adjusting the impact of edge information on classification results. When the edge density is high, the weight of edge information increases, and the classifier will pay more attention to the edge areas, thereby improving the classification accuracy of these areas; when the edge density is low, it relies more on the traditional Softmax classification results to avoid the inaccuracy that may be caused by over-reliance on edge information. This method can not only enhance the classification performance of edge areas, but also ensure the balance and robustness of the entire classification process, effectively improving the overall performance in image semantic segmentation tasks, especially in complex backgrounds and unclear edges, and can achieve more ideal classification results.
[0102] In order to verify the above method, the following experiment was designed for verification.
[0103] The experimental study area is Tai'an City, Shandong Province, which is located in the central part of Shandong Province. It is located in the transition zone between the East China Plain and the Central Shandong Mountains, with a total area of about 7,762 square kilometers. It belongs to the warm temperate monsoon climate zone, with four distinct seasons, rain and heat in the same season, and obvious wet and dry periods. The terrain in the region is diverse, and the terrain gradually slopes from northeast to southwest. The main type of cultivated land is dry land. Regular cultivated land is mainly distributed in the plains in the central and southern parts, and fragmented cultivated land is mainly distributed in the hilly areas in the west and north.
[0104] The deep learning framework used in the experimental environment is Pytorch, the development environment is Python3.9, and the calculation is performed through an NVIDIA GeForce RTX4060Ti GPU with 16G video memory.
[0105] In order to verify the effectiveness of the network structure, UNet, ERFNet, DeepLabv3+, and SegNet networks are trained based on the winter wheat field dataset we constructed, and the training results are tested and compared on the test set. DeepLabv3+ is improved on the basis of DeepLabv3. Its backbone network uses deep separable convolutions, and introduces deep separable convolutions in the dilated space pyramid pooling module and the encoder, which further improves the segmentation performance. UNet is a classic semantic segmentation network with a symmetrical encoder-decoder structure. The encoder gradually extracts features, and the decoder gradually restores the resolution. At the same time, features of different scales are fused through jump connections to achieve accurate segmentation effects. ERFNet is an efficient semantic segmentation network that adopts a step-by-step fusion strategy and uses a residual network for feature extraction.
[0106] At the same time, in order to verify the influence of the EDWC method designed in this paper on the extraction effect, on the basis of ensuring that the feature extractor, that is, the encoder and decoder remain unchanged, the softmax classification method is selected to accept the eigenvalues output by the last layer of the neural network for feature classification, so as to verify the effectiveness of the edge information on the feature classification constraints.
[0107] In the process of semantic segmentation, in order to evaluate the performance of the segmentation method, we used the following indicators: Accuracy, Precision, Recall and F1 score.
[0108]
[0109]
[0110]
[0111]
[0112] Among them, the real example Represents the number of pixels that are correctly predicted as positive examples, false positive examples Indicates the number of pixels that are incorrectly predicted as positive examples, true negative examples Indicates the number of pixels of correctly predicted counterexamples and false negatives Represents the number of pixels of incorrectly predicted counterexamples.
[0113] This paper uses the EGENet model to extract the spatial distribution of winter wheat planting in Tai'an City, Shandong Province. The winter wheat in Tai'an City is mainly concentrated in the plains in the central and southern parts, and the spatial distribution results are consistent with the suitable planting areas. It can be seen that for plains with flat terrain and flat plots, the EGENet model can extract the winter wheat planting areas relatively completely. For plots close to cities, roads, lakes and other land uses, the model can better identify their boundaries. In areas with complex terrain, the model can effectively distinguish between winter wheat planting areas in mountainous areas and other cultivated land, and more accurately extract irregularly shaped and scattered mountainous plots.
[0114] Figure 8 The results generated by five models in five regions are given. Figure 8 (a) is the original remote sensing image; Figure 8 (b) is the labeled image; Figure 8 (c) is the recognition result of EGENet model; Figure 8 (d) is the recognition result of ErfNet model; Figure 8 (e) is the recognition result of the UNet model; Figure 8 (f) is the recognition result of the segNet model; Figure 8 (g) is the recognition result of DeepLabv3+ model; Figure 8 As can be seen in (c), all the results of EGENet are relatively ideal. The edge contour of the winter wheat area is smoother and closer to the actual field boundary. At the same time, there are no broken points or holes in the extracted fields, and there are fewer misclassification phenomena. Figure 8 Middle (d) - Figure 8 As can be seen in (g), SegNet performs relatively complete in field segmentation, but has low edge segmentation accuracy. This is because it lacks the guidance of high-level features when fusing underlying detail information through jump connections, and is easily disturbed by irrelevant information. DeepLabv3+ uses the ASPP module to obtain multi-scale features with a large receptive field while retaining rich spatial information. It shows a high degree of completeness in winter wheat field segmentation, but there are still some deficiencies in the extraction of small roads. The results extracted by ERFNet have more small boundaries missed. The results extracted by UNet have rough edges, with more broken points and holes.
[0115] In order to more accurately analyze the performance of the EGENet model in the winter wheat extraction task, the accuracy of the field extraction results of EGENet and the comparison network was evaluated using four evaluation indicators: Accuracy, Precision, Recall and F1. As can be seen from Table 2, compared with the comparison network, the EGENet network has the highest segmentation accuracy on this dataset, which is 5.17%, 4.96%, 4.12% and 4.54% higher than UNet in terms of Accuracy, Precision, Recall and F1, respectively, showing good segmentation performance. Compared with DeepLabv3+ and ErfNet, the network structure of the EGENet model has a large improvement in various accuracy evaluation indicators. At the same time, compared with SegNet, the various indicators of the model have also been improved. In general, the EGENet model proposed in this paper shows good overall segmentation performance. It can not only completely segment the field, but also correct the field edge, which fully demonstrates the effectiveness of the proposed model.
[0116] Table 2 Comparison of evaluation indicators of different models
[0117]
[0118] In order to more intuitively compare the shape boundaries of winter wheat fields extracted by different networks, Python code is used to extract the boundaries of the predicted binary image, as shown below: Fig. 9 Winter wheat field boundaries shown. Fig. 9 It can be seen that compared with other networks, the field boundaries extracted from the prediction results of the EGENet network are clearer and smoother, and closer to the actual field boundary shape.
[0119] In order to verify the effectiveness and robustness of the HA-EAM module, FFM module and EDWC classifier in the proposed model for winter wheat field segmentation, this paper conducted an ablation experiment on the constructed winter wheat field dataset.
[0120] The proposed model is improved based on ResNet18, so it is used as the encoder in the experiment, layer-by-layer deconvolution upsampling is used as the decoder, and softmax is used as the classifier to build a U-Net model as the baseline network. The experimental results are shown in Fig.10 As shown: Fig.10 (a) is the original image; Fig.10 (b) is the label; Fig.10 (c) is the baseline network result; Fig.10 (d) shows the results of introducing EEM and HA-EAM; Fig.10 Column (e) shows the results of introducing EEM, HA-EAM and FFM; Fig.10Column (f) shows the results of introducing EEM, HA-EAM, FFM and EDWC.
[0121] from Fig.10 It can be clearly seen that after introducing EEM and HA-EAM on the basis of the baseline network, the integrity of the field has been significantly improved compared with the baseline network, and the roughness at the edge has also been improved. After the introduction of the FFM module branch, the integrity of the field segmentation is improved, and the boundaries of the field are clearer. Finally, in order to verify the influence of the EDWC method designed in this paper on the extraction effect, on the basis of ensuring that the feature extractor, that is, the encoder and decoder remain unchanged, the softmax classification method is selected to accept the feature values output by the last layer of the neural network for feature classification. That is, after the introduction of EDWC, the boundaries of the field are further corrected, the segmentation integrity is higher, and the obtained field shape is more consistent with the actual.
[0122] In order to quantitatively analyze the effectiveness of the introduced branches on the model, the performance index comparison results of the ablation experiment are shown in Table 3.
[0123] Table 3 Comparison of evaluation indicators of ablation experiments
[0124]
[0125] It can be seen from Table 3 that after the baseline model introduces EEM and HA-EAM, the Accuracy and F1 index increase by 2.31% and 1.36% respectively. After introducing FFM into the model, the Accuracy and Precision increase by 3.76% and 3.14% respectively. Finally, EDWC is introduced, that is, the softmax classifier is replaced by EDWC. The Accuracy, Precision, Recall and F1 of the model reach 90.91%, 91.86%, 92.26% and 92.06% respectively, and the segmentation accuracy is further improved.
[0126] From the above experiments, it can be seen that when using convolutional neural networks to extract the spatial distribution of winter wheat, the effective use of edge information has an important impact on the feature extraction ability. In the same winter wheat farmland area, the internal pixel feature consistency is relatively high, while the edge pixels are affected by many factors, such as interaction with adjacent objects, illumination changes, etc., resulting in significant differences between their features and internal pixels. When extracting features, traditional convolutional neural networks often focus on overall texture, color and other information, and it is difficult to accurately distinguish the subtle differences between edge pixels and internal pixels, making it inaccurate when identifying edge areas. Therefore, the feature extraction capability is improved by introducing edge information.
[0127] This effect can be clearly seen by comparing the segmentation results of different models. ErfNet adopts an end-to-end encoder-decoder structure, which has certain advantages in computational efficiency and learning ability. However, in the edge segmentation task, due to the lack of an effective integration mechanism specifically for edge information, the model cannot fully capture edge details, making the field boundary unclear. Unet constructs a network structure in which the first half is used for feature extraction, and the second half performs upsampling and adopts multi-scale feature fusion technology. Although edge features are retained to a certain extent, due to insufficient utilization of edge details, misclassification is easily caused in complex edge areas, such as the junction of field boundaries and adjacent objects such as roads and ditches, resulting in inaccurate distinction between the boundaries of winter wheat and other objects. Therefore, the EGENet method uses an edge feature extraction module and an edge perception module based on an attention mechanism to effectively extract edge features. Subsequently, in order to achieve effective utilization of edge information, the edge and semantic information are deeply fused through the feature fusion module, so that multiple features complement each other, improve the segmentation effect of the model on the edge, and achieve significant improvement. This fully demonstrates that the effective use of edge information can make up for the shortcomings of traditional convolutional neural networks in feature extraction in edge areas, greatly improving the overall feature extraction capability of the model, enabling it to more accurately identify winter wheat fields in complex scenarios.
[0128] The performance of the classifier is directly related to the accuracy of the final classification results. Although convolutional neural networks perform well in the feature extraction stage and can effectively capture various features in the image, if the classifier cannot reasonably use these features, it will still lead to deviations in the classification results.
[0129] When extracting winter wheat areas, high-resolution images contain rich ground object information and complex scenes, and the differences between different ground object categories are sometimes subtle. During the processing, although the Softmax classifier can calculate the probability distribution of each category based on the output of the neural network, it has obvious limitations. It mainly relies on the characteristic value of the pixel itself. When classifying pixel by pixel, it fails to fully consider the close connection between the pixel and the edge of the object, and the potential impact of the surrounding pixels on the current pixel classification. In the edge area of the winter wheat field, the pixel characteristics are greatly disturbed by the adjacent ground objects. Due to the lack of effective integration of edge information, the Softmax classifier is prone to misjudgment of the pixel category, thereby affecting the overall classification accuracy. Therefore, the edge density-based dynamic weighted classification method (EDWC) proposed in this paper is innovatively designed to address the shortcomings of Softmax. By calculating the edge density of the edge weight map to dynamically adjust the α value, an adaptive weighting mechanism is realized. In areas with high edge density, the α value is reduced to rely on edge information to accurately judge the edge pixel category. When the edge density is low, the α value is increased to focus on Softmax's own classification ability to avoid misjudgment. At the same time, the EW term further strengthens the role of edge information in the classification process by introducing edge strength weighted classification probability, making the classification result more in line with the actual situation. This paper designs a comparative experiment between Softmax and EDWC while keeping the feature extractor unchanged to verify the impact of the classifier on the segmentation results.
[0130] See also Fig.11 , Fig.11 (a) is the original image. Fig.11 (b) is the label map. Fig.11 (c) is the softmax classification result diagram. Fig.11 Middle (d) is the EWEC classification result diagram. Fig.11 In (c) and (d) of Figure 11, six segmentation details are marked in red. In the results of classification using Softmax, edges 1 and 2 are not fully recognized, edges 3 and 4 are misclassified, and edges 5, 6 and other edges are mostly not smooth and have jagged edges. Compared with Softmax, when EDWC is used for classification, edges 1 and 2 in the results can be accurately and completely recognized, edges 5 and 6 are smoother and clearer, and edges 3 and 4 are not misclassified, which is more consistent with the actual situation. It can be seen that when EDWC is used for classification, the recognition accuracy of the EGENet model in object edges is improved, which proves the effectiveness of the EDWC proposed in this embodiment.
[0131] In the embodiments of the present application, "at least one" refers to one or more, and "plurality" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent the existence of A alone, the existence of A and B at the same time, and the existence of B alone. Among them, A and B can be singular or plural. The character " / " generally indicates that the previous and subsequent associated objects are in an "or" relationship. "At least one of the following" and similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b and c can be represented by: a, b, c, ab, ac, bc, or abc, where a, b, c can be single or multiple.
[0132] The above is only a specific implementation of the present application. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. The protection scope of the present application should be based on the protection scope of the claims.
Claims
1. An edge-guided remote sensing extraction method for fine spatial distribution of winter wheat, characterized in that: include: Establish an edge-guided extraction network EGENet and train EGENet using the constructed training sample dataset, and adjust EGENet parameters through the validation sample set to obtain the best EGENet; The acquired remote sensing images are input into EGENet to obtain the spatial feature information, multi-scale semantic feature information and edge information of winter wheat fields. The EGENet includes: an encoder, a decoder and a classifier. The remote sensing image is input, and different feature information extracted by the encoder is input into the decoder for fusion and prediction. The fused features are restored to the same size as the original input image through bilinear interpolation. The encoder includes a first branch composed of multiple residual networks Res and a second branch composed of multiple edge extraction modules EEM, and the second branch is connected to an edge perception module HA-EAM based on an attention mechanism; The decoder includes a feature fusion module FFM and an upsampling module, and different feature information extracted by the encoder passes through the FFM and the upsampling module in sequence; The classifier includes a dynamic weighted classification module EWDC, which receives the output image from the upsampling module to implement classification; The step of inputting the acquired remote sensing image into EGENet to obtain spatial feature information, multi-scale semantic feature information and edge information of the winter wheat field, includes: Inputting the remote sensing image into the first branch and the second branch simultaneously to respectively obtain multi-scale semantic feature information and spatial feature information of the remote sensing image; The spatial feature information obtained by the second branch is input into the HA-EAM to obtain the edge information of the winter wheat field; Then, the spatial feature information, multi-scale semantic feature information and edge information of the winter wheat field are fused and predicted to retain more comprehensive feature information of the winter wheat field.
2. The edge-guided remote sensing extraction method for fine spatial distribution of winter wheat according to claim 1, characterized in that: The second branch obtains the spatial feature information of the remote sensing image, including: on the basis of the Sobel operator in the EEM detecting the image features in the horizontal and vertical directions, adding the gradient calculation in the two diagonal directions, detecting the edge features of the image in four directions, fusing the gradient features in each direction through the convolution layer, and compensating for the omission of edge features through the residual connection.
3. The edge-guided remote sensing extraction method for fine spatial distribution of winter wheat according to claim 1 or 2, characterized in that: The spatial feature information obtained by the second branch is input into the HA-EAM to obtain the edge information of the winter wheat field, including: HA-EAM uses the spatial feature image extracted by EEM , =2,3,4, make predictions; The spatial feature image is upsampled to make the resolution consistent with the input image, where: , , Two, four, and eight times upsampling are required respectively; Use mixed attention to further optimize edge features from both spatial and channel dimensions; Finally, the 1×1 convolution kernel is used to splice and output the final edge prediction image.
4. The edge-guided remote sensing extraction method for fine spatial distribution of winter wheat according to claim 3 is characterized in that: The spatial feature information, multi-scale semantic feature information and edge information of the winter wheat field are integrated and predicted to achieve information fusion, and more comprehensive feature information of the winter wheat field is retained, including: FFM obtains an edge weight map through the edge prediction image; The semantic feature map and the spatial detail feature map are weighted respectively, and the weighted feature maps are cascaded; Then through The intermediate results generated by the convolutional layer enter the residual structure combined with the channel attention mechanism. When learning the correlation between input feature channels, it emphasizes important feature channels in the form of weights; When performing feature classification, EWDC introduces edge density calculation to measure the richness of edge information in the image and dynamically adjusts the weighting coefficient of the classifier, thereby realizing an adaptive weighting mechanism.
5. The edge-guided remote sensing extraction method for fine spatial distribution of winter wheat according to claim 4, characterized in that: When performing feature classification, EWDC introduces edge density calculation to measure the richness of edge information in the image and dynamically adjusts the weighting coefficient of the classifier, thereby realizing an adaptive weighting mechanism, including: Dynamic weighted classification method based on edge density, adding weight term And the marginal weighted classification probability adjustment term EW, the formula is: , Define the formula: ,in, It is the logits output by the model, indicating that a certain pixel belongs to a category The score, represents the index score, Sum the index scores for all categories; The edge density is measured by calculating the ratio of edge pixels to total pixels in the image. The higher the edge density, the richer the edge details in the image. The definition formula is: , where: is the edge density, is the sum of the edge strength values in the region, is the width of the image, is the height of the image, is the edge graph; To enhance the classification of edge areas, The edge strength To weight the classification probability, more attention can be paid to edge information when classifying edge areas. Combined with logits, the edge information is integrated into the calculation of classification probability. The calculation formula is: .
6. The edge-guided remote sensing extraction method for fine spatial distribution of winter wheat according to claim 5, characterized in that: The calculation method of the weighting coefficient depends on the edge density of the image, and the calculation formula is: in, is a preset maximum value. Adaptively adjust according to the changes in edge information in the image.
7. The edge-guided remote sensing extraction method for fine spatial distribution of winter wheat according to claim 1, characterized in that: The weighted loss function combined with edge information is introduced in the training of EGENet, and its calculation formula is: in: , is the weight, is the binary cross entropy loss function, is the marginal loss function; The calculation formula is: in: is the total number of pixels, is the predicted image segmentation result, yes The pixels, is the actual label value, yes The pixels; The calculation formula is: in: is the total number of pixels, is the predicted image edge result, No. pixels, is the true edge label value, yes The pixels.
Citation Information
Patent Citations
High-resolution remote sensing image land coverage classification method based on local detail enhancement and edge constraint
CN113343789A
Winter wheat planting area image extraction method combining GF-6 and Sentinel-2
CN114842339A