A Provincial Winter Wheat Remote Sensing Identification Method Based on an Improved U2-Net Network Model
By adding a multi-scale channel attention module to the U2-Net network model, the accuracy of remote sensing recognition and microstructure processing capabilities of winter wheat are improved, and the problem of winter wheat distribution recognition at provincial scale is solved, and high-accurate winter wheat distribution image recognition is achieved.
Patent Information
- Application Number
- CN202411213598.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-30
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-08-30
AI Technical Summary
The prior art is difficult to accurately identify and extract winter wheat distribution on provincial scales, especially in large-scale areas, where there is a problem of lack of global consistency in the identification results, blurred boundaries, and insufficient microstructure processing.
Improve the U2-Net network model by adding a multi-scale channel attention module after the first three encoder modules, enhancing the model's ability to capture and fusion of different scale features, and retaining spatial information through depth separation convolution operations to reduce the amount of calculation.
The image recognition accuracy of remote sensing recognition of winter wheat is improved, and the distribution of winter wheat can be accurately extracted on a provincial scale, solving the problem of global consistency of recognition results and insufficient microstructure processing.
Smart Images

Figure CN119107563B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of satellite remote sensing technology, and particularly relates to a remote sensing identification method for winter wheat at the provincial level based on an improved U2-Net network model. Background Art
[0002] Winter wheat is an important food crop in China. Timely and accurately grasping the planting distribution and area change of winter wheat is of great significance for maintaining the food security bottom line of 1.4 billion people in China and formulating relevant policies and issuing agricultural subsidies in a timely manner. Remote sensing technology has the advantages of wide coverage, low cost, objectivity, etc., and has now become an important means for quickly extracting the spatial distribution of crops.
[0003] Traditional remote sensing extraction methods for winter wheat mainly rely on shallow features such as the spectrum, texture, and shape of winter wheat in remote sensing images. By extracting general information or feature information from the input data, the images are classified and predicted. For example, machine learning methods such as vegetation index threshold method, support vector machine, and random forest. Machine learning classification methods mainly involve the computer obtaining feature information such as color, texture, and space from known data and using classification rules extracted and set by professionals to predict and classify unknown data. The setting of eigenvalue and classification rules requires human participation. In addition, machine learning classification methods require a certain number of samples to be labeled for all categories in the study area, rather than only labeling the target ground object samples, resulting in a large workload for sample labeling.
[0004] Deep learning is a pattern analysis method in which a machine automatically learns the internal features and representation levels of sample data through multiple network structures. It can automatically obtain some features that humans cannot imagine and cannot be represented by combinations such as color or space, making up for the deficiencies of traditional remote sensing extraction methods to a certain extent. Convolutional neural network is the main network structure of deep learning, generally consisting of an input layer, a convolutional layer, a pooling layer, a fully connected layer, and an output layer. The input layer is a pixel matrix of an image; the convolutional layer is the inner product operation of the input sample and the convolutional kernel; the first convolution is performed on the input layer to obtain a feature map. In the second and subsequent convolutional layers, the feature map of the previous layer is convolved; the role of the pooling layer is to reduce the size of the feature map generated by the convolutional layer, reducing the image resolution while extracting image features; the fully connected layer calculates the activation value and calculates the output value corresponding to the feature map through the activation function; the output layer uses the likelihood function to calculate the probability that a pixel belongs to a certain category.
[0005] Convolutional neural networks can automatically obtain the semantic information features of target sample data through complex convolution processes and backpropagation. By treating unlabeled regions as negative samples, they can reduce the impact of undersampling on the model through data augmentation means even when the positive and negative samples are uneven. Currently, research on crop extraction using deep learning is basically only applied to small areas at the county level, and there are still few studies on large-scale crop extraction at the provincial scale. Summary of the Invention
[0006] To solve at least one technical problem existing in the prior art, the present application provides a method for remotely sensing and identifying winter wheat at the provincial level based on an improved U2-Net network model.
[0007] The present application discloses a method for remotely sensing and identifying winter wheat at the provincial level based on an improved U2-Net network model, including the following steps:
[0008] S1. Obtain the optical remote sensing image of winter wheat in the area to be measured, and preprocess the optical remote sensing image, where the area to be measured is an area at the provincial level;
[0009] S2. Select a predetermined number of original sample images from the preprocessed optical remote sensing image, then disperse the original sample images into a predetermined number of sample frames, and perform sliding window cropping processing. Subsequently, use the object-oriented method for image segmentation and combine visual interpretation to complete sample drawing to obtain a sample data set;
[0010] S3. Augment the sample data set, and divide the augmented sample data set into a training set and a validation set. Subsequently, build an improved U2-Net network model, and use the training set to train the improved U2-Net network model, where the improved U2-Net network model adds a multi-scale channel attention module after the first three encoder modules in the original U2-Net network model, and the multi-scale channel attention module includes:
[0011] Depth-wise convolution, which is used to aggregate local information, and its output will be used as attention weights to re-weight the input;
[0012] Multi-branch depth-wise convolution, which is used to capture multi-scale context information, and its output will be used as attention weights to re-weight the input;
[0013] Convolution, which is used to perform correlation modeling in the channel dimension, and its output will be used as attention weights to re-weight the input;
[0014] S4. Use the validation set to evaluate the accuracy of the improved U2-Net network model trained in step S3. If the accuracy requirement is met, proceed to step S5; if not, return to steps S2 and S3 to check the sample drawing and adjust and optimize the improved U2-Net network model until the requirement is met.
[0015] S5. Apply the improved U2-Net network model obtained in step S4 to the optical remote sensing images of winter wheat in the area to be measured for prediction scene by scene, extract the distribution area of winter wheat, and then splice the extraction results of all scene images to obtain the winter wheat distribution map of the area to be measured.
[0016] In an optional implementation manner, in step S1, the optical remote sensing image is L2-level surface reflectance data, and the preprocessing of the optical remote sensing image at least includes band combination, splicing, and cropping. Among them, the optical remote sensing image after band combination includes at least 5 bands: blue, green, red, near-infrared, and short-wave infrared.
[0017] In an optional implementation manner, in the improved U2-Net network model, the multi-scale channel attention module is used for:
[0018] Perform depthwise separable convolution operation on the input features through the depthwise convolution, retain the spatial information and reduce the computational amount, then input the features into the multi-branch depthwise convolution, and each branch uses convolution kernels of different sizes to capture features of different scales. Then, splice the outputs of multiple branches in the channel dimension to converge the multi-scale features, and then use a 1×1 convolution layer to fuse and adjust the channels of the spliced feature map to generate the final output.
[0019] In an optional implementation manner, in the improved U2-Net network model, each encoder and decoder uses a residual U block. Each residual U block uses a basic U-shaped structure and adds a shortcut branch for residual signal transmission on the basis of the U-shaped structure, which is used to directly transmit the residual to the subsequent network when the network propagates errors.
[0020] In an optional implementation manner, in the improved U2-Net network model, the input of each decoder is obtained by cascading the upsampling of the features of its previous stage and the symmetric encoder features after passing through the multi-scale channel attention module.
[0021] In an optional implementation manner, at the end of the model training in step S3, the side output saliency maps of 6 decoders are obtained through 3×3 convolution, and the 6 saliency maps are spliced and then passed through 1×1 convolution and sigmoid activation function to obtain the final prediction map.
[0022] In an alternative embodiment, when building the improved U2-Net network model in step S3, 6 residual U blocks are used in the encoding stage and 5 residual U blocks are used in the decoding stage. Among them, the first residual U block has one more downsampling and upsampling process respectively based on the second residual U block. The third and fourth residual U blocks gradually reduce one downsampling and upsampling process based on the second residual U block. The fifth and sixth residual U blocks use dilated convolution instead of pooling based on the second residual U block to avoid detail loss caused by excessive downsampling. Additionally, in the decoding stage, the connection of the upsampled feature map from the previous stage and the downsampled feature map from the symmetric encoder stage is used as the input to transmit context information.
[0023] In an alternative embodiment, in the step S3, the improved U2-Net network model is trained based on the pytorch framework, and the model parameters and the settings of each model parameter are as follows:
[0024] Optimization function, which is set as the stochastic gradient descent optimizer;
[0025] Loss function, which is set as the cross-entropy loss function;
[0026] Initial value of learning rate, which is set as 0.001;
[0027] Learning rate, which is set as the warmup strategy and the cosine annealing strategy;
[0028] Batch size, which is set as 4.
[0029] In an alternative embodiment, in the step S4, the model accuracy evaluation formula is as follows:
[0030]
[0031] Where P represents the accuracy rate, R represents the recall rate, F1 represents the accuracy of the winter wheat recognition model, TP represents the area of winter wheat correctly extracted by the model, FN represents the area of winter wheat missed by the model, and FP represents the area of winter wheat mis-extracted by the model;
[0032] Where when the F1 score is greater than or equal to 0.9, it is determined that the accuracy requirement is met.
[0033] In an alternative embodiment, the provincial winter wheat remote sensing recognition method further includes:
[0034] S6. Extract the winter wheat area by using the winter wheat distribution map obtained in the step S5, and verify the extraction result through the following formula:
[0035]
[0036] Among them, P a represents the extraction error of winter wheat, and A 1 represents the extracted area of winter wheat, and A 2 represents the true area of winter wheat, which is the area of winter wheat obtained by using optical remote sensing images with higher spatial resolution and combining computer automatic classification with visual refinement method.
[0037] The present application has at least the following beneficial technical effects:
[0038] 1) The provincial winter wheat remote sensing recognition method based on the improved U2-Net network model in the present application optimizes and improves the U2-Net network structure. To solve the problems of lack of global consistency, fuzzy boundaries, and insufficient processing of fine structures in the U2-Net network prediction results, a multi-scale channel attention mechanism is added to the U2-Net network structure, which can fuse features of different scales such as the space and spectrum of crops, and at the same time can adaptively perform feature recalibration on the multi-dimensional channel attention weights to learn richer high-level semantic information and improve the image recognition accuracy.
[0039] 2) The provincial winter wheat remote sensing recognition method based on the improved U2-Net network model in the present application constructs a provincial-scale winter wheat sample data set and recognition model. In the existing technology, there is no such literature research. The present application uses Sentinel data, considers the planting distribution characteristics of provincial winter wheat and the sample size requirements of the deep learning model, and constructs a provincial-scale winter wheat sample data set and a deep learning remote sensing recognition model.
[0040] 3) In the provincial winter wheat remote sensing recognition method based on the improved U2-Net network model in the present application, the multi-scale attention module can perform depthwise separable convolution operations on the input features, retain spatial information and reduce the amount of calculation, then input the features into multiple branches, each branch uses different-sized convolutional kernels to capture features of different scales, then splice the outputs of multiple branches in the channel dimension to converge multi-scale features, and finally use a 1*1 convolutional layer to fuse and adjust the channels of the spliced feature map to generate the final output, enhancing the overall feature expression ability of the model at different scales.
[0041] 4) In the provincial winter wheat remote sensing recognition method based on the improved U2-Net network model in the present application, the network model selects to use the cross-entropy loss function, and the learning rate selects the strategies of warmup and cosine annealing, making the adjustment of the learning rate more flexible.
[0042] 5) The provincial winter wheat remote sensing identification method based on the improved U2-Net network model of this application has been demonstrated and applied at the provincial scale in Henan Province. At the same time, promotion experiments have been carried out in the surrounding counties of Henan Province. Good winter wheat extraction effects have been achieved in both the demonstration application and promotion experiment areas. Moreover, the application is simple and the identification is fast, having the advantage of further promotion and application. Description of the Drawings
[0043] Figure 1 is the flow chart of the provincial winter wheat remote sensing identification method based on the improved U2-Net network model of this application;
[0044] Figure 2 is the winter wheat sample distribution map obtained after completing the sample drawing in an embodiment of the provincial winter wheat remote sensing identification method based on the improved U2-Net network model of this application;
[0045] Figure 3 is the schematic diagram of the network structure of the improved U2-Net network model in the provincial winter wheat remote sensing identification method based on the improved U2-Net network model of this application;
[0046] Figure 4 is the schematic diagram of the multi-scale attention module structure in the provincial winter wheat remote sensing identification method based on the improved U2-Net network model of this application;
[0047] Figure 5 is the schematic diagram of the network structure of the second residual U block in the provincial winter wheat remote sensing identification method based on the improved U2-Net network model of this application;
[0048] Figure 6 is the schematic diagram of the network structure of the fifth residual U block in the provincial winter wheat remote sensing identification method based on the improved U2-Net network model of this application;
[0049] Figure 7 is the comparison chart of the winter wheat identification effects using the existing U2-Net network model and the improved U2-Net network model of this application;
[0050] Figure 8 is the spatial distribution map of winter wheat in Henan Province extracted by using the provincial winter wheat remote sensing identification method based on the improved U2-Net network model of this application;
[0051] Figure 9 is the spatial distribution map of winter wheat in Xun County extracted by using the provincial winter wheat remote sensing identification method based on the improved U2-Net network model of this application;
[0052] Figure 10 is the spatial distribution map of winter wheat in Qi County extracted by using the provincial winter wheat remote sensing identification method based on the improved U2-Net network model of this application. Specific implementation manners
[0053] To make the purpose, technical solutions, and advantages of the implementation of this application clearer, the technical solutions in the embodiments of this application will be described in more detail below in conjunction with the accompanying drawings in the embodiments of this application. The described embodiments are part of the embodiments of this application, rather than all of the embodiments.
[0054] The following combines the attached Figures 1 - 10 To further elaborate in detail on the provincial winter wheat remote sensing identification method based on the improved U2-Net network model of this application.
[0055] This application discloses a provincial winter wheat remote sensing identification method based on an improved U2-Net network model, as Figure 1 shown, including the following steps:
[0056] Step S1: Data acquisition and preprocessing.
[0057] The data in this step mainly refers to the Sentinel-2 optical remote sensing images of winter wheat in the area to be measured, and this data can be obtained from the official website of the European Space Agency.
[0058] In this embodiment, Henan Province is taken as the research object. When selecting data, according to the remote sensing image characteristics of winter wheat in different growth stages, cloud-free images of Henan Province from March to April 2023 were selected for download. The downloaded images are L2-level surface reflectance data, and preprocessing such as band combination, mosaicking, and cropping is performed in remote sensing software such as envi. Among them, the images after band combination processing contain at least 5 bands: blue, green, red, near-infrared, and short-wave infrared.
[0059] Step S2: Sample quantity calculation and sample drawing.
[0060] In this step, considering that the deep convolutional neural network generally requires more than 5000 sample images, and means such as horizontal flipping, vertical flipping, and image rotation (selectively rotating 90°, 180°, 270°) for data augmentation to expand the sample set, approximately 1600 original sample images are required to meet the requirements.
[0061] In addition, to ensure sample diversity, referring to Figure 2As shown, according to the spatial distribution characteristics of winter wheat planting in Henan Province, this application arranges sample frames in 8 different regions, namely, 5 sample frames are arranged in 2 major production areas, namely, the central and northern wheat area (including Xuchang, Zhengzhou, Luoyang and irrigated land north of the Yellow River) and the central and eastern wheat area (including the central and northern Zhumadian, Luohe, Zhoukou, Shangqiu, and eastern Pingdingshan); and 1 sample frame is arranged in each of the wheat areas in the Nanyang Basin (including Nanyang City and Biyang County of Zhumadian), the rice stubble wheat area in southern Henan (including Xinyang, southern Zhumadian and Tongbai, Nanyang, etc.), and the dryland wheat area in western, southwestern, and northern Henan (including Sanmenxia, Jiyuan, Pingdingshan, Anyang and other low-mountain and hilly areas).
[0062] 1600 original sample images are distributed to 8 sample frames, so each sample frame needs to have about 200 sample images. According to the sample frame design as a square, each sample frame can have 14*14=196 sample images. Considering that the input sample image size of the convolutional neural network cannot be too large, it is generally 224*224 pixels in size. In order to make the sample image more sufficient during training and avoid the position error caused by subsequent splicing, this application adopts a sliding window cropping method in each sample frame, and the sliding window overlap rate is set to 40%. The horizontal and vertical sliding window steps are both 14 pixels (to ensure that there are 14*14 sample images). Taking the Sentinel-2 image as an example (spatial resolution is 10 meters), the side length of each sample frame is 1*224*10+13*(1-0.4)*224*10=19712m≈20km.
[0063] Finally, 8 approximately 400km long 2 sample area; further, Figure 2 As shown in the figure, the image in each sample frame is segmented using an object-oriented method, and the features are extracted as segmentation objects of winter wheat. Combined with manual visual correction (i.e. visual interpretation), winter wheat samples are accurately drawn in 8 sample frames, and sample drawing is completed, resulting in a total of 32,400 surface patches. Among them, 80% of the samples are used as training sets for model training, and 20% of the samples are used as validation sets for model evaluation.
[0064] Step S3: build an improved U2-Net network model and use the training data set obtained in step S2 to train the improved U2-Net network model.
[0065] Among them, the improved U2-Net (also called U 2 -Net) network model is an improvement based on the existing U2-Net network model.
[0066] The U2-Net network model structure has good anti-noise ability and has very good results in the saliency object detection task. It is often used to solve binary classification problems. Therefore, this application selects U2-Net as the basic network model structure.
[0067] Specifically, as Figure 3 shown, the improved U2-Net network model body has a U-shaped symmetric encoder-decoder structure. The left side is the convolution and pooling layers for extracting shallow and deep features, which is the feature downsampling process; the right side is the connection layer, which is the feature upsampling process. The feature maps obtained by each convolution layer of the network will be connected to the corresponding downsampling layer, so that the final obtained feature map results contain enough high-level and low-level features, realizing the fusion of features at different scales and improving the model recognition accuracy.
[0068] Furthermore, the improved U2-Net network model is also a two-level nested U structure, which contains 6 encoders and 6 decoders. Each encoder and decoder uses a residual U block (hereinafter referred to as RSU for short), and the RSU is also a U-shaped structure. In addition, at least one shortcut branch for residual signal transmission is set in the two-level nested U structure, so that when the network propagates errors, the residuals can be directly transmitted to the corresponding network behind through this shortcut branch, rather than being transmitted layer by layer backward.
[0069] In the encoding stage of building the model, each encoder contains 6 RSUs. Among them, the first 4 RSUs are set to reduce the dimension of the feature map, increase the receptive field, and obtain multi-scale information. As Figure 5 shown, it is the detailed network structure of the second RSU, and the first RSU has one more downsampling and upsampling process respectively on the basis of the second RSU. The third and fourth RSUs have one less downsampling and upsampling process respectively on the basis of the second RSU. The last two RSUs (i.e., the fifth and sixth) are set to avoid detail loss caused by excessive downsampling, and dilated convolution is used instead of pooling. Figure 6 shows the schematic diagram of the network structure of RSU5 (i.e., the fifth RSU). The structure of RSU6 (i.e., the sixth RSU) is the same as that of RSU5, and dilated convolutions with different dilation coefficients are used to retain feature information.
[0070] In the decoding stage, each decoder contains 5 RSUs. The network structure of this RSU is the same as the RSU network structure in the encoder (i.e., corresponding to the first 5 RSUs in the decoder). At the same time, in the decoding stage, the connection of the upsampled feature map from the previous stage and the downsampled feature map from the symmetric encoder stage is used as the input to transmit context information.
[0071] It should be noted that the existing U2-Net network model is an efficient object detection network, but there are also some limitations. Its receptive field is relatively small, and when dealing with images with a large range of semantic information, there is sometimes a problem of insufficient receptive field, resulting in a lack of global consistency in the object extraction results. At the same time, there are problems of blurred boundaries and insufficient processing of fine structures in the image segmentation task.
[0072] Therefore, the improved U2-Net network model of this application is improved and optimized by adding a multi-scale attention module to the U2-Net network, that is, a multi-scale channel attention module is added after the first three encoder modules.
[0073] As Figure 4 shown, the multi-scale channel attention module is mainly composed of three parts: depth-wise convolution, multi-branch depth-wise convolution, and convolution. Among them, the first part of the depth-wise convolution layer is used to aggregate local information, the second part of the multi-branch depth-wise convolution is used to capture multi-scale context information, and the last part of the convolution is used to perform correlation modeling in the channel dimension. Moreover, the output of each part of the convolution will be used as an attention weight to re-weight the input.
[0074] Furthermore, it should be noted that the core idea of the first part of the depth-wise convolution is to decompose a complete convolution operation into two steps, namely separable convolution and pointwise convolution. Compared with ordinary convolution, for the same input, to obtain the same output feature map, the number of parameters used by the depth-wise convolution is about 1 / 3 of that of ordinary convolution. Therefore, using depth-wise convolution can greatly reduce the number of model parameters.
[0075] This application uses a multi-scale attention module to capture features at different spatial scales and fuse these features in the channel dimension, enhancing the model's feature expression ability at different scales. The improved and optimized network not only increases the receptive field but also strengthens the ability to extract features at different scales.
[0076] Correspondingly, the structure of the decoder is similar to that of the encoder. However, the input of the decoder is obtained by cascading the upsampled feature map of the previous-level feature and the symmetric encoder feature after being processed by the attention module.
[0077] Finally, 6 side output saliency maps of the decoder are obtained through 3×3 convolution, and after the 6 saliency maps are concatenated, the final prediction map is obtained through 1×1 convolution and sigmiod activation function; see as Figure 7As shown, the left figure is a false color image, the middle figure is the recognition effect before improvement, and the right figure is the recognition effect after improving the U2-Net network model of the present application. The large black part is the recognition result. It can be clearly seen that the boundary of the rightmost figure is clearer and more accurate.
[0078] Furthermore, the model training of the present application is preferably carried out based on the pytorch framework, and the model parameters are set as follows: the optimization function selects the Stochastic Gradient Descent (SGD) optimizer; the loss function selects the cross-entropy loss function; the initial value of the learning rate is set to 0.001, and with the iteration of the training times; the learning rate selects the warmup and cosine annealing strategies; the batch size is set to 4.
[0079] Step S4: Evaluate the accuracy of the improved U2-Net network model trained in step S3. If the accuracy requirement is met, proceed to step S5; if the accuracy requirement is not met, return to steps S2 and S3 to check the sample drawing and adjust and optimize the improved U2-Net network model until the requirement is met.
[0080] Specifically, the accuracy evaluation formula is as follows:
[0081]
[0082] Among them, P represents the accuracy rate, R represents the recall rate, F1 represents the accuracy of the winter wheat recognition model, TP represents the area of winter wheat correctly extracted by the model, FN represents the area of winter wheat missed by the model, and FP represents the area of winter wheat mis-extracted by the model.
[0083] In this step, use the trained network model to evaluate the accuracy on the validation set. If the F1 score reaches more than 0.9, start the prediction of winter wheat. If the accuracy requirement is not met, re-check the sample drawing, parameter settings, etc., and adjust and optimize the model until the requirement is met.
[0084] Step S5: Extract the winter wheat area.
[0085] Apply the trained improved U2-Net network model to Henan Province, predict the preprocessed images scene by scene, extract the distribution area of winter wheat, and then splice the extraction results of all scene images to obtain the winter wheat distribution map of the whole province as shown in Figure 8 In this embodiment, the winter wheat area in Henan Province is extracted to be 77.625 million mu.
[0086] It should be noted that in order to verify the feasibility of the method proposed in this application, two verification methods are adopted. The first verification method is shown in step S5, which uses evaluation indicators such as the F1 score and the intersection over union (IOU) based on the confusion matrix, and calculates the accuracy of the winter wheat model in Henan Province by using the above formulas (1)-(3). In this embodiment, the F1 value of the winter wheat extraction model is 0.96.
[0087] Another verification method (i.e., step S6) is to use an image with a higher spatial resolution (e.g., 3 meters), take the winter wheat result obtained by combining computer automatic classification and visual refinement as the ground truth, and then substitute it into the following formula (4) to calculate the winter wheat extraction error:
[0088]
[0089] Where P a represents the winter wheat extraction error, A 1 represents the area of winter wheat extracted by the model, and A 2 represents the ground truth area extracted by the model.
[0090] In this embodiment, as Figure 9 shown, Xun County in Henan Province is selected, and the winter wheat extraction error in Xun County is calculated to be -1.34% by using the above formula (4).
[0091] At the same time, to test the scalability of the method of this application, as Figure 10 shown, in Qiuxian County, Hebei Province, the winter wheat in Qiuxian County in 2023 is directly extracted by using the winter wheat model in Henan Province (i.e., without going through the steps of drawing winter wheat samples in Qiuxian County and model training). The results show that the area of winter wheat in Qiuxian County is 337,000 mu. Then, the winter wheat extraction accuracy in Qiuxian County, Hebei Province is calculated to be -4.48% by using formula (4).
[0092] The above results show that the technical method of this application is feasible, and the deep learning model with the improved U2-Net network structure can be used for the remote sensing extraction of winter wheat in Henan Province and its surrounding areas.
[0093] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in this application should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claimed rights.
Claims
1. A provincial winter wheat remote sensing identification method based on an improved U2-Net network model, characterized in that: The following steps are involved: S1. Obtaining an optical remote sensing image of winter wheat in a region to be measured, and preprocessing the optical remote sensing image, wherein the region to be measured is a provincial-level region; S2, selecting a predetermined number of original sample images from the preprocessed optical remote sensing images, dispersing the original sample images into a predetermined number of sample frames, and performing sliding window cropping processing, and then using an object-oriented method to perform image segmentation and complete sample drawing in combination with visual interpretation to obtain a sample data set; S3, augmenting the sample data set, and dividing the augmented sample data set into a training set and a validation set, then building an improved U2-Net network model, and using the training set to train the improved U2-Net network model, wherein the improved U2-Net network model adds a multi-scale channel attention module after the first three encoder modules in the original U2-Net network model, and the multi-scale channel attention module includes: Depth-wise convolution, which is used to aggregate local information, and its output will be used as attention weights to reweight the input; Multi-branch depth-wise convolution, which is used to capture multi-scale contextual information, and its output will be used as attention weights to reweight the input; Convolution, which is used to model correlation in the channel dimension, and its output will be used as attention weights to reweight the input; S4, using the verification set to evaluate the accuracy of the improved U2-Net network model trained in step S3, if the accuracy requirement is met, proceed to step S5, if not, return to steps S2 and S3, check sample drawing and adjust and optimize the improved U2-Net network model until the requirement is met; S5. Apply the improved U2-Net network model obtained in step S4 to the optical remote sensing image of winter wheat in the area to be measured, perform prediction scene by scene, extract the winter wheat distribution area, and then splice the extraction results of all scene images to obtain the winter wheat distribution map of the area to be measured.
2. The provincial winter wheat remote sensing identification method according to claim 1, characterized in that: In step S1, the optical remote sensing image is L2 level surface reflectance data, and the preprocessing of the optical remote sensing image includes at least band combination, stitching and cropping, wherein the optical remote sensing image after band combination contains at least five bands: blue, green, red, near red and short-wave infrared.
3. The provincial winter wheat remote sensing identification method according to claim 1, characterized in that: In the improved U2-Net network model, the multi-scale channel attention module is used to: The depth-wise convolution is used to perform a depth-separable convolution operation on the input features to retain spatial information and reduce the amount of calculation. The features are then input into a multi-branch depth-wise convolution. Each branch uses convolution kernels of different sizes to capture features of different scales. The outputs of multiple branches are then concatenated in the channel dimension to aggregate multi-scale features. A 1*1 convolution layer is then used to fuse and adjust the channels of the concatenated feature maps to generate the final output.
4. The provincial winter wheat remote sensing identification method according to claim 3 is characterized in that: In the improved U2-Net network model, each encoder and decoder uses a residual U block. Each residual U block uses a basic U-shaped structure, and adds a shortcut branch for residual signal transmission on the basis of the U-shaped structure, which is used to transmit the residual directly to the subsequent network when the network propagates the error.
5. The provincial winter wheat remote sensing identification method according to claim 4, characterized in that: In the improved U2-Net network model, the input of each decoder is obtained by cascading the feature upsampling and symmetric encoder features of the previous stage through the multi-scale channel attention module.
6. The provincial winter wheat remote sensing identification method according to claim 4, characterized in that: At the end of the model training in step S3, 3×3 convolution is performed to obtain the side output saliency maps of 6 decoders, and the 6 saliency maps are concatenated and then subjected to 1×1 convolution and sigmoid activation function to obtain the final prediction map.
7. The provincial winter wheat remote sensing identification method according to claim 4, characterized in that: When building the improved U2-Net network model in step S3, 6 residual U blocks are used in the encoding stage and 5 residual U blocks are used in the decoding stage. Among them, the first residual U block has one more layer of downsampling and upsampling processes on the basis of the second residual U block, and the third and fourth residual U blocks gradually reduce one layer of downsampling and upsampling processes on the basis of the second residual U block. The fifth and sixth residual U blocks use hollow convolution instead of pooling on the basis of the second residual U block to avoid detail loss caused by excessive downsampling. In addition, in the decoding stage, the connection of the upsampled feature map from the previous stage and the downsampled feature map from the symmetric encoder stage is used as input to pass context information.
8. The provincial winter wheat remote sensing identification method according to claim 1, characterized in that: In S3, the improved U2-Net network model is trained based on the pytorch framework, and the model parameters and the settings of each model parameter are respectively: The optimization function, which is set to the stochastic gradient descent optimizer; Loss function, which is set to the cross entropy loss function; The initial value of the learning rate is set to 0.001; Learning rate, which is set to warmup strategy and cosine annealing strategy; Batch size, which is set to 4.
9. The provincial winter wheat remote sensing identification method according to claim 1, characterized in that: In step S4, the model accuracy evaluation formula is as follows: Among them, P represents the precision, R represents the recall rate, F1 represents the accuracy of the winter wheat recognition model, TP represents the winter wheat area correctly extracted by the model, FN represents the winter wheat area missed by the model, and FP represents the winter wheat area incorrectly extracted by the model; Among them, when the F1 score is greater than or equal to 0.9, it is determined that the accuracy requirement is met.
10. The provincial winter wheat remote sensing identification method according to claim 1, characterized in that: Also includes: S6, extracting the winter wheat area using the winter wheat distribution map obtained in step S5, and verifying the extraction result by the following formula: Among them, P a represents the winter wheat extraction error, A1 represents the extracted winter wheat area, and A2 represents the true area of winter wheat, which is the winter wheat area obtained by using optical remote sensing images based on higher spatial resolution, computer automatic classification and visual refinement methods.
Citation Information
Patent Citations
Winter wheat remote sensing recognition analysis method and system based on deep learning
CN114463637A
Winter wheat planting area image extraction method combining GF-6 and Sentinel-2
CN114842339A