An economic crop information recognition method and system based on deep learning

By optimizing the structure and adjusting the parameters of the U-Net model, the ISDU-Net model was generated, which solved the problem of insufficient accuracy of UAV remote sensing imagery in the classification of economic crops and achieved efficient and accurate identification of economic crops.

CN117115685BActive Publication Date: 2025-11-18CHINA AGRI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310914018.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-24
Publication Date
2025-11-18
Estimated Expiration
2043-07-24

AI Technical Summary

Technical Problem

In existing technologies, the classification accuracy of UAV remote sensing images for economic crop classification and identification is limited by classical deep learning models, resulting in unsatisfactory classification results.

Method used

An improved deep learning semantic segmentation model, ISDU-Net, is adopted. By optimizing the structure and adjusting the parameters of the classic U-Net model, adding the Inception structure and the channel attention mechanism module SE-Block, and deepening the network, an improved deep learning semantic segmentation model is generated for the identification of economic crop information in UAV remote sensing images.

Benefits of technology

This improved the accuracy and extraction efficiency of UAV remote sensing imagery in the classification of economic crops, enhanced the semantic segmentation model's recognition performance on small samples, and enabled scientific, convenient, and efficient identification of economic crops.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117115685B_ABST
    Figure CN117115685B_ABST
Patent Text Reader

Abstract

The application provides an economic crop information recognition method and system based on deep learning, comprising: cutting the unmanned aerial vehicle multispectral remote sensing image of a research area according to a prescribed shape to obtain a regular shape image; sample labeling is performed on the regular shape image, and the regular shape image is divided into a training set, a verification set and a test set; the U-Net model is trained according to the training set and the verification set; the trained U-Net model is used as a basic model, an Inception structure and an SE-Block are added, and the network is deepened; the hyperparameters are selected to optimize the parameters of the network-deepened model, an improved deep learning semantic segmentation ISDU-Net model is generated, and the ISDU-Net model is trained based on the training set and the verification set; and the trained ISDU-Net model is used for information recognition of economic crops. The improved deep learning semantic segmentation model is used for extracting unmanned aerial vehicle remote sensing images of economic crops and completing rapid mapping, and is outstanding in small sample category prediction, thereby providing a new idea for accurate recognition of economic crops.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing, and in particular to a method and system for identifying economic crop information based on deep learning. Background Technology

[0002] Cash crops are an important component of agriculture, possessing high economic value and ecological benefits. They serve as the material foundation for national economic and social development while also positively impacting the improvement of the ecological environment. my country has a vast territory with diverse topography and soil types. Some areas have steep slopes, relatively infertile soil, and harsh ecological environments, making them unsuitable for food crops, but providing a good foundation for cultivating economic fruit forests. Cash crops play a significant role in ensuring national ecological security, enriching forest product supply, increasing rural income, and improving quality of life; therefore, accurate classification of cash crops is crucial. However, traditional ground surveys of cash crops are time-consuming and labor-intensive. The advent of remote sensing imagery has made rapid identification of cash crops possible.

[0003] With the gradual development of technology, UAV remote sensing technology has become increasingly mature and is rapidly being applied in various fields, becoming one of the new methods for acquiring remote sensing images. UAV remote sensing has great potential, mainly due to its advantages of flexibility, high efficiency, and portability. It is being used in agricultural fields such as crop identification and resource surveys, gradually becoming a mainstay in the field of identification and classification. Lightweight design and intelligent operation significantly reduce manpower and material resources, making image acquisition more convenient and offering greater advantages compared to traditional satellite remote sensing images. In recent years, deep learning has gradually developed and matured, becoming a popular research direction in image processing tasks and finding wide application in various fields. Deep learning models do not rely on manual feature selection; by automatically extracting features and calculating class probabilities, they greatly improve classification efficiency. In UAV remote sensing image classification research, the classification accuracy has been further improved after adopting deep learning algorithms. Currently, some scholars are researching ways to improve deep learning models to process remote sensing image classification more efficiently, providing new ideas for UAV and other types of remote sensing image data processing, with classification results significantly improved compared to classical models.

[0004] While existing research has explored various methods for using UAV remote sensing technology in classification and identification, current research on the classification and identification of economic crops based on UAV remote sensing imagery is limited, and classification accuracy is constrained by classical deep learning models. Therefore, this invention selects economic crops as the research object, aiming to develop an improved deep learning model to study the characteristic differences between different economic crops, improve the classification accuracy and extraction efficiency of economic fruit orchards from UAV remote sensing imagery, improve the recognition performance of semantic segmentation models on small samples, and research a method for identifying economic crop information based on UAV remote sensing imagery data, providing a reference for the scientific, convenient, and efficient identification of the spatial distribution information of economic crops. Summary of the Invention

[0005] The purpose of this invention is to provide a method and system for identifying economic crop information based on deep learning, in order to solve the problem that the classification accuracy of existing technologies for classifying and identifying economic crops using UAV remote sensing images is limited by classical deep learning models, resulting in unsatisfactory classification results.

[0006] The present invention provides a deep learning-based method for identifying economic crop information, comprising:

[0007] The UAV multispectral remote sensing images of the study area are cropped according to the specified shape to obtain images with regular shapes.

[0008] The regular shape image is labeled to form sample data, and the sample data is divided into training set, validation set and test set;

[0009] The deep learning semantic segmentation model is trained based on the training set and validation set. The deep learning semantic segmentation model is the classic semantic segmentation model U-Net.

[0010] The trained U-Net model is used as the base model. An Inception structure and a channel attention mechanism module SE-Block are added, and the network is deepened to obtain a deep learning semantic segmentation model with a deeper network.

[0011] The hyperparameters are selected to optimize the parameters of the deep learning semantic segmentation model after the network is deepened, and the improved deep learning semantic segmentation model ISDU-Net is generated. The ISDU-Net model is trained based on the training set and validation set to obtain the trained ISDU-Net model.

[0012] The trained ISDU-Net model was used to identify information from multispectral remote sensing images of economic crops in the area to be detected.

[0013] According to the deep learning-based economic crop information identification method provided by the present invention, the step of cropping the UAV multispectral remote sensing image of the study area according to a prescribed shape to obtain a regular-shaped image includes:

[0014] Acquire UAV multispectral remote sensing images of the study area;

[0015] Automatic radiometric correction, deformation correction and image stitching are performed on the multispectral remote sensing images;

[0016] Using ArcGIS band compositing tools, the five single-band images of red, green, blue, red edge, and near-infrared were combined into a five-band image.

[0017] Using a raster data cropping tool, the study area was cropped by quadrat to obtain images with regular shapes.

[0018] According to the deep learning-based economic crop information identification method provided by the present invention, the step of annotating the regular shape image to form sample data, and dividing the sample data into a training set, a validation set, and a test set includes:

[0019] The obtained regular-shaped image is cropped and bit-transformed.

[0020] Sample annotation is performed on the image data after cropping and bit transformation.

[0021] The image data after sample annotation will be randomly divided into training set, validation set and test set according to a predetermined ratio.

[0022] According to the deep learning-based economic crop information identification method provided by the present invention, the sample annotation specifically includes:

[0023] A classification system was determined based on the types of economic crops in the study area, and RGB labels were assigned to each type.

[0024] Based on the information from actual sampling points, the boundaries of each crop were drawn and colored to create a label map, resulting in labeled samples.

[0025] According to the deep learning-based economic crop information identification method provided by the present invention, after randomly dividing the image data after sample annotation into a training set, a validation set, and a test set according to a predetermined ratio, the method further includes:

[0026] Data augmentation is performed on the training set, validation set, and test set through geometric transformations to obtain the data-augmented training set, validation set, and test set.

[0027] According to the deep learning-based economic crop information recognition method provided by the present invention, the step of using the trained U-Net model as the base model, adding an Inception structure and a channel attention mechanism module SE-Block, and deepening the network to obtain a deep learning semantic segmentation model with a deepened network, specifically includes:

[0028] The U-Net model is used as the base model, which includes a model encoding part and a model decoding part.

[0029] In the five-layer downsampling structure of the U-Net model encoding part, the first three layers each set an Inception structure, and the last two layers set three convolutional layers, three BN layers, three ReLU activation layers and one pooling layer to obtain the model encoding part after network deepening. The Inception module contains four parts, namely 2D convolution with 1*1, 3*3 and 5*5 convolutional kernels and a 3*3 kernel max pooling layer. Activation layer and BN layer operations are performed during the convolution.

[0030] In the five-layer upsampling structure of the U-Net model decoding part, a channel attention mechanism module SE-Block is set before each layer upsampling to obtain the model decoding part after network deepening. The channel attention mechanism module SE-Block is set with two parts: compression Squeeze and excitation. The improved model decoding part adopts the form of transposed convolution, decreases the number of filters layer by layer, and uses high and low level feature fusion to fuse with the feature layers obtained from the encoding part, and then uses 2D convolution for feature extraction.

[0031] Based on the improved encoding and decoding parts, a deep learning semantic segmentation model with enhanced network depth is generated.

[0032] According to the deep learning-based economic crop information recognition method provided by the present invention, the channel attention mechanism module's processing of the input image includes:

[0033] Based on the input UAV multispectral remote sensing image, global average pooling is performed through the compression part to obtain the global feature information of each channel of the multispectral remote sensing image, and the global feature information is then passed to the excitation part.

[0034] Based on the global feature information input from the compression part, the excitation part performs a nonlinear transformation on the global feature information through two fully connected layers to obtain the weight value of each channel of the multispectral remote sensing image. The weight value is then assigned to the global feature information of the multispectral remote sensing image to obtain the original features of each channel of the multispectral remote sensing image.

[0035] The compression portion uses a global average pooling method, defined as follows:

[0036]

[0037] Where H is the feature layer height, W is the feature layer width, and u c The expression for the output layer result obtained from convolution is as follows:

[0038]

[0039] Among them, v c This represents the c-th convolution kernel, s represents the number of channels, and * represents convolution calculation.

[0040] According to the deep learning-based economic crop information recognition method provided by the present invention, the step of selecting hyperparameters to optimize the parameters of the deep learning semantic segmentation model after the network is deepened includes: selecting two hyperparameters, learning rate and batch size, to optimize the parameters of the improved deep learning semantic segmentation model.

[0041] According to the deep learning-based economic crop information identification method provided by the present invention, after training the ISDU-Net model based on the training set and validation set to obtain the trained ISDU-Net model, the method further includes:

[0042] The trained ISDU-Net model is used to predict the test set to obtain the prediction results for the test set.

[0043] Based on the prediction results, the accuracy of the trained ISDU-Net model is evaluated using the confusion matrix, pixel accuracy, F1 score, crossover ratio, and Kappa coefficient.

[0044] This invention also provides a deep learning-based economic crop information identification system, comprising:

[0045] The image cropping module is used to crop the UAV multispectral remote sensing images of the study area according to a specified shape to obtain images with regular shapes.

[0046] The sample annotation module is used to annotate the regular shape image to form sample data, and divide the sample data into training set, validation set and test set.

[0047] The model training module is used to train the deep learning semantic segmentation model based on the training set and the validation set. The deep learning semantic segmentation model is the classic semantic segmentation model U-Net.

[0048] The structure optimization module is used to take the trained U-Net model as the base model, add the Inception structure and the channel attention mechanism module SE-Block, and deepen the network to obtain a deep learning semantic segmentation model with a deeper network.

[0049] The parameter optimization module is used to select hyperparameters to optimize the parameters of the deep learning semantic segmentation model after the network is deepened, and generate the improved deep learning semantic segmentation model ISDU-Net model. The ISDU-Net model is trained based on the training set and validation set to obtain the trained ISDU-Net model.

[0050] The information recognition module is used to perform information recognition on the multispectral remote sensing images of economic crops in the area to be detected using the trained ISDU-Net model.

[0051] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the deep learning-based economic crop information recognition method described above.

[0052] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the deep learning-based economic crop information identification method described above.

[0053] This invention selects economic crops as the research object and develops an improved deep learning semantic segmentation model based on the classic deep learning semantic segmentation model. This invention is mainly applied to study the feature differences between different economic crops, improving the classification accuracy and extraction efficiency of economic crops from UAV remote sensing imagery, and enhancing the recognition performance of the semantic segmentation model for small samples. The economic crop information identification method based on UAV remote sensing imagery data provided by this invention offers a reference for the scientific, convenient, and efficient identification of the spatial distribution information of economic crops, while also enabling automated identification and mapping, thereby extracting economic crops more accurately. Attached Figure Description

[0054] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0055] Figure 1 This is a flowchart illustrating the deep learning-based economic crop information identification method provided in an embodiment of the present invention.

[0056] Figure 2 This is a schematic diagram of structural optimization based on the U-Net model provided in an embodiment of the present invention;

[0057] Figure 3 This is a schematic diagram of the structure of the deep learning-based economic crop information identification system provided in an embodiment of the present invention;

[0058] Figure 4 This is a schematic diagram of the structure of the electronic device provided in an embodiment of the present invention. Detailed Implementation

[0059] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0060] This invention provides a deep learning-based method for identifying economic crops. First, the acquired UAV remote sensing images are preprocessed to create a sample set, which is then divided into training, validation, and test sets. By amplifying the samples, a dataset of economic crops for the study area is established. A classic semantic segmentation model is selected for initial classification of economic fruit orchards in the study area. After accuracy evaluation, the model U-Net, which has the highest classification accuracy and best performance, is chosen. The basic network is then deepened and widened to obtain an improved deep learning semantic segmentation model, ISDU-Net. The ISDU-Net model is then used to identify information from multispectral remote sensing images of economic crops, and the recognition accuracy is evaluated. The evaluation results show that the improved deep learning semantic segmentation model, ISDU-Net, significantly improves the classification and identification performance of economic crops.

[0061] The following is combined Figures 1-4 This invention describes the method, system, electronic device, and storage medium for identifying economic crop information based on deep learning.

[0062] Figure 1 This is a flowchart illustrating the deep learning-based economic crop information identification method provided by the present invention, as shown below. Figure 1 As shown in one specific embodiment, the present invention provides a method for identifying economic crop information based on deep learning, comprising the following steps:

[0063] Step S110: Crop the UAV multispectral remote sensing image of the study area according to the specified shape to obtain a regular shape image;

[0064] Step S120: Label the regular shape image to form sample data, and divide the sample data into training set, validation set and test set;

[0065] Step S130: Train the deep learning semantic segmentation model based on the training set and validation set. The deep learning semantic segmentation model is the classic semantic segmentation model U-Net.

[0066] Step S140: Using the trained U-Net model as the base model, add the Inception structure and the channel attention mechanism module SE-Block, and deepen the network to obtain the deep learning semantic segmentation model with deepened network.

[0067] Step S150: Select hyperparameters to optimize the parameters of the deep learning semantic segmentation model after the network is deepened, generate the improved deep learning semantic segmentation model ISDU-Net, and train the ISDU-Net model based on the training set and validation set to obtain the trained ISDU-Net model.

[0068] Step S160: Use the trained ISDU-Net model to identify information from the multispectral remote sensing images of economic crops in the area to be detected.

[0069] The following is about Figure 1 The steps in the text are explained in detail below:

[0070] Step S110: Crop the UAV multispectral remote sensing image of the study area according to the specified shape to obtain a regular shape image;

[0071] In this embodiment of the invention, the preprocessing of multispectral remote sensing images of economic crops to generate a regularly shaped image of the study area cropped according to a regular shape includes:

[0072] Acquire UAV multispectral remote sensing images of the study area;

[0073] Automatic radiometric correction, deformation correction and image stitching are performed on the multispectral remote sensing images;

[0074] Using ArcGIS band compositing tools, the five single-band images of red, green, blue, red edge, and near-infrared were combined into a five-band image.

[0075] Using a raster data cropping tool, the study area was cropped by quadrat to obtain images with regular shapes.

[0076] Specifically, the research area selected in this embodiment of the invention is an agricultural economic forest with a total area of ​​approximately 700 mu (about 46.7 hectares). This area has a variety of fruit trees, planted in strips, with main economic crops including grapes, peaches, apples, and cherries. Research on this area is of scientific value.

[0077] In this embodiment, UAV multispectral remote sensing images of a typical study area are first acquired and then preprocessed. The preprocessing process includes:

[0078] Acquiring multispectral remote sensing imagery from drones involves using drones to conduct aerial photography in the field to obtain high-resolution multispectral remote sensing images. Handheld GPS devices are used to select and record sample points, and the acquired image data is preprocessed. In this embodiment, a quadcopter drone is used to acquire images of the study area, resulting in multispectral remote sensing imagery of the study area.

[0079] Using ZhiTu software, automatic radiometric correction, deformation correction, and image stitching are performed on the obtained multispectral remote sensing images. The image stitching process includes: using images taken by multispectral UAVs containing geographic coordinate data, lens error data, DEM data, and radiometric data, etc., ZhiTu software is used to automatically perform radiometric correction, deformation correction, and image stitching to generate high-quality orthophotos of the study area in visible light and single bands, and save them in TIFF format.

[0080] Using ArcGIS's band composite tool, the five single-band images (red, green, blue, red-edge, and near-infrared) are combined into a five-band image. Since the orthophoto images are single-band images after stitching, ArcGIS's band composite tool is used to combine the five single-band images into a five-band image containing red, green, blue, red-edge, and near-infrared bands, and saves it in TIFF format for subsequent model training.

[0081] Secondly, the preprocessed image data undergoes further processing, including:

[0082] The preprocessed image is then cropped and bit-shifted.

[0083] In this embodiment, due to the irregular shape of the study area, in order to produce the subsequent dataset, the study area was cropped by quadrat using the raster data cropping tool of ArcGIS software to obtain 16 regular rectangular images with a size of 2136*2629 pixels, which were then stored in TIFF format.

[0084] Sixteen 32-bit images were batch converted to 8-bit. Since the obtained UAV multi-band orthophoto images are 32-bit, and the images need to be normalized before training the deep learning model, a Python script was written to batch convert the 16 32-bit images to 8-bit to facilitate normalization and reduce subsequent work.

[0085] Step S120: Label the regular shape image to form sample data, and divide the sample data into training set, validation set and test set;

[0086] In this embodiment of the invention, the step of annotating the regular shape image to form data samples, and dividing the data samples into a training set and a test set, includes:

[0087] First, a dataset of economic crops is established, and then the dataset is expanded. The specific steps for expanding and creating the dataset include:

[0088] The obtained regular-shaped image is cropped and bit-transformed.

[0089] Sample annotation is performed on the image data after cropping and bit transformation.

[0090] The image data after sample annotation will be randomly divided into training set, validation set and test set according to a predetermined ratio.

[0091] In this embodiment, the sample labeling specifically includes:

[0092] A classification system was determined based on the types of economic crops in the study area, and RGB labels were assigned to each type.

[0093] Based on the information from actual sampling points, the boundaries of each crop were drawn and colored to create a label map, resulting in labeled data samples, which were then saved as TIFF files.

[0094] In this embodiment, the training set, validation set, and test set are randomly divided in an 8:2 ratio.

[0095] After dividing the data samples into training set, validation set and test set, the method further includes: performing data augmentation on the training set, validation set and test set through geometric transformation to obtain augmented training set, validation set and test set.

[0096] Specifically, deep learning models often rely on large amounts of sample data, thus requiring data augmentation to obtain a large number of training samples to further improve the model's classification accuracy. Common data augmentation methods include geometric transformations, color transformations, and noise augmentation / reduction. Geometric transformations typically include translation, rotation, flipping, cropping, and scaling; color transformations involve adjusting color components in different color spaces to change the image's color; noise augmentation / reduction involves blurring the image, adding noise, or covering specific regions. Existing research shows that color transformations and noise augmentation / reduction do not perform well in semantic segmentation tasks in the remote sensing field, and may even lead to the erasure of local information needed by the model, reducing the model's learning ability. Therefore, this embodiment chooses geometric transformations for data augmentation.

[0097] In one specific embodiment, data augmentation is performed on the training set, validation set, and test set through geometric transformation, specifically including:

[0098] The 16 preprocessed images were divided into an 8:2 ratio. Two images were randomly selected as the test set for subsequent prediction and accuracy evaluation, while the remaining 14 images were used as the training and validation sets for model training. The 16 images were horizontally flipped, vertically flipped, and diagonally mirrored to obtain data augmentation. The augmented dataset contained 17,985 images, which were then randomly divided into training and validation sets in an 8:2 ratio.

[0099] Step S130: Train the deep learning semantic segmentation model based on the training set and validation set. The deep learning semantic segmentation model is the classic semantic segmentation model U-Net.

[0100] Specifically, the first step is to perform an initial classification of cash crops using a classic deep learning model. Since there are many existing deep learning semantic segmentation models for initial classification of cash crops, this embodiment selects several classic models and, through specific evaluation and comparative analysis, chooses the model with the best classification performance as the foundational model for this invention.

[0101] In this embodiment, a classic deep learning semantic segmentation model is selected as the initial classification model for extracting economic crops; and commonly used classic semantic segmentation models FCN, SegNet, and U-Net are selected as the basic classification models.

[0102] Furthermore, FCN, SegNet, and U-Net models were used to predict specific sample sets of data. The accuracy of the prediction results was evaluated, and the model with the highest accuracy was selected. This embodiment uses pixel accuracy, F1 score, intersection-over-union ratio (IoU), and Kappa coefficient as the evaluation metrics for model prediction accuracy, namely:

[0103]

[0104] In the formula, PA represents pixel accuracy;

[0105]

[0106]

[0107]

[0108] In the formula, CPA is the class pixel precision, Recall is the recall rate, and F1 is the F1 score;

[0109]

[0110] In the formula, MIoU is the crossover-union ratio;

[0111]

[0112] The Kappa coefficient in the formula.

[0113] Calculated using the above formulas, the pixel accuracy, average crossover ratio (CRO), frequency weight CRO, and Kappa coefficient of the U-Net model are 87.73%, 70.68%, 78.69%, and 0.84, respectively, making it the best among the three models. Through comparative analysis of accuracy evaluation, the U-Net network model was determined to have the best classification performance and high overall accuracy. Therefore, the U-Net network structure was selected as the basic network model.

[0114] Step S140: Using the trained U-Net model as the base model, add the Inception structure and the channel attention mechanism module SE-Block, and deepen the network to obtain the deep learning semantic segmentation model with deepened network.

[0115] Specifically, the process involves adding an Inception structure and a channel attention mechanism module SE-Block to the trained U-Net model, and then deepening the network to generate a deep learning semantic segmentation model with a deepened network. In essence, this is a structural optimization of the U-Net model.

[0116] Figure 2 This is a schematic diagram of the structural optimization based on the U-Net model provided by the present invention, such as... Figure 2 As shown, the specific process of structural optimization based on the U-Net model includes:

[0117] Using the U-Net model as the base model, the U-Net model is widened and deepened, including adding an Inception structure, a channel attention mechanism module SE-Block, and network deepening;

[0118] The U-Net model includes a model encoding part and a model decoding part. The model encoding part increases the number of filters layer by layer through 2D convolution to form a five-layer downsampling structure. The first three layers of the downsampling structure are each set as an Inception structure, and the last two layers are each set as a structure containing three convolutional layers, three batch normalization layers, three ReLU activation layers, and one pooling layer. The last two layers are both set as cubic convolutional structures. The model decoding part decreases the number of filters layer by layer through transposed convolution. It is fused with the feature layers obtained from the encoding part using a high-low level feature fusion method, and features are extracted using 2D convolution. A five-layer upsampling structure is adopted, and a channel attention mechanism module SE-Block is added before each upsampling layer.

[0119] Specifically, the Inception module's main principle is based on the sparse connectivity characteristics of biological neural systems. It replaces fully connected layers with sparse connections, increasing the network's depth and width without increasing computational cost through parallel convolutional structures and smaller kernels. The Inception module consists of four parallel convolutional structures: 1x1, 3x3, and 5x5 convolutions, and a 3x3 max pooling operation. For example, since the 5x5 convolution kernel requires significant computation, the Inception structure adds a 1x1 convolution operation before the 3x3 and 5x5 convolutions and after max pooling. This reduces the dimensionality of the feature layers, decreasing computational cost while simultaneously increasing the network's width and depth. Finally, concatenation merges the four results, ensuring feature diversity and thus enhancing the network's learning ability.

[0120] In this embodiment, the Inception module uses the ReLU activation function in each convolutional layer and adds a Batch Normalization (BN) layer to accelerate network training and convergence, and prevent gradient vanishing and overfitting. Since the Inception module can obtain more diverse sample feature information, this experiment replaces the first three downsampling layers of the original U-Net model with the Inception module, increasing the width of the downsampling portion to improve the model's learning ability for small sample types. This addresses the imbalance problem of some class samples, broadens the model's receptive field, obtains more effective information, and improves the prediction accuracy of orchard classification.

[0121] Understandably, deepening a network generally enhances its learning ability, but this isn't always the case. A deeper network requires more data, otherwise it can lead to overfitting. Therefore, appropriately deepening the model can improve its expressive power. This invention, based on the original U-Net model, appropriately deepens the downsampling part to improve its feature extraction capabilities. The basic U-Net model's downsampling part contains two convolutions per layer. In this invention, besides using an Inception structure instead of ordinary convolutions in the first three downsampling layers to make the network deeper and wider, the last two downsampling layers are also appropriately adjusted, changing two 2D convolutions into three, thereby improving the model's learning ability, obtaining more feature information, and achieving better classification results.

[0122] On the other hand, the channel attention mechanism module SE-Block, or Squeeze-and-Excitation module, originated from the convolutional neural network model SENet. In ordinary convolutional pooling, the importance of each channel is generally considered equal, that is, all channels of each feature map are fused. However, in reality, the importance of different channels to the classification result is often not the same. SE-Block is a module that focuses on the relationship and importance of different channels. Through SE-Block, the model can learn the importance of features of different channels, thereby making better judgments.

[0123] The channel attention mechanism module SE-Block consists of two parts: compression (Squeeze) and excitation (Excitation). In the compression part, global average pooling is performed, and the network can learn channel-level global features. These features are then fed into the excitation part, where two fully connected layers are used to perform a nonlinear transformation on the results obtained from the compression part to obtain the weight values ​​for each channel, which are then assigned to the feature layer.

[0124] Since convolution is a feature extraction operation performed only within a local space, it is usually difficult to obtain enough information to extract the relationships between channels. To expand the receptive field, the compression part encodes the entire spatial features of a channel into global features to meet the receptive field requirements. Therefore, global average pooling is chosen to achieve this purpose. The compression part is defined as follows:

[0125]

[0126] Where H is the feature layer height, W is the feature layer width, and u c The expression for the output layer result obtained from convolution is as follows:

[0127]

[0128] Among them, v c This represents the c-th convolution kernel, s represents the number of channels, and * represents convolution calculation.

[0129] The excitation portion specifically includes:

[0130] Receive the channel global feature information obtained by performing global average pooling on the compression section;

[0131] By using two fully connected layers, a nonlinear transformation is performed on the global feature information to obtain the weight coefficients of each channel;

[0132] Based on the weight coefficients of each channel, the original features of each channel are obtained.

[0133] In detail, the global features obtained through compression require another computational method to obtain the relationships between channels. This operation needs to flexibly learn the non-linear relationships between each channel, and the learned relationships are not mutually exclusive, so as to ensure that multi-channel features exist simultaneously. Therefore, an activation part was designed. The activation part first uses a fully connected layer for dimensionality reduction, which can improve the network's generalization ability and reduce model complexity. Then, it is activated by an activation function, and then restored to the original dimension by a fully connected layer.

[0134] The relationship between channels is obtained using the Gating mechanism, the principle of which is as follows:

[0135] s = F ex (z,W)=σ(g(z,W))=σ(W2ReLU(W1z))

[0136] In the formula, ReLU is the activation function.

[0137] The learned channel weight coefficients are then assigned to the original features of each channel, as shown in the following formula:

[0138]

[0139] This makes the model more able to distinguish the features of each channel, which is a channel attention mechanism.

[0140] In this embodiment of the invention, since the original images, based on a high-resolution multispectral remote sensing image dataset from a UAV, involve many bands, the resulting feature layers are complex after the initial Inception structure and deeper multi-layer convolutions. Therefore, when performing feature fusion in the decoding part, using SE-Block to consider channel importance before fusion can achieve better classification results, consider more comprehensive channel information, and improve model learning efficiency.

[0141] Step S150: Select hyperparameters to optimize the parameters of the deep learning semantic segmentation model after the network is deepened, generate the improved deep learning semantic segmentation model ISDU-Net, and train the ISDU-Net model based on the training set and validation set to obtain the trained ISDU-Net model.

[0142] Specifically, the selection of hyperparameters to optimize the parameters of the deep learning semantic segmentation model after network deepening is essentially an adjustment of the parameters of the U-Net model after structural optimization.

[0143] In detail, the parameter tuning of the structurally optimized U-Net model includes:

[0144] In this embodiment, based on the structurally optimized U-Net model, two hyperparameters, learning rate and batch size, are selected for parameter optimization, resulting in an improved U-Net model. It is understood that the accuracy of deep learning models, besides being affected by the model structure, largely depends on the adjustment of model parameters. Deep learning model parameters are generally divided into two categories: model parameters and hyperparameters. Model parameters are adjusted based on data; through continuous iterative learning, the model adjusts the weights within the model using the acquired data information. Hyperparameters, on the other hand, do not depend on data and are directly adjusted through human intervention. By setting initial values ​​and adjusting the model, then learning again, the optimal parameters are experimentally determined. This invention adjusts parameters based on the structurally optimized model. Since hyperparameter adjustment does not depend on data, this embodiment improves the structurally optimized model by adjusting model hyperparameters, primarily focusing on adjusting two common hyperparameters: learning rate and batch size.

[0145] In this embodiment of the invention, with other parameters remaining the same, the network was trained and tested with initial learning rates of 0.01, 0.001, 0.0001, and 0.00001, and finally, the initial learning rate of 0.0001 was determined to be the optimal setting.

[0146] This embodiment of the invention also conducted a comparative study on batch size settings. With memory allowing, and other parameters remaining unchanged, this embodiment compared batch sizes of 4, 6, 8, and 10, determining that a batch size of 8 yielded the optimal model performance.

[0147] Furthermore, after adjusting the parameters of the structurally optimized U-Net model, an improved deep learning semantic segmentation model ISDU-Net is generated. The ISDU-Net model is then trained based on the training set and validation set to obtain the trained ISDU-Net model.

[0148] After optimizing the structure of the U-Net model and adjusting the parameters of the optimized U-Net model to generate an improved deep learning semantic segmentation model, the improved deep learning semantic segmentation model is used to identify information from the multispectral remote sensing images of economic crops in the area to be detected.

[0149] After identifying information from multispectral remote sensing images of economic crops, it is also necessary to evaluate the recognition accuracy of the improved deep learning semantic segmentation model.

[0150] In this embodiment, the accuracy of the improved deep learning semantic segmentation model can be evaluated using the confusion matrix, pixel accuracy, F1 score, intersection-union ratio, and Kappa coefficient mentioned above.

[0151] Step S160: Use the trained ISDU-Net model to identify information from the multispectral remote sensing images of economic crops in the area to be detected.

[0152] The deep learning-based economic crop information identification method provided by this invention utilizes a UAV multispectral remote sensing source and an improved U-Net deep learning model to achieve remote sensing extraction of economic crops, achieving high accuracy and strong reliability. The classification model proposed in this invention shows a significant improvement in accuracy for identification of small samples compared to classic models, effectively solving the problem of insufficient samples for a certain category. It has the potential for spatiotemporal analysis of economic crops and processing of UAV remote sensing data, and can provide practical guidance for the management and allocation of economic crop resources.

[0153] The present invention also provides a deep learning-based economic crop information recognition system, including an image cropping module, a sample annotation module, a model training module, a structure optimization module, a parameter optimization module, and an information recognition module.

[0154] Figure 3 This is a schematic diagram of the structure of the deep learning-based economic crop information identification system provided in an embodiment of the present invention, as shown below. Figure 3 As shown, the deep learning-based economic crop information identification system provided by the present invention includes:

[0155] The image cropping module 310 is used to crop the UAV multispectral remote sensing image of the study area according to a specified shape to obtain a regular-shaped image;

[0156] The sample annotation module 320 is used to annotate the regular shape image to form sample data, and divide the sample data into a training set, a validation set and a test set.

[0157] The model training module 330 is used to train the deep learning semantic segmentation model based on the training set and the validation set. The deep learning semantic segmentation model is the classic semantic segmentation model U-Net.

[0158] The structure optimization module 340 is used to take the trained U-Net model as the base model, add the Inception structure and the channel attention mechanism module SE-Block, and deepen the network to obtain a deep learning semantic segmentation model with a deeper network.

[0159] The parameter optimization module 350 is used to select hyperparameters to optimize the parameters of the deep learning semantic segmentation model after the network is deepened, generate the improved deep learning semantic segmentation model ISDU-Net, and train the ISDU-Net model based on the training set and validation set to obtain the trained ISDU-Net model.

[0160] The information recognition module 360 ​​is used to perform information recognition on the multispectral remote sensing images of economic crops in the area to be detected using the trained ISDU-Net model.

[0161] The economic crop information recognition system based on deep learning provided by this invention includes: an image cropping module that crops UAV multispectral remote sensing images of a study area according to a specified shape to obtain regularly shaped images; a sample annotation module that annotates the regularly shaped images to form sample data, which is then divided into training, validation, and test sets; a model training module that trains a deep learning semantic segmentation model, specifically the classic U-Net model, based on the training and validation sets; and a structure optimization module that uses the trained U-Net model as a base model and adds an Inception structure and a channel attention mechanism module (SE-Blo) to it. The network is deepened to obtain a deep learning semantic segmentation model. The parameter optimization module selects hyperparameters to optimize the parameters of the deep learning semantic segmentation model after network deepening, generating an improved deep learning semantic segmentation model ISDU-Net. The ISDU-Net model is trained based on the training set and validation set to obtain the trained ISDU-Net model. The information recognition module uses the trained ISDU-Net model to perform information recognition on the multispectral remote sensing images of economic crops in the detection area, which solves the problem that the classification accuracy is limited by the classic deep learning model when classifying and recognizing UAV remote sensing images of economic crops, resulting in unsatisfactory classification results.

[0162] This invention addresses the shortcomings of existing research by providing a deep learning-based method for identifying economic crop information. By selecting economic crops as the research object, and building upon the classic deep learning semantic segmentation model, an improved deep learning model is developed to study the characteristic differences between different economic crops. This method improves the accuracy and extraction efficiency of UAV remote sensing imagery in economic crop classification, enhances the recognition performance of the semantic segmentation model on small samples, and provides a reference for the scientific, convenient, and efficient identification of the spatial distribution information of economic crops. Furthermore, it enables automated identification and mapping, and can extract economic crops more accurately.

[0163] Figure 4 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 4As shown, the electronic device may include: a processor 410, a communication interface 820, a memory 430, and a communication bus 440, wherein the processor 410, the communication interface 420, and the memory 430 communicate with each other through the communication bus 440. The processor 410 can call logical instructions in the memory 430 to execute a deep learning-based economic crop information recognition method. This method includes: cropping a UAV multispectral remote sensing image of the study area according to a prescribed shape to obtain a regular-shaped image; labeling the regular-shaped image to form sample data, and dividing the sample data into a training set, a validation set, and a test set; training a deep learning semantic segmentation model, which is a classic semantic segmentation model, U-Net, based on the training and validation sets; using the trained U-Net model as a base model, adding an Inception structure and a channel attention mechanism module (SE-Block), and deepening the network to obtain a deepened deep learning semantic segmentation model; selecting hyperparameters to optimize the parameters of the deepened deep learning semantic segmentation model, generating an improved deep learning semantic segmentation model, ISDU-Net; training the ISDU-Net model based on the training and validation sets to obtain a trained ISDU-Net model; and using the trained ISDU-Net model to perform information recognition on the multispectral remote sensing image of economic crops in the area to be detected.

[0164] Furthermore, the logical instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0165] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements the above-described deep learning-based economic crop information recognition method. This method includes: cropping UAV multispectral remote sensing images of a study area according to a predetermined shape to obtain regular-shaped images; labeling the regular-shaped images to form sample data, and dividing the sample data into a training set, a validation set, and a test set; training a deep learning semantic segmentation model based on the training set and validation set, wherein the deep learning semantic segmentation model is the classic semantic segmentation model U-Net; using the trained U-Net model as a base model, adding an Inception structure and a channel attention mechanism module SE-Block, and deepening the network to obtain a deepened deep learning semantic segmentation model; selecting hyperparameters to optimize the parameters of the deepened deep learning semantic segmentation model, generating an improved deep learning semantic segmentation model ISDU-Net; training the ISDU-Net model based on the training set and validation set to obtain a trained ISDU-Net model; and using the trained ISDU-Net model to perform information recognition on the multispectral remote sensing images of economic crops in the area to be detected.

[0166] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0167] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0168] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for identifying economic crop information based on deep learning, characterized in that, include: The UAV multispectral remote sensing images of the study area are cropped according to the specified shape to obtain images with regular shapes. The regular shape image is labeled to form sample data, and the sample data is divided into training set, validation set and test set; The deep learning semantic segmentation model is trained based on the training set and validation set. The deep learning semantic segmentation model is the classic semantic segmentation model U-Net. The trained U-Net model is used as the base model. An Inception structure and a channel attention mechanism module SE-Block are added, and the network is deepened to obtain a deep learning semantic segmentation model with a deeper network. The hyperparameters are selected to optimize the parameters of the deep learning semantic segmentation model after the network is deepened, and the improved deep learning semantic segmentation model ISDU-Net is generated. The ISDU-Net model is trained based on the training set and validation set to obtain the trained ISDU-Net model. The trained ISDU-Net model was used to identify information from multispectral remote sensing images of economic crops in the area to be detected. The process involves using the trained U-Net model as the base model, adding an Inception structure and a channel attention mechanism module SE-Block, and deepening the network to obtain a deep learning semantic segmentation model with a deepened network. Specifically, this includes: The U-Net model is used as the base model, which includes a model encoding part and a model decoding part. In the five-layer downsampling structure of the U-Net model's encoding part, the first three layers each have an Inception structure, and the last two layers each have three convolutional layers, three batch normalization layers, three ReLU activation layers, and one pooling layer, resulting in the deepened model encoding part. The Inception structure comprises four parts, namely 1 1, 3 3, 5 2D convolution performed with 5 convolution kernels and a 3 A max pooling layer with 3 kernels is used to perform activation and BN operations during convolution. In the five-layer upsampling structure of the U-Net model decoding part, a channel attention mechanism module SE-Block is set before each layer upsampling to obtain the model decoding part after network deepening. The channel attention mechanism module SE-Block is set to two parts: compression Squeeze and excitation. The model decoding part after network deepening adopts the form of transposed convolution, gradually decreasing the number of filters layer by layer, and uses high and low level feature fusion to fuse with the feature layers obtained from the encoding part, and then uses 2D convolution for feature extraction. Based on the deepened encoding and decoding parts of the network, a deep learning semantic segmentation model with a deepened network is generated.

2. The method for identifying economic crop information based on deep learning according to claim 1, characterized in that, The process of cropping the UAV multispectral remote sensing image of the study area according to a specified shape to obtain a regular-shaped image includes: Acquire UAV multispectral remote sensing images of the study area; Automatic radiometric correction, deformation correction and image stitching are performed on the multispectral remote sensing images; Using ArcGIS band compositing tools, the five single-band images of red, green, blue, red edge, and near-infrared were combined into a five-band image. Using a raster data cropping tool, the study area was cropped by quadrat to obtain images with regular shapes.

3. The method for identifying economic crop information based on deep learning according to claim 1, characterized in that, The step of annotating the regular shape image to form sample data, and dividing the sample data into a training set, a validation set, and a test set, includes: The obtained regular-shaped image is cropped and bit-transformed. Sample annotation is performed on the image data after cropping and bit transformation. The image data after sample annotation will be randomly divided into training set, validation set and test set according to a predetermined ratio.

4. The method for identifying economic crop information based on deep learning according to claim 3, characterized in that, The sample labeling process specifically includes: A classification system was determined based on the types of economic crops in the study area, and RGB labels were assigned to each type. Based on the information from actual sampling points, the boundaries of each crop were drawn and colored to create a label map, resulting in labeled samples.

5. The method for identifying economic crop information based on deep learning according to claim 3, characterized in that, After the image data with sample annotations is randomly divided into training, validation, and test sets according to a predetermined ratio, it also includes: Data augmentation is performed on the training set, validation set, and test set through geometric transformations to obtain the data-augmented training set, validation set, and test set.

6. The method for identifying economic crop information based on deep learning according to claim 1, characterized in that, The channel attention mechanism module processes the input image as follows: Based on the input UAV multispectral remote sensing image, global average pooling is performed through the compression part to obtain the global feature information of each channel of the multispectral remote sensing image, and the global feature information is then passed to the excitation part. Based on the global feature information input from the compression part, the excitation part performs a nonlinear transformation on the global feature information through two fully connected layers to obtain the weight value of each channel of the multispectral remote sensing image. The weight value is then assigned to the global feature information of the multispectral remote sensing image to obtain the original features of each channel of the multispectral remote sensing image. The compression portion uses a global average pooling method, defined as follows: ; in, H The height of the feature layer. W The width of the feature layer. u c The expression for the output layer result obtained from convolution is as follows: ; in, v c Indicates the first c One convolutional kernel, s Indicates the number of channels. This indicates convolution calculation.

7. The method for identifying economic crop information based on deep learning according to claim 1, characterized in that, The selection of hyperparameters to optimize the parameters of the deep learning semantic segmentation model after network deepening includes: selecting two hyperparameters, learning rate and batch size, to optimize the parameters of the improved deep learning semantic segmentation model.

8. The method for identifying economic crop information based on deep learning according to claim 7, characterized in that, After training the ISDU-Net model based on the training set and validation set to obtain the trained ISDU-Net model, the process further includes: The trained ISDU-Net model is used to predict the test set to obtain the prediction results for the test set. Based on the prediction results, the accuracy of the trained ISDU-Net model is evaluated using the confusion matrix, pixel accuracy, F1 score, crossover ratio, and Kappa coefficient.

9. A deep learning-based economic crop information identification system, characterized in that, include: The image cropping module is used to crop the UAV multispectral remote sensing images of the study area according to a specified shape to obtain images with regular shapes. The sample annotation module is used to annotate the regular shape image to form sample data, and divide the sample data into training set, validation set and test set. The model training module is used to train the deep learning semantic segmentation model based on the training set and the validation set. The deep learning semantic segmentation model is the classic semantic segmentation model U-Net. The structure optimization module is used to take the trained U-Net model as the base model, add the Inception structure and the channel attention mechanism module SE-Block, and deepen the network to obtain a deep learning semantic segmentation model with a deeper network. The parameter optimization module is used to select hyperparameters to optimize the parameters of the deep learning semantic segmentation model after the network is deepened, and generate the improved deep learning semantic segmentation model ISDU-Net model. The ISDU-Net model is trained based on the training set and validation set to obtain the trained ISDU-Net model. The information recognition module is used to recognize information from the multispectral remote sensing images of economic crops in the area to be detected using the trained ISDU-Net model. The process involves using the trained U-Net model as the base model, adding an Inception structure and a channel attention mechanism module SE-Block, and deepening the network to obtain a deep learning semantic segmentation model with a deepened network. Specifically, this includes: The U-Net model is used as the base model, which includes a model encoding part and a model decoding part. In the five-layer downsampling structure of the U-Net model's encoding part, the first three layers each have an Inception structure, and the last two layers each have three convolutional layers, three batch normalization layers, three ReLU activation layers, and one pooling layer, resulting in the deepened model encoding part. The Inception structure comprises four parts, namely 1 1, 3 3, 5 2D convolution performed with 5 convolution kernels and a 3 A max pooling layer with 3 kernels is used to perform activation and BN operations during convolution. In the five-layer upsampling structure of the U-Net model decoding part, a channel attention mechanism module SE-Block is set before each layer upsampling to obtain the model decoding part after network deepening. The channel attention mechanism module SE-Block is set to two parts: compression Squeeze and excitation. The model decoding part after network deepening adopts the form of transposed convolution, gradually decreasing the number of filters layer by layer, and uses high and low level feature fusion to fuse with the feature layers obtained from the encoding part, and then uses 2D convolution for feature extraction. Based on the deepened encoding and decoding parts of the network, a deep learning semantic segmentation model with a deepened network is generated.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the deep learning-based economic crop information identification method as described in any one of claims 1 to 8.

11. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the deep learning-based economic crop information identification method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Remote sensing image water area automatic extraction method and system based on deep learning

    CN111767801A

  • Remote sensing image semantic segmentation method with high accuracy

    CN113298817A