An unmanned aerial vehicle crop state visual identification method for air-ground cooperation

CN117409339BActive Publication Date: 2026-09-25SOUTHEAST UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311321928.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-13
Publication Date
2026-09-25
Estimated Expiration
2043-10-13

AI Technical Summary

Technical Problem

传统机器学习算法包括决策树、支持向量机、随机森林等模型,其性能高度依赖于所提取的特征的准确度,解释性较好,但鲁棒性较差,难以处理实际工作环境下复杂背景影响;而深度学习算法多采用语义分割算法,由网络直接提取深层特征信息并进行端到端倪的区域划分,模型规模较大、鲁棒性较强

Benefits of technology

[0036]本发明公开了一种用于空地协同的无人机作物状态视觉识别方法,该方法利用巡检无人机航拍获取目标田地RGB图片和无人机实时位置信息和姿态数据,构建基于密集连接和多尺度卷积块并联结构的Dense-GoogleNet语义特征提取结构、基于灰度共生矩阵和局部二值模式等特征的纹理特征提取模块和基于通道自注意力机制和编码器-解码器结构的特征图语义分割结构,将航拍图片作为输入,田地作物状态掩码作为输出,获得倒伏区域的像素坐标;根据无人机拍摄实时坐标和姿态角建立坐标转换模型,根据空中三角几何关系将神经网络输出的像素坐标转化为大地坐标下的位置坐标,得到倒伏区域的GPS定位信息。该方法缓解了基于深度学习的语义分割算法模型规模大、计算负荷高和分割精度冗余的问题,在保证实际应用需求的基础上大幅度减小了网络规模和计算量,可以实现倒伏区域的实时、准确监测。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117409339B_ABST
    Figure CN117409339B_ABST
Patent Text Reader

Abstract

The application discloses a kind of unmanned aerial vehicle crop state visual identification methods for air-ground cooperation, comprising the following steps: 1, based on dense connection and multiscale convolution block parallel structure realizes the semantic feature extraction of aerial image;2, based on gray level co-occurrence matrix, local binary pattern algorithm such as extraction aerial image shallow texture feature as the supplement of semantic feature;3, based on channel self-attention mechanism and encoder-decoder structure builds semantic segmentation structure, realizes the grid state judgment of aerial image;4, according to the real-time coordinate and attitude angle of unmanned aerial vehicle shooting constructs coordinate conversion model, according to the geometric relationship of air-ground converts the grid pixel coordinate output by neural network into position coordinate under geodetic coordinate, obtains the position information of crop lodging area.The method is suitable for crop lodging area positioning based on unmanned aerial vehicle inspection, can realize the real-time monitoring of crop lodging state, provides data support for the header parameter adjustment of automatic harvester.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of smart agriculture automatic inspection, specifically a method for visual recognition of crop status using drones for air-ground collaboration. Background Technology

[0002] Lodging significantly reduces crop quality and is a major factor limiting crop yield. Timely and accurate identification of lodged areas can provide technical support for determining the affected area and assessing losses after a disaster. Furthermore, regardless of whether harvesting is done mechanically or manually, lodging significantly increases harvesting difficulty, thereby reducing crop production efficiency. Therefore, there is an urgent need to develop a rapid and efficient crop lodging detection system to quickly obtain accurate information such as the area and location of lodged crops.

[0003] Currently, methods for extracting lodging areas mainly include traditional manual measurement and remote sensing. Manual measurement suffers from problems such as strong subjectivity, high randomness, and lack of unified standards, resulting in low efficiency and being time-consuming and labor-intensive. The rapid development of remote sensing technology, such as near-ground remote sensing, satellite remote sensing, and UAV remote sensing, has provided an effective way for large-scale and rapid detection of lodging information. However, the inefficiency of near-ground remote sensing limits its application at the farmland scale. Satellite remote sensing data has poor spatiotemporal resolution, and the imagery is easily affected by weather, making it difficult to meet the needs of precision agriculture. In contrast, UAV near-ground remote sensing data, with its high accuracy, less terrain constraint, low cost, and ease of operation, effectively bridges the gap between ground surveys and satellite remote sensing, gradually becoming an important method for agricultural information acquisition in the field of precision agriculture.

[0004] After acquiring high-precision near-ground remote sensing data, establishing a reasonable fitting model is crucial. Current crop lodging detection methods based on UAV near-ground remote sensing can be broadly categorized into two types: those based on traditional machine learning and those based on neural networks. Traditional machine learning algorithms, including decision trees, support vector machines, and random forests, are highly dependent on the accuracy of the extracted features. They offer good interpretability but lack robustness and struggle to handle complex background effects in real-world environments. Deep learning algorithms, on the other hand, often employ semantic segmentation, directly extracting deep feature information from the network and performing end-to-end region segmentation. These models are larger in scale and more robust. Considering the complex background environments and uneven distribution of target areas in real-world applications, the key to constructing a lodging area monitoring network lies in the rational design of the neural network and the development of efficient feature extraction modules and pixel classification methods. These factors also determine the network's accuracy and inference efficiency.

[0005] The differences compared to existing technologies are as follows:

[0006] Comparison with patent CN116437801A, which describes "operating vehicle, crop condition detection system, crop condition detection method, crop condition detection program, and recording medium containing the crop condition detection program".

[0007] 1. In patent CN116437801A, crop images are acquired by sensors installed on a harvester, which only detect the area in front of the harvester in the direction of travel. In contrast, we use a drone equipped with sensors to acquire images, which can obtain information about the entire field.

[0008] 2. Patent CN116437801A uses color information to determine crop status. This application uses a combination of color features and vegetation index to achieve status determination.

[0009] Technical comparison with patent CN116367708A "Method and apparatus for determining and mapping crop height"

[0010] 1. Patent CN116367708A uses the height of the cutting rod, the height of the drum, and the crop height information obtained by the height sensor to determine the lodging status, while this application uses the color features obtained by the image sensor to determine the crop status. Summary of the Invention

[0011] To address the aforementioned technical issues, this invention proposes a visual recognition method for crop status using unmanned aerial vehicles (UAVs) for air-ground collaboration. This method alleviates the problems of large semantic segmentation network size, high computational burden, and slow inference speed. It is suitable for detecting lodged areas of crops based on aerial images and can achieve real-time localization of lodged areas. It has low computational load and good real-time performance, thereby improving the efficiency of crop growth status monitoring.

[0012] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0013] A method for visual recognition of crop status using unmanned aerial vehicles (UAVs) for air-to-ground collaboration includes the following steps:

[0014] (1) Obtain visible light images of the target field taken by the inspection drone, and read the drone's real-time position information and attitude data, where the drone's flight altitude is... GPS coordinates are Camera field of view β, pitch angle heading angle ;

[0015] (2) A Dense-GoogleNet structure is constructed based on the parallel structure of dense connections and multi-scale convolutional blocks to realize the semantic feature extraction of aerial images, with an output size of [size missing]. 128-channel feature map;

[0016] The architecture consists of five Inception modules and four downsampling modules. Each Inception module contains four parallel convolutional branches, which reduce the number of model parameters through 1×1 convolutional layers and channel dimensionality reduction and expansion. They also use three different sizes of convolutional kernels and one pooling operation to extract multi-scale features. Each downsampling module consists of a 1×1 convolution to reduce channel dimensionality, a 3×3 convolution, and an average pooling layer with a stride of 2. The output of each Inception module is densely concatenated with the outputs of all preceding Inception modules.

[0017] A dropout layer with a probability of 0.5 is added after each inception module, and a batch normalization (BN) layer is added after each convolution. Large-scale convolution and pooling operations are performed before the image input to the first inception module to reduce the image size.

[0018] (3) A texture feature extraction module is established based on algorithms such as Gray-Level Co-occurrence Matrix (GLCM) and Local Binary Pattern (LBP). This module first uses a set of Gabor filters, including four dilation and four rotation, to filter the original image. Then, it extracts 192 texture features, including GLCM, LBP, frequency domain features, and basic color and intensity features. For GLCM, this module selects a gray level of 8, takes four distance values ​​(1, 2, 4, 5) and four direction values ​​(0°, 45°, 90°, and 135°), and calculates six types of texture feature statistics: energy, contrast, inverse variance, entropy, correlation, and homogeneity. For LBP, eight neighborhood sampling points are selected with a sampling radius of 1. The module uses five statistical measures as parameters: mean, variance, skewness, kurtosis, and entropy, for the feature histograms of basic LBP, rotation-invariant LBP, uniform LBP, and variance LBP. For basic color and intensity features, this module extracts the mean, standard deviation, kurtosis, skewness, average gradient, and Laplacian mean for each channel (r, g, b, h, s, v), as well as the frequency domain energy, mean, variance, entropy, center distance, standard moment, and Hu moment after Fourier transform. In addition to canopy structure and texture features, this module also extracts 10 visible vegetation indices from the RGB image. Finally, all features are concatenated according to grid pixel positions into 192 channels with the same size. Feature map;

[0019] (4) Construct a feature map semantic segmentation structure based on channel self-attention mechanism and encoder-decoder structure; This module introduces channel self-attention mechanism to autonomously learn the importance of deep features extracted by neural network and shallow texture features obtained by texture analysis, and assigns a weight value to each channel, so that the output results tend to depend on the features of key channels. The encoder-decoder architecture is an asymmetric feature fusion network. The encoder includes four downsampling processes, implemented through a 3×3 convolutional layer with a stride of 2, a batch normalization (BN) layer, and a ReLU activation function. The decoder includes four upsampling processes, implemented through a 2×2 transposed convolutional layer with a stride of 2, a concatenation operation, and a 3×3 convolutional block. Skip connections are used to fuse the feature maps before downsampling with those obtained from upsampling, preserving the pixel spatial information in the original shallow feature maps. Each downsampling module halves the feature map size and doubles the number of channels; each upsampling module expands the feature map size and halves the number of channels. The output structure includes a convolutional layer, a sigmoid function, and a rounding operation, responsible for converting the single-channel feature map values ​​output by the decoder into probability values ​​and performing binarization, ultimately obtaining output labels with pixel values ​​of only 0 or 1, achieving grid-level classification of the input image.

[0020] (5) Construct a convolutional neural network, add the outputs of the Dense-GoogleNet structure and the texture feature extraction module in the channel dimension, input the feature map semantic segmentation structure to realize the segmentation of the lodging area, the entire network takes aerial images as input and field crop state mask as output, train the constructed neural network to obtain a crop lodging recognition network for visible light images;

[0021] The Dense-GoogleNet structure and texture feature extraction module divides the original image into sections. The algorithm extracts deep semantic features and shallow texture features from each grid cell. After weighting all features of each grid cell, it classifies the cells into collapsed / normal states. The final output single-channel size is [size missing]. Crop state mask;

[0022] The Focal loss algorithm is used as the loss function. After each generation of training, the classification loss function is calculated by comparing the output mask and the ground truth mask. The formula is as follows, where p is the pixel value of the output mask and y is the corresponding pixel value of the ground truth mask:

[0023]

[0024] (6) The target positioning method based on UAV POS data is adopted. The camera attitude angle, field of view angle, UAV flight altitude, GPS coordinates and other information are obtained through the airborne GPS / INS system during image capture. The GPS coordinates of the target pixel are calculated based on the aerial triangulation relationship.

[0025] The lodging monitoring network obtains the horizontal and vertical coordinates of the lodging grid. According to the gridding scale Get the pixel coordinates of the center point of the region :

[0026]

[0027] The drone's flight altitude is GPS coordinates are Camera field of view β, pitch angle heading angle The camera's field of view is The GPS coordinates of the target pixel are The original aerial image size is ;

[0028] First, calculate the camera's field of view. middle Based on the triangular relationship in the air, we have:

[0029] ; ; ; ; Based on the similarity between the pixel coordinates of the visible light image and the actual GPS coordinates, the GPS coordinates of the center point of the collapsed area can be calculated using the following formula; ; .

[0030] As a further improvement to the identification method of the present invention, step (4) of training the constructed segmentation network is as follows:

[0031] (1) Gaussian noise and contrast, brightness and sharpness adjustment enhancement operations are performed on the dataset; 65% of the enhanced dataset is randomly selected as the training dataset, 15% of the images constitute the validation dataset, and the remaining 20% ​​constitute the test dataset.

[0032] (2) The semantic segmentation part of the feature map is randomly initialized; the Dense-GoogleNet part of the semantic feature extraction network uses pre-trained weights on the COCO dataset for transfer learning. In order to prevent the weights of the feature extraction network from being destroyed in the early stage of training, the parameters of the backbone network in the first 25 generations of training are frozen and do not participate in gradient update.

[0033] (3) Based on the backpropagation algorithm, the Adam optimizer and mini-batch stochastic gradient descent method are adopted. The learning rate decrease curve adopts the StepLR fixed step size decay strategy, and gamma is 0.9. The weights of the semantic feature extraction network and the feature map semantic segmentation structure are fine-tuned and updated respectively.

[0034] As a further improvement to the identification method of the present invention, step (6) uses the aerial triangular geometry of the drone's aerial photography posture to locate the target in the aerial image.

[0035] Beneficial effects:

[0036] This invention discloses a visual recognition method for crop status using drones for air-ground collaboration. The method utilizes aerial photography of target fields to acquire RGB images and real-time drone location and attitude data. It constructs a Dense-GoogleNet semantic feature extraction structure based on dense connections and multi-scale convolutional block parallel structures, a texture feature extraction module based on features such as gray-level co-occurrence matrix and local binary patterns, and a feature map semantic segmentation structure based on channel self-attention mechanism and encoder-decoder structure. Using the aerial images as input and the field crop status mask as output, the pixel coordinates of the lodged area are obtained. A coordinate transformation model is established based on the real-time coordinates and attitude angles captured by the drone. Based on aerial triangulation, the pixel coordinates output by the neural network are transformed into geodetic position coordinates, obtaining the GPS positioning information of the lodged area. This method alleviates the problems of large model size, high computational load, and redundant segmentation accuracy in deep learning-based semantic segmentation algorithms. While ensuring practical application requirements, it significantly reduces network size and computational load, enabling real-time and accurate monitoring of lodged areas. Attached Figure Description

[0037] Figure 1 This is a flowchart of the method disclosed in this invention;

[0038] Figure 2 This is a diagram of the semantic feature network structure in this invention;

[0039] Figure 3 This is a diagram of the feature fusion network structure in this invention;

[0040] Figure 4 This is a schematic diagram showing the flight parameters and camera field of view during drone inspection. Detailed Implementation

[0041] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments:

[0042] This invention discloses a visual recognition method for crop status using unmanned aerial vehicles (UAVs) for air-to-ground collaboration. The flowchart of the disclosed method is as follows: Figure 1 As shown, it includes the following steps:

[0043] Step 1: Acquire visible light images of the target field taken by the inspection drone, and read the drone's real-time position and attitude data. The drone's flight altitude is... GPS coordinates are Camera field of view β, pitch angle heading angle .

[0044] Step 2: Construct a Dense-GoogleNet structure based on dense connections and a parallel structure of multi-scale convolutional blocks to extract semantic features from aerial images. The output size is [size missing]. The 128-channel feature map, its structural diagram is as follows: Figure 2 As shown.

[0045] This module consists of five Inception modules and four downsampling modules. Each Inception module contains four parallel convolutional branches, which reduce the number of model parameters through 1×1 convolutional layers and channel dimensionality reduction and expansion, and extract multi-scale features using three different sized convolutional kernels and one pooling operation. Each downsampling module consists of a 1×1 convolution to reduce channel dimensionality, a 3×3 convolution, and an average pooling layer with a stride of 2. To further improve the model's accuracy and efficiency, the output of each Inception module is densely concatenated with the outputs of all preceding Inception modules to increase information flow and sharing, thereby improving accuracy and efficiency while maintaining a small number of model parameters.

[0046] To prevent overfitting, a dropout layer with a probability of 0.5 is added after each inception module, and a batch normalization (BN) layer is added after each convolution. To improve the speed of network training, large-scale convolution and pooling operations are performed before the image input to the first inception module to reduce the image size.

[0047] Step 3: Establish a texture feature extraction module based on algorithms such as gray-level co-occurrence matrix and local binary mode. This module first uses a set of Gabor filters, including four dilation and four rotation, to filter the original image, and then extracts 192 texture features, including gray-level co-occurrence matrix (GLCM), local binary mode (LBP), frequency domain features, and basic color and intensity features. For GLCM, this module selects 8 gray levels, takes four distance values ​​(1, 2, 4, 5) and four directional values ​​(0°, 45°, 90°, 135°), and calculates six types of texture feature statistics: energy, contrast, inverse variance, entropy, correlation, and homogeneity. For LBP, 8 neighborhood sampling points are selected with a sampling radius of 1, and five types of statistics—mean, variance, skewness, kurtosis, and entropy—are calculated as parameters for the feature histograms of basic LBP, rotation-invariant LBP, uniform LBP, and variance LBP. For basic color and intensity features, this module extracts the mean, standard deviation, kurtosis, skewness, average gradient, and Laplacian mean of each channel (r, g, b, h, s, v), as well as the frequency domain energy, frequency domain mean, frequency domain variance, frequency domain entropy, frequency domain center distance, frequency domain standard moment, and frequency domain Hu moment after Fourier transform. In addition to canopy structure and texture features, this module also extracts 10 visible vegetation indices from RGB images, as shown in Table 1. Finally, all features are concatenated according to their grid pixel positions to form a 192-channel array with the same size. The feature map.

[0048] Table 1

[0049]

[0050] Step 4: Construct a feature map semantic segmentation structure based on the channel self-attention mechanism and encoder-decoder structure, as shown in the schematic diagram below. Figure 3As shown, this module introduces a channel self-attention mechanism to autonomously learn the importance of deep features extracted by the neural network and shallow texture features obtained from texture analysis. Each channel is assigned a weight value, thus making the output biased towards the features of key channels. The encoder-decoder structure is an asymmetric feature fusion network. The encoder includes four downsampling processes, implemented through a 3×3 convolutional layer with a stride of 2, a BN layer, and a ReLU activation function. The decoder includes four upsampling processes, implemented through a 2×2 transposed convolutional layer with a stride of 2, a concatenation operation, and a 3×3 convolutional block. Skip connections are used to fuse the feature map before downsampling with the feature map obtained after upsampling, preserving the pixel spatial information in the original shallow feature map. Each downsampling module halves the feature map size and doubles the number of channels; each upsampling module expands the feature map size and halves the number of channels. The output structure includes a convolutional layer, a sigmoid function, and a rounding operation. It is responsible for converting the single-channel feature map values ​​output by the decoder into probability values ​​and performing binarization, ultimately obtaining output labels with pixel values ​​of only 0 or 1, thus achieving grid-level classification of the input image.

[0051] Step 5: Construct a lodging region detection neural network. Use Dense-GoogleNet from Step 2 to extract deep semantic features, use the texture feature module from Step 3 to obtain shallow texture features, and use the feature map semantic segmentation structure from Step 4 to extract lodging regions from aerial images. Design a loss function using the corresponding bit Focal loss method of mask image. Use the aerial RGB image as network input and the grid-level crop state mask image as output to train the constructed neural network, thus obtaining a neural network for real-time detection of crop lodging.

[0052] The Dense-GoogleNet structure and texture feature extraction module divides the original image into sections. The algorithm extracts deep semantic features and shallow texture features from each grid cell. After weighting all features of each grid cell, it classifies the cells into collapsed / normal states. The final output single-channel size is [size missing]. The crop state mask.

[0053] The Focal loss algorithm is used as the loss function. After each generation of training, the classification loss function is calculated by comparing the output mask and the ground truth mask. The formula is as follows, where p is the pixel value of the output mask and y is the corresponding pixel value of the ground truth mask:

[0054] (6)

[0055] The steps for training the constructed neural network are as follows:

[0056] (5-1) Enhance the dataset by adding Gaussian noise and adjusting contrast, brightness, and sharpness; randomly select 65% of the enhanced dataset as the training dataset, 15% of the images as the validation dataset, and the remaining 20% ​​as the test dataset.

[0057] (5-2) The semantic segmentation part of the feature map is randomly initialized; the Dense-GoogleNet part of the semantic feature extraction network uses pre-trained weights on the COCO dataset for transfer learning. In order to prevent the weights of the feature extraction network from being destroyed in the early stage of training, the parameters of the backbone network in the first 25 generations of training are frozen and do not participate in gradient update.

[0058] (3-3) Based on the backpropagation algorithm, the Adam optimizer and mini-batch stochastic gradient descent method are adopted. The learning rate decrease curve adopts the StepLR fixed step size decay strategy, and gamma is 0.9. The weights of the semantic feature extraction network and the feature map semantic segmentation structure are fine-tuned and updated respectively.

[0059] Step 6: Use the target localization method based on UAV POS data. Use the onboard GPS / INS system to obtain information such as camera attitude angle, field of view angle, UAV flight altitude, and GPS coordinates during image capture. Calculate the GPS coordinates of the target pixel based on aerial triangulation.

[0060] The lodging monitoring network obtains the horizontal and vertical coordinates of the lodging grid. According to the gridding scale Get the pixel coordinates of the center point of the region :

[0061] (7)

[0062] A diagram illustrating the flight parameters and camera field of view during drone inspection is shown below. Figure 4 As shown. The drone's flight altitude is... GPS coordinates are Camera field of view β, pitch angle heading angle The camera's field of view is The GPS coordinates of the target pixel are The original aerial image size is .

[0063] First, calculate the GPS coordinates of the four corner points within the camera's field of view, i.e., the four vertices of the image. Based on aerial triangulation, we have: ; ; ; ; Based on the similarity between the pixel coordinates of the visible light image and the actual GPS coordinates, the GPS coordinates of the center point of the collapsed area can be calculated using the following formula; ; .

[0064] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any other way. Any modifications or equivalent changes made based on the technical essence of the present invention shall still fall within the scope of protection claimed by the present invention.

Claims

1. A method for visual recognition of crop status using unmanned aerial vehicles (UAVs) for air-to-ground collaboration, characterized in that, Includes the following steps: (1) Obtain visible light images of the target field taken by the inspection drone, and read the drone's real-time position information and attitude data, where the drone's flight altitude is... GPS coordinates are Camera field of view β, pitch angle heading angle ; (2) A Dense-GoogleNet structure is constructed based on the parallel structure of dense connections and multi-scale convolutional blocks to realize the semantic feature extraction of aerial images, with an output size of [size missing]. 128-channel feature map; The architecture consists of five Inception modules and four downsampling modules. Each Inception module contains four parallel convolutional branches, which reduce the number of model parameters through 1×1 convolutional layers and channel dimensionality reduction and expansion. They also use three different sizes of convolutional kernels and one pooling operation to extract multi-scale features. Each downsampling module consists of a 1×1 convolution responsible for reducing channel dimensionality, a 3×3 convolution, and an average pooling layer with a stride of 2. The output of each Inception module is densely concatenated with the outputs of all the preceding Inception modules. A dropout layer with a probability of 0.5 is added after each inception module, and a batch normalization (BN) layer is added after each convolution. Large-scale convolution and pooling operations are performed before the image input to the first inception module to reduce the image size. (3) A texture feature extraction module is established based on the gray-level co-occurrence matrix and local binary mode algorithm. This module first uses a set of Gabor filters including four dilation and four rotation to filter the original image, and then extracts 192 texture features including gray-level co-occurrence matrix GLCM, local binary mode LBP, frequency domain features and basic color and intensity features. For GLCM, this module selects 8 gray levels, takes four distance values ​​of 1, 2, 4, 5 and four direction values ​​of 0°, 45°, 90° and 135°, and calculates six types of texture feature statistics: energy, contrast, inverse variance, entropy, correlation and homogeneity. For LBP (Local Beam Profile), eight neighborhood sampling points with a sampling radius of 1 were selected. Five statistical measures—mean, variance, skewness, kurtosis, and entropy—were calculated as parameters for the feature histograms of basic LBP, rotation-invariant LBP, uniform LBP, and variance LBP. For basic color and intensity features, this module extracted the mean, standard deviation, kurtosis, skewness, average gradient, and Laplacian mean for each channel (r, g, b, h, s, v), as well as the frequency domain energy, mean, variance, entropy, center distance, standard moment, and Hu moment after Fourier transform. In addition to canopy structure and texture features, this module also extracted 10 visible vegetation indices from the RGB image. Finally, all features were concatenated according to grid pixel positions into a 192-channel array with the same size. Feature map; (4) Construct a feature map semantic segmentation structure based on channel self-attention mechanism and encoder-decoder structure; This module introduces channel self-attention mechanism to autonomously learn the importance of deep features extracted by neural network and shallow texture features obtained by texture analysis, and assigns a weight value to each channel, so that the output results tend to depend on the features of key channels. The encoder-decoder structure is an asymmetric feature fusion network, in which the encoder contains four downsampling processes, implemented through a 3×3 convolutional layer with a stride of 2, a BN layer and a ReLU activation function; the decoder contains four upsampling processes, implemented through a 2×2 step The system is implemented using a transposed convolutional layer of length 2, a concatenation operation, and a 3×3 convolutional block. Skip connections are used to fuse the feature map before downsampling with the feature map obtained from upsampling, preserving the pixel spatial information in the original shallow feature map. Each downsampling module halves the feature map size and doubles the number of channels; each upsampling module expands the feature map size and halves the number of channels. The output structure contains a convolutional layer, a sigmoid function, and a rounding operation, which is responsible for converting the single-channel feature map values ​​output by the decoder into probability values ​​and performing binarization, ultimately obtaining output labels with pixel values ​​of only 0 or 1, achieving grid-level classification of the input image. (5) Construct a convolutional neural network, add the outputs of the Dense-GoogleNet structure and the texture feature extraction module in the channel dimension, input the feature map semantic segmentation structure to realize the segmentation of the lodging area, the entire network takes aerial images as input and field crop state mask as output, train the constructed neural network to obtain a crop lodging recognition network for visible light images; The Dense-GoogleNet structure and texture feature extraction module divides the original image into sections. The algorithm extracts deep semantic features and shallow texture features from each grid cell. After weighting all features of each grid cell, it classifies the cells into collapsed / normal states. The final output single-channel size is [size missing]. Crop state mask; The Focal loss algorithm is used as the loss function. After each generation of training, the classification loss function is calculated by comparing the output mask and the ground truth mask. The formula is as follows, where p is the pixel value of the output mask and y is the corresponding pixel value of the ground truth mask: ; (6) The target positioning method based on UAV POS data is adopted. The camera attitude angle, field of view angle, UAV flight altitude and GPS coordinate information are obtained through the airborne GPS / INS system during image capture. The GPS coordinates of the target pixel are calculated based on the aerial triangulation relationship. The lodging monitoring network obtains the horizontal and vertical coordinates of the lodging grid. According to the gridding scale Get the pixel coordinates of the center point of the region : ; The drone's flight altitude is Camera field of view β, pitch angle heading angle The camera's field of view is The GPS coordinates of the target pixel are The original aerial image size is ; First, calculate the camera's field of view. middle Based on the triangular relationship in the air, we have: ; ; ; ; Based on the similarity between the pixel coordinates of the visible light image and the actual GPS coordinates, the GPS coordinates of the center point of the collapsed area can be calculated using the following formula; ; 。 2. The method for visual recognition of crop status by unmanned aerial vehicles (UAVs) for air-to-ground coordination according to claim 1, characterized in that, Step (4) involves training the constructed segmentation network as follows: (1) Gaussian noise and contrast, brightness and sharpness adjustment enhancement operations are performed on the dataset; 65% of the enhanced dataset is randomly selected as the training dataset, 15% of the images constitute the validation dataset, and the remaining 20% ​​constitute the test dataset. (2) The semantic segmentation part of the feature map is randomly initialized; The semantic feature extraction network Dense-GoogleNet uses pre-trained weights on the COCO dataset for transfer learning. To prevent the weights of the feature extraction network from being destroyed in the early stages of training, the backbone network parameters in the first 25 generations of training are frozen and do not participate in gradient updates. (3) Based on the backpropagation algorithm, the Adam optimizer and mini-batch stochastic gradient descent method are adopted. The learning rate decrease curve adopts the StepLR fixed step size decay strategy, and gamma is 0.

9. The weights of the semantic feature extraction network and the feature map semantic segmentation structure are fine-tuned and updated respectively.

3. The method for visual recognition of crop status by unmanned aerial vehicles (UAVs) for air-to-ground coordination according to claim 1, characterized in that, In step (6), the aerial triangulation relationship of the drone's aerial photography attitude is used to locate the target in the aerial image.

Citation Information

Patent Citations

  • Method and apparatus for determining and mapping crop height

    CN116367708A

  • Work vehicle, crop state detection system, crop state detection method, crop state detection program, and recording medium on which crop state detection program is recorded

    CN116437801A