Adaptive prediction method for number of spikes per hectare of wheat based on variational autoencoder

The wheat ear count prediction model constructed by variational autoencoder, combined with image recognition and deep learning technology, automatically identifies the number of wheat ears and estimates the number of ears per mu over a wide range. This solves the problems of high labor intensity and poor accuracy in traditional methods, and achieves efficient and accurate prediction results.

CN121305334BActive Publication Date: 2026-05-12SHANDONG LUYAN SEED CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANDONG LUYAN SEED CO LTD
Filing Date
2025-09-22
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing methods for predicting the number of wheat ears per mu rely on manual counting and empirical estimation, which are labor-intensive and lack efficient and accurate technical support. They are particularly inaccurate in large-scale farmland, and traditional deep learning models suffer from overfitting, making it difficult to achieve efficient and accurate predictions in different regions.

Method used

An adaptive wheat ear count prediction method based on variational autoencoder is adopted. Image data of wheat fields are acquired by UAVs, and a prediction model is constructed using variational encoder, uncertainty module, multi-scale attention module and density decoding module. Combined with image recognition and deep learning technology, the wheat ear count is automatically identified and the ear count per mu is estimated over a large area.

Benefits of technology

It achieves high-precision prediction of wheat ear count per mu under different plots and environmental conditions, reduces human intervention, improves the model's generalization ability, reduces the risk of overfitting, and improves the accuracy and robustness of prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121305334B_ABST
    Figure CN121305334B_ABST
Patent Text Reader

Abstract

The application discloses a kind of self-adapting wheat ear number prediction methods based on variational autoencoder, it is related to machine learning technical field, and local and global image after wheat heading is collected by unmanned aerial vehicle, and data set is constructed by combining unmanned aerial vehicle flight parameter, and feature extraction and modeling are carried out.Then, the local image and global image of wheat are extracted by variational autoencoder structure Feature, and the uncertainty module is used to model the uncertainty of the model.Then, the local detail features and global spatial features are fused by the multi-scale attention mechanism, and finally the wheat ear number per mu is predicted by the multi-scale density decoding module.The method can improve the prediction accuracy of wheat ear number per mu under different conditions, and has good generalization ability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of machine learning technology, specifically to an adaptive wheat ear count prediction method based on variational autoencoders. Background Technology

[0002] Wheat, as one of the world's most important food crops, is widely cultivated around the world. Accurate yield prediction is crucial for agricultural management, resource allocation, and policy-making during wheat production. The number of ears per acre is a key factor influencing yield; therefore, accurate prediction of this number provides a valuable basis for agricultural decision-making. However, traditional methods for predicting wheat ear count largely rely on manual counting and empirical estimation. These methods are not only labor-intensive but also lack efficient and accurate technical support in large-scale planting areas.

[0003] Currently, with the development of modern agricultural technology, especially the continuous advancement of remote sensing and image processing technologies, crop monitoring and prediction using image recognition and deep learning technologies has become a research hotspot. Field image data acquired through devices such as drones and satellites can provide rich visual information for crop growth status and yield prediction. However, existing crop prediction methods based on image data, especially in predicting the number of wheat ears per acre, still face several significant challenges: First, how to efficiently and accurately identify the number of wheat ears per acre from field images; second, in large-scale farmland, traditional methods directly extrapolate the prediction results of local areas to the number of ears per acre, while the number of wheat ears varies greatly in different areas, resulting in poor accuracy; third, how to avoid the overfitting problem in traditional deep learning models and improve the generalization ability of the model when dealing with large and diverse datasets. Summary of the Invention

[0004] In order to overcome the shortcomings of the above technologies, this invention provides a method to improve the accuracy of predicting the number of wheat ears per mu under different plots and environmental conditions.

[0005] The technical solution adopted by this invention to overcome its technical problems is:

[0006] An adaptive wheat ear count prediction method based on variational autoencoder includes:

[0007] S1. Obtaining data from the wheat experimental field using drones. A collection of images of small square samples and A collection of images of a large square sample , ,in For the first An image of a small square sample, , ,in For the first An image of a large square sample;

[0008] S2. Obtain the actual set of ears per mu (unit of land area). , , For the first The actual number of ears per mu;

[0009] S3. Preprocess the images of the large square quadrats and the small square quadrats to obtain a set of preprocessed large square image sets. and the image set of preprocessed small sample plots According to the image set and image collection Create the WESN dataset and divide it into a training set and a test set;

[0010] S4. Construct a wheat ear number prediction model consisting of a first variational encoder, a second variational encoder, a third variational encoder, an uncertainty module, a global area calculation module, a multi-scale attention module, and a multi-scale density decoding module;

[0011] S5. Input the preprocessed images of small square quadrats from the training set into the first variational encoder, second variational encoder, and third variational encoder of the wheat ear number prediction model, and output the first encoded features respectively. Second coding features Third coding feature ;

[0012] S6. Combine the preprocessed images of large square quadrats in the training set with the second encoded features. The input is given to the uncertainty module of the wheat ear number prediction model, and the output is the uncertainty features. ;

[0013] S7. Input the preprocessed images of small square quadrats and the preprocessed images of large square quadrats from the training set into the global area calculation module of the wheat ear number prediction model, and output the local area estimates respectively. and global area estimation ;

[0014] S8. The third coding feature Uncertainty characteristics The input is fed into the multi-scale attention module of the wheat ear number prediction model, and the output is the attention feature. ;

[0015] S9. Attention Features The data is input into the multi-scale density decoding module of the wheat ear number prediction model, and the output is the wheat ear number per acre. ;

[0016] S10. Based on the number of wheat ears per mu (unit of land area) With the The actual number of ears per mu Construct a mean squared error loss function, use the Adam optimizer, and train the wheat ear number prediction model using the mean squared error loss function to obtain the optimized wheat ear number prediction model.

[0017] S11. Input the preprocessed images of small square quadrats and large square quadrats from the test set into the optimized wheat ear number prediction model, and output the wheat ear number per mu (unit of land area). .

[0018] Furthermore, step S1 includes the following steps:

[0019] S1-1. Select a wheat experimental field with an area of ​​M, and randomly set up [various species] within the experimental field. A 10-square-meter large-scale square plot ,in For the first A large square sample, ;

[0020] S1-2. During the critical growth period of wheat, 5-7 days after heading, data were collected using drones. A large square sample The image of a small square with a center area of ​​1 square meter is obtained. Image of a small square sample ;

[0021] S1-3. After the drone has taken images of all the square sample plots, adjust the drone's flight altitude so that its field of view covers the [number missing]. A large square sample The first one was collected. Image of a large square .

[0022] Preferably, in step S1-1, M is greater than or equal to 1000 square meters; during the critical growth period 5-7 days after wheat heading, on a sunny day with a wind speed of less than or equal to 5 m / s, between 9:00 and 11:00 AM, the wheat experimental field is photographed using a Zenmuse P1 camera mounted on a DJI M300 RTK drone with the gimbal tilt angle set to -90°; in step S1-1 The value is 100. The large square plots do not overlap.

[0023] Furthermore, in step S2, in the first A large square sample A central 1-square-meter area was designated as a statistical quadrat. The number of ears in each quadrat was manually counted, and the result was multiplied by 666.67 to obtain the number of ears. The actual number of ears per mu .

[0024] Furthermore, step S3 includes the following steps:

[0025] S3-1. The first Image of a small square sample The cv2.createCLAHE() function in OpenCV is used to enhance the contrast, resulting in the enhanced contrast level. Image of a small square sample , will the Image of a small square sample Smoothing and denoising are performed using a Gaussian filter with a kernel size of 3×3, resulting in the preprocessed i.e. Image of a small square sample ,all The images of the preprocessed square minima constitute the image set of the preprocessed minima. , ;

[0026] S3-2. The first Image of a large square The cv2.createCLAHE() function in OpenCV is used to enhance the contrast, resulting in the enhanced contrast level. Image of a large square , will the Image of a large square Smoothing and denoising are performed using a Gaussian filter with a kernel size of 3×3, resulting in the preprocessed i.e. Image of a large square ,all The images of the preprocessed square matrices constitute the image set of the preprocessed matrices. , ;

[0027] S3-3. Constructing Data Pairs ,all The WESN dataset consists of several data pairs. ;

[0028] S3-4. Divide the WESN dataset into training and test sets in a 7:3 ratio.

[0029] Furthermore, step S5 includes the following steps:

[0030] S5-1. The first variational encoder of the wheat ear number prediction model consists of a first convolutional layer, a first ReLU activation function, a second convolutional layer, a second ReLU activation function, a third convolutional layer, a third ReLU activation function, a fourth convolutional layer, and a fourth ReLU activation function.

[0031] S5-2. The preprocessed i-th element in the training set Image of a small square sample The input is fed into the first variational encoder, and the output is the first encoded feature. ;

[0032] S5-3. The second variational encoder of the wheat ear number prediction model consists of a spatial mapping layer, a first fully connected layer, a first ReLU activation function, a second fully connected layer, and a second ReLU activation function, which encodes the first feature. The input to the spatial mapping layer is flattened into a vector using the `flatten()` function of the PyTorch library. The flattened vector is then reduced to 512 dimensions using the `nn.Linear()` function of the PyTorch library to obtain the features. , will feature The inputs are sequentially fed into the first fully connected layer, the first ReLU activation function, the second fully connected layer, and the second ReLU activation function of the second variational encoder, and the output is the second encoded feature. ;

[0033] S5-4. The third variational encoder of the wheat ear number prediction model consists of a fully connected layer, a reshape function, a first convolutional layer, a second convolutional layer, and a global average pooling layer.

[0034] S5-5. The second coding feature The input is fed into the fully connected layer of the third variational encoder, and the output is a tensor. , tensor The input is fed into the Reshape function of the third variational encoder, and the output is a tensor. , the first encoded feature With tensor By concatenating the features along the channel dimension, a fused feature map is obtained. ;

[0035] S5-6. Merge feature maps The inputs are sequentially fed into the first convolutional layer, the second convolutional layer, and the global average pooling layer of the third variational encoder, and the output is the third encoded feature. .

[0036] Furthermore, step S6 includes the following steps:

[0037] S6-1. The uncertainty module of the wheat ear number prediction model consists of a feature fusion layer and an uncertainty distribution modeling layer;

[0038] S6-2. The preprocessed first... Image of a large square and second coding features The input is fed into the feature fusion layer of the uncertainty module, through the formula. Calculated features In the formula, For a randomly initialized extended weight matrix, As a bias term, the features The input is fed into the reshape function of the PyTorch library, and the output is the feature. , will feature Compared with the preprocessed first Image of a large square The channel-dimensional concat() function from the PyTorch library is used to perform concatenation operations to obtain fused features. ;

[0039] S6-3. Merging Features The uncertainty distribution modeling layer input into the uncertainty module uses the flatten() function from the PyTorch library to flatten and reduce dimensionality, resulting in a vector. Through formula Calculate the mean vector , It is a linear mapping matrix. As the bias vector, it is obtained through the formula The variance vector is calculated. In the formula It is a linear mapping matrix. It is the bias vector;

[0040] S6-4. Through formula Uncertainty characteristics were calculated. In the formula For element-wise multiplication, Standard normal noise generated for the randn() function in the PyTorch library.

[0041] Furthermore, step S7 includes the following steps:

[0042] S7-1. Through formula The local area estimate is calculated. In the formula For the preprocessed first Image of a small square sample The Middle Line number The pixel values ​​of the columns of pixels. For the preprocessed first Image of a small square sample pixel domain, For drones to collect the first Image of a small square sample Flight altitude relative to the ground For the first Image of a small square sample Horizontal pixel count, For the first Image of a small square sample Vertical pixel count, For drones to collect the first Image of a small square sample The focal length of the camera at that time. For drones to collect the first Image of a small square sample The roll angle at that time For drones to collect the first Image of a small square sample The pitch angle at that time;

[0043] S7-2. Through formula The global area estimate is calculated. In the formula For the processed first Image of a large square The Middle Line number The pixel values ​​of the columns of pixels. For the processed first Image of a large square pixel domain, For drones to collect the first Image of a large square Flight altitude relative to the ground For the first Image of a large square Horizontal pixel count, For the first Image of a large square Vertical pixel count, For drones to collect the first Image of a large square The focal length of the camera at that time. For drones to collect the first Image of a large square The roll angle at that time For drones to collect the first Image of a large square The pitch angle at that time.

[0044] Furthermore, step S8 includes the following steps:

[0045] S8-1. The third coding feature Uncertainty characteristics The input is fed into the multi-scale attention module of the wheat ear number prediction model via the formula.

[0046] Calculate attention features In the formula For the third coding feature Uncertainty characteristics the number of rows, For the third coding feature Uncertainty characteristics The number of columns, For the third coding feature The Line number The parameter values ​​of the column, , , Uncertainty characteristics The Line number The parameter values ​​of the column, Uncertainty characteristics The Line number The parameter values ​​of the column, , For smoothing parameters.

[0047] Furthermore, step S9 includes the following steps:

[0048] S9-1. The multi-scale density decoding module of the wheat ear number prediction model consists of a first fully connected layer, a first ReLU activation function, a second fully connected layer, a second ReLU activation function, a third fully connected layer, a third ReLU activation function, and a weight fusion layer.

[0049] S9-2. Attention Features The inputs are sequentially fed into the first fully connected layer, the first ReLU activation function, the second fully connected layer, the second ReLU activation function, the third fully connected layer, and the third ReLU activation function of the multi-scale density decoding module, and the output yields the initial number of ears per acre. ;

[0050] S9-3. Initial number of ears per mu The weight fusion layer of the multi-scale density decoding module is input into the formula. Calculate the number of wheat ears per mu .

[0051] The beneficial effects of this invention are: by combining image recognition technology and a deep learning model, it automatically identifies the number of ears of wheat after heading and, combined with the spatial information of the plot, further calculates the number of ears per acre over a larger area, and finally calculates the number of ears per acre. First, the model extracts features and reduces noise from images of wheat fields to identify the number of wheat ears. Then, through the combination of spatial data and adaptive adjustment of the model, the ear count information of local fields is extended to a larger area, thereby accurately estimating the number of ears per acre in large-area plots. Attached Figure Description

[0052] Figure 1 This is a flowchart of the method of the present invention;

[0053] Figure 2 This is a comparison image of the input image and the output image of the multi-scale attention module of this invention. Detailed Implementation

[0054] The following is in conjunction with the appendix Figure 1 The present invention will be further described below.

[0055] An adaptive wheat ear count prediction method based on variational autoencoder includes:

[0056] S1. Obtaining data from the wheat experimental field using drones. A collection of images of small square samples and A collection of images of a large square sample , ,in For the first An image of a small square sample, , ,in For the first An image of a large square.

[0057] S2. Obtain the actual set of ears per mu (unit of land area). , , For the first The actual number of ears per mu.

[0058] S3. Preprocess the images of the large square quadrats and the small square quadrats to obtain a set of preprocessed large square image sets. and the image set of preprocessed small sample plots According to the image set and image collection Create the WESN dataset and divide it into a training set and a test set.

[0059] S4. Construct a wheat ear number prediction model consisting of a first variational encoder, a second variational encoder, a third variational encoder, an uncertainty module, a global area calculation module, a multi-scale attention module, and a multi-scale density decoding module.

[0060] S5. Input the preprocessed images of small square quadrats from the training set into the first variational encoder, second variational encoder, and third variational encoder of the wheat ear number prediction model, and output the first encoded features respectively. Second coding features Third coding feature .

[0061] S6. Combine the preprocessed images of large square quadrats in the training set with the second encoded features. The input is given to the uncertainty module of the wheat ear number prediction model, and the output is the uncertainty features. .

[0062] S7. Input the preprocessed images of small square quadrats and the preprocessed images of large square quadrats from the training set into the global area calculation module of the wheat ear number prediction model, and output the local area estimates respectively. and global area estimation .

[0063] S8. The third coding feature Uncertainty characteristics The input is fed into the multi-scale attention module of the wheat ear number prediction model, and the output is the attention feature. .

[0064] S9. Attention Features The data is input into the multi-scale density decoding module of the wheat ear number prediction model, and the output is the wheat ear number per mu (unit of land area). .

[0065] S10. Based on the number of wheat ears per mu (unit of land area) With the The actual number of ears per mu A mean squared error loss function is constructed, and the Adam optimizer is used to train the wheat ear number prediction model, resulting in an optimized wheat ear number prediction model.

[0066] S11. Input the preprocessed images of small square quadrats and large square quadrats from the test set into the optimized wheat ear number prediction model, and output the wheat ear number per mu (unit of land area). .

[0067] Because the wheat ear number prediction model was effectively trained, the optimized wheat ear number prediction model outputs the number of wheat ears per mu (unit of land area). It is accurate. New images collected by the drone can be continuously input into the optimized wheat ear number prediction model to obtain accurate wheat ear number identification values. It realizes an efficient and robust prediction method that can be applied in a variety of agricultural environments, helping to realize precision agriculture.

[0068] To verify the effectiveness of the adaptive wheat ear count prediction method based on variational autoencoder described in this invention, a comparative experiment with traditional methods was conducted. All experiments were performed using the WESN dataset, and the test data included local and global images of wheat, as well as relevant flight parameters of the UAV. The experimental results are shown in Table 1:

[0069] Table 1 Comparison of Experimental Results

[0070]

[0071] All methods were trained using the same test set, which consisted of local and global images of wheat and relevant UAV flight parameters. The input features of all methods were standardized. The difference between the predicted number of ears per acre and the actual number of ears per acre reflects the prediction accuracy of each method under different scenarios.

[0072] The mean squared error (MSE) of the method in this invention (based on a variational autoencoder) is 0.123, significantly lower than the other three traditional methods (0.150 for traditional method A, 0.135 for traditional method B, and 0.145 for traditional method C). This indicates that the method in this invention has the smallest error when predicting the number of wheat ears per acre, demonstrating high prediction accuracy in practical applications and a better fit to the data. Traditional method A (linear regression) has the highest MSE, performing relatively poorly. This may be because linear regression is too simplistic and cannot capture the complex nonlinear characteristics of wheat growth. Although the MSEs of traditional methods B (neural network) and C (random forest) are lower than those of linear regression, they are still higher than those of the method in this invention. This shows that in modeling nonlinear features, the variational autoencoder-based model still outperforms traditional deep learning and ensemble learning methods.

[0073] The mean absolute error (MAE) reflects the average deviation between the predicted and actual results. The MAE of the method in this invention is 0.115, which is also lower than the other three methods. The lower MAE indicates that the variational autoencoder-based model can estimate the number of wheat ears per acre more accurately, with a relatively small prediction error. The MAE of traditional method A (linear regression) is 0.140, significantly higher than the other methods, indicating the disadvantage of linear regression in prediction accuracy. The MAEs of traditional methods B (neural network) and C (random forest) are 0.130 and 0.140, respectively. Although these are improvements over linear regression, they still cannot achieve the prediction effect of the variational autoencoder-based method.

[0074] The R² value represents the proportion of data variance explained by the model; the closer the value is to 1, the better the model fits the data. The R² value of the method in this invention is 0.95, significantly higher than the other three methods (traditional method A: 0.90, traditional method B: 0.92, and traditional method C: 0.91). This means that the method in this invention performs well in capturing the changing trend of wheat ear count per acre and can explain the variance of the actual data relatively well. The R² values ​​of the traditional methods are all lower than those of the method in this invention, indicating that their fit to the data is poor.

[0075] The comparison between the predicted and actual number of ears per mu (unit of land area) shows that the method of this invention has the smallest difference between the predicted and actual values, and the highest prediction accuracy. Other traditional methods deviate significantly from the actual values, especially the linear regression method, which has a difference of 12 units between the predicted and actual values, demonstrating its limitations in prediction tasks.

[0076] This invention discloses an adaptive wheat ear count prediction method based on variational autoencoders. In practical applications, the system acquires local images of wheat fields using a UAV, performs a series of data processing steps, and then accurately predicts wheat ears using a multi-scale attention module. The following section provides a detailed analysis of the input image and the output image after decoding by the multi-scale attention module.

[0077] Appendix Figure 2 The middle image (a) represents the input image, which shows a local view of a wheat field. The image contains wheat leaves and ears, exhibiting high complexity and detail. Despite some overlap and occlusion of the leaves and ears in the image, the model is still able to extract effective information from these complex features.

[0078] Appendix Figure 2Figure (b) shows the output of the multi-scale attention module. It can be seen that after processing by the multi-scale attention module, the wheat ear region in the image is significantly enhanced, with the red and orange areas accurately corresponding to the high-temperature regions of the wheat ear. This indicates that the model can successfully focus on key areas and highlight the temperature characteristics of the wheat ear. Even when parts of the leaves and ears overlap in the image, the model can still effectively distinguish most of the ear region, demonstrating the advantages of the multi-scale attention mechanism in detail capture.

[0079] The current module successfully identified and highlighted wheat ears in most cases, especially in densely growing environments, demonstrating strong robustness and adaptability. Although there were slight misidentifications in some detailed areas, the overall performance was excellent, with high predictive accuracy.

[0080] In one embodiment of the present invention, step S1 includes the following steps:

[0081] S1-1. Select a wheat experimental field with an area of ​​M, and randomly set up [various species] within the experimental field. A 10-square-meter large-scale square plot ,in For the first A large square sample, .

[0082] S1-2. During the critical growth period of wheat, 5-7 days after heading, data were collected using drones. A large square sample The image of a small square with a center area of ​​1 square meter is obtained. Image of a small square sample .

[0083] S1-3. After the drone has taken images of all the square sample plots, adjust the drone's flight altitude so that its field of view covers the [number missing]. A large square sample The first one was collected. Image of a large square .

[0084] In this embodiment, specifically, in step S1-1, M is greater than or equal to 1000 square meters; during the critical growth period of wheat (5-7 days after heading), on a sunny day with a wind speed of less than or equal to 5 m / s, between 9:00 and 11:00 AM, the wheat experimental field is photographed using a Zenmuse P1 camera mounted on a DJI M300 RTK drone with the gimbal tilt angle set to -90°; in step S1-1... The value is 100. The large square plots do not overlap.

[0085] In one embodiment of the present invention, in step S2, the first A large square sample A central 1-square-meter area was designated as a statistical quadrat. The number of ears in each quadrat was manually counted, and the result was multiplied by 666.67 to obtain the number of ears. The actual number of ears per mu .

[0086] In one embodiment of the present invention, step S3 includes the following steps:

[0087] S3-1. The first Image of a small square sample The cv2.createCLAHE() function in OpenCV is used to enhance the contrast, resulting in the enhanced contrast level. Image of a small square sample , will the Image of a small square sample A Gaussian filter with a kernel size of 3×3 was used for smoothing and denoising to ensure that the wheat ear features were clearly visible, resulting in the preprocessed [preprocessed] [earth ear]. Image of a small square sample ,all The images of the preprocessed square minima constitute the image set of the preprocessed minima. , .

[0088] S3-2. The first Image of a large square The cv2.createCLAHE() function in OpenCV is used to enhance the contrast, resulting in the enhanced contrast level. Image of a large square , will the Image of a large square Smoothing and denoising are performed using a Gaussian filter with a kernel size of 3×3, resulting in the preprocessed i.e. Image of a large square ,all The images of the preprocessed square matrices constitute the image set of the preprocessed matrices. , .

[0089] S3-3. Constructing Data Pairs ,all The WESN dataset consists of several data pairs. .

[0090] S3-4. Divide the WESN dataset into training and test sets in a 7:3 ratio.

[0091] In one embodiment of the present invention, step S5 includes the following steps:

[0092] S5-1. The first variational encoder of the wheat ear number prediction model consists of a first convolutional layer, a first ReLU activation function, a second convolutional layer, a second ReLU activation function, a third convolutional layer, a third ReLU activation function, a fourth convolutional layer, and a fourth ReLU activation function. Specifically, the first convolutional layer has a kernel size of 3×3, 32 channels, and a stride of 1; the second convolutional layer has a kernel size of 3×3, 64 channels, and a stride of 2; and the third convolutional layer has a kernel size of 3×3, 128 channels, and a stride of 2.

[0093] S5-2. The preprocessed i-th training set Image of a small square sample The input is fed into the first variational encoder, and the output is the first encoded feature. .

[0094] S5-3. The second variational encoder of the wheat ear number prediction model consists of a spatial mapping layer, a first fully connected layer, a first ReLU activation function, a second fully connected layer, and a second ReLU activation function, which encodes the first feature. The input to the spatial mapping layer is flattened into a vector using the `flatten()` function of the PyTorch library. The flattened vector is then reduced to 512 dimensions using the `nn.Linear()` function of the PyTorch library to obtain the features. , will feature The inputs are sequentially fed into the first fully connected layer, the first ReLU activation function, the second fully connected layer, and the second ReLU activation function of the second variational encoder, and the output is the second encoded feature. Specifically, the first fully connected layer has an input dimension of 64 and an output dimension of 128, and the second fully connected layer has an input dimension of 128 and an output dimension of 128.

[0095] S5-4. The third variational encoder of the wheat ear number prediction model consists of a fully connected layer, a reshape function, a first convolutional layer, a second convolutional layer, and a global average pooling layer.

[0096] S5-5. The second coding feature The input is fed into the fully connected layer of the third variational encoder, and the output is a tensor. , tensor The input is fed into the Reshape function of the third variational encoder, and the output is a tensor of size 16×16×32. , the first encoded feature With tensor By concatenating the features along the channel dimension, a fused feature map is obtained. .

[0097] S5-6. Merge feature maps The inputs are sequentially fed into the first convolutional layer, the second convolutional layer, and the global average pooling layer of the third variational encoder, and the output is the third encoded feature. Specifically, the first convolutional layer has a kernel size of 3×3, 128 channels, and a stride of 1, while the second convolutional layer has a kernel size of 3×3, 64 channels, and a stride of 1.

[0098] In one embodiment of the present invention, step S6 includes the following steps:

[0099] S6-1. The uncertainty module of the wheat ear number prediction model consists of a feature fusion layer and an uncertainty distribution modeling layer.

[0100] S6-2. The preprocessed first... Image of a large square and second coding features The input is fed into the feature fusion layer of the uncertainty module, through the formula. Calculated features In the formula, For a randomly initialized extended weight matrix, As a bias term, the features The input is fed into the reshape function of the PyTorch library, and the output is the feature. , will feature Compared with the preprocessed first Image of a large square The channel-dimensional concat() function from the PyTorch library is used to perform concatenation operations to obtain fused features. Specifically, , For the real number space, For the preprocessed first Image of a large square of high, , For the preprocessed first Image of a large square width, , For the preprocessed first Image of a large square The number of channels, , , , , For bias terms The number of channels, .

[0101] S6-3. Merging Features The uncertainty distribution modeling layer input into the uncertainty module uses the flatten() function from the PyTorch library to flatten and reduce dimensionality, resulting in a vector. Through formula Calculate the mean vector , It is a linear mapping matrix. As the bias vector, it is obtained through the formula The variance vector is calculated. In the formula It is a linear mapping matrix. This is the bias vector. Specifically, , , , , , , , , .

[0102] S6-4. Through formula Uncertainty characteristics were calculated. In the formula For element-wise multiplication, Standard normal noise generated for the randn() function in the PyTorch library.

[0103] Because drones are affected by wind speed and gimbal stability during filming, especially when filming large sample plots of 10m², the error is significant. Due to the large area, even slight shifts in the camera's perspective can lead to large deviations in the final captured area. Since the area of ​​the sample plot directly affects the calculation of the final number of ears per acre, a global area calculation module is needed to accurately calculate the area of ​​the filmed sample plot. Therefore, in one embodiment of this invention, step S7 includes the following steps:

[0104] S7-1. Through formula The local area estimate is calculated. In the formula For the preprocessed first Image of a small square sample The Middle Line number The pixel values ​​of the columns of pixels. For the preprocessed first Image of a small square sample pixel domain, For drones to collect the first Image of a small square sample Flight altitude relative to the ground For the first Image of a small square sample Horizontal pixel count, For the first Image of a small square sample Vertical pixel count, For drones to collect the first Image of a small square sample The focal length of the camera at that time. For drones to collect the first Image of a small square sample The roll angle at that time For drones to collect the first Image of a small square sample The pitch angle at that time.

[0105] S7-2. Through formula The global area estimate is calculated. In the formula For the processed first Image of a large square The Middle Line number The pixel values ​​of the columns of pixels. For the processed first Image of a large square pixel domain, For drones to collect the first Image of a large square Flight altitude relative to the ground For the first Image of a large square Horizontal pixel count, For the first Image of a large square Vertical pixel count, For drones to collect the first Image of a large square The focal length of the camera at that time. For drones to collect the first Image of a large square The roll angle at that time For drones to collect the first Image of a large square The pitch angle at that time.

[0106] In one embodiment of the present invention, step S8 includes the following steps:

[0107] S8-1. The third coding feature Uncertainty characteristics The input is fed into the multi-scale attention module of the wheat ear number prediction model via the formula.

[0108] Calculate attention features In the formula For the third coding feature Uncertainty characteristics the number of rows, For the third coding feature Uncertainty characteristics The number of columns, For the third coding feature The Line number The parameter values ​​of the column, , , Uncertainty characteristics The Line number The parameter values ​​of the column, Uncertainty characteristics The Line number The parameter values ​​of the column, , The smoothing parameter controls the distribution of the weighting matrix, preventing overly sharp weighting differences. Specifically... , For The index with a base.

[0109] In one embodiment of the present invention, step S9 includes the following steps:

[0110] The multi-scale density decoding module of the S9-1 wheat ear number prediction model consists of a first fully connected layer, a first ReLU activation function, a second fully connected layer, a second ReLU activation function, a third fully connected layer, a third ReLU activation function, and a weight fusion layer. Specifically, the output dimension of the first fully connected layer is 256, the output dimension of the second fully connected layer is 64, and the output dimension of the third fully connected layer is 1.

[0111] S9-2. Attention Features The inputs are sequentially fed into the first fully connected layer, the first ReLU activation function, the second fully connected layer, the second ReLU activation function, the third fully connected layer, and the third ReLU activation function of the multi-scale density decoding module, and the output yields the initial number of ears per acre. That is, the predicted number of ears.

[0112] S9-3. Initial number of ears per mu The weight fusion layer of the multi-scale density decoding module is input through the formula Calculate the number of wheat ears per mu 666.67 represents the area of ​​one mu (approximately 0.16 acres). The logarithmic scale of the global area represents the weights that need to be considered when calculating the number of ears per mu (unit of land area) in the global image state based on the local ear count in the current local image.

[0113] In summary, the process of predicting the number of ears per acre from an image begins with extracting features from the image using a wheat ear number prediction model. These features are then processed through multiple layers, and the model maps them to an output value—the predicted number of ears per acre. The wheat ear number prediction model is optimized using a loss function to ensure the predicted value is as close as possible to the true value. Ultimately, the trained model can output an accurate prediction of the number of ears per acre based on the information in the image.

[0114] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. An adaptive wheat ear count prediction method based on variational autoencoder, characterized in that, include: S1. Obtaining data from the wheat experimental field using drones. A collection of images of small square samples and A collection of images of a large square sample , ,in For the first An image of a small square sample, , ,in For the first An image of a large square sample; S2. Obtain the actual set of ears per mu (unit of land area) data. , , For the first The actual number of ears per mu; S3. Preprocess the images of the large square quadrats and the small square quadrats to obtain a set of preprocessed large square image sets. and the image set of preprocessed small sample plots According to the image set and image collection Create the WESN dataset and divide it into a training set and a test set; S4. Construct a wheat ear number prediction model consisting of a first variational encoder, a second variational encoder, a third variational encoder, an uncertainty module, a global area calculation module, a multi-scale attention module, and a multi-scale density decoding module; S5. Input the preprocessed images of small square quadrats from the training set into the first variational encoder, second variational encoder, and third variational encoder of the wheat ear number prediction model, and output the first encoded features respectively. Second coding features Third coding feature ; S6. Combine the preprocessed images of large square quadrats in the training set with the second encoded features. The input is given to the uncertainty module of the wheat ear number prediction model, and the output is the uncertainty features. ; S7. Input the preprocessed images of small square quadrats and the preprocessed images of large square quadrats from the training set into the global area calculation module of the wheat ear number prediction model, and output the local area estimates respectively. and global area estimation ; S8. Encode the third feature Uncertainty characteristics The input is fed into the multi-scale attention module of the wheat ear number prediction model, and the output is the attention feature. ; S9. Attention Features The data is input into the multi-scale density decoding module of the wheat ear number prediction model, and the output is the wheat ear number per mu (unit of land area). ; S10. Based on the number of wheat ears per mu (unit of land area) With the The actual number of ears per mu Construct a mean squared error loss function, use the Adam optimizer, and train the wheat ear number prediction model using the mean squared error loss function to obtain the optimized wheat ear number prediction model. S11. Input the preprocessed images of small square quadrats and large square quadrats from the test set into the optimized wheat ear number prediction model, and output the wheat ear number per mu (unit of land area). ; S9-1. The multi-scale density decoding module of the wheat ear number prediction model consists of a first fully connected layer, a first ReLU activation function, a second fully connected layer, a second ReLU activation function, a third fully connected layer, a third ReLU activation function, and a weight fusion layer. S9-2. Attention Features The inputs are sequentially fed into the first fully connected layer, the first ReLU activation function, the second fully connected layer, the second ReLU activation function, the third fully connected layer, and the third ReLU activation function of the multi-scale density decoding module, and the output yields the initial number of ears per acre. ; S9-3. Initial number of ears per mu The weight fusion layer of the multi-scale density decoding module is input into the formula. Calculate the number of wheat ears per mu .

2. The adaptive wheat ear count prediction method based on variational autoencoder according to claim 1, characterized in that, Step S1 includes the following steps: S1-1. Select a wheat experimental field with an area of ​​M, and randomly set up [various species] within the experimental field. A 10-square-meter large-scale square plot ,in For the first A large square sample, ; S1-2. During the critical growth period of wheat, 5-7 days after heading, data were collected using drones. A large square sample The image of a small square with a center area of ​​1 square meter is obtained. Image of a small square sample ; S1-3. After the drone has taken images of all the square sample plots, adjust the drone's flight altitude so that its field of view covers the [number missing]. A large square sample The first one was collected. Image of a large square .

3. The adaptive wheat ear count prediction method based on variational autoencoder according to claim 2, characterized in that: In step S1-1, M is greater than or equal to 1000 square meters; during the critical growth period of wheat (5-7 days after heading), on a sunny day with a wind speed of less than or equal to 5 m / s, between 9:00 and 11:00 AM, the wheat experimental field is photographed using a Zenmuse P1 camera mounted on a DJI M300 RTK drone with the gimbal tilt angle set to -90°; in step S1-1 The value is 100. The large square plots do not overlap.

4. The adaptive wheat ear count prediction method based on variational autoencoder according to claim 2, characterized in that: In step S2, at the A large square sample A central 1-square-meter area was designated as a statistical quadrat. The number of ears in each quadrat was manually counted, and the result was multiplied by 666.67 to obtain the number of ears. The actual number of ears per mu .

5. The adaptive wheat ear count prediction method based on variational autoencoder according to claim 1, characterized in that, Step S3 includes the following steps: S3-1. The first Image of a small square sample The cv2.createCLAHE() function in OpenCV is used to enhance the contrast, resulting in the enhanced contrast level. Image of a small square sample , will the Image of a small square sample Smoothing and denoising are performed using a Gaussian filter with a kernel size of 3×3, resulting in the preprocessed i.e. Image of a small square sample ,all The images of the preprocessed square minima constitute the image set of the preprocessed minima. , ; S3-2. The first Image of a large square The cv2.createCLAHE() function in OpenCV is used to enhance the contrast, resulting in the enhanced contrast level. Image of a large square , will the Image of a large square Smoothing and denoising are performed using a Gaussian filter with a kernel size of 3×3, resulting in the preprocessed i.e. Image of a large square ,all The images of the preprocessed square matrices constitute the image set of the preprocessed matrices. , ; S3-3. Constructing Data Pairs ,all The WESN dataset consists of several data pairs. ; S3-4. Divide the WESN dataset into training and test sets in a 7:3 ratio.

6. The adaptive wheat ear count prediction method based on variational autoencoder according to claim 5, characterized in that, Step S5 includes the following steps: S5-1. The first variational encoder of the wheat ear number prediction model consists of a first convolutional layer, a first ReLU activation function, a second convolutional layer, a second ReLU activation function, a third convolutional layer, a third ReLU activation function, a fourth convolutional layer, and a fourth ReLU activation function. S5-2. The preprocessed i-th element in the training set Image of a small square sample The input is fed into the first variational encoder, and the output is the first encoded feature. ; S5-3. The second variational encoder of the wheat ear number prediction model consists of a spatial mapping layer, a first fully connected layer, a first ReLU activation function, a second fully connected layer, and a second ReLU activation function, which encodes the first feature. The input to the spatial mapping layer is flattened into a vector using the `flatten()` function of the PyTorch library. The flattened vector is then reduced to 512 dimensions using the `nn.Linear()` function of the PyTorch library to obtain the features. , will feature The inputs are sequentially fed into the first fully connected layer, the first ReLU activation function, the second fully connected layer, and the second ReLU activation function of the second variational encoder, and the output is the second encoded feature. ; S5-4. The third variational encoder of the wheat ear number prediction model consists of a fully connected layer, a reshape function, a first convolutional layer, a second convolutional layer, and a global average pooling layer. S5-5. The second coding feature The input is fed into the fully connected layer of the third variational encoder, and the output is a tensor. , tensor The input is fed into the Reshape function of the third variational encoder, and the output is a tensor. , the first encoded feature With tensor By concatenating the features along the channel dimension, a fused feature map is obtained. ; S5-6. Merge feature maps The inputs are sequentially fed into the first convolutional layer, the second convolutional layer, and the global average pooling layer of the third variational encoder, and the output is the third encoded feature. .

7. The adaptive wheat ear count prediction method based on variational autoencoder according to claim 1, characterized in that, Step S6 includes the following steps: S6-1. The uncertainty module of the wheat ear number prediction model consists of a feature fusion layer and an uncertainty distribution modeling layer; S6-2. The preprocessed first... Image of a large square and second coding features The input is fed into the feature fusion layer of the uncertainty module, through the formula. Calculated features In the formula, For a randomly initialized extended weight matrix, As a bias term, the features The input is fed into the reshape function of the PyTorch library, and the output is the feature. , will feature Compared with the preprocessed first Image of a large square The channel-dimensional concat() function from the PyTorch library is used to perform concatenation operations to obtain fused features. ; S6-3. Merging Features The uncertainty distribution modeling layer input into the uncertainty module uses the flatten() function from the PyTorch library to flatten and reduce dimensionality, resulting in a vector. Through formula Calculate the mean vector , It is a linear mapping matrix. As the bias vector, it is obtained through the formula The variance vector is calculated. In the formula It is a linear mapping matrix. It is the bias vector; S6-4. Through formula Uncertainty characteristics were calculated. In the formula For element-wise multiplication, Standard normal noise generated for the randn() function in the PyTorch library.

8. The adaptive wheat ear count prediction method based on variational autoencoder according to claim 5, characterized in that, Step S7 includes the following steps: S7-1. Through formula The local area estimate is calculated. In the formula For the preprocessed first Image of a small square sample The Middle Line 1 The pixel values ​​of the columns of pixels. For the preprocessed first Image of a small square sample pixel domain, For drones to collect the first Image of a small square sample Flight altitude relative to the ground For the first Image of a small square sample Horizontal pixel count, For the first Image of a small square sample Vertical pixel count, For drones to collect the first Image of a small square sample The focal length of the camera at that time. For drones to collect the first Image of a small square sample The roll angle at that time For drones to collect the first Image of a small square sample The pitch angle at that time; S7-2. Through formula The global area estimate is calculated. In the formula For the processed first Image of a large square The Middle Line 1 The pixel values ​​of the columns of pixels. For the processed first Image of a large square pixel domain, For drones to collect the first Image of a large square Flight altitude relative to the ground For the first Image of a large square Horizontal pixel count, For the first Image of a large square Vertical pixel count, For drones to collect the first Image of a large square The focal length of the camera at that time. For drones to collect the first Image of a large square The roll angle at that time For drones to collect the first Image of a large square The pitch angle at that time.

9. The adaptive wheat ear count prediction method based on variational autoencoder according to claim 1, characterized in that, Step S8 includes the following steps: S8-1. The third coding feature Uncertainty characteristics The input is fed into the multi-scale attention module of the wheat ear number prediction model via the formula. Calculate attention features In the formula For the third coding feature Uncertainty characteristics the number of rows, For the third coding feature Uncertainty characteristics The number of columns, For the third coding feature The Line 1 The parameter values ​​of the column, , , Uncertainty characteristics The Line 1 The parameter values ​​of the column, Uncertainty characteristics The Line 1 The parameter values ​​of the column, , For smoothing parameters.