Potato yield prediction model construction method based on multispectrum and machine learning

Through multi-spectral and machine learning methods, combining spectral, texture and structural characteristics, a potato yield prediction model is constructed, which solves the problem of inaccurate yield estimation in traditional methods and achieves more efficient yield prediction.

CN120373571APending Publication Date: 2025-07-25INSTITUTE OF VEGETABLES & FLOWERS CHINESE ACADEMY OF AGRICULTURAL SCIENCES
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510788759.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Traditional potato yield estimation methods ignore crop growth status, resulting in incomplete spatial coverage, time-consuming and labor-intensive, and can only obtain single point information, which cannot effectively improve land output and resource utilization.

Method used

A potato yield prediction model based on multi-spectral and machine learning is adopted. Through multi-source data acquisition and processing, spectral features, texture features and structural features are extracted, combined with canopy coverage, a stable yield prediction model is constructed, and uncertainty analysis is performed.

Benefits of technology

It effectively reduces the uncertainty caused by model input, improves the stability and accuracy of the model, and improves the accuracy of potato yield prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373571A_ABST
    Figure CN120373571A_ABST
Patent Text Reader

Abstract

The invention discloses a potato yield prediction model construction method based on multispectrum and machine learning, and belongs to the technical field of crop yield prediction, and the method comprises the following steps: S1, collecting and processing multi-source data; s2, preliminarily constructing a potato yield prediction model, and performing uncertainty analysis on the potato yield prediction model; and S3, based on the uncertainty analysis result obtained in the S2, fusing the multi-source data collected in the S1 to construct an optimal potato yield prediction model. According to the potato yield prediction model construction method based on multispectral and machine learning, uncertainty caused by model input can be effectively reduced, model stability can be improved by combining spectral features with texture features, and model precision can be further improved by fusing canopy coverage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of crop yield prediction, and in particular to a method for constructing a potato yield prediction model based on multispectral and machine learning. Background Art

[0002] Currently, humanity faces serious problems such as climate change, population growth, and resource shortages. As the world's third-largest food crop, improving the yield and resource utilization efficiency of potatoes is of great significance for ensuring food security and sustainable agricultural development. Agriculture faces multiple pressures such as resource shortages, deteriorating ecological environment, and aging labor force (Liao et al, 2019), and it is necessary to significantly improve land productivity, labor productivity, and resource utilization rate. Traditional methods estimate yields through planting area and limited field sampling. The former ignores the impact of the crop's own growth status on yield, and the latter is time-consuming, laborious, destructive, and can only obtain single-point information with incomplete spatial coverage. Summary of the Invention

[0003] The purpose of the present invention is to provide a method for constructing a potato yield prediction model based on multispectral and machine learning, which can effectively reduce the uncertainty caused by model input. The combination of spectral features and texture features can improve the stability of the model, and the fusion of canopy coverage can further improve the model accuracy.

[0004] To achieve the above purpose, the present invention provides a method for constructing a potato yield prediction model based on multispectral and machine learning, including the following steps: S1. Acquisition and processing of multi-source data; S1.1. Determination of the experimental site and experimental design; S1.2. Ground data measurement; S1.3. UAV data acquisition and processing; S1.4. Spectral feature extraction; In spectral feature extraction, a soil mask is removed from the multispectral background by optimizing the Optimized Soil-Adjusted Vegetation Index (OSAVI), and the calculation formula is as follows: ; Wherein, is the reflectance in the near-infrared band, is the reflectance in the red light band, is an adjustment factor used to adjust the influence of the soil, and its value is 0.16; S1.5. Texture feature extraction; S1.6. Structure feature extraction; S2. Initially construct a potato yield prediction model and perform uncertainty analysis on the potato yield prediction model; S3. Based on the uncertainty analysis results obtained in S2, fuse the multi-source data collected in S1 to construct the best potato yield prediction model.

[0005] Preferably, the structural features in S1.6 include plant height CH, canopy cover CC, and canopy growth curve CGCI; The extraction of CH includes three steps: constructing a digital terrain model DTM, calculating a crop height model CHM, and calculating plant height CH; CC is the ratio of vegetation cover per unit land area. Use OSAVI to segment soil and vegetation, and calculate CC by statistically analyzing the ratio of vegetation pixels to total pixels. Construct a mask through the vegetation extraction color index to remove the soil background. The calculation formula is as follows: ; Among them, , , represent the DN values of the red, green, and blue channels of the digital camera respectively; In CGCI, the growth dynamics of the potato canopy conforms to the three-stage law. Use the thermal days TD to describe the crop growth and development process driven by accumulated temperature. The calculation formula is as follows: ; Among them, is the daily average temperature, is the minimum temperature for crop growth, is the optimum temperature for crop growth, is the maximum temperature for crop growth, is the temperature response curvature coefficient.

[0006] Preferably, the specific operation steps of S2 include: S2.1. Modeling sample division; S2.2. Modeling feature selection; S2.3. Model structure screening.

[0007] Preferably, the specific operation steps of S2.2 include: S2.2.1. Feature selection algorithm; The filters include the Pearson correlation coefficient PCC, mutual information MI, grey relational analysis GRA based on single variables, and RReliefF based on multiple variables. The recursive feature elimination algorithm RFE based on support vector regression is also selected as the representative of the wrapper; PCC is the quotient of the covariance and standard deviation between two variables. The calculation formula is as follows: ; Among them, and respectively represent the One sample is the number of samples, and the value range of PCC is . When PCC is closer to 1, it indicates a strong positive correlation between two variables. When PCC is closer to -1, it indicates a strong negative correlation between two variables. When the value of PCC is 0, the two variables are independent of each other; MI is a measure of the mutual dependence between two variables, used to quantify the "information content" obtained about another random variable by observing one random variable. The calculation formula is as follows: ; where is and 's joint probability density distribution function, and are respectively and 's marginal probability density functions, and represent two different variables respectively. The value range of MI is . When the value of MI is 0, the two variables are independent of each other. The larger the MI, the more information the feature gives about the target variable; GRA is a multi-factor statistical analysis method, which judges whether the connection between two variables is close by the similarity degree of the geometric shapes of the sequence curves, and is evaluated by the correlation coefficient. The calculation formula is as follows: ; where and are respectively the minimum value and the maximum value of the reference sequence, and are respectively the minimum value and the maximum value of the comparison sequence, is the point of the reference sequence, is the point of the comparison sequence, is an adjustable coefficient, and its value is , used to adjust the gap size of the output results of different features. When takes the value of 0.5, the correlation coefficient takes values between 0 and 1, the larger the value of , the closer the connection between the feature and the target variable in the given feature; RFE is a wrapper-based feature selection method that iteratively removes the least relevant features based on feature importance from all features until a specific number of feature sets are selected; S2.2.2. Feature selection uncertainty analysis; The extended Quincheva similarity metric index EK is used to evaluate the stability between two groups of feature selection results: ; where is the number of common features, and are the feature numbers of the two groups of features respectively, and is the total number of features.

[0008] Preferably, the specific operation steps of S2.3 include: S2.3.1. Data preprocessing; S2.3.2. Hyperparameter tuning; S2.3.3. Model operation and uncertainty analysis.

[0009] Therefore, the method for constructing a potato yield prediction model based on multispectral and machine learning of the present invention can effectively reduce the uncertainty caused by model input, and the combination of spectral features and texture features can improve the model stability, and the fusion of canopy coverage can further improve the model accuracy.

[0010] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings

[0011] Figure 1 is the flowchart of the present invention; Figure 2 is the UAV remote sensing system of the present invention, wherein Figure 2 in (a) is a quadcopter UAV, Figure 2 in (b) is a multispectral camera, Figure 2 in (c) is a ground reference board, Figure 2 in (d) is a radiometric calibration board; Figure 3 is the process diagram of remote sensing image preprocessing of the present invention; Figure 4 is the growth dynamic curve diagram of the potato canopy of the present invention; Figure 5 is the yield diagram of different years of the present invention, wherein Figure 5 in (a) is the yield of each variety in 2022, Figure 5 in (b) is the yield of each variety in 2023, Figure 5 in (c) is the yield in 2022 and 2023; Figure 6 is the R of different sampling methods in the P3 period of the present invention 2 and the probability density function of RMSE, where Figure 6 in (a) is the R of different sampling methods in the P3 period 2 probability density function Figure 6 in (b) is the probability density function of RMSE of different sampling methods in the P3 period; Figure 7 is the modeling R of different sampling methods based on SLR of the present invention 2 distribution comparison, where Figure 7 in (a) is the model R 2 median Figure 7 in (b) is the model R 2 interquartile range Figure 7 in (c) is the model R 2 coefficient of variation; Figure 8 is the uncertainty analysis of feature selection based on data perturbation of the present invention; Figure 9 is the test result of different feature selection algorithms Lasso of the present invention; Figure 10 is the yield prediction accuracy of different machine learning models of the present invention, where Figure 10 in (a) is the yield prediction accuracy of SLR Figure 10 in (b) is the yield prediction accuracy of Lasso Figure 10 in (c) is the yield prediction accuracy of RF Figure 10 in (d) is the yield prediction accuracy of PLSR Figure 10 in (e) is the yield prediction accuracy of Ridge Figure 10 in (f) is the yield prediction accuracy of SVM; Figure 11 is the R under different feature quantity conditions of the present invention 2 uncertainty comparison, where Figure 11 in (a) is the model R 2 median Figure 11 in (b) is the model R 2 interquartile range Figure 11 in (c) is the model R 2 coefficient of variation; Figure 12 is the RMSE uncertainty comparison under different feature quantity conditions of the present invention, where Figure 12 in (a) is the average value of model RMSE Figure 12 in (b) is the interquartile range of model RMSE Figure 12 in (c) is the coefficient of variation of model RMSE Figure 13 They are the vegetation indices ranked in the top 20 in terms of the importance score of each period's modeling features of the present invention. Among them, Figure 13 (a) in it is the vegetation index ranked in the top 20 in terms of the importance score in the P1 period, Figure 13 and (b) in it is the vegetation index ranked in the top 20 in terms of the importance score in the P2 period, Figure 13 and (c) in it is the vegetation index ranked in the top 20 in terms of the importance score in the P3 period, Figure 13 and (d) in it is the vegetation index ranked in the top 20 in terms of the importance score in the P4 period, Figure 13 and (e) in it is the vegetation index ranked in the top 20 in terms of the importance score in the P5 period, Figure 13 and (f) in it is the vegetation index ranked in the top 20 in terms of the importance score in the P6 period. Detailed implementation mode

[0012] The technical solution of the present invention will be further described below with reference to the drawings and embodiments.

[0013] Unless otherwise defined, the technical terms or scientific terms used in the present invention shall have the ordinary meanings understood by those of ordinary skill in the field to which the present invention belongs.

[0014] Embodiment 1 As Figure 1 shown, the present invention provides a method for constructing a potato yield prediction model based on multispectral and machine learning, including the following steps: S1. Acquisition and processing of multi-source data; The specific operation steps include: S1.1. Determination of the experimental site and experimental design; Experimental site: From 2022 to 2023, three field experiments were carried out at the potato test and demonstration base of the Institute of Vegetables and Flowers, Chinese Academy of Agricultural Sciences, Chabei Management District, Zhangjiakou City, Hebei Province (115°3'38"E; 41°28'50"N). The base is located in the northwest of Hebei Province, belonging to the northern single-cropping area, with an East Asian continental monsoon climate, an altitude of 1390 m, an annual rainfall of 381.4 mm, and 2931.7 hours of sunshine throughout the year.

[0015] Three field cultivation experiments were designed and managed by using the precise micro-drop irrigation technology of high-ridge water and fertilizer integration.

[0016] Experiment 1: The "variety × nitrogen form" experiment adopted a split-plot design. The tested varieties were the mid-early variety Xiabodi and the late variety Zhongshu 18. Two nitrogen forms, nitrate nitrogen (calcium magnesium nitrate) and ammonium nitrogen (ammonium sulfate), were set, and three pure nitrogen (N) concentration gradients were set for each nitrogen form, namely 0 kg / ha (F1), 150 kg / ha (F2, NO3 - ; F4, NH4 +), and 300 kg / ha (F3, NO3 - ; F5, NH4 + ), with 4 replicates, totaling 40 plots. Each plot has 6 ridges, with a ridge spacing of 0.9 m. In 2022, the ridge length is 8 m, and in 2023, the ridge length is 10 m.

[0017] Experiment 2: The "Variety × Nitrogen Application Rate" experiment adopted a split-plot design. The tested varieties were the early-maturing variety Zhongshu 5 and the mid-late-maturing variety Zhongshu 49. Five nitrogen levels were set, namely N1 (0 kg / ha), N2 (50 kg / ha), N3 (100 kg / ha), N4 (200 kg / ha), and N5 (300 kg / ha), with 4 replicates, totaling 40 plots. Each plot has 6 ridges, with a ridge spacing of 0.9 m. In 2022, the ridge length is 14 m, and in 2023, the ridge length is 8 m.

[0018] Experiment 3: The "Variety × Density" experiment adopted a randomized block design. The tested varieties were the mid-maturing variety Zhongshu 27 and the late-maturing variety Zhongshu 19. Three density gradients of 3500 plants / mu, 4000 plants / mu, and 4500 plants / mu were set, with 3 replicates. The ridge length is 14 m, and the ridge spacing is 0.9 m, totaling 18 plots. The seed potatoes were original seeds or first-class seeds.

[0019] S1.2, Ground data measurement; When harvesting after the potatoes mature, during yield measurement, the plot protection rows were excluded, and the potatoes in the middle 5 m of each plot were picked up and weighed for fresh weight. The yield per unit area was calculated based on the harvested area.

[0020] S1.3, UAV data collection and processing; The quadcopter UAV DJI Inspire 2 was used to continuously collect canopy remote sensing images throughout the growth period, equipped with the RGB camera Zenmuse X5S and the multispectral sensor RedEdge-P respectively. The RGB camera has a resolution of up to 20.88 million pixels. When the flight altitude is 20 m, the spatial resolution is 0.44 cm / pixel. The multispectral includes 5 narrowband spectral bands and 1 panchromatic band, with a relatively low spatial resolution. When stitching images, the other narrowband spectral bands can be filtered by the higher-resolution panchromatic band to increase the number of pixels to 5.06 million, as shown in Table 1. When the flight altitude is 30 m, the spatial resolution of the multispectral is 1.89 cm / pixel. In addition, RedEdge-P is equipped with a sunlight sensor (Downwelling Light Sensor 2, DLS2) to record the solar altitude angle and ambient light during flight and correct the local brightness differences caused by changes in ambient light during flight.

[0021] Table 1 Related parameters of Zenmuse X5S and RedEdge-P ;

[0022] To ensure the image quality of each flight and reduce the influence of shadows, the flight mission is carried out at noon from 11:00 to 13:00 on a sunny day without clouds, wind or with gentle breeze. The unmanned aerial vehicle (UAV) remote sensing system is as Figure 2 shown, where Figure 2 in (a) is a quadrotor UAV, Figure 2 in (b) is a multispectral camera, Figure 2 in (c) is a ground reference board, Figure 2 and in (d) is a radiometric calibration board. When collecting RGB images, the flight altitude is set to 20 m, the forward overlap is 75%, the side overlap is 75%, and the camera angle is vertically downward (90°); when collecting multispectral images, the flight altitude is 30 m, and the radiometric calibration board is manually photographed before and after the flight. Other parameter settings are the same as those for RGB.

[0023] The preprocessing steps of remote sensing images include radiometric calibration, geometric calibration, image mosaicking and rotation, and cropping of the region of interest (ROI). It is completed using Agisoft Metashape 2.0.2 and ArcGIS 10.8, as Figure 3 shown.

[0024] First, align the photos according to the UAV attitude and position information, generate a mesh and a digital orthophoto map (DOM), and complete a rough mosaic. Then import the real-time kinematic GPS (RTK-GPS) data of the ground control points (GCPs) for manual point piercing to complete geometric calibration. After optimizing the alignment method, generate a dense point cloud, a digital elevation model (DEM), and a DOM in sequence. The geometric calibration method for multispectral images is the same as that for RGB. Before mosaicking, radiometric calibration is required according to the reference board manually photographed during the flight mission, and the panchromatic band image is set as the main channel to improve the resolution of each band. When exporting the DOM and DEM, calculate the reflectance of each band through the digital number (DN) value of the image, and set the projection coordinate system as "WGS1984 UTM Zone 50N" according to the geographical location of the base. Use the rotation tool in ArcGIS 10.8 to rotate the raster image counterclockwise by a certain angle for batch cropping and feature extraction. Construct the vector graphic file of each plot, create a new field and name each plot, and after separating each plot element, construct a batch cropping tool containing an iterator to crop the ROI image.

[0025] S1.4, Spectral feature extraction; The spectral features of remote sensing images mainly include the reflectance of each band collected by multispectral imaging and the vegetation index (VI) calculated based on band operations. After obtaining the ROI images of the plots in S1.3, the threshold segmentation method is used to remove the soil background. In spectral feature extraction, an optimized soil-adjusted vegetation index OSAVI is used to construct a soil mask to remove the multispectral background, and the threshold is adjusted according to the background removal effect. The final threshold is set to 0.45. The calculation formula is as follows: ; where, is the reflectance of the near-infrared band, is the reflectance of the red band, is the adjustment factor used to adjust the influence of the soil, usually taking 0.16; the reflectance of the five bands of each plot ROI is calculated by the average value of the vegetation pixels in each channel.

[0026] S1.5, Texture feature extraction; After converting each band image into a Grey Level Co-occurrence Matrix (GLCM), eight texture features are constructed through second-order statistics, as shown in Table 2.

[0027] Table 2 Texture features and their calculation formulas ;

[0028] Table 2 Texture features and their calculation formulas ;

[0029] S1.6, Structural feature extraction; Structural features include plant height CH, canopy coverage CC, and canopy growth curve CGCI; The extraction of CH includes three steps: constructing a digital terrain model DTM, calculating a crop height model CHM, and calculating plant height CH; first, point features are created at the bare soil locations and assigned values based on the DEM, and the DTM is interpolated by the Inverse Distance Weighting (IDW) method. CHM is obtained by subtracting the DTM from the DEM, and by calculating the statistics of CHM, such as the 99% quantile, the plant height can be obtained. The extraction of plant height requires high-precision spatial coordinate information, and the accuracy of the DTM directly affects the simulation accuracy. The DTM is constructed by combining the pre-emergence bare soil images with GCPs to calculate CHM, and the plant height is calculated by the 99% quantile of CHM.

[0030] CC is the ratio of vegetation cover per unit land area. OSAVI (threshold 0.45) is used to segment soil and vegetation, and CC is calculated by statistically analyzing the ratio of vegetation pixels to the total number of pixels. Since the RGB image does not have a near-infrared band, only digital indices within the visible light range can be selected to segment soil and vegetation. The Color Index of Vegetation Extraction (CIVE) relies on color information to distinguish vegetation from other land cover types and is applicable to high-resolution color images. A mask is constructed using CIVE to remove the soil background, and the calculation formula is as follows: ; where, 、 、 are the DN values of the red, green, and blue channels of the digital camera, respectively; In CGCI, the growth dynamics of the potato canopy follows a three-stage pattern. The growing degree days (TD) are used to describe the crop growth and development process driven by accumulated temperature, and the calculation formula is as follows: ; where, is the daily average temperature, 、 and are the minimum, optimum, and maximum temperatures for crop growth, respectively (the temperature three key points for potatoes are 5.5 °C, 23.4 °C, and 34.6 °C), is the temperature response curvature coefficient, which is set to 1.7 according to previous studies.

[0031] Using the growing degree days as the horizontal axis, the canopy growth curve is fitted in the form of a piecewise function, as shown in Figure 4 . T0 - T1 is the canopy morphogenesis stage, and the CC growth conforms to a logistic curve; at T1, CC reaches its peak (Vmax) and enters a plateau phase until T2. During this stage, the canopy maintains the maximum canopy coverage and the largest light interception area. The integral area of this stage curve is defined as Smax; in the T2 - Te stage, CC begins to decline until the aboveground part is completely senescent. T1, T2, Te, Vmax, Smax, the total integral area S, and the integral area of the canopy growth curve at a specific time (Canopy Growth Curve Integral, CGCI) are extracted from the canopy growth curve.

[0032] S2. Preliminary construction of a potato yield prediction model and uncertainty analysis of the potato yield prediction model; The specific operation steps include: S2.1. Division of modeling samples; Two different sampling methods, random sampling and stratified sampling, are used to evaluate the uncertainty of model inputs. The output result distributions of the two methods under the condition of fixed modeling features are analyzed using the mixed samples of 2022, 2023, and the two years respectively. To avoid the influence of modeling features and model structure, SLR is selected to compare the influence of the two sampling methods on model uncertainty. SLR is the simplest machine learning method, without the need to adjust hyperparameters, and establishes the relationship between a single feature and the target variable through ordinary least squares (OLS). 70% of the total samples are divided into the training set through the two sampling methods, and the remaining 30% are divided into the test set. The feature with the highest correlation is selected for modeling, and the coefficient of determination (R 2 ) and root mean square error (RMSE) are used to evaluate the prediction accuracy. The calculation formulas are as follows: ; ; where, is the true value of the th sample of the random variable, and is the predicted value of the th sample of the random variable.

[0033] Each sampling method is repeated 100 times. The median (MID), average (AVG), interquartile range (IQR), and coefficient of variation (CV) of model evaluation indicators such as R 2 and RMSE are used to evaluate the uncertainty of modeling sample selection, and the best sampling method is determined.

[0034] The Kolmogorov-Smirnov test (K-S test) and Shapiro-Wilk test (S-W test) are used to analyze the normality of the data distribution of model evaluation indicators. The K-S test is a non-parametric statistical test method used to compare the differences between two sample distributions or between a sample and a theoretical distribution. The K-S test does not require the data to conform to specific distribution assumptions, so it is a non-parametric test and can be used for the comparison of continuous and discrete data. The Shapiro-Wilk test is a statistical method used to test whether sample data conforms to a normal distribution. It is particularly suitable for the normality test of small sample data and is one of the very commonly used normality test methods in practice. It tests whether the sample data comes from a normal distribution. In the case of small samples, the Shapiro-Wilk test is more effective and powerful than other normality tests.

[0035] S2.2. Feature selection for modeling; The specific operation steps are as follows: S2.2.1. Feature selection algorithm; Redundant features will reduce the performance of the model. Before running the machine learning model, it is necessary to preprocess and screen the modeling features. The filters include the Pearson correlation coefficient PCC, mutual information MI, and grey relational analysis GRA based on single variables, and RReliefF based on multiple variables. The recursive feature elimination algorithm RFE based on support vector regression is also selected as a representative of the wrapper.

[0036] PCC is the quotient of the covariance and standard deviation between two variables, which can better describe the linear correlation degree between two sets of variables. The calculation formula is as follows: ; where and represent the th samples of two random variables respectively, is the number of samples, and the value range of PCC is . When PCC is closer to 1, it indicates a high positive correlation between the two variables. When PCC is closer to -1, it indicates a high negative correlation between the two variables. When the value of PCC is 0, the two variables are independent of each other. The correlation between each feature and the target variable is judged by the absolute value of PCC, and feature selection is carried out according to the correlation ranking.

[0037] MI is a measure of the mutual dependence between two variables, used to quantify the "information amount" obtained about another random variable by observing one random variable. The calculation formula is as follows: ; where is and 's joint probability density distribution function, and are the marginal probability density functions of and respectively, and represent two different variables respectively, and the value range of MI is . When the value of MI is 0, the two variables are independent of each other. The larger the MI, the more information the feature gives about the target variable.

[0038] GRA is a multi-factor statistical analysis method, which judges whether the connection between two variables is close by the similarity degree of the geometric shapes of the sequence curves, and is evaluated by the correlation coefficient. The calculation formula is as follows: ; Among them, and are the minimum and maximum values of the reference sequence respectively, and are the minimum and maximum values of the comparison sequence respectively, is the point of the reference sequence, is the point of the comparison sequence, is an adjustable coefficient, and its value is , which is used to adjust the difference in the output results of different features. When takes the value of 0.5, the correlation coefficient takes values between 0 and 1, the larger the value of , the closer the relationship between the feature and the target variable in the given feature.

[0039] The RReliefF algorithm is evolved from the binary classification feature selection algorithm Relief, and evaluates the importance of candidate features by identifying the feature value differences between adjacent samples. On this basis, the RReliefF algorithm can handle multi-class classification tasks, and the RReliefF algorithm is further extended and can be used for feature selection in regression tasks.

[0040] RFE is a wrapper-based feature selection method. Based on all features, the least relevant features are cyclically removed according to feature importance until a specific number of feature sets are selected; the support vector machine SVM is selected as the built-in model of RFE to evaluate feature importance.

[0041] S2.2.2. Uncertainty analysis of feature selection; Evaluate the uncertainty of modeling feature selection by testing the stability of different feature selection algorithms under data perturbation. The bootstrap method is used to draw a specified number of samples from the total samples with replacement, and the feature selection algorithm is run 1000 times in a loop. The extended Quencheva similarity metric index EK is used to evaluate the stability between two groups of feature selection results: ; Among them, is the number of common features, and are the feature numbers of the two groups of features respectively, is the total number of features.

[0042] All feature sets are compared pairwise, the EK value is calculated, and the stability is measured by the average value of all results. Run the feature selection algorithm through the total samples, and compare the consistency of different algorithm results. Run the Lasso model using the features selected by each algorithm, compare the prediction accuracy and uncertainty, and select the optimal feature selection algorithm.

[0043] S2.3, Model Structure Screening; Construct a yield prediction model to quantify and compare the uncertainties caused by differences in model structures. The specific operation steps include: S2.3.1, Data Preprocessing; Remove abnormal samples outside the mean plus or minus 3 standard deviations according to the "±3σ principle" to improve the stability of the model. At the same time, to avoid the influence of feature scales on the model, scale the features to the interval with a mean of 0 and a standard deviation of 1, that is, standardize. Select the optimal modeling sample division and feature selection method, screen the modeling features, and divide 70% of the total samples into the modeling set and 30% into the test set.

[0044] S2.3.2, Hyperparameter Tuning; Select samples from the modeling set through stratified sampling, and perform hyperparameter tuning on 5 machine learning algorithms, namely Lasso regression, Ridge regression, Partial Least Squares Regression (PLSR), Support Vector Machine (SVM), and Random Forest (RF). Among them, for Lasso and Ridge, determine the L1 and L2 regularization parameters alpha through 5-fold cross-validation. The dataset is randomly divided into 5 subsets. Each time, 4 of the subsets are used to train the model, and the remaining 1 subset is used to validate the model. Use the mean of the 5 validation results to evaluate the model performance. After PLSR transforms the original features into new variables (i.e., principal components), perform regression analysis between the newly constructed variables and the yield to reduce the correlation between independent variables. Select the optimal number of principal components by traversing the number of principal components from 1 to 20. SVM selects different kernel function types, penalty factors (C) of the error term, and tolerance parameters ( ) to construct a hyperparameter search space, and use Bayesian optimization to tune the model. Bayesian optimization is a global optimization method that uses Gaussian processes to simulate the unknown validation error function. Each search in the hyperparameter space is optimized based on the previously selected sample points, which greatly saves the tuning performance compared to grid search and random search. To balance the model running time and model performance, select linear and radial basis kernel functions for the kernel function type, and set the range of C to 0.1 to 1000. Used to control the model complexity, with the range set to 0.01 to 10. RF constructs a hyperparameter search space based on the number of decision trees, maximum depth, minimum number of samples per node, minimum number of samples for splitting, and maximum number of features, and uses Bayesian optimization to tune the model. The number of trees is set between 100 - 500. To prevent overfitting, select a depth of 2 - 15. The minimum number of samples per node is 2 - 7, the minimum number of samples for splitting is 5 - 20, and the value of the maximum number of features is set between 1 and the total number of features N.

[0045] S2.3.3, Model Running and Uncertainty Analysis.

[0046] Run the yield prediction model 100 times repeatedly, analyze the median, mean, IQR, and CV of the model accuracy evaluation indicators, and compare the uncertainties caused by the differences in the model structures of different algorithms.

[0047] S3. Based on the uncertainty analysis results obtained in S2, fuse the multi-source data collected in S1 to construct the best potato yield prediction model.

[0048] After determining the best sampling method, feature selection algorithm, and machine learning model through uncertainty analysis, the best strategy for constructing a potato yield prediction model by fusing multi-source data of spectra, textures, structures, meteorology, and varieties at different growth stages was studied. In addition, a maturity evaluation index (MEI) was constructed by combining the growth period process with prior knowledge as a variety parameter.

[0049] Results 1. Tuber yield distribution The yields in 2022 and 2023 are as Figure 5 shown, where Figure 5 (a) in it is the yield of each variety in 2022, Figure 5 and (b) in it is the yield of each variety in 2023, Figure 5 and (c) in it is the yield in 2022 and 2023. Due to the frost in 2022, the growth period of most varieties was shortened by 15 - 20 days compared to 2023. Coupled with insufficient rainfall during the critical growth period, the yield in 2022 was significantly lower than that in 2023. The distribution of the yield data of each experiment in 2022 was more concentrated, and the yield differences between varieties were not large. The average yield was 22.29 ton / ha, and the standard deviation was 5.09 ton / ha. The yields of Zhong 19, Zhong 27, and Zhong 49 were relatively higher, with means of 26.51, 25.00, and 25.86 ton / ha respectively. The mean yields of Zhong 5, Shepody, and Zhong 18 were relatively lower, being 21.81, 19.32, and 19.04 ton / ha respectively. The yields of each variety in 2023 were higher, and the data distribution range was wider, with a mean of 36.36 ton / ha and a standard deviation of 10.34 ton / ha. The average yields of Zhong 49 and Zhong 5 were the highest, being 46.45 and 45.12 ton / ha respectively. Followed by Zhong 27, Zhong 19, and Zhong 18, which were 35.51, 34.61, and 33.44 ton / ha respectively. The average yield of Shepody was the lowest, only 21.60 ton / ha. For a specific variety, the yield of Shepody in 2023 had no significant difference compared to 2022, and the yields of other varieties were all significantly higher than those in 2022.

[0050] 2. Correlation analysis between features and yield Select all samples in two years to construct a yield estimation model. For each year, select a total of 6 periods of remote sensing image data, namely P1 - P6, which are evenly distributed throughout the growth period, to participate in the model construction. The specific acquisition times are shown in Table 3.

[0051] Table 3 Acquisition Dates of Remote Sensing Images in Each Period ;

[0052] 2.1. Spectral Features The correlations between spectral features of different classes and yield vary. Among them, the reflectance in the near-infrared band (NIR) is positively correlated with yield, with the PCC ranging from 0.43 to 0.70, and the correlations in P2 - P4 are higher, with the PCC all greater than 0.67. The yield is positively correlated with the reflectance in the visible light (R, G, B) and red edge (RE) ranges during P3 and P4, but the highest PCC is only 0.56, and there is a negative correlation during P2 and P5. Among them, the correlation between the red light reflectance in P2 and the yield is the strongest, with the PCC being -0.69. Vegetation indices calculated based on NIR, such as the Difference Vegetation Index (DVI), Enhanced Vegetation Index (EVI), and Transformed Difference Vegetation Index (TDVI), have a similar correlation with yield as NIR, are positively correlated with yield, and have a stronger correlation during the P2 - P4 stage, with the highest PCC reaching 0.71. Digital indices such as ExG, GBDI, and Gray calculated based on the visible light band are similar to the reflectance in the visible light band. Vegetation indices such as the Transformed Chlorophyll Absorption in Reflectance Index (TCARI) and Structure-Insensitive Pigment Index (SIPI) are negatively correlated with yield, and the correlation is relatively high, with the PCC reaching -0.73. The results show that there are significant differences in the correlations between different types of spectral features and yield, and the correlations are generally higher during the P2 - P4 period (about 70 - 90 DAP).

[0053] 2.2. Texture Features The correlation laws between the texture features of different bands in the same period and the yield are similar. ASM, HOM, MIN, and VAR of each band are negatively correlated with the yield, and the correlation is the strongest in the P2 - P4 period. In the P3 and P4 periods, the PCC of HOM can reach -0.77. The correlations of CON, COR, DIS, and ENT with the yield vary with different periods. The correlations of CON and COR with the yield are weak. Only the CON in the P3 and P5 periods and the COR in the P1 period have slightly stronger correlations with the yield. Among them, the PCC of CON in the P3 period is between 0.45 and 0.64. DIS and ENT in the P2 - P4 periods have a positive correlation with the yield. Among them, DIS and ENT in the P3 period have a strong positive correlation with the yield, and the PCC is between 0.65 and 0.75.

[0054] 2.3 Structural characteristics The correlation between the structural characteristics and the yield varies with different periods, and CC and CGCI have relatively strong correlations with the yield. CC has a positive correlation with the yield, and the correlation is the highest in the P2 - P4 periods, with the PCC between 0.47 and 0.74. As the growth period progresses, the correlation between CGCI and the yield gradually increases from 0.46 in the P1 period to 0.73 in the P5 and P6 periods. The correlations of CH and COV in the P2 and P4 periods with the yield are higher than those in other periods. The PCCs of CH and COV in the P2 period are 0.62 and 0.64 respectively. The correlations of CH and COV in the P4 period are higher, and the PCCs are 0.74 and 0.65 respectively.

[0055] 3. Uncertainty analysis 3.1. Selection of modeling samples Simple Random Sampling (SRS) and Stratified Random Sampling (STR) are used to select modeling samples, and the vegetation indices with the highest correlations are selected to construct the SLR model. Repeat 100 times to analyze the distribution of the evaluation indices R 2 and RMSE to measure the uncertainty of the selection of modeling samples.

[0056] 3.1.1. Normality test of accuracy evaluation indices The Kolmogorov - Smirnov test (K - S test) and the Shapiro - Wilk test (S - W test) are used to analyze the normality of the data distribution of the model evaluation indices. The results show that when using the K - S test, the P - values of the vast majority of simulation results of R 2 obtained by the two sampling methods are greater than 0.05, and the data conforms to the normal distribution. However, when using the S - W test, which is more sensitive to small samples, the results show that the P - values are generally less than 0.05, and the data does not conform to the normal distribution. Taking the P3 period as an example, from R 2The Probability Density Function (PDF) has a certain deviation from the standard normal distribution as Figure 6 shown, where Figure 6 in (a) is the R 2 probability density function of different sampling methods in the P3 period, Figure 6 and in (b) is the RMSE probability density function of different sampling methods in the P3 period. The skewness is used to evaluate the degree of deviation from the standard normal curve. The results show that the skewness of R 2 in the test set of each period is less than -0.5, indicating a certain degree of skewed distribution. The median of the multiple-sampled R 2 is used to represent that the sampling result is more representative than the average value.

[0057] The same method is used to evaluate the normality of RMSE. The results show that the P-values of both test methods are greater than 0.05, the RMSE distribution conforms to the normal distribution, and the skewness is between -0.5 and 0.5, without skewed distribution. It can also be seen from the PDF that there is no obvious difference between the median and the mean of RMSE, and the average value of RMSE can be directly used to evaluate the overall performance of the model.

[0058] 3.1.2. Uncertainty Analysis of Sampling Methods Based on SLR The SLR model is constructed by fusing two-year data. The R 2 of both random sampling and stratified sampling are between 0.27 and 0.54.

[0059] Table 4 Uncertainty Analysis of Sampling Methods Based on SLR 2 ; ;

[0060] Table 4 Uncertainty Analysis of Sampling Methods Based on SLR 2 ; ;

[0061] There are significant differences in the SLR modeling accuracy between 2022 and 2023, indicating that the modeling input data is the main factor affecting the model accuracy. When fusing two-year data and using only the 2022 data for modeling, the highest test R 2 of the random sampling SLR model are 0.54 and 0.66 respectively, and the R 2 of the stratified sampling in the same period are 0.55 and 0.71 respectively. It can be seen that stratified sampling can improve the model accuracy. Except that the IQR of the P2 and P4 stratified sampling R 2 in the fused two-year data increases by 0.01 compared with random sampling, other IQRs and CV(%) are lower than random sampling.

[0062] From the prediction results of multiple samplings of SRS and STR at different times in 2022, it can be seen that the prediction result R of STR 2 is greater than that of SRS, and its IQR and CV are both smaller than those of SRS. The output results are more concentrated and the uncertainty is lower. As Figure 7 shown, among them, Figure 7 in (a) is the median of model R 2 , Figure 7 in (b) is the interquartile range of model R 2 , Figure 7 in (c) is the coefficient of variation of model R 2 . The results show that stratified sampling can effectively reduce the uncertainty of the results.

[0063] 3.2. Feature selection for modeling From the modeling results of SLR in different periods, it can be seen that when fusing data for modeling in 2022 and 2023, the R of the P2 - P5 stratified sampling test 2 is between 0.48 - 0.55, and the RMSE is 6.70 - 7.28 ton / ha, which is relatively stable. Select the P3 period among them for the uncertainty study of modeling features. At the same time, to control the complexity of the feature composition, only 58 spectral features are selected for feature screening.

[0064] 3.2.1. Comparative study on feature selection methods based on data perturbation The uncertainty analysis of modeling features based on data perturbation performs multiple sampling with replacement on the overall sample through Bootstrap to construct 1000 sample sets of the same size as the original data set. Evaluate the stability of feature selection methods GRA, Pearson correlation coefficient, MI, RReliefF, and RFE using the extended Kuncheva stability as Figure 8 shown. The number of features is often selected based on modeling experience. Considering that the total number of features is 58, the number of features from 1 to 20 is traversed for feature selection.

[0065] When only one feature is selected, the stability of the features selected by RFE is the highest, and the Kuncheva stability is 0.88. When selecting one feature each time in the 1000 datasets formed by Bootstrap, the Structure Insensitive Pigment Index (SIPI) was selected 937 times. In addition, GRA, PCC, MI, and RReliefF selected SIPI in 4.2%, 57.9%, 34.0%, and 0.0% of the datasets respectively. As the number of target features increases, the stability of different feature selection algorithms shows different trends. The stability of the features screened by RFE drops sharply as the number of features increases, but it is still higher than that of the filter-based feature selection methods GRA, MI, and PCC. These three methods are all single-variable-based feature selection methods and have weak resistance to data interference. The rate of decline in the stability of RFE slows down after the number of features increases to 5. When screening 1 - 5 features, the stability trend of RReliefF is opposite to that of RFE. As the number of target features increases, the selected features become more stable. At the same time, when selecting 4 - 16 features, the extended Kuncheva stability is between 0.50 - 0.74, which is higher than that of other methods, indicating that among the feature selection algorithms used, RReliefF has the strongest resistance to data interference.

[0066] 3.2.2 Modeling Performance Test of Different Feature Selection Methods Compare the prediction accuracies of different feature selection algorithms through Lasso, and repeat the modeling 100 times. Use different feature selection algorithms R 2 The median represents the model performance as Figure 9 shown. When the number of selected features is less than or equal to 7, different feature selection methods R 2 show an upward trend with the increase in the number of features. At this stage, the results of RFE are better than those of other feature selection methods, while the accuracy of RReliefF is the lowest. When the number of features is 7, the results are similar. When selecting 8 - 16 features, GRA, RFE, and MI show large fluctuations, while RReliefF and Pearson have entered a plateau period. When the number of features is greater than 16, different feature selection methods R 2 remain stable. When the number of selected features is greater than 7, the results of RReliefF are better than those of other selected feature selection methods, and R 2 remains at around 0.70. The results show that RFE is a better choice when a small number of modeling features need to be selected, while RReliefF can fully mine important features and avoid selecting redundant features.

[0067] 4 Model Structure Differences 4.1 Comparison of Yield Prediction Accuracies of Different Machine Learning Models The same modeling and validation samples are selected using a fixed random seed, and the yield prediction accuracies of different algorithms are compared in a single modeling, as Figure 10 shown, where Figure 10 in (a) is the yield prediction accuracy of SLR, Figure 10 in (b) is the yield prediction accuracy of Lasso, Figure 10 in (c) is the yield prediction accuracy of RF, Figure 10 in (d) is the yield prediction accuracy of PLSR, Figure 10 in (e) is the yield prediction accuracy of Ridge, Figure 10 in (f) is the yield prediction accuracy of SVM. SLR uses the vegetation index with the highest correlation for modeling, and other machine learning algorithms use the top 10 features selected by RReliefF for modeling. As Figure 10 can be seen, PLSR has the highest prediction accuracy (R 2 = 0.75, RMSE = 5.34 ton / ha). Followed by Ridge (R 2 = 0.74, RMSE = 5.41 ton / ha). The accuracies of RF, SVM, and Lasso are slightly lower, and R 2 is around 0.68 - 0.70. SLR has the lowest accuracy, R 2 = 0.49.

[0068] 4.2 Uncertainty analysis caused by model structure differences Modeling features are screened through RReliefF, and 5 different models are run 100 times. The results show that PLSR and Ridge have higher accuracies and lower uncertainties. SLR has only one feature, which is quite different from other algorithms and does not participate in the uncertainty analysis of model structure. Analyze the R 2 uncertainty under different numbers of features as Figure 11 shown, where Figure 11 in (a) is the median of model R 2 , Figure 11 in (b) is the interquartile range of model R 2 , Figure 11 in (c) is the coefficient of variation of model R 2 .

[0069] Analyze the RMSE uncertainty under different numbers of features Figure 12 shown, where Figure 12 in (a) is the average value of model RMSE, Figure 12 in (b) is the interquartile range of model RMSE, Figure 12 in (c) is the coefficient of variation of model RMSE. When the number of features is less than or equal to 3, for the two algorithms of SVM and RF, R 2has a higher median, lower CV, and lower uncertainty caused by the model structure. At the same time, the mean of RMSE is reduced by 0.8 - 1.5 ton / ha compared with the other three methods. When there are 4 - 10 features, the differences between various algorithms are not significant, but the IQR and CV of RF and SVM for R 2 and RMSE are higher than those of the other three methods, with greater uncertainty. After the number of features is greater than or equal to 10, the model accuracy tends to be stable. The R of Ridge and PLSR 2 is higher than that of SVM, RF, and Lasso, while the RMSE is lower, and the model prediction performance is better. At the same time, the uncertainty of Ridge and PLSR is also lower. When the number of features is 10, the test R of Ridge 2 has a median of 0.76 and a mean RMSE of 4.89 ton / ha. PLSR has slightly higher accuracy, and the R 2 has a median of 0.77 and a mean RMSE of 4.81 ton / ha.

[0070] 5. Construction of Potato Yield Prediction Model Based on Multi-source Data 5.1 Modeling Feature Selection The modeling features of each period are screened by the RReliefF algorithm. The top 20 vegetation indices with importance scores are as Figure 13 shown, where Figure 13 (a) in is the top 20 vegetation indices with importance scores in the P1 period, Figure 13 (b) in is the top 20 vegetation indices with importance scores in the P2 period, Figure 13 (c) in is the top 20 vegetation indices with importance scores in the P3 period, Figure 13 (d) in is the top 20 vegetation indices with importance scores in the P4 period, Figure 13 (e) in is the top 20 vegetation indices with importance scores in the P5 period, Figure 13Among them, (f) are the top 20 vegetation indices in terms of importance score during the P6 period. BNDVI is the blue normalized difference vegetation index, CCCI is the canopy chlorophyll content index, DVI is the difference vegetation index, EVI is the enhanced vegetation index, EVI2 is the two-band enhanced vegetation index, GBNDVI is the green-blue normalized difference vegetation index, GDVI is the green difference vegetation index, GNDVI is the green normalized difference vegetation index, GRNDVI is the green-red normalized difference vegetation index, GRVI is the green ratio vegetation index, MCARI is the modified chlorophyll absorption reflectance index, MCARI1 is the improved chlorophyll absorption index 1, NDRE is the red-edge normalized index, NDVI is the normalized difference vegetation index, OSAVI is the optimized soil-adjusted vegetation index, RDVI is the re-normalized difference vegetation index, SR is the simple ratio vegetation index, SAVI is the soil-adjusted vegetation index, WDRVI is the wide dynamic range vegetation index, MSR is the improved simple ratio vegetation index, SIPI is the structure-insensitive pigment index, GCI is the green chlorophyll index, RECI is the red-edge chlorophyll index, NDREI is the normalized difference red-edge index, GSAVI is the green soil-adjusted vegetation index, and B, G, R, RE, and NIR are the reflectances of the blue, green, red, red-edge, and near-infrared bands respectively. The characteristic importance of P2 to P5 is generally higher than the results of other periods. Among them, the rankings of TCARI, TCARI / OSAVI, NDREI, etc. are all in the front.

[0071] 5.2 Yield Estimation Based on Spectral and Texture Features Spectral features (VI), texture features (TEX), and the fusion of spectral and texture features are used for modeling. The top 10 features in terms of score are selected for each period, and PLSR is used to construct a potato yield prediction model based on multi-source data. The modeling samples and test samples are divided by stratified sampling, and a fixed random seed is set to ensure that each model uses the same modeling samples for training. The results show that during the P2 and P3 periods (70 - 80 DAP), the highest accuracy of yield estimation using texture and spectral features respectively is as shown in Table 5. P3 is the period with the highest prediction accuracy, R 2 = 0.81, RMSE = 4.71 ton / ha. The prediction accuracy of texture features is better during the P2 period, R 2 = 0.80, RMSE = 4.80 ton / ha. The prediction accuracy of fusing texture and spectral features during the P2 period does not improve compared with using spectral features alone, but compared with texture features, R 2 increases by 0.11 and RMSE decreases by 1.10 ton / ha. The prediction accuracy of fusing texture and spectral features during the P3 period does not improve compared with using texture features alone, but compared with spectral features, R 2 increases by 0.06 and RMSE decreases by 0.62 ton / ha.

[0072] Table 5 Evaluation of Yield Estimation Accuracy Based on Spectral and Textural Features of PLSR ;

[0073] Table 5 Evaluation of Yield Estimation Accuracy Based on Spectral and Textural Features of PLSR ;

[0074] In summary, under the condition of the same number of features, the modeling capabilities of textural features and spectral features vary with different periods. In the P2 and P3 periods, fusing spectral and textural information cannot improve the model accuracy, but the prediction results in different periods are more stable.

[0075] 5.3 Yield Estimation by Fusing Multi-source Data According to the yield estimation results of spectral and textural features, the P2 and P3 periods with better performance were selected to evaluate the accuracy of the potato yield prediction model by fusing multi-source data. Since the yield estimation accuracy fluctuates greatly in different periods when using spectral and textural features alone, R 2 is between 0.66 - 0.81, and the yield estimation results by fusing spectral and textural features are more stable, with R 2 between 0.74 - 0.77. Based on fusing spectral and textural features, structural and meteorological parameters were added for yield estimation. The structural features include canopy cover (CC) and canopy growth curve area (CGCI), and the meteorological features include growing degree days (GDDs), accumulated heat days (TD), cumulative rainfall during the critical growth period (KCP), and maturity evaluation index (MEI). The results show that adding the current canopy cover on the basis of VI and TEX can improve the model accuracy, with R 2 improved by 0.04 and RMSE reduced by 0.47 ton / ha. However, other structural and meteorological indicators cannot improve the model prediction performance as shown in Table 6.

[0076] Table 6 Evaluation of Yield Estimation Accuracy by Fusing Multi-source Data of Spectral and Textural Features of PLSR ;

[0077] Therefore, by adopting the above method for constructing a potato yield prediction model based on multi-spectral and machine learning, the uncertainty caused by model input can be effectively reduced. Combining spectral features with textural features can improve the model stability, and fusing canopy cover can further improve the model accuracy.

[0078] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions of the present invention or make equivalent replacements, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for constructing a potato yield prediction model based on multispectral and machine learning, characterized in that: It includes the following steps: S1. Multi-source data acquisition and processing; S1.

1. Determine the experimental site and experimental design; S1.

2. Ground data measurement; S1.

3. UAV data acquisition and processing; S1.

4. Spectral feature extraction; In spectral feature extraction, a soil mask is constructed by optimizing the Optimized Soil-Adjusted Vegetation Index (OSAVI) to remove the multispectral background. The calculation formula is as follows: ; Among them, is the reflectance in the near-infrared band, is the reflectance in the red band, is an adjustment factor used to adjust the influence of the soil, and its value is 0.16; S1.

5. Texture feature extraction; S1.

6. Structure feature extraction; S2. Initially construct a potato yield prediction model and conduct uncertainty analysis on the potato yield prediction model; S3. Based on the uncertainty analysis results obtained in S2, fuse the multi-source data collected in S1 to construct the best potato yield prediction model.

2. The method for constructing a potato yield prediction model based on multispectral and machine learning according to claim 1, wherein: The structure features in S1.6 include plant height CH, canopy coverage CC, and canopy growth curve CGCI; The extraction of CH includes three steps: constructing a digital terrain model DTM, calculating a crop height model CHM, and calculating the plant height CH; CC is the ratio of vegetation coverage on a unit land area. Use OSAVI to segment the soil and vegetation, and calculate CC by statistically analyzing the ratio of vegetation pixels to total pixels. A mask is constructed through the vegetation extraction color index to remove the soil background. The calculation formula is as follows: ; Among them, , , are the DN values of the red, green, and blue channels of the digital camera, respectively; In CGCI, the growth dynamics of the potato canopy conform to the three-stage law. Use the thermal days TD to describe the crop growth and development process driven by accumulated temperature. The calculation formula is as follows: ; Among them, is the daily average temperature, is the minimum temperature for crop growth, is the optimum temperature for crop growth, is the maximum temperature for crop growth, is the temperature response curvature coefficient.

3. The method for constructing a potato yield prediction model based on multispectral and machine learning according to claim 1, wherein: The specific operation steps of S2 include: S2.

1. Modeling sample division; S2.

2. Modeling feature selection; S2.

3. Model structure screening.

4. The method for constructing a potato yield prediction model based on multispectral and machine learning according to claim 3, wherein: The specific operation steps of S2.2 include: S2.2.

1. Feature selection algorithm; The filters include the Pearson correlation coefficient PCC based on a single variable, mutual information MI, grey relational analysis GRA, and RReliefF based on multiple variables. The recursive feature elimination algorithm RFE based on support vector regression is also selected as a representative of the wrapper; PCC is the quotient of the covariance and standard deviation between two variables. The calculation formula is as follows: ; Among them, and respectively represent the th samples of two random variables, is the number of samples. The value range of PCC is . When PCC is closer to 1, it indicates a strong positive correlation between the two variables. When PCC is closer to -1, it indicates a strong negative correlation between the two variables. When the value of PCC is 0, the two variables are independent of each other; MI is a measure of the mutual dependence between two variables, used to quantify the "information content" obtained about another random variable by observing one random variable. The calculation formula is as follows: ; Among them, is and 's joint probability density distribution function, and are respectively and 's marginal probability density functions, and represent two different variables respectively. The value range of MI is . When the value of MI is 0, the two variables are independent of each other. The larger the MI, the more information the feature gives about the target variable. GRA is a multi-factor statistical analysis method that determines whether the relationship between two variables is close by judging the similarity of the geometric shapes of the sequence curves, and is evaluated by the correlation coefficient. The calculation formula is as follows: ; Among them, and are the minimum and maximum values of the reference sequence respectively, and are the minimum and maximum values of the comparison sequence respectively, is a point of the reference sequence, is a point of the comparison sequence, is an adjustable coefficient, and its value is , which is used to adjust the difference in the output results of different features. When takes the value of 0.5, the correlation coefficient ranges from 0 to 1, the larger the value of, the closer the relationship between the feature and the target variable in the given feature; The RReliefF algorithm is evolved from the binary classification feature selection algorithm Relief, and evaluates the importance of candidate features by identifying the eigenvalue differences between adjacent samples; RFE is a wrapper-based feature selection method that cyclically eliminates the least relevant features based on feature importance among all features until a specific number of feature sets are selected; S2.2.

2. Feature selection uncertainty analysis; The extended Kuncheva similarity metric index EK is used to evaluate the stability between two groups of feature selection results: ; Among them, is the number of common features, and are the feature quantities of two groups of features respectively, is the total number of features.

5. The method for constructing a potato yield prediction model based on multispectral and machine learning according to claim 3, wherein: The specific operation steps of S2.3 include: S2.3.

1. Data preprocessing; S2.3.

2. Hyperparameter tuning; S2.3.

3. Model operation and uncertainty analysis.