A method and system for predicting and classifying the risk of saline-alkali soil wheat yield
Patent Information
- Application Number
- CN202611205867.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-10
- Publication Date
- 2026-09-29
AI Technical Summary
[0007]本发明旨在克服上述现有技术的至少一种缺陷,提供一种盐碱地小麦产量预测及风险分级方法,以解决盐碱地理化特征评价指标单一,信息利用不充分;难以处理多因子耦合作用,缺乏关键特征自动筛选机制等问题
(1)本发明提供的一种盐碱地小麦产量预测及风险分级方法和系统,区别于基于遥感植被指数、盐分指数和土壤盐含量的产量估算方法,构建了土壤理化性质、酶活性和微生物指标融合的地下土壤生态功能特征体系,可反映耕层土壤生态功能对小麦产量形成的影响。
Smart Images

Figure CN122840694A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the technical field of agricultural informatization, and more specifically, relates to a method and system for predicting wheat yield and risk classification in saline-alkali land. Background Technology
[0002] Saline-alkali land is a significant obstacle to improving agricultural productivity. It is typically characterized by high salinity, high pH levels, nutrient imbalances, and altered microbial ecosystems. These factors collectively affect soil fertility and the rhizosphere environment, significantly restricting the normal growth and yield of grain crops such as wheat. With the increasing scarcity of arable land and the growing demand for the development and utilization of saline-alkali land, establishing scientific evaluation and prediction methods for wheat production capacity in saline-alkali land is crucial for ensuring food security and improving land use efficiency.
[0003] Existing crop yield prediction technologies for saline-alkali land mainly fall into three categories: The first category is remote sensing index-driven methods, such as using the Sentinel-2 vegetation index, salinity index, and soil salt content, combined with a random forest model to estimate winter wheat yield in saline areas. This type of method is suitable for regional-scale monitoring, but it mainly relies on remote sensing spectral characteristics and salinity information, and it is difficult to directly reflect the impact of topsoil enzyme activity, microbial communities, and underground ecological functions on wheat yield formation.
[0004] The second category is crop phenotypic trait screening methods. For example, in the coastal saline-alkali land rice scenario, Pearson correlation analysis, LASSO regression, and leave-one-out cross-validation are used to screen phenotypic traits related to rice yield and eating quality. This type of method focuses on crop phenotypic traits and quality objectives, but does not establish a soil ecological function characteristic system that synergistically characterizes soil physicochemical properties, enzyme activity, and microbial indicators for saline-alkali land wheat.
[0005] The third category involves methods that couple remote sensing inversion with crop process models. For example, this involves inverting salinity, moisture, and leaf area indices based on vegetation and salinity indices, and then predicting crop yields in saline-alkali farmland using a HYDRUS-1D-EPIC coupled model and data assimilation methods. The HYDRUS-1D-EPIC coupled model is a crop process model formed by coupling the one-dimensional soil water and salt transport model HYDRUS-1D with the crop growth and yield simulation model EPIC, used to simulate soil water and salt dynamics and their impact on crop growth and yield formation. This type of method relies on process models, remote sensing inversion, and assimilation algorithms, resulting in high modeling complexity and strong dependence on parameter calibration and multi-source time-series data.
[0006] Existing machine learning prediction schemes often use feature selection and candidate model comparison as a general process, lacking a feature construction method for saline-alkali wheat soil ecological processes. This results in model inputs remaining at the level of raw indicators, failing to fully express the synergistic relationships between salt stress, nutrient supply, enzymatic transformation, and microbial responses. Furthermore, existing technologies lack a clear process for determining the optimal set of key features that is stable, low in redundancy, and has excellent predictive performance from multiple candidate feature sets. Summary of the Invention
[0007] This invention aims to overcome at least one of the defects of the prior art and provide a method for predicting wheat yield and risk classification in saline-alkali land, so as to solve the problems of single evaluation index of saline-alkali land geochemical characteristics, insufficient information utilization, difficulty in handling multi-factor coupling effects, and lack of automatic screening mechanism for key features.
[0008] The detailed technical solution of this invention is as follows: A method for predicting wheat yield and classifying risks in saline-alkali land, the method comprising: S1. Collect topsoil samples at a predetermined depth in the target saline-alkali land wheat planting plots, obtain the actual wheat yield of the corresponding plots, and detect to form a set of original soil ecological function indicators. S2. Perform missing value imputation, outlier processing and standardization on the original soil ecological function index set to obtain the standardized original feature matrix; S3. Construct soil ecological function derived features based on the standardized original feature matrix, and merge the standardized original feature matrix with the soil ecological function derived features to obtain the enhanced feature matrix; The derived characteristics of soil ecological functions include: salt stress regulation characteristics, comprehensive nutrient supply characteristics, enzyme-catalyzed nutrient transformation characteristics, and microbial ecological response characteristics; S4. Construct a continuous yield target variable based on the measured wheat yield of each plot, and construct a relative yield reduction target variable based on the reference yield of the target area; S5. Using the enhanced feature matrix as the input feature and the continuous output target variable as the supervised learning target, LASSO feature selection, random forest feature importance selection and K-optimal feature selection based on F test are performed under multiple resampling conditions to obtain multiple candidate feature sets, namely LASSO selected feature set, random forest selected feature set and K-optimal selected feature set based on F test. S6. For each candidate feature set, calculate the comprehensive score of the candidate feature set based on the selection frequency of each candidate feature in the multiple resampling, the cross-validation error of the prediction model built based on the candidate feature set, and the redundancy between each candidate feature within the candidate feature set, and determine the candidate feature set with the highest comprehensive score as the optimal key feature set. S7. Based on the optimal set of key features, establish a wheat yield prediction model for saline-alkali land, and optimize the yield prediction model. The model optimization includes: optimization of the number of latent variables, optimization of model parameters, and residual correction. S8. Input the optimal set of key features of the wheat planting plot in the saline-alkali land to be tested into the optimized yield prediction model to obtain the predicted wheat yield, and output the production capacity level and the yield reduction risk level according to the predicted wheat yield and the target variable of relative yield reduction rate.
[0009] Furthermore, the missing value filling specifically involves: when the proportion of missing soil indicators in the candidate indicators, i.e. the original soil ecological function indicator set, is not higher than the preset missing threshold, the median of the corresponding candidate indicators within the same salinity level plot is used for filling. When the missing percentage exceeds the preset missing threshold, the candidate indicator is removed. The outlier handling is specifically as follows: sort the sample values of each candidate indicator in ascending order, and calculate the first quartile Q1, the third quartile Q3, and the interquartile range IQR of the sample values. Here, Q1 represents the value of the candidate indicator sample at the 25th quartile, Q3 represents the value of the candidate indicator sample at the 75th quartile, and IQR represents the difference between Q3 and Q1, i.e., IQR = Q3 - Q1. When a candidate index value is less than Q1-1.5IQR or greater than Q3+1.5IQR, it is marked as an outlier and truncated using the boundary value within the corresponding salinity level. The standardization process uses Z-score standardization to eliminate dimensional interference for subsequent feature selection and PLSR modeling.
[0010] Furthermore, the salt stress regulation characteristics include: a salt-alkali stress index and a salt-nutrient inhibition characteristic; wherein, the salt-alkali stress index SAI=Z(SC)+Z(pH), where SC represents soil salt content, pH represents soil acidity / alkalinity, and Z represents standardized treatment; wherein, the salt-nutrient inhibition characteristic SNI=Z(SC) / [Z(OM)+Z(TN)+Z(TP)+Z(TK)+ε], where OM represents organic matter, TN represents total nitrogen, TP represents total phosphorus, TK represents total potassium, and ε is a constant to prevent the denominator from being zero; The comprehensive nutrient supply characteristics include: Comprehensive Nutrient Supply Index NSI = w1Z(NH4) + -N)+w2Z(NO3 - -N)+w3Z(OM)+w4Z(TN)+w5Z(TP)+w6Z(TK), where NH4 + -N indicates ammonium nitrogen, NO3 --N represents nitrate nitrogen, and w1 to w6 are the set weighting coefficients; The enzymatic nutrient conversion characteristics include: a comprehensive enzyme activity index and a carbon-nitrogen conversion balance characteristic; wherein, the comprehensive enzyme activity index EAI = v1Z(BG) + v2Z(NAG) + v3Z(LAP), where BG represents β-glucosidase, NAG represents acetamidoglucosidase, LAP represents L-leucine aminopeptidase, and v1 to v3 are set weighting coefficients; wherein, the carbon-nitrogen conversion balance characteristic CNE = Z(BG) / [Z(NAG) + Z(LAP) + ε]; The microbial ecological response characteristics include: microbial stability characteristics and microbial-enzyme synergy characteristics; wherein, the microbial stability characteristic MSI=Z(DIV)-Z(KSA), where DIV represents microbial diversity and KSA represents the abundance of key species related to salt stress; wherein, the microbial-enzyme synergy characteristic MES=Z(DIV)×Z(EAI).
[0011] Furthermore, the comprehensive score S of the candidate feature set core as follows: S core =αF stab +βF acc -γF red (1) Among them, F stab The stability score is determined based on the average selection frequency of each feature in the candidate feature set; F acc For accuracy scores, the coefficient of determination R², root mean square error (RMSE), and mean absolute error (MAE) are determined based on leave-one-out cross-validation or K-fold cross-validation; F red The redundancy score is determined based on the average of the pairwise absolute correlation coefficients of features within the candidate feature set; α, β, and γ are weighting coefficients.
[0012] Furthermore, the model optimization specifically includes: First, the number of latent variables in the production forecast model is optimized: the model is built in the range of 1 to the preset maximum number of latent variables Amax, and the prediction error under different numbers of latent variables is calculated by cross-validation. The number of latent variables with the smallest prediction error or that meets the unified standard error rule is determined as the optimal number of latent variables. Then, the model parameters are optimized: the model regression coefficients, intercept term, training sample prediction residuals and cross-validation error are calculated based on the training samples, and the final model parameters are determined with the optimization objectives of minimizing RMSE, lower MAE and higher R². Furthermore, residual correction is performed: the prediction residuals of the training samples are calculated, a residual correction model is established that combines the residuals with soil salinity, pH and microbial ecological response characteristics, and the corrected residual values are superimposed with the output values of the yield prediction model to obtain the corrected predicted yield.
[0013] In another aspect of the present invention, a system for predicting wheat yield and risk classification in saline-alkali land is provided. The system includes: a sample collection module, an indicator detection module, a data transmission module, a sample database module, a data preprocessing module, a derived feature construction module, a target variable construction module, a key feature stability screening module, a model construction and optimization module, a yield prediction and risk classification module, and a user display terminal. The sample collection module is used to collect soil samples from the 0-20cm topsoil layer of the target plot and obtain the measured wheat yield. The index detection module is used to detect and obtain the original set of soil ecological function indicators; The data transmission module is used to upload the original set of soil ecological function indicators to the sample database module in real time; The sample database module is used to store the original set of soil ecological function indicators; The data preprocessing module is used to perform missing value imputation, outlier processing and standardization on the original soil ecological function index set to obtain a standardized original feature matrix. The derived feature construction module is used to construct soil ecological function derived features and generate an enhanced feature matrix; The target variable construction module is used to construct continuous output target variables and relative production reduction rate target variables; The key feature stability screening module is used to determine the optimal set of key features through multi-method screening and comprehensive scoring under multiple resampling. The model building and optimization module is used to establish a production prediction model and perform latent variable number optimization, parameter optimization and residual correction. The yield prediction and risk classification module is used to output the predicted wheat yield, production capacity level, and yield reduction risk level. The user display terminal is used to display the prediction and classification results.
[0014] In another aspect of the invention, an electronic device is also provided, comprising: At least one processor; and The memory stores instructions that, when executed by the at least one processor, cause the at least one processor to perform a method for predicting wheat yield and risk classification in saline-alkali land as described above.
[0015] In another aspect of the invention, a computer-readable storage medium is also provided, which stores executable instructions that, when executed, cause the machine to perform a method for predicting wheat yield and risk classification in saline-alkali land as described above.
[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) The present invention provides a method and system for predicting wheat yield and risk classification in saline-alkali land. It differs from the yield estimation method based on remote sensing vegetation index, salinity index and soil salt content. It constructs a subsurface soil ecological function characteristic system that integrates soil physicochemical properties, enzyme activity and microbial indicators, which can reflect the influence of topsoil ecological function on wheat yield formation.
[0017] (2) The present invention provides a method and system for predicting wheat yield and risk classification in saline-alkali land, which further constructs salt stress regulation characteristics, nutrient supply comprehensive characteristics, enzyme-catalyzed nutrient transformation characteristics and microbial ecological response characteristics, so that the model input can express the key soil ecological processes in the formation of wheat yield in saline-alkali land.
[0018] (3) The present invention provides a method and system for predicting wheat yield and risk classification in saline-alkali land. By comprehensively determining the optimal set of key features through the selection frequency, cross-validation error and feature redundancy under multiple resampling, the impact of accidental selection and multicollinearity on the model can be reduced, which is different from using a single feature screening algorithm.
[0019] (4) The present invention provides a method and system for predicting wheat yield and risk classification in saline-alkali land. When establishing the partial least squares regression model, the method adds the optimization of latent variable number, model parameter optimization and residual correction process to improve the model’s adaptability to differences in salinity, pH and microbial ecological response. At the same time, it constructs continuous yield target variable and relative yield reduction rate target variable so that the model can output not only the yield prediction value, but also the plot-level production capacity level and yield reduction risk level, thereby improving the practicality of field management and saline-alkali land zoning regulation. Attached Figure Description
[0020] Figure 1 This is a schematic diagram of the process for predicting wheat yield and risk classification in saline-alkali land as described in this invention.
[0021] Figure 2 This is a Spearman correlation heatmap of soil ecological function indicators and wheat yield in Embodiment 1 of the present invention.
[0022] Figure 3 This is a comparison chart of different key feature sets and model performance in Embodiment 1 of the present invention.
[0023] Figure 4 This is a schematic diagram of the predicted yield and production capacity level of wheat for each sample in Embodiment 1 of the present invention.
[0024] Figure 5 This is a schematic diagram of the relative production reduction rate and production reduction risk level of each sample in Embodiment 1 of the present invention.
[0025] Figure 6is a process diagram of determining the optimal key feature set in Embodiment 1 of the present invention.
[0026] Figure 7 is a system deployment architecture diagram in Embodiment 2 of the present invention. Detailed Description of the Embodiments
[0027] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0028] It should be noted that the following detailed descriptions are all exemplary and are intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0029] It should be noted that the terms used herein are only for describing specific embodiments, and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should also be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate that there are features, steps, operations, devices, components, and / or combinations thereof.
[0030] Without conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other.
[0031] Embodiment 1 Referring to Figure 1 , this embodiment provides a method for wheat yield prediction and risk grading in saline-alkali land. Instead of simply using a general machine learning model to predict wheat yield, the present invention adds a process for constructing derived features of soil ecological functions before yield prediction, adds a process of stability screening and redundancy control in the key feature screening stage, and adds a model optimization process in the model establishment stage, thereby improving the accuracy and stability of wheat yield prediction and risk grading in saline-alkali land. The method comprises: S1, soil sample collection and index detection; Collect 0-20 cm plow layer soil samples from the target saline-alkali land wheat planting plots, obtain the actually measured wheat yield of the corresponding plots, and detect to form an original soil ecological function index set.
[0032] In this embodiment, Spearman correlation analysis is performed on the original soil ecological function indicators and wheat yield, as shown in Figure 2 , a Spearman correlation heatmap of soil indicators and wheat yield is obtained, where the color indicates the direction and strength of the correlation, a positive value indicates a positive correlation, a negative value indicates a negative correlation, and a larger absolute value indicates a stronger correlation. From Figure 2It can be seen that different soil physicochemical properties, soil enzyme activity, and microbial indicators are correlated with wheat yield in different directions and with varying strengths. Furthermore, there are also certain correlations among these indicators, indicating that wheat yield in saline-alkali land is influenced by multiple soil factors, including salinity, pH, nutrients, enzyme activity, and microbial community. Using only a single indicator or a single category of indicators is insufficient to fully characterize the wheat yield formation process in saline-alkali land.
[0033] Based on the correlation feature map, the original soil ecological function index set can be made to include: soil physicochemical property index, soil enzyme activity index and microbial index; The soil physicochemical properties include: pH, salt content, ammonium nitrogen, nitrate nitrogen, organic matter, total nitrogen, total phosphorus, and total potassium. The soil enzyme activity indicators include: β-glucosidase, acetamidoglucosidase and L-leucine aminopeptidase; The microbial indicators include: microbial diversity, abundance of key species, and microbial community characteristics.
[0034] S2, Data Preprocessing; The original soil ecological function index set was subjected to missing value imputation, outlier processing and standardization to obtain the standardized original feature matrix.
[0035] Specifically, the missing value filling is as follows: when the missing proportion of each soil indicator in the candidate indicator, i.e. the original soil ecological function indicator set, is not higher than the preset missing threshold, the median of the corresponding candidate indicator in the same salinity level plot is used for filling. When the proportion of missing candidate indicators is higher than the preset missing threshold, the candidate indicator is removed.
[0036] The outlier handling is specifically as follows: sort the sample values of each candidate indicator in ascending order, and calculate the first quartile Q1, the third quartile Q3, and the interquartile range IQR of the sample values. Here, Q1 represents the value of the candidate indicator sample at the 25th quartile, Q3 represents the value of the candidate indicator sample at the 75th quartile, and IQR represents the difference between Q3 and Q1, i.e., IQR = Q3 - Q1. When a candidate index value is less than Q1-1.5IQR or greater than Q3+1.5IQR, it is marked as an outlier and truncated using the boundary value within the corresponding salinity level.
[0037] The standardization process uses Z-score standardization to eliminate dimensional interference for subsequent feature selection and PLSR modeling.
[0038] S3. Construct the enhanced feature matrix; Soil ecological function derived features are constructed based on the standardized original feature matrix, and the standardized original feature matrix and the soil ecological function derived features are merged to obtain the enhanced feature matrix; The derived characteristics of soil ecological functions include: salt stress regulation characteristics, comprehensive nutrient supply characteristics, enzyme-catalyzed nutrient transformation characteristics, and microbial ecological response characteristics.
[0039] In this embodiment, based on the correlation analysis of the original indicators, derived characteristics of soil ecological function are further constructed. These derived characteristics are not... Figure 2 Instead of directly displaying the original detection indicators, the model is based on the original soil ecological function indicators and is obtained through standardization and combination calculations to enhance the model's ability to express salt stress, nutrient supply, enzymatic nutrient transformation and microbial ecological response.
[0040] Preferably, in this example, the salt stress regulation characteristics include: salt-alkali stress index (SAI) and salt-nutrient inhibition characteristics (SNI); the comprehensive nutrient supply characteristics include: comprehensive nutrient supply index (NSI); the enzyme-catalyzed nutrient transformation characteristics include: enzyme activity comprehensive index (EAI) and carbon-nitrogen transformation balance characteristics (CNE); and the microbial ecological response characteristics include: microbial stability characteristics (MSI) and microbial-enzyme activity synergy characteristics (MES), as detailed below: Construct the salt-alkali stress index SAI: SAI = Z(SC) + Z(pH); Where SC represents soil salt content, pH represents soil acidity / alkalinity, and Z represents standardized treatment.
[0041] Construct the salt-nutrient inhibition feature SNI: SNI=Z(SC) / [Z(OM)+Z(TN)+Z(TP)+Z(TK)+ε]; Where OM represents organic matter, TN represents total nitrogen, TP represents total phosphorus, TK represents total potassium, and ε is a constant to prevent the denominator from being zero.
[0042] Construct the comprehensive nutrient supply index NSI: NSI = w1Z(NH4) + -N)+w2Z(NO3 - -N)+w3Z(OM)+w4Z(TN)+w5Z(TP)+w6Z(TK); Among them, NH4 + -N indicates ammonium nitrogen, NO3 - -N represents nitrate nitrogen, and w1 to w6 are weighting coefficients.
[0043] Construct the enzyme activity comprehensive index EAI: EAI = v1Z(BG) + v2Z(NAG) + v3Z(LAP); Where BG represents β-glucosidase, NAG represents acetamidoglucosidase, LAP represents L-leucine aminopeptidase, and v1 to v3 are weighting coefficients.
[0044] Construct the carbon-nitrogen conversion equilibrium characteristic CNE: CNE=Z(BG) / [Z(NAG)+Z(LAP)+ε].
[0045] Construct the microbial stability characteristic (MSI): MSI = Z(DIV) - Z(KSA).
[0046] Here, DIV represents microbial diversity, and KSA represents the abundance of key species related to salt stress.
[0047] Constructing a microbial-enzyme synergistic feature MES: MES = Z(DIV) × Z(EAI).
[0048] By combining standardized original features with soil ecological function-derived features in the above manner, an enhanced feature matrix is formed. This enhanced feature matrix can... Figure 2 Based on the correlation of the original indicators, we further express the salt stress intensity, nutrient supply level, enzymatic nutrient conversion capacity, and microbial ecological response status, providing input for subsequent key feature screening and wheat yield prediction model construction.
[0049] S4. Construct continuous output target variables and relative production reduction rate target variables; A continuous yield target variable was constructed based on the measured wheat yield of each plot, and a relative yield reduction target variable was constructed based on the reference yield of the target area. The continuous yield target variable is used for key feature screening and wheat yield prediction model construction in subsequent steps; the relative yield reduction rate target variable is used for subsequent capacity level and yield reduction risk level classification.
[0050] In this embodiment, the continuous yield target variable is directly adopted as the measured wheat yield per unit area of each plot, and the continuous yield target variable = measured wheat yield of the plot; The target variable for relative yield reduction rate = (reference yield of target area - actual wheat yield of corresponding plot) / reference yield of target area; The reference yield for the target area is usually taken as the average wheat yield of local non-saline-alkali normal cultivated land, the regional multi-year average yield, or the management target yield, as the "no yield reduction benchmark".
[0051] S5. Filter multiple candidate feature sets; Using the enhanced feature matrix as input features and the continuous output target variable as supervised learning target, LASSO feature selection, random forest feature importance selection, and K-optimal feature selection based on F-test are performed under multiple resampling conditions to select features in the enhanced feature matrix and obtain multiple candidate feature sets.
[0052] Specifically, in this embodiment, all features in the enhanced feature matrix are used as a full feature comparison set. This forms multiple candidate feature sets for subsequent model evaluation, including: a full feature comparison set, a LASSO-selected feature set, a random forest-selected feature set, and a K-optimal feature set based on the F-test.
[0053] Then, multiple candidate prediction models are established, including: partial least squares regression model, random forest regression model, LASSO regression model, linear regression model, support vector regression model and multilayer perceptron regression model; Each of the above candidate feature sets is input into multiple candidate prediction models. Then, for each feature-model combination formed by the candidate feature set and the candidate prediction model, leave-one-out cross-validation or K-fold cross-validation is used for training and testing. The cross-validation error, i.e., the coefficient of determination R², the root mean square error RMSE, and the mean absolute error MAE are calculated.
[0054] like Figure 3 As shown, bar charts or subplots are used to illustrate the performance of different feature-model combinations on R², RMSE, and MAE. A higher R² indicates a stronger model fit and predictive ability, while lower RMSE and MAE indicate smaller prediction errors. Significant differences in fit exist between different candidate feature sets and different prediction models. This result provides an evaluation basis for determining the optimal key feature set and the preferred prediction model.
[0055] S6. Select the optimal set of key features; For each candidate feature set, a comprehensive score is calculated based on the selection frequency of each candidate feature in multiple resampling, the cross-validation error of the prediction model built based on the candidate feature set, and the redundancy between each candidate feature within the candidate feature set. The candidate feature set with the highest comprehensive score is then determined as the optimal key feature set. Specifically, the combination of the optimal key feature set selected for stability screening and the partial least squares regression model showed better prediction accuracy and stability in cross-validation, such as... Figure 6As shown, the process includes: inputting the enhanced feature matrix - B resampling iterations - performing three types of filtering in each resampling iteration (including: LASSO feature filtering, random forest importance filtering, and K-optimal feature filtering) - calculating the selection frequency of candidate features - calculating the redundancy of the candidate feature set - comprehensive score (S core - Determine the optimal set of key features.
[0056] Preferably, in this embodiment, the optimal set of key features selected by stability screening and the partial least squares regression model are used as the preferred combination for subsequent yield prediction.
[0057] In this embodiment, the calculation of candidate feature set redundancy-comprehensive score means calculating a comprehensive score S for each candidate feature set. core : S core =αF stab +βF acc -γF red (1) Among them, F stab F represents the stability score. acc F represents the prediction accuracy score. red The redundancy score is represented by α, β, and γ, which are weighting coefficients. Stability score F stab The selection frequency P is determined based on the average selection frequency of each feature in the candidate feature set, and the selection frequency P is calculated for each candidate feature. j :P j =m j / B, where m j Let be the number of times the j-th candidate feature is selected in the B resampling processes, where B is the total number of resampling processes.
[0058] Prediction accuracy score F acc The candidate feature set is determined based on the coefficient of determination R², root mean square error RMSE, and mean absolute error MAE under leave-one-out cross-validation or K-fold cross-validation. In a preferred embodiment, the prediction accuracy score F acc =a·Norm(R²)-b·Norm(RMSE)-c·Norm(MAE), where a, b, and c are weighting coefficients, and Norm represents the normalization process.
[0059] Redundancy score F red It is determined based on the average of the pairwise absolute correlation coefficients of the features within the candidate feature set.
[0060] The overall score S coreThe set of highest-ranking candidate features is determined as the optimal set of key features. This process avoids the randomness of determining key features based on a single LASSO screening, a single random forest importance ranking, or a single K-optimal screening, while also reducing the impact of strongly correlated redundant features on model stability.
[0061] The process further clarifies how the present invention determines the optimal set of key features, that is, it does not simply adopt a certain general feature selection algorithm, but comprehensively considers the stability of feature selection, prediction accuracy and feature redundancy.
[0062] S7. Construct a wheat yield prediction model based on the optimal set of key features; A wheat yield prediction model for saline-alkali land is established based on the optimal set of key features, and the model is optimized, including optimization of the number of latent variables, optimization of model parameters, and residual correction.
[0063] Preferably, after determining the optimal set of key features, a wheat yield prediction model is established based on the optimal set of key features. The feature values corresponding to the optimal set of key features are used as inputs and the continuous yield target variable is used as outputs. The optimal prediction model is trained and optimized to obtain the wheat yield prediction model. In this embodiment, a partial least squares regression (PLSR) model is used as the prediction model.
[0064] First, the partial least squares regression model is optimized for the number of latent variables: partial least squares regression models are established in the range of 1 to Amax, and cross-validation is used to calculate the prediction error under different numbers of latent variables. The number of latent variables with the smallest prediction error or that meets the one standard error rule is determined as the optimal number of latent variables. Amax represents the preset maximum number of candidate latent variables, and Amax is not greater than the smaller value between the number of training samples minus 1 and the number of input key features.
[0065] Then, the model parameters are optimized: the model regression coefficients, intercept term, training sample prediction residuals and cross-validation error are calculated based on the training samples, and the final model parameters are determined with the optimization objectives of minimizing RMSE, lower MAE and higher R².
[0066] Furthermore, residual correction is performed: the prediction residual e of the training samples is calculated. i e i =Y i - i Y i Let i be the measured wheat yield of the i-th sample. i Predict the wheat yield for the i-th sample; Considering that salt content, pH, and microbial ecological response in saline-alkali land may cause systematic bias in the model, this embodiment establishes a residual correction model e. i =f(SC,pH,SAI,MSI,MES), where SC is soil salt content, SAI is salt stress index, MSI is microbial stability characteristics, and MES is microbial-enzyme synergistic characteristics. Finally, the output value of the partial least squares regression model is added to the residual correction value to obtain the corrected predicted wheat yield: corr= PLSR+ê, where, corr is the corrected predicted wheat yield. PLSR is the output value of the partial least squares regression model, and ê is the residual correction value of the residual correction model.
[0067] In establishing a wheat yield prediction model, this invention not only selects the prediction model type, but also further optimizes the number of latent variables, optimizes model parameters, and corrects residuals, thereby improving the model's adaptability to saline-alkali stress conditions and differences in microbial ecological responses.
[0068] S8. Output the predicted wheat yield, its capacity level, and the risk level of yield reduction; The optimal set of key features of the wheat-growing plots in the saline-alkali land to be tested is input into the optimized yield prediction model to obtain the predicted wheat yield. Based on the predicted wheat yield and the target variable of relative yield reduction rate, the model outputs the production capacity level and the yield reduction risk level. like Figure 4 The model displays the predicted wheat yield and production capacity level for each sample, as well as the relative yield reduction rate and yield reduction risk level for each sample. In this embodiment, the key features of the plot to be tested are input into the optimized wheat yield prediction model, which outputs the predicted wheat yield and classifies the production capacity level and yield reduction risk level.
[0069] Specifically, soil samples from the 0-20 cm topsoil layer of the wheat-growing plot in the saline-alkali land to be tested were collected. Soil indicators corresponding to the optimal key feature set were detected, and preprocessing and derived feature calculation were performed according to the method described in Example 2. The optimal key feature set of the plot was input into the optimized yield prediction model to obtain the predicted wheat yield of the plot.
[0070] Based on the predicted wheat yield and relative yield reduction rate, the production capacity level and yield reduction risk level are classified. Let the first yield threshold be T1, the second yield threshold be T2, the first yield reduction rate threshold be R1, and the second yield reduction rate threshold be R2, where T1 is greater than T2, and R2 is greater than R1.
[0071] When the predicted wheat yield is not lower than T1 and the relative yield reduction rate is lower than R1, it is judged as high capacity and low risk; When the predicted wheat yield is lower than T1 but not lower than T2, or the relative yield reduction rate is not lower than R1 but lower than R2, it is judged as medium capacity and medium risk. When the predicted wheat yield is lower than T2, or the relative yield reduction rate is not lower than R2, it is judged as low production capacity and high risk.
[0072] In one specific implementation, T1 is set to 8000 kg / ha, T2 is set to 3000 kg / ha, R1 is set to 0.20, and R2 is set to 0.70. These thresholds can be adjusted based on the multi-year wheat yield distribution of the target area, the yield of non-saline reference plots, or management needs.
[0073] Forecast production map Figure 4 The different colored bars in the chart represent the production capacity levels of different samples, and the dashed lines represent the production capacity level thresholds. The relative production reduction rate chart is... Figure 5 Different colors represent low, medium, and high risk, while dashed lines represent yield reduction thresholds R1 and R2. This visually reflects the predicted wheat yield, production capacity level, and yield reduction risk level for each plot of land being tested.
[0074] Example 2 This embodiment provides a system for predicting wheat yield and risk classification in saline-alkali land, such as... Figure 7 As shown, the system includes: a sample acquisition module, an indicator detection module, a data transmission module, a sample database module, a data preprocessing module, a derived feature construction module, a target variable construction module, a key feature stability screening module, a model construction and optimization module, a yield prediction and risk classification module, and a user display terminal; wherein, the sample database module, data preprocessing module, derived feature construction module, target variable construction module, key feature stability screening module, model construction and optimization module, and yield prediction and risk classification module are deployed on the server / analysis platform.
[0075] The sample collection module is used to collect soil samples from the 0-20cm topsoil layer of the target plot and obtain the measured wheat yield. The index detection module is used to detect and obtain the original set of soil ecological function indicators; The data transmission module is used to upload the original set of soil ecological function indicators to the sample database module in real time; The sample database module is used to store the original set of soil ecological function indicators; The data preprocessing module is used to perform missing value imputation, outlier processing and standardization on the original soil ecological function index set to obtain a standardized original feature matrix. The derived feature construction module is used to construct soil ecological function derived features and generate an enhanced feature matrix; The target variable construction module is used to construct continuous output target variables and relative production reduction rate target variables; The key feature stability screening module is used to determine the optimal set of key features through multi-method screening and comprehensive scoring under multiple resampling. The model building and optimization module is used to establish a production prediction model and perform latent variable number optimization, parameter optimization and residual correction. The yield prediction and risk classification module is used to output the predicted wheat yield, production capacity level, and yield reduction risk level. The user display terminal is used to display the prediction and classification results.
[0076] This embodiment can realize a complete process from soil sample collection, index detection, data preprocessing, derived feature construction, key feature screening, model optimization, yield prediction to risk classification output, and is applicable to wheat planting management in saline-alkali land, plot productivity evaluation and yield reduction risk warning.
[0077] Example 3 This embodiment also provides an electronic device, including: At least one processor; and The memory stores instructions that, when executed by the at least one processor, cause the at least one processor to perform a method for predicting wheat yield and risk classification in saline-alkali land as described above.
[0078] In this embodiment, the electronic device may include, but is not limited to: personal computer, server computer, workstation, desktop computer, laptop computer, notebook computer, mobile computing device, smartphone, tablet computer, cellular phone, personal digital assistant (PDA), handheld device, messaging device, wearable computing device, consumer electronic device, etc.
[0079] Example 4 This embodiment also provides a machine-readable storage medium storing executable instructions that, when executed, cause the machine to perform a method for predicting wheat yield and risk classification in saline-alkali land as described above.
[0080] Specifically, a system or apparatus equipped with a readable storage medium may be provided, on which software program code implementing the functions of any of the embodiments described above is stored, and the computer or processor of the system or apparatus can read and execute the instructions stored in the readable storage medium.
[0081] In this case, the program code read from the readable medium itself can perform the functions of any of the above embodiments, and therefore the machine-readable code and the readable storage medium storing the machine-readable code constitute a part of this specification.
[0082] Examples of readable storage media include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD-RW), magnetic tapes, non-volatile memory cards, and ROMs. Alternatively, program code can be downloaded from a server computer or the cloud via a communication network.
[0083] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0084] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, systems, and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0085] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0086] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0087] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the technical solutions of the present invention, and are not intended to limit the specific implementation of the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the claims of the present invention should be included within the protection scope of the claims of the present invention.
Claims
1. A method for predicting wheat yield and classifying risks in saline-alkali land, characterized in that, The method includes: S1. Collect topsoil samples at a predetermined depth in the target saline-alkali land wheat planting plots, obtain the actual wheat yield of the corresponding plots, and detect to form a set of original soil ecological function indicators. S2. Perform missing value imputation, outlier processing and standardization on the original soil ecological function index set to obtain the standardized original feature matrix; S3. Construct soil ecological function derived features based on the standardized original feature matrix, and merge the standardized original feature matrix with the soil ecological function derived features to obtain the enhanced feature matrix; The derived characteristics of soil ecological functions include: salt stress regulation characteristics, comprehensive nutrient supply characteristics, enzyme-catalyzed nutrient transformation characteristics, and microbial ecological response characteristics; S4. Construct a continuous yield target variable based on the measured wheat yield of each plot, and construct a relative yield reduction target variable based on the reference yield of the target area; S5. Using the enhanced feature matrix as the input feature and the continuous output target variable as the supervised learning target, LASSO feature selection, random forest feature importance selection and K-optimal feature selection based on F test are performed under multiple resampling conditions to obtain multiple candidate feature sets. S6. For each candidate feature set, calculate the comprehensive score of the candidate feature set based on the selection frequency of each candidate feature in the multiple resampling, the cross-validation error of the prediction model built based on the candidate feature set, and the redundancy between each candidate feature within the candidate feature set, and determine the candidate feature set with the highest comprehensive score as the optimal key feature set. S7. Based on the optimal set of key features, establish a wheat yield prediction model for saline-alkali land, and optimize the yield prediction model. The model optimization includes: optimization of the number of latent variables, optimization of model parameters, and residual correction. S8. Input the optimal set of key features of the wheat planting plot in the saline-alkali land to be tested into the optimized yield prediction model to obtain the predicted wheat yield, and output the production capacity level and yield reduction risk based on the predicted wheat yield and the target variable of relative yield reduction rate.
2. The method for predicting wheat yield and classifying risks in saline-alkali land according to claim 1, characterized in that, The missing value filling is specifically as follows: when the missing proportion of each soil indicator in the candidate indicator, i.e. the original soil ecological function indicator set, is not higher than the preset missing threshold, the median of the corresponding candidate indicator in the same salinity level plot is used for filling. When the missing percentage exceeds the preset missing threshold, the candidate indicator is removed. The outlier handling is specifically as follows: sort the sample values of each candidate indicator in ascending order, and calculate the first quartile Q1, the third quartile Q3, and the interquartile range IQR of the sample values. Here, Q1 represents the value of the candidate indicator sample at the 25th quartile, Q3 represents the value of the candidate indicator sample at the 75th quartile, and IQR represents the difference between Q3 and Q1, i.e., IQR = Q3 - Q1. When a candidate index value is less than Q1-1.5IQR or greater than Q3+1.5IQR, it is marked as an outlier and truncated using the boundary value within the corresponding salinity level. The standardization process used is Z-score standardization.
3. The method for predicting wheat yield and classifying risks in saline-alkali land according to claim 1, characterized in that, The salt stress regulation characteristics include: salt-alkali stress index and salt-nutrient inhibition characteristics; The salt stress index SAI is calculated as Z(SC) + Z(pH), where SC represents soil salt content, pH represents soil acidity / alkalinity, and Z represents standardized treatment. The salt-nutrient inhibition characteristic SNI is calculated as follows: SNI = Z(SC) / [Z(OM) + Z(TN) + Z(TP) + Z(TK) + ε], where OM represents organic matter, TN represents total nitrogen, TP represents total phosphorus, TK represents total potassium, and ε is a constant to prevent the denominator from being zero. The comprehensive nutrient supply characteristics include: Comprehensive Nutrient Supply Index NSI = w1Z(NH4) + -N)+w2Z(NO3 - -N)+w3Z(OM)+w4Z(TN)+w5Z(TP)+w6Z(TK), where NH4 + -N indicates ammonium nitrogen, NO3 - -N represents nitrate nitrogen, and w1 to w6 are the set weighting coefficients; The enzymatic nutrient conversion characteristics include: comprehensive enzyme activity index and carbon-nitrogen conversion balance characteristics; The enzyme activity comprehensive index EAI is calculated as follows: EAI = v1Z(BG) + v2Z(NAG) + v3Z(LAP), where BG represents β-glucosidase, NAG represents acetamidoglucosidase, LAP represents L-leucine aminopeptidase, and v1 to v3 are set weighting coefficients. Among them, the carbon-nitrogen conversion equilibrium characteristic is CNE=Z(BG) / [Z(NAG)+Z(LAP)+ε]; The microbial ecological response characteristics include: microbial stability characteristics and microbial-enzyme synergistic characteristics; Among them, the microbial stability characteristic MSI = Z(DIV) - Z(KSA), where DIV represents microbial diversity and KSA represents the abundance of key species related to salt stress; Among them, the microbial-enzyme activity synergistic feature MES = Z(DIV) × Z(EAI).
4. The method for predicting wheat yield and classifying risks in saline-alkali land according to claim 1, characterized in that, The comprehensive score S of the candidate feature set core as follows: S core =αF stab +βF acc -γF red (1) Among them, F stab The stability score is determined based on the average selection frequency of each feature in the candidate feature set; F acc For accuracy scores, the coefficient of determination R², root mean square error (RMSE), and mean absolute error (MAE) are determined based on leave-one-out cross-validation or K-fold cross-validation; F red The redundancy score is determined based on the average of the pairwise absolute correlation coefficients of features within the candidate feature set; α, β, and γ are weighting coefficients.
5. The method for predicting wheat yield and classifying risks in saline-alkali land according to claim 1, characterized in that, The model optimization specifically includes: First, the number of latent variables in the production forecast model is optimized: the model is built in the range of 1 to the preset maximum number of latent variables Amax, and the prediction error under different numbers of latent variables is calculated by cross-validation. The number of latent variables with the smallest prediction error or that meets the unified standard error rule is determined as the optimal number of latent variables. Then, the model parameters are optimized: the model regression coefficients, intercept term, training sample prediction residuals and cross-validation error are calculated based on the training samples, and the final model parameters are determined with the optimization objectives of minimizing RMSE, lower MAE and higher R². Furthermore, residual correction is performed: the prediction residuals of the training samples are calculated, a residual correction model is established that combines the residuals with soil salinity, pH and microbial ecological response characteristics, and the corrected residual values are superimposed with the output values of the yield prediction model to obtain the corrected predicted yield.
6. The method for predicting wheat yield and classifying risks in saline-alkali land according to claim 1, characterized in that, The original soil ecological function index set includes: soil physicochemical properties, soil enzyme activity, and microbial indicators; The soil physicochemical properties include: pH, salt content, ammonium nitrogen, nitrate nitrogen, organic matter, total nitrogen, total phosphorus, and total potassium. The soil enzyme activity indicators include: β-glucosidase, acetamidoglucosidase and L-leucine aminopeptidase; The microbial indicators include: microbial diversity, abundance of key species, and microbial community characteristics.
7. A system for predicting wheat yield and risk classification in saline-alkali land, characterized in that, The system includes: a sample acquisition module, an indicator detection module, a data transmission module, a sample database module, a data preprocessing module, a derived feature construction module, a target variable construction module, a key feature stability screening module, a model construction and optimization module, a yield prediction and risk classification module, and a user display terminal; The sample collection module is used to collect soil samples from the cultivated layer at a preset depth in the target plot and obtain the measured wheat yield. The index detection module is used to detect and obtain the original set of soil ecological function indicators; The data transmission module is used to upload the original set of soil ecological function indicators to the sample database module in real time; The sample database module is used to store the original set of soil ecological function indicators; The data preprocessing module is used to perform missing value imputation, outlier processing and standardization on the original soil ecological function index set to obtain a standardized original feature matrix. The derived feature construction module is used to construct soil ecological function derived features and generate an enhanced feature matrix; The target variable construction module is used to construct continuous output target variables and relative production reduction rate target variables; The key feature stability screening module is used to determine the optimal set of key features through multi-method screening and comprehensive scoring under multiple resampling. The model building and optimization module is used to establish a production prediction model and perform latent variable number optimization, parameter optimization and residual correction. The yield prediction and risk classification module is used to output the predicted wheat yield, production capacity level, and yield reduction risk level. The user display terminal is used to display the prediction and classification results.
8. An electronic device, characterized in that, The electronic device includes: processor; A memory on which computer programs that can run on the processor are stored; When the computer program is executed by the processor, it implements the steps of a method for predicting wheat yield and risk classification in saline-alkali land as described in any one of claims 1 to 6.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 6.