A method, system, device and medium for evaluating corn achievable yield

By constructing a set of feature indicators and an optimal machine learning model, combined with an interpretable artificial intelligence framework, the problem of multi-dimensional factor fusion and dynamic yield gap prediction under future climate scenarios in existing maize yield assessment methods has been solved, achieving high-precision prediction of maize yield potential and optimization of planting management.

CN122509740APending Publication Date: 2026-08-04CHINA AGRI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA AGRI UNIV
Filing Date
2026-03-31
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

Existing methods for assessing maize yield are difficult to base on technical upper limits, lack the integration of multi-dimensional factors, have poor model interpretability, and cannot couple future climate scenarios and agronomic strategies on a large-scale grid scale to predict dynamic yield gaps.

Method used

By acquiring historical multi-source data, we conduct agronomic mechanism analysis to construct a set of feature indicators, use multiple machine learning models for cross-validation to select the optimal model, and introduce an interpretable artificial intelligence framework for analysis to identify yield limiting and promoting factors. Combined with sowing date regulation strategies, we predict yield potential and evaluate agronomic strategies under multiple future climate scenarios.

Benefits of technology

It achieves a closed-loop process from historical data mining to future scenario prediction, possesses model interpretability and dynamic evaluation capabilities, improves the accuracy and robustness of maize yield potential prediction, and provides optimal planting management solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122509740A_ABST
    Figure CN122509740A_ABST
Patent Text Reader

Abstract

The present application relates to a corn available yield evaluation method, system, device and medium, the method comprising: obtaining historical multi-source data; performing agronomic mechanism analysis on the historical multi-source data, and constructing a feature index set according to the analysis result; training a plurality of machine learning models respectively by taking the feature index set as an input variable and taking corn available yield data as a target variable, and evaluating an optimal machine learning model according to cross-validation; analyzing the optimal model by using an explainable artificial intelligence framework, inputting a planting management strategy generated based on the analysis result into the optimal model for prediction; comparing the predicted corn available yield with a benchmark yield of the same period, and screening an optimal planting management scheme according to the yield performance of different planting management strategies. The present application realizes high-resolution, explainable quantitative evaluation of corn available yield at a regional scale, and provides a scientific basis for agricultural policy making and climate change adaptation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of crop yield evaluation technology, and in particular to a method, system, equipment and medium for evaluating the yield of maize. Background Technology

[0002] As an important food crop, increasing maize yield is crucial for ensuring national food security. Currently, there is a significant gap between actual maize yield and the achievable yield under optimal varieties and management conditions. Accurately quantifying this yield gap is key to tapping production potential and optimizing resource allocation.

[0003] Currently, methods for assessing crop yield potential have the following shortcomings: First, process-based crop growth models have complex parameters, are difficult to calibrate locally, and have limited ability to simulate extreme climate stresses. Second, traditional statistical models are unable to effectively characterize the complex nonlinear relationship between yield and multiple factors. Third, recent machine learning yield prediction studies have mostly focused on farmers' actual yields, failing to use "available yield" as an assessment benchmark representing the upper limit of technology, and the models lack interpretability, making it difficult to identify yield-limiting factors to guide production practices. In addition, for large-scale assessments, existing technologies do not involve high-resolution grid scales and lack the ability to dynamically predict and assess gaps by coupling future multi-path climate scenarios with specific agronomic adaptation strategies.

[0004] Therefore, there is an urgent need to establish an assessment method that integrates multiple driving factors such as variety, meteorology, soil, and management, has model interpretability, and can quantitatively predict yield potential and development gaps under multiple future scenarios. Summary of the Invention

[0005] (a) Technical problems to be solved

[0006] In view of the above-mentioned shortcomings and deficiencies of the prior art, the present invention provides a method, system, equipment and medium for evaluating the yield of maize, which solves the technical problems of existing yield assessment methods that are difficult to base on technical upper limit benchmarks, lack multi-dimensional factor system integration, have poor model interpretability, and cannot couple future climate scenarios and agronomic strategies on a large-scale grid scale to predict dynamic yield gaps.

[0007] (II) Technical Solution

[0008] To achieve the above objectives, the main technical solutions adopted by the present invention include:

[0009] In a first aspect, embodiments of the present invention provide a method for evaluating the available yield of maize, comprising:

[0010] Historical multi-source data on corn yield can be obtained by acquiring related corn varieties;

[0011] Agronomic mechanism analysis was conducted on historical multi-source data, and a set of characteristic indicators including varietal characteristic indicators, environmental condition indicators, and planting management characteristic indicators was constructed based on the analysis results.

[0012] Using a set of feature indicators as input variables and corn yield data from historical multi-source data as target variables, we trained multiple preset machine learning models and selected the optimal machine learning model based on the cross-validation evaluation results.

[0013] The optimal machine learning model is analyzed using an interpretable artificial intelligence framework, and the planting management strategy generated based on the analysis results is input into the optimal machine learning model to predict the yield of maize at the target spatiotemporal scale.

[0014] The predicted yield of maize is compared with the baseline yield of the same period. The yield performance under different planting management strategies is evaluated based on the comparison results, and the optimal planting management plan for each environmental condition is selected.

[0015] Optionally, obtaining historical multi-source data on the yield of associated corn includes:

[0016] Field measurement data from multiple test sites in the target area were collected, and the collected data were preprocessed by data cleaning and normalization to obtain historical maize yield data for the target area.

[0017] Extract relevant data from a pre-set database that matches the time window and geographical range of historical corn yield data. The relevant data includes meteorological data, soil data, corn growth period phenological data, and planting management data.

[0018] Optionally, agronomic mechanism analysis is performed on historical multi-source data, and a feature index set including varietal characteristic indicators, environmental condition indicators, and planting management characteristic indicators is constructed based on the analysis results.

[0019] Agronomic mechanism analysis was conducted on historical multi-source data. Based on the analysis results, ecological factors related to maize yield were extracted from the historical multi-source data to determine the initial feature set, which includes varietal characteristics, environmental conditions, and planting management measures.

[0020] The initial feature set is subjected to correlation analysis and importance ranking, redundant and low-contribution features are removed, and the feature index set is optimized and constructed.

[0021] Among them, the varietal characteristic indicators include the length of the growing season and the effective accumulated temperature during the growing season, which reflect the heat demand characteristics of the variety; the environmental condition indicators include meteorological characteristic indicators and soil characteristic indicators; and the planting management characteristic indicators include planting density.

[0022] Optionally, using a set of feature indicators as input variables and corn yield data from historical multi-source data as target variables, multiple preset machine learning models are trained separately, and the optimal machine learning model is selected based on cross-validation evaluation results, including:

[0023] The feature index set and the corn yield data from historical multi-source data are divided into training set and test set according to a preset ratio;

[0024] The training set is used to train and cross-validate a variety of pre-defined machine learning models, including random forest, XGBoost, LightGBM, multilayer perceptron, support vector machine and partial least squares regression.

[0025] The test set is input into each trained machine learning model for testing. The evaluation metrics for each machine learning model are obtained based on the test results. The evaluation metrics include the coefficient of determination and the normalized root mean square error.

[0026] Based on the evaluation metrics of each machine learning model, a comprehensive ranking was conducted, and the machine learning model with the highest comprehensive ranking was selected as the optimal model for predicting the yield of corn.

[0027] Optionally, the optimal machine learning model is analyzed using an interpretable artificial intelligence framework, and the planting management strategy generated based on the analysis results is input into the optimal machine learning model to predict the achievable maize yield at the target spatiotemporal scale, including:

[0028] The optimal machine learning model is analyzed using an interpretable artificial intelligence framework to obtain the SHAP values ​​of each feature index in the feature index set for all predicted samples.

[0029] Based on the sign and absolute value of the SHAP value, each characteristic indicator is identified as a promoting factor or a limiting factor, and the contribution and direction of influence of each characteristic indicator on the yield prediction results of maize are quantified.

[0030] The identified promoting or restricting factors are used as regulatory targets to adjust the sowing date in the planting management strategy. The planting management strategy includes a static strategy that keeps the sowing date unchanged and an adaptive strategy that dynamically adjusts the sowing date based on historical trends.

[0031] By adjusting the planting date to modify environmental factors during the maize growth period, the optimal machine learning model is driven to predict the yield of maize under different environmental conditions in the future, and to evaluate the potential value of different planting management strategies to improve maize yield.

[0032] Optionally, the predicted yield of maize is compared with the baseline yield for the same period. Based on the comparison results, the yield performance under different planting management strategies is evaluated, and the optimal planting management plan for each environmental condition is selected, including:

[0033] The difference between the predicted maize yield under different planting management strategies and the benchmark yield in the same period is calculated to obtain the yield change under each strategy.

[0034] Based on the magnitude of yield change and combined with the planting management parameters corresponding to each strategy, a yield performance evaluation matrix is ​​constructed. The yield performance evaluation matrix is ​​used to quantify the yield response characteristics of each planting management strategy under different combinations of environmental conditions.

[0035] With the goal of maximizing yield, the planting management strategy with the best yield performance is selected as the recommended solution from the yield performance evaluation matrix for each target environmental condition.

[0036] Secondly, embodiments of the present invention provide a corn yield assessment system, comprising:

[0037] The data acquisition module is used to acquire historical multi-source data on the yield of related corn.

[0038] The feature processing module is used to perform agronomic mechanism analysis on historical multi-source data and construct a feature index set that includes variety characteristic indicators, environmental condition indicators, and planting management characteristic indicators based on the analysis results.

[0039] The model training and selection module is used to train multiple preset machine learning models with feature index set as input variable and corn yield data from historical multi-source data as target variable, and select the optimal machine learning model based on cross-validation evaluation results.

[0040] The yield prediction module is used to analyze the optimal machine learning model using an interpretable artificial intelligence framework, and input the planting management strategy generated based on the analysis results into the optimal machine learning model to predict the yield of corn at the target spatiotemporal scale.

[0041] The evaluation and scheme selection module is used to compare the predicted corn yield with the benchmark yield of the same period, evaluate the yield performance under different planting management strategies based on the comparison results, and select the optimal planting management scheme for each environmental condition.

[0042] Thirdly, embodiments of the present invention provide an electronic device, comprising:

[0043] At least one processor;

[0044] and memory that is communicatively connected to at least one processor;

[0045] The memory stores instructions that can be executed by at least one processor, which enables the at least one processor to perform the corn yield evaluation method described above.

[0046] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the corn yield evaluation method described above.

[0047] (III) Beneficial Effects

[0048] The beneficial effects of this invention are as follows: The proposed method for evaluating the available yield of maize firstly constructs an interpretable feature index set integrating variety characteristics, environmental conditions, and planting management dimensions through agronomic mechanism analysis, overcoming the problem of traditional black-box models lacking agronomic basis. Secondly, it utilizes cross-validation of multiple machine learning models to screen the optimal prediction model, ensuring the accuracy and robustness of the available yield prediction. Thirdly, it introduces an interpretability framework to analyze the optimal model, identifying available yield limiting and promoting factors, and based on this, achieves dynamic prediction of yield potential under multiple future climate scenarios and quantitative evaluation of agronomic strategies through sowing date control strategies. Finally, by constructing a yield performance evaluation matrix and multi-objective screening rules, it selects the optimal planting management scheme for different environmental conditions. Compared to existing technologies, this invention integrates multi-dimensional factors such as variety, meteorology, soil, and management, realizing a closed-loop chain from historical data mining to future scenario prediction, and possesses model interpretability and dynamic evaluation capabilities. Attached Figure Description

[0049] Figure 1 This is a flowchart illustrating a method for evaluating the yield of corn according to an embodiment of the present invention.

[0050] Figure 2 A flowchart illustrating a method for evaluating the yield of maize according to an embodiment of the present invention;

[0051] Figure 3 The prediction accuracy (R) of six candidate machine learning models provided in an embodiment of the present invention 2 Comparison of bar chart and line chart of error (NRMSE) and total error (NRMSE);

[0052] Figure 4 This is a schematic diagram illustrating the global importance ranking of various feature indicators obtained by SHAP analysis and their individual contributions to yield, provided as an embodiment of the present invention.

[0053] Figure 5This is a time trend chart of the national maize yield from 2015 to 2050 under different CMIP6 climate scenarios and planting strategies, provided as an embodiment of the present invention.

[0054] Figure 6 A graph showing the evolution trend of the gap between historical actual output, simulated available output, and exploitable output from 2001 to 2050, provided as an embodiment of the present invention.

[0055] Figure 7 This is a schematic diagram of the composition of a corn yield assessment system provided in an embodiment of the present invention. Detailed Implementation

[0056] To better explain and facilitate understanding of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0057] refer to Figures 1 to 7 As shown in the embodiment of the present invention, a method for evaluating the available yield of maize includes: acquiring historical multi-source data related to maize yield; acquiring historical multi-source data related to the available yield of maize; performing agronomic mechanism analysis on the historical multi-source data, and constructing a feature index set including variety characteristic indicators, environmental condition indicators, and planting management characteristic indicators based on the analysis results; using the feature index set as input variables and the available yield data of maize in the historical multi-source data as target variables, training multiple preset machine learning models respectively, and selecting the optimal machine learning model based on the cross-validation evaluation results; using an interpretable artificial intelligence framework to analyze the optimal machine learning model, and inputting the planting management strategy generated based on the analysis results into the optimal machine learning model to predict the available yield of maize at the target spatiotemporal scale; comparing the predicted available yield of maize with the benchmark yield of the same period, evaluating the yield performance under different planting management strategies based on the comparison results, and selecting the optimal planting management scheme for each environmental condition.

[0058] This embodiment first constructs an interpretable feature index set integrating variety characteristics, environmental conditions, and planting management from multiple dimensions through agronomic mechanism analysis, overcoming the problem of traditional black-box models lacking agronomic basis. Next, it uses cross-validation of multiple machine learning models to screen the optimal prediction model, ensuring the accuracy and robustness of yield prediction. Then, it introduces an interpretability framework to analyze the optimal model, identifying yield-limiting and promoting factors, and based on this, achieves dynamic prediction of yield potential under multiple future climate scenarios and quantitative evaluation of agronomic strategies through sowing date control strategies. Finally, by constructing a yield performance evaluation matrix and multi-objective screening rules, it selects the optimal planting management scheme for different environmental conditions. Compared with existing technologies, this invention integrates multi-dimensional factors such as variety, meteorology, soil, and management, realizing a closed-loop chain from historical data mining to future scenario prediction, and possesses model interpretability and dynamic evaluation capabilities.

[0059] To better understand the above technical solutions, exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that the present invention can be understood more clearly and thoroughly, and that the scope of the present invention can be fully conveyed to those skilled in the art.

[0060] Specifically, refer to Figure 1 and 2 As shown, the method for evaluating the yield of maize proposed in this embodiment of the invention may include the following steps S100 to S500:

[0061] S100: Obtain historical multi-source data on the yield of related corn.

[0062] In this embodiment, step S100 may include the following sub-steps S110 to S120:

[0063] S110. Field measurement data from multiple test sites in the target area were collected, and the sampled data were preprocessed by data cleaning and normalization to obtain historical maize yield data for the target area.

[0064] In one specific embodiment, regional experimental data covering five major maize-producing areas in China—the spring-sown maize region in northern China, the summer-sown maize region in the Huang-Huai-Hai Plain, the mountainous maize region in southwestern China, the irrigated maize region in northwestern China, and the hilly maize region in southern China—were extracted from 2001 to 2014. The highest yield records under optimal cultivation treatments were selected for each experimental site and each year. After data cleaning and normalization preprocessing, 3764 valid samples were obtained. This sample set was defined as the historical maize yield data.

[0065] S120. Extract relevant data from the preset database that matches the time window and geographical range of historical corn yield data. The relevant data includes meteorological data, soil data, corn growth period phenological data, and planting management data.

[0066] In one specific embodiment, the historical daily data (2001-2014) in the meteorological data was obtained through the China Meteorological Data Sharing Service System and supplemented with historical simulation grid data (0.5°×0.5°) from five global climate models (GFDL-ESM4, IPSL-CM6A-LR, MPI-ESM1-2-HR, MRI-ESM2-0, and UKESM1-0-LL) provided by ISIMIP (Inter-Sectoral Impact Model Intercomparison Project). Multi-model ensemble averaging was employed to reduce uncertainty. The future scenario data (2015-2050) in the meteorological data directly adopted the results of simulations using the aforementioned models within the CMIP6 (Sixth Coupled Model Intercomparison Project) framework, based on three different greenhouse gas emission scenarios (SSP1-2.6, SSP3-7.0, and SSP5-8.5).

[0067] Soil data: Global soil organic carbon content data were obtained from the SoilGrids2.0 platform. Four standard depth raster layers of 0~5cm, 5~15cm, 15~30cm and 30~60cm were extracted for the Chinese region. The original spatial resolution was 250 meters.

[0068] Phenological data for maize growth period: The 2001-2020 China maize phenological dataset released by the National Data Center for Ecological Sciences, with a spatial resolution of 30m, was used. This dataset includes the start and end dates of the growing season, as well as the length of the growing season. The 2001-2024 China maize planting distribution dataset released by the National Data Center for Ecological Sciences, with a spatial resolution of 30m, was also used. To focus on analyzing climate and management effects, the planting area in future simulations after 2020 was fixed at the average value from 2015-2020.

[0069] Planting management data: Planting density is directly adopted from the values ​​recorded in the trial reports of each region. In future scenario simulations, optimized recommended densities determined based on agronomic trials and tailored to different regions will be used.

[0070] S200. Perform agronomic mechanism analysis on historical multi-source data, and construct a feature index set that includes varietal characteristic indicators, environmental condition indicators, and planting management characteristic indicators based on the analysis results.

[0071] In this embodiment, step S200 may include the following sub-steps S210 to S220:

[0072] S210. Conduct agronomic mechanism analysis on historical multi-source data, and extract ecological factors related to maize yield from historical multi-source data based on the analysis results, and determine the initial feature set including variety characteristics, environmental conditions and planting management measures.

[0073] S220. Perform correlation analysis and importance ranking on the initial feature set, remove redundant and low-contribution features, and optimize the construction of the feature index set.

[0074] Furthermore, varietal characteristic indicators include the length of the growing season and the effective accumulated temperature during the growing season, which reflect the heat demand characteristics of the variety; environmental condition indicators include meteorological characteristic indicators and soil characteristic indicators; and planting management characteristic indicators include planting density. In this embodiment, for each sample unit (historical site or future grid), the following indicators are dynamically calculated based on its geographical coordinates, year, phenological period, and corresponding meteorological sequence:

[0075] 1. Variety characteristic indicators include:

[0076] Growth period length: The number of days from sowing to maturity is calculated directly based on phenological data.

[0077] Effective accumulated temperature during the reproductive period:

[0078] (1)

[0079] In equation (1), Let T be the average daily temperature on day i, and n be the entire growth period of maize. b This is the biological lower limit temperature for maize growth, set at 10℃.

[0080] 2. Meteorological characteristic indicators include:

[0081] The cumulative total solar radiation, average temperature, average maximum temperature, average minimum temperature, cumulative precipitation, and average daily temperature range are all cumulative or average values ​​within the reproductive period.

[0082] Accumulated high temperature days:

[0083] (2)

[0084] Equation (2), Let T be the daily maximum temperature on day i. h The critical temperature for high-temperature stress in maize is set at 30℃. (HDD) i A day with a high temperature of over 30°C.

[0085] Cumulative low temperature days:

[0086] (3)

[0087] Equation (3), Let T be the minimum daily temperature on day i. l The critical temperature for low-temperature stress in maize is 8℃, CDD i A day with a low temperature of less than 8°C.

[0088] 3. Soil characteristic indicators include:

[0089] Soil organic carbon content: The soil organic carbon content of the 0-60cm soil layer is used. It is calculated by depth-weighted averaging of the soil organic carbon content of four standard soil layers: 0-5cm, 5-15cm, 15-30cm, and 30-60cm. The calculation formula is as follows:

[0090] (4)

[0091] In equation (4), These represent the soil organic carbon content of the four soil layers mentioned above.

[0092] 4. Planting management characteristic indicators include:

[0093] Planting density: The corresponding data was used directly.

[0094] S300. Using the feature index set as the input variable and the corn yield data from historical multi-source data as the target variable, train multiple preset machine learning models respectively, and select the optimal machine learning model based on the cross-validation evaluation results.

[0095] In this embodiment, step S300 may include the following sub-steps S310 to S340:

[0096] S310. Divide the feature index set and the corn yield data in the historical multi-source data into a training set and a test set according to a preset ratio.

[0097] S320 uses a training set to train and cross-validate a variety of preset machine learning models, including random forest, XGBoost, LightGBM, multilayer perceptron, support vector machine, and partial least squares regression.

[0098] S330. Input the test set into each trained machine learning model and run the test. Based on the test results, obtain the evaluation index of each machine learning model. The evaluation index includes the coefficient of determination and the normalized root mean square error.

[0099] S340. Based on the evaluation metrics of each machine learning model, a comprehensive ranking is performed, and the machine learning model with the highest comprehensive ranking is selected as the optimal corn yield prediction model.

[0100] In one specific embodiment, the obtained production data from 3764 stations are paired with the corresponding feature index set calculated in step S200 to form a complete training set. Then, the data is randomly divided into a training set and an independent test set at a ratio of 80% to 20%, with the training set containing approximately 3011 samples and the test set containing approximately 753 samples. Next, the six machine learning models mentioned above are trained using Python libraries such as scikit-learn and lightGBM. For each model, grid search under five-fold cross-validation is used to optimize key hyperparameters (such as learning_rate and num_leaves in LightGBM). Finally, the coefficient of determination (R²) on the test set is used to calculate the training set. 2 Using normalized root mean square error (NRMSE) as the evaluation criteria, the evaluation results are as follows: Figure 3 As shown, the LightGBM model achieved the highest R-value on the test set. 2 With a mean of 0.74 and the lowest NRMSE (12%), LightGBM significantly outperformed other models. Therefore, LightGBM was ultimately selected as the final model for predicting available maize yields under current conditions.

[0101] S400: The optimal machine learning model is analyzed using an interpretable artificial intelligence framework, and the planting management strategy generated based on the analysis results is input into the optimal machine learning model to predict the yield of corn at the target spatiotemporal scale.

[0102] In this embodiment, the SHAP interpretability framework is introduced to analyze the optimal model, identify the limiting and promoting factors of each characteristic index for the obtainable yield, and design a sowing date regulation strategy accordingly. This enables dynamic prediction of yield potential under multiple future climate scenarios and quantitative evaluation of agronomic strategies. Specifically, step S400 may include the following sub-steps S410 to S440:

[0103] S410. Use an interpretable artificial intelligence framework to analyze the optimal machine learning model and obtain the SHAP value of each feature index in the feature index set for all predicted samples.

[0104] S420. Based on the sign and absolute value of the SHAP value, identify each characteristic indicator as a promoting factor or a restrictive factor, and quantify the contribution and direction of influence of each characteristic indicator on the yield prediction results of corn.

[0105] S430. Using the identified promoting or restricting factors as regulatory targets, the planting management strategy is adjusted to adjust the sowing date. The planting management strategy includes a static strategy that keeps the sowing date unchanged and an adaptive strategy that dynamically adjusts the sowing date based on historical trends.

[0106] S440. By changing the sowing date, environmental factors during the maize growth period are adjusted to drive the optimal machine learning model to predict the yield of maize under different environmental conditions in the future, and to evaluate the potential value of different planting management strategies to improve maize yield.

[0107] In one specific embodiment, the SHAP library is first used to perform in-depth analysis of the selected LightGBM model. The SHAP value for each feature across all predicted samples is calculated. For example... Figure 4 As shown, the results indicate that: (1) Cumulative total solar radiation is the primary positive factor affecting the yield, with the highest average absolute SHAP value, contributing 25.81% of the relative importance, indicating that photosynthetically active radiation is the energy basis for potential formation. (2) The length of the growing season and planting density are the second and third largest positive factors, contributing 16.74% and 15.41% respectively, reflecting the importance of making full use of the growing season and optimizing the population structure. (3) The diurnal temperature range is the fourth positive factor, contributing 11.38%, reflecting the importance of heat resources during maize growth.

[0108] Then, by selecting either a historical reconstruction model or a future prediction model, and using either promoting or restrictive factors as regulatory targets in the future prediction model, the sowing date is adjusted according to the planting management strategy:

[0109] Historical reconstruction model: The optimal maize yield prediction model (LightGBM) is deployed on a 0.5°×0.5° latitude and longitude grid. For each year from 2001 to 2014, 12 characteristic indicators are calculated for each grid with maize planting, and then input into the model to obtain the yield prediction value for that grid in that year.

[0110] Future prediction model: For 2015-2050, a multi-scenario analysis is implemented. (1) Climate scenario: such as using three paths: SSP1-2.6 (low carbon), SSP3-7.0 (medium to high carbon) and SSP5-8.5 (high carbon). (2) Management strategy: set two types of planting: static planting and adaptive planting. The latter dynamically adjusts the planting date of each year in the future based on the linear change trend of the phenological period of each grid from 2001 to 2020, so as to avoid the negative effects of global warming. (3) Combine the three climate scenarios with the two management strategies to form six future scenarios. For each scenario, calculate the future grid characteristics year by year and input them into the model to predict the yield. Figure 5 The study shows the changing trend of the national average yield. Among them, the SSP1-2.6+ adaptive planting combination showed the most positive growth potential, while the SSP3-7.0+ static planting combination showed the slowest growth, thus highlighting the importance of proactive adaptation measures.

[0111] S500: Compare the predicted yield of corn with the benchmark yield of the same period, evaluate the yield performance under different planting management strategies based on the comparison results, and select the optimal planting management plan for each environmental condition.

[0112] In this embodiment, step S500 may include the following sub-steps S510 to S530:

[0113] S510. Calculate the difference between the predicted yield of corn under different planting management strategies and the benchmark yield in the same period to obtain the yield change range under each strategy.

[0114] S520. Based on the magnitude of yield change and combined with the planting management parameters corresponding to each strategy, a yield performance evaluation matrix is ​​constructed. The yield performance evaluation matrix is ​​used to quantify the yield response characteristics of each planting management strategy under different combinations of environmental conditions.

[0115] S530. With the goal of maximizing yield, the planting management strategy with the best yield performance is selected as the recommended solution from the yield performance evaluation matrix for each target environmental condition.

[0116] In one specific embodiment, the gridded historical and future achievable maize yields obtained from the optimal maize yield prediction model are compared with the actual statistical yields of the same period to calculate the time series of exploitable yield gaps at the national scale. Figure 6 As shown, the evaluation results indicate that:

[0117] (1) Huge historical gap: Between 2001 and 2014, China’s actual corn yield averaged only 36.6% of the simulated yield, meaning that 63.4% of the yield potential was not tapped.

[0118] (2) The gap will remain in the future: Even by 2050, the gap between actual output and potential output is expected to remain in the range of 43.6% to 44.7%.

[0119] (3) Clear path to increase production: This assessment conclusion shows with data that by promoting the best existing varieties and optimizing agronomic management (such as reasonable dense planting and adaptive adjustment of sowing period), China’s total maize production has the scientific basis and practical feasibility to double in the future without significantly increasing arable land.

[0120] Secondly, refer to Figure 7 As shown, this embodiment also proposes a maize yield assessment system, which can adopt a B / S (browser / server) or C / S (client / server) architecture and be deployed on a high-performance server cluster. The system includes the following five core functional modules, which communicate through internal standardized data interfaces (such as JSON and NetCDF):

[0121] The data acquisition module 101 is used to acquire historical multi-source data related to corn yield. This module connects to external databases (such as meteorological databases and soil databases) and the file system, automatically performing all data preprocessing and sampling operations in step S100. This module provides visualization tools for monitoring data quality and coverage.

[0122] The feature processing module 102 is used to perform agronomic mechanism analysis on historical multi-source data and construct a feature index set containing varietal characteristic indicators, environmental condition indicators, and planting management characteristic indicators based on the analysis results. Users only need to specify the research area and time range, and the feature processing module can automatically call the corresponding basic data to generate gridded feature index datasets in batches.

[0123] The model training and selection module 103 is used to train multiple preset machine learning models using a set of feature indicators as input variables and corn yield data from historical multi-source data as the target variable. The module then selects the optimal machine learning model based on cross-validation evaluation results. This module integrates parallel computing versions of six algorithms, including Random Forest, XGBoost, and LightGBM. Users can import feature data and yield labels through the interface, and the module automatically completes data splitting, model training, hyperparameter tuning, cross-validation, and performance evaluation.

[0124] The yield prediction module 104 is used to analyze the optimal machine learning model using an interpretable artificial intelligence framework, and input the planting management strategy generated based on the analysis results into the optimal machine learning model to predict the yield of corn at the target spatiotemporal scale. After the model is trained, this module starts the SHAP analysis function to analyze the model, and then uses the analyzed optimal model to select either historical reconstruction or future prediction mode to predict the yield of corn at the target spatiotemporal scale.

[0125] The evaluation and scheme selection module 105 compares the predicted yield of maize with the benchmark yield for the same period, evaluates the yield performance under different planting management strategies based on the comparison results, and selects the optimal planting management scheme for each environmental condition. This module connects to the actual yield database, and automatically calculates the potential yield gap after specifying the comparison area and time period.

[0126] Thirdly, this embodiment also proposes an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to execute the corn yield evaluation method described above.

[0127] Fourthly, this embodiment also proposes a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the corn yield evaluation method described above.

[0128] In summary, this invention presents a method, system, equipment, and medium for evaluating maize yield. By integrating agronomic mechanism analysis and machine learning modeling, it constructs an interpretable set of feature indicators and uses cross-validation to select the optimal prediction model. Combined with the SHAP interpretability framework, it achieves quantitative analysis of yield-influencing factors and dynamic optimization of planting management strategies. This method not only improves the accuracy and robustness of maize yield potential prediction but also realizes a closed-loop evaluation chain from historical data mining to future multi-climate scenario prediction, providing a scientific basis for selecting planting management schemes and exploring yield potential at the regional scale. Compared with existing technologies, this invention has significant advantages such as strong model interpretability, adaptability to different spatiotemporal scales, and support for multi-objective strategy optimization.

[0129] Since the systems / devices described in the above embodiments of the present invention are systems / devices used to implement the methods of the above embodiments of the present invention, those skilled in the art can understand the specific structure and modifications of the systems / devices based on the methods described in the above embodiments of the present invention, and therefore will not be repeated here. All systems / devices used in the methods of the above embodiments of the present invention fall within the scope of protection of the present invention.

[0130] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0131] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, as well as combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions.

[0132] It should be noted that in the description of this invention, the word "a" or "an" preceding a component does not exclude the existence of multiple such components. This invention can be implemented by means of hardware comprising several different components and by means of a suitably programmed computer. The use of terms such as first, second, third, etc., is merely for convenience and does not indicate any order. These terms can be understood as part of the component names.

[0133] Furthermore, it should be noted that in the description of this specification, the terms "one embodiment," "some embodiments," "embodiment," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Furthermore, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0134] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning of the basic inventive concept, can make other changes and modifications to these embodiments.

[0135] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from the spirit and scope of the invention.

Claims

1. A method for evaluating the yield of maize, characterized in that, include: Historical multi-source data on corn yield can be obtained by acquiring related corn varieties; Agronomic mechanism analysis was conducted on historical multi-source data, and a set of characteristic indicators including varietal characteristic indicators, environmental condition indicators, and planting management characteristic indicators was constructed based on the analysis results. Using a set of feature indicators as input variables and corn yield data from historical multi-source data as target variables, we trained multiple preset machine learning models and selected the optimal machine learning model based on the cross-validation evaluation results. The optimal machine learning model is analyzed using an interpretable artificial intelligence framework, and the planting management strategy generated based on the analysis results is input into the optimal machine learning model to predict the yield of maize at the target spatiotemporal scale. The predicted yield of maize is compared with the baseline yield of the same period. The yield performance under different planting management strategies is evaluated based on the comparison results, and the optimal planting management plan for each environmental condition is selected.

2. The method as described in claim 1, characterized in that, Historical multi-source data on corn yield obtained from related sources includes: Field measurement data from multiple test sites in the target area were collected, and the collected data were preprocessed by data cleaning and normalization to obtain historical maize yield data for the target area. Extract relevant data from a pre-set database that matches the time window and geographical range of historical corn yield data. The relevant data includes meteorological data, soil data, corn growth period phenological data, and planting management data.

3. The method as described in claim 1, characterized in that, Agronomic mechanism analysis was conducted on historical multi-source data, and a feature index set was constructed based on the analysis results, including varietal characteristic indicators, environmental condition indicators, and planting management characteristic indicators. Agronomic mechanism analysis was conducted on historical multi-source data. Based on the analysis results, ecological factors related to maize yield were extracted from the historical multi-source data to determine the initial feature set, which includes varietal characteristics, environmental conditions, and planting management measures. The initial feature set is subjected to correlation analysis and importance ranking, redundant and low-contribution features are removed, and the feature index set is optimized and constructed. Among them, the varietal characteristic indicators include the length of the growing season and the effective accumulated temperature during the growing season, which reflect the heat demand characteristics of the variety; the environmental condition indicators include meteorological characteristic indicators and soil characteristic indicators; and the planting management characteristic indicators include planting density.

4. The method as described in claim 1, characterized in that, Using a set of feature indicators as input variables and corn yield data from historical multi-source data as the target variable, several pre-defined machine learning models were trained, and the optimal machine learning model was selected based on cross-validation evaluation results. These models included: The feature index set and the corn yield data from historical multi-source data are divided into training set and test set according to a preset ratio; The training set is used to train and cross-validate a variety of pre-defined machine learning models, including random forest, XGBoost, LightGBM, multilayer perceptron, support vector machine and partial least squares regression. The test set is input into each trained machine learning model for testing. The evaluation metrics for each machine learning model are obtained based on the test results. The evaluation metrics include the coefficient of determination and the normalized root mean square error. Based on the evaluation metrics of each machine learning model, a comprehensive ranking was conducted, and the machine learning model with the highest comprehensive ranking was selected as the optimal model for predicting the yield of corn.

5. The method as described in claim 1, characterized in that, The optimal machine learning model is analyzed using an interpretable artificial intelligence framework, and the planting management strategy generated based on the analysis results is input into the optimal machine learning model to predict the achievable maize yield at the target spatiotemporal scale, including: The optimal machine learning model is analyzed using an interpretable artificial intelligence framework to obtain the SHAP values ​​of each feature index in the feature index set for all predicted samples. Based on the sign and absolute value of the SHAP value, each characteristic indicator is identified as a promoting factor or a limiting factor, and the contribution and direction of influence of each characteristic indicator on the yield prediction results of maize are quantified. The identified promoting or restricting factors are used as regulatory targets to adjust the sowing date in the planting management strategy. The planting management strategy includes a static strategy that keeps the sowing date unchanged and an adaptive strategy that dynamically adjusts the sowing date based on historical trends. By adjusting the planting date to modify environmental factors during the maize growth period, the optimal machine learning model is driven to predict the yield of maize under different environmental conditions in the future, and to evaluate the potential value of different planting management strategies to improve maize yield.

6. The method as described in claim 1, characterized in that, The predicted yield of maize is compared with the baseline yield for the same period. Based on the comparison results, the yield performance under different planting management strategies is evaluated, and the optimal planting management plan for various environmental conditions is selected, including: The difference between the predicted maize yield under different planting management strategies and the benchmark yield in the same period is calculated to obtain the yield change under each strategy. Based on the magnitude of yield change and combined with the planting management parameters corresponding to each strategy, a yield performance evaluation matrix is ​​constructed. The yield performance evaluation matrix is ​​used to quantify the yield response characteristics of each planting management strategy under different combinations of environmental conditions. With the goal of maximizing yield, the planting management strategy with the best yield performance is selected as the recommended solution from the yield performance evaluation matrix for each target environmental condition.

7. A system for evaluating the yield of maize, characterized in that, include: The data acquisition module is used to acquire historical multi-source data on the yield of related corn. The feature processing module is used to perform agronomic mechanism analysis on historical multi-source data and construct a feature index set that includes variety characteristic indicators, environmental condition indicators, and planting management characteristic indicators based on the analysis results. The model training and selection module is used to train multiple preset machine learning models with feature index set as input variable and corn yield data from historical multi-source data as target variable, and select the optimal machine learning model based on cross-validation evaluation results. The yield prediction module is used to analyze the optimal machine learning model using an interpretable artificial intelligence framework, and input the planting management strategy generated based on the analysis results into the optimal machine learning model to predict the yield of corn at the target spatiotemporal scale. The evaluation and scheme selection module is used to compare the predicted corn yield with the benchmark yield of the same period, evaluate the yield performance under different planting management strategies based on the comparison results, and select the optimal planting management scheme for each environmental condition.

8. An electronic device, characterized in that, include: At least one processor; and memory that is communicatively connected to at least one processor; The memory stores instructions that can be executed by at least one processor, which enables the at least one processor to perform a method for evaluating the yield of corn as described in any one of claims 1-6.

9. A computer-readable storage medium storing computer-executable instructions thereon, characterized in that, When executed by a processor, the computer-executable instructions implement a method for evaluating the yield of corn as described in any one of claims 1-6.