Grain yield prediction model determination method, application method and related system

By integrating multi-source data and combining multi-algorithm model training with dynamic weighted integration methods, a grain yield prediction model is constructed, which solves the problems of difficulty in capturing multi-factor interactions and data integration limitations in existing technologies, and achieves high-precision grain yield prediction and scientific agricultural policy support.

CN120671910APending Publication Date: 2025-09-19INSTITUTE OF ENVIRONMENT AND SUSTAINABLE DEVELOPMENT IN AGRICULTURE CAAS
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510768204.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing technologies for predicting China's future grain production have problems such as a single model being unable to capture the complex interactions between multiple factors, insufficient adaptability of traditional models, and limitations in data integration, resulting in the need to improve the accuracy and reliability of prediction results.

Method used

By acquiring spatiotemporal dynamic data on key influencing indicators, integrating data from multiple sources, including climate, soil, management, and economics, and combining multi-algorithm model training with a dynamic weighted integration approach, a grain yield forecasting model was constructed. This model improves forecast accuracy and adaptability through multi-dimensional indicator screening, multi-algorithm training, and dynamic weighted integration.

Benefits of technology

It has achieved high-precision prediction of grain production, which can provide a scientific basis for grain production under different climatic conditions and management strategies, and help formulate more accurate agricultural policies and management measures, thereby increasing grain production capacity and ensuring national food security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120671910A_ABST
    Figure CN120671910A_ABST
Patent Text Reader

Abstract

The invention discloses a determination method, an application method and a related system of a grain yield prediction model, and relates to the field of grain yield prediction, and the determination method comprises the steps: obtaining time-space dynamic data; preprocessing and normalizing climate data, management data and soil data in the spatial-temporal dynamic data to obtain a spatial-temporal data set; respectively inputting the spatio-temporal data set into a plurality of machine learning models to obtain outputs of the plurality of machine learning models; respectively calculating a root-mean-square error and a decision coefficient of each machine learning model based on grain annual output data corresponding to the spatio-temporal data set and the output of each machine learning model; abandoning the machine learning model with the determination coefficient smaller than a preset value to obtain a screened machine learning model; and based on the screened machine learning model and the corresponding root-mean-square error, constructing a weighted average grain yield prediction model. According to the invention, a scientific basis can be provided for grain production under different climate scenes and management strategies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of regional future grain production forecasting, and in particular to a method for determining a grain production forecasting model, an application method, and a related system. Background Art

[0002] In recent years, with the ever-changing global food market and the increasing uncertainty of climate change, it has become imperative and urgent for China to continuously enhance its food production capacity and ensure food security. The importance of exploring the impact of climate change and management on food production has become increasingly apparent. Currently, most forecasting techniques for China's future food production focus on climate, soil, management, and economic factors, using single analyses such as crop models, ecological models, economic models, or machine learning models to analyze their impact on food production. However, these approaches have limitations: first, a single model often struggles to fully capture the complex interactions between multiple factors; second, traditional models have limited adaptability to dynamic changes and are unable to effectively address the uncertainties associated with climate fluctuations and adjustments to management strategies; furthermore, existing technologies remain insufficient in multidimensional data integration, model optimization, and scenario simulation, resulting in a need to improve the accuracy and reliability of forecast results. Summary of the Invention

[0003] The purpose of this application is to provide a method for determining, applying, and related systems for a grain yield prediction model that can integrate multi-dimensional indicators, fully combine multi-source data such as climate, soil, management, and economy, and combine multi-algorithm model training and prediction to provide a scientific basis for grain production under different climate scenarios and management strategies.

[0004] To achieve the above objectives, this application provides the following solutions:

[0005] In a first aspect, the present application provides a method for determining a grain yield prediction model, the method comprising:

[0006] Obtain spatiotemporal dynamic data of key influencing indicators; the spatiotemporal dynamic data include: annual grain production data, climate data, management data and soil data.

[0007] The climate data, the management data and the soil data are preprocessed to obtain preprocessed data.

[0008] The preprocessed data is normalized to obtain a spatiotemporal data set.

[0009] The spatiotemporal data sets are input into several machine learning models respectively to obtain outputs of several machine learning models.

[0010] Based on the annual grain production data corresponding to the spatiotemporal data set and the output of each machine learning model, the root mean square error and determination coefficient of each machine learning model are calculated respectively.

[0011] The machine learning models with a determination coefficient smaller than a preset value are discarded to obtain the screened machine learning models.

[0012] Based on the screened machine learning models and the corresponding root mean square errors, a weighted average grain yield prediction model is constructed.

[0013] In a second aspect, the present application provides an application method of a grain yield prediction model, the application method of the grain yield prediction model comprising:

[0014] Obtain future scenario data; the future scenario data include: climate scenario data, soil scenario data, and management scenario data; the climate scenario data are selected from the Global Climate Models (GCMs) and Shared Socioeconomic Pathways (SSPs) emission scenarios in the Sixth Coupled Model Intercomparison Project Phase 6 (CMIP6); the soil scenario data are output from ecosystem process models based on simulations of different climate scenarios; and the management scenario data are data set based on tillage scenarios, irrigation scenarios, fertilization scenarios, relevant policy recommendations, and historical trends;

[0015] The future scenario data is input into a grain production prediction model to obtain future grain production; the grain production prediction model is a model obtained based on the determination method of the grain production prediction model.

[0016] In a third aspect, the present application provides a computer device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described method for determining the grain yield prediction model or the above-described method for applying the grain yield prediction model.

[0017] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for determining the grain yield prediction model described above or the method for applying the grain yield prediction model described above.

[0018] In a fifth aspect, the present application provides a computer program product, comprising a computer program, which, when executed by a processor, implements the method for determining the grain yield prediction model described above or the method for applying the grain yield prediction model described above.

[0019] According to the specific embodiments provided in this application, this application discloses the following technical effects:

[0020] The present application provides a method for determining a grain yield prediction model, an application method, and a related system. By obtaining spatiotemporal dynamic data of key influencing indicators, it can cover natural (climate, soil) and human (management) factors, avoid prediction bias caused by a single indicator, and provide comprehensive and real input features for subsequent modeling to ensure the reliability of the prediction results. By preprocessing the climate data, the management data, and the soil data, preprocessed data are obtained; normalizing the preprocessed data to obtain a spatiotemporal data set, the reliability of the data can be further improved, and at the same time, the format and standard are unified to ensure that data from different sources can be fused and analyzed. By inputting the spatiotemporal data set into several machine learning models respectively, the outputs of several machine learning models are obtained; based on the annual grain yield data corresponding to the spatiotemporal data set and the output of each machine learning model, the root mean square error and determination coefficient of each machine learning model are calculated respectively; the machine learning model with a determination coefficient less than a preset value is discarded to obtain a screened machine learning model; based on the screened machine learning model and the corresponding root mean square error, a weighted average grain yield prediction model is constructed. This application integrates multi-dimensional indicators, fully utilizing data from multiple sources, including climate, soil, management, and economics. Combining multi-algorithm model training with a dynamic weighted ensemble approach, it achieves high-precision forecasts of grain yields. Furthermore, through multi-scenario simulation and forecasting, it can provide a scientific basis for grain production under different climate conditions and management strategies, assisting in the formulation of more precise agricultural policies and management measures, thereby increasing grain production capacity and ensuring national food security. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0022] Figure 1 A flow chart of a method for dynamic prediction of grain yield based on multi-source data fusion and machine learning provided in one embodiment of the present application.

[0023] Figure 2This is an application environment diagram of a method for determining a grain yield prediction model in one embodiment of the present application.

[0024] Figure 3 A flowchart of a method for determining a grain yield prediction model provided in one embodiment of the present application.

[0025] Figure 4 A flowchart of an application method of a grain yield prediction model provided in one embodiment of the present application.

[0026] Figure 5 A schematic diagram of corn production in the three northeastern provinces from 2022 to 2050 is provided for one embodiment of the present application.

[0027] Figure 6 A schematic diagram of the structure of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0028] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0029] like Figure 1As shown, the present application aims to provide a method and system for simulating and predicting future grain production based on the integration of multi-dimensional indicators, multi-source data fusion and multi-type machine learning models. The method includes: using Pearson's Product-Moment Correlation Coefficient Analysis (Pearson's r), Spearman's Rank Correlation Coefficient Analysis (Spearman's rho), Grey Relational Analysis (GRA) and Mutual Information Analysis (MIA) to screen the core indicators affecting grain production, quantify the correlation coefficient and dynamic contribution of each indicator to grain production; collect grain crop yield, climate, management and soil data, and perform preprocessing, including spatial scale unification, outlier removal and interpolation to fill missing data to ensure the integrity of the data set; divide the data set into 80% training set and 20% validation set, input the back propagation neural network (BP), convolutional neural network (CNN), long short-term memory network (Long Short-Term Memory (LSTM), Random Forest (RF), eXtreme Gradient Boosting (XGBoost), Support Vector Machine (SVM), Histogram-based Gradient Boosting (HGB) and Gradient Boosting Decision Tree (GBDT) are machine learning models, using the root mean square error (RMSE) and the coefficient of determination (R 2 ) Evaluate model performance; Use dynamic weighting methods to build an integrated model for outstanding machine learning models to improve model robustness and prediction accuracy; Set future climate and management scenarios, input future food production prediction models, and simulate future food production trends.

[0030] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0031] The method for determining the grain yield prediction model provided in the embodiment of the present application can be applied to Figure 2In the application environment shown, the terminal 102 communicates with the server 104 via a network. The data storage system can store data that the server 104 needs to process. The data storage system can be set up separately, integrated on the server 104, or placed on the cloud or other servers. Terminal 102 may send spatiotemporal dynamic data of key influencing indicators to server 104, the spatiotemporal dynamic data including annual grain production data, climate data, management data, and soil data. After receiving the spatiotemporal dynamic data of key influencing indicators, server 104 may preprocess the climate data, management data, and soil data to obtain preprocessed data. The preprocessed data may be normalized to obtain a spatiotemporal data set. The spatiotemporal data set may be input into a plurality of machine learning models to obtain outputs of the plurality of machine learning models. Based on the annual grain production data corresponding to the spatiotemporal data set and the outputs of the machine learning models, the root mean square error (RMSE) and coefficient of determination (CDR) of each machine learning model may be calculated. Machine learning models with CDRs less than a preset value may be discarded to obtain selected machine learning models. A weighted average grain production prediction model may be constructed based on the selected machine learning models and their corresponding RMS errors. Server 104 may provide feedback to terminal 102 on the obtained grain production prediction model. In addition, in some embodiments, the method for determining the grain yield prediction model can also be implemented independently by the server 104 or the terminal 102. For example, the terminal 102 can directly determine the grain yield prediction model based on the spatiotemporal dynamic data of key influencing indicators, or the server 104 can obtain the spatiotemporal dynamic data of key influencing indicators from the data storage system and determine the grain yield prediction model based on the spatiotemporal dynamic data of key influencing indicators.

[0032] The terminal 102 may be, but is not limited to, various desktop computers, laptop computers, smart phones, and tablet computers. The server 104 may be implemented as an independent server or a server cluster consisting of multiple servers, or a cloud server.

[0033] In an exemplary embodiment, Figure 3 As shown, a method for determining a grain yield prediction model is provided. The method is executed by a computer device, specifically a computer device such as a terminal or a server, or a terminal and a server. In the embodiment of the present application, the method is applied to Figure 2 The server 104 in the example is used for explanation, and the steps include the following steps S1 to S7.

[0034] S1: Obtain spatiotemporal dynamic data of key influencing indicators; the spatiotemporal dynamic data include: annual grain production data, climate data, management data and soil data.

[0035] S2: Preprocessing the climate data, the management data, and the soil data to obtain preprocessed data.

[0036] S3: Normalize the preprocessed data to obtain a spatiotemporal data set.

[0037] S4: Input the spatiotemporal data set into several machine learning models respectively to obtain outputs of several machine learning models.

[0038] S5: Based on the annual grain production data corresponding to the spatiotemporal data set and the output of each machine learning model, the root mean square error and determination coefficient of each machine learning model are calculated respectively.

[0039] S6: Discard the machine learning models whose determination coefficient is less than the preset value to obtain the screened machine learning models.

[0040] S7: Based on the screened machine learning model and the corresponding root mean square error, a weighted average grain yield prediction model is constructed.

[0041] Implementing steps S1 through S7 above can address existing issues such as the difficulty of single models in capturing multi-factor relationships, the lack of adaptability of traditional models, and data integration limitations. By screening multi-dimensional indicators, integrating multi-source data, and training multiple algorithms, a dynamic weighted integrated model can be constructed to improve forecast accuracy and adaptability. Combined with climate and management scenario simulations, this model provides a scientific basis for food production and enhances the reliability of agricultural decision-making.

[0042] As an optional implementation, before step S1, the method for determining the grain yield prediction model further includes:

[0043] A1: Based on literature analysis and analysis of influencing factors, determine preliminary indicators that affect grain production.

[0044] A2: The Pearson product-moment correlation coefficient analysis method, the Spearman rank correlation coefficient analysis method, the grey correlation analysis method and the mutual information analysis method were used to conduct correlation analysis on the preliminary indicators and obtain several analysis results.

[0045] A3: Based on the above analysis results, determine the correlation coefficient and dynamic contribution of each preliminary indicator to grain production.

[0046] A4: Based on the correlation coefficient and dynamic contribution, key influencing indicators are determined; the key influencing indicators include: climate indicators, management indicators and soil indicators.

[0047] Specifically, first, based on literature analysis and influencing factor analysis, targeting the specific research background and purpose, we preliminarily selected indicators that have an impact on grain production. Then, we conducted correlation analysis on the indicators that affect grain production through analysis methods such as Pearson product-moment correlation coefficient analysis, Spearman rank correlation coefficient analysis, grey correlation analysis, and mutual information analysis. We comprehensively analyzed and quantified the correlation coefficient and dynamic contribution of each indicator with grain production, and screened out a set of influencing indicators with significant correlation.

[0048] As an optional implementation, in step S1, it specifically includes:

[0049] Obtain spatiotemporal dynamic data on influencing indicators, including annual grain production data such as corn production, wheat production, and rice production; climate data such as annual average temperature and annual average precipitation; management data such as total power of agricultural machinery, pure fertilizer application amount, pesticide use, effective irrigation area, mulch coverage area, and drought- and flood-resistant area; and soil data such as soil organic matter and soil pH.

[0050] As an optional implementation, in step S2, it specifically includes:

[0051] S21: using Anuspline interpolation software to interpolate the climate data to the spatial data to obtain interpolated climate data.

[0052] S22: Using Arcgis software, upscaling or downscaling the soil data and the interpolated climate data to obtain upscaled or downscaled data.

[0053] S23: performing interpolation and filling processing on the management data and the data after the scaling processing to obtain spatial grid data with unified spatiotemporal resolution and complete time series for all data.

[0054] Specifically, the climate data, management data, and soil data are preprocessed, including interpolating the climate data from the site data to the spatial data using the professional interpolation software Anuspline; to address the problem of spatial resolution mismatch between the soil data and the interpolated climate data, the original data are upscaled or downscaled using Arcgis software; and to address the problem of missing values ​​in time or space in the management data and the data after upscaling and downscaling, interpolation and filling are performed on them to obtain spatial grid data with unified spatiotemporal resolution and complete time series for all indicators.

[0055] Normalization was performed on the preprocessed climate data, preprocessed management data, and preprocessed soil data. Taking into account the distribution characteristics, dimensional differences, and model requirements of different data sources, the preprocessed climate data was normalized using the Z-Score method.

[0056]

[0057] Among them, x1′ is the normalized climate data; x1 is the preprocessed climate data; μ is the mean of the preprocessed climate data; σ is the standard deviation of the preprocessed climate data.

[0058] The preprocessed soil data were normalized using the Min-Max normalization method.

[0059]

[0060] Among them, x2 is the preprocessed soil data; x max is the maximum value of the preprocessed soil data; x min is the minimum value of the preprocessed soil data.

[0061] The preprocessed management data were normalized using the logarithmic transformation + robust scaling method.

[0062]

[0063] Among them, x3′ is the normalized management data, x3 is the preprocessed management data, and x max ′ is the maximum value of the management data after preprocessing.

[0064] Then conduct multi-algorithm model training and performance evaluation.

[0065] The processed climate, management, and soil spatiotemporal data sets are combined as input data for multi-algorithm model training, and the yields of different types of crops are used as feature variables to train different types of machine learning models. Machine learning models include BP, CNN, LSTM, RF, XGBOOST, SVM, HGB, GBDT and other machine learning and deep learning models, allowing the machine learning models to learn and simulate the intrinsic relationship between grain yield and influencing indicators. Before training different types of machine learning models, the historical data is divided into 80% training sets and 20% test sets. After training, the root mean square error (RMSE) and determination coefficient (R) of each machine learning model are calculated. 2 ), which is used to evaluate the performance of different models in making predictions.

[0066] The calculation formula of the root mean square error is:

[0067]

[0068] Where RMSE is the root mean square error; n is the sample size (the number of measured values ​​in the test set); y i is the measured value (measured annual grain production data); is the predicted value (predicted annual grain production data).

[0069] The calculation formula of the determination coefficient is:

[0070]

[0071] Among them, R 2 is the coefficient of determination; It is the average of the measured values ​​(the average of the measured annual grain production data).

[0072] Finally, dynamic weighted ensemble learning is used to establish a grain yield prediction model.

[0073] By setting model screening criteria, the machine learning models with the highest coefficient of determination are selected, while the remaining machine learning models are discarded, thus selecting the best performing models. By calculating the weight 1 / RMSE of these selected models, a weighted average ensemble model is constructed, which becomes the grain yield prediction model.

[0074] like Figure 4 As shown, a method for applying a grain yield prediction model is provided, and the method for applying the grain yield prediction model includes:

[0075] B1: Obtain future scenario data; the future scenario data include: climate scenario data, soil scenario data, and management scenario data. The climate scenario data are selected from the global climate model and shared economy pathway emission scenarios in the Sixth Coupled Model Intercomparison Project. The soil scenario data are output from ecosystem process models based on simulations of different climate scenarios. The management scenario data are based on tillage scenarios, irrigation scenarios, fertilization scenarios, relevant policy recommendations, and historical trends.

[0076] B2: Inputting the future scenario data into a grain production prediction model to obtain future grain production; the grain production prediction model is a model obtained based on the above-mentioned method for determining the grain production prediction model.

[0077] Specifically, by setting different future scenarios, including climate, soil, and management scenarios, the data calculated based on these scenarios were fed into a previously constructed grain yield prediction model for rapid simulation to determine future grain yields. Climate scenarios refer to future climate change scenarios, selected from the GCMs and SSPs in CMIP6; management scenarios refer to tillage, irrigation, and fertilization scenarios, as well as corresponding scenarios set based on relevant policy recommendations and historical trends; and soil scenarios are derived from simulation outputs of ecosystem process models based on different climate scenarios.

[0078] This application also provides an application scenario, which applies the above-mentioned application method of the grain yield prediction model. Specifically: The application method of the grain yield prediction model provided in this embodiment can be applied in the grain yield prediction scenario. The food production forecasting scenario includes: a food production forecasting model determination link, a future scenario data acquisition link and a food production forecasting link; first, obtaining the spatiotemporal dynamic data of key influencing indicators; the spatiotemporal dynamic data includes: annual food production data, climate data, management data and soil data; second, preprocessing the climate data, the management data and the soil data to obtain preprocessed data; normalizing the preprocessed data to obtain a spatiotemporal data set; then, inputting the spatiotemporal data set into several machine learning models respectively to obtain the outputs of several machine learning models; based on the annual food production data corresponding to the spatiotemporal data set and the outputs of each machine learning model, calculating the root mean square error and determination coefficient of each machine learning model respectively; discarding the machine learning model with a determination coefficient less than a preset value to obtain a screened machine learning model; constructing a weighted average food production forecasting model based on the screened machine learning model and the corresponding root mean square error; finally, obtaining future scenario data; inputting the future scenario data into the food production forecasting model to obtain future food production.

[0079] The present application is described below with reference to specific embodiments.

[0080] The research content of this embodiment is the prediction of corn production in the three northeastern provinces by 2035 and 2050. Figure 5 As shown, the specific steps include:

[0081] 1. The study covers the three provinces of Northeast China (Liaoning, Jilin, and Heilongjiang) from 2017 to 2021. Corn yield, climate, management, and soil data are collected at the county level. Climate data include annual mean temperature and annual mean precipitation. Management data include total agricultural machinery power, fertilizer application, pesticide use, effective irrigation area, mulch film coverage, agricultural water use, improved seed coverage, and the proportion of organic fertilizer application area. Soil data include soil organic matter content. To quantify the impact of influencing indicators on corn yield, these data require preprocessing. This typically involves upscaling historical climate and soil data, translating high-resolution grid data into county-level units, and interpolating missing values ​​due to the large data span.

[0082] 2. The integrated data set is used as input data for different types of machine learning models. In this embodiment, six machine learning and deep learning models, including BP, CNN, RF, XGBOOST, HGB, and Ridge, are used to allow the machine learning model to learn and simulate the intrinsic relationship between corn yield and influencing indicators. Before model training, the data set is divided into 80% training set and 20% validation set. After training, the root mean square error (RMSE) and determination coefficient (R) of each machine learning model are calculated. 2 ), which is used to evaluate the performance of different models in making predictions.

[0083] 3. Screen different machine learning models by setting screening criteria. 2 The ranking was: RF (0.9867) > XGBOOST (0.9851) > HGB (0.9848) > BP (0.9823) > CNN (0.9776) > Ridge (0.9715). RF, XGBOOST, and HGB models performed well and were retained, while the others were discarded. Parameter optimization was used to calculate weights, with RF, XGBOOST, and HGB weights set to 0.2, 0.7, and 0.1, respectively. Finally, through weighted integration, a model for predicting future corn yields in the three northeastern provinces was derived.

[0084] 4. By setting different future scenarios, including climate scenarios, management scenarios, and soil scenarios. Here, climate scenarios and soil scenarios refer to future scenarios of average annual temperature, average annual precipitation, and soil organic matter content. The above future scenario data uses experimental data of future scenarios composed of different shared socioeconomic pathways (SSPs) and representative concentration pathways (RCPs): SSP1-2.6, SSP2-4.5, SSP3-7.0, and SSP5-8.5. Among them, future climate data are simulated by GCMs in CMIP6, including Access-CM2, Access-ESM1-5, CanESM5, CMCC-ESM2, GFDL-ESM4, INM-CM4-8, INM-CM5-0, MIROC, MPI-ESM1-2-LR, MRI-ESM2-0, and TaiESM1; future soil data are simulated by the CEVSA2 model. Management scenarios refer to tillage scenarios, irrigation scenarios, fertilization scenarios, and corresponding scenarios set based on relevant policy recommendations and historical trends, including the use of fertilizers based on 2021 as the base value, which will decrease at an average annual growth rate of -1.3546% by 2035 and 2050; the use of pesticides based on 2021 as the base value, which will decrease at an average annual growth rate of -0.9444% by 2035 and 2050; and the effective irrigation area based on 2021 as the base value, which will increase at an average annual growth rate of 1.1728% by 2035 and 2050. Based on the future scenarios set above, the future scenario data are calculated as the model input data, and a rapid simulation is performed to obtain the corn production data for China's three northeastern provinces (Liaoning, Jilin, and Heilongjiang) by 2035 and 2050. Through result analysis ( Figure 5 ), under the four future climate scenarios, the average corn production in Liaoning, Jilin and Heilongjiang provinces in 2035 will be 18.3294 million tons, 33.1300 million tons and 48.4730 million tons respectively; and the average corn production in Liaoning, Jilin and Heilongjiang provinces in 2050 will be 19.0649 million tons, 29.2188 million tons and 45.9640 million tons respectively.

[0085] It can be seen that the method of this application can predict future regional grain production changes more quickly and efficiently.

[0086] In summary, this application integrates multi-dimensional indicators, fully utilizing data from multiple sources, including climate, soil, management, and economics, and combining multi-algorithm model training with a dynamic weighted ensemble approach to achieve high-precision forecasts of grain yields. Furthermore, through multi-scenario simulation and forecasting, it can provide a scientific basis for grain production under different climate conditions and management strategies, assisting in the formulation of more precise agricultural policies and management measures, thereby increasing grain production capacity and ensuring national food security.

[0087] Specifically, this application can screen out key indicators through multi-dimensional indicator screening to avoid single indicator bias and overfitting caused by irrelevant indicators; multi-algorithm model training can capture data characteristics from different angles, reduce the dependence of a single algorithm on data, enhance model robustness, and better handle nonlinear relationships; dynamic weighted integration can reduce the impact of single model fluctuations on the overall situation, and improve the stability and reliability of predictions by dynamically adjusting weights; multi-scenario simulation and prediction can reduce uncertainty and provide more comprehensive predictions by considering multiple possible climate scenarios and soil scenarios.

[0088] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal. The internal structure diagram thereof may be as follows: Figure 6 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store spatiotemporal dynamic data or future scenario data of key influencing indicators. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a method for determining a grain yield prediction model or an application method of a grain yield prediction model.

[0089] Those skilled in the art will understand that Figure 6 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0090] In an exemplary embodiment, a computer device is further provided, including a memory and a processor. The memory stores a computer program, and the processor implements the above method embodiments when executing the computer program.

[0091] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, which implements the above-mentioned method embodiments when executed by a processor.

[0092] In an exemplary embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the above method embodiments are implemented.

[0093] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0094] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0095] The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may include, but are not limited to, general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic units, data processing logic units based on quantum computing, and the like.

[0096] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0097] This document uses specific examples to illustrate the principles and implementation methods of this application. The description of the above examples is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.

Claims

1. A method for determining a grain yield prediction model, characterized in that: The method for determining the grain yield prediction model includes: Obtaining spatiotemporal dynamic data of key influencing indicators; the spatiotemporal dynamic data includes: annual grain production data, climate data, management data, and soil data; Preprocessing the climate data, the management data, and the soil data to obtain preprocessed data; Normalizing the preprocessed data to obtain a spatiotemporal data set; Inputting the spatiotemporal data sets into a plurality of machine learning models respectively to obtain outputs of the plurality of machine learning models; Based on the annual grain production data corresponding to the spatiotemporal data set and the output of each machine learning model, respectively calculating the root mean square error and the coefficient of determination of each machine learning model; The machine learning models with a coefficient of determination smaller than a preset value are discarded to obtain the selected machine learning models; Based on the screened machine learning models and the corresponding root mean square errors, a weighted average grain yield prediction model is constructed.

2. The method for determining a grain yield prediction model according to claim 1, wherein: Before the step of obtaining the spatiotemporal dynamic data of key influencing indicators, the method for determining the grain yield prediction model further includes: Based on literature analysis and analysis of influencing factors, preliminary indicators that affect grain production were identified; Pearson product-moment correlation coefficient analysis, Spearman rank correlation coefficient analysis, grey relational analysis and mutual information analysis were used to conduct correlation analysis on the preliminary indicators, and several analysis results were obtained; Based on the above analysis results, the correlation coefficient and dynamic contribution of each preliminary indicator to grain output are determined; Based on the correlation coefficient and the dynamic contribution, key influencing indicators are determined; the key influencing indicators include: climate indicators, management indicators and soil indicators.

3. The method for determining a grain yield prediction model according to claim 1, wherein: Preprocessing the climate data, the management data, and the soil data to obtain preprocessed data specifically includes: interpolating the climate data to spatial data using Anuspline interpolation software to obtain interpolated climate data; Using Arcgis software to upscale or downscale the soil data and the interpolated climate data to obtain upscaled or downscaled data; Interpolation and filling processing is performed on the management data and the data after the scaling processing to obtain spatial grid data with unified spatiotemporal resolution and complete time series for all data.

4. The method for determining a grain yield prediction model according to claim 1, wherein: The expression for normalizing the preprocessed climate data is: The expression for normalizing the preprocessed soil data is: The expression for normalizing the preprocessed management data is: Among them, x1′ is the normalized climate data; x1 is the preprocessed climate data; μ is the mean of the preprocessed climate data; σ is the standard deviation of the preprocessed climate data; x2′ is the normalized soil data; x2 is the preprocessed soil data; x max is the maximum value of the preprocessed soil data; x min is the minimum value of the pre-processed soil data; x3′ is the normalized management data, x3 is the pre-processed management data, x max ′ is the maximum value of the management data after preprocessing.

5. The method for determining a grain yield prediction model according to claim 1, wherein: The calculation formula of the root mean square error is: Where RMSE is the root mean square error; n is the sample size; y i is the measured value; is the predicted value.

6. The method for determining a grain yield prediction model according to claim 1, wherein: The calculation formula of the determination coefficient is: Among them, R 2 is the coefficient of determination; n is the sample size; y i is the measured value; is the predicted value; is the average of the measured values.

7. An application method of a grain yield prediction model, characterized in that: The application method of the grain yield prediction model includes: Obtain future scenario data; the future scenario data includes: climate scenario data, soil scenario data, and management scenario data; the climate scenario data is selected from the global climate model and shared economy pathway emission scenarios in the Sixth Coupled Model Intercomparison Project; the soil scenario data is data output by ecosystem process models based on simulations of different climate scenarios; and the management scenario data is data set based on tillage scenarios, irrigation scenarios, fertilization scenarios, relevant policy recommendations, and historical trends; The future scenario data is input into a grain yield prediction model to obtain future grain yield; the grain yield prediction model is a model obtained based on the method for determining the grain yield prediction model according to any one of claims 1 to 6.

8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the method for determining the grain yield prediction model described in any one of claims 1 to 6 or the method for applying the grain yield prediction model described in claim 7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the method for determining the grain yield prediction model described in any one of claims 1 to 6 or the method for applying the grain yield prediction model described in claim 7.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, it implements the method for determining the grain yield prediction model described in any one of claims 1 to 6 or the method for applying the grain yield prediction model described in claim 7.

Citation Information

Cited By

  • Climate mode optimization and set estimation method based on interpretable and machine learning

    CN121350498A

  • Explainable and machine learning based climate model preference and ensemble prediction method

    CN121350498B

  • Crop yield prediction method based on multi-source data space-time fusion

    CN121615821A