Crop planting layout optimization method based on crop mechanism model, machine learning and life cycle assessment
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-07
- Publication Date
- 2026-08-11
AI Technical Summary
然而,现有作物布局优化技术普遍存在三方面不足:其一,对产量稳定性的系统评估不足,现有研究多采用均值导向的评估方法,侧重于单一年份或多年平均气候条件下的产量模拟与提升,难以捕捉作物产量在不利气候年份的波动特征,缺乏对高产与稳产双重目标的协同优化能力;其二,次适宜种植区的识别与优化不足,当前对次适宜种植区的划分多基于单一作物的平均产量水平,未能将产量稳定性作为核心判别依据,导致部分产量虽不高但波动较小的区域被忽视,而产量虽较高但极易受气候波动影响的区域却未被识别为优化对象,限制了布局调整的精准性;其三,环境成本考量缺失,在碳中和目标日益强化的背景下,现有技术未能将作物产量的高产稳产指标与不同作物的碳排放量有效耦合,也未在次适宜区种植布局优化中系统纳入环境成本考量,难以实现产量提升与碳减排的协同调控
[0021]基于上述技术方案可知,本申请的基于作物机理模型、机器学习与生命周期评价的作物种植布局优化方法,相对于现有技术,至少具备如下有益效果之一:
Smart Images

Figure CN122549671A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of crop planting layout technology, and in particular relates to a crop planting layout optimization method based on crop mechanism models, machine learning and life cycle assessment. Background Technology
[0002] Optimizing crop planting layout is a crucial technical means to ensure regional food security and address climate change. Against the backdrop of climate change, many traditional agricultural production areas face the challenge of declining crop production stability. Rising temperatures alter crop growth and development rhythms, increasing the frequency and intensity of extreme weather events such as heat waves and droughts, threatening crop yield stability. Among widely distributed agricultural planting areas, there are numerous suboptimal planting areas with relatively low crop yields and significant interannual fluctuations. Existing research indicates that climate variability is the main factor contributing to production instability in these regions, and the main constraints to achieving high and stable yields stem from unfavorable meteorological conditions, such as insufficient or excessive accumulated temperature, uneven precipitation distribution, and extreme temperature events. For some regions, temperature stress during the critical reproductive growth stage has a particularly pronounced impact on yield; while in some arid and semi-arid regions, crop production is primarily constrained by insufficient water supply. In recent decades, the increased frequency and intensity of heat waves have further exacerbated these risks, and the combined effects of heat stress and drought can lead to significant yield losses. Therefore, mitigating the vulnerability of crops to climate disasters during their reproductive growth stage is crucial for enhancing yield stability and ensuring regional food security.
[0003] To mitigate climate risks, various adaptation strategies have been employed in existing technologies, including adjusting sowing dates and planting densities, selecting climate-adaptive varieties, and optimizing water and nitrogen management. While these measures have proven effective in core production areas, for less suitable growing regions, management adjustments alone cannot fundamentally address the inherent constraints of water and heat resources. Furthermore, such management adjustments must consider socio-economic constraints and weigh economic benefits against environmental costs. Therefore, replanting ecologically more adaptable alternative crops in previously less suitable areas, thereby adjusting crop planting patterns, becomes a feasible path to mitigate climate risks and enhance the climate resilience of regional agricultural systems.
[0004] Currently, crop layout optimization technology mainly follows two paths: one is statistical analysis, and the other is process-based crop simulation. However, existing crop layout optimization technologies generally suffer from three shortcomings: First, there is insufficient systematic assessment of yield stability. Existing studies mostly adopt mean-oriented assessment methods, focusing on the simulation and improvement of yield under single-year or multi-year average climatic conditions. This makes it difficult to capture the fluctuation characteristics of crop yield in unfavorable climatic years and lacks the ability to synergistically optimize the dual objectives of high and stable yield. Second, there is insufficient identification and optimization of suboptimal planting areas. Current classification of suboptimal planting areas is mostly based on the average yield level of a single crop, failing to use yield stability as the core criterion. This leads to the neglect of some areas with low yields but small fluctuations, while areas with high yields but highly susceptible to climate fluctuations are not identified as optimization targets, limiting the accuracy of layout adjustments. Third, there is a lack of consideration for environmental costs. With the increasing emphasis on carbon neutrality, existing technologies have failed to effectively couple high and stable crop yield indicators with the carbon emissions of different crops, and have not systematically incorporated environmental cost considerations into the optimization of planting layout in suboptimal areas, making it difficult to achieve synergistic regulation of yield improvement and carbon emission reduction.
[0005] Therefore, there is an urgent need for a crop planting layout optimization method that can systematically assess crop yield stability, accurately identify suboptimal planting areas, and synergistically optimize crop yield and carbon emissions. This method can compensate for the shortcomings of existing technologies in adjusting the planting structure in suboptimal areas and achieve synergistic regulation of improved yield stability and reduced carbon emissions. Summary of the Invention
[0006] This application aims to at least partially address one of the technical problems in related technologies. To this end, the crop planting layout optimization method based on crop mechanism models, machine learning, and life cycle assessment provided in this application can accurately identify suboptimal areas with unstable yields by constructing a high and stable yield index (HSI) and simultaneously measuring average yield and interannual fluctuations. By coupling the APSIM mechanism model with a machine learning prediction model, it retains the mechanism constraints while enhancing the ability to capture complex nonlinear relationships, breaking through the limitations of single models in terms of accuracy and efficiency, and achieving a synergistic closed loop of the goals of "high and stable yield" and "low carbon" in crop layout optimization.
[0007] To achieve the above objectives, this application provides a method for optimizing crop planting layout based on crop mechanism models, machine learning, and life cycle assessment, including: S1. Construct a gridded multi-source database for the target agricultural region, use crop mechanism models to simulate the yield sequence of the target crop in each grid unit under multi-year historical meteorological conditions, and calculate the high and stable yield index of each grid based on the average yield and the degree of yield fluctuation. S2. Identify the sub-suitable planting area grids for each target crop based on the high and stable yield index threshold, and determine the current crop planted in the sub-suitable planting area grid as the original crop; S3. Extract environmental feature variables related to the high and stable yield capacity of crops, use the high and stable yield index as the target variable and the environmental feature variables as the input variables, and train machine learning prediction models for different target crops to predict the high and stable yield index of the target crops on any grid. S4. For the sub-suitable planting area grid, select other target crops besides the original crop as candidate alternative crops, obtain the predicted high and stable yield index of each candidate alternative crop based on the crop mechanism model and the machine learning prediction model, and use the life cycle assessment method to calculate the greenhouse gas emissions and emission intensity of each crop planting scheme. S5. Based on the high yield stability index, greenhouse gas emissions and emission intensity, perform multi-objective comparisons of the original crop and each candidate alternative crop, determine the crop substitution scheme according to the preset optimization rules, and output the optimized crop planting spatial layout.
[0008] Preferably, the multi-source database includes crop distribution data, historical meteorological data, soil attribute data, and crop management data with uniform spatial resolution.
[0009] Preferably, the crop distribution data is used to characterize the current planting range and planting area of each target crop; The historical meteorological data includes daily maximum and minimum temperatures, solar radiation, wind speed, relative humidity, and precipitation over many years; The soil property data include soil depth, soil bulk density, saturated water content, field capacity, wilting water content, organic carbon content, total nitrogen content, and pH value. The crop management data includes crop variety parameters, sowing date, and planting density.
[0010] Preferably, step S1 uses a crop mechanism model to simulate the yield sequence of the target crop in each grid cell under multi-year historical meteorological conditions, wherein the target crop includes at least two of corn, rice and soybean; and the yield sequence includes a rainfed yield sequence.
[0011] Preferably, the crop mechanism model is an APSIM model, including at least two of the APSIM-Maize model, APSIM-Rice model, and APSIM-Soybean model.
[0012] Preferably, in step S3, the formula for calculating the high and stable production index is: ; in, For the first i The high stable production index of each grid For the first i Average annual rain-fed yield per grid S i For the first i Standard deviation of multi-year average rain-fed yield per grid The average rainfed yield is the total yield of all grids within the study area.
[0013] Preferably, the environmental characteristic variables in step S3 include three categories: target crop temperature index, target crop moisture index, and soil index for each grid unit.
[0014] Preferably, the target crop temperature indicators include the multi-year average accumulated temperature of heat injury, the multi-year average accumulated temperature of cold injury, the multi-year average total number of days of heat injury, and the multi-year average total number of days of cold injury. The target crop moisture indicators include the multi-year average total precipitation during the crop growth period, the longest consecutive wet days during the crop growth period, and the longest consecutive dry days during the crop growth period. The soil indicators include organic carbon content, total nitrogen content, soil bulk density, and pH value.
[0015] Preferably, the target crop temperature index is calculated based on the upper and lower temperature thresholds of the target crop, and the target crop moisture index is calculated based on the precipitation during the crop's growth period.
[0016] Preferably, step S1, which involves simulating the yield sequence of the target crop under multi-year historical meteorological conditions within each grid cell using a crop mechanism model, includes: For each target crop, multiple maturity type variety parameters are set, and the multiple maturity types include at least early-maturing varieties, mid-maturing varieties, and late-maturing varieties; Based on soil property data, historical meteorological data, crop management data, and variety parameters for each maturity type in each grid cell, the crop mechanism model is used to simulate the multi-year yield sequence of each target crop in each grid cell.
[0017] Preferably, the machine learning prediction models for different target crops in step S3 include a high and stable yield index prediction model for maize, a high and stable yield index prediction model for rice, and a high and stable yield index prediction model for soybean. The machine learning prediction model is selected from any one of the following: multiple linear regression model, ridge regression model, decision tree model, support vector machine model, random forest model, Light GBM model, and XG Boost model.
[0018] Preferably, the preset optimization rule is configured as follows: The high yield stability index of the candidate alternative crop is higher than that of the original crop, and the greenhouse gas emissions or emission intensity of the candidate alternative crop are lower than those of the original crop; or The rate of increase in the high-yield stability index of the candidate alternative crop relative to the original crop exceeds a first preset threshold, and the rate of increase in the greenhouse gas emissions or emission intensity of the candidate alternative crop relative to the original crop does not exceed a second preset threshold, wherein the value of the first preset threshold is greater than the second preset threshold; or The rate of decrease in greenhouse gas emissions or emission intensity of the candidate alternative crop relative to the greenhouse gas emissions or emission intensity of the original crop exceeds a third preset threshold, and the rate of decrease in the high yield stability index of the candidate alternative crop relative to the high yield stability index of the original crop does not exceed a fourth preset threshold, wherein the value of the third preset threshold is greater than the value of the fourth preset threshold.
[0019] Preferably, after obtaining the predicted high and stable yield index of the alternative crop in step S4, the method further includes: using the SHAP interpretability analysis method to quantify the marginal contribution of each environmental characteristic variable to the high and stable yield index, and obtaining the influence direction and threshold response characteristics of each environmental characteristic variable.
[0020] Preferably, in step S4, the greenhouse gas emissions include the sum of indirect emissions from agricultural inputs, N2O emissions from farmland, and CH4 emissions from farmland; the agricultural inputs include pesticides, seeds, fertilizers, labor, mulch film, and agricultural machinery; and the greenhouse gas emission intensity is determined based on the greenhouse gas emissions and yield of crops.
[0021] Based on the above technical solutions, the crop planting layout optimization method based on crop mechanism models, machine learning, and life cycle assessment proposed in this application has at least one of the following beneficial effects compared to existing technologies: 1. This application uses a crop mechanism model to simulate yield sequences and calculates a high-yield stability index based on the mean yield and the degree of fluctuation, thereby achieving a comprehensive characterization of yield level and interannual stability, overcoming the shortcomings of traditional methods that ignore the risk of climate fluctuations; and identifies sub-suitable planting areas based on the high-yield stability index threshold, so that optimization is precisely focused on the target areas that need adjustment; at the same time, by extracting environmental characteristic variables and training machine learning prediction models, the process interpretability of the mechanism model and the nonlinear fitting ability of machine learning are organically coupled, breaking through the limitations of the accuracy and efficiency of a single model.
[0022] 2. This application introduces the SHAP interpretability analysis method after the output of the machine learning prediction model to quantify the marginal contribution of each environmental characteristic variable to the high and stable yield index and its threshold response characteristics, making the black-box decision-making process of traditional machine learning transparent. This step not only provides a traceable causal explanation for each alternative, enhancing the credibility of the decision, but also makes the regional customization of model feature selection and optimization rules based on evidence when this method is promoted to other agricultural regions, significantly improving the transferability of the technical solution.
[0023] 3. This application introduces life cycle assessment to calculate greenhouse gas emissions and emission intensity, and conducts multi-objective comparative decision-making with the high and stable yield index to achieve a synergistic closed loop between the goals of "high and stable yield" and "low carbon," making the output planting layout scheme both climate-resilient and environmentally friendly. By decomposing greenhouse gas emissions into three parts—indirect emissions from agricultural inputs, N2O emissions from farmland, and CH4 emissions from farmland—it covers the complete emission chain from upstream production to direct emissions in the field during the crop planting stage, making carbon emission constraints a quantifiable and operable means in layout optimization. By calculating the greenhouse gas emission intensity combining crop yield and emissions, a high-yield and green agricultural production mode is determined. Attached Figure Description
[0024] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0025] Figure 1 This is a flowchart of the crop planting layout optimization method based on crop mechanism model, machine learning and life cycle assessment provided in this application. Detailed Implementation
[0026] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0027] The terms “first,” “second,” “third,” “fourth,” “fifth,” “sixth,” “seventh,” and “eighth,” etc. (if present), in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein.
[0028] Furthermore, the terms “comprising” and “having”, and any variations thereof, are intended to cover non-exclusive inclusion, such that a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to such process, method, product, or apparatus.
[0029] Existing crop layout optimization technologies generally suffer from three main shortcomings: First, there is insufficient systematic assessment of yield stability. Current research often employs mean-oriented assessment methods, focusing on the simulation and improvement of yield under single-year or multi-year average climatic conditions. This makes it difficult to capture the fluctuation characteristics of crop yield in unfavorable climatic years and lacks the ability to synergistically optimize both high and stable yields. Second, there is insufficient identification and optimization of suboptimal planting areas. Current classification of suboptimal planting areas is mostly based on the average yield level of a single crop, failing to use yield stability as the core criterion. This leads to the neglect of some areas with low yields but small fluctuations, while areas with high yields but highly susceptible to climatic fluctuations are not identified as optimization targets, limiting the accuracy of layout adjustments. Third, there is a lack of consideration for environmental costs. With the increasing emphasis on carbon neutrality, existing technologies have failed to effectively couple high and stable crop yield indicators with the carbon emissions of different crops, and have not systematically incorporated environmental cost considerations into the optimization of planting layouts in suboptimal areas, making it difficult to achieve synergistic regulation of yield improvement and carbon emission reduction.
[0030] To address the aforementioned issues, this application proposes a crop planting layout optimization method based on crop mechanism models, machine learning, and life cycle assessment. The technical concept of this method is as follows: by coupling crop mechanism models and machine learning prediction models, a four-stage technical framework is constructed: "mechanism simulation—machine learning prediction model prediction—life cycle assessment—multi-objective decision-making." First, the crop mechanism model is used to simulate multi-year yield sequences and calculate a high-yield stability index that integrates the average yield and the degree of fluctuation, achieving accurate identification of suboptimal planting areas. Second, environmental characteristic variables are extracted to train the machine learning prediction model, overcoming the limitations of the mechanism model's large computational load and difficulty in large-scale spatial extrapolation. Third, the life cycle assessment method is introduced to quantify the greenhouse gas emissions and emission intensity of each crop planting scheme. Finally, the high-yield stability index and carbon emissions are used as dual constraints for multi-objective decision-making, achieving synergistic optimization of yield stability improvement and carbon emission reduction within suboptimal crop areas, while considering assessment accuracy and decision interpretability, thereby compensating for the shortcomings of existing technologies in adjusting planting structures in suboptimal areas.
[0031] like Figure 1 As shown in the embodiments of this application, a method for optimizing crop planting layout based on crop mechanism models, machine learning, and life cycle assessment is provided, including the following steps: S1. Construct a gridded multi-source database for the target agricultural region, use crop mechanism models to simulate the yield sequence of the target crop in each grid unit under multi-year historical meteorological conditions, and calculate the high and stable yield index of each grid based on the average yield and the degree of yield fluctuation. S2. Identify the sub-suitable planting area grids for each target crop based on the high and stable yield index threshold, and determine the current crop planted in the sub-suitable planting area grid as the original crop; S3. Extract environmental feature variables related to the high and stable yield capacity of crops, use the high and stable yield index as the target variable and the environmental feature variables as the input variables, and train machine learning prediction models for different target crops to predict the high and stable yield index of the target crops on any grid. S4. For the sub-suitable planting area grid, select other target crops besides the original crop as candidate alternative crops, obtain the predicted high and stable yield index of each candidate alternative crop based on the crop mechanism model and the machine learning prediction model, and use the life cycle assessment method to calculate the greenhouse gas emissions and emission intensity of each crop planting scheme. S5. Based on the high yield stability index, greenhouse gas emissions and emission intensity, perform multi-objective comparisons of the original crop and each candidate alternative crop, determine the crop substitution scheme according to the preset optimization rules, and output the optimized crop planting spatial layout.
[0032] The preset optimization rules are configured as follows: The high yield stability index of the candidate alternative crop is higher than that of the original crop, and the greenhouse gas emissions or emission intensity of the candidate alternative crop are lower than those of the original crop; or The rate of increase in the high yield stability index of the candidate alternative crop relative to the original crop exceeds a first preset threshold, and the rate of increase in the greenhouse gas emissions or emission intensity of the candidate alternative crop relative to the original crop does not exceed a second preset threshold. The first preset threshold is greater than the second preset threshold (i.e., with the high yield stability index as the main optimization objective, a slight increase in greenhouse gas emissions is allowed when the yield stability of the candidate alternative crop is significantly improved); or The rate of change of the greenhouse gas emissions or emission intensity of the candidate alternative crop relative to the greenhouse gas emissions or emission intensity of the original crop exceeds the third preset threshold, and the rate of change of the high and stable yield index of the candidate alternative crop relative to the high and stable yield index of the original crop does not exceed the fourth preset threshold. The value of the third preset threshold is greater than the fourth preset threshold (that is, with greenhouse gas emission reduction as the main optimization objective, when the carbon emissions of the candidate alternative crop are significantly reduced, a slight decrease in the high and stable yield index is allowed to a certain extent).
[0033] In one possible implementation, this application takes the grid unit of the target agricultural area as the basic decision object, and by constructing a multi-source database, coupling crop mechanism model and machine learning model, and introducing life cycle assessment, it quantitatively evaluates the high and stable yield capacity and greenhouse gas emissions of different crops on different grids, and outputs the optimal planting layout scheme.
[0034] Specifically, in step S1, a gridded multi-source database of the target agricultural region is first constructed. Gridding refers to dividing the study area into several grid cells with a uniform spatial resolution, with each grid cell serving as an independent evaluation and decision-making object. Data sources in the multi-source database include, but are not limited to, satellite remote sensing data, reanalysis meteorological data, soil survey data, and agricultural statistics. In one specific implementation, the spatial resolution of the grid is set to 0.1° × 0.1°. This resolution strikes a balance between computational efficiency and spatial accuracy, capturing the spatial variability of climate and soil conditions within the region without causing excessive computational burden in subsequent simulations (the choice of resolution is not limited to 0.1° and can be flexibly adjusted according to the size of the target area and data availability; for example, it can be set to 0.05°, 0.25°, or higher).
[0035] Then, crop mechanism models were used to simulate the yield sequence of target crops under multi-year historical meteorological conditions within each grid cell. Specifically, APSIM mechanism models were established for the target crops (rice, soybean, and maize), including APSIM-Maize, APSIM-Rice, and APSIM-Soybean models. Multiple maturity types were set for each crop, preferably including early-maturing, mid-maturing, and late-maturing varieties (the purpose of setting multiple maturity types is to cover a variety population with different growth periods that may be planted in the region, making the mechanism simulation more representative). Based on the collected variety parameters of the three maturity types of different crops, the rainfed yield sequence of each grid cell was simulated. The crop mechanism model is a dynamic simulation model based on crop physiological and ecological processes, which can simulate the growth, development, and yield formation process of crops on a daily basis according to meteorological conditions, soil properties, and management measures. The time span of multi-year historical meteorological conditions is usually chosen to be a long time series in order to fully capture the interannual variability of climate. In one specific implementation, the historical meteorological period was set to 40 years from 1981 to 2020, which is long enough to include a variety of typical and atypical climate years. It should be noted that the selection of historical meteorological periods is not limited to this specific period; different time spans, such as from 1991 to 2020, can also be selected.
[0036] Finally, a high-stability-yield index is calculated for each grid based on the average yield and the degree of yield fluctuation. The high-stability-yield index is a comprehensive indicator that simultaneously reflects the yield level and the stability of yield fluctuations across different years. A higher average yield indicates better production potential for the grid; a lower degree of yield fluctuation indicates stronger production stability. By combining these two factors, the high-stability-yield index effectively overcomes the limitations of traditional methods that rely solely on average yield for evaluation.
[0037] Specifically, in step S2, the sub-suitable planting area grids for each target crop are identified based on the high yield stability index threshold. The threshold divides the grids within the study area into two categories: "suitable planting areas" and "sub-suitable planting areas." The threshold is typically determined based on the statistical distribution of the high yield stability index across all grids. The current crop planted within the sub-suitable planting area grid is identified as the original crop. The original crop refers to the main crop type currently or historically planted in that grid, and it is also the crop that needs to be replaced in subsequent assessments.
[0038] The formula for calculating the above-mentioned high and stable production index is as follows: in, For the first i The high stable production index of each grid For the first i Average annual rain-fed yield per grid S i For the first i Standard deviation of multi-year average rain-fed yield per grid The average rainfed yield is the total yield of all grids within the study area.
[0039] For each crop, a High Yield Stability Index (HSI) is calculated on each grid cell based on multi-year rainfed yield sequences to simultaneously characterize average yield level and interannual fluctuation risk. Suboptimal planting areas are identified according to the quantile thresholds of the HSI for each crop. Preferably, grid cells with an HSI below the 25th percentile are defined as suboptimal planting areas for that crop; these areas serve as candidate areas for subsequent planting structure adjustments. From an optimization feasibility perspective, defining grid cells with an HSI below the 25th percentile as suboptimal areas achieves a reasonable balance between identification scale and optimization feasibility. It should be noted that the selection of the 25th percentile threshold is only a preferred implementation method and does not constitute a limitation on the scope of protection.
[0040] Specifically, in step S3, environmental characteristic variables related to the high and stable yield capacity of crops are first extracted. These environmental characteristic variables are statistical quantities derived from meteorological and soil data, used to characterize environmental factors affecting crop yield and its stability. Then, using the high and stable yield index as the target variable and the environmental characteristic variables as input variables, machine learning prediction models for different target crops are trained respectively. The target variable is the high and stable yield index that the model needs to predict, and the input variables are the environmental characteristic variables on which the model makes predictions. The role of the machine learning prediction model is to learn the mapping relationship between environmental characteristic variables and the high and stable yield index, so that even for grids without historical yield records or where the crop has not yet been planted, the high and stable yield index of the crop in that grid can be predicted based on its environmental characteristics.
[0041] The reason for training models for different target crops is that different crops have significantly different response mechanisms to environmental conditions. For example, rice is a warm- and humid crop, and its yield stability is mainly constrained by water supply and low-temperature stress; while corn and soybeans are dryland crops, and their response patterns to temperature and water are different. Therefore, training a dedicated prediction model for each target crop can more accurately characterize the crop's environmental response features. In one implementation, the target crops include corn, rice, and soybeans, and separate high-yield stability index prediction models are trained for corn, rice, and soybeans. After training, the machine learning prediction model can be used to predict the high-yield stability index of the target crop on any grid, including grids where the crop is not currently planted.
[0042] Specifically, in step S4, for the suboptimal planting area grid, other target crops besides the original crop are selected as candidate alternative crops. "Candidate alternative crops" refer to other target crop types that have the potential to replace the original crop in a given grid. If the original crop of the grid is soybean, the candidate alternative crop could be corn or rice; and so on.
[0043] Then, the predicted high and stable yield index of each candidate alternative crop is obtained based on the crop mechanism model and the machine learning prediction model. The specific process is as follows: first, the crop mechanism model is used to simulate the yield sequence of the candidate alternative crop on the grid to obtain its mean yield and fluctuation parameters; then, the meteorological variables, soil variables and the simulation results of the crop mechanism model are input together into the machine learning prediction model of the crop to obtain the predicted value of the high and stable yield index of the candidate alternative crop on the grid.
[0044] Simultaneously, the greenhouse gas emissions and emission intensity of each crop cultivation scheme are calculated using life cycle assessment (Life Cycle Assessment) methods. Life cycle assessment is a method for systematically evaluating the environmental impact of a product, process, or service throughout its entire life cycle, from raw material acquisition and production to final disposal. In this application, the Life Cycle Assessment focuses on greenhouse gas emissions and emission intensity during the crop cultivation stage.
[0045] Specifically, in step S5, a multi-objective comparison is performed on the original crop and each candidate alternative crop based on the high yield stability index, greenhouse gas emissions, and emission intensity. Multi-objective comparison refers to simultaneously considering at least two potentially contradictory objective dimensions, which in this application are the "yield stability index (high yield stability index)" and the "environmental impact index (greenhouse gas emissions or emission intensity)." Crop substitution schemes are determined according to preset optimization rules. These preset optimization rules are a set of pre-defined criteria used to judge whether candidate alternative crops are superior to the original crop. The general idea of the rules is that the substitution scheme should improve yield stability without significantly worsening environmental emissions; or, the substitution scheme should improve environmental emissions without significantly worsening yield stability. The optimized crop planting spatial layout is output, which can be one or more crop planting layout maps, marking the optimal crop type recommended for planting in each grid unit.
[0046] In one possible implementation, after obtaining the predicted high and stable yield index of the alternative crop in step S4, the step further includes an interpretability analysis of the optimal machine learning prediction model.
[0047] This interpretability analysis employs the SHAP interpretability analysis method. SHAP is a machine learning prediction model interpretation method based on Shapley values in cooperative game theory. Its core idea is to decompose the predicted value of a machine learning prediction model for a specific input sample into the sum of the contributions of each input feature to that predicted value. By calculating the Shapley value of each environmental feature variable, the following three types of output information can be obtained: The first type of output information is the global importance ranking of each environmental characteristic variable. The average absolute value of each variable's Shapley value across all samples is used as its importance measure, and these variables are ranked from largest to smallest. This visually reveals which environmental variables have the most significant impact on the crop yield stability index. For example, in one embodiment, the maize yield stability index is most significantly affected by growing season precipitation, multi-year average accumulated heat temperature, and soil bulk density.
[0048] The second type of output information is the direction of the positive or negative influence of the values of each environmental characteristic variable. By observing the distribution of the positive and negative signs of a certain characteristic's Shapley value, it can be determined whether an increase in that characteristic value promotes or inhibits the high yield stability index. For example, in one embodiment, an increase in the low temperature stress index generally inhibits the high yield stability index.
[0049] The third type of output information is the threshold response characteristics of each environmental characteristic variable. By plotting a scatter plot or dependency graph between the value of a variable and its Shapley value, the nonlinear response relationship between the variable and the high-yield stability index can be observed, and the threshold points that cause a significant change in the function's shape can be identified.
[0050] The role of this interpretability analysis step is manifested in the following ways: First, it ensures that the prediction results of the machine learning prediction model are no longer an untraceable "black box," but can be explained from a physical mechanism perspective regarding the model's decision-making logic. Second, the interpretability analysis results provide scientific guidance for the selection of environmental characteristic variables and the regional customization of optimization rules when extending the patent to other agricultural regions. Finally, the interpretability analysis results can be cross-validated with expert knowledge in the crop field, enhancing the scientific credibility of the "mechanism-machine learning" hybrid modeling framework of this application. It should be noted that the interpretability analysis method is not limited to SHAP; other model interpretation methods such as LIME (Locally Interpretable Model-Independent Interpretation) or partial dependency graphs can also achieve similar purposes.
[0051] In one possible implementation, the specific numerical ranges of the first to fourth preset thresholds in the pre-set optimization rules of crop substitution schemes can be refined based on the output of the SHAP interpretability Analysis method. Specifically, the SHAP analysis outputs three types of information: the global importance ranking of each environmental characteristic variable, the positive and negative influence direction of each variable, and the threshold response characteristics of each variable. Based on the global importance ranking of each environmental characteristic variable, the main controlling environmental factors affecting the high and stable yield index of the original crop within the grid of the sub-suitable planting area are identified. When the main controlling environmental factor belongs to the variable type characterizing interannual climate fluctuations (such as the number of days with low temperatures or the number of consecutive days with low temperatures), it indicates that the main cause of yield fluctuations is uncontrollable climate risk. In this case, the second or fourth preset threshold is increased to relax the constraints on the increase in greenhouse gas emissions or emission intensity or the absolute value of the rate of change of the high and stable yield index. This is because the core benefit of the substitution scheme at this time lies in avoiding specific climate risk windows by changing crop types, thereby improving the climate resilience of the planting system.
[0052] When the controlling environmental factor belongs to the variable type characterizing soil basic fertility or soil physical structure (such as soil organic carbon content, high soil bulk density), it indicates that yield fluctuations are related to long-term farming management practices. Therefore, the second or fourth preset threshold should be reduced to strengthen the constraint on the increase in greenhouse gas emissions or the absolute value of the rate of change in the high-yield stability index. Simply changing crop types without altering high-input management habits may perpetuate or even exacerbate dependence on high-carbon inputs. In this case, the preset optimization rules can be supplemented with additional conditions, such as requiring alternative solutions to be accompanied by low-carbon measures like conservation tillage or soil testing-based fertilization before being deemed feasible. Caution should be exercised when changing crops; they must be cleaner and more sustainable. Fundamental soil health issues should not be ignored for the sake of short-term stable yields or reduced carbon emissions.
[0053] Through the above methods, SHAP interpretability analysis not only provides causal explanations for the prediction results of machine learning prediction models, but also upgrades the preset optimization rules from a set of fixed threshold judgments to a flexible decision-making framework that can be adaptively adjusted, thus taking into account both the scientific nature and flexibility of decision-making, and significantly improving the applicability and scalability of this method in different agricultural regions.
[0054] In one possible implementation, the specific composition of the gridded multi-source database constructed in step S1 is further defined.
[0055] Crop distribution data is used to characterize the current planting range and area of each target crop. This data can be derived from global or regional spatial crop distribution datasets, or from agricultural statistical yearbooks or remote sensing interpretation data. The current planting range refers to the geographical boundary of the target crop's current actual distribution, and the planting area refers to the actual sown area of the crop within each grid cell.
[0056] Historical meteorological data includes daily maximum and minimum temperatures, solar radiation, wind speed, relative humidity, and precipitation over many years. These variables are the basic meteorological inputs required to drive crop mechanism models to simulate crop growth processes. Solar radiation is measured in MJ / m² / d, wind speed in m / s, and precipitation in mm. Data sources can be reanalysis datasets or interpolated observation products from surface meteorological stations. Reanalysis datasets refer to a set of spatiotemporally continuous, physically consistent, and multi-element complete historical meteorological gridded data products generated by reprocessing and integrating various historically scattered and discontinuous meteorological observation data (such as observations from surface meteorological stations, weather balloons, satellites, aircraft, and ships) using numerical weather prediction models and data assimilation techniques.
[0057] Soil property data includes soil depth, soil bulk density, saturated water content, field capacity, wilting water content, organic carbon content, total nitrogen content, and pH value. Among these, soil depth determines the physical space available for crop root extension; saturated water content, field capacity, and wilting water content collectively describe the soil's water-holding capacity and the range of available water for crops, and are key parameters in the water balance module of the mechanistic model; organic carbon content and total nitrogen content reflect the soil's basic fertility level; and pH value affects nutrient availability. Data can be sourced from global or regional gridded soil information databases.
[0058] Crop management data includes crop variety parameters, sowing date, and planting density. Variety parameters should include at least the growth and development characteristics of varieties with different maturity periods; the sowing date should be the typical sowing date for this crop in this region; and the planting density should be the number of plants or hills per unit area.
[0059] By unifying and aligning the multi-source data mentioned above, a comprehensive data foundation is constructed to support subsequent mechanistic simulation and machine learning prediction model building. The data sources are not limited to the datasets specifically listed above; other available data products can be selected based on the specific circumstances of the target region.
[0060] In one possible implementation, the target crop type, crop mechanism model type, and yield sequence type involved in step S1 are further defined.
[0061] The target crops include at least two of corn, rice, and soybeans. These three crops are the main food crops in the black soil region of Northeast China and are the focus of this invention. It should be noted that the method framework of this invention is also applicable to other crops such as wheat, sorghum, and rapeseed, and is not limited to the three crops mentioned above.
[0062] The yield series includes a rainfed yield series. Rainfed yield refers to crop yield under conditions of no irrigation supplementation and complete reliance on natural precipitation. Rainfed yield is chosen as the evaluation indicator because, in regions like the Northeast Black Soil Region where rainfed agriculture is dominant, precipitation and its interannual fluctuations are the leading factors affecting yield stability. Multi-year rainfed yield series can accurately reflect the actual production potential and risks of crops under the grid climate conditions. For regions with irrigation, irrigated yield or potential yield can also be simulated, and this method is not the only option.
[0063] In one possible implementation, the composition of the environmental characteristic variables in step S3 is further defined.
[0064] Environmental characteristic variables include three main categories: target crop temperature index, target crop moisture index, and soil index for each grid cell. These three categories of indicators systematically characterize the environmental conditions of the grid from three dimensions: heat, water, and soil.
[0065] The target crop temperature indicators include: multi-year average accumulated temperature of heat injury, defined as the cumulative average of the temperature at which the daily maximum temperature exceeds the upper limit temperature threshold of the crop during the crop's growth period; and multi-year average accumulated temperature of cold injury, defined as the cumulative average of the temperature difference between the daily minimum temperature and the lower limit temperature threshold of the crop during the crop's growth period.
[0066] Preferably, in this embodiment, the multi-year average total number of days of heat injury is calculated using the multi-year average total number of days of heat injury for ≥3 consecutive days, defined as the multi-year average of the number of days during the crop growth period when the daily maximum temperature reaches or exceeds the upper limit temperature threshold for 3 consecutive days; the multi-year average total number of days of cold injury is calculated using the multi-year average total number of days of cold injury for ≥3 consecutive days, defined as the multi-year average of the number of days during the crop growth period when the daily minimum temperature reaches or exceeds the lower limit temperature threshold for 3 consecutive days.
[0067] The target crop temperature index is calculated based on the upper and lower temperature thresholds of the target crop. The upper temperature threshold is the highest temperature above which the crop's photosynthetic rate begins to decline and respiration consumption increases; the lower temperature threshold is the lowest temperature below which crop growth and development are significantly inhibited or even suffer chilling injury. Different crops have different temperature thresholds due to their different origin environments and physiological characteristics. The selection of the above thresholds is based on the sensitivity of the corresponding crop to temperature stress during critical reproductive growth stages (such as the tasseling and silking stage of maize, the heading and booting stage of rice, and the flowering and pod-setting stage of soybean).
[0068] The target crop moisture indicators include: the multi-year average total precipitation during the crop growth period, which is the average of the total precipitation from sowing to maturity over a certain number of years; the longest consecutive wet days during the crop growth period, specifically the multi-year average of the number of days with consecutive daily precipitation exceeding a certain threshold (e.g., 5 mm) for more than 3 consecutive days; and the longest consecutive dry days during the crop growth period, specifically the multi-year average of the number of days with consecutive daily precipitation below a certain threshold (e.g., 0.1 mm).
[0069] Soil indicators include organic carbon content, total nitrogen content, soil bulk density, and pH value. Organic carbon content reflects the level of soil organic matter and affects the soil's water and fertilizer retention capacity; total nitrogen content directly affects the nitrogen supply to crops; and pH value affects nutrient availability. Soil bulk density affects root penetration resistance and soil aeration and permeability.
[0070] It should be noted that the composition of the above-mentioned environmental characteristic variables is not limited to the items listed. They can be appropriately added or removed according to the specific environmental characteristics of the target area and the availability of data. For example, wind speed can be added to reflect evapotranspiration intensity, or soil effective water capacity can be added to comprehensively reflect soil moisture conditions.
[0071] In one possible implementation, the type and training method of the machine learning prediction model in step S3 are further specified.
[0072] Machine learning prediction models for different target crops include a high-yield and stable-yield index prediction model for maize, a high-yield and stable-yield index prediction model for rice, and a high-yield and stable-yield index prediction model for soybeans. Although these three models share the basic model structure and training process, their target variables and input variables have different statistical distribution characteristics, so they are trained and evaluated independently.
[0073] Model selection uses root mean square error and coefficient of determination as evaluation metrics, and the predictive performance of each model is compared on an independent validation set. In one specific embodiment, the random forest model showed the best performance in HSI prediction for maize, rice, and soybean. It should be noted that the determination of the optimal model depends on the specific data characteristics and study area, and is not limited to the choice of random forest alone.
[0074] In one possible implementation, the specific composition of greenhouse gas emissions and emission intensity in step S4 is further defined.
[0075] Greenhouse gas emissions include indirect emissions from agricultural inputs, N2O emissions from farmland, and CH4 emissions from farmland. This composition covers the main sources of greenhouse gas emissions during the crop growing season, and the greenhouse gas emission intensity is determined based on the greenhouse gas emissions and yield of the crop.
[0076] Indirect emissions from agricultural inputs refer to greenhouse gas emissions generated during the production, manufacturing, and transportation of agricultural inputs used in crop production. These emissions are upstream indirect emissions. Agricultural inputs include pesticides, seeds, fertilizers, labor, mulch film, and agricultural machinery. The indirect emissions of each input are calculated by multiplying its usage per unit area by the corresponding emission factor. The emission factor can be selected by referring to authoritative life cycle databases or relevant national standards.
[0077] Agricultural N2O emissions refer to direct N2O emissions generated during soil microbial nitrification and denitrification processes of nitrogen applied to farmland (including chemical fertilizer nitrogen and organic fertilizer nitrogen), as well as indirect N2O emissions into the atmosphere and water bodies after ammonia volatilization and nitrogen leaching. N2O is a potent greenhouse gas with a global warming potential 298 times that of carbon dioxide.
[0078] CH4 emissions from farmland refer to the CH4 emissions produced during crop cultivation, especially in paddy fields, from the decomposition of soil organic matter by methanogenic bacteria. The global warming potential of CH4 is 25 times that of carbon dioxide.
[0079] Classifying greenhouse gas emissions into indirect emissions from inputs, N2O emissions, and CH4 emissions helps identify the main contributors to emissions from each crop, providing targeted guidance for the development of emission reduction measures.
[0080] Greenhouse gas emission intensity is calculated based on crop yield per unit area and greenhouse gas emissions, which is helpful for quantifying the environmental impact of crop production per unit area.
[0081] In one possible implementation, the black soil region of Northeast China is taken as the research area, and a grid cell with a spatial resolution of 0.1°×0.1° is used as the basic decision object, covering the three major target crops of maize, rice and soybean. All steps of the above method are fully executed to achieve synergistic optimization of improving yield stability and reducing carbon emissions in sub-suitable areas.
[0082] First, a gridded multi-source database for the study area was constructed. Crop planting spatial distribution data were sourced from globally accepted spatial crop distribution datasets (e.g., SPAM2020, Spatial Production Allocation Model 2020, the 2020 version of the spatial yield allocation model). Multi-year daily meteorological data were sourced from reanalysis datasets (e.g., ERA5, ECMWF Reanalysis v5, the fifth generation of reanalysis data from the European Centre for Medium-Range Weather Forecasts), with a spatial resolution of 0.1°, including maximum and minimum temperatures, solar radiation, wind speed, relative humidity, and precipitation, with complete and uninterrupted station data. Soil property data were sourced from a global gridded soil information database (e.g., Soil Grids), with a spatial resolution of 0.08°, including soil depth (0-100cm), soil bulk density, field capacity, wilting moisture content, organic carbon content, total nitrogen content, and pH value.
[0083] Crop management data includes crop variety parameters, sowing date, and planting density. Maize was sown on April 28th with a density of 60,000 plants / hm²; rice was sown on April 20th with a density of 30 hills / m²; soybeans were sown on May 5th with a density of 350,000 plants / hm²; each crop was categorized into early-maturing, mid-maturing, and late-maturing varieties.
[0084] Soil data was integrated into the geographic information system environment and resampled to a resolution of 0.1°×0.1° using the nearest neighbor interpolation method to maintain spatial consistency with meteorological and crop data.
[0085] Then, mechanistic models of the target crops were established and yield sequences were simulated. APSIM-Maize (maize module), APSIM-Rice (rice module), and APSIM-Soybean (soybean module) mechanistic models were established separately. Variety parameters for three maturity types (early, medium, and late) were set for each crop. Based on soil property data, historical meteorological data (1981 to 2020), crop management data, and variety parameters for each maturity type in each grid cell, the rainfed yield sequences of maize, rice, and soybean in each grid cell were simulated using the crop mechanistic models.
[0086] Taking a representative grid cell as an example, the multi-year simulation results for the corresponding grid cell are as follows: average rainfed yield of maize is 8.7 t / hm², average rainfed yield of rice is 4.2 t / hm², and average rainfed yield of soybean is 3.2 t / hm². Based on the aforementioned yield sequence, the high-yield stability index (HSI) for each grid unit was calculated using the formula for the high-yield stability index. Statistically, in this embodiment, the HSI ranges from -0.07 to 1.44 for maize, -0.92 to 1.41 for rice, and 0.01 to 1.53 for soybeans. Based on the threshold of HSI below the 25th percentile, suboptimal planting areas were identified, ultimately determining the suboptimal planting areas for maize, rice, and soybeans to be 2.4 million hectares, 1.1 million hectares, and 0.7 million hectares, respectively.
[0087] Environmental characteristic variables were extracted, and a feature variable set including temperature, moisture, and soil indicators was constructed. Temperature thresholds were set as follows: upper limit 34℃ and lower limit 8℃ for maize; upper limit 32℃ and lower limit 8℃ for rice; and upper limit 32℃ and lower limit 8℃ for soybeans. In this embodiment, the values of each environmental characteristic variable are distributed as follows: Regarding temperature indicators, the cumulative accumulated temperature exceeding the upper temperature threshold (multi-year average accumulated temperature for heat injury) in different grid units within the study area was as follows: maize 0–16.7℃·d (degrees Celsius per day), rice 0–22.7℃·d, and soybean 0–21.2℃·d. The total number of days exceeding the upper temperature threshold for three consecutive days or more (multi-year average total number of days for heat injury) was as follows: maize 0–2.5 days, rice 0–3.7 days, and soybean 0–2.2 days. The cumulative accumulated temperature below the lower temperature threshold (multi-year average accumulated temperature for chilling injury) was as follows: maize 0.3–630.8℃·d, rice 24.7–452.5℃·d, and soybean 10–718.8℃·d. The total number of days below the lower temperature threshold for three consecutive days or more (multi-year average total number of days for chilling injury) was as follows: maize 0–70.2 days, rice 5.9–59.9 days, and soybean 1.8–68.6 days.
[0088] Regarding moisture indicators, the multi-year average total precipitation during the crop growth period in different grid units within the study area was: 36 mm to 918 mm for maize, 82 mm to 831 mm for rice, and 44 mm to 839 mm for soybeans. The longest consecutive wet days during the crop growth period were: 12 to 50 days for maize, 15 to 46 days for rice, and 10 to 45 days for soybeans. The longest consecutive dry days during the crop growth period were: 3 to 83 days for maize, 7 to 76 days for rice, and 7 to 84 days for soybeans.
[0089] Regarding soil indicators, the total nitrogen content in different grid units within the study area ranged from 0.1 g / kg to 1.18 g / kg; the organic carbon content ranged from 0.87 g / kg to 17.7 g / kg; the soil bulk density ranged from 1.0 g / cm³ to 1.51 g / cm³; and the pH value ranged from 5.7 to 8.8.
[0090] Using the high and stable yield index as the target variable and the aforementioned environmental characteristic variables as input variables, the training and validation sets were randomly divided at an 8:2 ratio (with a fixed random seed to ensure reproducibility). Machine learning prediction models for corn, rice, and soybean were trained separately. Seven types of models—multiple linear regression, ridge regression, decision tree, support vector machine, random forest, Light GBM, and XG Boost—were optimized using grid search and 5-fold cross-validation. Model performance was comprehensively evaluated using the root mean square error (RMSE) of the validation set and the coefficient of determination (R²). Results showed that the random forest model exhibited the best performance in modeling corn, rice, and soybean, with a high R². 2 The mean square error was 0.89 to 0.94, the root mean square error was 0.11 to 0.12, and the normalized root mean square error was 5.4% to 7.0%, which was finally determined as the prediction model.
[0091] The SHAP interpretability analysis method was used to conduct global and local interpretability analysis on the random forest model, quantifying the marginal contribution, direction of influence, and threshold response of each environmental characteristic variable on the high and stable yield index. The ranking of the importance of global variables is as follows: The high and stable yield index of maize is most significantly affected by growing season precipitation, multi-year average accumulated temperature for heat injury, soil bulk density, and multi-year average accumulated temperature for cold injury, with average SHAP values of 0.09, 0.08, 0.08, and 0.06, respectively. The high and stable yield index of rice is most significantly affected by multi-year average accumulated temperature for cold injury, the total number of days below the lower limit temperature threshold for three consecutive days, and soil organic carbon content, with average SHAP values of 0.30, 0.06, and 0.04, respectively. The high and stable yield index of soybean is most significantly affected by multi-year average accumulated temperature for cold injury, the total number of days below the lower limit temperature threshold for three consecutive days, and the total number of days above the upper limit temperature threshold for three consecutive days, with average SHAP values of 0.19, 0.09, and 0.07, respectively.
[0092] Taking a suitable planting area grid as an example, an evaluation of alternative crops was conducted. The original crop of this grid was soybean, and rice and corn were selected as candidate alternative crops. After simulation by the APSIM model, the results were input into the random forest prediction model, and the predicted high and stable yield index (HSI) for each crop were obtained: the HSI for the original crop soybean was 0.44, the HSI for candidate corn was 0.53, and the HSI for candidate rice was 0.36.
[0093] Greenhouse gas emissions for various crop planting schemes were calculated using life cycle assessment methods. The formulas for calculating greenhouse gas emissions (GHG) and emission intensity (GHGI) are as follows: ; ; ; in: Indirect emissions from agricultural inputs; For N2O-related emissions; For CH4-related emissions, Let i be the amount of agricultural input of type i used per unit area. The corresponding emission coefficients are used; the selection of emission coefficients can refer to authoritative life cycle databases or relevant national standards; the inputs mentioned include at least pesticides, seeds, fertilizers, labor, mulch film, and agricultural machinery, etc. Y This refers to the yield per unit area of crop.
[0094] N2O emissions are expressed as follows: ; Among them, N2O t 44 / 28 represents the total N2O emissions from farmland, 44 / 28 represents the molecular weight ratio of N2O to N, and 298 represents the global warming potential equivalent coefficient of N2O (kg CO2eq kg). -1 N2O emissions from farmland originate from direct emissions of nitrogen (including chemical fertilizer nitrogen and organic fertilizer nitrogen) applied to farmland during soil microbial nitrification and denitrification processes, as well as indirect emissions into the atmosphere and water bodies through ammonia volatilization and nitrogen leaching.
[0095] CH4-related emissions The calculation method is as follows: ; Where 25 is the global warming potential equivalent coefficient of CH4 (kg CO2eq kg). -1 CH4).
[0096] ; in, CH4 emission rate (kg CH4 ha)-1 d -1 ), A Planting area (hectares, ha). D The length of the crop growing season (in days, days).
[0097] The calculated greenhouse gas emissions for each crop planting scheme in this grid are as follows: GHG emissions from soybeans are 4.60 tCO2 eq ha. -1 (tons of CO2 equivalent per hectare); GHG emissions from corn are 2.89 t CO2 eq ha. -1 The GHG emissions from rice were 10.64 t CO2 eq ha -1 The GHG emission intensity of soybeans is 2.16tCO2eq t. -1 (Total carbon dioxide equivalent per ton); Corn GHG emission intensity is 0.48t CO2eq t -1 The GHG emission intensity of rice is 3.92 t CO2eq t. -1 .
[0098] A multi-objective comparison was conducted based on the high yield stability index, greenhouse gas emissions, and emission intensity. The preset optimization rule was: the high yield stability index of the candidate alternative crop was higher than that of the original crop, and the greenhouse gas emissions or emission intensity decreased or showed no significant difference. The comparison results are as follows: for the soybean → rice scheme, the HSI decreased from 0.44 to 0.36 (a reduction of 18.2%), and GHG emissions decreased from 4.60 tCO2eq ha -1 Increased to 5.90t CO2eq ha -1 (Increased by 28.3%), GHG emission intensity increased from 2.16 t CO2eq t -1 Increased to 15.74tCO2eq t -1 (628% increase) Does not meet optimization conditions. For the soybean-corn route, HSI increased from 0.44 to 0.53 (a 20.5% increase), and GHG emission intensity increased from 2.16t CO2eq t. -1 Reduced to 0.68t CO2eq t -1 (Reduced by 69%), GHG emissions decreased from 4.60 tCO2 eq ha -1 Increased to 4.75t CO2eq ha -1 (The increase of 3.3% is within the range of no significant difference), which meets all the optimization conditions and is determined to be the optimal alternative.
[0099] Therefore, the suitable planting area grid was optimized from soybean planting to corn planting, achieving synergistic optimization of improved yield stability and reduced agricultural carbon emissions, and completing the optimal adjustment of spatial crop layout. This embodiment fully demonstrates the entire process of this method, from data construction, mechanism simulation, index calculation, subsuitable area identification, machine learning prediction model prediction, SHAP interpretable analysis, life cycle carbon emission assessment to multi-objective decision output, verifying the feasibility and effectiveness of this method in practical applications.
[0100] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0101] The foregoing has described specific embodiments of the present application. In some cases, the described actions or steps may be performed in a different order than those shown in the embodiments and the desired results may still be achieved. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are also possible or may be advantageous.
[0102] In the description of the embodiments of this application, the terms "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the embodiments of this application. In the embodiments of this application, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in a suitable manner in any one or more embodiments or examples. Furthermore, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in the embodiments of this application, as well as the features of different embodiments or examples.
[0103] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing custom logic functions or processes, and the scope of preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order according to the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.
[0104] The above embodiments are merely preferred embodiments of this application and are not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A method for crop planting layout optimization based on crop mechanism model, machine learning and life cycle assessment, characterized in that, include: S1. Construct a gridded multi-source database for the target agricultural region, use crop mechanism models to simulate the yield sequence of the target crop in each grid unit under multi-year historical meteorological conditions, and calculate the high and stable yield index of each grid based on the average yield and the degree of yield fluctuation. S2. Identify the sub-suitable planting area grids for each target crop based on the high and stable yield index threshold, and determine the current crop planted in the sub-suitable planting area grid as the original crop; S3. Extract environmental feature variables related to the high and stable yield capacity of crops, use the high and stable yield index as the target variable and the environmental feature variables as the input variables, and train machine learning prediction models for different target crops to predict the high and stable yield index of the target crops on any grid. S4. For the sub-suitable planting area grid, select other target crops besides the original crop as candidate alternative crops, obtain the predicted high and stable yield index of each candidate alternative crop based on the crop mechanism model and the machine learning prediction model, and use the life cycle assessment method to calculate the greenhouse gas emissions and emission intensity of each crop planting scheme. S5. Based on the high yield stability index, greenhouse gas emissions and emission intensity, perform multi-objective comparisons of the original crop and each candidate alternative crop, determine the crop substitution scheme according to the preset optimization rules, and output the optimized crop planting spatial layout.
2. The crop planting layout optimization method according to claim 1, characterized by, The multi-source database includes crop distribution data, historical meteorological data, soil property data, and crop management data with uniform spatial resolution.
3. The crop planting layout optimization method according to claim 2, characterized by, The crop distribution data is used to characterize the current planting range and planting area of each target crop; The historical meteorological data includes daily maximum and minimum temperatures, solar radiation, wind speed, relative humidity, and precipitation over many years; The soil property data include soil depth, soil bulk density, saturated water content, field capacity, wilting water content, organic carbon content, total nitrogen content, and pH value. The crop management data includes crop variety parameters, sowing date, and planting density.
4. The crop planting layout optimization method according to claim 1, characterized by, Step S1 uses a crop mechanism model to simulate the yield sequence of the target crop in each grid cell under multi-year historical meteorological conditions, wherein the target crop includes at least two of corn, rice and soybean; the yield sequence includes rainfed yield sequence.
5. The crop planting layout optimization method according to claim 4, characterized by, The crop mechanism model is an APSIM model, including at least two of the APSIM-Maize model, APSIM-Rice model, and APSIM-Soybean model.
6. The crop planting layout optimization method according to claim 4, characterized by, In step S3, the formula for calculating the high and stable production index is: ; in, For the first i The high stable production index of each grid For the first i Average annual rain-fed yield per grid S i For the first i Standard deviation of multi-year average rain-fed yield per grid The average rainfed yield is the total yield of all grids within the study area.
7. The crop planting layout optimization method according to claim 1, characterized by, The environmental characteristic variables mentioned in step S3 include three categories: target crop temperature index, target crop moisture index, and soil index for each grid cell.
8. The crop planting layout optimization method according to claim 7, characterized by, The target crop temperature indicators include the multi-year average accumulated temperature of heat injury, the multi-year average accumulated temperature of cold injury, the multi-year average total number of days of heat injury, and the multi-year average total number of days of cold injury. The target crop moisture indicators include the multi-year average total precipitation during the crop growth period, the longest consecutive wet days during the crop growth period, and the longest consecutive dry days during the crop growth period. The soil indicators include organic carbon content, total nitrogen content, soil bulk density, and pH value.
9. The crop planting layout optimization method according to claim 8, characterized by, The target crop temperature index is calculated based on the upper and lower temperature thresholds of the target crop, and the target crop moisture index is calculated based on the precipitation during the crop's growth period.
10. The crop planting layout optimization method according to claim 4, characterized by, Step S1, which involves simulating the yield sequence of the target crop under multi-year historical meteorological conditions within each grid cell using a crop mechanism model, includes: For each target crop, multiple maturity type variety parameters are set, and the multiple maturity types include at least early-maturing varieties, mid-maturing varieties, and late-maturing varieties; Based on soil property data, historical meteorological data, crop management data, and variety parameters for each maturity type in each grid cell, the crop mechanism model is used to simulate the multi-year yield sequence of each target crop in each grid cell.
11. The crop planting layout optimization method according to claim 4, characterized by, The machine learning prediction models for different target crops in step S3 include the high and stable yield index prediction model for maize, the high and stable yield index prediction model for rice, and the high and stable yield index prediction model for soybean. The machine learning prediction model is selected from any one of the following: multiple linear regression model, ridge regression model, decision tree model, support vector machine model, random forest model, Light GBM model, and XG Boost model.
12. The crop planting layout optimization method according to claim 1, characterized by, The preset optimization rules are configured as follows: The high yield stability index of the candidate alternative crop is higher than that of the original crop, and the greenhouse gas emissions or emission intensity of the candidate alternative crop are lower than those of the original crop. or The rate of increase in the high-yield stability index of the candidate alternative crop relative to the original crop exceeds a first preset threshold, and the rate of increase in the greenhouse gas emissions or emission intensity of the candidate alternative crop relative to the original crop does not exceed a second preset threshold, wherein the value of the first preset threshold is greater than the second preset threshold; or The rate of decrease in greenhouse gas emissions or emission intensity of the candidate alternative crop relative to the greenhouse gas emissions or emission intensity of the original crop exceeds a third preset threshold, and the rate of decrease in the high yield stability index of the candidate alternative crop relative to the high yield stability index of the original crop does not exceed a fourth preset threshold, wherein the value of the third preset threshold is greater than the value of the fourth preset threshold.
13. The crop planting layout optimization method according to claim 1, characterized by, After obtaining the predicted high and stable yield index of alternative crops in step S4, the method further includes: using the SHAP interpretability analysis method to quantify the marginal contribution of each environmental characteristic variable to the high and stable yield index, and obtaining the influence direction and threshold response characteristics of each environmental characteristic variable.
14. The crop planting layout optimization method according to claim 1, characterized by, In step S4, the greenhouse gas emissions include the sum of indirect emissions from agricultural inputs, N2O emissions from farmland, and CH4 emissions from farmland; the agricultural inputs include pesticides, seeds, fertilizers, labor, mulch film, and agricultural machinery; the greenhouse gas emission intensity is determined based on the greenhouse gas emissions and yield of crops.