A soybean planting suitable area assessment method integrated with yield prediction
By constructing a multi-temporal remote sensing vegetation index status and change trend factor set and a machine learning model, combined with the maximum entropy model, suitable areas for high soybean yields were identified, solving the problem of incomplete soybean planting adaptability assessment and achieving efficient yield prediction and resource optimization.
Patent Information
- Application Number
- CN202510855803.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-06-25
AI Technical Summary
Existing technologies fail to effectively assess the yield-increasing potential of soybean cultivation and lack quantitative analysis of the impact of ecological and environmental factors, resulting in an incomplete assessment of soybean cultivation adaptability.
By collecting soybean planting spatial layout and yield data, MOD09A1 surface reflectance products during the growth period, bioclimatic data and soil data, a multi-temporal remote sensing vegetation index status and change trend factor set was constructed. Machine learning models and maximum entropy models were used to identify suitable soybean high-yield areas and draw spatial distribution maps.
It has achieved accurate identification of suitable areas for high-yield soybeans, simplified the yield estimation model, improved the accuracy and efficiency of soybean planting adaptability assessment, and avoided waste of resources.
Smart Images

Figure CN120355272B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of data processing technology, and in particular relates to a soybean planting suitable area assessment method integrated with yield prediction. Background Art
[0002] Soybeans are an important food and oil crop, as well as a nitrogen-fixing plant, and are relatively environmentally friendly. However, China faces a severe dependence on imports. While the country is promoting the expansion of soybean cultivation to increase planted area and yield, current technologies have yet to fully analyze the adaptability of soybean cultivation based on its potential for increased yields.
[0003] Because crop growth is selective for ecological and environmental factors such as climate, soil, and topography, crop suitability assessments are used to evaluate the suitability of specific land areas for crop cultivation. Upgrading the impact of agricultural resource and environmental factors on crop distribution from traditional qualitative to quantitative evaluations also provides new insights into the response of crop spatial patterns to climate change. By comprehensively analyzing the agricultural resource and environmental conditions that crop growth depends on and their matching characteristics, the extent to which these conditions influence crop cultivation can be effectively assessed. Summary of the Invention
[0004] To solve the above technical problems, the present invention proposes a soybean planting suitable area assessment method integrated with yield prediction to identify suitable growth areas with high soybean yield.
[0005] To achieve the above objectives, the present invention provides a method for evaluating suitable areas for soybean planting integrated with yield prediction, comprising:
[0006] Collect soybean planting spatial layout and yield data, MOD09A1 surface reflectance products during the growth period, bioclimatic data, and soil data;
[0007] The soybean growing season was divided into seven typical periods based on phenology. A multi-temporal remote sensing vegetation index state and change trend factor set was constructed for soybean planting areas. Two combinations were used to input into a machine learning model. Yield was used as the predictor variable, and the model was selected through accuracy evaluation to obtain a grid-scale spatial distribution map of soybean yield.
[0008] The identified high-yield grid units were used as point data to construct a maximum entropy model, and multi-source environmental variables were integrated to carry out suitability modeling analysis. Through variable contribution analysis and response curves, the key ecological and environmental factors affecting the distribution of soybean high-yield suitability were identified, and a spatial distribution map of soybean high-yield suitability was drawn.
[0009] Optionally, the collected data include: soybean yield data, bioclimatic variables, soil variables, topographic variables and variables related to human activities;
[0010] The soybean yield data is obtained based on the long-term statistical yearbook to obtain the soybean yield data of the district or county or the soybean yield of the sampling point;
[0011] The bioclimatic variables include bioclimatic data representing annual trends, seasonality, and extreme or limiting environmental factors;
[0012] The soil variables include soil bulk density, total nitrogen content, total phosphorus content, soil cation exchange capacity, calcium carbonate content, organic matter, pH value, exchangeable Ca 2+ ;
[0013] The terrain variables include altitude, slope, aspect, terrain relief, and distance to rivers and reservoirs;
[0014] The variables related to human activities include distance to impervious surface, GDP at grid scale, and population.
[0015] Optionally, the seven typical periods include sowing period, emergence period, three-leaf period, flowering period, pod-setting period, grain-filling period and maturity period.
[0016] Optionally, the process of constructing a multi-temporal remote sensing vegetation index state and change trend factor set includes: using the MODIS land level 3 standard data product surface reflectance product to extract the normalized difference vegetation index, enhanced vegetation index, surface moisture index, and Green vegetation index; and calculating the average value and Theil-Sen slope of the vegetation index in different growth periods respectively.
[0017] Optionally, the process of inputting the machine learning model includes:
[0018] Correlation analysis was conducted between county yield and the average and slope values of vegetation indices during the seven growth periods, excluding variables that were not significantly correlated with soybean yield.
[0019] Two different combinations were used to input the machine learning model. The first combination used county yield as the prediction target and vegetation indices of different growth periods with high correlation significance as prediction factors. The second combination included the slopes of vegetation indices of different growth periods with high correlation significance with yield in addition to vegetation indices with high correlation significance.
[0020] Optionally, the process of constructing the maximum entropy model includes: screening grids with soybean yields higher than 2000 kg / ha as sample points; performing Pearson correlation analysis on environmental variables and eliminating environmental variables with correlation coefficients greater than 0.8; and converting planting records and environmental variables into CSV format and inputting them into the model.
[0021] Optionally, the parameterization process of the maximum entropy model includes: 75% of the distribution point data are used for modeling, 25% are used for model testing, and the operation is repeated 20 times; the background data is selected as 1,000 distribution points; the jackknife method is used to test the contribution rate of each environmental variable to the high-yield distribution of soybeans; the R software installation package is used to increase the control frequency of the model from 1 to 3 by 0.5 every time, and cross-combination verification is performed with the feature combination, and a parameter combination with significance and △AICc equal to 0 is selected.
[0022] Optionally, the process of drawing the spatial distribution map of soybean high-yield suitability includes: using the area value under the subject characteristic working curve to evaluate the model accuracy; inputting the model calculation results into the GIS platform, and dividing the distribution of soybean high yield into four levels: unsuitable, low suitability, medium suitability, and high suitability according to the probability of occurrence through reclassification and natural break point classification method.
[0023] Technical effect of the present invention: The present invention discloses a method for evaluating suitable areas for soybean planting that integrates yield prediction, a simplified yield estimation model, and a random forest machine learning method combined with vegetation index and known crop growth rules, making full use of different vegetation indices and their change rates to characterize yield, which has good practical effects. The constructed MaxEnt model is optimized, and the degree of fit between the model and the data is evaluated based on the modified AICc through the regularization multiplier RM and the feature combination FC parameters, with priority given to the model with the smallest AIC value. In addition, the application of this model is limited to recording the occurrence points of species, and in the adaptability assessment of soybeans, the area may be suitable for soybean growth, but it cannot be guaranteed that the yield of soybeans in this area is high. The present invention provides a more convenient way to identify suitable growth areas with high soybean yields. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of this application. The exemplary embodiments and descriptions of this application are intended to explain this application and do not constitute an improper limitation on this application. In the accompanying drawings:
[0025] Figure 1 This is a flow chart of a method for evaluating suitable soybean planting areas integrated with yield prediction according to an embodiment of the present invention;
[0026] Figure 2 This is a spatial diagram of soybean yield simulation based on the optimal model combination according to an embodiment of the present invention;
[0027] Figure 3 This is the accuracy verification of the MaxEnt model in the embodiment of the present invention;
[0028] Figure 4Response curves of the dominant environmental factors of the embodiment of the present invention, where (a) is the soybean high yield distribution probability corresponding to the average temperature in the wettest season, and (b) is the soybean high yield distribution probability corresponding to the precipitation in the wettest month;
[0029] Figure 5 This is the adaptive distribution of high-yield soybeans in the embodiments of the present invention. DETAILED DESCRIPTION
[0030] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0031] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0032] like Figure 1 As shown, this embodiment provides a method for evaluating suitable areas for soybean planting integrated with yield prediction, including:
[0033] S1. Collect regional data related to crop planting;
[0034] S2. Based on phenology, the soybean growing season is divided into seven typical periods. A multi-temporal remote sensing vegetation index state and change trend factor set is constructed for the soybean planting area. Two combinations are input into the machine learning model, with yield as the predictor variable. Model selection is carried out through accuracy evaluation to obtain a grid-scale spatial distribution map of soybean yield.
[0035] S3. Using the high-yield grid units identified in stage S2 as “existence point” data, a maximum entropy model was constructed, and multi-source environmental variables were integrated to conduct suitability modeling analysis, draw a spatial distribution map of soybean high-yield suitability, and identify potential advantageous planting areas.
[0036] The following steps are involved:
[0037] In step S1, regional data are collected, including (1) soybean yield data: based on the long-term statistical yearbook, soybean yield data of districts and counties are obtained, which can also be soybean yield of sampling points, and the number of sample points must be more than 400; (2) bioclimatic variables: including bioclimatic data representing annual trends (e.g., annual average temperature, annual precipitation), seasonality (e.g., annual range of temperature and precipitation), and extreme or limiting environmental factors (e.g., temperature of the coldest and hottest months, and precipitation in the wet and dry seasons); (3) soil parameters are obtained, including soil bulk density, total nitrogen content, total phosphorus content, cation exchange capacity of soil, calcium carbonate content, organic matter, pH value, exchangeable Ca 2+ etc.; (4) obtain terrain data, including altitude, slope, aspect, terrain relief, and distance to water bodies such as rivers and reservoirs; (5) collect data related to human activities, including distance to impervious surfaces, GDP at grid scale, population, etc.; (6) obtain a spatial distribution map of regional soybean planting.
[0038] Preferably, in step S1, points (1)-(4) in the collected data are used as pre-selected input factors of the MaxEnt model.
[0039] In step S2, soybean yield estimation is performed at a grid scale using a machine learning method based on remote sensing data and county yield data. Step S2 further includes the following steps:
[0040] Divide soybean growth period and calculate vegetation index and change trend:
[0041] To better explain the impact of predictors on crop yield at different growth stages, the soybean growing season was divided into seven major growth stages based on data from agricultural meteorological stations. For Heilongjiang Province, these stages are: T1: sowing (April 23-May 9); T2: seedling emergence (May 10-June 5); T3: three-leaf stage (June 6-24); T4: flowering (June 25-July 13); T5: pod setting (July 14-August 12); T6: grain filling (August 13-September 10); and T7: maturity (September 11-30).
[0042] The MODIS land level 3 standard data product, surface reflectance product (MOD09A1), was used as the main data source. The Normalized Difference Vegetation Index (NDVI), Enhanced Vegetation Index (EVI), Land Surface Water Index (LSWI), and Green Vegetation Index (VIgreen) were extracted from the Google Earth Engine (GEE) platform. The average values and Theil-Sen slopes of vegetation indices at different growth stages were calculated.
[0043] (1);
[0044] (2);
[0045] (3);
[0046] (4);
[0047] (5);
[0048] Where, It is the red band (0.620-0.670μm), It is the near-infrared band (0.841-0.876μm), It is the blue band (0.459-0.479μm), It is the green band (0.545-0.565μm), It is the shortwave infrared band (1.628-1.625μm). is the vegetation index at time and value. is an estimate of Sen's slope, if , it means that the series is on an upward trend during the growth period. On the contrary, if , it means that the data shows a downward trend during the growth period.
[0049] S22. Invert soybean yield at grid scale:
[0050] The county yield was correlated with the average value and sen slope of the vegetation index of the seven growth periods, and the variables that were not significantly correlated with soybean yield (P>0.001) were excluded. Two different combinations of input machine learning were used to analyze the impact of various factors on crop yield. Machine learning included support vector machine classifier (SVM), random forest (RF), deep neural network (DNN), multi-layer perceptron model (MLP), and extreme gradient boosting (XGBoost). The first input model M1 used county yield as the prediction target and vegetation index of different growth periods with high correlation as predictors. In the second input model M2, yield was again used as the target variable. However, in addition to the vegetation index with high correlation with yield, the predictor of this model also included the sen slope of vegetation index of different growth periods with high correlation with yield. The model runs used 70% of the total data for training and 30% for verification, covering the entire study area. The determination coefficient ( R ²) and the root mean square error ( RMSE ) was used to evaluate model performance. The model with the best performance was selected as the prediction model to obtain the spatial distribution of soybean yield at the grid scale.
[0051] (6);
[0052] (7);
[0053] in, County The output, County The predicted production, is the mean of the observed values, is the number of counties.
[0054] In step S3, S31, collect soybean high-yield planting records and environmental variables:
[0055] Environmental variables are crucial parameters for constructing species distribution models. Based on the principles of scientificity, representativeness, and accessibility, we selected 35 environmental variables, including the aforementioned bioclimate, soil, and topography variables, as well as those related to human activities. To avoid overfitting of the model due to multicollinearity among environmental variables, we performed a Pearson correlation analysis using the cor function in R software. After eliminating environmental variables with correlation coefficients greater than 0.8, we selected the final set of environmental variables for model construction.
[0056] For soybean yields above 2000 kg / ha, the sample points were selected. In order to eliminate the duplication, the ‘cellFromXY’ in the R software installation package ‘raster’ and the ‘duplicated’ in ‘data.table’ were combined to retain the distribution point closest to the center point in each grid (2 km × 2 km). Finally, the effective distribution points were obtained. The planting records and environmental variables were converted into CSV format.
[0057] S32, Mode Parameterization:
[0058] High-yield soybean distribution data and environmental variables under current climate conditions were imported into MaxEnt 3.4.4 software for modeling. 75% of the data points were used for modeling, and 25% for model validation. This was repeated 20 times, with the average of these 20 runs serving as the final result. Background data consisted of 1000 distribution points. The jackknife method was used to examine the contribution of each environmental variable to the distribution of high soybean yields, and the logistic output was .asc. Training gains were calculated for different environmental variables to examine their importance in constraining the geographic distribution of high soybean yields.
[0059] Because the MaxEnt model is sensitive to sampling bias and model complexity is significantly correlated with feature combination (FC) and regularization multiplier (RM), to avoid the impact of overfitting on model transferability, the 'ENMevaluate' function in the R software installation package "ENMeval" was used to cross-validate the RM of the model with FC (linear-L, quadratic-Q, hinge-H, product-P, and feature combinations LQ, LP, QP, and LQP) starting from 1 and increasing by 0.5. The optimal parameter combination with a significant ΔAICc (the difference between the AICc of the model and the AICc of the model with the lowest AICc) equal to 0 was selected to run the final model.
[0060] S33. Accuracy verification and classification of habitat suitability:
[0061] During crop distribution simulation, the area under the receiver operating characteristic curve (AUC) is used to evaluate model accuracy. A larger AUC value indicates higher model confidence. It is generally considered that an AUC of 0.5-0.7 indicates average predictive ability, 0.7-0.9 indicates good predictive ability, and 0.9-1.0 indicates excellent predictive ability.
[0062] The results obtained from the MaxEnt model were substituted into Arcgis 10.8. Through reclassification and natural break point classification, the distribution of high soybean yield was divided into four levels according to the probability of occurrence, namely unsuitable, low suitability, moderate suitability, and high suitability.
[0063] like Figure 2 In a certain region, the random forest model and M2 combination were used to invert the soybean yield at the grid scale. The regional soybean yield showed significant spatial heterogeneity and certain latitudinal differences. The soybean yield at low latitudes was slightly higher than that at high latitudes.
[0064] Figure 3 The distribution data and environmental variables for the selected high-yield soybeans were imported into MaxEnt software. The initial model parameters were: logistic output format, random seed selected, 25% random test percentage, and 20 replicates. The MaxEnt model with LQHPT selected and RM = 0.5 had an AUC of 0.876, indicating good simulation accuracy for the regional model.
[0065] Figure 4Among them, temperature (average temperature in the wettest season) and precipitation (precipitation in the wettest month) are the most important meteorological factors restricting the high-yield distribution of soybeans. The suitable precipitation range for soybeans in the wettest month is 95.4-155.3 mm. Too high or too low will cause a decrease in soybean yield. The wettest month is July, which corresponds to the flowering and podding period of soybeans. The flowering stage is the most sensitive stage to water deficit. When water stress occurs during the flowering period, the shedding of flowers and pods will increase, thereby reducing the yield. Figure 4 (a) is the probability of high soybean yield distribution corresponding to the average temperature in the wettest season, Figure 4 (b) is the probability of high soybean yield distribution corresponding to the precipitation in the wettest month.
[0066] Figure 5 Using reclassification and natural breakpoint classification, the distribution of high-yield soybeans was divided into four levels based on the probability of occurrence: unsuitable (0, 0.16), low suitable (0.16, 0.32), moderately suitable (0.32, 0.52), and highly suitable (0.52, 0.83). Currently, the suitable area is mainly distributed in the southwest of the region.
[0067] like Figure 5 As shown, this paper combines remote sensing data, phenological information, machine learning, and the MaxEnt model to develop a method for assessing suitable areas for soybean cultivation. This method can address the misalignment between resource endowment and crop distribution, which can lead to yield reduction and waste of agricultural resources. Constructing an adaptability assessment system based on soybean productivity can help reveal the spatial distribution of high-yield soybean cultivation and promote the establishment of a soybean crop planting layout tailored to local conditions.
[0068] The above are merely preferred embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A method for evaluating suitable areas for soybean planting integrated with yield prediction, characterized in that: include: Collect soybean planting spatial layout and yield data, MOD09A1 surface reflectance products during the growth period, bioclimatic data, and soil data; The soybean growing season was divided into seven typical periods based on phenology. A multi-temporal remote sensing vegetation index state and change trend factor set was constructed for soybean planting areas. Two combinations were used to input into a machine learning model. Yield was used as the predictor variable, and the model was selected through accuracy evaluation to obtain a grid-scale spatial distribution map of soybean yield. Using the identified high-yield grid cells as point data, a maximum entropy model was constructed, and multi-source environmental variables were integrated to conduct suitability modeling analysis. Through variable contribution analysis and response curves, the key ecological and environmental factors affecting the distribution of soybean high-yield suitability were identified, and a spatial distribution map of soybean high-yield suitability was drawn. The data collected include: soybean yield data, bioclimatic variables, soil variables, topographic variables and variables related to human activities; The soybean yield data is obtained based on the long-term statistical yearbook to obtain the soybean yield data of the district or county or the soybean yield of the sampling point; The bioclimatic variables include bioclimatic data representing annual trends, seasonality, and extreme or limiting environmental factors; The soil variables include soil bulk density, total nitrogen content, total phosphorus content, soil cation exchange capacity, calcium carbonate content, organic matter, pH value, exchangeable Ca 2+ ; The terrain variables include altitude, slope, aspect, terrain relief, and distance to rivers and reservoirs; The variables related to human activities include distance to impervious surface, GDP at grid scale, and population; The seven typical periods include sowing period, seedling emergence period, three-leaf period, flowering period, pod setting period, grain filling period and maturity period; The process of constructing a multi-temporal remote sensing vegetation index state and change trend factor set includes: extracting the normalized difference vegetation index, enhanced vegetation index, surface moisture index, and Green vegetation index using the MODIS land level 3 standard data product surface reflectance product; and calculating the average value and Theil-Sen slope of the vegetation index in different growth stages respectively; The process of inputting the machine learning model includes: Correlation analysis was conducted between county yield and the average and slope values of vegetation indices during the seven growth periods, excluding variables that were not significantly correlated with soybean yield. Two different combinations of inputs were used into the machine learning model. The first combination used county yield as the prediction target and the vegetation indices with high correlations at different growth stages as predictors. The second combination included the slopes of the vegetation indices with high correlations with yield in addition to the vegetation indices with high correlations at different growth stages. The process of constructing the maximum entropy model includes: selecting grids with soybean yields higher than 2000 kg / ha as sample points; performing Pearson correlation analysis on environmental variables and eliminating environmental variables with correlation coefficients greater than 0.8; converting planting records and environmental variables into CSV format and inputting them into the model; The parameterization process of the maximum entropy model includes: 75% of the distribution point data are used for modeling, 25% are used for model testing, and the operation is repeated 20 times; 1000 distribution points are selected as the background data; the contribution rate of each environmental variable to the high-yield distribution of soybean is tested by the jackknife method; the control frequency of the model is increased from 1 to 3 by 0.5 using the R software installation package, and cross-combination verification is performed with the feature combination, and the parameter combination with significance and △AICc equal to 0 is selected.
2. The soybean planting suitable area assessment method integrated with yield prediction according to claim 1, characterized in that: The process of drawing the spatial distribution map of soybean high-yield suitability includes: using the area value under the subject characteristic working curve to evaluate the model accuracy; inputting the model calculation results into the GIS platform, and dividing the distribution of soybean high-yield into four levels of unsuitable, low suitability, medium suitability, and high suitability according to the probability of occurrence through reclassification and natural break point classification method.
Citation Information
Patent Citations
Drought assessment method considering phenological characteristics of vegetation
CN119558536A
Method for rapidly searching suitable area of garcinia pauciflora
CN119886519A
Geographically identified agricultural product suitability evaluation method, device, equipment and medium
CN119940953A