Soybean planting suitable area evaluation method fused with yield prediction

Through machine learning and remote sensing technology combined with maximum entropy model, we can identify suitable areas for high yields of soybeans, solve the problem of insufficient assessment of soybean planting and increasing yield potential in the existing technology, and achieve efficient yield prediction and optimized resource allocation.

CN120355272AActive Publication Date: 2025-07-22INST OF AGRI RESOURCES & REGIONAL PLANNING CHINESE ACADEMY OF AGRI SCI

Patent Information

Application Number
CN202510855803.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-07-22
Estimated Expiration
2045-06-25

AI Technical Summary

Technical Problem

The existing technology has failed to effectively evaluate the yield-increasing potential of soybean cultivation, and lacks quantitative analysis on the impact of ecological and environmental factors, resulting in waste of resources and reduced yields.

Method used

The machine learning model is used to combine remote sensing vegetation index and multi-source environmental variables to identify the suitable areas for high-yield soybeans through time-series division and maximum entropy model, and a spatial distribution map for high-yield soybeans are constructed.

Benefits of technology

Identifying suitable growth areas with high yields of soybeans simplifies the yield estimation model, improves the adaptive evaluation efficiency of soybean cultivation, reduces resource waste, and improves the yield prediction accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355272A_ABST
    Figure CN120355272A_ABST
Patent Text Reader

Abstract

The invention discloses a soybean planting suitable area evaluation method fused with yield prediction, and relates to the technical field of data processing, and the method comprises the steps: collecting soybean planting related data; the method comprises the following steps: carrying out time sequence division on soybean growth seasons based on phenology, refining the whole growth process into seven typical periods, constructing a multi-temporal remote sensing vegetation index state and change trend factor set in a soybean planting area, respectively inputting two combination forms into a machine learning model, taking the yield as a predictive variable, and carrying out model selection through precision evaluation. Obtaining a grid-scale soybean yield spatial distribution diagram; and constructing a maximum entropy model by taking the identified high-yield grid units as existence point data, carrying out suitability modeling analysis by integrating multi-source environment variables, identifying key ecological environment factors influencing soybean high-yield suitability distribution through variable contribution analysis and a response curve, and drawing a soybean high-yield suitability spatial distribution diagram. According to the invention, the soybean high-yield suitable growth area is identified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data processing, and particularly relates to a method for evaluating suitable areas for soybean planting by integrating yield prediction. Background Art

[0002] Soybeans, as important food and oil crops, are also important nitrogen-fixing plants and are relatively friendly to the environment. However, China faces a serious dependence on imports. Although the country is promoting the expansion of soybean planting to increase the planting area and yield, the current technology has not comprehensively analyzed the adaptability of soybean planting based on the yield increase potential.

[0003] Since the growth of crops is selective to ecological environment factors such as climate, soil, and terrain, crop suitability evaluation is used to assess whether a specific land area is suitable for crop planting. Upgrading the impact of agricultural resource and environmental factors on crop layout from traditional qualitative evaluation to quantitative evaluation also provides a new idea for clarifying the response process of crop spatial pattern to climate change. By comprehensively analyzing the agricultural resource and environmental conditions on which crops depend and their matching characteristics, the advantages and disadvantages of the impact of resource and environmental conditions on crop planting can be effectively evaluated. Summary of the Invention

[0004] To solve the above technical problems, the present invention proposes a method for evaluating suitable areas for soybean planting by integrating yield prediction, and identifies suitable growth areas for high-yield soybeans.

[0005] To achieve the above object, the present invention provides a method for evaluating suitable areas for soybean planting by integrating yield prediction, including: Collecting soybean planting spatial layout and yield data, MOD09A1 surface reflectance products during the growth period, bioclimatic data, and soil data; Based on phenology, the soybean growth season is divided into time series, the entire growth process is refined into seven typical periods, a factor set of multi-temporal remote sensing vegetation index status and change trends in the soybean planting area is constructed, and two combination forms are respectively input into a machine learning model. Taking the yield as the prediction variable, the model is selected through accuracy evaluation to obtain a spatial distribution map of soybean yield at the grid scale; Using the identified high-yield grid cells as presence point data, a maximum entropy model is constructed, and multi-source environmental variables are integrated to carry out suitability modeling analysis. Through variable contribution degree analysis and response curves, the key ecological environment factors affecting the distribution of high-yield suitability of soybeans are identified, and a spatial distribution map of high-yield suitability of soybeans is drawn.

[0006] Optionally, the collected data includes: soybean yield data, bioclimatic variables, soil variables, terrain variables, and variables related to human activities; The soybean yield data is obtained based on the long-term statistical yearbook for the soybean yield data of districts and counties or the soybean yield at sampling points. The bioclimatic variables include bioclimatic data representing annual trends, seasonality, and extreme or limiting environmental factors. The soil variables include soil bulk density, total nitrogen content, total phosphorus content, cation exchange capacity of the soil, calcium carbonate content, organic matter, pH value, exchangeable Ca 2+ ; The terrain variables include altitude, slope, aspect, terrain relief, and distance to rivers and reservoirs. The variables related to human activities include distance to impervious surfaces, GDP at the grid scale, and population.

[0007] Optionally, the seven typical periods include the sowing period, emergence period, trifoliate stage, flowering stage, pod-setting stage, grain-filling stage, and maturity stage.

[0008] Optionally, the process of constructing the multi-temporal remote sensing vegetation index status and change trend factor set includes: using the MODIS land level-3 standard data product surface reflectance product to extract the normalized difference vegetation index, enhanced vegetation index, land surface water index, and Green vegetation index; calculating the average value and Theil-Sen slope of the vegetation index at different growth stages respectively.

[0009] Optionally, the process of inputting into the machine learning model includes: Performing a correlation analysis on the county-level yield with the average value and slope of the vegetation index at seven growth stages, and excluding variables that are not significantly correlated with the soybean yield; Using two different combinations to input into the machine learning model. The first combination uses the county-level yield as the prediction target and the vegetation indices at different growth stages with high correlation significance as the prediction factors. The second combination includes, in addition to the vegetation indices with high correlation significance, the slopes of the vegetation indices at different growth stages with high correlation with the yield.

[0010] Optionally, the process of constructing the maximum entropy model includes: screening the grids with a soybean yield higher than 2000 kg / ha as sample points; performing a Pearson correlation analysis on the environmental variables and excluding the environmental variables with a correlation coefficient greater than 0.8; converting the planting records and environmental variables into CSV format and inputting them into the model.

[0011] Optionally, the parameterization process of the maximum entropy model includes: 75% of the distribution point data is used for modeling, 25% is used for model verification, and the process is repeated 20 times; 1000 distribution points are selected as background data; the jackknife method is used to test the contribution rate of each environmental variable to the high-yield distribution of soybeans; the R software installation package is used to cross-validate the regulation multiple frequency of the model starting from 1 and increasing by 0.5 up to 3 with the feature combination, and the parameter combination with significance and △AICc equal to 0 is selected.

[0012] Optionally, the process of drawing the spatial distribution map of soybean high-yield suitability includes: using the area value under the receiver operating characteristic curve to evaluate the model accuracy; inputting the model operation results into the GIS platform, and through reclassification and natural breakpoint grading method, the distribution of high-yield soybeans is divided into four grades: unsuitable, low suitable, medium suitable, and high suitable according to the occurrence probability.

[0013] Technical effects of the present invention: The present invention discloses an evaluation method for soybean planting suitable areas integrating yield prediction, a simplified yield estimation model, which uses the random forest machine learning method combined with vegetation indices and known crop growth rules, and fully utilizes different vegetation indices and their change rates to characterize yield, having good practical effects. The constructed MaxEnt model is optimized, and the fitting degree of the model and data is evaluated based on the corrected AICc through the regularization multiplier RM and the feature combination FC parameters, and the model with the smallest AIC value is given priority. In addition, the application of this model is limited to recording the sample points where the species appear. In the evaluation of soybean adaptability, this area may be suitable for soybean growth, but it does not guarantee high soybean yield in this area. The present invention provides a relatively convenient way to identify the suitable growth areas for high-yield soybeans. Description of the Drawings

[0014] The drawings forming a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation of this application. In the drawings: Figure 1 It is a schematic flowchart of an evaluation method for soybean planting suitable areas integrating yield prediction according to an embodiment of the present invention; Figure 2 It is a spatial map of soybean single yield simulation based on the optimal model combination according to an embodiment of the present invention; Figure 3 It is the accuracy verification of the MaxEnt model according to an embodiment of the present invention; Figure 4 It is the response curve of the dominant environmental factors according to an embodiment of the present invention, where (a) is the high-yield distribution probability of soybeans corresponding to the average temperature in the wettest season, and (b) is the high-yield distribution probability of soybeans corresponding to the precipitation in the wettest month; Figure 5This is the adaptive distribution of high soybean yield in the embodiments of the present invention. Detailed implementation manners

[0015] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will describe this application in detail with reference to the drawings and in combination with the embodiments.

[0016] It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0017] As Figure 1 shown, in this embodiment, a method for evaluating the suitable area for soybean planting by integrating yield prediction is provided, including: S1. Collect regional data related to crop planting; S2. Divide the soybean growth season into seven typical periods based on phenology, construct a multi-temporal remote sensing vegetation index status and change trend factor set for the soybean planting area, input it into the machine learning model in two combination forms respectively, use the yield as the prediction variable, select the model through accuracy evaluation, and obtain the spatial distribution map of soybean yield at the grid scale; S3. Use the high-yield grid cells identified in the S2 stage as the "presence point" data, construct the maximum entropy model, and integrate multi-source environmental variables to carry out suitability modeling analysis, draw the spatial distribution map of soybean high-yield suitability, and clarify the potential advantageous planting areas.

[0018] It includes the following steps: In step S1, collect regional related data, including (1) soybean yield data: based on the long-term statistical yearbook, obtain the soybean yield data of the district or county, or it can also be the soybean yield of the sampling points, and it is necessary to ensure that the number of sample points exceeds 400; (2) bioclimatic variables: including bioclimatic data representing annual trends (such as annual average temperature, annual precipitation), seasonality (such as annual range of temperature and precipitation), and extreme or limiting environmental factors (such as temperature in the coldest and hottest months, and precipitation in the wet and dry quarters); (3) obtain soil parameters, including soil bulk density, total nitrogen content, total phosphorus content, cation exchange capacity of the soil, calcium carbonate content, organic matter, pH value, exchangeable Ca 2+ etc.; (4) obtain terrain data, including altitude, slope, aspect, terrain undulation, distance to water bodies such as rivers and reservoirs; (5) collect data related to human activities, including distance to impervious surfaces, GDP at the grid scale, population, etc.; (6) obtain the spatial distribution map of regional soybean planting.

[0019] Preferably, in step S1, the data points (1)-(4) in the collected data are used as the input factors for the preselected MaxEnt model.

[0020] In step S2, based on the remote sensing data and the county-level yield data, machine learning methods are used for soybean yield estimation at the raster scale. Step S2 further includes the following steps: Divide the soybean growth period and calculate the vegetation index and its change trend: To better explain the impact of the prediction factors on crop yield at different growth stages, according to the agricultural meteorological station data, the soybean growth season is divided into seven main growth periods. Taking Heilongjiang Province as an example, they are T1: sowing period (April 23 - May 9); T2: emergence period (May 10 - June 5); T3: trifoliate stage (June 6 - June 24); T4: flowering period (June 25 - July 13); T5: pod-setting stage (July 14 - August 12); T6: grain-filling stage (August 13 - September 10); T7: maturity stage (September 11 - September 30).

[0021] Using the MODIS Level 3 standard data product for land - surface reflectance product (MOD09A1) as the main data source, extract the Normalized Difference Vegetation Index (NDVI), Enhanced Vegetation Index (EVI), Land Surface Water Index (LSWI), and Green Vegetation Index (VIgreen) from the Google Earth Engine (GEE) platform, and calculate the average value and Theil - Sen slope (slope) of the vegetation index for different growth periods respectively: (1); (2); (3); (4); (5); In the formula, is the red band (0.620 - 0.670μm), is the near - infrared band (0.841 - 0.876μm), is the blue band (0.459 - 0.479μm), is the green band (0.545 - 0.565μm), is the short - wave infrared band (1.628 - 1.625μm). is the value of the vegetation index at time and respectively. is the estimated value of the Sen slope. If , it indicates that the sequence shows an upward trend during the growth period. On the contrary, if , it indicates that the data shows a downward trend during the growth period.

[0022] S22. Inversion of soybean yield at the grid scale: The correlation analysis was carried out on the county-level yield, the average value of the vegetation indices at 7 growth stages, and the sen slope, and the variables that were not significantly correlated with the soybean yield (P>0.001) were excluded. Two different combinations were used as input for machine learning to analyze the influence of various factors on crop yield. The machine learning methods included support vector machine classifier (SVM), random forest (RF), deep neural network (DNN), multi-layer perceptron model (MLP), and extreme gradient boosting (XGBoost). In the first input model M1, the county-level yield was used as the prediction target, and the vegetation indices at different growth stages with high correlation significance were used as predictors. In the second input model M2, the yield was again used as the target variable. However, in addition to the vegetation indices with high correlation significance with the yield, the sen slopes of the vegetation indices at different growth stages with high correlation significance with the yield were also included as predictors in this model. 70% of the total data was used for training and 30% for validation during the model operation, covering the entire study area. The coefficient of determination ( R ²) and the root mean square error ( RMSE ) were used to evaluate the performance of the model. The model with good performance was selected as the prediction model to obtain the spatial distribution of soybean yield at the grid scale.

[0023] (6); (7); Among them, is the yield of county , is the predicted yield of county , is the average value of the observed values, is the number of counties.

[0024] In step S3, S31. Collect soybean high-yield planting records and environmental variables: Environmental variables are important parameters for constructing species distribution models. Based on the principles of scientificity, representativeness, and accessibility, 35 environmental variables were selected, including the above-mentioned bioclimatic, soil, terrain, and variables related to human activities. To avoid overfitting of the model caused by multicollinearity among environmental variables, the Pearson correlation analysis was carried out on the environmental variables using the cor function in R software. After excluding the environmental variables with a correlation coefficient greater than 0.8, the environmental variables were finally selected for constructing the model.

[0025] Regarding the sample points with soybean yields higher than 2000 kg / ha, in order to eliminate the 'cellFromXY' in the 'raster' package of the R software installation package and the 'duplicated' in the 'data.table', so that the nearest distribution point to the center point is retained within each grid (2 km × 2 km), and finally the effective distribution points are obtained, and the planting records and environmental variables are converted into CSV format.

[0026] S32. Model parameterization: Import the distribution data of high-yield soybeans and environmental variables under the current climate into MaxEnt 3.4.4 software for modeling operations. 75% of the distribution point data is used for modeling, and 25% is used for model verification. It is run 20 times repeatedly, and the average result of the 20 operations is used as the final result. The background data is selected as 1000 distribution points, and the Jackknife method is used to test the contribution rate of each environmental variable to the high-yield distribution of soybeans, and the Logistic output is in.asc format. Calculate the test gain (training gain) of different environmental variables to detect the importance of environmental variables restricting the geographical distribution of high-yield soybeans.

[0027] Since the MaxEnt model is sensitive to sampling bias and the complexity of the model is significantly correlated with the Feature Combination (FC) and the Regularization Multiplier (RM), to avoid the impact of overfitting on the model migration ability, the 'ENMevaluate' function in the 'ENMeval' package of the R software installation package is used to cross-combine and verify the RM of the model from 1 incrementing upward by 0.5 to 3 and the FC (linear - L, quadratic - Q, hinge - H, product - P, and the feature combinations LQ, LP, QP, LQP) respectively. Finally, the optimal parameter combination with significance, where ΔAICc (the difference between the AICc of this model and the AICc of the lowest AICc model) is equal to 0, is selected to run the final model.

[0028] S33. Accuracy verification and habitat suitability division: During the crop distribution simulation process, the area under the receiver operating characteristic curve (AUC) value is used to evaluate the model accuracy. The larger the AUC value, the higher the credibility of the model. Generally, it is considered that when the AUC value is between 0.5 and 0.7, the model prediction ability is average; when it is between 0.7 and 0.9, the model prediction ability is good; when it is between 0.9 and 1.0, the model prediction ability is excellent.

[0029] Substitute the results obtained from the MaxEnt model operation into ArcGIS 10.8. Through reclassification and the natural breaks classification method, according to the occurrence probability, the distribution of high soybean yields is divided into four grades, namely unsuitable, low suitability, medium suitability, and high suitability.

[0030] For example Figure 2 , a certain area used the random forest model and the M2 combination to invert soybean yields at the raster scale. The regional soybean yields showed significant spatial heterogeneity characteristics and certain latitude differences. The soybean yields in the low-latitude areas were slightly higher than those in the high-latitude areas.

[0031] Figure 3 In, the screened distribution data of high soybean yields and environmental variables were imported into the MaxEnt software. The initial model setting parameters were: the output format was Logistic, the random seed was checked, the random test percentage was 25%, and the number of repetitions was 20 times. For the MaxEnt model with LQHPT and RM = 0.5, the AUC was 0.876, indicating that the model in the region had good simulation accuracy.

[0032] Figure 4 In, temperature (average temperature in the wettest season) and precipitation (precipitation in the wettest month) are the most important meteorological factors restricting the distribution of high soybean yields. The suitable range of precipitation in the wettest month for soybeans is 95.4 - 155.3 mm. Too high or too low will cause a decrease in soybean yields. The wettest month is July, corresponding to the flowering and pod-setting period of soybeans. The flowering stage is the most sensitive stage to water deficit. When water stress occurs during the flowering period, it will increase the abscission of flowers and pods, thus reducing yields. Figure 4 Figure (a) shows the probability of high soybean yields corresponding to the average temperature in the wettest season. Figure 4 Figure (b) shows the probability of high soybean yields corresponding to the precipitation in the wettest month.

[0033] Figure 5 In, through reclassification and the natural breaks classification method, according to the occurrence probability, the distribution of high soybean yields is divided into four grades, namely unsuitable [0, 0.16), low suitability [0.16, 0.32), medium suitability [0.32, 0.52), and high suitability [0.52, 0.83]. The current suitable areas are mainly distributed in the southwestern part of the region.

[0034] For example Figure 5 As shown, the present invention couples remote sensing data, phenological information, machine learning, and the MaxEnt model to develop an evaluation method for suitable areas for soybean planting. This method can focus on the mismatch between resource endowments and crop layouts, resulting in reduced yields and wasted agricultural resources. Constructing an adaptability evaluation system based on soybean productivity can help reveal the high-yield spatial distribution of soybean planting and promote the establishment of a soybean crop planting layout adapted to local conditions.

[0035] The above are only the preferred specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for evaluating suitable areas for soybean cultivation by integrating yield prediction, characterized in that, Including: Collecting soybean planting spatial layout and yield data, MOD09A1 surface reflectance products during the growth period, bioclimatic data, and soil data; Dividing the soybean growth season into time series based on phenology, refining the entire growth process into seven typical periods, constructing a multi-temporal remote sensing vegetation index status and change trend factor set for the soybean planting area, inputting it into a machine learning model in two combination forms respectively, using yield as the prediction variable, and selecting the model through accuracy evaluation to obtain the soybean yield spatial distribution map at the grid scale; Taking the identified high-yield grid cells as presence point data, constructing a maximum entropy model, and integrating multi-source environmental variables to carry out suitability modeling analysis. Through variable contribution analysis and response curves, identifying the key ecological environment factors affecting the high-yield suitability distribution of soybeans, and drawing the high-yield suitability spatial distribution map of soybeans.

2. The soybean planting suitable area evaluation method integrating yield prediction according to claim 1, wherein The collected data includes: soybean yield data, bioclimatic variables, soil variables, topographic variables, and variables related to human activities; The soybean yield data is based on the long-term statistical yearbook to obtain the soybean yield data of the county or the soybean yield of the sampling points; The bioclimatic variables include bioclimatic data representing annual trends, seasonality, and extreme or restrictive environmental factors; The soil variables include soil bulk density, total nitrogen content, total phosphorus content, cation exchange capacity of the soil, calcium carbonate content, organic matter, pH value, exchangeable Ca 2+ ; The topographic variables include altitude, slope, aspect, terrain relief, and distance to rivers and reservoirs; The variables related to human activities include distance to impervious surfaces, GDP at the grid scale, and population.

3. The soybean planting suitable area evaluation method integrating yield prediction according to claim 1, wherein The seven typical periods include sowing period, emergence period, trifoliate period, flowering period, pod-setting period, grain-filling period, and maturity period.

4. The soybean planting suitable area evaluation method integrating yield prediction according to claim 1, characterized in that, The process of constructing the multi-temporal remote sensing vegetation index status and change trend factor set includes: using the MODIS land level-3 standard data product surface reflectance product to extract the normalized difference vegetation index, enhanced vegetation index, surface water index, and Green vegetation index; calculating the average value and Theil-Sen slope of the vegetation index at different growth periods respectively.

5. The soybean planting suitable area evaluation method integrating yield prediction according to claim 1, characterized in that The process of inputting into the machine learning model includes: Conducting a correlation analysis between the county yield and the average value and slope of the vegetation index in seven growth periods, and excluding variables that are not significantly correlated with the soybean yield; Inputting into the machine learning model in two different combinations. The first combination takes the county yield as the prediction target and the vegetation indices in different growth periods with high correlation significance as the prediction factors. The second combination includes, in addition to the vegetation indices with high correlation significance, the slopes of the vegetation indices in different growth periods with high correlation with the yield.

6. The soybean planting suitability zone evaluation method integrating yield prediction according to claim 1, characterized in that, The process of constructing the maximum entropy model includes: screening the grid cells with soybean yield higher than 2000 kg / ha as sample points; conducting a Pearson correlation analysis on the environmental variables and removing the environmental variables with a correlation coefficient greater than 0.8; converting the planting records and environmental variables into CSV format and inputting them into the model.

7. The soybean planting suitable area evaluation method integrating yield prediction according to claim 1, characterized in that The parameterization process of the maximum entropy model includes: 75% of the distribution point data is used for modeling, 25% is used for model testing, and the process is repeated 20 times; the background data is selected as 1000 distribution points; the jackknife method is used to test the contribution rate of each environmental variable to the high-yield distribution of soybeans; the R software installation package is used to cross-validate the regulatory frequency of the model starting from 1 and increasing by 0.5 up to 3 with the feature combinations, and the parameter combination with significance and △AICc equal to 0 is selected.

8. The soybean planting suitable area evaluation method integrating yield prediction according to claim 1, wherein, The process of drawing the spatial distribution map of high-yield suitability of soybeans includes: evaluating the model accuracy by using the area value under the receiver operating characteristic curve; inputting the model operation results into the GIS platform, and through reclassification and natural break classification method, the distribution of high-yield soybeans is divided into four grades: unsuitable, low suitable, medium suitable, and high suitable according to the occurrence probability.

Citation Information

Patent Citations

  • Multi-time-phase remote sensing inversion method of soybean biomass in black soil area by introducing topographic factor

    CN109325433A

  • Drought assessment method considering phenological characteristics of vegetation

    CN119558536A

  • Method for rapidly searching suitable area of garcinia pauciflora

    CN119886519A

  • Geographically identified agricultural product suitability evaluation method, device, equipment and medium

    CN119940953A

  • Combination play facility with irregular tree

    KR102631283B1

Cited By

  • Locust suitable growing area monitoring method, device and equipment, storage medium and program product

    CN121543034A

  • A method and related device for predicting suitable chestnut planting areas

    CN122509423A