Regional rice planting density recommendation method based on multi-source data
Through the combination of nonlinear model and machine learning, the problem of insufficient spatial applicability and accuracy of rice planting density construction methods is solved, ecological zone division and customized planting density recommendations are realized, and the accuracy and adaptability of rice planting density are improved.
Patent Information
- Application Number
- CN202510609920.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-08-29
AI Technical Summary
The existing rice planting density construction methods have problems such as poor spatial applicability, lack of quantification of genotype and environmental interaction effects, and insufficient analytical accuracy of multi-factor nonlinear relationships, making it difficult to achieve regional planting density recommendations with high precision and strong generalization capabilities.
By collecting rice experimental data, fitting the relationship between density and yield using nonlinear models, combining cluster analysis and machine learning models, integrating multi-source data such as meteorology and soil, we will build a regional rice planting density recommendation method, and conduct ecological zone division and customized planting density recommendation.
It realizes high-precision regionalized rice planting density recommendations, improves the adaptability and accuracy of the model, provides reliable agricultural decision-making support, and is suitable for customized planting density guidance in different ecological areas.
Smart Images

Figure CN120561633A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of agricultural production technology, and in particular to a regional rice planting density recommendation method based on multi-source data. Background Art
[0002] Food security is a core challenge to global sustainable development. As one of the world's three major staple crops, rice feeds more than half of the world's population, and its stable yield is directly linked to national food security strategies. Against the backdrop of limited arable land resources and intensifying climate change, optimizing rice planting density through precision agricultural management and creating high-yielding and stable populations has become a key breakthrough in increasing rice yields. Insufficient planting density can deplete land resources, while overly dense planting can lead to increased plant competition and increased pest and disease risks, ultimately limiting yield potential. Therefore, scientifically determining regional rice planting density is a key path to achieving the goal of increasing yields through density.
[0003] Currently, existing methods for constructing rice planting density face three limitations: poor spatial applicability, making it difficult to extrapolate results from small-scale experiments to large-scale regions; a lack of quantification and integration of genotype-environment interaction effects; and insufficient precision in analyzing multi-factor nonlinear relationships. Therefore, there is an urgent need to develop a regionalized rice planting density recommendation method based on multi-source data, incorporating multi-dimensional variables such as climate, soil, and crop variety, to construct a highly accurate and generalizable decision-making model to address these issues. Summary of the Invention
[0004] The purpose of the present invention is to provide a regional rice planting density recommendation method based on multi-source data to solve the problems raised in the prior art.
[0005] To achieve the above objectives, the present invention provides the following technical solution: a regional rice planting density recommendation method based on multi-source data, the recommendation method comprising:
[0006] Step S100: collecting rice test data from a rice test station;
[0007] Step S200: fitting the planting density and yield data through a nonlinear model to obtain a density-yield relationship curve, calculating the recommended planting density, and performing data cleaning on the calculated recommended planting density;
[0008] Step S300: Dividing the target area into rice-growing ecological zones using a cluster analysis algorithm, and extracting meteorological data and soil data for each rice-growing ecological zone;
[0009] Step S400: constructing a multi-source dataset based on the extracted meteorological data and soil data; dividing the multi-source dataset into a training set and a test set, and constructing a machine learning model;
[0010] Step S500: Input the test set into the machine learning model, calculate the correlation coefficient, root mean square error, and relative root mean square error to evaluate the model;
[0011] Step S600: Inputting the representative value of the rice-growing ecological zone into the evaluated machine learning model to predict the recommended planting density of the rice-growing ecological zone.
[0012] Furthermore, step S100 includes:
[0013] Collect rice test data from rice test stations, which are several test sites for rice cultivation; the rice test data includes test site, test year, crop variety, planting density, and yield; the collected rice test data is obtained in the form of articles or tables; if the rice test data includes graphical data, it is extracted using image digitization software;
[0014] The rice trial data in the above steps include information such as independent trial sites, trial years, crop varieties, planting density and yield, and should include at least three density gradients to ensure the accuracy of planting density calculation.
[0015] Furthermore, step S200 includes:
[0016] The collected rice test data are divided into independent tests according to the test site, test year, and crop variety, and the planting density and yield corresponding to each independent test are extracted; the planting density and yield corresponding to each independent test are fitted using a nonlinear model, and a density-yield relationship curve is output; the planting density corresponding to the highest yield in the density-yield relationship curve is defined as the recommended planting density; data with recommended planting density exceeding the test density range are eliminated, and the recommended planting density data of the remaining fitting results are counted as a recommended planting density data set; outlier removal and normality test are performed on the obtained recommended planting density data set to obtain recommended planting density data that conforms to the normal distribution;
[0017] In the above steps, the planting density and yield corresponding to each independent experiment are fitted using a nonlinear model, such as quadratic function fitting, linear plus platform fitting, etc., to output a curve showing the relationship between density and yield. By eliminating data that cannot reflect the actual recommended planting density, the recommended planting density is ensured to reflect the actual density response law. The data quality is ensured by removing outliers and performing normality tests on the obtained recommended planting density data set.
[0018] Furthermore, step S300 includes:
[0019] Step S301: collecting meteorological data of a target area over several years and soil data of the target area over several years or soil data of the current year; the target area is an area corresponding to the recommended rice planting density; selecting a corresponding resolution according to the area of the target area; if the area of the target area is greater than a threshold, selecting a first resolution; if the area of the target area is less than the threshold, selecting a second resolution; selecting a corresponding coordinate system according to the geographical location of the target area; unifying the soil data, meteorological data, and recommended planting density data into the same coordinate system, and resampling the soil data and meteorological data to the same resolution;
[0020] Due to the large interannual variations in meteorological data, multi-year average meteorological data is more representative of the ecological region; because soil properties are stable, multi-year data or current year data can be selected according to needs; the target area in the above steps can be a province, city, county, etc.; if the target area is at the provincial scale, the changes in soil types and weather patterns may be large, and a coarser resolution can capture these changes without being bothered by local details that may be irrelevant at this scale; if the target area is a smaller area such as a city or county, medium to high resolution is more suitable for capturing local changes related to planting density; in smaller areas, differences in soil composition and microclimate may be more localized and have a significant impact on the optimal planting density, and higher resolution data is required to distinguish these differences; the above coordinate system selection is mainly determined by the geographical location of the target area to ensure compatibility with available data sources and regional standards (for example, WGS84 is used for global applications and CGCS2000 is used for applications within China).
[0021] Step S302: Using meteorological factors of the target area as clustering features, the meteorological factors include accumulated temperature, annual precipitation, and solar radiation; clustering using a k-means clustering algorithm to divide a number of rice-growing ecological zones with the same environmental conditions; and calculating the average value of the meteorological factors of each divided rice-growing ecological zone as a representative value of each rice-growing ecological zone.
[0022] Step S303: extracting independent test data of the independent test sites from the coordinate system according to the collected independent test site coordinates; the independent test data of each independent test includes density-yield data, soil data of the rice planting ecological zone, and meteorological data of the rice planting ecological zone.
[0023] Furthermore, the soil data includes soil organic carbon, pH, total nitrogen, total phosphorus, total potassium, cation exchange capacity, soil texture type, bulk density and soil thickness; the meteorological data adopts the average meteorological data provided by TerraClimate, including maximum temperature, minimum temperature, solar radiation and precipitation.
[0024] Furthermore, step S400 includes:
[0025] Feature selection was performed on meteorological data and soil data through LASSO regression, including selecting organic carbon content, total nitrogen content, cation exchange capacity, and soil pH as soil factors, and selecting rainfall, accumulated temperature, solar radiation, minimum temperature during the growing period, and maximum temperature during the growing period as meteorological factors; the soil factors and meteorological factors were combined with the calculated recommended planting density and yield to form a multi-source data set, 80% of the multi-source data set was divided into a training set for training a machine learning model, and 20% of the multi-source data set was divided into a test set for evaluating the machine learning model.
[0026] Furthermore, meteorological and soil characteristic data are extracted from the meteorological and soil data through LASSO regression, and the characteristic data are subjected to a multicollinearity test, including: removing highly linearly correlated features from the characteristic data, using the remaining features as covariates, and constructing a machine learning model based on the covariates.
[0027] Furthermore, constructing a machine learning model based on the covariates includes: taking the obtained recommended planting density data that conforms to the normal distribution as the dependent variable, taking the covariates corresponding to the recommended planting density data as the independent variables, constructing a random forest model for planting density recommendation, and inputting the training set into the random forest model to train the model.
[0028] Furthermore, step S500 includes:
[0029] The formula for calculating the correlation coefficient is as follows:
[0030]
[0031] The formula for calculating the root mean square error is as follows:
[0032]
[0033] The formula for calculating the relative root mean square error is as follows:
[0034]
[0035] Among them, R 2 represents the correlation coefficient; RMSE represents the root mean square error; RRMSE represents the relative root mean square error; i represents the recommended planting density data of the i-th independent experiment in the test set; S i represents the model-predicted recommended planting density data for the i-th independent trial in the test set; Represents the average value of the recommended planting density data of all independent trials in the test set; It represents the average value of the model-predicted recommended planting density data of all independent trials in the test set, and n represents the number of independent trials contained in the test set.
[0036] Compared with the prior art, the present invention has the following beneficial effects:
[0037] 1. Ecological zone division and customized planting density recommendations: This invention divides the study area into ecological zones and makes targeted recommendations based on the meteorological and soil characteristics of each ecological zone, providing regionalized and customized rice planting density recommendations. This avoids the universal limitations of traditional methods and can provide the most suitable rice planting density for each region based on the specific environmental conditions. This is not only of great significance for large-scale agricultural production, but also can provide tailored planting density guidance for rice production in different ecological zones.
[0038] 2. Multi-source data fusion to comprehensively reflect influencing factors: This invention integrates multi-dimensional data such as meteorological, soil, and crop varieties, breaking through the limitation of traditional methods that rely solely on field test data. By comprehensively considering various environmental factors that affect rice planting density, it can more accurately reflect the complex situations in actual agricultural production. This data fusion method not only improves the accuracy of the model, but also enhances its ability to adapt to different environmental conditions.
[0039] 3. High-precision prediction to improve agricultural decision support: The established machine learning model R 2 Up to 0.6, RMSE ≤ 3.03×10 4 holes / hectare, which is more accurate than the linear model; by adopting advanced machine learning algorithms such as random forest, XGBoost and LightGBM, a large amount of multi-source data is learned and trained, which can effectively capture the complex nonlinear relationship in the data, thereby improving the accuracy of rice planting density recommendations and providing agricultural producers with a more reliable decision-making basis. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 This is a flow chart of a method for recommending regional rice planting density based on multi-source data according to the present invention;
[0041] Figure 2 A schematic diagram of the relationship between planting density and yield in a regional rice planting density recommendation method based on multi-source data of the present invention;
[0042] Figure 3 A spatial distribution map of sample points of a regional rice planting density recommendation method based on multi-source data according to the present invention;
[0043] Figure 4 A rice planting area division map for a regional rice planting density recommendation method based on multi-source data according to the present invention;
[0044] Figure 5 This is a comparative analysis chart of the predicted values and actual values of the recommended density values of independent experiments using a random forest model of a regionalized rice planting density recommendation method based on multi-source data in the present invention. DETAILED DESCRIPTION
[0045] Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative work shall fall within the scope of protection of the present invention.
[0046] Example: Figure 1-Figure 5 As shown, the present invention provides a technical solution, a regional rice planting density recommendation method based on multi-source data, the recommendation method comprising:
[0047] Step S100: collecting rice test data from a rice test station;
[0048] Wherein, step S100 includes:
[0049] Collect rice test data from rice test stations, which are several test sites for rice cultivation; the rice test data includes test site, test year, crop variety, planting density, and yield; the collected rice test data is obtained in the form of articles or tables; if the rice test data includes graphical data, it is extracted using image digitization software;
[0050] In an embodiment of the present invention, the collection rules for collecting rice test data include:
[0051] (1) The experiment should be conducted in the field. Data from greenhouse and pot experiments are not considered.
[0052] (2) more than three levels of plant density were evaluated in a given field trial;
[0053] (3) provide specific yield information, excluding model-simulated yields;
[0054] (4) Production data includes mean and standard deviation
[0055] (5) Information on the trial location, year, and variety;
[0056] (6) Stress and intercropping experiments are not taken into account;
[0057] Planting density data for the Northeast region were extracted, including specific test site location information, test year, and test variety. Data were mainly obtained from tables in the articles. For data displayed only in graphs, GetData GraphDigitizer software was used to extract the corresponding values.
[0058] Step S200: fitting the planting density and yield data through a nonlinear model to obtain a density-yield relationship curve, calculating the recommended planting density, and performing data cleaning on the calculated recommended planting density;
[0059] Wherein, step S200 includes:
[0060] The collected rice test data are divided into independent tests according to the test site, test year, and crop variety, and the planting density and yield corresponding to each independent test are extracted; the planting density and yield corresponding to each independent test are fitted using a nonlinear model, and a density-yield relationship curve is output; the planting density corresponding to the highest yield in the density-yield relationship curve is defined as the recommended planting density; data with recommended planting density exceeding the test density range are eliminated, and the recommended planting density data of the remaining fitting results are counted as a recommended planting density data set; outlier removal and normality test are performed on the obtained recommended planting density data set to obtain recommended planting density data that conforms to the normal distribution;
[0061] In the embodiment of the present invention, independent trials are divided according to the trial site location information, trial year, and trial variety. The density and yield data of each independent trial are fitted with a quadratic function, and the density corresponding to the maximum yield is used as the recommended planting density. The quadratic function fitting results used in this embodiment are as follows: Figure 2 As shown in the figure, the recommended planting density and yield data were processed, and the data with recommended planting density exceeding the experimental range were eliminated. The isolation forest algorithm was used to remove outliers to ensure data quality. The cleaned data were in line with the normal distribution. The final data used included 207 available independent experiments obtained from 150 literatures. The test point distribution results are shown in the figure. Figure 3 As shown;
[0062] Step S300: Dividing the target area into rice-growing ecological zones using a cluster analysis algorithm, and extracting meteorological data and soil data for each rice-growing ecological zone;
[0063] Wherein, step S300 includes:
[0064] Step S301: collecting meteorological data of a target area over several years and soil data of the target area over several years or soil data of the current year; the target area is an area corresponding to the recommended rice planting density; selecting a corresponding resolution according to the area of the target area; if the area of the target area is greater than a threshold, selecting a first resolution; if the area of the target area is less than the threshold, selecting a second resolution; selecting a corresponding coordinate system according to the geographical location of the target area; unifying the soil data, meteorological data, and recommended planting density data into the same coordinate system, and resampling the soil data and meteorological data to the same resolution;
[0065] Step S302: Using meteorological factors of the target area as clustering features, the meteorological factors include accumulated temperature, annual precipitation, and solar radiation; clustering using a k-means clustering algorithm to divide a number of rice-growing ecological zones with the same environmental conditions; and calculating the average value of the meteorological factors of each divided rice-growing ecological zone as a representative value of each rice-growing ecological zone.
[0066] In the embodiment of the present invention, the number of clusters centers is set to 6, the number of initial point selections nstart is set to 50 to divide the rice planting ecological zones, and the divided ecological zones are as follows: Figure 4 As shown;
[0067] Step S303: extracting independent test data of the independent test sites from the coordinate system based on the collected independent test site coordinates; the independent test data of each independent test includes density-yield data, soil data of the rice planting ecological zone, and meteorological data of the rice planting ecological zone;
[0068] The soil data includes soil organic carbon, pH, total nitrogen, total phosphorus, total potassium, cation exchange capacity, soil texture type, bulk density, and soil thickness; the meteorological data adopts the average meteorological data provided by TerraClimate, including maximum temperature, minimum temperature, solar radiation, and precipitation;
[0069] Step S400: constructing a multi-source dataset based on the extracted meteorological data and soil data; dividing the multi-source dataset into a training set and a test set, and constructing a machine learning model;
[0070] Wherein, step S400 includes:
[0071] Performing feature selection on meteorological data and soil data through LASSO regression, including selecting organic carbon content, total nitrogen content, cation exchange capacity, and soil pH as soil factors, and selecting rainfall, accumulated temperature, solar radiation, minimum temperature during the growing period, and maximum temperature during the growing period as meteorological factors; combining the soil factors and meteorological factors with the calculated recommended planting density and yield to form a multi-source dataset, dividing 80% of the multi-source dataset into a training set for training a machine learning model, and dividing 20% of the multi-source dataset into a test set for evaluating the machine learning model;
[0072] Extracting meteorological and soil characteristic data from meteorological and soil data by LASSO regression and performing a multicollinearity test on the characteristic data includes: removing highly linearly correlated features from the characteristic data, using the remaining features as covariates, and constructing a machine learning model based on the covariates;
[0073] In the embodiment of the present invention, a total of 9 covariates are selected for random forest model training; in the climate data, the calculation time of the growth period is from April to October;
[0074] The method of constructing a machine learning model based on the covariate includes: using the obtained recommended planting density data that conforms to a normal distribution as a dependent variable, using the covariate corresponding to the recommended planting density data as an independent variable, constructing a random forest model for planting density recommendation, and inputting the training set into the random forest model to train the model;
[0075] In the embodiment of the present invention, a random forest model is selected for training using a training set, and the parameters are selected through ten-fold cross validation and are set as follows: the number of decision trees ntree in the random forest is set to 500, and the number of variables mtry is set to 5;
[0076] Step S500: Input the test set into the machine learning model, calculate the correlation coefficient, root mean square error, and relative root mean square error to evaluate the model;
[0077] Wherein, step S500 includes:
[0078] The formula for calculating the correlation coefficient is as follows:
[0079]
[0080] The formula for calculating the root mean square error is as follows:
[0081]
[0082] The formula for calculating the relative root mean square error is as follows:
[0083]
[0084] Among them, R 2 represents the correlation coefficient; RMSE represents the root mean square error; RRMSE represents the relative root mean square error; i represents the recommended planting density data of the i-th independent experiment in the test set; S i represents the model-predicted recommended planting density data for the i-th independent trial in the test set; Represents the average value of the recommended planting density data of all independent trials in the test set; represents the average value of the model-predicted recommended planting density data of all independent trials in the test set, and n represents the number of independent trials contained in the test set;
[0085] Step S600: Inputting the representative value of the rice-growing ecological zone into the evaluated machine learning model to predict the recommended planting density of the rice-growing ecological zone.
[0086] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art will be able to modify the technical solutions described in the aforementioned embodiments or substitute equivalents for some of the technical features. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A regional rice planting density recommendation method based on multi-source data, characterized in that: The recommended methods include: Step S100: collecting rice test data from a rice test station; Step S200: fitting the planting density and yield data through a nonlinear model to obtain a density-yield relationship curve, calculating the recommended planting density, and performing data cleaning on the calculated recommended planting density; Step S300: Dividing the target area into rice-growing ecological zones using a cluster analysis algorithm, and extracting meteorological data and soil data for each rice-growing ecological zone; Step S400: constructing a multi-source dataset based on the extracted meteorological data and soil data; dividing the multi-source dataset into a training set and a test set, and constructing a machine learning model; Step S500: Input the test set into the machine learning model, calculate the correlation coefficient, root mean square error, and relative root mean square error to evaluate the model; Step S600: Inputting the representative value of the rice-growing ecological zone into the evaluated machine learning model to predict the recommended planting density of the rice-growing ecological zone.
2. The method for recommending regional rice planting density based on multi-source data according to claim 1, characterized in that: Step S100 includes: collecting rice test data from a rice test station, wherein the rice test station is a plurality of test sites for rice cultivation; the rice test data includes the test site, test year, crop variety, planting density, and yield; the collected rice test data is obtained in the form of articles or tables; if the rice test data includes graphical data, it is extracted using image digitization software.
3. The method for recommending regional rice planting density based on multi-source data according to claim 1, characterized in that: Step S200 includes: dividing the collected rice test data into independent tests according to the test site, test year, and crop variety, and extracting the planting density and yield corresponding to each independent test; fitting the planting density and yield corresponding to each of the independent tests using a nonlinear model, and outputting a density-yield relationship curve; defining the planting density corresponding to the highest yield in the density-yield relationship curve as the recommended planting density; eliminating data whose recommended planting density exceeds the test density range, and counting the recommended planting density data of the remaining fitting results as a recommended planting density data set; performing outlier removal and normality test on the obtained recommended planting density data set to obtain recommended planting density data that conforms to the normal distribution.
4. The method for recommending regional rice planting density based on multi-source data according to claim 1, characterized in that: Step S300 includes: Step S301: collecting meteorological data of a target area over several years and soil data of the target area over several years or soil data of the current year; the target area is an area corresponding to the recommended rice planting density; selecting a corresponding resolution according to the area of the target area; if the area of the target area is greater than a threshold, selecting a first resolution; if the area of the target area is less than the threshold, selecting a second resolution; selecting a corresponding coordinate system according to the geographical location of the target area; unifying the soil data, meteorological data, and recommended planting density data into the same coordinate system, and resampling the soil data and meteorological data to the same resolution; Step S302: Using meteorological factors of the target area as clustering features, the meteorological factors include accumulated temperature, annual precipitation, and solar radiation; clustering using a k-means clustering algorithm to divide a number of rice-growing ecological zones with the same environmental conditions; and calculating the average value of the meteorological factors of each divided rice-growing ecological zone as a representative value of each rice-growing ecological zone. Step S303: extracting independent test data of the independent test sites from the coordinate system according to the collected independent test site coordinates; the independent test data of each independent test includes density-yield data, soil data of the rice planting ecological zone, and meteorological data of the rice planting ecological zone.
5. The method for recommending regional rice planting density based on multi-source data according to claim 4, characterized in that: The soil data include soil organic carbon, pH, total nitrogen, total phosphorus, total potassium, cation exchange capacity, soil texture type, bulk density and soil thickness; the meteorological data adopts the average meteorological data provided by TerraClimate, including maximum temperature, minimum temperature, solar radiation and precipitation.
6. The method for recommending regional rice planting density based on multi-source data according to claim 1, characterized in that: Step S400 includes: performing feature selection on meteorological data and soil data through LASSO regression, including selecting organic carbon content, total nitrogen content, cation exchange capacity, and soil pH as soil factors, and selecting rainfall, accumulated temperature, solar radiation, minimum temperature during the growing period, and maximum temperature during the growing period as meteorological factors; the soil factors and meteorological factors are combined with the calculated recommended planting density and yield to form a multi-source data set, 80% of the multi-source data set is divided into a training set for training a machine learning model, and 20% of the multi-source data set is divided into a test set for evaluating the machine learning model.
7. The method for recommending regional rice planting density based on multi-source data according to claim 6, characterized in that: Extracting meteorological and soil characteristic data from meteorological and soil data by LASSO regression, and performing a multicollinearity test on the characteristic data includes: removing highly linearly correlated features from the characteristic data, using the remaining features as covariates, and constructing a machine learning model based on the covariates.
8. The method for recommending regional rice planting density based on multi-source data according to claim 7, characterized in that: Constructing a machine learning model based on the covariate includes: taking the obtained recommended planting density data that conforms to the normal distribution as the dependent variable, taking the covariate corresponding to the recommended planting density data as the independent variable, constructing a random forest model for planting density recommendation, and inputting the training set into the random forest model to train the model.
9. The method for recommending regional rice planting density based on multi-source data according to claim 1, characterized in that: Step S500 includes: The formula for calculating the correlation coefficient is as follows: The formula for calculating the root mean square error is as follows: The formula for calculating the relative root mean square error is as follows: Among them, R 2 represents the correlation coefficient; RMSE represents the root mean square error; RRMSE represents the relative root mean square error; i represents the recommended planting density data of the i-th independent experiment in the test set; S i represents the model-predicted recommended planting density data for the i-th independent trial in the test set; Represents the average value of the recommended planting density data of all independent trials in the test set; It represents the average value of the model-predicted recommended planting density data of all independent trials in the test set, and n represents the number of independent trials contained in the test set.