Method for predicting water productivity of irrigated area based on CatBoost
By using the CatBoost model and Bayesian optimization technology, combined with SHAP analysis, the accuracy and adaptability issues of irrigation water productivity prediction in the context of climate change and water conservation were solved, and efficient and stable irrigation area water productivity prediction and key factor identification were achieved.
Patent Information
- Application Number
- CN202511258384.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-04
- Publication Date
- 2025-10-17
AI Technical Summary
Existing irrigation water productivity prediction methods have problems of insufficient accuracy and weak adaptability in the context of climate change and water conservation, especially in arid and semi-arid areas, where it is difficult to effectively handle complex nonlinear relationships and high-dimensional data.
A CatBoost-based machine learning model is used, combined with the Bayesian optimization algorithm to automatically adjust hyperparameters, and the SHAP method is used to analyze feature contributions. A water productivity prediction model for irrigation areas is constructed to process multi-source data and make accurate predictions.
It significantly improves the accuracy and adaptability of irrigation water productivity prediction, solves the limitations of traditional models in dynamic nonlinear environments and data requirements, achieves more efficient and stable prediction results, and provides quantitative identification and visual analysis of key factors.
Smart Images

Figure CN120806283A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of agricultural water management, and particularly relates to a design of an irrigation area water productivity prediction method based on CatBoost. BACKGROUND
[0002] Water scarcity in irrigated agriculture is a serious challenge to global food security and sustainable use of water resources. According to the data of the United Nations Food and Agriculture Organization, irrigated agriculture accounts for about 20% of the world's cultivated land, but it produces 40% of the world's food. However, this field consumes more than 70% of the world's freshwater resources. The dependence on irrigation in arid and semi-arid regions is particularly serious, and the supply of water resources in these regions directly determines the agricultural productivity. Limited water resources are difficult to meet the needs of agricultural production, directly threatening food security and the stability of the ecological system. Climate change, population growth and increasing water demand exacerbate the contradiction between ensuring agricultural water and ensuring broader water security. In particular, in arid and semi-arid regions, the inefficient use of agricultural water resources and climate change have seriously threatened the improvement of China's food production capacity. For example, the agricultural production in Hetao Irrigation District is facing the severe challenges of water shortage and low crop water productivity; water scarcity exacerbates the pressure on agricultural production and regional water resources management. It is urgent to improve water use efficiency to address these challenges, which is also the focus of agricultural research.
[0003] Irrigation water productivity (IWP) measures the crop yield per unit of irrigation water and is an important indicator of water use efficiency. Improving IWP plays an important role in solving water shortage and improving agricultural productivity, especially in areas with limited water resources. This is particularly important in arid and semi-arid regions, where water scarcity and harsh conditions pose greater challenges to the agricultural production condition system (a collection of all natural environments, agricultural management measures and infrastructure that affect crop growth, yield and water use efficiency, i.e. irrigation water productivity IWP).
[0004] Various methods have been employed in the prior art to assess and improve irrigation water productivity, including traditional models, physical models, and statistical models, which provide valuable insights into water productivity and irrigation optimization. Traditional models, such as the Penman-Monteith equation and the CROPWAT model, rely on empirical data and linear assumptions to estimate crop water requirements and irrigation water use. While these models have been widely applied, they fall short in considering nonlinear and dynamic environmental variables related to climate change. In addition, the calibration of physical models requires extensive field data, which limits their scalability in broader and long-term applications. In recent years, advances in remote sensing technology have enabled the combination of satellite data with hydrological models to more accurately estimate agricultural water productivity. While this approach improves spatial resolution, its practical application is challenged by the uncertainty of meteorological and geographical inputs. Statistical models (such as regression analysis models) provide an alternative approach by considering complex environmental factors, including climate and soil characteristics. However, their predictive power largely depends on the quality and quantity of available data sets. In high-dimensional data sets, these models tend to overfit, thereby reducing their reliability. Given the limitations of traditional statistical methods, there is an increasing need for advanced tools that can handle complex nonlinear relationships and provide accurate predictions under dynamic environmental conditions. To address these challenges, it is crucial to develop robust agricultural water management solutions.
[0005] The rapid development of data science and artificial intelligence has revolutionized the prediction of agricultural irrigation water productivity, with ML (machine learning) algorithms playing a key role. Compared to traditional empirical and statistical methods, ML models have higher accuracy and adaptability, with significant advantages in handling the complex nonlinear relationships inherent in agricultural systems. However, simple ML models have limitations in handling large-scale data sets and effectively optimizing hyperparameters. SUMMARY
[0006] The purpose of the present invention is to address the lack of precision and weak adaptability of existing irrigation water productivity prediction methods in the context of climate change and water conservation. A CatBoost-based irrigation water productivity prediction method is proposed.
[0007] The technical solution of the present invention is: a CatBoost-based irrigation water productivity prediction method, comprising the following steps: S1, collecting multi-source data of the target irrigation area, and pre-processing the multi-source data to obtain input features.
[0008] S2, calculating the target variable IWP of the target irrigation area.
[0009] S3, constructing a modeling data set based on the target variable IWP and the input features, and dividing the modeling data set into a training set and a test set.
[0010] S4, training CatBoost through the training set and automatically adjusting hyperparameters of CatBoost in a parameter range by a Bayesian optimization algorithm to obtain an irrigation district water productivity prediction model.
[0011] S5, analyzing contribution degrees of each input feature to IWP prediction by a SHAP method, and taking an input feature with a contribution degree greater than a preset threshold as a key factor.
[0012] S6, inputting the key factor in the test set into the irrigation district water productivity prediction model to predict water productivity of the target irrigation district, and visualizing a result.
[0013] Further, the multi-source data in step S1 includes climate factors, soil and groundwater characteristics and crop management parameters.
[0014] The climate factors include precipitation PREC, average temperature AT, maximum temperature MaxT, minimum temperature MinT, average relative humidity AH, average wind speed AWS, sunshine duration SH and net radiation NR.
[0015] The soil and groundwater characteristics include initial soil moisture content ISMC, initial soil salinity ISS, initial groundwater salinity IGS, initial groundwater depth IGWD, groundwater depth GWD, soil conductivity SEC, groundwater salinity GS, capillary evaporation CE, soil moisture SM and soil type code STC.
[0016] The crop management parameters include bare land proportion BL, wheat planting proportion WHT, corn planting proportion MAIZ, sunflower planting proportion SUNF, irrigation water amount IW, leaf area index LAI, normalized difference vegetation index NDVI, temperature vegetation drought index TVDI, channel seepage CS and channel drainage CD.
[0017] Further, step S1 includes the following sub-steps: S11, collecting multi-source data of a target irrigation district.
[0018] S12, dividing the target irrigation district into 1km×1km simulation units, generating spatial grid data by using a Kriging interpolation method, and unifying temporal and spatial resolutions of the multi-source data.
[0019] S13, performing standardization processing on the multi-source data to obtain standardized data.
[0020] S14, processing missing points in the standardized data by using a missing value interpolation method, and removing abnormal values exceeding 3 times of a standard deviation to obtain input features.
[0021] S15, Pearson correlation analysis is used to screen the relationship between input features and target variable IWP, and the main load of the first principal component is obtained.
[0022] Further, the formula for calculating the target variable IWP of the target irrigation area in step S2 is: wherein represents the irrigation water productivity, represents the maximum potential yield, represents the irrigation water amount, represents the crop yield response coefficient, represents the actual transpiration amount, represents the potential transpiration amount.
[0023] Further, the initial hyperparameters of CatBoost in step S4 are set as follows: learning rate is 0.05, maximum tree depth is 8, L2 regularization coefficient is 3, bin boundary number is 254, and iteration number is 1000.
[0024] Further, the specific method for automatically adjusting the hyperparameters of CatBoost within the parameter range in step S4 is as follows: A1, define the search space of the hyperparameters of CatBoost.
[0025] A2, establish a Gaussian process as a proxy model to approximate the relationship between the objective function of CatBoost and the hyperparameters.
[0026] A3, take the root mean square error of CatBoost under five-fold cross-validation as the objective function, and randomly select five groups of hyperparameter combinations as initial points.
[0027] A4, perform 25 rounds of iterative optimization on the hyperparameters of CatBoost. In each round of iterative optimization, based on the data of all the evaluated points at present, the proxy model is updated, and the EI value of all points in the search space is calculated using the expected improvement acquisition function. The hyperparameter combination with the maximum EI value is selected as the point to be evaluated in the next round, and the hyperparameter combination with the maximum EI value is used to train CatBoost, and the root mean square error of the hyperparameter combination with the maximum EI value under five-fold cross-validation is calculated.
[0028] A5, complete a total of 30 evaluations including 5 initial points and 25 iterations, and select the hyperparameter combination with the minimum root mean square error as the optimal hyperparameter combination.
[0029] Further, the expected improvement acquisition function in step A4 is: wherein denotes the desired improvement, denotes the hyperparameter combination, denotes the predicted mean of the surrogate model, denotes the currently observed optimal performance value, denotes the currently optimal hyperparameter combination, denotes the adjustable exploration parameter, is an intermediate parameter, denotes the standard deviation of the surrogate model, denotes the cumulative distribution function of the standard normal distribution, denotes the probability density function of the standard normal distribution.
[0030] Further, the calculation formula of the contribution of each input feature to the IWP prediction in step S5 is: wherein denotes the contribution of the input feature i , denotes the feature subset excluding the input feature i , denotes the number of features in the feature subset , denotes the number of input features, denotes the output value when the irrigation water productivity prediction model only uses the feature subset for prediction, denotes the output value when the irrigation water productivity prediction model uses the feature subset and the input feature i for prediction.
[0031] The beneficial effects of the present application are: (1) Compared with traditional crop models (such as Penman-Monteith, CROPWAT) and statistical models (such as regression analysis), the CatBoost machine learning model used in the present application can more effectively handle the complex nonlinear relationships and high-dimensional data between climate, soil, groundwater and management measures, significantly improving the accuracy and adaptability of irrigation water productivity prediction, and overcoming the limitations of traditional models in dynamic nonlinear environment and data requirements.
[0032] (2) Compared with existing mainstream machine learning models (such as random forest, deep neural network, support vector machine, etc.), the CatBoost model used in the present application shows better comprehensive performance and faster computing efficiency when dealing with large-scale, multi-feature agricultural data sets, and can more stably and efficiently complete the prediction task.
[0033] (3) The application innovatively combines the Bayesian optimization technology with the CatBoost model, realizes the automatic and intelligent optimization of the model super parameter, effectively avoids the blindness of manual parameter adjustment, further improves the prediction accuracy and generalization ability of the model, solves the problem of easy overfitting of the machine learning model, and ensures the robustness of the model in a complex environment.
[0034] (4) The application solves the problem of difficult quantification and identification of the key factors affecting the IWP by analyzing the feature importance through the SHAP method. BRIEF DESCRIPTION OF DRAWINGS
[0035] Figure 1 Fig. 1 shows a CatBoost-based irrigation area water productivity prediction method flowchart provided by an embodiment of the application.
[0036] Figure 2 Fig. 2 shows a performance schematic diagram of the CatBoost model in wheat, corn and sunflower training and testing provided by an embodiment of the application.
[0037] Figure 3 Fig. 3 shows a feature importance ranking schematic diagram based on SHAP values in the corn IWP prediction model provided by an embodiment of the application.
[0038] Figure 4 Fig. 4 shows a sensitivity schematic diagram of the IWP of the three crop models to the key variables provided by an embodiment of the application.
[0039] Figure 5 Fig. 5 shows an annual spatial distribution schematic diagram of the irrigation water productivity of the three crops provided by an embodiment of the application. DETAILED DESCRIPTION
[0040] Exemplary embodiments of the application will now be described in detail with reference to the accompanying drawings. It should be understood that the embodiments shown and described in the drawings are merely exemplary and are intended to illustrate the principles and spirit of the application, and are not intended to limit the scope of the application.
[0041] An embodiment of the application provides a CatBoost-based irrigation area water productivity prediction method, as shown in Fig. 1, which comprises the following steps S1-S6: Figure 1 S1, collecting multi-source data of a target irrigation area, and pre-processing the multi-source data to obtain input features.
[0042] Step S1 comprises the following sub-steps S11-S15: S11, collecting multi-source data of a target irrigation area.
[0043] In an embodiment of the application, the multi-source data includes climate factors, soil and groundwater characteristics, and crop management parameters.
[0044] Wherein, the climate factors include precipitation PREC (mm), average temperature AT (℃), maximum temperature MaxT (℃), minimum temperature MinT (℃), average relative humidity AH (%), average wind speed AWS (m / s), sunshine hours SH (h) and net radiation NR (MJ / m²).
[0045] The soil and groundwater characteristics include initial soil moisture content ISMC (%), initial soil salinity ISS (g / kg), initial groundwater salinity IGS (g / L), initial groundwater depth IGWD (m), groundwater depth GWD (m), soil electrical conductivity SEC (dS / m), groundwater salinity GS (g / L), capillary evaporation CE (mm), soil moisture SM (%) and soil type code STC.
[0046] The crop management parameters include bare land proportion BL (%), wheat planting proportion WHT (%), corn planting proportion MAIZ (%), sunflower planting proportion SUNF (%), irrigation water amount IW (mm), leaf area index LAI, normalized difference vegetation index NDVI, temperature vegetation drought index TVDI, channel seepage CS (m / s) and channel drainage CD (m / s). 3 3
[0047] In the embodiment of the application, the multi-source data sources include a national earth system science data center, a national ecological science data center, an earth resource data cloud platform, a CMIP climate model data set and a MODIS remote sensing data set.
[0048] S12, the target irrigation area is divided into 1km*1km simulation units, and the Kriging interpolation method is used to generate spatial grid data, and the temporal and spatial resolutions of the multi-source data are unified.
[0049] In the embodiment of the application, the "create fishing net" tool of ArcGIS is used to divide the target irrigation area.
[0050] S13, the multi-source data is standardized to obtain standardized data.
[0051] S14, the missing value interpolation method is used to process the missing points in the standardized data, and the abnormal values exceeding 3 times the standard deviation are removed to obtain input features.
[0052] S15, Pearson correlation analysis is used to screen the relationship between the input features and the target variable IWP to obtain the main load of the first principal component.
[0053] In the embodiment of the present application, Pearson correlation analysis results show that the absolute value of the correlation coefficient of irrigation water quantity IW, soil conductivity SEC and sunshine duration SH and IWP is large, so the three input features are taken as the main load of the first principal component.
[0054] In the embodiment of the present application, the main load of the first principal component is used to reveal the relationship between input features and the overall structure of data before model training, and provides a reference for feature selection and model understanding. It is a data exploration tool, not a component directly involved in model prediction.
[0055] These load values reveal which original features contribute most to the formation of the first principal component, that is, these features have the strongest influence on the main change direction of the data set. For example, if the load absolute values of precipitation and temperature are high, it means that these two factors jointly dominate most of the variation of the data set.
[0056] S2, calculate the target variable IWP of the target irrigation area.
[0057] In the embodiment of the present application, IWP is defined as the crop yield per unit irrigation water quantity, and its calculation formula is: Among them represents the actual crop yield (kg / m²), represents the irrigation water quantity (mm). The actual crop yield is calculated by the crop water production function model proposed by FAO: Among them represents the maximum potential yield (kg / ha), which is determined by local statistical data and FAO recommended value, represents the crop yield response coefficient (dimensionless), and the typical values of different crops are, for example, 1.15 for wheat, 1.25 for corn, and 1.05 for sunflower, represents the actual transpiration (mm), represents the potential transpiration (mm), wherein derived from MODIS MOD16A2 product, calculated according to FAO56 method. Combining the two formulas: S3, construct a modeling data set according to the target variable IWP and the input features, and divide the modeling data set into a training set and a test set.
[0058] In the embodiment of the present application, the modeling data set is divided by year, wherein the training set accounts for 75%, and the test set accounts for 25%.
[0059] S4, training CatBoost through the training set, and automatically adjusting the hyperparameters of CatBoost in the parameter range by using a Bayesian optimization algorithm to obtain a farmland water productivity prediction model.
[0060] In the embodiment of the application, six kinds of machine learning models, CatBoost, AdaBoost, random forest (RF), deep neural network (DNN), K nearest neighbor regression (KNN) and support vector regression (SVM), are used to construct the farmland water productivity prediction model. The coefficient of determination R 2 , root mean square error (RMSE), mean absolute percentage error (MAPE) and relative absolute error (RAE) are used to evaluate the test set. The results show that the R 2 of CatBoost model in the prediction of wheat, corn and sunflower reaches 0.956, 0.952 and 0.949 respectively, and the RMSE is less than 0.1, which is better than the other five models, and the training time is short and the generalization ability is strong, so CatBoost is selected as the final prediction model in the embodiment of the application.
[0061] In the embodiment of the application, the initial hyperparameters of CatBoost are set as follows: the learning rate is 0.05, the maximum tree depth (Max_Depth, which determines the maximum depth of a single decision tree, is a key parameter for controlling model complexity and preventing overfitting) is 8, the L2 regularization coefficient is 3, the number of bin boundaries is 254, and the number of iterations is 1000. The Bayesian optimization algorithm is used to automatically adjust the hyperparameters in the parameter range to improve the model performance and prevent overfitting.
[0062] In the embodiment of the application, the specific method of automatically adjusting the hyperparameters of CatBoost in the parameter range by using the Bayesian optimization algorithm is as follows: A1, define the search space of the hyperparameters of CatBoost, that is, set the value range of the CatBoost hyperparameters to be optimized.
[0063] In the embodiment of the application, the specific range is set as follows: the learning rate is (0.01, 0.5), the maximum tree depth is (3, 16), the L2 regularization coefficient is (1, 20), the number of bin boundaries is (32, 255), and the number of iterations is (100, 1000).
[0064] A2, establish a Gaussian process (Gaussian Process) as a surrogate model to approximate the relationship between the objective function of CatBoost and the hyperparameters.
[0065] A3, taking the root mean square error (RMSE) of CatBoost under 5-fold cross-validation as the objective function, and randomly selecting 5 sets of hyperparameter combinations as initial points.
[0066] A4, performing 25 rounds of iterative optimization on the hyperparameters of CatBoost, in each round of iterative optimization, updating the surrogate model based on the data of all evaluated points at present, and calculating the EI values of all points in the search space using the expected improvement (EI) acquisition function, selecting the hyperparameter combination with the largest EI value as the point to be evaluated in the next round, and using the hyperparameter combination with the largest EI value to train CatBoost, while calculating the root mean square error of the hyperparameter combination with the largest EI value under 5-fold cross-validation.
[0067] In the embodiment of the application, the expected improvement acquisition function is: wherein represents the expected improvement, represents the hyperparameter combination, represents the predicted mean of the surrogate model, represents the currently observed optimal performance value, represents the currently optimal hyperparameter combination, represents the adjustable exploration parameter, is an intermediate parameter, represents the standard deviation of the surrogate model, represents the cumulative distribution function of the standard normal distribution, represents the probability density function of the standard normal distribution. In the embodiment of the application, the performance value refers to the model performance indicator obtained on the validation set after training CatBoost using the current optimal hyperparameter combination. Specifically, the performance value refers to the root mean square error of the model under 5-fold cross-validation.
[0068] A5, completing a total of 30 evaluations including 5 initial points and 25 iterations, and selecting the hyperparameter combination with the smallest root mean square error as the optimal hyperparameter combination.
[0069] In the embodiment of the application, for the wheat crop, the optimal hyperparameter combination is: learning rate = 0.086, maximum tree depth = 14 (taking the nearest integer of 13.784), L2 regularization coefficient = 11.241, bin boundary number = 113 (taking the nearest integer of 113.410), and iteration number = 761 (taking the nearest integer of 761.025).
[0070] S5, analyze the contribution of each input feature to the IWP prediction by the SHAP (Shapley Additive Explanations) method, and input features with a contribution greater than a preset threshold as key factors.
[0071] In the embodiments of the present application, the calculation formula of the contribution of each input feature to the IWP prediction is: wherein represents the contribution of the input feature i , represents a feature subset excluding the input feature i , represents the number of features in the feature subset , represents the number of input features, represents the output value when the irrigation water productivity prediction model only uses the feature subset for prediction, represents the output value when the irrigation water productivity prediction model uses the feature subset and the input feature i for prediction.
[0072] In the embodiments of the present application, the SHAP method is used for feature importance analysis, and the results show that the irrigation water amount IW, soil electrical conductivity SEC, sunshine hours SH and net radiation NR are key factors affecting the IWP. Among them, IW is positively correlated with the IWP of corn below 800 mm, and has a greater negative impact above 800 mm; the negative effect of IW on the IWP of wheat is obvious, and the negative impact of SEC on the IWP of corn is the largest. SHAP value analysis also reveals the nonlinear relationship between the key factors and the IWP, for example, for sunflowers, when SEC is greater than 4 dS / m, the IWP decreases intensively.
[0073] S6, input the key factors in the test set into the irrigation water productivity prediction model, predict the water productivity of the target irrigation area, and visualize the results.
[0074] In the embodiments of the present application, field water-saving irrigation measures (reducing SEC by 20%), channel lining measures (reducing CS by 50%) and combined scenarios are constructed. The corresponding key factors are replaced in the scenario data set, and other conditions remain unchanged. The IWP spatio-temporal distribution under each scenario in 2023-2030 is predicted by using the irrigation district water productivity prediction model. The prediction results show that the average IWP of the whole region can be increased by 8.5% by water-saving measures alone, by 5.2% by salt control measures alone, and by 13.1% by the combination of the two. The IWP spatial distribution map and change trend map with a resolution of 1km x 1km are generated by using ArcGIS and visualized, and crop layout adjustment suggestions are proposed for different regions, such as planting salt-tolerant crops (sunflowers) in high-salt areas and planting high-yield water-consuming crops (corn) in high-water-efficiency areas.
[0075] As shown in Figure 2 , the comparison between the predicted values and the true values of the CatBoost model for the irrigation water productivity (IWP) of wheat, corn and sunflower in the training set and the test set is shown in the form of scatter plots and fitted curves. Figure 2 In each scatter plot, each dot represents the IWP value of a 1km x 1km grid cell in a certain year. The diagonal line (y=x) is the ideal perfect prediction line. Figure 2 The specific values of the determination coefficient (R 2 ) are clearly marked.
[0076] As can be seen from Figure 2 , the prediction R 2 of the CatBoost model in the present application for the three crops on the test set is stable at more than 0.95. This means that the model can explain more than 95% of the variance in IWP changes, and its predicted values almost coincide with the true values. The RMSE is less than 0.1, further proving that the prediction error is minimal. This directly verifies the excellent ability of the CatBoost algorithm used in the present application in handling high-dimensional, nonlinear agricultural data, and its accuracy is much higher than that of the traditional empirical model and ordinary statistical model mentioned in the background art.
[0077] In addition, Figure 2 the scatter plots of the training set (train) and the test set (test) are closely distributed on both sides of the diagonal line, and the R 2 and RMSE values of the two are very close. There is no overfitting phenomenon that the training set performs very well while the test set performs significantly worse. This shows that the present application successfully learns the universal rules in the data through the technical means of Bayesian optimization of hyperparameters, rather than memorizing noise, ensuring the reliability and stability of the model in different years and climate conditions.
[0078] As shown in Figure 3As shown in the figure, a horizontal bar chart shows the average absolute value of the SHAP value of each input feature in the corn IWP prediction model. The length of the bar represents the average impact of the feature on the model output, that is, the quantitative ranking of feature importance.
[0079] Depend on Figure 3 As can be seen, this invention goes beyond simply providing a prediction value. Instead, it uses SHAP, a unified framework based on game theory, to clearly quantify the contribution of each feature to the final prediction result. This enables agricultural experts and managers to understand how the model makes decisions, greatly enhancing decision makers' trust in the model's predictions.
[0080] Figure 3 The study clearly shows that for corn, irrigation water (IW), soil electrical conductivity (SEC, representing salinity), sunshine hours (SH), and minimum temperature (MinT) are the most critical factors affecting its water productivity. This finding is highly consistent with agricultural physics, validates prior knowledge in agriculture from a data science perspective, and provides a quantitative priority ranking.
[0081] At the same time by Figure 3 The analysis directly indicates that to improve corn IWP, priority should be given to ensuring an adequate irrigation water supply (positive correlation with IW below 800 mm) and focusing on controlling soil salinization (negative correlation with SEC). This enables managers to invest limited water resources and funds in the most effective management measures, achieving a leap from "flood irrigation" management to "targeted" precision management.
[0082] like Figure 4 As shown in the figure, the four most critical features (SEC, IW, SH, and NR) were selected, and a bar chart showing the sensitivity of IWP to the key features was drawn. Figure 4 By calculating the maximum impact of each feature within the observed range (ΔIWP = IWP(max_feature) - IWP(min_feature)) and normalizing it, we can visually demonstrate the overall impact and direction of different features on irrigation water productivity (IWP). The horizontal axis represents the four key input features (IW, SH, NR, and SEC), and the vertical axis represents the normalized ΔIWP value of each feature. Positive values indicate that increasing the feature increases IWP, while negative values indicate a decrease. The different colored bars represent wheat, corn, and sunflower.
[0083] The analysis shows that the irrigation water amount (IW) is significantly negatively correlated with the IWP of all crops, among which the negative impact on corn is the strongest, indicating that corn has the highest dependence on irrigation water, and excessive irrigation will significantly reduce its water use efficiency; as energy input, sunshine hours (SH) and net radiation (NR) are positively correlated with the IWP of all crops, and sufficient light and energy input are beneficial to improve crop yield and water productivity; as a salt indicator, soil electrical conductivity (SEC) has a negative impact on the IWP of all crops, among which corn is the most sensitive, followed by wheat, and sunflower is the least affected, showing its strong salt tolerance.
[0084] Figure 4 The average influence trend of each factor on IWP and the differences between crops are clearly revealed, providing a scientific basis for formulating precise farmland management strategies. For example, for corn planting areas, irrigation water amount should be strictly controlled to avoid excessive irrigation leading to a decrease in water productivity, and attention should be paid to soil salt monitoring and regulation to prevent salinization from causing serious threats to yield. For wheat, although its dependence on irrigation water is also high, it is more stable than corn. For sunflower, its good performance in high-salt environments makes it an ideal crop for rotation or replacement in saline-alkali irrigation areas or water resource scarce areas. This quantitative analysis based on machine learning models converts complex agricultural system relationships into comparable and operable decision-making information, effectively supporting the "water-based and soil-based planting" of fine agricultural management.
[0085] Figure 5 is a box plot for comparing the overall distribution of the irrigation water productivity (IWP) values of wheat, corn, and sunflower during the study period from 2006 to 2013. Figure 5 Each box represents the middle 50% of the IWP data for each crop (i.e., the 25th percentile to the 75th percentile), and the horizontal line within the box represents the median. The upper and lower "whiskers" outside the box represent the maximum and minimum values within the normal range, respectively.
[0086] Figure 5 It is clearly shown that the median IWP of the three crops is in the order of sunflower > wheat > corn. This is highly consistent with the biological and physiological characteristics of the three crops: sunflower is more drought and salt tolerant, and can produce relatively high economic benefits under the same irrigation water amount, so its water productivity is the highest; corn is a high-yield but high-water-consuming crop, so its yield per unit water amount (IWP) is relatively low. Figure 5 First, it is proved from a macro perspective that the calculation results of the invention conform to the basic laws of agriculture, verifying the reliability of the model's basic data and the correctness of the calculation process.
[0087] At the same time, from Figure 5It can be seen that the present invention not only gives the trend but also provides an accurate numerical range: the IWP of wheat is the most stable, concentrated in the range of 0.61 - 2.89 kg / m 3 The IWP of maize was the widest (0.11-3.86 kg / m 3 ), and there are a large number of low-value outliers. This sharply points out the extreme inefficiency in current corn production, which means that there is huge potential and space to improve its water productivity through measures such as optimizing irrigation and preventing and controlling salinization. Sunflower has the highest IWP lower limit (0.89 kg / m 3 ), and the distribution is concentrated, indicating that it can still maintain a high stability of water use under adverse conditions.
[0088] Figure 5 The conclusions from this analysis can directly guide agricultural production: In areas with extremely scarce water resources or high soil salinity, sunflower cultivation should be prioritized to ensure stable water productivity and farmer income. In areas with favorable water and soil conditions, corn can be planted to achieve high yields, but this must be accompanied by precision irrigation and salt management to avoid the inefficiencies shown in the figure and raise its IWF to the upper half of the range. Wheat can be a conservative option for stable yields.
[0089] Through such a concise graph, the present invention makes the complex issue of agricultural production efficiency measurable and comparable, providing a crucial quantitative decision-making basis for the adjustment of planting structure of "growing grain when it is suitable for grain, and growing economic crops when it is suitable for economic crops".
[0090] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of the present invention, and it should be understood that the scope of protection of the present invention is not limited to such specific descriptions and embodiments. Those skilled in the art can make various other specific variations and combinations based on the technical teachings disclosed in the present invention without departing from the essence of the present invention, and such variations and combinations are still within the scope of protection of the present invention.
Claims
1. A method for predicting irrigation area water productivity based on CatBoost, characterized in that: The following steps are involved: S1. Collect multi-source data of the target irrigation area and pre-process the multi-source data to obtain input features; S2. Calculate the target variable IWP of the target irrigation area; S3. Construct a modeling dataset based on the target variable IWP and input features, and divide the modeling dataset into a training set and a test set; S4. CatBoost is trained using the training set, and the Bayesian optimization algorithm is used to automatically adjust the hyperparameters of CatBoost within the parameter range to obtain the irrigation area water productivity prediction model; S5. Analyze the contribution of each input feature to IWP prediction using the SHAP method, and take the input features whose contribution is greater than the preset threshold as the key factors; S6. Input the key factors in the test set into the irrigation area water productivity prediction model, predict the water productivity of the target irrigation area, and visualize the results.
2. The CatBoost-based irrigation area water productivity prediction method according to claim 1, characterized in that: The multi-source data in step S1 include climate factors, soil and groundwater characteristics, and crop management parameters; The climate factors include precipitation PREC, average temperature AT, maximum temperature MaxT, minimum temperature MinT, average relative humidity AH, average wind speed AWS, sunshine hours SH and net radiation NR; The soil and groundwater characteristics include initial soil water content ISMC, initial soil salinity ISS, initial groundwater salinity IGS, initial groundwater depth IGWD, groundwater depth GWD, soil electrical conductivity SEC, groundwater salinity GS, capillary evaporation CE, soil moisture SM and soil type code STC; The crop management parameters include wasteland ratio BL, wheat planting ratio WHT, corn planting ratio MAIZ, sunflower planting ratio SUNF, irrigation water volume IW, leaf area index LAI, normalized difference vegetation index NDVI, temperature vegetation drought index TVDI, channel leakage CS and channel drainage CD.
3. The CatBoost-based irrigation area water productivity prediction method according to claim 1, characterized in that: The step S1 includes the following sub-steps: S11. Collect multi-source data of the target irrigation area; S12. Divide the target irrigation area into 1km×1km simulation units, use Kriging interpolation to generate spatial grid data, and unify the spatiotemporal resolution of multi-source data; S13, performing standardization processing on multi-source data to obtain standardized data; S14. Use missing value interpolation to process missing points in the standardized data, remove outliers that exceed 3 times the standard deviation, and obtain input features; S15. Pearson correlation analysis was used to screen the relationship between the input features and the target variable IWP, and the main loading of the first principal component was obtained.
4. The CatBoost-based irrigation area water productivity prediction method according to claim 1, characterized in that: The formula for calculating the target variable IWP of the target irrigation area in step S2 is: in represents irrigation water productivity, represents the maximum potential output, Indicates the amount of irrigation water, represents the crop yield response coefficient, The actual transpiration rate, represents potential transpiration.
5. The CatBoost-based irrigation area water productivity prediction method according to claim 1, characterized in that: The initial hyperparameters of CatBoost in step S4 are set as follows: learning rate of 0.05, maximum tree depth of 8, L2 regularization coefficient of 3, number of bin boundaries of 254, and number of iterations of 1000.
6. The CatBoost-based irrigation area water productivity prediction method according to claim 1, characterized in that: The specific method of automatically adjusting the hyperparameters of CatBoost within the parameter range using the Bayesian optimization algorithm in step S4 is: A1. Define the search space of CatBoost hyperparameters. A2. Establish a Gaussian process as a surrogate model to approximate the relationship between CatBoost’s objective function and hyperparameters. A3. Use the root mean square error of CatBoost under five-fold cross validation as the objective function and randomly select five sets of hyperparameter combinations as the initial points. A4. Perform 25 rounds of iterative optimization of CatBoost's hyperparameters. In each round, update the surrogate model based on the data of all currently evaluated points. Calculate the EI value of all points in the search space using the expected improvement acquisition function. Select the hyperparameter combination with the largest EI value as the point to be evaluated in the next round. Use this hyperparameter combination to train CatBoost. Calculate the root mean square error of this hyperparameter combination under five-fold cross-validation. A5. Complete a total of 30 evaluations, including 5 initial points and 25 iterations, and select the hyperparameter combination with the smallest root mean square error as the optimal hyperparameter combination.
7. The CatBoost-based irrigation area water productivity prediction method according to claim 6, characterized in that: The expected improved acquisition function in step A4 is: in Expressing the expectation of improvement, represents a hyperparameter combination, represents the predicted mean of the surrogate model, represents the best performance value currently observed, represents the current optimal hyperparameter combination, represents an adjustable exploration parameter, is the intermediate parameter, represents the standard deviation of the surrogate model, represents the cumulative distribution function of the standard normal distribution, Represents the probability density function of the standard normal distribution.
8. The CatBoost-based irrigation area water productivity prediction method according to claim 1, characterized in that: The calculation formula for the contribution of each input feature to IWP prediction in step S5 is: in Represents input features i The contribution of Indicates that input features are not included i The feature subset of Representing feature subsets The number of features in represents the number of input features, Indicates that the irrigation area water productivity prediction model only uses a subset of features The output value when making predictions, Representation of irrigation area water productivity prediction model using feature subsets and input features i The output value when making predictions together.
Citation Information
Cited By
River and lake backflow recognition and driving mechanism analysis method based on interpretable machine learning
CN121901900A
A method for identifying and analyzing driving mechanism of river-lake reverse inflow based on interpretable machine learning
CN121901900B