Soil water content prediction method and system based on canopy-atmospheric environment information

Through the soil moisture content prediction method based on canopy-atmospheric environmental information, multi-source data and machine learning algorithms are used to solve the problems of insufficient prediction accuracy and lack of decision support in the existing technology, and the accurate prediction of soil moisture and scientific support for irrigation decisions are achieved.

CN120146328AInactive Publication Date: 2025-06-13INSTITUTE OF CROP SCIENCE CHINESE ACADEMY OF AGRICULTURAL SCIENCES +1

Patent Information

Application Number
CN202510629138.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-06-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing soil moisture content prediction technology has problems such as insufficient prediction accuracy, extensive data processing, weak generalization of models and lack of decision-making support, which is difficult to meet the requirements of modern agriculture for water resource utilization efficiency and irrigation accuracy.

Method used

The soil moisture content prediction method based on canopy-atmospheric environment information is adopted to build a high-precision model through multi-source data acquisition, advanced data processing and machine learning algorithms to achieve accurate prediction of soil relative moisture content. The method includes the application of environmental data acquisition, data preprocessing, model training and dynamic irrigation decision modules.

Benefits of technology

Real-time and accurate prediction of soil moisture conditions is achieved, scientific basis for irrigation decisions, water resource utilization efficiency is improved, crop yield is guaranteed, and the characteristics of different regions are adapted through transfer learning technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146328A_ABST
    Figure CN120146328A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of intelligent agriculture, and discloses a canopy-atmospheric environment information-based soil water content prediction method and system, and the method comprises the steps: setting a canopy and atmospheric environment monitoring unit in a field, and synchronously collecting the canopy temperature, wind speed, relative humidity, and six-dimensional environment parameters of atmospheric temperature, wind speed and relative humidity; the crown temperature difference is calculated after preprocessing; the method comprises the following steps: constructing a training set by utilizing measured data of soil relative water content, constructing a soil water content prediction model by adopting a random forest algorithm, screening and confirming an optimal prediction model through grid search and cross validation, and screening key features by combining an SHAP value; the water demand is automatically calculated according to a pre-established crop whole-growth-period drought stress threshold table and a dynamic irrigation decision formula, and meanwhile, the adaptation of a general model to a regional specific model is realized by constructing an environmental parameter-actually measured water content database and transfer learning, so that the soil water content state is accurately predicted, and the soil quality is improved. And a scientific basis is provided for precise irrigation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of smart agriculture, and particularly relates to a method for predicting soil relative water content based on canopy temperature information and machine learning, which is used to guide the efficient water-saving irrigation production of field crops. Background Art

[0002] Currently, in agricultural production, soil moisture, as an important factor affecting crop growth and yield, its monitoring and management have always been the key links in precision irrigation technology. Traditionally, soil water content, as an important indicator to measure the soil moisture status in the field, is mainly obtained through two major methods: direct measurement and indirect prediction. The direct measurement methods mainly include the oven-drying method, electrical methods, and radiation methods, while the indirect prediction methods are mostly based on the statistical relationships between meteorological data and related parameters. Although these methods meet the needs of basic moisture monitoring to a certain extent, with the continuous improvement of the requirements for water resource utilization efficiency and irrigation accuracy in modern agriculture, the existing technologies have great limitations in terms of accuracy, real-time performance, and on-site applicability.

[0003] Among them, the oven-drying method, as an international standard method, calculates the water content by measuring the mass difference of soil samples before and after drying. Although it has high measurement accuracy, it has obvious defects such as destroying the soil structure, being time-consuming and laborious, and being unable to achieve in-situ continuous monitoring. This method is not only cumbersome to operate, but also difficult to support large-scale field real-time monitoring, thus restricting the dynamic management of soil moisture status and timely irrigation decision-making.

[0004] In electrical methods, the resistivity method and the capacitance method indirectly infer the soil water content through the electrical properties of the soil. The resistivity method needs to rely on electrodes to form a closed loop for measurement, and its measurement results are easily interfered by the salt content, ion concentration, and other physical and chemical properties in the soil, resulting in unstable data accuracy; while the capacitance method depends on the dielectric constant of the soil at a specific frequency band, but its measurement results are greatly affected by factors such as soil texture and water content uniformity. At the same time, it also requires frequent calibration of the equipment and is difficult to maintain precise stability in a changing field environment.

[0005] The radiation methods include the neutron method and the gamma-ray method, which determine the soil water content by detecting the degree of radiation attenuation. Although this method can theoretically achieve non-destructive measurement, the neutron method has a high radiation safety risk, while the gamma-ray method requires complex shielding equipment and high instrument costs, so it is difficult to be widely applied in large-area fields. In addition, these direct measurement methods all rely on artificial sampling or fixed-installed professional equipment, and cannot achieve dynamic and in-situ continuous monitoring, and the equipment cost is relatively high, resulting in great limitations in real-time performance and universality.

[0006] To overcome the deficiencies of direct measurement methods in real-time monitoring and on-site dynamic regulation, the industry has started to attempt indirect prediction using meteorological data. The most common indirect prediction method is represented by the Canopy Temperature Deficit (CTD) model, which estimates soil water content based on the difference between canopy temperature and atmospheric temperature. However, this method has many problems: First, the CTD model assumes a simple linear relationship between canopy temperature deficit and soil moisture. In reality, affected by weather changes, crop types, and growth stages, the change in canopy temperature has obvious uncertainty; on cloudy or overcast days, the solar radiation weakens, resulting in a reduced daily variation range of canopy temperature, making the linear relationship no longer valid. Second, for different crops and different growth stages, there are significant differences in their canopy structures, leading to limited applicability of the same temperature index under different conditions. In addition, local microclimate factors in the field, such as the regulating effect of terrain undulation or adjacent water bodies on the temperature and humidity fields, also make the regional adaptability of the CTD model poor, ultimately affecting the prediction accuracy.

[0007] Therefore, the existing techniques for measuring and predicting soil relative water content have the following problems:

[0008] (1) Although the traditional direct measurement method has high accuracy, due to the need to destroy the soil structure, the operation is cumbersome, and it relies on professional equipment, it cannot meet the requirements of on-site continuous monitoring and real-time decision-making.

[0009] (2) Indirect prediction methods mainly rely on single or linear relationships such as canopy temperature deficit and atmospheric temperature to establish models. In practical applications, due to the influence of weather, crop growth stages, and local environments, it is often difficult to maintain stable and accurate prediction effects.

[0010] (3) Data processing methods mainly use traditional means such as linear interpolation, ignoring the spatio-temporal autocorrelation of soil water content data at the field scale. And due to the single data source and the lack of integration of high-resolution remote sensing data and unmanned aerial vehicle data, the regional adaptability and generalization ability of the model are poor.

[0011] (4) Most of the existing systems are for offline data analysis, lacking automated decision support linked to actual irrigation facilities, unable to dynamically adjust the irrigation volume according to real-time monitoring data, and lacking an effective early warning mechanism in the case of long-term prediction and extreme climates.

[0012] A Chinese patent with the publication number CN116304524B discloses a soil water content monitoring method, device, storage medium and apparatus. This monitoring method obtains soil water content influencing factors (such as albedo, surface temperature, vegetation index, surface net radiation flux, soil heat flux and improved temperature vegetation drought index) from multi-source remote sensing data and observation station data, and determines these influencing factors based on the thermal inertia theory, thereby constructing a preset soil water content monitoring model. Based on this model, a multiple linear regression relationship is established with the actual observed values, and then Gaussian process regression (GPR) analysis is used to perform probability inference on the output of the multiple linear regression model. Furthermore, the GPR model is optimized through Bayes' theorem, and finally an optimal soil water content monitoring model is obtained. Through the above technical means, although this method can relatively accurately invert the regional surface soil water distribution by multi-source data fusion, combining multiple statistical regression and probability models, and reduce the errors caused by the environment, data noise, etc. However, this solution mainly relies on remote sensing images and fixed meteorological observation station data, and the spatial resolution and time resolution are relatively limited, making it difficult to fully reflect the subtle changes and local dynamic environment during the growth period of field crops. Moreover, this monitoring method involves relatively complex processes such as obtaining remote sensing data, establishing a preset model, performing multi-level regression and probability optimization, with a large model complexity and limited applicability.

[0013] The above problems not only affect the accurate monitoring of soil water status, but also restrict the popularization and application of irrigation decision-making and water resource regulation measures based on real-time data in modern agriculture. Therefore, it is urgent to find new solutions based on the existing technology. Summary of the Invention

[0014] In view of this, the present invention aims to solve the technical problems existing in the existing soil water content prediction technologies, such as insufficient prediction accuracy, rough data processing, weak model generalization and lack of decision support, and provides a soil water content prediction method and system based on canopy-atmosphere environment information. By collecting multi-source data, advanced processing and machine learning algorithms, a high-precision model is constructed to achieve accurate prediction of soil relative water content, provide a scientific basis for precise irrigation, improve water resource utilization efficiency and ensure crop yield.

[0015] To achieve the above object, the technical solution of the present invention is realized as follows:

[0016] A soil water content prediction method based on canopy-atmosphere environment information, comprising:

[0017] S1: Synchronously collect six-dimensional environmental parameters of canopy temperature, canopy wind speed and canopy relative humidity, as well as atmospheric temperature, ambient wind speed and atmospheric relative humidity through a canopy environment monitoring unit and an atmospheric environment monitoring unit;

[0018] S2: Clean the data of the six - dimensional environmental parameters collected above, fill in the missing values, and perform normalization processing to form a standardized data set, and calculate the canopy - air temperature difference;

[0019] S3: Based on the measured data of the collected soil relative water content, form a training data set with the six - dimensional environmental parameters collected, construct a non - linear soil relative water content prediction model using a machine - learning algorithm, and optimize it by using grid search to optimize the hyperparameter combination. Evaluate the reliability of the model according to the coefficient of determination R 2 and the root - mean - square error, determine the optimal prediction model. The input parameters of the optimal prediction model are the six - dimensional environmental parameters, and the output is the predicted value RC of the soil relative water content pred ;

[0020] S4: According to the predicted value of the soil relative water content output by the optimal prediction model, and referring to the pre - established drought stress threshold table for the whole growth period of the crop, when < , calculate the water requirement according to the following formula:

[0021]

[0022] where d is the root layer depth (m), S is the irrigation unit area (m 2 ), k is the soil type correction coefficient, is the predicted value of the soil relative water content, is the set threshold in the drought stress threshold table for the whole growth period of the crop.

[0023] Further, it also includes step S5:

[0024] Construct a "environmental parameter - measured water content" control database, use the transfer learning algorithm, freeze the underlying decision tree structure of the optimal prediction model, and fine - tune the weights of the last layer to adapt the general model to a region - specific model.

[0025] Further, in step S1, the canopy environment monitoring unit is fixed to the ear position layer of the crop canopy through an adjustable bracket, so that the data collection of the canopy environment monitoring unit is aligned with the crop respiration area, and the atmospheric environment monitoring unit is set at a height of 2 - 5 meters from the ground.

[0026] Further, in step S2, the cubic spline interpolation method is used to fill in the missing values of the canopy temperature and climate environment information data, construct a continuous cubic polynomial to fit the data points according to the data records before and after the missing values, and use the Min - Max normalization method to scale the canopy temperature, climate environment data, and soil relative water content data proportionally to the range of [0, 1].

[0027] Further, in step S3, it includes the following steps:

[0028] S31: Take the collected canopy temperature, meteorological environment information, and the corresponding relative soil water content as inputs. Select three machine learning algorithms, namely support vector machine (SVM), partial least squares regression (PLSR), and random forest regression (RFR), to predict the relative water content. Use the grid search method to screen the best prediction model in each type of model.

[0029] S32: Compare the prediction accuracies of the best prediction models of each type to screen out the best prediction model with the best prediction performance.

[0030] S33: Conduct a feature importance analysis of the input features based on the best prediction model with the best prediction performance to obtain multiple features with high feature importance.

[0031] S34: Use these multiple feature values as the input values of the model with the best prediction performance to construct a non - linear relative soil water content prediction model.

[0032] Further, in step S3, when constructing a non - linear relative soil water content prediction model based on machine learning algorithms, directly adopt the machine learning model of the random forest algorithm. After completing the model training, evaluate the contribution degree of the input features based on the SHAP value, screen out the three key feature variables with the greatest impact on the prediction results as the model inputs, and use the integrated random forest regression algorithm to predict the screened features. The prediction results of each decision tree are integrated by averaging or weighting. The grid search optimizes the hyperparameter combination, including the number of decision trees, the maximum tree depth, and the minimum number of samples in leaf nodes.

[0033] Further, in step S4, the method for generating the crop whole - growth - period drought stress threshold table includes the following steps:

[0034] S41: Set up drought - gradient experimental fields in the target area and divide them into multiple groups of soil water content gradients.

[0035] S42: Measure the leaf water potential and stomatal conductance of each group of crops. When the leaf water potential ≤ - 1.5 MPa, record the corresponding soil water content as the threshold reference value.

[0036] S43: Weight - correct the reference value according to the water - demand sensitivity in the growth stage. The correction formula is:

[0037]

[0038] Among them, is the growth - stage coefficient, is the crop - variety coefficient, is the set threshold in the crop whole - growth - period drought stress threshold table, is the threshold reference value of soil moisture.

[0039] Furthermore, the drought stress threshold table for the entire growth period of the crop dynamically sets thresholds according to growth stages, where the threshold is 55% at the sowing stage, 60% at the jointing stage, and 65% at the heading stage.

[0040] Furthermore, in step S4, the calculation of the soil type correction coefficient k value includes:

[0041] a: Monitoring nodes and TDR devices are arranged in clay, sandy soil, and saline-alkali soil areas to collect data for at least 3 growth cycles;

[0042] b: Construct a soil-specific dataset and optimize hyperparameters using a random forest regression model for grid search;

[0043] c: Select a model with R 2 ≥0.85 and RMSE ≤ 5% as the regional-specific model through ten-fold cross-validation;

[0044] d: Calculate the k value based on the results of the model feature importance ranking:

[0045]

[0046] where is the contribution degree of the i-th feature, is the preset weight, is the feature contribution degree of the preset standard soil.

[0047] The second objective of this application is to disclose a soil water content prediction system based on canopy-atmosphere environment information, including:

[0048] An environmental monitoring device, including a canopy environment monitoring unit and an atmospheric environment monitoring unit, and the canopy monitoring unit is installed through an adjustable bracket;

[0049] A data processing and transmission module, including a microprocessor and a communication unit, for preprocessing the data collected by the environmental monitoring device and uploading it to the cloud server;

[0050] A water content prediction model, using an integrated random forest regression algorithm, with the input parameters being a six-dimensional environmental parameter dataset composed of canopy temperature, canopy wind speed, canopy relative humidity, atmospheric temperature, environmental wind speed, and atmospheric relative humidity, and outputting the predicted value of soil relative water content. The water content prediction model optimizes the hyperparameter combination of the number of decision trees, the maximum tree depth, and the minimum number of samples at leaf nodes through grid search, and uses SHAP values for feature importance analysis to screen key input features;

[0051] The dynamic irrigation decision-making module includes a storage medium and an irrigation trigger module. The storage medium has a built-in drought stress threshold table for the entire growth period of crops. The irrigation trigger module can automatically calculate the water requirement according to the water requirement calculation formula when the predicted value of the relative soil water content is lower than the threshold value of the corresponding growth stage.

[0052] The model optimization interface supports the transfer learning algorithm and can adapt the general model to a region-specific model.

[0053] Compared with the prior art, the soil water content prediction method based on canopy-atmosphere environment information of the present invention has the following advantages:

[0054] 1. In this application, six-dimensional environmental parameters including canopy temperature, wind speed, relative humidity, and atmospheric temperature, wind speed, and relative humidity are synchronously collected by the canopy and atmosphere environment monitoring unit. After data cleaning, missing value filling, and normalization processing, the canopy-air temperature difference is calculated to fully capture the dynamic coupling relationship between crop leaf transpiration and the atmospheric temperature and humidity field. Then, combined with the measured soil water content data, a training set is formed. A non-linear model is constructed using a machine learning algorithm, the hyperparameters are optimized through grid search, the optimal model is determined based on the coefficient of determination and the root mean square error. Finally, the water requirement is calculated according to the model output and the crop drought stress threshold table, realizing the intelligent triggering and automatic calculation of irrigation water volume, being able to predict the soil moisture status in real time and accurately, providing a scientific basis for irrigation decision-making, helping to improve the water resource utilization efficiency, and realizing the efficient water-saving production of field crops.

[0055] 2. This application also realizes the regional adaptation of the model by constructing an "environmental parameter - measured water content" control database and introducing the transfer learning technology. In this step, the underlying decision tree structure of the pre-trained random forest model is frozen, only the top regression layer is fine-tuned, and the data in the poorly performing areas are weighted and trained through the adaptive boosting sampling technology, thereby effectively reducing the risk of model overfitting and improving the adaptability of the model to different soil types and regional microenvironments.

[0056] 3. This application combines the pre-established drought stress threshold table for the entire growth period of crops with the dynamic irrigation decision-making module, and automatically calculates the water requirement according to the root layer depth, irrigation plot area, and soil type correction coefficient when the predicted relative soil water content is lower than the threshold, realizing the closed-loop intelligent management from data collection, model prediction to irrigation decision-making and result verification, greatly improving the water resource utilization efficiency and reducing the risks caused by insufficient water or over-irrigation of crops. Description of the Drawings

[0057] Figure 1 It is a diagram showing the change of the relative soil water content with irrigation and rainfall in the embodiment of the present invention;

[0058] Figure 2Performance of predicting soil relative water content of different soil types under the optimal model in Specific Embodiment 1 of the present invention;

[0059] Figure 3 Prediction accuracy diagram of irrigation amount and rainfall by the model in Specific Embodiment 1 of the present invention. Detailed implementation manners

[0060] Next, the technical solutions in the embodiments of the present application will be clearly described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art belong to the scope protected by the present application.

[0061] In the description of the present application, it should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present application. For the convenience of description, the dimensions of each part shown in the drawings are not drawn according to the actual proportional relationship. Technologies, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods, and devices should be regarded as part of the authorization specification. In all the examples shown and discussed here, any specific value should be interpreted as merely exemplary and not as a limitation. Therefore, other examples of the exemplary embodiments may have different values. It should be noted that similar reference numerals and letters indicate similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.

[0062] It should be noted that the terms "first", "second", etc. in the description and claims of the present application are used to distinguish similar objects and are not used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. generally belong to the same category, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, "and / or" in the description and claims means at least one of the connected objects, and the character " / " generally means an "or" relationship between the associated objects before and after.

[0063] It should be noted that in the description of this application, the orientation or positional relationships indicated by the orientation terms such as "front, rear, upper, lower, left, right", "lateral, vertical, perpendicular, horizontal" and "top, bottom", etc. are usually based on the orientation or positional relationships shown in the drawings. This is only for the convenience of describing this application and simplifying the description. Without contrary explanations, these orientation terms do not indicate or imply that the devices or elements referred to must have a specific orientation or be constructed and operated in a specific orientation. Therefore, they should not be construed as limiting the protection scope of this application; the orientation terms "inside, outside" refer to the inside and outside relative to the contour of each component itself.

[0064] It should be noted that in this application, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitations, the element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, article or device comprising that element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of this application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the reverse order according to the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, the features described with reference to certain examples may be combined in other examples.

[0065] This application discloses a method for predicting soil moisture content based on canopy-atmosphere environment information, including:

[0066] S1: Environmental data collection

[0067] Simultaneously collect six-dimensional environmental parameters of canopy temperature (CT), canopy wind speed (CWS), canopy relative humidity (CRH), as well as atmospheric temperature (AT), ambient wind speed (WS), and atmospheric relative humidity (RH) through a canopy environment monitoring unit and an atmospheric environment monitoring unit;

[0068] S2: Data transmission and preprocessing

[0069] Calculate the canopy-air temperature difference (CTD = CT – AT), and perform data cleaning, missing value filling, and normalization processing on the above six-dimensional environmental parameters collected to form a standardized data set;

[0070] S3: Model training

[0071] Based on the measured data of the relative soil water content collected, a training dataset is formed with the six-dimensional environmental parameters collected. A non-linear prediction model of the relative soil water content is constructed using a machine learning algorithm, and the hyperparameter combination is optimized using grid search. According to the coefficient of determination R 2 and the root mean square error (RMSE), the reliability of the model is evaluated to determine the optimal prediction model. The input parameters of the optimal prediction model are the six-dimensional environmental parameters, and the predicted value of the relative soil water content is output ( );

[0072] S4: Dynamic irrigation decision

[0073] According to the predicted value of the relative soil water content output by the optimal prediction model ( ), and referring to the pre-established drought stress threshold table for the whole growth period of the crop, when is lower than the corresponding threshold ( ), the water requirement is calculated according to the following formula:

[0074]

[0075] where d is the root layer depth (0.3 - 0.8 m), S is the irrigation unit area (m 2 ), and k is the soil type correction coefficient.

[0076] The soil water content prediction method based on canopy-atmosphere environment information disclosed in this application synchronously collects the crop canopy temperature (CT), canopy wind speed (CWS), and canopy relative humidity (CRH) through the canopy environment monitoring unit and the atmosphere environment monitoring unit, and obtains the atmospheric temperature (AT), ambient wind speed (WS), and atmospheric relative humidity (RH) above the canopy at a height of 2-5 meters from the ground, constituting a six-dimensional environmental parameter set. After the collected data is cleaned, missing value filling (using appropriate interpolation methods), and Min–Max normalization processing by the data processing module, the canopy-air temperature difference (CTD = CT – AT) is calculated to capture the dynamic coupling relationship between crop leaf transpiration and the atmospheric temperature and humidity field. Subsequently, this characteristic parameter and other environmental parameters are used as input data to construct a non-linear soil relative water content prediction model based on machine learning algorithms. The model uses grid search to automatically optimize hyperparameters such as the number of decision trees, the maximum tree depth, and the minimum number of samples at leaf nodes to ensure high accuracy and stability of the prediction output in complex field environments. At the same time, based on a pre-established drought stress threshold table for the entire growth period of the crop, which is dynamically set according to the sensitivity of the crop to water demand at different times, when the predicted value is lower than the threshold, the irrigation decision module automatically calculates the required irrigation water volume based on the preset water demand calculation formula, considering the root layer depth, irrigation unit area, and soil type correction coefficient. According to the prediction results and the water demand of the crop, a reasonable irrigation plan is formulated to achieve precise irrigation, thus realizing the full-process intelligent management from environmental monitoring, feature extraction, model prediction, decision execution to model evolution.

[0077] The soil water content prediction method based on canopy-atmosphere environment information described in this application uses real-time environmental data with high spatio-temporal resolution to synchronously obtain crop canopy and atmosphere environment information, breaking through the limitations of traditional reliance on static remote sensing images and fixed meteorological station data, effectively overcoming the problem of unstable prediction accuracy of traditional canopy-air temperature difference models in cloudy days, overcast days, or crop growth stage changes, and leveraging the non-linear modeling advantages of the random forest algorithm and the automated grid search method to achieve precise prediction of soil relative water content in complex and variable field environments. Based on the significant increase in the determination coefficient of the prediction model and the significant reduction in the root mean square error, the optimal prediction model is determined. At the same time, the preset drought stress threshold table can be dynamically adjusted according to the crop growth stage, realizing intelligent triggering and automatic calculation of irrigation water volume, being able to predict the soil moisture status in real-time and accurately, providing a scientific basis for irrigation decision-making, helping to improve water resource utilization efficiency, and achieving high-efficiency water-saving production of field crops.

[0078] As a preferred example of this application, the soil water content prediction method based on canopy-atmosphere environment information further includes step S5: model optimization

[0079] Construct a "environmental parameter - measured water content" control database, and adopt a transfer learning algorithm to freeze the underlying decision tree structure of the random forest model and fine-tune the weight parameters of the last layer to adapt the general model to a region-specific model.

[0080] In the example of this application, by first collecting measured data of environmental factors such as temperature, humidity, light intensity, rainfall, etc. and corresponding soil water content in different climate zones to construct a multi-source heterogeneous control database, and then using a pre-trained random forest model as the basic model and freezing its underlying decision tree structure to maintain the general environmental parameter extraction ability. On this basis, only the weight of the top regression layer is fine-tuned to adapt to the environmental characteristics of a specific region. This application uses an adaptive boosting sampling technique to perform weighted training on the regional data with poor performance in the dataset to further improve the local prediction accuracy. The transfer learning framework realizes the transfer and rapid adjustment of knowledge from the global model to the regional model throughout the process. Finally, the soil water content prediction based on canopy-atmosphere environment information described in this application can not only make full use of global general features, but also perform precise re-learning for regional microclimate, soil texture, and crop growth differences, thus forming a closed-loop collaborative intelligent system among real-time data collection, environmental parameter preprocessing, random forest prediction, and model fine-tuning, which can quickly adapt to new regional environmental changes and output high-precision soil water content prediction results in real time without large-scale re-training.

[0081] As a preferred example of this application, in step S1, the canopy environment monitoring unit is fixed to the ear layer of the crop canopy through an adjustable bracket, and the atmospheric environment monitoring unit is set at a height of 2 - 5 meters from the ground. In the example of this application, a modular adjustable bracket system and an intelligent environment monitoring device are used to realize real-time collection of crop canopy and atmospheric environment parameters in the field. Specifically, the adjustable bracket includes a lower base bracket and an extension bracket. The height of the lower base bracket is set to 2m, and a vertical guide rail is built-in. The extension bracket is selected with a height of 2m or 3m and is connected to the lower base bracket through a threaded interface. The canopy environment monitoring unit is installed on the guide rail through a slider and moves along the guide rail to the ear layer based on the received height instruction to ensure that the data collection is aligned with the crop respiration area. At the same time, the atmospheric environment monitoring unit adopts a multi-position positioning mechanism. When the canopy height ≤ 1m, the atmospheric environment monitoring unit is fixed to the top of the lower base bracket; when the canopy height > 1m, the atmospheric environment monitoring unit is disassembled and the extension bracket is installed to keep its height always higher than the top of the canopy. In the example of this application, the canopy environment monitoring unit can also be fixed on a bamboo pole through nylon ties. The direction of the bamboo pole is parallel to the crop growth direction. As the ear position changes during the crop growth period, the nylon ties are replaced in real time and the position of the canopy environment monitoring unit is manually adjusted to realize dynamic adjustment of the instrument height.

[0082] This application realizes the precise installation and real-time height adjustment of the crop canopy and atmospheric environment monitoring unit by adopting a modular adjustable bracket system, effectively overcoming the measurement errors caused by crop growth changes during data collection in traditional monitoring technologies. Its double-layer partitioned acquisition structure can simultaneously obtain canopy and upper atmospheric environment parameters, thereby accurately capturing the dynamic coupling information between the two, laying a solid and reliable foundation for subsequent data processing and analysis.

[0083] As a preferred example of this application, in step S2, the cubic spline interpolation method is used to fill in the missing values of the canopy temperature and climate environment information data. A continuous cubic polynomial is constructed to fit the data points based on the data records before and after the missing values. Using the Min-Max normalization method, the canopy temperature, climate environment data, and soil relative water content data are scaled proportionally to the range of [0,1]. The normalization formula is as follows:

[0084]

[0085] In the formula, X is the original data, and are the minimum and maximum values in the original data respectively.

[0086] In the example of this application, the data preprocessing link adopts a two-stage processing architecture. First, the cubic spline interpolation method is used to perform high-precision fitting on the missing values in the canopy temperature and climate environment information. By constructing a locally continuous cubic polynomial and solving the optimal polynomial coefficients based on the data records and slope information before and after the missing values, data continuity and second-order derivative smoothing are achieved, effectively reducing the interpolation error in scenarios of sudden changes in temperature or humidity. Subsequently, the Min–Max normalization (MinMax Scaling) method is applied to all environmental parameters and soil relative water content data. According to the normalization formula, the original data is mapped to a unified [0,1] interval. This normalization process uses a dynamic extreme value tracking mechanism to update the minimum and maximum values in the original data in real time, ensuring that the normalization parameters can adapt to changes in the data distribution during long-term data monitoring, thereby providing a data basis with consistent dimensions and high quality for machine learning models such as random forests, minimizing the adverse effects caused by data noise interference and dimensional differences, and being particularly significant in areas with large spatial variations in soil texture, thus laying a solid data foundation for constructing a highly robust and accurate soil water content prediction model.

[0087] As a preferred example of this application, in step S3, the following steps are included:

[0088] S31: Use the collected canopy temperature, meteorological environment information, and the corresponding relative soil water content as inputs. Select three machine learning algorithms, namely support vector machine (SVM), partial least squares regression (PLSR), and random forest regression (RFR), to predict the relative water content. Use the grid search method to screen the best prediction models in each type of model;

[0089] S32: Compare the prediction accuracies of the best prediction models of each type to screen out the best prediction model with the best prediction performance;

[0090] S33: Conduct feature importance analysis of the input features based on the best prediction model with the best prediction performance to obtain multiple features with high feature importance;

[0091] S34: Use these multiple feature values as the input values of the model with the best prediction performance to construct a non - linear relative soil water content prediction model.

[0092] In the example of this application, by using the collected canopy temperature, meteorological environment information, and the corresponding relative soil water content as input content, select three machine learning algorithms suitable for soil moisture monitoring, namely support vector machine, partial least squares regression, and random forest regression, to predict the relative water content. The selection of the above three algorithms is based on their characteristics. Theoretically, these three algorithms can break through the limitations of traditional linear relationship methods and more effectively predict the relative soil water content. And use grid search to evaluate each hyperparameter combination through cross - validation, and screen out the best prediction model corresponding to the best hyperparameter combination with the negative mean squared error as the measurement standard. Then, train and predict through multiple groups of input training data sets, and compare and screen out the best prediction model with the best prediction performance by the coefficient of determination (R 2 ), and root mean square error (RMSE). Then conduct feature importance analysis on the prediction model with the best prediction performance, select the top three feature variables in terms of contribution degree, and use the selected feature subset as the input value of the best prediction model with the best prediction performance to construct a non - linear soil water content prediction model with a reduced dimension.

[0093] Specifically, in the examples of this application, in step S31, for the machine learning algorithms applicable to soil moisture monitoring, three machine learning algorithms are mainly selected from support vector machine (SVM), partial least squares regression (PLSR), and random forest regression (RFR). Among them, partial least squares regression (PLSR) is a multivariate statistical method that combines principal component analysis, canonical correlation analysis, and linear regression analysis; support vector machine (SVM) is a supervised learning algorithm that performs classification or regression analysis by constructing hyperplanes or a set of hyperplanes in high-dimensional or infinite-dimensional spaces; random forest (RFR) also belongs to the category of supervised learning, makes predictions by constructing multiple decision trees, and makes the final prediction decision by integrating the results of multiple decision trees. Theoretically, these three machine learning algorithms can all break through the limitations of traditional linear relationship methods and more effectively predict the relative soil moisture content. In specific applications, after data preprocessing, six-dimensional environmental parameter information such as canopy temperature, canopy wind speed, canopy relative humidity, atmospheric temperature, ambient wind speed, and atmospheric relative humidity is used as input data, and three different machine learning models of support vector machine, partial least squares regression, and random forest regression are constructed to predict the relative soil moisture content. The grid search method is used to screen the hyperparameter combinations with the best soil relative moisture content prediction performance among the three models. The model corresponding to the best hyperparameter combination is the best prediction model in this type of model. Among them, the hyperparameters involved in partial least squares regression for screening are the number of components n_components (1, 2, 3, 4, 5) that control the model dimension and the maximum number of iterations max_iter (100, 200, 500); the hyperparameters involved in support vector machine for screening are the kernel function type (linear kernel function linear, radial basis kernel function rbf, and polynomial kernel function poly), penalty parameter C (0.1, 1, 10, 100), tolerance epsilon (0.01, 0.1, 0.5), and gamma parameter that controls the influence range of data points; the hyperparameters involved in random forest for screening include the number of decision trees n_estimators (50, 100, 150, 200), the maximum depth max_depth of each tree (None, 1, 3, 5, 7, 9, 10), the minimum number of samples min_samples_split required for further partitioning of each internal node (2, 3, 4, 5, 6, 7, 8, 9, 10), and the minimum number of samples min_samples_leaf at the leaf node (1, 2, 3, 4, 5). The grid search evaluates each hyperparameter combination through cross-validation and uses the negative mean squared error (Neg MSE) as a measure to select the optimal hyperparameter combination, that is, to select the hyperparameter combination with the smallest negative mean squared error (closest to zero) in the same type of model. The calculation formula of the negative mean squared error (Negative MSE) is as follows:

[0094]

[0095] Where n is the number of samples, is the i-th true value, is the i-th predicted value.

[0096] In this application, by using the parallel construction of three models, namely support vector machine, partial least squares regression and random forest regression, and the hyperparameter automatic optimization technology of grid search and cross-validation, the full capture of complex non-linear relationships in soil relative water content prediction is realized. By finely screening different models and their hyperparameters, the prediction error can be significantly reduced, and the generalization ability and robustness of the model can be improved. On the one hand, by combining multiple machine learning methods and making full use of their respective advantages, information such as temperature, humidity, and wind speed hidden in environmental parameters can be comprehensively analyzed during the prediction process. On the other hand, cross-validation is used to ensure that each hyperparameter combination can perform consistently under different data partitions, avoiding the risk of model overfitting. In addition, the negative mean square error is selected as the measurement standard to ensure the minimization of the deviation between the predicted value and the true value. Finally, the model can stably output high-precision soil water content prediction results under different crop growth stages and complex environmental conditions, thus providing reliable data support and technical guarantee for agricultural precise irrigation decision-making.

[0097] In the example of this application, the step S32 includes:

[0098] S321: Randomly divide the dataset for model training into ten parts;

[0099] S322: Select seven of them as training data to construct the best soil relative water content prediction model, where the hyperparameters of the model are set as the best hyperparameter combination determined through grid search;

[0100] S323: Use the remaining three parts of data as the validation data in the process of this model construction, and use the best prediction model obtained to make predictions, so as to obtain the prediction results of the three models for the soil relative water content;

[0101] S324: Compare the prediction results of the three best prediction models, and use the coefficient of determination (R 2 ) and the root mean square error (RMSE) as evaluation indicators. Compare and screen out the best prediction model with the largest coefficient of determination and the smallest root mean square error as the best prediction model with the best prediction performance. The calculation formulas of R 2 and RMSE are as follows:

[0102]

[0103]

[0104] Where n is the number of samples, is the i-th true value, is the i-th predicted value, is the average value of the actual values.

[0105] In the example of this application, the training data set is randomly divided into ten parts, seven of which are taken for model training, and the preset hyperparameter combinations in various machine learning algorithms are automatically traversed by means of grid search technology to construct the best soil relative water content prediction model. The remaining three parts of data are used as verification data to conduct prediction tests on the trained model. The coefficient of determination is used to evaluate the degree of explanation of the model for data fluctuations (the closer to 1, the better the model fits the data), and the root mean square error is used to measure the average deviation between the predicted value and the true value (the smaller the value, the more accurate the prediction). By comprehensively comparing the two, the model with the best performance in the training and verification processes is selected. This process effectively utilizes the cross-validation technology, reduces the bias that may be caused by random data division, and at the same time ensures that the model can maintain high accuracy and stability when facing complex and changeable field environment data, and provides reliable soil relative water content prediction data for subsequent irrigation decisions, so as to achieve the optimal balance and closed-loop feedback management between data utilization and model prediction. In the example of this application, the best prediction model with the best prediction performance selected in step S32 is the Random Forest Regression (RFR) model.

[0106] In the example of this application, in step S33, based on the best prediction model with the best prediction performance obtained in S32, the SHAP values of the input features such as canopy information data and meteorological information data are calculated using the Shapley Additive Explanations (SHAP) method. By sorting the magnitudes of the SHAP values, the importance of each feature is determined. Specifically, the larger the SHAP value, the higher the importance of the feature. The calculation formula of the SHAP value is as follows:

[0107]

[0108] In the formula, is the given prediction model, x represents the input sample; is the SHAP value of feature x i ; Q is a subset of the feature set, indicating that feature i is not in Q when calculating the SHAP value; N is the set of all features; f(Q) is the predicted value of the model on the feature subset Q; f(Q∪{i}) is the predicted value of the model on the feature subset Q∪{i}.

[0109] In the example of this application, based on the best prediction model with the best prediction performance determined in step S32, starting from game theory, the SHAP method considers, for each input feature (canopy and meteorological information data), the change in the model prediction value after adding the feature in the case of all subsets Q without this feature, and uses a formula to perform a weighted sum of these changes to obtain the SHAP value of each feature. By sorting the SHAP values, the importance of each feature for predicting the relative soil water content is clarified, and multiple feature variables with high feature importance are screened out, which not only explains how features affect the prediction in a single sample, but also globally shows which features have the greatest impact on the model decision-making, revealing the internal correlation logic between features and prediction values.

[0110] In the example of this application, in step S34, through the SHAP values of the input features calculated in step S33, three feature variables with the greatest feature importance are selected and input into the best prediction model (optimal prediction model) with the best prediction performance for prediction, so as to obtain the predicted value of the relative soil water content. In the example of this application, the best prediction model with the best prediction performance determined in S32 is random forest regression. By inputting the three feature variables screened out in step S33, the random forest algorithm is used for prediction. This algorithm estimates the target variable through the prediction output of each tree; then, by combining the prediction results of each tree, its average value or weighted average value is calculated to obtain the final prediction result. The prediction of the random forest model is represented by the following formula:

[0111]

[0112]

[0113] In the formula, T is the number of trees in the random forest, and f t (X) is the prediction result of the t-th decision tree, generated based on the input feature X.

[0114] In some examples of this application, in step S3, when constructing a non-linear relative soil water content prediction model using a machine learning algorithm, a machine learning model directly using the random forest algorithm is adopted. After completing the model training, the contribution degree of the input features is evaluated based on the SHAP values, and three key feature variables with the greatest impact on the prediction result are screened out as the model input, and the integrated random forest regression algorithm is used to predict the screened features, where the prediction results of each decision tree are integrated by averaging or weighting. The grid search optimizes the hyperparameter combination including the number of decision trees (50 - 200), the maximum tree depth (5 - 20), and the minimum number of samples at the leaf node (1 - 5).

[0115] In the example of this application, after data preprocessing and model training are completed, by calculating the SHAP values of each environmental parameter in the input data, the contribution degree of different features to the prediction of soil relative water content is quantitatively evaluated. Through this process, three key feature variables with the greatest impact on the prediction are selected as the input parameters of the model. The independent prediction results of each decision tree in the random forest regression model are averaged or weighted averaged to achieve the prediction. This method first uses each decision tree to perform independent non-linear mapping on the input features, and then synthesizes the results of each tree to form an overall prediction output. During the whole process, the optimal hyperparameters determined by cross-validation and grid search ensure that the model reaches a balance in data fitting and generalization ability. Then, the key features screened by the SHAP value effectively eliminate noise and redundant information, further improving the high-precision prediction ability of the model for soil relative water content. This method combines the effective screening of input features by the SHAP value and the integrated prediction of the random forest regression model to achieve accurate prediction of soil relative water content. Using the SHAP value to identify key features greatly weakens the interference of irrelevant data, making the model input more concise and efficient. At the same time, the random forest algorithm reduces the prediction error and overfitting risk of a single tree through the integration of multiple decision trees, ensuring the stability and accuracy of the overall model in the complex and changeable field environment.

[0116] As a preferred example of this application, when predicting the irrigation amount and rainfall through the optimal prediction model in step S3, through the collection of irrigation and rainfall data, a canopy temperature and humidity recorder is used to record the canopy temperature before irrigation and rainfall; the volumetric soil water content before and after irrigation and rainfall is obtained through a soil moisture sensor and converted into soil relative water content; the irrigation amount and rainfall amount of irrigation and rainfall events are recorded in real time through a precision water meter and a small weather station respectively. The conversion formula between the volumetric soil water content and the soil relative water content is as follows:

[0117]

[0118] In the formula, RC represents the soil relative water content, SC represents the volumetric soil water content, and FC is the field capacity.

[0119] Then, in order to evaluate the performance of the best soil relative water content prediction model in predicting the irrigation amount and rainfall amount, the prediction method in step S3 is adopted. The specific steps are as follows: input three feature data collected before irrigation and rainfall, and predict the soil relative water content before irrigation and rainfall; then, according to the actually recorded irrigation amount and rainfall amount data, calculate the soil relative water content after irrigation and rainfall; by comparing the prediction result with the actually measured soil relative water content, calculate the R 2 value. The larger the R 2 value, the better the prediction performance of the model. The calculation formula for the soil relative water content after irrigation and rainfall is as follows:

[0120]

[0121] In the formula, I and R respectively represent the irrigation amount and rainfall amount (m 3 ), d is the soil layer depth (m), SC 1 is the soil volume water content after irrigation and after rainfall, SC 2 is the soil volume water content before irrigation and before rainfall, and S represents the area of the irrigation unit (m 2 ).

[0122] Finally, the relative soil water contents before irrigation and before rainfall, and after irrigation and after rainfall predicted by the optimal prediction model are used to calculate the irrigation amount and rainfall amount. The calculation formula is as follows:

[0123]

[0124] In this application, devices such as a canopy temperature and humidity recorder, a soil moisture sensor, a precision water meter, and a small weather station are used to collect data such as the canopy temperature, soil volume water content, irrigation amount, and rainfall amount before irrigation and rainfall. The soil volume water content is converted into the relative soil water content by using the conversion formula between the soil volume water content and the relative soil water content, and then the three characteristic data collected before irrigation and rainfall are input. The optimal prediction model is used to predict the relative soil water content before irrigation and before rainfall. Then, according to the actually recorded irrigation amount and rainfall amount, the relative soil water content after irrigation and after rainfall is calculated through the calculation formula of the relative soil water content after irrigation and after rainfall, and R is calculated by comparing the predicted value with the actual value 2 to evaluate the model performance; finally, the relative soil water contents before irrigation, before rainfall, after irrigation, and after rainfall predicted by the best prediction model are substituted into the calculation formulas of the irrigation amount and rainfall amount to calculate the irrigation amount and rainfall amount, forming a complete process from data collection, model prediction to result verification and calculation. The optimal prediction model described in this application can not only assist in irrigation decision-making, utilize the prediction of the relative soil water content before and after irrigation, but also effectively predict the irrigation amount and rainfall amount.

[0125] This application realizes the accurate prediction of irrigation volume and rainfall by adopting multi-sensor collaborative data acquisition and advanced data preprocessing technologies, using SHAP values to screen key features, and constructing an optimal soil relative water content prediction model (non-linear soil relative water content prediction model) based on the ensemble learning method of a random forest regression model. This method can fully capture the complex non-linear relationship between environmental parameters and soil moisture, and optimize the model hyperparameters through cross-validation and grid search to ensure the robustness and generalization ability of the model in different environments. At the same time, the performance of the model is evaluated and corrected by comparing the predicted soil relative water content with the actual monitoring data, so as to ensure the high accuracy and stability of the prediction results. Finally, the accurate irrigation volume and rainfall are automatically calculated through closed-loop data feedback, effectively reducing water resource waste and ensuring the water required for crop growth, thus promoting the realization of precise irrigation and intelligent management in agricultural production.

[0126] As a preferred example of this application, in step S4, the method for generating the crop's whole growth period drought stress threshold table includes the following steps:

[0127] S41: Set up drought gradient test fields in the target area, and divide them into 5 groups of soil water content gradients (50%, 60%, 70%, 80%, 90%);

[0128] S42: Measure the leaf water potential and stomatal conductance of each group of crops. When the leaf water potential ≤ -1.5 MPa, record the corresponding soil water content as the threshold reference value;

[0129] S43: Weight and correct the reference value according to the water requirement sensitivity at the growth stage. The correction formula is:

[0130]

[0131] Where is the growth stage coefficient (such as 0.05 at the jointing stage and 0.08 at the heading stage); is the crop variety coefficient (such as 1.0 for corn and 0.95 for wheat); is the set threshold in the crop's whole growth period drought stress threshold table; is the threshold reference value of soil moisture, that is, the soil water content recorded when the leaf water potential < -1.5 MPa.

[0132] This application innovatively constructs a dynamic threshold system for drought stress coupling multiple factors. First, the system obtains the benchmark water content value of the crop at the critical drought stage by setting experimental fields with different soil water content gradients. Subsequently, correction coefficients for growth stages and crop varieties are introduced to dynamically weight and correct the benchmark value, thereby generating a drought stress threshold table that not only has physiological mechanism support but also reflects the actual field situation. This dynamic threshold system realizes the quantitative mapping of the relationship among crop leaf water potential, stomatal conductance, and soil water content, effectively improving the accuracy of drought diagnosis and enabling irrigation decisions guided by this threshold table.

[0133] As a preferred example of this application, in step S4, the drought stress threshold table for the entire growth period of the crop dynamically sets thresholds according to growth stages. Among them, the threshold is 55% at the sowing stage, 60% at the jointing stage, and 65% at the heading stage.

[0134] This application accurately divides the key nodes of the entire growth period of the crop, uses a dynamic irrigation decision-making engine to achieve accurate drought stress judgment during the crop growth period, and sets three-level stepped thresholds: 55% at the sowing stage, 60% at the jointing stage, and 65% at the heading stage. In the actual production process, the dynamic irrigation trigger module compares the real-time predicted relative soil water content ( ) with the threshold corresponding to the growth stage. When the predicted value is lower than the threshold, the system automatically calculates the water requirement according to the water requirement formula to guide irrigation. This application constructs a stepped risk warning system by quantifying the drought sensitivity of different growth stages, enabling the irrigation amount to not only compensate for soil water deficit but also avoid water resource waste caused by over-irrigation.

[0135] As a preferred example of this application, in step S4, the value of the soil type correction coefficient k is obtained by a machine learning method driven by multi-period environmental parameters and measured RC data, including:

[0136] a: Monitoring nodes and TDR devices are arranged in clay, sandy soil, and saline-alkali soil areas to collect data for at least 3 growth cycles;

[0137] b: Construct a soil-specific dataset and use a random forest regression model to optimize hyperparameters through grid search;

[0138] c: After ten-fold cross-validation, select the model with R 2 ≥0.85 and RMSE≤5% as the region-specific model;

[0139] d: Calculate the value of k based on the ranking result of model feature importance:

[0140]

[0141] Among them, is the contribution degree of the i-th feature, is the preset weight (e.g., CTD: 0.6, WS: 0.3, RH: 0.1), which is the contribution degree of the characteristics of the preset standard soil.

[0142] This application uses continuous measured data and environmental parameters under different soil conditions, screens out a high-precision region-specific model through grid search and ten-fold cross-validation of a random forest regression model, and combines SHAP values to effectively screen and weight-correct key features, so that the water demand can be accurately calculated in dynamic irrigation decision-making, realizing a closed-loop intelligent decision-making process from the collection of multi-period environmental data, data preprocessing, construction of region-specific models, screening of key features to dynamic calculation of correction coefficients. This process makes full use of continuous monitoring data to ensure a solid data foundation, combines ensemble learning and cross-validation to reduce the risk of model overfitting, and through SHAP value analysis, while ensuring the general non-linear mapping ability, it carefully depicts the key influencing factors in the regional microenvironment, realizing a hybrid modeling paradigm that combines mechanism and data.

[0143] In the example of this application, a soil water content prediction system based on canopy-atmosphere environment information is also disclosed, including:

[0144] An environmental monitoring device, including a canopy environment monitoring unit and an atmosphere environment monitoring unit, and the canopy monitoring unit is installed through an adjustable bracket;

[0145] A data processing and transmission module, including a microprocessor and a communication unit, for preprocessing the data collected by the environmental monitoring device and uploading it to a cloud server;

[0146] A water content prediction model, adopting an integrated random forest regression algorithm, with the input parameters being a six-dimensional environmental parameter dataset composed of canopy temperature, canopy wind speed, canopy relative humidity, atmospheric temperature, environmental wind speed, and atmospheric relative humidity, and outputting the predicted value of soil relative water content ( ), and the water content prediction model optimizes the hyperparameter combination of the number of decision trees, the maximum tree depth, and the minimum number of samples at leaf nodes through grid search, and uses SHAP values for feature importance analysis to screen key input features;

[0147] A dynamic irrigation decision-making module, including a storage medium and an irrigation trigger module, the storage medium internally stores a crop whole growth period drought stress threshold table ( ), and the irrigation trigger module can automatically calculate the water demand according to the formula when the predicted value of soil relative water content ( ) is lower than the threshold value of the corresponding growth stage;

[0148] A model optimization interface, supporting transfer learning algorithms and capable of adapting a general model to a region-specific model.

[0149] Specifically, the soil water content prediction system based on canopy-atmosphere environment information disclosed in this application includes:

[0150] An environmental monitoring device, including:

[0151] A canopy environment monitoring unit, fixed to the ear position layer of the crop canopy through an adjustable bracket, and real-time collecting canopy temperature (CT), canopy wind speed (CWS), and canopy relative humidity (CRH);

[0152] An atmospheric environment monitoring unit, set at a height of 2 - 5 meters from the ground, collecting atmospheric temperature (AT), ambient wind speed (WS), and relative humidity (RH) above the canopy;

[0153] A data processing and transmission module, including:

[0154] An embedded microprocessor, used to calculate the canopy-air temperature difference (CTD = CT - AT);

[0155] A 5G / LoRa communication unit, uploading the six-dimensional environmental parameters of CT, AT, WS, CWS, RH, and CRH to the cloud server;

[0156] A water content prediction model, which is a machine learning model based on the random forest algorithm. The input parameters are the six-dimensional environmental parameter data, and the output is the predicted value of soil relative water content ( ); The model is trained in the following way:

[0157] Combining with a TDR soil moisture meter to synchronously collect environmental parameters and measured data of soil relative water content to form a training dataset;

[0158] Using grid search to optimize the hyperparameter combination, including the number of decision trees (50 - 200), the maximum tree depth (5 - 20), and the minimum number of samples at leaf nodes (1 - 5);

[0159] A dynamic irrigation decision-making module, including:

[0160] A storage medium, with a drought stress threshold table for the entire growth period of the crop built-in ( ), dynamically setting thresholds according to the growth stage (55% at the sowing stage, 60% at the jointing stage, 65% at the heading stage);

[0161] An irrigation trigger module, when < When, calculate the water demand according to the formula:

[0162]

[0163] Among them, d is the root layer depth (0.3 - 0.8m), S is the irrigation unit area (m 2 ), and k is the soil type correction coefficient;

[0164] The model optimization interface supports external TDR / FDR soil moisture detection devices and adapts the general model to a region-specific model through a transfer learning algorithm, including:

[0165] Construct a control database of "environmental parameters - measured water content";

[0166] Freeze the underlying decision tree of the general model and only fine-tune the weight parameters of the last layer.

[0167] The soil water content prediction method and system based on canopy-atmosphere environment information disclosed in this application use multi-sensor collaboration of the canopy and the atmosphere environment to collect real-time environment data. Advanced data preprocessing technology is used to clean, fill in missing values, and perform Min–Max normalization on six-dimensional environmental parameters such as canopy temperature, canopy wind speed, canopy relative humidity, atmospheric temperature, ambient wind speed, and atmospheric relative humidity, and then calculate the canopy-air temperature difference to capture the dynamic coupling relationship between crop leaf transpiration and the atmospheric temperature and humidity field. Based on this, a non-linear soil relative water content prediction model is constructed by screening the model with the best prediction performance using various machine learning algorithms such as random forest, support vector machine, and partial least squares regression, or directly construct a non-linear soil relative water content prediction model based on the random forest machine learning algorithm. The prediction model automatically optimizes hyperparameters through grid search and cross-validation and quantitatively evaluates the contribution degree of input features using SHAP values, screening out the key feature variables that have the greatest impact on the prediction effect, thereby realizing the reduction of model input features and noise elimination. At the same time, in the model optimization stage of this application, a control database of "environmental parameters - measured water content" is constructed, and the transfer learning technology is used to freeze the underlying decision tree structure of the random forest model and only fine-tune the weight of the top regression layer to achieve the rapid adaptation of the general model to the region-specific model, so as to ensure high-precision prediction under different soil types and regional conditions. In addition, this application also dynamically adjusts the threshold according to the built-in crop whole growth period drought stress threshold table of the system, based on different growth stages (for example, the sowing stage is set to 55%, the jointing stage is 60%, and the heading stage is 65%). When the predicted soil relative water content is lower than the corresponding threshold, the irrigation trigger module automatically calculates the water demand according to the preset irrigation amount calculation formula, which combines the root layer depth, irrigation unit area, and soil type correction coefficient, forming a closed-loop intelligent management system from data collection, model prediction to irrigation decision-making and result verification.

[0168] A person of ordinary skill in the art can understand that all or part of the steps of implementing the above-mentioned method embodiments can be completed by hardware related to the computer program. The aforementioned computer program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above-mentioned method embodiments; and the aforementioned storage medium includes: ROM, RAM, FLASH, floppy disk, magnetic disk or optical disk, mechanical hard disk, cloud and other media that can store program codes. The computer processor can be a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), that is, the steps of the soil moisture prediction method based on canopy-atmospheric environment information can be implemented by the above-mentioned processor.

[0169] The soil moisture prediction method and system based on canopy-atmospheric environmental information described in this application significantly improves the accuracy and real-time performance of soil moisture status prediction through multiple optimization measures such as high spatiotemporal resolution data acquisition, fine data preprocessing, nonlinear model prediction, and dynamic threshold control, effectively reducing crop yield losses and water resource waste caused by insufficient or excessive irrigation, and through the transfer learning mechanism, the model has the ability to quickly adapt to different regions and soil conditions, providing high-precision, low-cost intelligent irrigation solutions for large-scale farmland. The present invention supports the atmospheric environment monitoring unit to adjust its height as the canopy grows by designing an expandable bracket structure, and optimizes the regional specificity model by combining TDR detection data with a transfer learning algorithm. Field tests have shown that the system prediction accuracy R 2 ≥0.85, which can adapt to multiple soil types such as loam and sand, and provide high-precision, low-cost intelligent irrigation solutions for large-scale farmland.

[0170] Example 1

[0171] Taking the corn field experiments in Luohe City, Henan Province and Tongliao City, Inner Mongolia Autonomous Region as examples, the specific steps of the soil moisture prediction method based on canopy-atmosphere environmental information are explained.

[0172] By setting up field test points in Luohe Farm and Liaohe Town, Horqin District, Tongliao City respectively, the soil types of the two places are loam and sand, and the climate characteristics are temperate transitional monsoon climate and mid-temperate semi-arid continental monsoon climate respectively.

[0173] In the experimental design, a split plot experiment was used with three irrigation treatments: 75%-85% FC (W1), 55%-65% FC (W2), and ≤50% FC (W3). Each treatment had three replicates and the plot size was no less than 96m. 2 . Install high-precision water meters at the main water inlet of each plot to accurately control the irrigation volume. Figure 1The graph showing the variation of relative soil water content with irrigation and rainfall.

[0174] In terms of crop management, select corn varieties suitable for local cultivation, adopt wide-narrow row planting, shallow buried drip irrigation and integrated fertilization system. Conduct drip irrigation one day after sowing to improve the uniformity of seedling emergence; during the growth period, ensure the effective control of weeds and pests in the experimental plot through chemical control.

[0175] In terms of data collection and processing, collect soil moisture data at fixed intervals, calculate the relative water content, and record meteorological data such as canopy temperature, atmospheric temperature, wind speed, relative humidity, etc. Preprocess the collected data, including operations such as data cleaning, missing value filling, and normalization.

[0176] In terms of model training and prediction, divide the preprocessed data into a training set and a test set, select three machine learning algorithms suitable for soil moisture monitoring, namely support vector machine, partial least squares regression, and random forest regression, to predict the training of the relative water content model. Optimize the model performance by adjusting the model parameters. Use the trained model to predict the test set data, and evaluate the model prediction results through indicators such as the coefficient of determination (R 2 ), and root mean square error (RMSE). Then conduct a feature importance analysis on the prediction model with the best prediction performance, select the top three feature variables in terms of contribution, and use the selected feature subset as the input value of the best prediction model with the best prediction performance to construct a non-linear soil water content prediction model with a reduced dimension.

[0177] In terms of irrigation amount decision-making, based on the optimal model, input the real-time collected canopy temperature, canopy-air temperature difference, and atmospheric temperature data to predict the relative soil water content. Then, according to the prediction results and the water requirements of the crops, formulate a reasonable irrigation plan to achieve precise irrigation.

[0178] In specific implementation, the ability of the soil water content prediction method based on canopy-atmosphere environment information to predict the relative soil water content under two soil types. The method for the optimal prediction model to predict the relative soil water content under two soil types includes the following steps:

[0179] 1 Data collection. The collected data mainly comes from two regions, namely the experimental area in Luohe City, Henan Province (loam) and the experimental area in Tongliao City, Inner Mongolia Autonomous Region (sandy soil). The two regions have typical different soil types;

[0180] 2 The collected data is divided into two categories according to soil types, and the dataset input into the prediction model is randomly divided into ten parts; seven of them are selected as training data, and the hyperparameters of the prediction model are set as the best hyperparameter combination determined through grid search; the remaining three parts of data are used as the validation data of the prediction model, and the constructed prediction model is used for prediction, so as to obtain the prediction results of the soil relative water content by the model under two soil types; the evaluation indexes for the prediction results are the coefficient of determination (R 2 ) and the root mean square error (RMSE).

[0181] The results show that, as Figure 2 shown Figure 2 , for the two soil types of sandy soil and loam soil, the figure shows the comparison between the actual soil relative water content (Actual RC) and the predicted soil relative water content (Predicted RC). The training accuracy of the prediction model for sandy soil reaches 92%, and the RMSE deviation is only 3.75%; the accuracy of loam soil reaches 94%, and the RMSE deviation is 3.09%. The prediction accuracy of the validation set data for sandy soil reaches 86%, and the RMSE deviation is 5.18%; the accuracy of loam soil reaches 85%, and the RMSE deviation is 5.58%.

[0182] The method for verifying the prediction ability of the best model for irrigation amount and rainfall amount in this embodiment includes the following steps:

[0183] 1 Data acquisition The acquired data mainly includes canopy-meteorological data obtained through environmental monitoring devices before and after irrigation and rainfall in two regions, as well as the recorded actual irrigation amount and rainfall amount;

[0184] 2 Input the data before and after irrigation and rainfall obtained through the environmental monitoring device into the model to obtain the soil relative water content before and after irrigation and rainfall. Calculate the corresponding irrigation amount and rainfall amount according to the two predicted soil relative water contents before and after. Compare the calculation results with the actual implemented and occurred irrigation and rainfall amounts, and calculate the coefficient of determination. The larger the coefficient of determination, the better the prediction performance of the model is proved.

[0185] The results show that, referring to Figure 3 , Figure 3 , for the prediction accuracy of the optimal prediction model for irrigation amount and rainfall amount in the soil water content prediction method based on canopy-atmosphere environment information described in this application, the correlation R 2 between the predicted irrigation and rainfall amounts and the actual irrigation and rainfall amounts reaches 83%.

[0186] The embodiments of the present application have been described above in conjunction with the accompanying drawings. Without conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application is not limited to the specific embodiments described above. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them belong to the protection scope of the present application.

Claims

1. A soil moisture prediction method based on canopy-atmosphere environmental information, characterized in that: include: S1: The canopy environment monitoring unit and the atmospheric environment monitoring unit synchronously collect the canopy temperature, canopy wind speed and canopy relative humidity, as well as the six-dimensional environmental parameters of atmospheric temperature, ambient wind speed and atmospheric relative humidity; S2: Data cleaning, missing value filling and normalization of the above-collected six-dimensional environmental parameters are performed to form a standardized data set and calculate the canopy air temperature difference; S3: Based on the measured data of soil relative moisture content, a training data set is formed with the collected six-dimensional environmental parameters. A nonlinear soil relative moisture content prediction model is constructed based on machine learning algorithm, and the hyperparameter combination is optimized using grid search. The determination coefficient R 2 The reliability of the model is evaluated by the root mean square error, and the optimal prediction model is determined. The input parameters of the optimal prediction model are the six-dimensional environmental parameters, and the output is the predicted value of soil relative moisture content RC pred ; S4: Based on the predicted soil relative moisture content output by the optimal prediction model and referring to the pre-established crop drought stress threshold table for the entire growth period, < The water requirement is calculated according to the following formula: ; Where d is the root layer depth (m), S is the irrigation unit area (m 2 ), k is the soil type correction factor, is the predicted value of soil relative water content, Set thresholds in the drought stress threshold table for the entire crop growth period.

2. The soil moisture prediction method based on canopy-atmosphere environmental information according to claim 1 is characterized in that: The step S5 is also included: A comparison database of "environmental parameters-measured water content" was constructed, and the transfer learning algorithm was used to freeze the underlying decision tree structure of the optimal prediction model, and the weight parameters of the last layer were fine-tuned to adapt the general model to a regional-specific model.

3. The soil moisture prediction method based on canopy-atmosphere environmental information according to claim 1 is characterized in that: In step S1, the canopy environment monitoring unit is fixed to the crop canopy ear layer through an adjustable bracket so that the data collection of the canopy environment monitoring unit is aligned with the crop breathing area, and the atmospheric environment monitoring unit is set at a height of 2-5 meters from the ground.

4. The soil moisture prediction method based on canopy-atmosphere environmental information according to claim 1, characterized in that: In step S2, the cubic spline interpolation method is used to fill the missing values ​​of the canopy temperature and climate environment information data. Continuous cubic polynomial fitting data points are constructed according to the data records before and after the missing values. The Min-Max normalization method is used to scale the canopy temperature, climate environment data and soil relative moisture content data to the range of [0,1].

5. The soil moisture prediction method based on canopy-atmosphere environmental information according to claim 1, characterized in that: In step S3, the following steps are included: S31: The collected canopy temperature and meteorological environment information and the corresponding soil relative moisture content are used as inputs, and three machine learning algorithms, support vector machine, partial least squares regression and random forest regression, are selected to predict relative moisture content. The grid search method is used to screen the best prediction model in each type of model. S32: comparing the prediction accuracies of various best prediction models to select the best prediction model with the best prediction performance; S33: performing feature importance analysis of input features based on the best prediction model with the best prediction performance, and obtaining multiple features with high feature importance; S34: Use these multiple eigenvalues ​​as input values ​​of the model with the best prediction performance to construct a nonlinear soil relative moisture content prediction model.

6. The soil moisture prediction method based on canopy-atmosphere environmental information according to claim 1, characterized in that: In step S3, when a nonlinear soil relative moisture content prediction model is constructed based on a machine learning algorithm, a machine learning model of a random forest algorithm is directly used. After the model training is completed, the contribution of the input features is evaluated based on the SHAP value, and the three key feature variables that have the greatest impact on the prediction results are screened out as model inputs. An integrated random forest regression algorithm is used to predict the screened features, wherein the prediction results of each decision tree are integrated by averaging or weighting, and the grid search optimization hyperparameter combination includes the number of decision trees, the maximum tree depth, and the minimum number of leaf node samples.

7. The soil moisture prediction method based on canopy-atmosphere environmental information according to claim 1, characterized in that: In step S4, the method for generating the crop full growth period drought stress threshold table comprises the following steps: S41: Set up drought gradient test fields in the target area, with multiple groups of soil moisture gradients; S42: Measure the leaf water potential and stomatal conductance of each group of crops. When the leaf water potential is ≤-1.5MPa, record the corresponding soil moisture content as the threshold reference value; S43: Weighted correction of the benchmark value according to the water sensitivity of the growth stage. The correction formula is: ; in, is the reproductive stage coefficient, is the crop variety coefficient, Set the threshold value in the drought stress threshold table for the whole growth period of crops. is the soil moisture threshold value.

8. The soil moisture prediction method based on canopy-atmosphere environmental information according to claim 1, characterized in that: The drought stress threshold table for the entire growth period of crops dynamically sets thresholds according to the growth stage, including 55% at the sowing stage, 60% at the jointing stage, and 65% at the heading stage.

9. The soil moisture prediction method based on canopy-atmosphere environmental information according to claim 1, characterized in that: In step S4, the soil type correction coefficient k value is calculated, including: a: Deploy monitoring nodes and TDR equipment in clay, sandy soil, and saline-alkali soil areas to collect data for at least three growth cycles; b: Construct soil-specific datasets and use random forest regression models to perform grid search hyperparameter optimization; c: R is screened by ten-fold cross validation 2 Models with ≥0.85 and RMSE ≤5% were considered as region-specific models; d: Calculate the k value based on the model feature importance ranking results: ; in, is the contribution of the i-th feature, is the preset weight, It is the characteristic contribution of the preset standard soil.

10. A soil moisture prediction system based on canopy-atmosphere environmental information, characterized in that: include: An environmental monitoring device, including a canopy environment monitoring unit and an atmospheric environment monitoring unit, wherein the canopy monitoring unit is installed via an adjustable bracket; The data processing and transmission module includes a microprocessor and a communication unit, which is used to pre-process the data collected by the environmental monitoring device and upload it to the cloud server; The moisture content prediction model adopts an integrated random forest regression algorithm, and the input parameters are a six-dimensional environmental parameter data set consisting of canopy temperature, canopy wind speed, canopy relative humidity, atmospheric temperature, ambient wind speed, and atmospheric relative humidity. The output is a predicted value of soil relative moisture content. The moisture content prediction model optimizes the hyperparameter combination of the number of decision trees, the maximum tree depth, and the minimum number of leaf node samples through grid search, and uses SHAP value to perform feature importance analysis to screen key input features; A dynamic irrigation decision module includes a storage medium and an irrigation trigger module. The storage medium has a built-in drought stress threshold table for the entire growth period of crops. The irrigation trigger module can automatically calculate the water requirement according to the water requirement calculation formula when the predicted value of the relative moisture content of the soil is lower than the threshold of the corresponding growth stage; Model optimization interface that supports transfer learning algorithms and can adapt general models to region-specific models.

Citation Information

Patent Citations

  • Soil moisture monitoring method, equipment, storage medium and device

    CN116304524B

  • Rain collection and loss regulation crop water demand prediction method based on improved neural network

    CN109934400A

  • Soil water content influence factor sensitive interval judgment method based on interpretable ensemble learning model

    CN116205310A

  • Data processing method for measuring soil moisture content through time domain reflectometry based on mixed deep learning

    CN118747289A

  • Advanced Systems Providing Irrigation Optimization Using Sensor Networks and Soil Moisture Modeling

    US20230068574A1

Cited By

  • Planting field soil temperature control system and method with intelligent temperature sensing regulation and control

    CN120631085A

  • Dynamic calculation method for water demand of crops and intelligent irrigation system

    CN120805065A

  • Rural drought-caused drinking difficulty risk dynamic assessment and early warning method

    CN120875129A

  • Self-adaptive feature optimization soil salinity estimation method and device

    CN120912975A

  • Intelligent codling moth monitoring system based on multi-source trapping data fusion

    CN120951226A