Tea garden drought grading prediction method and system based on machine learning
By building a tea garden drought classification prediction system based on machine learning and utilizing multimodal feature extraction and fusion technology, the accuracy of tea garden drought prediction and the refinement of irrigation decisions have been improved, solving the problems of inaccurate tea garden drought prediction and imprecise irrigation decisions, and achieving efficient water resource management in tea gardens.
Patent Information
- Application Number
- CN202510782202.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-09-23
AI Technical Summary
Existing technologies are not very accurate in predicting drought conditions in tea gardens and irrigation decisions lack refinement, resulting in waste of water resources.
By building a tea garden drought classification prediction system based on machine learning, using sensor arrays to collect data, combining Limma algorithm, COX survival regression model and GLM generalized linear model analysis, constructing multimodal feature extraction layer, feature fusion layer and soil entropy prediction layer, outputting tea garden drought prediction results, and realizing refined irrigation decision-making.
It has improved the accuracy of drought predictions in tea gardens and the refinement of irrigation decisions, saving water resources and ensuring the healthy growth of tea trees.
Smart Images

Figure CN120687959A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of smart agricultural technology, and in particular to a method and system for grading and predicting drought conditions in tea gardens based on machine learning. Background Art
[0002] Drought has a significant impact on tea plant growth, physiological metabolism, tea quality, and yield. Drought stress can have at least the following effects on tea plants: 1. It causes tea leaves to become smaller and thinner, with shortened leaf length and width, fewer leaves, and dwarfed plants. 2. It increases the number of lateral roots, but restricts overall root development, causing root hairs to dry out and die, and reducing absorption capacity. 3. Long-term drought can stagnate tea plant growth, leading to slow shoot growth, bud and leaf shrinkage, and even death. 4. It inhibits photosynthesis, resulting in a reduced photosynthetic rate. 5. It can alter the content of caffeine, amino acids, tea polyphenols, and other quality components in tea leaves. Therefore, it is necessary to predict drought conditions in tea plantations for better management. Soil moisture, which refers to the content, distribution, and variation of soil water, is a key environmental indicator reflecting soil drought conditions and directly affects tea plant growth. Predicting soil moisture can help predict drought conditions and allow irrigation to be implemented at the appropriate time, saving water resources and ensuring healthy tea plant growth.
[0003] Traditionally, drought prediction in tea gardens involves simply collecting monitoring data through sensors. This data is then used to train a neural network model, which is then used to determine whether a drought exists. This approach has the following shortcomings: 1. Because the factors with the highest correlation with drought in the tea garden monitoring data are not screened, the model is affected by factors with low correlation with drought during the prediction process, resulting in poor prediction accuracy. 2. Drought is simply predicted without implementing hierarchical management of the drought, making it impossible to carry out refined irrigation of the tea garden and wasting water resources.
[0004] Therefore, how to provide a tea garden drought classification prediction method and system based on machine learning to improve the accuracy of tea garden drought prediction and the refinement of irrigation decision-making has become a technical problem that needs to be solved urgently. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a tea garden drought classification prediction method and system based on machine learning, so as to improve the accuracy of tea garden drought prediction and the refinement of irrigation decision-making.
[0006] In a first aspect, the present invention provides a method for predicting drought conditions in tea gardens based on machine learning, comprising the following steps:
[0007] Step S1, collecting a large amount of historical monitoring data from the tea garden through a sensor array including a humidity sensor, a light intensity sensor, a temperature sensor, a carbon dioxide sensor, an air pressure sensor, a particulate matter sensor, a global solar radiation sensor, a pH sensor, a salinity sensor, a conductivity sensor, a rainfall sensor, and a wind speed and direction sensor;
[0008] Step S2: constructing a data set after preprocessing each of the historical monitoring data, and performing a sample expansion operation on the data set;
[0009] Step S3, analyzing the data set using the Limma algorithm, the COX survival regression model, and the GLM generalized linear model to obtain an environmental factor correlation graph;
[0010] Step S4: creating a tea garden drought prediction model based on the sequentially connected multimodal feature extraction layer, feature fusion layer, soil entropy prediction layer, and drought mapping output layer, and setting a loss function of the tea garden drought prediction model;
[0011] The multimodal feature extraction layer is used to extract time-dependent features, soil dynamic features, and spatiotemporal features from the monitoring data; the feature fusion layer is used to adjust the weights of time-dependent features, soil dynamic features, and spatiotemporal features according to the environmental factor correlation diagram, and fuse them to obtain multimodal features; the soil entropy prediction layer is used to calculate the soil entropy value based on the multimodal features; the drought mapping output layer is used to map the soil entropy value according to the set drought classification rules and output the tea garden drought prediction result;
[0012] Step S5: training a tea garden drought prediction model using the data set, and deploying the trained tea garden drought prediction model;
[0013] Step S6: predicting the tea garden drought by using the deployed tea garden drought prediction model.
[0014] Furthermore, in step S1, the historical monitoring data at least include maximum light intensity, average light intensity, sunshine duration, maximum temperature, minimum temperature, average temperature, daytime average temperature, nighttime average temperature, daily temperature difference, maximum air humidity, minimum air humidity, average relative humidity, daytime average air humidity, nighttime average air humidity, maximum carbon dioxide concentration, minimum carbon dioxide concentration, average carbon dioxide concentration, daytime average carbon dioxide concentration, nighttime average carbon dioxide concentration, maximum air pressure value, minimum air pressure value, average air pressure value, daytime average air pressure value, nighttime average air pressure value, air pressure difference, daily average PM2.5, daily average PM10, maximum TBQ total radiation, minimum TBQ total radiation, daily average TBQ total radiation, maximum soil temperature, minimum soil temperature, daily average soil temperature, soil temperature difference, maximum soil pH, minimum soil pH, average soil pH, maximum soil salinity value, minimum soil salinity value, average soil salinity value, maximum soil conductivity, minimum soil conductivity, average soil conductivity, rainfall, wind direction, wind speed, sampling location and sampling time.
[0015] Furthermore, the step S2 is specifically as follows:
[0016] After filling missing values and repairing outliers in each of the historical monitoring data, each of the historical monitoring data is labeled with soil entropy, drought level, and time segment labels to complete the preprocessing of each of the historical monitoring data. A data set is constructed based on the preprocessed historical monitoring data, and a sample expansion operation is performed on the data set through a generative adversarial network; the time segment label is the month of the sampling time;
[0017] The missing value filling is based on the K-nearest neighbor filling method; the outlier repair is based on the Z-score algorithm and the K-means clustering algorithm, specifically:
[0018] Perform Z-score standardization on each of the historical monitoring data, calculate the Z-score value of each of the historical monitoring data after Z-score standardization, and compare the Z-score value with a preset abnormality threshold to screen out abnormal values; cluster the historical monitoring data after Z-score standardization using a K-means clustering algorithm to obtain several cluster centers, and repair each abnormal value in turn based on the nearest cluster center.
[0019] Furthermore, the step S3 is specifically as follows:
[0020] Segmenting the data set based on time segmentation labels to obtain a number of data subsets, and screening significant factors between the data subsets whose change rates exceed a preset change threshold using the Limma algorithm;
[0021] Calculating a first drought risk level for each of the significant factors using a COX survival regression model, calculating a second drought risk level for each of the significant factors using a GLM generalized linear model, cross-validating the first drought risk level and the second drought risk level to screen N environmental factors with the greatest impact on drought conditions under different time segment labels from each of the significant factors, where N is a positive integer; and constructing an environmental factor correlation graph based on the screened environmental factors and time segment labels.
[0022] The significant factors are maximum light intensity, average light intensity, sunshine duration, maximum temperature, minimum temperature, average temperature, average daytime temperature, average nighttime temperature, diurnal temperature difference, maximum air humidity, minimum air humidity, average relative humidity, average daytime air humidity, average nighttime air humidity, maximum carbon dioxide concentration, minimum carbon dioxide concentration, average carbon dioxide concentration, average daytime carbon dioxide concentration, average nighttime carbon dioxide concentration, maximum air pressure, minimum air pressure, average air pressure, average daytime air pressure, average nighttime air pressure, air pressure difference, daily average PM2.5, daily average PM10, maximum TBQ global radiation, minimum TBQ global radiation, daily average TBQ global radiation, maximum soil temperature, minimum soil temperature, daily average soil temperature, soil temperature difference, maximum soil pH, minimum soil pH, average soil pH, maximum soil salinity, minimum soil salinity, average soil salinity, maximum soil electrical conductivity, minimum soil electrical conductivity, average soil electrical conductivity, rainfall, wind direction, wind speed or sampling location.
[0023] Furthermore, in step S4, the multimodal feature extraction layer is constructed based on the meteorological time series encoding module, the soil dynamic encoding module and the spatiotemporal embedding module;
[0024] The meteorological time series encoding module is used to extract time-dependent features from the monitoring data through a bidirectional LSTM network; the soil dynamic encoding module is used to extract soil dynamic features from the monitoring data through a fully connected network with residual connections; the spatiotemporal embedding module is used to extract spatiotemporal features from the monitoring data through an embedding layer;
[0025] The feature fusion layer is used to adjust the weights of time-dependent features, soil dynamic features, and spatiotemporal features by combining the environmental factor correlation graph with a multi-head attention mechanism, and output multimodal features through 1D convolutional layer fusion;
[0026] The soil entropy prediction layer is used to infer multimodal features through gated recurrent units and feature decoupling heads to calculate soil entropy values;
[0027] The specific drought classification rules are as follows:
[0028] The quantile of soil entropy value is (75,100], which corresponds to the first level of drought; the quantile is (25,75], which corresponds to the second level of drought; the quantile is (5,25], which corresponds to the third level of drought; the quantile is (0,5], which corresponds to the fourth level of drought.
[0029] In a second aspect, the present invention provides a tea garden drought classification prediction system based on machine learning, comprising the following modules:
[0030] A historical monitoring data acquisition module is used to collect a large amount of historical monitoring data from the tea garden through a sensor array including a humidity sensor, a light intensity sensor, a temperature sensor, a carbon dioxide sensor, an air pressure sensor, a particulate matter sensor, a global solar radiation sensor, a pH sensor, a salinity sensor, a conductivity sensor, a rainfall sensor, and a wind speed and direction sensor;
[0031] A data set construction module, which constructs a data set after preprocessing each of the historical monitoring data, and performs a sample expansion operation on the data set;
[0032] An environmental factor correlation graph generation module is used to analyze the data set using the Limma algorithm, the COX survival regression model, and the GLM generalized linear model to obtain an environmental factor correlation graph;
[0033] A tea garden drought prediction model creation module is used to create a tea garden drought prediction model based on the sequentially connected multimodal feature extraction layer, feature fusion layer, soil entropy prediction layer, and drought mapping output layer, and set the loss function of the tea garden drought prediction model;
[0034] The multimodal feature extraction layer is used to extract time-dependent features, soil dynamic features, and spatiotemporal features from the monitoring data; the feature fusion layer is used to adjust the weights of time-dependent features, soil dynamic features, and spatiotemporal features according to the environmental factor correlation diagram, and fuse them to obtain multimodal features; the soil entropy prediction layer is used to calculate the soil entropy value based on the multimodal features; the drought mapping output layer is used to map the soil entropy value according to the set drought classification rules and output the tea garden drought prediction result;
[0035] A tea garden drought prediction model deployment module is used to train the tea garden drought prediction model using the data set and deploy the trained tea garden drought prediction model;
[0036] The tea garden drought prediction module is used to predict the tea garden drought by using the deployed tea garden drought prediction model.
[0037] Furthermore, in the historical monitoring data acquisition module, the historical monitoring data at least include maximum light intensity, average light intensity, sunshine duration, maximum temperature, minimum temperature, average temperature, average daytime temperature, average nighttime temperature, daily temperature difference, maximum air humidity, minimum air humidity, average relative humidity, average daytime air humidity, average nighttime air humidity, maximum carbon dioxide concentration, minimum carbon dioxide concentration, average carbon dioxide concentration, average daytime carbon dioxide concentration, average nighttime carbon dioxide concentration, maximum air pressure value, minimum air pressure value, average air pressure value, average daytime air pressure value, average nighttime air pressure value, air pressure difference, daily average PM2.5, daily average PM10, maximum TBQ total radiation, minimum TBQ total radiation, daily average TBQ total radiation, maximum soil temperature, minimum soil temperature, daily average soil temperature, soil temperature difference, maximum soil pH, minimum soil pH, average soil pH, maximum soil salinity value, minimum soil salinity value, average soil salinity value, maximum soil conductivity, minimum soil conductivity, average soil conductivity, rainfall, wind direction, wind speed, sampling location and sampling time.
[0038] Furthermore, the dataset construction module is specifically used to:
[0039] After filling missing values and repairing outliers in each of the historical monitoring data, each of the historical monitoring data is labeled with soil entropy, drought level, and time segment labels to complete the preprocessing of each of the historical monitoring data. A data set is constructed based on the preprocessed historical monitoring data, and a sample expansion operation is performed on the data set through a generative adversarial network; the time segment label is the month of the sampling time;
[0040] The missing value filling is based on the K-nearest neighbor filling method; the outlier repair is based on the Z-score algorithm and the K-means clustering algorithm, specifically:
[0041] Perform Z-score standardization on each of the historical monitoring data, calculate the Z-score value of each of the historical monitoring data after Z-score standardization, and compare the Z-score value with a preset abnormality threshold to screen out abnormal values; cluster the historical monitoring data after Z-score standardization using a K-means clustering algorithm to obtain several cluster centers, and repair each abnormal value in turn based on the nearest cluster center.
[0042] Furthermore, the environmental factor correlation graph generation module is specifically used to:
[0043] Segmenting the data set based on time segmentation labels to obtain a number of data subsets, and screening significant factors between the data subsets whose change rates exceed a preset change threshold using the Limma algorithm;
[0044] Calculating a first drought risk level for each of the significant factors using a COX survival regression model, calculating a second drought risk level for each of the significant factors using a GLM generalized linear model, cross-validating the first drought risk level and the second drought risk level to screen N environmental factors with the greatest impact on drought conditions under different time segment labels from each of the significant factors, where N is a positive integer; and constructing an environmental factor correlation graph based on the screened environmental factors and time segment labels.
[0045] The significant factors are maximum light intensity, average light intensity, sunshine duration, maximum temperature, minimum temperature, average temperature, average daytime temperature, average nighttime temperature, diurnal temperature difference, maximum air humidity, minimum air humidity, average relative humidity, average daytime air humidity, average nighttime air humidity, maximum carbon dioxide concentration, minimum carbon dioxide concentration, average carbon dioxide concentration, average daytime carbon dioxide concentration, average nighttime carbon dioxide concentration, maximum air pressure, minimum air pressure, average air pressure, average daytime air pressure, average nighttime air pressure, air pressure difference, daily average PM2.5, daily average PM10, maximum TBQ global radiation, minimum TBQ global radiation, daily average TBQ global radiation, maximum soil temperature, minimum soil temperature, daily average soil temperature, soil temperature difference, maximum soil pH, minimum soil pH, average soil pH, maximum soil salinity, minimum soil salinity, average soil salinity, maximum soil electrical conductivity, minimum soil electrical conductivity, average soil electrical conductivity, rainfall, wind direction, wind speed or sampling location.
[0046] Furthermore, in the tea garden drought prediction model creation module, the multimodal feature extraction layer is constructed based on the meteorological time series coding module, the soil dynamic coding module and the spatiotemporal embedding module;
[0047] The meteorological time series encoding module is used to extract time-dependent features from the monitoring data through a bidirectional LSTM network; the soil dynamic encoding module is used to extract soil dynamic features from the monitoring data through a fully connected network with residual connections; the spatiotemporal embedding module is used to extract spatiotemporal features from the monitoring data through an embedding layer;
[0048] The feature fusion layer is used to adjust the weights of time-dependent features, soil dynamic features, and spatiotemporal features by combining the environmental factor correlation graph with a multi-head attention mechanism, and output multimodal features through 1D convolutional layer fusion;
[0049] The soil entropy prediction layer is used to infer multimodal features through gated recurrent units and feature decoupling heads to calculate soil entropy values;
[0050] The specific drought classification rules are as follows:
[0051] The quantile of soil entropy value is (75,100], which corresponds to the first level of drought; the quantile is (25,75], which corresponds to the second level of drought; the quantile is (5,25], which corresponds to the third level of drought; the quantile is (0,5], which corresponds to the fourth level of drought.
[0052] The advantages of the present invention are:
[0053] 1. A large amount of historical monitoring data is collected from the tea garden through a sensor array including humidity sensors, light intensity sensors, temperature sensors, carbon dioxide sensors, air pressure sensors, particulate matter sensors, global solar radiation sensors, pH sensors, salinity sensors, conductivity sensors, rainfall sensors and wind speed and direction sensors. After preprocessing each historical monitoring data, a data set is constructed and sample expansion operations are performed; then the data set is analyzed by the Limma algorithm, COX survival regression model and GLM generalized linear model to obtain an environmental factor correlation graph; then a tea garden drought prediction model is created based on the sequentially connected multimodal feature extraction layer, feature fusion layer, soil entropy prediction layer and drought mapping output layer; the multimodal feature extraction layer is used to extract time-dependent features, soil dynamic features and spatiotemporal features from the monitoring data; the feature fusion layer is used to adjust the time-dependent features, soil dynamic features and spatiotemporal features according to the environmental factor correlation graph. The weights of spatiotemporal features are combined and multimodal features are obtained; the soil entropy prediction layer is used to calculate the soil entropy value based on the multimodal features; the drought mapping output layer is used to map the soil entropy value according to the set drought classification rules and output the tea garden drought prediction result; then the tea garden drought prediction model is trained through the data set, the trained tea garden drought prediction model is deployed, and finally the tea garden drought is predicted by the deployed tea garden drought prediction model; that is, the tea garden drought is predicted by the pre-trained tea garden drought prediction model. Before the prediction, the feature fusion layer of the tea garden drought prediction model calls the environmental factor correlation map to adjust the weight of each feature, that is, the weights of the features corresponding to the N environmental factors that have the greatest impact on the drought prediction under the current time segment label are increased, and the tea garden drought prediction result is output in combination with the drought classification rules, which facilitates the refined irrigation of the tea garden, and ultimately greatly improves the accuracy of the tea garden drought prediction and the refinement of the irrigation decision.
[0054] 2. By integrating 11 types of sensors, including meteorological (temperature, humidity, light, and rainfall), soil (pH, salinity, conductivity), and air quality (particulate matter and carbon dioxide), the physical, chemical, and biological dynamics of the tea garden environment are covered, providing multi-source heterogeneous data support for the model. By collecting multi-dimensional data covering extreme values (such as maximum / minimum temperature), mean (daily average temperature), difference (daily temperature difference), time segment mean (daytime / nighttime humidity), etc., the dynamic changes of environmental factors can be accurately portrayed, effectively improving the generalization ability of the tea garden drought prediction model.
[0055] 3. Through the hybrid outlier repair mechanism, that is, combining Z-score anomaly detection and K-means clustering repair, it can not only identify statistical outliers but also use cluster centers for spatial distribution calibration, which is more robust than a single method.
[0056] 4. By adopting generative adversarial networks (GANs) for sample generation, the data can be expanded while maintaining time series characteristics (such as seasonal cycles) and spatial correlations (such as the spatial distribution of soil parameters), effectively alleviating the problem of overfitting of small samples.
[0057] 5. In addition to the drought level, soil entropy value (continuous variable) and month label are additionally marked, which not only retains fine-grained supervision information but also provides prior knowledge for time segment analysis.
[0058] 6. The Limma algorithm was used to screen significant factors with significant fluctuations between time segments (months) to capture seasonally sensitive indicators. The COX survival regression model was used to assess the temporal impact of each significant factor on drought risk and identify key drought-causing factors. The GLM generalized linear model was used to quantify the strength of the association between significant factors and drought severity and identify key drought-causing factors (screening indicators with strong spatial correlation). The output results of the COX survival regression model and the GLM generalized linear model were then cross-verified to ensure the temporal and spatial universality of the screened environmental factors and avoid bias from a single model.
[0059] 7. By using the environmental factor correlation map as prior knowledge, the multi-head attention mechanism is guided to dynamically adjust feature weights (such as strengthening the sunshine weight in the drought season) to achieve knowledge-guided feature fusion; local feature interaction is carried out through the 1D convolution layer to enhance the collaborative expression ability of features in adjacent time steps.
[0060] 8. The temporal context features are extracted through the GRU gated unit, and the decoupling head is combined to separate the deterministic component (such as current humidity) and the uncertain component (such as future rainfall forecast) of the soil entropy value to improve the robustness of the prediction.
[0061] 9. By combining the Limma algorithm in biostatistics and the COX model in the medical field with deep learning, a new paradigm for agricultural data analysis is opened up; by complementing the advantages of traditional statistical models (GLM) and deep learning (GAN, Attention), both interpretability and predictive performance are achieved.
[0062] 10. Holographic collection of tea garden environmental data is achieved by deploying a multi-dimensional sensor array. The data enhancement strategy of hybrid outlier repair and generative adversarial network is combined to improve data quality. The Limma algorithm, COX regression and GLM model are innovatively integrated to screen key environmental factors and construct a spatiotemporal correlation map. A multimodal fusion prediction model is further designed. Bidirectional LSTM, residual network and spatiotemporal embedding layer are used to extract heterogeneous features. Dynamic weighted fusion is achieved through the attention mechanism and combined with the dynamic quantile grading rule of soil entropy value to finally achieve high-precision drought level prediction. It combines the innovation of cross-domain technology integration, model interpretability and engineering deployment practicality, and significantly improves the real-time performance and warning reliability of tea garden drought monitoring. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0064] Figure 1 The present invention is a flow chart of a method for predicting tea garden drought conditions based on machine learning.
[0065] Figure 2 It is a structural schematic diagram of a tea garden drought classification prediction system based on machine learning of the present invention. DETAILED DESCRIPTION
[0066] The technical solution in the embodiments of the present application has the following overall idea: a pre-trained tea garden drought prediction model is used to predict tea garden drought. Before the prediction, the feature fusion layer of the tea garden drought prediction model calls the environmental factor correlation graph to adjust the weight of each feature, that is, the weights of the features corresponding to the N environmental factors that have the greatest impact on drought prediction under the current time segment label are increased, and the tea garden drought prediction results are output in combination with the drought classification rules, which facilitates refined irrigation of the tea garden, thereby improving the accuracy of tea garden drought prediction and the refinement of irrigation decisions.
[0067] Please refer to Figures 1 to 2 As shown, a preferred embodiment of a tea garden drought classification prediction method based on machine learning of the present invention includes the following steps:
[0068] Step S1, collecting a large amount of historical monitoring data from the tea garden through a sensor array including a humidity sensor, a light intensity sensor, a temperature sensor, a carbon dioxide sensor, an air pressure sensor, a particulate matter sensor, a global solar radiation sensor, a pH sensor, a salinity sensor, a conductivity sensor, a rainfall sensor, and a wind speed and direction sensor; the sensor array is also provided with a locator;
[0069] Step S2: constructing a data set after preprocessing each of the historical monitoring data, and performing a sample expansion operation on the data set;
[0070] Step S3, analyzing the data set using the Limma algorithm, the COX survival regression model, and the GLM generalized linear model to obtain an environmental factor correlation graph;
[0071] Step S4: creating a tea garden drought prediction model based on the sequentially connected multimodal feature extraction layer, feature fusion layer, soil entropy prediction layer, and drought mapping output layer, and setting a loss function of the tea garden drought prediction model;
[0072] The multimodal feature extraction layer is used to extract time-dependent features, soil dynamic features and spatiotemporal features from the monitoring data; the feature fusion layer is used to adjust the weights of time-dependent features, soil dynamic features and spatiotemporal features according to the environmental factor correlation diagram, and fuse them to obtain multimodal features; the soil entropy prediction layer is used to calculate the soil entropy value based on the multimodal features; the drought mapping output layer is used to map the soil entropy value according to the set drought classification rules, and output the tea garden drought prediction result; the drought mapping output layer outputs the drought level probability (level 1 drought / level 2 drought / level 3 drought / level 4 drought) through the fully connected layer + Softmax, and introduces temperature scaling technology to calibrate the prediction confidence; the tea garden drought prediction result carries at least the drought period, drought level and drought location;
[0073] Step S5: training the tea garden drought prediction model using the data set, and deploying the trained tea garden drought prediction model; using the AdamW optimizer in the training process, combined with cosine annealing learning rate scheduling;
[0074] Step S6: Predicting the drought in the tea garden using the deployed tea garden drought prediction model. A closed loop is formed from data collection, preprocessing, model training to online prediction, supporting real-time drought monitoring and early warning.
[0075] By deploying a multi-dimensional sensor array, holographic collection of tea garden environmental data is achieved, and data enhancement strategies using hybrid outlier repair and generative adversarial networks are combined to improve data quality. The Limma algorithm, COX regression, and GLM models are innovatively integrated to screen key environmental factors and construct a spatiotemporal correlation map. A multimodal fusion prediction model is further designed, and heterogeneous features are extracted using bidirectional LSTM, residual networks, and spatiotemporal embedding layers. Dynamic weighted fusion is achieved through an attention mechanism and combined with dynamic quantile grading rules for soil entropy values to ultimately achieve high-precision drought level prediction. This model combines innovative cross-domain technology integration, model interpretability, and engineering deployment practicality, significantly improving the real-time performance and warning reliability of tea garden drought monitoring.
[0076] In step S1, the historical monitoring data at least includes maximum light intensity, average light intensity, sunshine duration, maximum temperature, minimum temperature, average temperature, daytime average temperature, nighttime average temperature, daily temperature difference, maximum air humidity, minimum air humidity, average relative humidity, daytime average air humidity, nighttime average air humidity, maximum carbon dioxide concentration, minimum carbon dioxide concentration, average carbon dioxide concentration, daytime average carbon dioxide concentration, nighttime average carbon dioxide concentration, maximum air pressure value, minimum air pressure value, average air pressure value, daytime average air pressure value, nighttime average air pressure value, air pressure difference, daily average PM2.5, daily average PM10, maximum TBQ total radiation, minimum TBQ total radiation, daily average TBQ total radiation, maximum soil temperature, minimum soil temperature, daily average soil temperature, soil temperature difference, maximum soil pH value, minimum soil pH value, average soil pH value, maximum soil salinity value, minimum soil salinity value, average soil salinity value, maximum soil conductivity, minimum soil conductivity, average soil conductivity, rainfall, wind direction, wind speed, sampling location and sampling time.
[0077] By integrating 11 types of sensors, including meteorology (temperature, humidity, light, and rainfall), soil (PH, salinity, conductivity), and air quality (particulate matter and carbon dioxide), the physical, chemical, and biological dynamics of the tea garden environment are covered, providing multi-source heterogeneous data support for the model; by collecting multi-dimensional data such as extreme values (such as maximum / minimum temperature), average values (daily average temperature), difference values (daily temperature difference), and time segment average values (daytime / nighttime humidity), the dynamic changes of environmental factors can be accurately portrayed, effectively improving the generalization ability of the tea garden drought prediction model.
[0078] The step S2 is specifically as follows:
[0079] After filling missing values and repairing outliers in each of the historical monitoring data, each of the historical monitoring data is labeled with soil entropy, drought level, and time segment labels to complete the preprocessing of each of the historical monitoring data. A data set is constructed based on the preprocessed historical monitoring data, and a sample expansion operation is performed on the data set using a generative adversarial network (GAN); the time segment label is the month of the sampling time;
[0080] The missing value filling is based on the K-nearest neighbor filling method; the outlier repair is based on the Z-score algorithm and the K-means clustering algorithm, specifically:
[0081] Perform Z-score standardization on each of the historical monitoring data, calculate the Z-score value of each of the historical monitoring data after Z-score standardization, and compare the Z-score value with a preset abnormality threshold to screen out abnormal values; cluster the historical monitoring data after Z-score standardization using a K-means clustering algorithm to obtain several cluster centers, and repair each abnormal value in turn based on the nearest cluster center.
[0082] Through the hybrid outlier repair mechanism, that is, combining Z-score anomaly detection and K-means clustering repair, it can not only identify statistical outliers, but also use cluster centers for spatial distribution calibration, which is more robust than a single method.
[0083] By adopting generative adversarial networks (GANs) for sample generation, the data is expanded while maintaining time series characteristics (such as seasonal cycles) and spatial correlations (such as the spatial distribution of soil parameters), effectively alleviating the problem of overfitting of small samples.
[0084] In addition to the drought level, soil entropy values (continuous variables) and month labels are additionally annotated to retain fine-grained supervision information and provide prior knowledge for time segment analysis.
[0085] The step S3 is specifically as follows:
[0086] Segmenting the data set based on time segmentation labels to obtain a number of data subsets, and screening significant factors between the data subsets whose change rates exceed a preset change threshold using the Limma algorithm;
[0087] Calculating a first drought risk level for each of the significant factors using a COX survival regression model, calculating a second drought risk level for each of the significant factors using a GLM generalized linear model, cross-validating the first drought risk level and the second drought risk level to screen N environmental factors with the greatest impact on drought conditions under different time segment labels from each of the significant factors, where N is a positive integer; and constructing an environmental factor correlation graph based on the screened environmental factors and time segment labels.
[0088] The Limma algorithm was used to screen significant factors with significant fluctuations between time segments (months) to capture seasonally sensitive indicators. The COX survival regression model was used to evaluate the temporal impact of each significant factor on the risk of drought and identify key drought-causing factors. The GLM generalized linear model was used to quantify the strength of the association between significant factors and drought levels and identify key drought-causing factors (screening indicators with strong spatial correlation). The output results of the COX survival regression model and the GLM generalized linear model were then cross-verified to ensure the temporal and spatial universality of the screened environmental factors and avoid single model bias.
[0089] The significant factors are maximum light intensity, average light intensity, sunshine duration, maximum temperature, minimum temperature, average temperature, average daytime temperature, average nighttime temperature, diurnal temperature difference, maximum air humidity, minimum air humidity, average relative humidity, average daytime air humidity, average nighttime air humidity, maximum carbon dioxide concentration, minimum carbon dioxide concentration, average carbon dioxide concentration, average daytime carbon dioxide concentration, average nighttime carbon dioxide concentration, maximum air pressure, minimum air pressure, average air pressure, average daytime air pressure, average nighttime air pressure, air pressure difference, daily average PM2.5, daily average PM10, maximum TBQ global radiation, minimum TBQ global radiation, daily average TBQ global radiation, maximum soil temperature, minimum soil temperature, daily average soil temperature, soil temperature difference, maximum soil pH, minimum soil pH, average soil pH, maximum soil salinity, minimum soil salinity, average soil salinity, maximum soil electrical conductivity, minimum soil electrical conductivity, average soil electrical conductivity, rainfall, wind direction, wind speed or sampling location.
[0090] In step S4, the multimodal feature extraction layer is constructed based on the meteorological time series encoding module, the soil dynamic encoding module and the spatiotemporal embedding module;
[0091] The meteorological time series encoding module is used to extract time-dependent features (light, temperature and humidity, carbon dioxide, air pressure, PM, radiation, etc.) from the monitoring data through a bidirectional LSTM network. The soil dynamic encoding module is used to extract soil dynamic features (temperature, pH, salinity, and conductivity) from the monitoring data through a fully connected network with residual connections, introducing residual connections to enhance gradient propagation. The spatiotemporal embedding module is used to extract spatiotemporal features from the monitoring data through an embedding layer (mapping longitude and latitude into high-dimensional vectors and decomposing the sampling time into periodic codes of year / month / day / time period).
[0092] The bidirectional LSTM network is used to capture the long-term and short-term temporal dependencies of meteorological factors (such as the cumulative effects of continuous droughts); the residual fully connected network is used to model the dynamic nonlinear relationships of soil parameters (such as the coupled changes in salinity and electrical conductivity); and the spatiotemporal embedding module is used to encode the joint distribution characteristics of spatial position and time (such as the impact of slope aspect on water evaporation).
[0093] The feature fusion layer is used to adjust the weights of time-dependent features, soil dynamic features, and spatiotemporal features by combining the environmental factor correlation graph with a multi-head attention mechanism, and output multimodal features through 1D convolutional layer fusion;
[0094] By using the environmental factor correlation map as prior knowledge, the multi-head attention mechanism is guided to dynamically adjust feature weights (such as strengthening the sunshine weight in the dry season), thereby realizing knowledge-guided feature fusion; local feature interaction is carried out through the 1D convolution layer to enhance the collaborative expression ability of features in adjacent time steps, that is, the 1D convolution layer is used to reduce the dimensionality of the fused high-dimensional features, retaining key information and reducing redundancy.
[0095] The soil entropy prediction layer is used to infer multimodal features through a gated recurrent unit (GRU) and a feature decoupling head to calculate soil entropy values. The feature decoupling head outputs soil parameter prediction values through independent fully connected branches, achieving multi-task joint optimization.
[0096] The temporal context features are extracted through the GRU gating unit, and the decoupling head is combined to separate the deterministic component (such as current humidity) and the uncertain component (such as future rainfall forecast) of the soil entropy value to improve the robustness of the prediction.
[0097] By combining the Limma algorithm in biostatistics and the COX model in the medical field with deep learning, a new paradigm for agricultural data analysis is opened up; by complementing the advantages of traditional statistical models (GLM) and deep learning (GAN, Attention), it has both interpretability and predictive performance.
[0098] The specific drought classification rules are as follows:
[0099] The soil entropy value quantile is between (75,100], corresponding to the first level of drought; the quantile is between (25,75], corresponding to the second level of drought; the quantile is between (5,25], corresponding to the third level of drought; the quantile is between (0,5], corresponding to the fourth level of drought. In specific implementation, the quantile is based on the distribution characteristics of the soil entropy value to adaptively define the level threshold (rather than a fixed threshold) to adapt to the baseline differences in different climate zones.
[0100] A preferred embodiment of a tea garden drought classification prediction system based on machine learning of the present invention includes the following modules:
[0101] A historical monitoring data acquisition module is used to collect a large amount of historical monitoring data from the tea garden through a sensor array including a humidity sensor, a light intensity sensor, a temperature sensor, a carbon dioxide sensor, an air pressure sensor, a particulate matter sensor, a global solar radiation sensor, a pH sensor, a salinity sensor, a conductivity sensor, a rainfall sensor, and a wind speed and direction sensor; the sensor array is also equipped with a locator;
[0102] A data set construction module, which constructs a data set after preprocessing each of the historical monitoring data, and performs a sample expansion operation on the data set;
[0103] An environmental factor correlation graph generation module is used to analyze the data set using the Limma algorithm, the COX survival regression model, and the GLM generalized linear model to obtain an environmental factor correlation graph;
[0104] A tea garden drought prediction model creation module is used to create a tea garden drought prediction model based on the sequentially connected multimodal feature extraction layer, feature fusion layer, soil entropy prediction layer, and drought mapping output layer, and set the loss function of the tea garden drought prediction model;
[0105] The multimodal feature extraction layer is used to extract time-dependent features, soil dynamic features and spatiotemporal features from the monitoring data; the feature fusion layer is used to adjust the weights of time-dependent features, soil dynamic features and spatiotemporal features according to the environmental factor correlation diagram, and fuse them to obtain multimodal features; the soil entropy prediction layer is used to calculate the soil entropy value based on the multimodal features; the drought mapping output layer is used to map the soil entropy value according to the set drought classification rules, and output the tea garden drought prediction result; the drought mapping output layer outputs the drought level probability (level 1 drought / level 2 drought / level 3 drought / level 4 drought) through the fully connected layer + Softmax, and introduces temperature scaling technology to calibrate the prediction confidence; the tea garden drought prediction result carries at least the drought period, drought level and drought location;
[0106] A tea garden drought prediction model deployment module is used to train the tea garden drought prediction model using the data set and deploy the trained tea garden drought prediction model; the AdamW optimizer is used in the training process, combined with cosine annealing learning rate scheduling;
[0107] The tea garden drought prediction module is used to predict tea garden drought conditions using the deployed tea garden drought prediction model. This module forms a closed loop from data collection, preprocessing, model training, to online prediction, supporting real-time drought monitoring and early warning.
[0108] By deploying a multi-dimensional sensor array, holographic collection of tea garden environmental data is achieved, and data enhancement strategies using hybrid outlier repair and generative adversarial networks are combined to improve data quality. The Limma algorithm, COX regression, and GLM models are innovatively integrated to screen key environmental factors and construct a spatiotemporal correlation map. A multimodal fusion prediction model is further designed, and heterogeneous features are extracted using bidirectional LSTM, residual networks, and spatiotemporal embedding layers. Dynamic weighted fusion is achieved through an attention mechanism and combined with dynamic quantile grading rules for soil entropy values to ultimately achieve high-precision drought level prediction. This model combines innovative cross-domain technology integration, model interpretability, and engineering deployment practicality, significantly improving the real-time performance and warning reliability of tea garden drought monitoring.
[0109] In the historical monitoring data acquisition module, the historical monitoring data at least includes maximum light intensity, average light intensity, sunshine duration, maximum temperature, minimum temperature, average temperature, daytime average temperature, nighttime average temperature, daily temperature difference, maximum air humidity, minimum air humidity, average relative humidity, daytime average air humidity, nighttime average air humidity, maximum carbon dioxide concentration, minimum carbon dioxide concentration, average carbon dioxide concentration, daytime average carbon dioxide concentration, nighttime average carbon dioxide concentration, maximum air pressure value, minimum air pressure value, average air pressure value, daytime average air pressure value, nighttime average air pressure value, air pressure difference, daily average PM2.5, daily average PM10, maximum TBQ total radiation, minimum TBQ total radiation, daily average TBQ total radiation, maximum soil temperature, minimum soil temperature, daily average soil temperature, soil temperature difference, maximum soil pH, minimum soil pH, average soil pH, maximum soil salinity value, minimum soil salinity value, average soil salinity value, maximum soil conductivity, minimum soil conductivity, average soil conductivity, rainfall, wind direction, wind speed, sampling location and sampling time.
[0110] By integrating 11 types of sensors, including meteorology (temperature, humidity, light, and rainfall), soil (PH, salinity, conductivity), and air quality (particulate matter and carbon dioxide), the physical, chemical, and biological dynamics of the tea garden environment are covered, providing multi-source heterogeneous data support for the model; by collecting multi-dimensional data such as extreme values (such as maximum / minimum temperature), average values (daily average temperature), difference values (daily temperature difference), and time segment average values (daytime / nighttime humidity), the dynamic changes of environmental factors can be accurately portrayed, effectively improving the generalization ability of the tea garden drought prediction model.
[0111] The dataset construction module is specifically used for:
[0112] After filling missing values and repairing outliers in each of the historical monitoring data, each of the historical monitoring data is labeled with soil entropy, drought level, and time segment labels to complete the preprocessing of each of the historical monitoring data. A data set is constructed based on the preprocessed historical monitoring data, and a sample expansion operation is performed on the data set using a generative adversarial network (GAN); the time segment label is the month of the sampling time;
[0113] The missing value filling is based on the K-nearest neighbor filling method; the outlier repair is based on the Z-score algorithm and the K-means clustering algorithm, specifically:
[0114] Perform Z-score standardization on each of the historical monitoring data, calculate the Z-score value of each of the historical monitoring data after Z-score standardization, and compare the Z-score value with a preset abnormality threshold to screen out abnormal values; cluster the historical monitoring data after Z-score standardization using a K-means clustering algorithm to obtain several cluster centers, and repair each abnormal value in turn based on the nearest cluster center.
[0115] Through the hybrid outlier repair mechanism, that is, combining Z-score anomaly detection and K-means clustering repair, it can not only identify statistical outliers, but also use cluster centers for spatial distribution calibration, which is more robust than a single method.
[0116] By adopting generative adversarial networks (GANs) for sample generation, the data is expanded while maintaining time series characteristics (such as seasonal cycles) and spatial correlations (such as the spatial distribution of soil parameters), effectively alleviating the problem of overfitting of small samples.
[0117] In addition to the drought level, soil entropy values (continuous variables) and month labels are additionally annotated to retain fine-grained supervision information and provide prior knowledge for time segment analysis.
[0118] The environmental factor correlation graph generation module is specifically used for:
[0119] Segmenting the data set based on time segmentation labels to obtain a number of data subsets, and screening significant factors between the data subsets whose change rates exceed a preset change threshold using the Limma algorithm;
[0120] Calculating a first drought risk level for each of the significant factors using a COX survival regression model, calculating a second drought risk level for each of the significant factors using a GLM generalized linear model, cross-validating the first drought risk level and the second drought risk level to screen N environmental factors with the greatest impact on drought conditions under different time segment labels from each of the significant factors, where N is a positive integer; and constructing an environmental factor correlation graph based on the screened environmental factors and time segment labels.
[0121] The Limma algorithm was used to screen significant factors with significant fluctuations between time segments (months) to capture seasonally sensitive indicators. The COX survival regression model was used to evaluate the temporal impact of each significant factor on the risk of drought and identify key drought-causing factors. The GLM generalized linear model was used to quantify the strength of the association between significant factors and drought levels and identify key drought-causing factors (screening indicators with strong spatial correlation). The output results of the COX survival regression model and the GLM generalized linear model were then cross-verified to ensure the temporal and spatial universality of the screened environmental factors and avoid single model bias.
[0122] The significant factors are maximum light intensity, average light intensity, sunshine duration, maximum temperature, minimum temperature, average temperature, average daytime temperature, average nighttime temperature, diurnal temperature difference, maximum air humidity, minimum air humidity, average relative humidity, average daytime air humidity, average nighttime air humidity, maximum carbon dioxide concentration, minimum carbon dioxide concentration, average carbon dioxide concentration, average daytime carbon dioxide concentration, average nighttime carbon dioxide concentration, maximum air pressure, minimum air pressure, average air pressure, average daytime air pressure, average nighttime air pressure, air pressure difference, daily average PM2.5, daily average PM10, maximum TBQ global radiation, minimum TBQ global radiation, daily average TBQ global radiation, maximum soil temperature, minimum soil temperature, daily average soil temperature, soil temperature difference, maximum soil pH, minimum soil pH, average soil pH, maximum soil salinity, minimum soil salinity, average soil salinity, maximum soil electrical conductivity, minimum soil electrical conductivity, average soil electrical conductivity, rainfall, wind direction, wind speed or sampling location.
[0123] In the tea garden drought prediction model creation module, the multimodal feature extraction layer is constructed based on the meteorological time series coding module, the soil dynamic coding module and the spatiotemporal embedding module;
[0124] The meteorological time series encoding module is used to extract time-dependent features (light, temperature and humidity, carbon dioxide, air pressure, PM, radiation, etc.) from the monitoring data through a bidirectional LSTM network. The soil dynamic encoding module is used to extract soil dynamic features (temperature, pH, salinity, and conductivity) from the monitoring data through a fully connected network with residual connections, introducing residual connections to enhance gradient propagation. The spatiotemporal embedding module is used to extract spatiotemporal features from the monitoring data through an embedding layer (mapping longitude and latitude into high-dimensional vectors and decomposing the sampling time into periodic codes of year / month / day / time period).
[0125] The bidirectional LSTM network is used to capture the long-term and short-term temporal dependencies of meteorological factors (such as the cumulative effects of continuous droughts); the residual fully connected network is used to model the dynamic nonlinear relationships of soil parameters (such as the coupled changes in salinity and electrical conductivity); and the spatiotemporal embedding module is used to encode the joint distribution characteristics of spatial position and time (such as the impact of slope aspect on water evaporation).
[0126] The feature fusion layer is used to adjust the weights of time-dependent features, soil dynamic features, and spatiotemporal features by combining the environmental factor correlation graph with a multi-head attention mechanism, and output multimodal features through 1D convolutional layer fusion;
[0127] The soil entropy prediction layer is used to infer multimodal features through a gated recurrent unit (GRU) and a feature decoupling head to calculate soil entropy values. The feature decoupling head outputs soil parameter prediction values through independent fully connected branches, achieving multi-task joint optimization.
[0128] The temporal context features are extracted through the GRU gating unit, and the decoupling head is combined to separate the deterministic component (such as current humidity) and the uncertain component (such as future rainfall forecast) of the soil entropy value to improve the robustness of the prediction.
[0129] By combining the Limma algorithm in biostatistics and the COX model in the medical field with deep learning, a new paradigm for agricultural data analysis is opened up; by complementing the advantages of traditional statistical models (GLM) and deep learning (GAN, Attention), it has both interpretability and predictive performance.
[0130] The specific drought classification rules are as follows:
[0131] The soil entropy value quantile is between (75,100], corresponding to the first level of drought; the quantile is between (25,75], corresponding to the second level of drought; the quantile is between (5,25], corresponding to the third level of drought; the quantile is between (0,5], corresponding to the fourth level of drought. In specific implementation, the quantile is based on the distribution characteristics of the soil entropy value to adaptively define the level threshold (rather than a fixed threshold) to adapt to the baseline differences in different climate zones.
[0132] In summary, the advantages of the present invention are:
[0133] 1. A large amount of historical monitoring data is collected from the tea garden through a sensor array including humidity sensors, light intensity sensors, temperature sensors, carbon dioxide sensors, air pressure sensors, particulate matter sensors, global solar radiation sensors, pH sensors, salinity sensors, conductivity sensors, rainfall sensors and wind speed and direction sensors. After preprocessing each historical monitoring data, a data set is constructed and sample expansion operations are performed; then the data set is analyzed by the Limma algorithm, COX survival regression model and GLM generalized linear model to obtain an environmental factor correlation graph; then a tea garden drought prediction model is created based on the sequentially connected multimodal feature extraction layer, feature fusion layer, soil entropy prediction layer and drought mapping output layer; the multimodal feature extraction layer is used to extract time-dependent features, soil dynamic features and spatiotemporal features from the monitoring data; the feature fusion layer is used to adjust the time-dependent features, soil dynamic features and spatiotemporal features according to the environmental factor correlation graph. The weights of spatiotemporal features are combined and multimodal features are obtained; the soil entropy prediction layer is used to calculate the soil entropy value based on the multimodal features; the drought mapping output layer is used to map the soil entropy value according to the set drought classification rules and output the tea garden drought prediction result; then the tea garden drought prediction model is trained through the data set, the trained tea garden drought prediction model is deployed, and finally the tea garden drought is predicted by the deployed tea garden drought prediction model; that is, the tea garden drought is predicted by the pre-trained tea garden drought prediction model. Before the prediction, the feature fusion layer of the tea garden drought prediction model calls the environmental factor correlation map to adjust the weight of each feature, that is, the weights of the features corresponding to the N environmental factors that have the greatest impact on the drought prediction under the current time segment label are increased, and the tea garden drought prediction result is output in combination with the drought classification rules, which facilitates the refined irrigation of the tea garden, and ultimately greatly improves the accuracy of the tea garden drought prediction and the refinement of the irrigation decision.
[0134] 2. By integrating 11 types of sensors, including meteorological (temperature, humidity, light, and rainfall), soil (pH, salinity, conductivity), and air quality (particulate matter and carbon dioxide), the physical, chemical, and biological dynamics of the tea garden environment are covered, providing multi-source heterogeneous data support for the model. By collecting multi-dimensional data covering extreme values (such as maximum / minimum temperature), mean (daily average temperature), difference (daily temperature difference), time segment mean (daytime / nighttime humidity), etc., the dynamic changes of environmental factors can be accurately portrayed, effectively improving the generalization ability of the tea garden drought prediction model.
[0135] 3. Through the hybrid outlier repair mechanism, that is, combining Z-score anomaly detection and K-means clustering repair, it can not only identify statistical outliers but also use cluster centers for spatial distribution calibration, which is more robust than a single method.
[0136] 4. By adopting generative adversarial networks (GANs) for sample generation, the data can be expanded while maintaining time series characteristics (such as seasonal cycles) and spatial correlations (such as the spatial distribution of soil parameters), effectively alleviating the problem of overfitting of small samples.
[0137] 5. In addition to the drought level, soil entropy value (continuous variable) and month label are additionally marked, which not only retains fine-grained supervision information but also provides prior knowledge for time segment analysis.
[0138] 6. The Limma algorithm was used to screen significant factors with significant fluctuations between time segments (months) to capture seasonally sensitive indicators. The COX survival regression model was used to assess the temporal impact of each significant factor on drought risk and identify key drought-causing factors. The GLM generalized linear model was used to quantify the strength of the association between significant factors and drought severity and identify key drought-causing factors (screening indicators with strong spatial correlation). The output results of the COX survival regression model and the GLM generalized linear model were then cross-verified to ensure the temporal and spatial universality of the screened environmental factors and avoid bias from a single model.
[0139] 7. By using the environmental factor correlation map as prior knowledge, the multi-head attention mechanism is guided to dynamically adjust feature weights (such as strengthening the sunshine weight in the drought season) to achieve knowledge-guided feature fusion; local feature interaction is carried out through the 1D convolution layer to enhance the collaborative expression ability of features in adjacent time steps.
[0140] 8. The temporal context features are extracted through the GRU gated unit, and the decoupling head is combined to separate the deterministic component (such as current humidity) and the uncertain component (such as future rainfall forecast) of the soil entropy value to improve the robustness of the prediction.
[0141] 9. By combining the Limma algorithm in biostatistics and the COX model in the medical field with deep learning, a new paradigm for agricultural data analysis is opened up; by complementing the advantages of traditional statistical models (GLM) and deep learning (GAN, Attention), both interpretability and predictive performance are achieved.
[0142] 10. Holographic collection of tea garden environmental data is achieved by deploying a multi-dimensional sensor array. The data enhancement strategy of hybrid outlier repair and generative adversarial network is combined to improve data quality. The Limma algorithm, COX regression and GLM model are innovatively integrated to screen key environmental factors and construct a spatiotemporal correlation map. A multimodal fusion prediction model is further designed. Bidirectional LSTM, residual network and spatiotemporal embedding layer are used to extract heterogeneous features. Dynamic weighted fusion is achieved through the attention mechanism and combined with the dynamic quantile grading rule of soil entropy value to finally achieve high-precision drought level prediction. It combines the innovation of cross-domain technology integration, model interpretability and engineering deployment practicality, and significantly improves the real-time performance and warning reliability of tea garden drought monitoring.
[0143] Although the specific embodiments of the present invention are described above, those skilled in the art should understand that the specific embodiments described are merely illustrative and are not intended to limit the scope of the present invention. Equivalent modifications and changes made by those skilled in the art in accordance with the spirit of the present invention should be included within the scope of protection of the claims of the present invention.
Claims
1. A tea garden drought classification prediction method based on machine learning, characterized by: The steps include: Step S1, collecting a large amount of historical monitoring data from the tea garden through a sensor array including a humidity sensor, a light intensity sensor, a temperature sensor, a carbon dioxide sensor, an air pressure sensor, a particulate matter sensor, a global solar radiation sensor, a pH sensor, a salinity sensor, a conductivity sensor, a rainfall sensor, and a wind speed and direction sensor; Step S2: constructing a data set after preprocessing each of the historical monitoring data, and performing a sample expansion operation on the data set; Step S3, analyzing the data set using the Limma algorithm, the COX survival regression model, and the GLM generalized linear model to obtain an environmental factor correlation graph; Step S4: creating a tea garden drought prediction model based on the sequentially connected multimodal feature extraction layer, feature fusion layer, soil entropy prediction layer, and drought mapping output layer, and setting a loss function of the tea garden drought prediction model; The multimodal feature extraction layer is used to extract time-dependent features, soil dynamic features, and spatiotemporal features from the monitoring data; the feature fusion layer is used to adjust the weights of time-dependent features, soil dynamic features, and spatiotemporal features according to the environmental factor correlation diagram, and fuse them to obtain multimodal features; the soil entropy prediction layer is used to calculate the soil entropy value based on the multimodal features; the drought mapping output layer is used to map the soil entropy value according to the set drought classification rules and output the tea garden drought prediction result; Step S5: training a tea garden drought prediction model using the data set, and deploying the trained tea garden drought prediction model; Step S6: predicting the tea garden drought by using the deployed tea garden drought prediction model.
2. The method for predicting tea garden drought conditions based on machine learning according to claim 1, wherein: In step S1, the historical monitoring data at least includes maximum light intensity, average light intensity, sunshine duration, maximum temperature, minimum temperature, average temperature, daytime average temperature, nighttime average temperature, daily temperature difference, maximum air humidity, minimum air humidity, average relative humidity, daytime average air humidity, nighttime average air humidity, maximum carbon dioxide concentration, minimum carbon dioxide concentration, average carbon dioxide concentration, daytime average carbon dioxide concentration, nighttime average carbon dioxide concentration, maximum air pressure value, minimum air pressure value, average air pressure value, daytime average air pressure value, nighttime average air pressure value, air pressure difference, daily average PM2.5, daily average PM10, maximum TBQ total radiation, minimum TBQ total radiation, daily average TBQ total radiation, maximum soil temperature, minimum soil temperature, daily average soil temperature, soil temperature difference, maximum soil pH value, minimum soil pH value, average soil pH value, maximum soil salinity value, minimum soil salinity value, average soil salinity value, maximum soil conductivity, minimum soil conductivity, average soil conductivity, rainfall, wind direction, wind speed, sampling location and sampling time.
3. The method for predicting tea garden drought conditions based on machine learning according to claim 1, wherein: The step S2 is specifically as follows: After filling missing values and repairing outliers for each of the historical monitoring data, each of the historical monitoring data is labeled with soil entropy, drought level, and time segment labels to complete the preprocessing of each of the historical monitoring data. A data set is constructed based on the preprocessed historical monitoring data, and a sample expansion operation is performed on the data set through a generative adversarial network; the time segment label is the month of the sampling time; The missing value filling is based on the K-nearest neighbor filling method; the outlier repair is based on the Z-score algorithm and the K-means clustering algorithm, specifically: Perform Z-score standardization on each of the historical monitoring data, calculate the Z-score value of each of the historical monitoring data after Z-score standardization, and compare the Z-score value with a preset abnormality threshold to screen out abnormal values; cluster the historical monitoring data after Z-score standardization using a K-means clustering algorithm to obtain several cluster centers, and repair each abnormal value in turn based on the nearest cluster center.
4. The method for predicting tea garden drought conditions based on machine learning according to claim 1, wherein: The step S3 is specifically as follows: Segmenting the data set based on time segmentation labels to obtain a number of data subsets, and screening significant factors between the data subsets whose change rates exceed a preset change threshold using the Limma algorithm; Calculating a first drought risk level for each of the significant factors using a COX survival regression model, calculating a second drought risk level for each of the significant factors using a GLM generalized linear model, cross-validating the first drought risk level and the second drought risk level to screen N environmental factors with the greatest impact on drought conditions under different time segment labels from each of the significant factors, where N is a positive integer; and constructing an environmental factor correlation graph based on the screened environmental factors and time segment labels. The significant factors are maximum light intensity, average light intensity, sunshine duration, maximum temperature, minimum temperature, average temperature, average daytime temperature, average nighttime temperature, diurnal temperature difference, maximum air humidity, minimum air humidity, average relative humidity, average daytime air humidity, average nighttime air humidity, maximum carbon dioxide concentration, minimum carbon dioxide concentration, average carbon dioxide concentration, average daytime carbon dioxide concentration, average nighttime carbon dioxide concentration, maximum air pressure, minimum air pressure, average air pressure, average daytime air pressure, average nighttime air pressure, air pressure difference, daily average PM2.5, daily average PM10, maximum TBQ global radiation, minimum TBQ global radiation, daily average TBQ global radiation, maximum soil temperature, minimum soil temperature, daily average soil temperature, soil temperature difference, maximum soil pH, minimum soil pH, average soil pH, maximum soil salinity, minimum soil salinity, average soil salinity, maximum soil electrical conductivity, minimum soil electrical conductivity, average soil electrical conductivity, rainfall, wind direction, wind speed or sampling location.
5. The method for predicting tea garden drought conditions based on machine learning according to claim 1, wherein: In step S4, the multimodal feature extraction layer is constructed based on the meteorological time series encoding module, the soil dynamic encoding module and the spatiotemporal embedding module; The meteorological time series encoding module is used to extract time-dependent features from the monitoring data through a bidirectional LSTM network; the soil dynamic encoding module is used to extract soil dynamic features from the monitoring data through a fully connected network with residual connections; the spatiotemporal embedding module is used to extract spatiotemporal features from the monitoring data through an embedding layer; The feature fusion layer is used to adjust the weights of time-dependent features, soil dynamic features, and spatiotemporal features by combining the environmental factor correlation graph with a multi-head attention mechanism, and output multimodal features through 1D convolutional layer fusion; The soil entropy prediction layer is used to infer multimodal features through gated recurrent units and feature decoupling heads to calculate soil entropy values; The specific drought classification rules are as follows: The quantile of soil entropy value is (75,100], which corresponds to the first level of drought; the quantile is (25,75], which corresponds to the second level of drought; the quantile is (5,25], which corresponds to the third level of drought; the quantile is (0,5], which corresponds to the fourth level of drought.
6. A tea garden drought classification prediction system based on machine learning, characterized by: Includes the following modules: A historical monitoring data acquisition module is used to collect a large amount of historical monitoring data from the tea garden through a sensor array including a humidity sensor, a light intensity sensor, a temperature sensor, a carbon dioxide sensor, an air pressure sensor, a particulate matter sensor, a global solar radiation sensor, a pH sensor, a salinity sensor, a conductivity sensor, a rainfall sensor, and a wind speed and direction sensor; A data set construction module, which constructs a data set after preprocessing each of the historical monitoring data, and performs a sample expansion operation on the data set; An environmental factor correlation graph generation module is used to analyze the data set using the Limma algorithm, the COX survival regression model, and the GLM generalized linear model to obtain an environmental factor correlation graph; A tea garden drought prediction model creation module is used to create a tea garden drought prediction model based on the sequentially connected multimodal feature extraction layer, feature fusion layer, soil entropy prediction layer, and drought mapping output layer, and set the loss function of the tea garden drought prediction model; The multimodal feature extraction layer is used to extract time-dependent features, soil dynamic features, and spatiotemporal features from the monitoring data; the feature fusion layer is used to adjust the weights of time-dependent features, soil dynamic features, and spatiotemporal features according to the environmental factor correlation diagram, and fuse them to obtain multimodal features; the soil entropy prediction layer is used to calculate the soil entropy value based on the multimodal features; the drought mapping output layer is used to map the soil entropy value according to the set drought classification rules and output the tea garden drought prediction result; A tea garden drought prediction model deployment module is used to train the tea garden drought prediction model using the data set and deploy the trained tea garden drought prediction model; The tea garden drought prediction module is used to predict the tea garden drought by using the deployed tea garden drought prediction model.
7. The tea garden drought classification prediction system based on machine learning according to claim 6, characterized in that: In the historical monitoring data acquisition module, the historical monitoring data at least includes maximum light intensity, average light intensity, sunshine duration, maximum temperature, minimum temperature, average temperature, daytime average temperature, nighttime average temperature, daily temperature difference, maximum air humidity, minimum air humidity, average relative humidity, daytime average air humidity, nighttime average air humidity, maximum carbon dioxide concentration, minimum carbon dioxide concentration, average carbon dioxide concentration, daytime average carbon dioxide concentration, nighttime average carbon dioxide concentration, maximum air pressure value, minimum air pressure value, average air pressure value, daytime average air pressure value, nighttime average air pressure value, air pressure difference, daily average PM2.5, daily average PM10, maximum TBQ total radiation, minimum TBQ total radiation, daily average TBQ total radiation, maximum soil temperature, minimum soil temperature, daily average soil temperature, soil temperature difference, maximum soil pH, minimum soil pH, average soil pH, maximum soil salinity value, minimum soil salinity value, average soil salinity value, maximum soil conductivity, minimum soil conductivity, average soil conductivity, rainfall, wind direction, wind speed, sampling location and sampling time.
8. The tea garden drought classification prediction system based on machine learning according to claim 6, characterized in that: The dataset construction module is specifically used for: After filling missing values and repairing outliers for each of the historical monitoring data, each of the historical monitoring data is labeled with soil entropy, drought level, and time segment labels to complete the preprocessing of each of the historical monitoring data. A data set is constructed based on the preprocessed historical monitoring data, and a sample expansion operation is performed on the data set through a generative adversarial network; the time segment label is the month of the sampling time; The missing value filling is based on the K-nearest neighbor filling method; the outlier repair is based on the Z-score algorithm and the K-means clustering algorithm, specifically: Perform Z-score standardization on each of the historical monitoring data, calculate the Z-score value of each of the historical monitoring data after Z-score standardization, and compare the Z-score value with a preset abnormality threshold to screen out abnormal values; cluster the historical monitoring data after Z-score standardization using a K-means clustering algorithm to obtain several cluster centers, and repair each abnormal value in turn based on the nearest cluster center.
9. The tea garden drought classification prediction system based on machine learning according to claim 6, characterized in that: The environmental factor correlation graph generation module is specifically used for: Segmenting the data set based on time segmentation labels to obtain a number of data subsets, and screening significant factors between the data subsets whose change rates exceed a preset change threshold using the Limma algorithm; Calculating a first drought risk level for each of the significant factors using a COX survival regression model, calculating a second drought risk level for each of the significant factors using a GLM generalized linear model, cross-validating the first drought risk level and the second drought risk level to screen N environmental factors with the greatest impact on drought conditions under different time segment labels from each of the significant factors, where N is a positive integer; and constructing an environmental factor correlation graph based on the screened environmental factors and time segment labels. The significant factors are maximum light intensity, average light intensity, sunshine duration, maximum temperature, minimum temperature, average temperature, average daytime temperature, average nighttime temperature, diurnal temperature difference, maximum air humidity, minimum air humidity, average relative humidity, average daytime air humidity, average nighttime air humidity, maximum carbon dioxide concentration, minimum carbon dioxide concentration, average carbon dioxide concentration, average daytime carbon dioxide concentration, average nighttime carbon dioxide concentration, maximum air pressure, minimum air pressure, average air pressure, average daytime air pressure, average nighttime air pressure, air pressure difference, daily average PM2.5, daily average PM10, maximum TBQ global radiation, minimum TBQ global radiation, daily average TBQ global radiation, maximum soil temperature, minimum soil temperature, daily average soil temperature, soil temperature difference, maximum soil pH, minimum soil pH, average soil pH, maximum soil salinity, minimum soil salinity, average soil salinity, maximum soil electrical conductivity, minimum soil electrical conductivity, average soil electrical conductivity, rainfall, wind direction, wind speed or sampling location.
10. The tea garden drought classification prediction system based on machine learning according to claim 6, characterized in that: In the tea garden drought prediction model creation module, the multimodal feature extraction layer is constructed based on the meteorological time series coding module, the soil dynamic coding module and the spatiotemporal embedding module; The meteorological time series encoding module is used to extract time-dependent features from the monitoring data through a bidirectional LSTM network; the soil dynamic encoding module is used to extract soil dynamic features from the monitoring data through a fully connected network with residual connections; the spatiotemporal embedding module is used to extract spatiotemporal features from the monitoring data through an embedding layer; The feature fusion layer is used to adjust the weights of time-dependent features, soil dynamic features, and spatiotemporal features by combining the environmental factor correlation graph with a multi-head attention mechanism, and output multimodal features through 1D convolutional layer fusion; The soil entropy prediction layer is used to infer multimodal features through gated recurrent units and feature decoupling heads to calculate soil entropy values; The specific drought classification rules are as follows: The quantile of soil entropy value is (75,100], which corresponds to the first level of drought; the quantile is (25,75], which corresponds to the second level of drought; the quantile is (5,25], which corresponds to the third level of drought; the quantile is (0,5], which corresponds to the fourth level of drought.
Citation Information
Patent Citations
Biomarker and model for prognosis risk prediction of colorectal cancer and application of biomarker and model
CN116030880A
Tea garden state monitoring method fusing visual time sequence text pre-training model
CN119046673A
Tea garden drought prediction method based on climate environment monitoring data
CN119204726A