Wheat scab risk prediction system and method based on deep learning and random forest

By combining the multimodal data fusion prediction system with LSTM neural network and random forest algorithm, nonlinear modeling and real-time problems of wheat gibberellosis prediction in the existing technology are solved, accurate prediction and prevention and control suggestions for wheat gibberellosis are realized, and efficiency and accuracy of agricultural disaster prediction are improved.

CN120543316APending Publication Date: 2025-08-26HENAN ZHAODI ELECTRONIC TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510610222.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

The existing agricultural disaster prediction methods have insufficient modeling capabilities for nonlinear time series, lack of dynamic weight allocation capabilities for a single deep learning model, joint analysis of risk assessment without combining historical data and prediction data, low model deployment efficiency, and cannot meet the real-time prediction needs. The occurrence of wheat gibberellia is closely related to meteorological conditions, and it is necessary to accurately predict and timely prevent and control it.

Method used

A multimodal data fusion prediction system based on LSTM neural network and random forest algorithm is adopted, combined with meteorological conditions and historical data, long-term and short-term memory model is modeled through the LSTM model, and risk assessment is performed, and the attention mechanism is used to improve feature capture ability to form an accurate wheat gibberellia risk prediction system.

Benefits of technology

Accurate prediction of wheat gibberellia risks is achieved, targeted prevention and control suggestions are provided, model deployment efficiency and prediction accuracy are improved, and a closed-loop feedback system is formed to optimize agricultural management strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120543316A_ABST
    Figure CN120543316A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of agricultural intelligent monitoring, in particular to a wheat scab risk prediction system and method based on deep learning and random forest. Data acquisition: transmitting historical meteorological data of past 30 days through an interface; preprocessing, wherein the system cleans and standardizes the data to ensure the data quality; model prediction and result output: combining a prediction result of the LSTM model with original historical data, and calculating a future risk probability through a random forest model; the invention discloses a multi-modal data fusion prediction system based on an LSTM neural network, an attention mechanism and a random forest algorithm. The multi-modal data fusion prediction system is used for wheat scab risk assessment and weather parameter prediction. And in combination with an LSTM neural network, an attention mechanism and a random forest algorithm, accurate prediction of the risk of the wheat scab is realized, and targeted prevention and control suggestions are provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of agricultural intelligent monitoring technology, and specifically to a wheat fusarium head blight risk prediction system and method based on deep learning and random forest. Background Art

[0002] Existing agricultural disaster prediction methods have the following defects:

[0003] 1. Traditional statistical models (such as ARIMA) are not capable of modeling nonlinear time series;

[0004] 2. Single deep learning models lack the ability to dynamically assign weights to key features;

[0005] 3. Risk assessments fail to incorporate joint analysis of historical and forecast data;

[0006] 4. Model deployment efficiency is low and cannot meet real-time prediction needs.

[0007] Furthermore, the occurrence of wheat fusarium head blight is closely related to meteorological conditions. Wheat in the middle and lower reaches of the Yangtze River reaches its heading and flowering stage from mid-to-late April to early May, a critical stage for wheat yield formation and a period of susceptibility to fusarium head blight. Therefore, accurately predicting fusarium head blight risks and implementing timely prevention and control measures are crucial. Summary of the Invention

[0008] To address the challenges of the existing technology, this paper provides a wheat fusarium head blight risk prediction system and method based on deep learning and random forests. Specifically, it involves a multimodal data fusion prediction system based on an LSTM neural network, an attention mechanism, and a random forest algorithm, for wheat fusarium head blight risk assessment and weather parameter forecasting.

[0009] To achieve the above objectives, the present invention employs the following technical solution: Fusarium head blight spores favor meteorological conditions with daily average temperatures exceeding 15°C, frequent rainy days with humidity exceeding 80%, and illumination less than 2000 Lux. When these conditions persist for more than three days, the likelihood of a major outbreak increases significantly.

[0010] A wheat scab risk prediction system based on deep learning and random forest, comprising: a user end;

[0011] Forecast API: used to receive historical meteorological data submitted by users and pass this data to the data preprocessing module; historical meteorological data of the past 30 days is input through the interface, and users submit historical meteorological data and other data to the forecast API.

[0012] Data preprocessing:

[0013] Clean and standardize the received data and convert the processed data into a 30-day feature sequence suitable for model input;

[0014] LSTM model:

[0015] Use the long short-term memory network to model the 30-day feature sequence and output the forecast data for the next 7 days;

[0016] Random Forest Model:

[0017] The random forest model is an ensemble learning method that improves prediction accuracy by building multiple decision trees and combining their prediction results, performing risk assessment, and outputting risk probability distribution;

[0018] Model deployment module:

[0019] Integrate the results from different models to form the final forecast result, merge the 7-day forecast data output by the LSTM model with historical data, and convert the model output into a readable form for users to understand and apply.

[0020] Return result:

[0021] The prediction results of the LSTM model are combined with the original historical data, and the future risk probability is calculated through the random forest model. The system returns the structured prediction results to the user end.

[0022] Users can take appropriate measures based on these predictions, such as adjusting agricultural management strategies to reduce the risk of wheat scab;

[0023] Feedback loop:

[0024] After receiving the prediction results, the user can submit new data again through the prediction API to form a closed-loop system.

[0025] The preferred LSTM state update equation is:

[0026] i t =σ(W xi x t +W hi h t-1 +b i )

[0027] f t =σ(W xf x t +W hf h t-1 +b f )

[0028] o t =σ(W xo xt +W ho h t-1 +b o )

[0029]

[0030] h t =o t tanh(C t )

[0031] in:

[0032] x t : Input feature vector (input at time step t).

[0033] h t-1 : The hidden state at the previous time step.

[0034] C t-1 : The cell state at the previous time step.

[0035] W xi ,W xf ,W xo ,W xc : Weight matrix input to the gating unit.

[0036] W hi ,W hf ,W ho ,W hc : The weight matrix from hidden state to gate unit.

[0037] b i ,b f ,b o ,b c : Bias term.

[0038] σ: Sigmoid activation function, used to calculate the gate value.

[0039] tanh: Hyperbolic tangent activation function, used to generate candidate cell states.

[0040] Preferred two-layer LSTM structure:

[0041] The output dimension of the first layer is D, and the output dimension of the second layer is D / 2.

[0042] The preferred attention mechanism formula is as follows:

[0043] Q=AW Q ,K=AW K ,V=AW V

[0044]

[0045] in:

[0046] A: Hidden state matrix output by LSTM.

[0047] W Q ,W K ,W V : A learnable parameter matrix used to generate queries, keys, and values, respectively.

[0048] Q: Query Matrix

[0049] K: bond matrix.

[0050] V: value matrix.

[0051] d k : The dimension of the key vector, used to scale the dot product result.

[0052] The preferred risk probability calculation formula is as follows:

[0053]

[0054] Where: Φ(x): Feature representation after random forest transforms the input features.

[0055] w: weight vector for logistic regression.

[0056] This formula is a standard logistic regression probability calculation formula.

[0057] The method of the wheat fusarium wilt risk prediction system based on deep learning and random forest includes the following steps: data collection, which involves inputting historical meteorological data from the past 30 days through an interface; preprocessing, in which the system cleans and standardizes the data to ensure data quality; model prediction and result output, in which the prediction results of the LSTM model are combined with the original historical data, and the future risk probability is calculated through a random forest model.

[0058] The method of the wheat fusarium wilt risk prediction system based on deep learning and random forest takes the seven meteorological parameters of the last five days of historical data, standardizes the historical data, merges the standardized historical data and the predicted data, and then marks the merged data to see whether it meets the single-day conditions of average temperature exceeding 15°C, rainy weather, air humidity greater than 80% RH, and illumination less than 2000 Lux. If these conditions are met, it is marked as risk data.

[0059] The method of the wheat fusarium scabra risk prediction system based on deep learning and random forest counts the number of consecutive days that meet the conditions starting from the current prediction day. If the conditions are not met, the counting is terminated. The random forest model algorithm is used to predict the risk probability. The risk level is divided into three levels: high, medium, and low. If the predicted data meets all meteorological conditions, it will be judged as medium or high risk. The longer the consecutive rainy days, the greater the probability of risk. The risk prediction probability is greater than 0.5 and less than 0.8 for medium risk, and less than 0.5 for low risk.

[0060] Compared with the existing technology, the beneficial effect of the invention is: the present invention proposes a dual-model collaborative prediction system, which combines the LSTM neural network, attention mechanism and random forest algorithm to achieve accurate prediction of wheat fusarium wilt risk and provide targeted prevention and control recommendations. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Other features, objects and advantages of the present invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings.

[0062] Figure 1 : Flow chart of meteorological data input of the present invention;

[0063] Figure 2 : System architecture diagram of the present invention;

[0064] Figure 3 : LSTM-Attention model structure diagram of the present invention;

[0065] Figure 4 : Data preprocessing flow chart of the present invention;

[0066] Figure 5 : Risk assessment flow chart of the present invention;

[0067] Figure 6 : The overall implementation architecture diagram of the system of the present invention.

[0068] Figure 7 : The present invention reads the original data graph.

[0069] Figure 8 : The final output result diagram of the present invention.

[0070] Figure 9 : The present invention normalizes the complete 30-day feature data.

[0071] Figure 10 : The weight matrix diagram of the present invention input to the gate control unit.

[0072] Figure 11 : Weight matrix diagram from hidden state to gate unit of the present invention.

[0073] Figure 12 : The weight distribution diagram of the second layer LSTM input to the gate and hidden to the gate.

[0074] Figure 13 : The output gate activation value distribution histogram of the present invention Figure 1 .

[0075] Figure 14 : The output gate activation value distribution histogram of the present invention Figure 2 .

[0076] Figure 15 : Input map of the attention layer of the present invention.

[0077] Figure 16 : The query weight W of the present invention Q picture.

[0078] Figure 17 : The query matrix Q graph of the present invention.

[0079] Figure 18 : Key weight W of the present invention K [64,64]Fig.

[0080] Figure 19 : The bond matrix K[30,64] diagram of the present invention.

[0081] Figure 20 : The weight of the present invention W V [64,64]Fig.

[0082] Figure 21 : The value matrix V[30,64] diagram of the present invention.

[0083] Figure 22 : The Q matrix value distribution, K matrix value distribution and V matrix value distribution diagram of the present invention.

[0084] Figure 23 : Example diagram of the prediction results of the present invention. DETAILED DESCRIPTION

[0085] The present invention is further described in detail below by way of examples. The examples are only used to illustrate the present invention and are not intended to limit the scope of the present invention.

[0086] In the solution of the present invention, Figure 2 The overall architecture of the system is demonstrated, including data collection, which uses an interface to input historical meteorological data from the past 30 days; preprocessing, in which the system cleans and standardizes the data to ensure data quality; model prediction and result output, in which the prediction results of the LSTM model are combined with the original historical data, and the random forest model is used to calculate future risk probabilities.

[0087] Figure 3This demonstrates how LSTM is combined with the attention mechanism. For each time step's hidden state, attention weights are calculated, and the weighted sum is used to obtain a context vector, which is then fed into the next layer or used for prediction.

[0088] Figure 4 This section demonstrates the data completion and feature extraction process. Organizing historical data into time series requires normalization because different features have significantly different dimensions, such as illuminance and temperature.

[0089] Figure 5 Demonstrate the complete process from meteorological data to ergot risk assessment.

[0090] Figure 6 Demonstrates the workflow of the system in actual application scenarios. The user submits historical meteorological data, which is processed through the interface data and then passed through the prediction model network to finally return the predicted risk result data to the user.

[0091] The wheat scab risk prediction system based on deep learning and random forest includes: a user end, which is responsible for submitting historical meteorological data and processing the returned results;

[0092] Forecast API: used to receive historical meteorological data submitted by users and pass this data to the data preprocessing module; historical meteorological data of the past 30 days is input through the interface, and users submit historical meteorological data and other data to the forecast API.

[0093] Data preprocessing:

[0094] The received data is cleaned and standardized to ensure its accuracy and consistency.

[0095] The processed data is converted into a feature sequence suitable for model input, such as a 30-day feature sequence.

[0096] LSTM model:

[0097] A long short-term memory (LSTM) network is used to model the 30-day feature sequence.

[0098] The LSTM model is able to capture long-term dependencies in time series, thereby predicting future trends more accurately.

[0099] Output forecast data for the next 7 days.

[0100] Multi-layer LSTM attention prediction model:

[0101] The state update equation of LSTM is:

[0102] i t =σ(W xi x t+W hi h t-1 +b i )

[0103] f t =σ(W xf x t +W hf h t-1 +b f )

[0104] o t =σ(W xo x t +W ho h t-1 +b o )

[0105]

[0106] h t =o t tanh(C t )

[0107] illustrate:

[0108] x t : Input feature vector (input at time step t), such as weather data such as temperature and humidity on a certain day.

[0109] h t-1 : The hidden state at the previous time step; this is the information that the LSTM model retains after processing the previous day's data.

[0110] C t-1 : The cell state at the previous time step; this is the long-term memory of the LSTM model, which retains past information.

[0111] W xi ,W xf ,W xo ,W xc : Weight matrix input to the gating unit.

[0112] W hi ,W hf ,W ho ,W hc : The weight matrix from hidden state to gate unit.

[0113] b i ,b f ,b o ,b c : Bias term.

[0114] σ: Sigmoid activation function, used to calculate the gate value; decide which information needs to be retained or discarded.

[0115] Tanh: Hyperbolic tangent activation function, used to generate candidate cell states, i.e. new information. Weather data is processed and memorized through the state update equation, leading to more accurate predictions.

[0116] Two-layer LSTM structure:

[0117] The output dimension of the first layer is D, and the output dimension of the second layer is D / 2.

[0118] Attention mechanism calculation

[0119] The core formula of the attention mechanism is as follows:

[0120] Q=AW Q ,K=AW K ,V=AW V

[0121]

[0122] illustrate:

[0123] A: Hidden state matrix output by LSTM.

[0124] W Q ,W K ,W V : A learnable parameter matrix used to generate queries, keys, and values, respectively.

[0125] Q: Query Matrix

[0126] K: bond matrix.

[0127] V: value matrix.

[0128] d k : The dimension of the key vector, used to scale the dot product results. The attention mechanism determines the importance of each key's corresponding value by calculating the similarity between the query and the key (usually using a dot product). These similarities are then normalized using a softmax function to obtain a weight for each value. Finally, these weighted values ​​are summed to obtain the output of the attention mechanism.

[0129] This method uses a 4-head attention mechanism; the attention mechanism is calculated independently multiple times (4 times), each time using a different learnable parameter matrix. The purpose of this is to allow the model to focus on information from multiple different perspectives, thereby capturing richer features. Finally, the outputs of these 4 attention mechanisms are concatenated or averaged as the final attention output;

[0130] The multi-head attention mechanism works as follows:

[0131] Split and Transform: First, the input sequence is split into multiple heads, each of which has a different transformation. This means that each head focuses on different aspects of the input sequence.

[0132] Scaled Dot-Product Attention: In each head, a scaled dot-product attention mechanism is used to calculate the attention weights between each word and every other word. This is achieved by calculating the dot product between the query, key, and value, and then applying a softmax function.

[0133] Head merging: Concatenate the attention outputs of all heads and then merge them into a single attention tensor through a linear transformation.

[0134] Residual Connection and Layer Normalization: Finally, the output of the multi-head attention mechanism is residually connected to the input, and layer normalization is applied to improve the stability of the model.

[0135] Random Forest Model: The random forest model receives the prediction results of the LSTM model and historical data, performs risk assessment through 100 decision trees with a maximum depth of 10, and outputs the risk probability distribution.

[0136] The random forest model is an ensemble learning method that improves prediction accuracy by building multiple decision trees and combining their predictions. This was done to evaluate the wheat fusarium head blight risk predicted by the LSTM model.

[0137] The random forest model performs risk assessment and outputs a risk probability distribution.

[0138] Risk probability calculation, where the risk probability calculation formula is as follows:

[0139]

[0140] Where: Φ(x): Feature representation after random forest transforms the input features.

[0141] w: weight vector for logistic regression.

[0142] This formula is a standard logistic regression probability calculation formula.

[0143] Model deployment module: responsible for deploying the trained model to the production environment and providing continuous services.

[0144] The model deployment module integrates the results from different models, merging the 7-day forecast data output by the LSTM model with historical data; this data is then input into the random forest model for risk assessment to form the final forecast result.

[0145] The model deployment module converts the output of the model into a readable form for users to understand and apply, and can obtain the probability that each sample belongs to a certain category (such as the high-risk group for wheat fusarium wilt).

[0146] A model deployment module typically consists of the following key components, which work together to enable a seamless transition of models from development to production environments:

[0147] Model Repository:

[0148] Used to store different versions of trained models.

[0149] Often integrated with a version control system such as Git to track the change history of the model.

[0150] Build and Test Automation:

[0151] Automated scripts to build the model environment, including installing necessary dependent libraries and frameworks.

[0152] Automatically run unit and integration tests to ensure models meet expected performance and quality standards before deployment.

[0153] Deployment Strategies and Tools:

[0154] Tools for implementing strategies such as blue-green deployment, rolling updates, and canary releases.

[0155] Return result:

[0156] The LSTM model's prediction results are combined with the original historical data, and the random forest model is used to calculate future risk probabilities. The system returns the structured prediction results to the user.

[0157] Users can take appropriate measures based on these prediction results, such as adjusting agricultural management strategies to reduce the risk of wheat fusarium wilt.

[0158] Feedback loop:

[0159] After receiving the forecast results, the user can submit new data again through the forecast API, forming a closed-loop system. Users can continuously optimize their agricultural management strategies based on previous forecast results and actual conditions, and update the model by submitting new data to further improve forecast accuracy.

[0160] This feedback mechanism helps to continuously optimize model performance and improve prediction accuracy. As shown in Table 1, taking the meteorological data for April 2023 as an example, Table 1 shows the meteorological data for April 2023:

[0161]

[0162] Table 1

[0163] Read historical data. The data of 2023-04-03, 2023-04-04 and 2023-04-10 are missing data. Use linear interpolation data completion algorithm to fill in the missing data. Read the original data as follows Figure 7 As shown:

[0164] Data completion algorithm, in which the linear interpolation formula

[0165]

[0166] illustrate:

[0167] f(x - ): The previous valid value of a known data point

[0168] f(x + ): The next valid value of a known data point.

[0169] x - : The time point of the previous known data.

[0170] x + : The time point at which the last data is known.

[0171] x: The time point that needs to be completed.

[0172] Calculate missing values ​​according to the linear interpolation formula:

[0173] x - =2023-04-02,f(x - )=18.8;

[0174] x + =2023-04-05,f(x + )=12.7;

[0175] Time interval:

[0176] x + -x - =3 days;

[0177] Linearly interpolate the missing dates x = 2023-04-03 and x = 2023-04-04

[0178] Table 2 shows the date and air temperature data.

[0179]

[0180] Table 2

[0181] Similarly, the interpolation of other features can be calculated, and the final output result is as follows Figure 8 As shown:

[0182] After the date completion and padding operations are completed, such as Figure 9 The complete 30-day feature data is normalized as shown. The data is fed into the neural network to obtain:

[0183] LSTM1 output dimension: [1, 30, 128]

[0184] LSTM1 last time step cell state: [1,128]

[0185] LSTM2 output dimension: [1, 30, 64]

[0186] LSTM2 cell state: [1,64]

[0187] Attention weight shape: [1, 30, 30]

[0188] The weight matrix input to the gate unit is [512,7], such as Figure 10 shown.

[0189] The weight matrix from hidden state to gate unit: [512,128], such as Figure 11 shown.

[0190] The second layer LSTM input to gate and hidden to gate weight distribution diagram, such as Figure 12 shown.

[0191] The output gate activation value distribution histogram is as follows: Figure 13 、 14 As shown;

[0192] The calculation of the attention mechanism needs to adjust the attention layer input to (seq_len, batch, embed_dim)

[0193] seq_len: sequence length. The input data length here is 30 days of historical data

[0194] Batch: Batch size. Here we use 1 batch input

[0195] embed_dim: The length of the feature vector of each element. Here, half of the hidden layer is 128 / 2

[0196] The input of the attention layer is: [30,1,64], such as Figure 15 shown.

[0197] Query weight W Q is [64,64] and the query matrix Q is [30,64], such as Figure 16 、 Figure 17 shown.

[0198] Key weight W K [64,64] and the key matrix K[30,64], such as Figure 18 、 Figure 19 shown.

[0199] Value weight W V [64,64] and value matrix V[30,64], such as Figure 20 、 Figure 21 shown.

[0200] The Q matrix value distribution, K matrix value distribution, and V matrix value distribution are as follows: Figure 22 As shown;

[0201] According to the characteristics of meteorological data, relative humidity, wind speed, and illuminance cannot have negative values. Therefore, constraints are added to the network structure and denormalization, and the output of the model is controlled in the range of 0 to 1. The prediction results are as follows: Figure 23 As shown;

[0202] Risk assessment of the forecast data reveals that the meteorological conditions that influence the occurrence of wheat scab are favorable for scab spores: daily average temperatures exceeding 15°C, frequent rainy days, humidity exceeding 80% RH, and illumination below 2000 Lux. Rainy weather is particularly favorable for the development of scab. When these conditions persist for more than three days, spores become active, and the longer they persist, the greater the likelihood of a major outbreak.

[0203] Take the seven meteorological parameters characteristic of the last five days of historical data, standardize the historical data, merge the standardized historical data and forecast data, and then mark the merged data to see whether it meets the daily conditions of average temperature exceeding 15°C, rainy weather, air humidity greater than 80% RH, and illumination less than 2000 Lux. If these conditions are met, it is marked as risk data.

[0204] The number of consecutive days that meet the conditions starting from the current forecast day is counted. If the conditions are not met, the counting is stopped. The random forest model algorithm is used to predict the risk probability. The risk level is divided into three levels: high, medium, and low. If the forecast data meets all meteorological conditions, it will be judged as medium or high risk. The longer the consecutive rainy days, the greater the probability of risk. The risk prediction probability is greater than 0.5 and less than 0.8 for medium risk, and less than 0.5 for low risk.

[0205] Assume that the model predicts the weather data for May 1 as follows:

[0206] Air temperature (air_temp) = 16°C

[0207] Rainfall = 1 mm

[0208] Air humidity (air_humidity) = 85%

[0209] Light = 1500 Lux

[0210] risk assessment

[0211] Condition check:

[0212] Air temperature 16>15→Satisfied

[0213] Rainfall 1>0→Satisfied

[0214] Humidity 85>80→Satisfied

[0215] Light 1500<2000→Satisfied

[0216] →Marked as risk (1)

[0217] Random Forest Classification:

[0218] Input feature vector: [16, 1, 85, 1500]

[0219] Each tree determines whether the condition is met, and the majority vote result is 1.

[0220] This solution: 1. Dual-model collaborative architecture

[0221] LSTM model: used to predict weather data for the next 7 days, and the output is used as one of the inputs of the random forest model.

[0222] Random Forest Model: Assess risk based on historical data and LSTM prediction data.

[0223] Through the feature sharing layer: realize joint training of two models and improve prediction accuracy.

[0224] 2. Dynamic data completion mechanism

[0225] Linear interpolation: used to handle date breakpoints (such as when there are data breaks).

[0226] Exponential decay filling: used to fill in the missing parts of historical data to ensure that the sequence length of the input data is constant at 30 days.

[0227] 3. Precise prevention and control recommendations

[0228] Provide targeted prevention and control recommendations based on the forecast results:

[0229] The weather is sunny and the air temperature is high during the heading period: the medicine can be used when the heading is complete.

[0230] The air temperature is low and the sunshine is little during the heading period: it is advisable to use the medicine during the initial flowering period.

[0231] If there is continuous rainy weather during the heading period: it is better to spray the pesticide early rather than late, and spray the pesticide multiple times during the rain breaks for prevention and control.

[0232] The various modules of this solution work together to ensure that the system can efficiently process and analyze data related to wheat scab and provide accurate prediction results. This process has broad application value in agricultural production, disease control, and other fields.

Claims

1. A wheat scab risk prediction system based on deep learning and random forest, which is characterized by: include: User end; Forecast API: used to receive historical meteorological data submitted by users and pass this data to the data preprocessing module; Data preprocessing: Clean and standardize the received data and convert the processed data into a 30-day feature sequence suitable for model input; LSTM model: Use the long short-term memory network to model the 30-day feature sequence and output the forecast data for the next 7 days; Random Forest Model: The random forest model is an ensemble learning method that improves prediction accuracy by building multiple decision trees and combining their prediction results, performing risk assessment, and outputting risk probability distribution; Model deployment module: Integrate the results from different models, merge the 7-day forecast data output by the LSTM model with historical data, and then input the data into the random forest model for risk assessment to form the final forecast results. Convert the model output into a readable form for users to understand and apply; Return result: The prediction results of the LSTM model are combined with the original historical data, and the future risk probability is calculated through the random forest model. The system returns the structured prediction results to the user end. Users can take appropriate measures based on these predictions, such as adjusting agricultural management strategies to reduce the risk of wheat scab; Feedback loop: After receiving the prediction results, the user can submit new data again through the prediction API to form a closed-loop system.

2. The wheat scab risk prediction system based on deep learning and random forest according to claim 1, characterized in that: The state update equation of LSTM is: i t =σ(W xi x t +W hi h t-1 +b i ) f t =σ(W xf x t +W hf h t-1 +b f ) o t =σ(W xo x t +W ho h t-1 +b o ) h t =o t ·tanh(C t ) in: x t : Input feature vector; h t-1 : The hidden state of the previous time step; C t-1 : cell state at the previous time step; W xi ,W xf ,W xo ,W xc : The weight matrix input to the gate unit; W hi ,W hf ,W ho ,W hc : The weight matrix from hidden state to gate unit; b i ,b f ,b o ,b c : bias term; σ: Sigmoid activation function, used to calculate the gate value; tanh: Hyperbolic tangent activation function, used to generate candidate cell states.

3. The wheat scab risk prediction system based on deep learning and random forest according to claim 2, characterized in that: Two-layer LSTM structure: The output dimension of the first layer is D, and the output dimension of the second layer is D / 2.

4. The wheat scab risk prediction system based on deep learning and random forest according to claim 3, characterized in that: The formula of the attention mechanism is as follows: Q=AW Q ,K=AW K ,V=AW V in: A: hidden state matrix output by LSTM; W Q ,W K ,W V : Learnable parameter matrix, used to generate query, key, and value respectively; Q: Query Matrix K: bond matrix; V: value matrix; d k : The dimension of the key vector, used to scale the dot product result.

5. The wheat scab risk prediction system based on deep learning and random forest according to claim 4, characterized in that: The formula for calculating risk probability is as follows: Where: Φ(x): Feature representation after random forest transforms the input features; w: weight vector for logistic regression; This formula is a standard logistic regression probability calculation formula.

6. A method for wheat scab risk prediction system based on deep learning and random forest according to claim 5, characterized in that: The following steps are included: data collection, which involves inputting historical meteorological data from the past 30 days through an interface; Preprocessing: The system cleans and standardizes the data to ensure data quality; model prediction and result output: The prediction results of the LSTM model are combined with the original historical data, and the future risk probability is calculated through the random forest model.

7. The method of the wheat scab risk prediction system based on deep learning and random forest according to claim 6, wherein: Take the seven meteorological parameters characteristic of the last five days of historical data, standardize the historical data, merge the standardized historical data and forecast data, and then mark the merged data to see whether it meets the daily conditions of average temperature exceeding 15°C, rainy weather, air humidity greater than 80% RH, and illumination less than 2000 Lux. If these conditions are met, it is marked as risk data.

8. The method of the wheat scab risk prediction system based on deep learning and random forest according to claim 7, wherein: The number of consecutive days that meet the conditions starting from the current forecast day is counted. If the conditions are not met, the counting is stopped. The random forest model algorithm is used to predict the risk probability. The risk level is divided into three levels: high, medium, and low. If the forecast data meets all meteorological conditions, it will be judged as medium or high risk. The longer the consecutive rainy days, the greater the probability of risk. The risk prediction probability is greater than 0.5 and less than 0.8 for medium risk, and less than 0.5 for low risk.

Citation Information

Cited By

  • Intelligent monitoring and early warning warehouse for agricultural products

    CN121635595A