Agricultural real-time early warning method and system based on Flink stream computing and deep learning
By constructing a bidirectional long short-term memory neural network model using Flink streaming computing and deep learning, the problem of greenhouse monitoring systems being unable to analyze multi-point greenhouse data in real time was solved, achieving minute-level early warning capabilities and resource optimization.
Patent Information
- Application Number
- CN202210513503.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-12
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2042-05-12
AI Technical Summary
Existing greenhouse monitoring systems cannot analyze data from multiple greenhouses in real time, and environmental data collected by sensors is not processed in real time, resulting in the inability to issue timely warnings.
By employing Flink streaming computing and deep learning methods, a bidirectional long short-term memory neural network model is constructed, which, combined with time-series prediction and anomaly detection, enables minute-level early warning capabilities.
It improves the accuracy and reliability of early warning, enables real-time monitoring and minute-level early warning of environmental factors in greenhouses, and reduces the waste of computing resources.
Smart Images

Figure CN114911831B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of smart agriculture technology, and particularly relates to a smart agriculture real-time early warning method and system based on Flink stream computing and deep learning. BACKGROUND
[0002] Stream computing is a frontier technology in the field of big data, focusing on real-time data streams generated on the Internet of Things or the Internet. Through technologies such as storing local states of data streams, processing out-of-order data, automatic back pressure, and save points, stream computing has the capabilities of low latency, high throughput, and high fault tolerance. Smart agriculture is the application of emerging Internet technologies, big data, cloud computing, Internet of Things, artificial intelligence, and expert knowledge to traditional agriculture. It relies on various sensor nodes and software in agricultural production to achieve intelligent sensing, intelligent early warning, intelligent analysis, and expert guidance, providing precise planting, visual management, and intelligent decision-making for agricultural production. Due to the seasonal nature of agricultural production, using greenhouse cultivation mode can overcome seasonal difficulties. Factors such as temperature, humidity, light intensity, carbon dioxide concentration, and soil moisture in the greenhouse have a significant impact on crop growth. To improve the efficiency of agricultural production, real-time monitoring of crop growth environment and precise control are essential.
[0003] Most current greenhouse monitoring systems only display and query environmental factors, and cannot make real-time decisions and warnings. The main problems are: first, only the environmental factors in a single greenhouse are analyzed, and the data from multiple geographically dispersed greenhouses are not aggregated, and no prediction model is constructed using relevant algorithms and expert decisions. Second, the environmental data collected by sensors cannot be processed and analyzed in real time, and can only be adjusted after the environment deteriorates, which cannot provide early warning. SUMMARY
[0004] The present application provides a smart agriculture real-time early warning method and system based on Flink stream computing and deep learning, which has good prediction effect and can achieve minute-level early warning capability.
[0005] To achieve the above-mentioned purpose, the present application provides a smart agriculture real-time early warning method based on Flink stream computing and deep learning, comprising:
[0006] Obtaining environmental factor data from a time series database and preprocessing the environmental factor data;
[0007] Determining the hyperparameters of a bidirectional long short-term memory neural network (Bi-LSTM) using an improved sparrow search algorithm (SSA) to construct a time series prediction model;
[0008] Deploying the time series prediction model to a Flink runtime environment for anomaly detection.
[0009] Further, the environmental factor data is obtained from the time series database, and the environmental factor data is preprocessed, specifically:
[0010] The historical environmental factor data in a period of time is extracted from the time series database, and the historical environmental factor data includes temperature, humidity, illumination, carbon dioxide concentration, etc.
[0011] The noise of the historical environmental factor data is filtered using a noise filter;
[0012] The filtered historical environmental factor data is normalized, and the normalization formula is:
[0013]
[0014] In the formula, x t n is the normalized environmental factor data, x t is the historical environmental factor data, x avg , x sd are the average value and standard deviation of the historical environmental factor data, respectively.
[0015] Further, the hyperparameters of the bidirectional long short-term memory neural network Bi-LSTM are determined by the improved sparrow search algorithm SSA, and a time series prediction model is constructed, specifically:
[0016] The environmental factor data is initialized by selecting the Sin chaotic method, and the one-dimensional mapping formula is:
[0017]
[0018] In the formula, x n is the environmental factor data;
[0019] The sparrow search algorithm abstracts the process of sparrow foraging into a finder model and a joiner model, and the position update formula of the finder model is:
[0020]
[0021] In the formula, t is the current iteration number, x i,j represents the position information of the i-th sparrow in the j-th dimension; V2 is the warning value, ST is the safety threshold; Q is a random number obeying normal distribution, L is a row of all 1 matrix obtained by multi-dimensional, and randn(0, 1) represents Gaussian distribution with mean 0 and standard deviation 1;
[0022] In the joiner model, a certain probability of finder is introduced, and the position update formula is:
[0023]
[0024] where k is the discoverer, FL∈[0,1] represents the probability of the joiner following the discoverer;
[0025] The input layer of the bidirectional long short-term memory neural network is a greenhouse environmental factor data sequence:
[0026] T n =(a1,a2,…,a t ,…,a n )
[0027] a t =(temp t ,Hum t ,Lux t ,Cdc t )
[0028] where temp t , Hum t , Lux t , Cdc t represent the temperature, humidity, illumination, and carbon dioxide concentration at time t, respectively;
[0029] The hidden layer of the bidirectional long short-term memory neural network is H=(h1,h2,…,h n ), the updated state of each memory cell in the hidden layer is C(t), and the output layer is the predicted value y^ at time t+1. The weights between the input layer and the hidden layer are represented as W ih , the weights of the hidden layer itself are represented as W hh , and the weights between the hidden layer and the output layer are represented as W ho . The Sigmoid function is selected as the activation function σ, b h is used to represent the hidden layer bias, and b y is used to represent the output layer bias. The output of the bidirectional long short-term memory neural network is:
[0030] H(t) = σ(W ih T i (t) + W hh T i (t-1) + b h ) ⊙ tanh C(t)
[0031] y^ = σ(W ho H(t) + b y ).
[0032] Further, the hidden layer is divided into a forward hidden layer and a reverse hidden layer, which are:
[0033]
[0034]
[0035] n is the amount of environmental factor data, and w is the length of the time window.
[0036] Furthermore, the attention mechanism is used to assign attention weights to the temporal information carried by historical moments to distinguish the impact of different moments on the prediction of the current moment t: the input of the temporal attention mechanism is the output state H of the hidden layer at moment t t =(h 1,t ,j 2,t ,…,h l,t ), where L is the time window of the input sequence; set W d is the training weight matrix, b d is the bias vector, and the ReLU activation function is used to increase the difference. Then the attention weight vector E corresponding to each historical moment at time t is t =(e 1,t ,e 2,t ,…,e l,t )for:
[0037] E t =ReLU(W d H t +b d )
[0038] The temporal attention weights are normalized by the Softmax function to obtain B t =(β 1,t ,β 2,t ,…,β l,t ), it is multiplied and accumulated with the hidden layer state of the corresponding historical moment to obtain the prediction H after temporal attention weighting. t ',Right now:
[0039]
[0040]
[0041] The bidirectional long short-term memory neural network was trained and tested using the TensorFlow deep learning framework to obtain a time series prediction model.
[0042] Furthermore, the time series prediction model is deployed to the Flink runtime environment, specifically:
[0043] The Flink runtime environment provides a window mechanism, which uses sliding processing time windows combined with a time series prediction model to obtain the predicted value at time t. The point to be predicted is defined as a t , the sliding neighbor window is L t , using the prediction point at The first L points are input to the time series prediction model as a sliding time window size; wherein the sliding time window is defined as follows:
[0044] L xt = (a t-L ,a t-L+1 ,…,a t-1 ).
[0045] Further, the anomaly detection is specifically implemented as follows:
[0046] The prediction data is smoothed by a moving average method using a Flink window mechanism: the prediction data is a1', a2',..., a n ′, m groups of data (m is less than n) are taken, and k times of bidirectional moving average are performed;
[0047] The forward moving average is as follows:
[0048]
[0049] The reverse moving average is as follows:
[0050]
[0051] The forward and reverse moving average is cycled k times;
[0052] If the prediction value x i is not within the warning threshold ε, and exceeds the threshold frequency f within a certain time period, an alarm is started.
[0053] The application also provides an agricultural real-time warning system based on Flink streaming calculation and deep learning, comprising a data access layer, a message intermediate layer, a real-time processing layer, a data storage layer and an application management layer:
[0054] The data access layer adopts an MQTT protocol, classifies and summarizes Internet of Things sensors in a gateway, converts into a lightweight Json format, sets a field to represent environmental factor data, and sends to a local EMQ proxy server;
[0055] The message intermediate layer adopts a Kafka message queue to form a cluster, and Kafka subscribes to multiple environmental factor data sources of EMQ in each place, sets a topic Topic according to the place;
[0056] The real-time processing layer is a distributed stream processing system Flink, Flink pulls data in Kafka through the addSource() method, uses its water line and window mechanism to ensure the consistency and order of the data when consuming the Kafka data stream; there are two kinds of nodes in the Flink cluster, the JobManager node is responsible for the central scheduling management of the task, and the TaskManager node is responsible for the execution of the specific Flink task; the data stream flowing into the Flink cluster contains the environmental factor data of multiple greenhouse greenhouses, the keyBy operator is used to partition the environmental factor data, and a sliding window is selected for data windowing operation; when the time series prediction model is combined with Flink, FAE (Flink-Ai-Extend) is selected as the intermediate layer to connect the Flink operator of the Java process and the deep learning model of the Python process, and data exchange is completed through shared memory; the Flink cluster undertakes the tasks of offline training and real-time prediction of the time series prediction model in the form of stream computing, and simultaneously utilizes the aggregation window function to output the maximum value, the minimum value and the average value of the aggregation time window to the distributed database InfluxDB; when the real-time processing layer detects an anomaly, the application management layer is sent an early warning information.
[0057] The data storage layer adopts a distributed database InfluxDB and a relational database MySQL, the distributed database InfluxDB focuses on storing the environmental factor time series data processed by Flink with a timestamp as the primary key; InfluxDB has multiple built-in data analysis functions, which provide data statistics and real-time analysis as a data source to the application layer; the relational database MySQL stores warning information and configuration information.
[0058] The application management layer includes a management console, a visual large screen and a warning push module; the management console is used by system operation and maintenance personnel, and is responsible for setting the warning threshold, the time series prediction model training period and the warning push parameters; the management console uses the Prometheus monitoring system in combination with the Grafana instrument panel to display the machine CPU, memory, network and thread usage of each node in the cluster in real time; the visual large screen is responsible for displaying the environmental factor data of the greenhouse greenhouses in various places; the warning push module converts the warning data into a textual description and notifies the user through an API such as a short message, a Dingding group and a WeChat group.
[0059] Compared with the prior art, the above technical scheme adopted by the present application has the advantages that: (1) the present application uses the time series data of the environmental factors in the greenhouse, and learns by using a bidirectional long short-term memory neural network model, which greatly improves the accuracy and has good prediction effect compared with the ARIMA model of the traditional machine learning.
[0060] (2) The time series prediction model of the present invention is trained periodically and has good self-learning ability. It also supports the adjustment of the warning threshold, thereby achieving the accuracy and reliability of the warning judgment.
[0061] (3) This paper uses the Flink streaming system to process data streams and leverages the DataStream API to simultaneously perform offline model training and real-time prediction. With only one processing engine, it can uniformly process both streaming data and batch historical data, reducing significant logical redundancy and wasted computing resources compared to traditional Lambda architectures.
[0062] (4) This invention is guided by real-time data streams, taking into account their real-time, random, disordered, and infinite characteristics. The Kafka message queue ensures high throughput and recoverability of the data stream, the Flink stream processing system ensures the order and consistency of the data stream, and the InfluxDB database provides efficient real-time collection and storage. By combining these technologies, the system achieves minute-level early warning capabilities. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 This is a flowchart of a real-time agricultural early warning method based on Flink streaming computing and deep learning;
[0064] Figure 2 To improve the flow chart of the sparrow search algorithm;
[0065] Figure 3 It is a warning chart of the curve data of the time series prediction model;
[0066] Figure 4 This is a diagram showing the relationship between the time series prediction model and the Flink cluster system.
[0067] Figure 5 This diagram shows the structure of the agricultural real-time early warning system based on Flink streaming computing and deep learning. DETAILED DESCRIPTION
[0068] In order to make the purpose, technical solutions and advantages of this application more clearly understood, this application is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application. That is, the embodiments described are only part of the embodiments of this application, not all of them.
[0069] Example 1
[0070] like Figure 1 As shown, this embodiment proposes a real-time agricultural early warning method based on Flink streaming computing and deep learning, which specifically includes the following steps:
[0071] S1: Obtain environmental factor data from a time series database and preprocess the environmental factor data, specifically:
[0072] S11: Extract historical environmental factor data within a period of time (such as 1 week) from the time series database, which includes temperature, humidity, illumination, carbon dioxide concentration, etc.
[0073] S12: Since environmental factor data has periodicity and seasonality, noise has isolation and transience, so noise filter can be used to filter noise.
[0074] S13: To eliminate the influence of large feature value range on model gradient update, statistical method is used to normalize the filtered historical environmental factor data, and the normalization formula is:
[0075]
[0076] where x t n is the normalized environmental factor data, x t is the historical environmental factor data, x avg , x sd are the mean and standard deviation of the historical environmental factor data, respectively.
[0077] S2: Determine the hyperparameters of bidirectional long short-term memory neural network Bi-LSTM through improved sparrow search algorithm SSA to build time series prediction model, specifically:
[0078] As shown in Figure 2 , the hyperparameters of Bi-LSTM neural network are determined using improved sparrow search algorithm. The sparrow search algorithm abstracts the process of sparrow foraging into finder model and joiner model, updates the positions of finder and joiner during simulation, and obtains the optimal solution.
[0079] To obtain better chaos, Sin chaos is selected to initialize environmental factor data, and the one-dimensional mapping formula is:
[0080]
[0081] The position update formula in the finder model is:
[0082]
[0083] where t is the current iteration number, x i,j represents the position information of the i-th sparrow in the j-th dimension; V2 is the warning value, ST is the safety threshold; Q is a random number following normal distribution, L is a row of all 1 matrix obtained by multi-dimensional, and randn(0,1) represents Gaussian distribution with mean 0 and standard deviation 1.
[0084] In the joiner model, a certain probability of discoverer is introduced to improve the problem of easy falling into local optimum, and the position update formula is:
[0085]
[0086] where k is the discoverer, and FL∈[0,1] represents the probability of the joiner following the discoverer;
[0087] The input layer of the bidirectional long short-term memory neural network is the greenhouse environmental factor data sequence:
[0088] T n =(a1,a2,…,a t ,…,a n )
[0089] a t =(temp t ,Hum t ,Lux t ,Cdc t )
[0090] Where temp t , Hum t , Lux t , Cdc t represent temperature, humidity, light intensity, and carbon dioxide concentration at time t, respectively.
[0091] The hidden layer of the bidirectional long short-term memory neural network is H=(h1,h2,…,h n ), and the updated state of each memory unit in the hidden layer is C(t). The output layer is the predicted value y^ at time t+1. The weights between the input layer and the hidden layer are represented as W ih , the weights of the hidden layer itself are represented as W hh , and the weights between the hidden layer and the output layer are represented as W ho . The Sigmoid function is selected as the activation function σ, and b h is used to represent the hidden layer bias, and b y is used to represent the output layer bias. The output of the bidirectional long short-term memory neural network is:
[0092] H(t)=σ(W ih T i (t)+W hh T i (t-1)+b h )⊙tanhC(t)
[0093] y^=σ(W ho H(t)+b y )
[0094] The hidden layer is divided into a forward hidden layer and a backward hidden layer, which are based on past time data and consider future time data. They are:
[0095]
[0096]
[0097] n is the amount of environmental factor data, and w is the length of the time window.
[0098] The attention mechanism is used to assign attention weights to the temporal information carried by historical moments to distinguish the impact of different moments on the prediction of the current moment t: the input of the temporal attention mechanism is the output state H of the hidden layer at moment t t =(h 1,t ,h 2,t ,…,h l,t ), where L is the time window of the input sequence; set W d is the training weight matrix, b d is the bias vector, and the ReLU activation function is used to increase the difference. Then the attention weight vector E corresponding to each historical moment at time t is t =(e 1,t ,e 2,t ,…,e l,t )for:
[0099] E t =ReLU(W d H t +b d )
[0100] The temporal attention weights are normalized by the Softmax function to obtain B t =(β 1,t ,β 2,t ,…,β l,t ), it is multiplied and accumulated with the hidden layer state of the corresponding historical moment to obtain the prediction H after temporal attention weighting. t ',Right now:
[0101]
[0102]
[0103] The bidirectional long short-term memory neural network was trained and tested using the TensorFlow deep learning framework to obtain a time series prediction model.
[0104] The square loss function MSE is selected as the loss function, and the L2 regularization term is introduced to prevent overfitting of the model. The Adam optimizer is selected for gradient descent optimization. During the training process, the two most important hyperparameters are the lag feature Lag and the number of hidden layer nodes N h . The Lag value and N h value that minimizes the error are selected.
[0105] S3: Deploy the time series prediction model in the Flink running environment for anomaly detection, specifically:
[0106] The Flink running environment provides a window mechanism, which combines a sliding time window with a time series prediction model to make predictions and obtain predicted values at time t. The predicted point is defined as a t , the sliding neighbor window is L t , and the L previous points of the predicted point a t are input into the time series prediction model as the sliding time window size. The sliding time window is defined as follows:
[0107] L xt = (a t-L , a t-L+1 ,..., a t-1 )
[0108] As shown in Figure 3 , the data obtained through real-time prediction will have up and down fluctuations. The Flink window mechanism is used to smooth the predicted data using a moving average method: the predicted data is a1', a2',..., a n ', m groups of data are taken, and k times of bidirectional moving average are performed.
[0109] The forward moving average is:
[0110]
[0111] The reverse moving average is:
[0112]
[0113] The forward and reverse moving averages are cycled k times.
[0114] If the predicted value x i is not within the warning threshold ε and exceeds the threshold frequency f within a certain time period, an alarm is triggered.
[0115] Both offline training and real-time prediction use the DataStream streaming API of Flink to form a Kappa architecture. The offline training frequency of the time series prediction model can be set to once a week, and the offline training is triggered once a week using the historical data excluding the abnormal data. The new model trained in this week is used as the latest prediction model, so that the model can be self-optimized. The early warning threshold can also be modified later, and the balance between accuracy and reliability is maintained.
[0116] As shown in Figure 5 The embodiment also provides an agricultural real-time early warning system based on Flink streaming calculation and deep learning, which comprises a data access layer, a message intermediate layer, a real-time processing layer, a data storage layer and an application management layer.
[0117] The data access layer is the part of the system communicating with the sensor. The sensor needs to collect four environmental factors, namely temperature, humidity, illumination and carbon dioxide concentration. The data access layer uses the MQTT protocol to collect the Internet of Things sensors in the gateway, classifies them into Json format, sets the field to represent the environmental factor data, and sends them to the local EMQ proxy server. At the same time, the data access layer is also responsible for obtaining the mapping parameters from the message intermediate layer as instructions.
[0118] The message intermediate layer uses a Kafka message queue to form a cluster. Kafka subscribes to multiple data sources of EMQ in various places and sets topics according to the places. Kafka, as a message middleware, has the functions of decoupling, caching and peak shaving. At the same time, because Kafka writes data to the disk and saves multiple partitions, it has recoverability and high fault tolerance.
[0119] The real-time processing layer is based on the distributed stream processing system Flink. Flink pulls data from Kafka through the addSource() method, and uses its water line and window mechanism to ensure the consistency and order of the data when consuming the Kafka data stream. There are two kinds of nodes in the Flink cluster, the JobManager node is responsible for the central scheduling management of the task, and the TaskManager node is responsible for the execution of the specific Flink task. The data stream flowing into the Flink cluster contains environmental factor data of multiple greenhouse greenhouses. In order to ensure that the data does not interfere with each other, the keyBy operator is used to partition the data, and the sliding window is used for data windowing operation. When the time series prediction model is combined with Flink, FAE (Flink-Ai-Extend) is selected as the intermediate layer to connect the Flink operator of the Java process with the deep learning model of the Python process, and the data exchange is completed through shared memory. The combination relationship of the time series prediction model and the Flink cluster system is as shown in Figure 4The Flink cluster undertakes the task of offline training and real-time prediction of the time series prediction model in the form of stream computing. At the same time, Flink reuses the aggregation window function to output the maximum, minimum and average values of the aggregated time window to the distributed database InfluxDB. When the real-time processing layer detects an anomaly, it sends an early warning message to the application management layer.
[0120] The data storage layer adopts the distributed database InfluxDB and the relational database MySQL. InfluxDB focuses on time series data scenarios and stores the environment factor time series data processed by Flink with time as the primary key. InfluxDB has various built-in data analysis functions, such as standard deviation, which can be used as a data source to provide data statistics and real-time analysis to the application layer. MySQL is responsible for storing early warning information and configuration information.
[0121] The application management layer includes a management console, a visual large screen and an early warning push module; the management console is used by system operation and maintenance personnel and is responsible for setting the real-time processing layer parameters such as early warning threshold and time series prediction model training period, and can also set early warning push parameters. The management console uses the Prometheus monitoring system combined with the Grafana dashboard to display the CPU, memory, disk, network and thread usage of each node in the cluster in real time. The visual large screen is responsible for displaying the change curves of the environmental factors of the greenhouse in different places, such as temperature, humidity, illumination and carbon dioxide concentration. The early warning push module converts the early warning data into a textual description and notifies the user through SMS, DingTalk group, WeChat group and other API.
[0122] The foregoing description of specific exemplary embodiments of the application is intended to be illustrative only and is not intended to limit the application to the precise forms described. Many modifications and variations are possible in light of the above teachings without departing from the spirit or essential characteristics of the application. The exemplary embodiments were chosen and described in order to explain the principles of the application and its practical application and to allow others skilled in the art to understand the application for various exemplary embodiments with various modifications being suited to the particular use contemplated. The scope of the application is intended to be defined by the claims and their equivalents.
Claims
1. An agricultural real-time early warning method based on Flink stream computing and deep learning, characterized in that, The method comprises the following steps: obtaining environment factor data from a time series database and preprocessing the environment factor data; determining hyperparameters of a bidirectional long short-term memory neural network Bi-LSTM through an improved sparrow search algorithm SSA to construct a time series prediction model; deploying the time series prediction model to a Flink running environment for anomaly detection; determining hyperparameters of a bidirectional long short-term memory neural network Bi-LSTM through an improved sparrow search algorithm SSA to construct a time series prediction model, specifically: initializing the environment factor data in a Sin chaotic mode, and a one-dimensional mapping formula thereof is: wherein is environmental factor data; The sparrow search algorithm abstracts the process of sparrow foraging into a finder model and a joiner model, and the position update formula of the finder model is: where t is the current iteration number, represents the position information of the i-th sparrow in the j-th dimension; is the early warning value, ST is the safety threshold; Q is a random number subject to normal distribution, and L is a row of all-1 matrix obtained in multi-dimensions, represents a Gaussian distribution with a mean of 0 and a standard deviation of 1; In the joiner model, a certain probability of finder is introduced, and the position update formula is: where k is the discoverer, represents the probability of the joiner following the discoverer; The sequence of greenhouse environmental factor data in the input layer of the bidirectional long short-term memory neural network is: wherein respectively represent temperature, humidity, light intensity, carbon dioxide concentration at time t. The hidden layer of the bidirectional long short-term memory neural network is , the updated state of each memory unit in the hidden layer is , and the output layer is the prediction value at time t+1 . The weights of the input layer and the hidden layer are represented as The weights of the hidden layer itself are represented as The weights of the hidden layer and the output layer are represented as The Sigmoid function is selected as the activation function The hidden layer bias is represented as The output layer bias is represented as The output of the bidirectional long short-term memory neural network is: The time sequence information carried by the historical time is allocated with attention weight by the attention mechanism, so as to distinguish the influence of different times on the prediction of the current t time: the input of the time sequence attention mechanism is the output state of the hidden layer of the t time , wherein L is the time window of the input sequence; set is the training weight matrix, is the bias vector, and an activation function is used to increase the difference, and the attention weight vector of each historical time corresponding to the t time is : The time sequence attention weight is normalized by a Softmax function Then, the hidden layer state corresponding to the historical moment is multiplied and accumulated to obtain the prediction weighted by the time sequence attention That is, training and testing the bidirectional long short-term memory neural network using a TensorFlow deep learning framework to obtain a time series prediction model; deploying the time series prediction model to a Flink running environment, specifically: A window mechanism is provided in the Flink running environment, a prediction value at time t is obtained by combining a sliding time window with a time series prediction model, and a to-be-predicted point , a sliding neighbor window , adopts L previous points as a sliding time window size input to a time series prediction model; wherein the sliding time window is defined as follows: The above method is implemented by the following system, including a data access layer, a message intermediate layer, a real-time processing layer, a data storage layer and an application management layer: The data access layer uses the MQTT protocol to classify and summarize the Internet of Things sensors in the gateway into a lightweight Json format, sets a field to represent the environment factor data, and sends it to the local EMQ proxy server; The message intermediate layer uses a Kafka message queue to form a cluster, and Kafka subscribes to multiple environment factor data sources of EMQ in each place and sets a topic Topic according to the place; The real-time processing layer is a distributed stream processing system Flink, which pulls data from Kafka through the addSource() method, uses its water line and window mechanism to ensure the consistency and order of the data when consuming the Kafka data stream; The Flink cluster has two types of nodes, the JobManager node is responsible for the central scheduling management of the task, and the TaskManager node is responsible for the execution of the specific Flink task; The data stream flowing into the Flink cluster contains environmental factor data of multiple greenhouse greenhouses, which is partitioned using the keyBy operator, and a sliding window is selected for data windowing operation; When the time series prediction model is combined with Flink, FAE is selected as the intermediate layer to connect the Flink operator of the Java process and the deep learning model of the Python process, and data exchange is completed through shared memory; The Flink cluster undertakes the tasks of offline training and real-time prediction of the time series prediction model in the form of stream computing, and simultaneously outputs the maximum, minimum and average values of the aggregated time window to the distributed database InfluxDB using the aggregation window function; When the real-time processing layer detects an anomaly, it sends an early warning message to the application management layer. The data storage layer adopts a distributed database InfluxDB and a relational database MySQL, the distributed database InfluxDB focuses on storing environment factor time series data processed by Flink with a timestamp as a primary key; InfluxDB has multiple built-in data analysis functions, providing data statistics and real-time analysis as a data source to the application layer; the relational database MySQL stores early warning information and configuration information; The application management layer includes a management console, a visual large screen and an early warning push module; the management console is used by system operation and maintenance personnel, responsible for setting early warning thresholds, time series prediction model training periods and early warning push parameters; the management console uses a Prometheus monitoring system combined with a Grafana dashboard to display the CPU, memory, network and thread usage of each node in the cluster in real time; the visual large screen is responsible for displaying greenhouse environmental factor data in different places; The early warning push module converts early warning data into a text description and notifies users through SMS, DingTalk groups and WeChat groups. 2.The real-time warning method based on Flink stream computing and deep learning for agriculture according to claim 1, wherein, The environment factor data is obtained from the time series database and preprocessed, specifically: Extract historical environment factor data within a period of time from the time series database, including temperature, humidity, illumination and carbon dioxide concentration; Use a noise filter to filter the noise of the historical environment factor data; The filtered historical environment factor data is normalized, and the normalization formula is: In the formula is the normalized environmental factor data, is the historical environmental factor data, are the average value and the standard deviation of the historical environmental factor data, respectively. 3.The real-time warning method based on Flink stream computing and deep learning for agriculture according to claim 1, wherein, The hidden layer is divided into a forward hidden layer and a reverse hidden layer, respectively: n is the amount of environment factor data, and w is the length of the time window. 4.The real-time warning method based on Flink stream computing and deep learning for agriculture according to claim 1, wherein, The specific implementation of anomaly detection is: The prediction data is smoothed by a moving average method using a Flink window mechanism: , m groups of data are taken, and k times of bidirectional moving average is performed; The forward moving average is: The reverse moving average is: The forward and reverse moving averages are cycled k times. If the predicted value If the predicted value is not within the warning threshold ε and exceeds the threshold frequency f within a certain time period, an alert is initiated.
Citation Information
Patent Citations
Road motor vehicle exhaust emission prediction method based on improved attention bidirectional long-short-term memory network
CN111612254A
Mechanical equipment residual life prediction method for optimizing BiLSTM based on DRSN and sparrow search
CN113723007A