Lake and reservoir water bloom prediction intelligent system and method based on DTW algorithm
By using the DTW algorithm to calculate the similarity of the data sequence in the Huku Shuhua prediction intelligent system and establish a prediction matrix, the problem of large data demand and calculation of the existing Shuhua prediction model is solved, and effective prediction of the growth status of the waterhua and early warning of the waterhua growth are achieved.
Patent Information
- Application Number
- CN202510005640.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-02
AI Technical Summary
The existing water flower prediction model requires the collection of large amounts of data, which requires high heterogeneity of data, has a large amount of calculation, and cannot effectively and flexibly predict the growth status of water flower.
The Huku Shuhua prediction intelligent system based on the DTW algorithm is adopted. This system collects the concentration data of chlorophyll a by setting sensor monitoring points in the lake, uses the DTW algorithm to calculate the similarity of the data sequence, establishes a prediction matrix, and dynamically adjusts the weights through the weight allocation module to obtain the prediction results.
It reduces data demand and calculation volume, improves the accuracy and robustness of the prediction results, and can effectively predict the growth status of the water bath and trigger the water bath warning.
Smart Images

Figure CN119917872A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of lake and reservoir algae bloom prediction, and in particular to an intelligent system and method for lake and reservoir algae bloom prediction based on a DTW algorithm. Background Art
[0002] At present, under the background of increasing eutrophication of global water bodies and global warming, the frequency of algal bloom outbreaks in surface water bodies has increased year by year, seriously threatening the production and life of human society and bringing severe challenges to the management of surface water resources. As a technical means to support the management of surface water resources, algal bloom prediction technology has made great progress in recent years. This technology often establishes mathematical models to simulate the response relationship between the growth state of phytoplankton and other main organisms constituting algal blooms and environmental variables (meteorology, hydrology, water quality), and predicts the corresponding phytoplankton growth state, so as to realize the judgment of algal bloom outbreak risk. Typical algal bloom prediction statistical models include: algal bloom prediction model based on linear statistics, algal bloom prediction model based on generalized linear regression, algal bloom prediction model based on machine learning algorithm, etc.
[0003] The above algal bloom prediction models are implemented based on different statistical algorithms, avoiding the risk of insufficient understanding of natural laws. When there is sufficient data, they have high prediction accuracy and robustness. However, the above models need to collect a large amount of data, have high requirements on data heterogeneity, and have a large amount of calculation. Since the water ecological conditions of natural water bodies such as lakes and reservoirs are affected by multiple factors such as hydrology, meteorology, exogenous material input, and human activity interference, and the mapping relationship between variables has strong spatiotemporal variability, the models established using large amounts of data and complex mechanisms cannot effectively and flexibly predict the growth status of algal blooms. Therefore, in order to be able to use a relatively simple model framework and less data to predict the growth status of algal blooms, we propose an intelligent system and method for lake and reservoir algal bloom prediction based on the DTW algorithm. Summary of the invention
[0004] The purpose of the present invention is to solve the problem that the reservoir algae bloom prediction model needs to collect a large amount of data, has high requirements on data heterogeneity, and has a large amount of calculation. At the same time, the model established using a large amount of data and complex mechanisms cannot effectively and flexibly predict the growth status of the algae bloom. In order to be able to use a relatively simple model framework and less data to predict the growth status of the algae bloom, at the same time, ensure the accuracy and robustness of the prediction results.
[0005] To achieve the above-mentioned purpose, the present invention provides an intelligent system for predicting lake and reservoir algae bloom based on the DTW algorithm, comprising a data acquisition module, a matrix construction module, a weight distribution module and a display alarm module;
[0006] The data acquisition module sets a plurality of continuously arranged sensor monitoring points in the lake, collects the concentration value of chlorophyll a continuously at fixed time intervals for a long time, arranges and integrates the collected data in chronological order to form a time series, and divides the monitoring data of chlorophyll a into two parts, the first part is the training data set Xt, which records the continuous detection values of chlorophyll a concentration from the start time to the nth time; the second part is the comparison sequence Xc, which records the chlorophyll a concentration detection values from the nth time to the last time;
[0007] The matrix construction module calculates the similarity between the comparison sequence Xc divided by the data acquisition module and the training data set Xt through the DTW algorithm, selects the m subsequences with the largest similarity, and sorts them to obtain a similarity vector Sm containing m similarities and arranged from large to small, and then selects the m subsequences with the largest similarity to the trajectory of the sequence Xc from each subsequence of Xt, and composes a prediction matrix Xnm in descending order of similarity;
[0008] The weight distribution module calculates the average relative error between the m subsequences determined by the matrix construction module and the comparison sequence Xc to obtain the error vector Am, obtains the corresponding distribution coefficient vector according to the distribution coefficient formula, normalizes the elements thereof, determines the weight vector of each subsequence, and uses the weight vector to perform weighted summation on the adjusted prediction matrix, integrates the information of each subsequence according to its importance weight, and determines the concentration prediction value of chlorophyll a;
[0009] The display alarm module compares the predicted chlorophyll a concentration value determined by the weight allocation module with the chlorophyll a concentration threshold of the reservoir. If the chlorophyll a concentration exceeds the threshold range, an algal bloom warning is triggered to remind staff to take corresponding preventive measures.
[0010] Preferably, the matrix building module includes a similarity calculation unit and a matrix building unit;
[0011] The similarity calculation unit calculates the DTW-adjusted Euclidean distance between the comparison sequence Xc and the training data set Xt by using the DTW algorithm to determine the similarity between the comparison sequence Xc and the training data set Xt;
[0012] The matrix building unit builds a prediction matrix using the time step as a row vector and the similarity subsequence as a column vector.
[0013] Preferably, the formula for calculating the DTW-adjusted Euclidean distance between the comparison sequence Xc and the training data set Xt by the similarity calculation unit is:
[0014]
[0015]
[0016] Where d is the Euclidean distance, D is the Euclidean distance after DTW adjustment, Plan a path for a specific DTW, x i To compare the data values in the sequence Xc, y i is the data value in the training data set Xt.
[0017] Preferably, the matrix building unit fills the prediction matrix by filling in columns and taking similar data sequences as a guide, and for each similar data sequence after sorting, fills its related data into the corresponding column of the prediction matrix.
[0018] Preferably, the weight allocation module includes a weight determination unit and a prediction result unit;
[0019] The weight determination unit determines the distribution coefficient vector from the error vector and the similarity vector according to the distribution coefficient formula, and then normalizes the elements thereof to determine the weight vector of each subsequence;
[0020] The prediction result unit performs weighted summation on the adjusted prediction matrix by using the weight vector, and integrates the information of each subsequence according to its importance weight, so as to obtain the prediction result of the future chlorophyll a concentration.
[0021] Preferably, the distribution coefficient formula in the weight determination unit is:
[0022]
[0023] Among them, α m is the distribution coefficient vector, β is the parameter value, A m is the error vector, S m is the similarity vector.
[0024] Preferably, when calculating the prediction result of the future chlorophyll a concentration, the prediction result unit determines whether the result of the prediction model is accurate by comparing the average relative error between the predicted value and the chlorophyll a concentration value obtained by subsequent actual measurement.
[0025] Preferably, the data acquisition module includes a data acquisition unit and a data division unit;
[0026] The data collection unit sets a plurality of monitoring points at different locations of the lake, collects the concentration value of chlorophyll a continuously at fixed time intervals for a long time, and arranges and integrates the collected data in chronological order to form a time series;
[0027] The data division unit performs data cleaning on the collected data sequence, processes missing values and abnormal values in the data sequence, and then divides the data sequence.
[0028] Preferably, the display alarm module compares the predicted chlorophyll a concentration value with the chlorophyll a concentration threshold of the reservoir, determines that the current chlorophyll a concentration change trend of the lake is similar to the situation before the historical algal bloom outbreak, triggers the algal bloom warning, and displays the warning time, warning location and predicted time range of the algal bloom outbreak.
[0029] The second object of the present invention is to provide a lake and reservoir algae bloom prediction method based on the DTW algorithm, including any one of the above-mentioned lake and reservoir algae bloom prediction intelligent systems based on the DTW algorithm, comprising the following steps:
[0030] S1. Collect the chlorophyll a concentration value in the reservoir through the data acquisition module, and arrange it in time series, which is divided into two parts;
[0031] S2, using the matrix construction module to calculate the similarity between the comparison sequence Xc and the training data set Xt, and establish a prediction matrix with the time step n as the row vector and the subsequence of similarity m as the column vector;
[0032] S3, calculating the average relative error between the subsequence and the comparison sequence Xc through the weight distribution module to obtain the error vector Am, obtaining the weight vector from the error vector and the similarity vector, and performing weighted summation of the weight vector on the adjusted prediction matrix to obtain the prediction result of the future chlorophyll a concentration;
[0033] S4. The display alarm module compares the predicted chlorophyll a concentration value with the reservoir chlorophyll a concentration threshold value, determines that the current chlorophyll a concentration change trend of the lake is similar to the situation before the historical algal bloom outbreak, and triggers an algal bloom warning.
[0034] Compared with the prior art, the present invention has the following beneficial effects:
[0035] 1. The intelligent system method for predicting algae bloom in lakes and reservoirs based on the DTW algorithm sets multiple continuously arranged sensor monitoring points in the lake through the data acquisition module, collects the concentration value of chlorophyll a at fixed time intervals for a long time, and arranges them in time series, which are divided into two parts. The DTW algorithm is used to calculate the similarity between the comparison sequence Xc and the training data set Xt through the matrix construction module, and the prediction matrix is established with the time step as the row vector and the similarity subsequence as the column vector. The model only uses chlorophyll a as the monitoring data, which greatly reduces the demand for original data, and uses the DTW algorithm to reduce the calculation amount of the model;
[0036] 2. The average relative error between the subsequence and the comparison sequence Xc is calculated through the weight distribution module to obtain the error vector Am. The weight vector is obtained from the error vector and the similarity vector. The weighted sum of the weight vector to the adjusted prediction matrix is used to obtain the prediction result of the future chlorophyll a concentration. The accuracy and robustness of the prediction result are guaranteed by dynamically adjusting the weights between different similarities.
[0037] 3. Use the display alarm module to compare the predicted chlorophyll a concentration value with the reservoir chlorophyll a concentration threshold to determine that the current chlorophyll a concentration change trend of the lake is similar to the situation before the historical algal bloom outbreak, triggering an algal bloom warning and displaying the warning time, warning location and predicted time range of the algal bloom outbreak. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 It is the overall flow chart of the present invention;
[0039] Figure 2 It is the overall detailed flow chart of the present invention;
[0040] Figure 3 This is a flow chart of the weight allocation module of the present invention;
[0041] Figure 4 The figure is a flow chart of the method of the present invention.
[0042] The meaning of each number in the figure is:
[0043] 100, data acquisition module; 110, data acquisition unit; 120, data division unit; 200, matrix construction module; 210, similarity calculation unit; 220, matrix establishment unit; 300, weight allocation module; 310, weight determination unit; 320, prediction result unit; 400, alarm display module. DETAILED DESCRIPTION
[0044] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0045] At present, reservoir algal bloom prediction models need to collect a large amount of data, have high requirements on data heterogeneity, and require a large amount of calculation. At the same time, models established using a large amount of data and complex mechanisms cannot effectively and flexibly predict the growth status of algal blooms. In order to use a relatively simple model framework and less data to predict the growth status of algal blooms, while ensuring the accuracy and robustness of the prediction results.
[0046] Therefore, the present invention proposes that the concentration values of chlorophyll a at different positions of the reservoir are collected by the data collection module, and arranged in time series and divided into two parts, and the similarity between the comparison sequence Xc and the training data set Xt is calculated by the matrix construction module, and the prediction matrix is established with the time step n as the row vector and the subsequence of similarity m as the column vector, and the average relative error between the subsequence and the comparison sequence Xc is calculated by the weight distribution module to obtain the error vector Am, and the weight vector is obtained by the error vector and the similarity vector. The prediction result of the future chlorophyll a concentration is obtained by weighted summing the weight vector to the adjusted prediction matrix, and the display alarm module compares the chlorophyll a concentration threshold of the reservoir with the chlorophyll a concentration prediction value, and determines that the chlorophyll a concentration change trend of the current lake reservoir is similar to the situation before the historical algal bloom outbreak, and triggers the algal bloom warning, as follows:
[0047] Embodiment 1, as Figure 1 As shown, one of the purposes of the present invention is to provide an intelligent system for predicting lake blooms based on the DTW algorithm, including a data acquisition module 100, a matrix construction module 200, a weight distribution module 300 and a display alarm module 400;
[0048] The data acquisition module 100 sets a plurality of continuously arranged sensor monitoring points in the lake, collects the concentration value of chlorophyll a continuously at fixed time intervals for a long time, arranges and integrates the collected data in chronological order to form a time series, and divides the monitoring data of chlorophyll a into two parts, namely, the front part and the back part, according to the time series. The first part is the training data set Xt, which records the continuous detection values of the chlorophyll a concentration from the start time to the nth time; the second part is the comparison sequence Xc, which records the detection values of the chlorophyll a concentration from the nth time to the last time.
[0049] The matrix construction module 200 calculates the similarity between the comparison sequence Xc divided by the data acquisition module 100 and the training data set Xt through the DTW algorithm, selects the m subsequences with the largest similarity, and sorts them to obtain a similarity vector Sm containing m similarities and arranged from large to small, and then selects the m subsequences with the largest similarity to the trajectory of the sequence Xc from each subsequence of Xt, and composes a prediction matrix Xnm in descending order of similarity;
[0050] The weight distribution module 300 calculates the average relative error between the m subsequences determined by the matrix construction module 200 and the comparison sequence Xc to obtain the error vector Am, obtains the corresponding distribution coefficient vector according to the distribution coefficient formula, normalizes the elements thereof, determines the weight vector of each subsequence, and uses the weight vector to perform weighted summation on the adjusted prediction matrix, integrates the information of each subsequence according to its importance weight, and determines the concentration prediction value of chlorophyll a;
[0051] The display alarm module 400 compares the predicted chlorophyll a concentration value determined by the weight allocation module 300 with the chlorophyll a concentration threshold of the reservoir. If the chlorophyll a concentration exceeds the threshold range, an algal bloom warning is triggered to remind the staff to take corresponding preventive measures.
[0052] The concentration of chlorophyll a can reflect the biomass of phytoplankton in the water body, and the time series formed by its continuous monitoring data contains the trend information of water body ecological changes. By collecting chlorophyll a concentration values at multiple monitoring points in a research area (such as lakes, reservoirs, etc.) for a long time and at fixed intervals, these chronologically arranged data are integrated to form a time series, which provides basic data support for subsequent analysis of changes in water body ecological status and prediction;
[0053] When dividing the comparison sequence Xc and the training data set Xt, the comparison sequence Xc is used as a reference standard for comparison and analysis with each subsequence in the training data set Xt to find similar patterns. The length that can cover the short-term fluctuation cycle or reflect the stage-by-stage water quality changes is used as the length of the comparison sequence Xc. Assuming that we have collected a time series consisting of 1,000 chlorophyll a concentration detection values, if n=50 is determined, the first 50 values can be used as the comparison sequence Xc, and the remaining 950 values can be divided into groups of 50. In this way, k=950÷50=19 subsequences of length 50 can be obtained, which constitute the training data set Xt.
[0054] like Figure 2 As shown, the matrix building module 200 includes a similarity calculation unit 210 and a matrix building unit 220;
[0055] The similarity calculation unit 210 calculates the DTW-adjusted Euclidean distance between the comparison sequence Xc and the training data set Xt by using the DTW algorithm to determine the similarity between the comparison sequence Xc and the training data set Xt;
[0056] A matrix building unit 220 builds a prediction matrix with the time step n as a row vector and the subsequences of similarity m as a column vector;
[0057] The DTW (Dynamic Time Warping) algorithm is used to calculate the trajectory similarity between each subsequence in the sequence Xt and the sequence Xc, and the similarity vector S = [s1, s2, s3, ..., sk] is obtained. The m subsequences with the largest similarity are selected and sorted to obtain a similarity vector Sm containing m similarities arranged from large to small. Then, the m subsequences with the largest similarity to the trajectory of the sequence Xc are selected from each subsequence of Xt, and the prediction matrix Xn×m is formed in descending order of similarity.
[0058] Among them, the DTW algorithm is the dynamic time warping algorithm. It is a method to measure the similarity of two time series. It is mainly used to deal with the expansion and distortion of time series on the time axis, so that two sequences that are not completely aligned in the time dimension can also be effectively compared for similarity.
[0059] Considering that time series may be stretched or distorted on the time axis, the DTW algorithm can better find the optimal matching path to measure similarity. It constructs a distance matrix and then uses dynamic programming ideas to gradually calculate the cumulative distance from the starting position to the end position. The final cumulative distance can reflect the degree of similarity (which can also be appropriately converted into a similarity value).
[0060] In order to more accurately calculate the similarity between the comparison sequence Xc and the training data set Xt, the formula for calculating the DTW-adjusted Euclidean distance between the comparison sequence Xc and the training data set Xt by the similarity calculation unit 210 is:
[0061]
[0062] Where d is the Euclidean distance, D is the Euclidean distance after DTW adjustment, For a specific DTW planning path, xi is the data value in the comparison sequence Xc, y i is the data value in the training data set Xt.
[0063] For each subsequence in the training data set Xt, the distance with the comparison sequence Xc is calculated in this way, and a data set containing all distance values is obtained, which can be used as a representation of the similarity data set S (the inverse of the distance and other forms can also be converted into similarity measurement);
[0064] Based on the Euclidean distance, we focus on the most similar data. These training data subsequences that are highly similar to the comparison sequence are more likely to contain valuable information for subsequent predictions and can more accurately reflect the trend and characteristics of water ecological changes related to the current comparison sequence. By sorting by similarity, we can further organize the data, making it easier to process the most similar data in order of importance, making subsequent operations such as building a prediction model more logical and reasonable.
[0065] In order to build the prediction matrix more quickly, the matrix building unit 220 fills the prediction matrix by filling in columns and taking similar data sequences as a guide. For each similar data sequence after sorting, its related data is filled into the corresponding column of the prediction matrix.
[0066] For each similar data sequence after sorting, fill its related data into the corresponding column of the prediction matrix. If it is time series data, fill it in order according to the time step. For example, suppose a prediction matrix with 3 rows and 5 columns has been determined, and there are 3 similar data sequences sorted by similarity, each sequence has 5 time steps of data, fill the 5 time steps of the first similar data sequence into the first column of the prediction matrix, fill the second sequence data into the second column, and fill the third sequence data into the third column;
[0067] The m most similar training data subsequences are integrated to form a matrix structure, which is convenient for subsequent batch data analysis and processing, such as calculating their errors with the comparison sequence, and arranging them in the order corresponding to the similarity vector Sm, ensuring the relevance of the data and the logic of the processing, so that the information of each subsequence can be integrated more orderly and efficiently when using these data for prediction in the future.
[0068] like Figure 3 As shown, the weight allocation module 300 includes a weight determination unit 310 and a prediction result unit 320;
[0069] The weight determination unit 310 determines the distribution coefficient vector from the error vector and the similarity vector according to the distribution coefficient formula, and then normalizes the elements thereof to determine the weight vector of each subsequence;
[0070] The prediction result unit 320 performs weighted summation on the adjusted prediction matrix by applying the weight vector, and integrates the information of each subsequence according to its importance weight, to obtain the prediction result of the future chlorophyll a concentration;
[0071] The mean relative error can measure the difference between each training data subsequence and the comparison sequence, reflecting the accuracy of similarity from another perspective. It can further understand the deviation between these similar subsequences and the comparison benchmark, and provide an important reference for subsequent operations such as reasonable weight allocation, so that the final prediction model can more accurately consider the reliability and accuracy of each subsequence.
[0072] The formula for the mean relative error is:
[0073]
[0074] Among them, x ij is the jth subsequence in the prediction matrix, x ic is the i-th subsequence in the comparison sequence.
[0075] In order to better calculate the allocation coefficient, the allocation coefficient formula in the weight determination unit 310 is:
[0076]
[0077] Among them, α m is the distribution coefficient vector, β is the parameter value, A m is the error vector, S m is the similarity vector.
[0078] Using the similarity vector S m With the error vector A m , calculate the distribution coefficient vector α according to the above formula m , β is a parameter with a value between (1,2) and plays a regulating role, so that the allocation coefficient can comprehensively consider the similarity and error factors, reasonably allocate weights, and more accurately use the information of different subsequences in the subsequent prediction process, avoiding the prediction bias caused by relying solely on similarity or error single factors, making the model more scientific and reasonable.
[0079] Ensure that the sum of the distribution coefficients is 1 after processing, so that the weight vector can reflect the relative importance of each subsequence in the prediction in a reasonable proportion, which conforms to the basic definition and properties of weights, and facilitates the accurate weighted summation and other operations according to the relative importance of each subsequence in subsequent prediction calculations, ensuring that the prediction result is a reasonable output after comprehensive consideration of all relevant factors;
[0080] The weight vector is used to perform weighted summation on the adjusted prediction matrix, and the information of each subsequence is integrated according to its importance weight to obtain a comprehensive prediction value. This value is an estimate of the future chlorophyll a concentration based on the previous series of data analysis and processing. It reflects the prediction judgment after comprehensive consideration of multiple similar data and their weights, and is in line with the basic logic and method of prediction based on data similarity.
[0081] In order to better judge the accuracy of the prediction model, when the prediction result unit 320 calculates the prediction result of the future chlorophyll a concentration, it determines whether the result of the prediction model is accurate by comparing the average relative error between the predicted value and the chlorophyll a concentration value obtained by subsequent actual measurement;
[0082] By comparing the predicted value with the chlorophyll a concentration value obtained by subsequent actual measurement and calculating the average relative error between them, we can intuitively understand the accuracy and reliability of the prediction method. The smaller the average relative error, the closer the predicted value is to the actual value, which means that the prediction model has a better prediction effect on the chlorophyll a concentration in the current study area and can be further applied and optimized; otherwise, it is necessary to adjust and improve the model parameters, data processing methods, etc. to improve the accuracy and practicality of the prediction.
[0083] In order to determine the growth status of water bloom in the reservoir according to the predicted chlorophyll a concentration value, the display alarm module 400 compares the predicted chlorophyll a concentration value with the reservoir chlorophyll a concentration threshold, determines that the current chlorophyll a concentration change trend of the lake is similar to the situation before the water bloom outbreak in history, triggers the water bloom warning, and displays the warning time, warning location and water bloom outbreak time range prediction;
[0084] By comparing the predicted chlorophyll a concentration value with the reservoir chlorophyll a concentration threshold, if the concentration is less than the threshold, it is determined that the current chlorophyll a concentration change trend of the lake is similar to the situation before the historical algal bloom outbreak, triggering an algal bloom warning. The warning information should include the warning level (such as mild warning, moderate warning, and severe warning, which can be divided according to the proximity of the DTW distance to the threshold), warning time, warning location (i.e., monitoring point), and possible algal bloom outbreak time range prediction. The predicted algal bloom outbreak time range can be determined based on the statistical distribution of the time interval of algal bloom outbreaks under similar DTW distances in historical data. For example, if 7 of the past 10 similar situations have algal blooms that broke out within 2-4 days, this predicted algal bloom may occur within the next 2-4 days. The warning information is promptly conveyed to relevant managers through the user interaction and visualization module so that they can take corresponding countermeasures, such as the use of algaecides, increasing water aeration, adjusting the sluice flow, etc., to reduce the risk and harm of algal bloom outbreaks.
[0085] The second object of the present invention is to provide a lake and reservoir water bloom prediction method based on the DTW algorithm, including any one of the above-mentioned lake and reservoir water bloom prediction intelligent systems based on the DTW algorithm, comprising the following steps:
[0086] S1. Collect the chlorophyll a concentration value in the reservoir through the data acquisition module 100, and arrange it in time series, dividing it into two parts;
[0087] S2, using the matrix construction module 200 to calculate the similarity between the comparison sequence Xc and the training data set Xt, and establishing a prediction matrix with the time step n as the row vector and the subsequences of similarity m as the column vector;
[0088] S3, calculating the average relative error between the subsequence and the comparison sequence Xc through the weight distribution module 300 to obtain the error vector Am, obtaining the weight vector from the error vector and the similarity vector, and performing weighted summation on the adjusted prediction matrix by the weight vector to obtain the prediction result of the future chlorophyll a concentration;
[0089] S4. The display alarm module 400 compares the predicted chlorophyll a concentration value with the reservoir chlorophyll a concentration threshold value, determines that the current chlorophyll a concentration change trend of the lake is similar to the situation before the historical algal bloom outbreak, and triggers an algal bloom warning.
[0090] In summary, the working principle of this solution is as follows:
[0091] The lake bloom prediction intelligent system method based on the DTW algorithm sets a plurality of continuously arranged sensor monitoring points in the lake through the data acquisition module 100, collects the concentration value of chlorophyll a at fixed time intervals for a long time, and arranges them in time series, which are divided into two parts, and calculates the similarity between the comparison sequence Xc and the training data set Xt by using the DTW algorithm through the matrix construction module 200, and establishes a prediction matrix with the time step n as the row vector and the subsequence of the similarity m as the column vector. The model only uses chlorophyll a as the monitoring data, which greatly reduces the demand for original data, and uses the DTW algorithm to reduce the calculation amount of the model. 00 calculates the average relative error between the subsequence and the comparison sequence Xc to obtain the error vector Am, obtains the weight vector from the error vector and the similarity vector, and obtains the prediction result of the future chlorophyll a concentration by weighted summing the weight vector to the adjusted prediction matrix. By dynamically adjusting the weights between different similarities, the accuracy and robustness of the prediction result are guaranteed. The display alarm module 400 compares the predicted chlorophyll a concentration value with the chlorophyll a concentration threshold of the reservoir to determine that the current chlorophyll a concentration change trend of the lake is similar to the situation before the historical algal bloom outbreak, triggers the algal bloom warning, and displays the warning time, warning location and algal bloom outbreak time range forecast.
[0092] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and descriptions are only preferred examples of the present invention and are not intended to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, which fall within the scope of the present invention. The scope of protection of the present invention is defined by the attached claims and their equivalents.
Claims
1. An intelligent system for predicting lake and reservoir algae blooms based on the DTW algorithm, characterized by: It comprises a data collection module (100), a matrix construction module (200), a weight distribution module (300) and a display alarm module (400); The data acquisition module (100) sets a plurality of continuously arranged sensor monitoring points in the lake to continuously collect chlorophyll a concentration values at fixed time intervals for a long time, arranges and integrates the collected data in chronological order to form a time series, and divides the chlorophyll a monitoring data into two parts, namely, a front part and a back part, according to the time series. The first part is a training data set Xt, which records the continuous detection values of the chlorophyll a concentration from the start time to the nth time; and the second part is a comparison sequence Xc, which records the chlorophyll a concentration detection values from the nth time to the last time. The matrix construction module (200) calculates the similarity between the comparison sequence Xc divided by the data acquisition module (100) and the training data set Xt by using the DTW algorithm, selects m subsequences with the greatest similarity, and sorts them to obtain a similarity vector Sm containing m similarities and arranged from large to small, then selects m subsequences with the greatest similarity to the trajectory of the sequence Xc from each subsequence of Xt, and composes a prediction matrix Xnm in descending order of similarity; The weight distribution module (300) calculates the average relative error between the m subsequences determined by the matrix construction module (200) and the comparison sequence Xc to obtain an error vector Am, obtains a corresponding distribution coefficient vector according to a distribution coefficient formula, normalizes the elements thereof, determines a weight vector for each subsequence, performs weighted summation on the adjusted prediction matrix using the weight vector, integrates the information of each subsequence according to its importance weight, and determines a predicted value of chlorophyll a concentration; The display alarm module (400) compares the predicted chlorophyll a concentration value determined by the weight allocation module (300) with the chlorophyll a concentration threshold of the reservoir, and triggers an algal bloom warning if the chlorophyll a concentration exceeds the threshold range, reminding staff to take corresponding preventive measures.
2. The intelligent system for predicting lake and reservoir algae blooms based on the DTW algorithm according to claim 1 is characterized by: The matrix building module (200) comprises a similarity calculation unit (210) and a matrix building unit (220); The similarity calculation unit (210) calculates the DTW-adjusted Euclidean distance between the comparison sequence Xc and the training data set Xt by using a DTW algorithm to determine the similarity between the comparison sequence Xc and the training data set Xt; The matrix building unit (220) builds a prediction matrix using the time step as a row vector and the similarity subsequence as a column vector.
3. The intelligent system for predicting lake and reservoir algae blooms based on the DTW algorithm according to claim 2 is characterized by: The formula for calculating the DTW-adjusted Euclidean distance between the comparison sequence Xc and the training data set Xt by the similarity calculation unit (210) is: Where d is the Euclidean distance, D is the Euclidean distance after DTW adjustment, Plan a path for a specific DTW, x i To compare the data values in the sequence Xc, y i is the data value in the training data set Xt.
4. The intelligent system for predicting lake and reservoir algae blooms based on the DTW algorithm according to claim 2 is characterized by: The matrix building unit (220) fills the prediction matrix by filling in columns and taking similar data sequences as a guide. For each similar data sequence after sorting, its related data is filled into the corresponding column of the prediction matrix.
5. The intelligent system for predicting lake and reservoir algae blooms based on the DTW algorithm according to claim 1 is characterized by: The weight allocation module (300) includes a weight determination unit (310) and a prediction result unit (320); The weight determination unit (310) determines a distribution coefficient vector from the error vector and the similarity vector according to a distribution coefficient formula, and then normalizes the elements thereof to determine a weight vector for each subsequence; The prediction result unit (320) performs weighted summation on the adjusted prediction matrix by applying the weight vector, and integrates the information of each subsequence according to its importance weight, so as to obtain a prediction result of the future chlorophyll a concentration.
6. The intelligent system for predicting lake and reservoir algae blooms based on the DTW algorithm according to claim 5 is characterized by: The distribution coefficient formula in the weight determination unit (310) is: Among them, α m is the distribution coefficient vector, β is the parameter value, A m is the error vector, S m is the similarity vector.
7. The intelligent system for predicting lake and reservoir algae blooms based on the DTW algorithm according to claim 5 is characterized by: When calculating the prediction result of the future chlorophyll a concentration, the prediction result unit (320) determines whether the result of the prediction model is accurate by comparing the average relative error between the predicted value and the chlorophyll a concentration value obtained by subsequent actual measurement.
8. The intelligent system for predicting lake and reservoir algae blooms based on the DTW algorithm according to claim 1 is characterized by: The data acquisition module (100) comprises a data acquisition unit (110) and a data division unit (120); The data collection unit (110) sets a plurality of monitoring points at different locations of the lake, collects the concentration value of chlorophyll a continuously at fixed time intervals for a long period of time, and arranges and integrates the collected data in chronological order to form a time series; The data division unit (120) performs data cleaning on the collected data sequence, processes missing values and abnormal values in the data sequence, and then divides the data sequence.
9. The intelligent system for predicting lake and reservoir algae blooms based on the DTW algorithm according to claim 1 is characterized by: The display alarm module (400) compares the predicted chlorophyll a concentration value with the reservoir chlorophyll a concentration threshold value, determines that the current chlorophyll a concentration change trend of the lake is similar to the situation before the algal bloom outbreak in history, triggers the algal bloom warning, and displays the warning time, warning location and the predicted time range of the algal bloom outbreak.
10. A method for implementing lake and reservoir algae bloom prediction based on the DTW algorithm, comprising the lake and reservoir algae bloom prediction intelligent system based on the DTW algorithm according to any one of claims 1 to 9, comprising the following steps: S1, collecting chlorophyll a concentration values in the reservoir through a data collection module (100), arranging them in time series, and dividing them into two parts; S2, using the matrix construction module (200) to calculate the similarity between the comparison sequence Xc and the training data set Xt, and to establish a prediction matrix with the time step n as the row vector and the subsequences of similarity m as the column vector; S3, calculating the average relative error between the subsequence and the comparison sequence Xc through the weight distribution module (300), obtaining the error vector Am, obtaining the weight vector from the error vector and the similarity vector, and performing weighted summation on the adjusted prediction matrix by the weight vector to obtain the prediction result of the future chlorophyll a concentration; S4. The display alarm module (400) compares the predicted chlorophyll a concentration value with the chlorophyll a concentration threshold of the reservoir, determines that the current chlorophyll a concentration change trend of the lake is similar to the situation before the algal bloom outbreak in history, and triggers an algal bloom warning.
Citation Information
Cited By
Water eutrophication monitoring and pollution early warning system
CN120971678A