Plateau lake water quality and water quantity prediction method based on machine learning
By employing a machine learning-based hybrid neural network architecture and dynamic adjustment mechanism, the problems of data instability and complexity in the prediction of water quality and quantity in plateau lakes are solved, achieving high-precision and adaptive prediction results that are suitable for water quality and quantity management in plateau lakes.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-11
- Publication Date
- 2026-04-03
AI Technical Summary
Existing methods for predicting water quality and quantity in plateau lakes cannot effectively address the instability, incompleteness, and inaccuracy of monitoring data. They are also difficult to dynamically adjust models and monitoring plans according to different prediction accuracies, and lack comprehensive analysis and optimization of data from multiple monitoring points, thus failing to meet the need for accurate prediction in the complex environment of plateau lakes.
A machine learning-based hybrid neural network architecture is adopted, combining long short-term memory networks, attention mechanisms, and convolutional layers. Multiple logical judgment conditions are used to determine prediction accuracy information, dynamically adjust the monitoring plan and model parameters, generate comprehensive prediction results, and generate the final prediction result through a weighted fusion algorithm, while updating the model parameters in real time.
It improves the accuracy and adaptability of water quality and quantity prediction for plateau lakes, enabling flexible responses to changes in data quality and dynamic environmental changes. It solves the problem of insufficient accuracy of traditional methods when dealing with complex data, and realizes dynamic optimization and adaptability of the model.
Smart Images

Figure CN121786536A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of environmental science and artificial intelligence, and more specifically to a machine learning-based method for predicting the water quality and quantity of plateau lakes. Background Technology
[0002] Accurate prediction is crucial for ecological protection and rational water resource utilization in the management of water quality and quantity in plateau lakes. Traditional methods mainly rely on empirical models or simple statistical analysis. These methods often have limitations when dealing with complex plateau lake systems because they cannot fully consider the dynamic interactions and nonlinear relationships of multiple factors. In recent years, with the development of machine learning technology, its application in water quality and quantity prediction has gradually attracted attention. However, existing methods are mostly focused on the application of single models and have failed to fully integrate the special geographical and hydrological characteristics of plateau lakes, resulting in insufficient prediction accuracy and difficulty in adapting to the complex situations of different monitoring points.
[0003] In implementing the embodiments of the present invention, the prior art has at least the following problems or defects: the existing prediction methods cannot effectively handle the instability, incompleteness and inaccuracy of monitoring data, it is difficult to dynamically adjust the model and monitoring plan according to different prediction accuracies, and there is a lack of comprehensive analysis and optimization of data from multiple monitoring points, which cannot meet the needs of accurate prediction in the complex environment of plateau lakes. Summary of the Invention
[0004] This invention provides a machine learning-based method for predicting water quality and quantity in plateau lakes, comprising: Obtain monitoring data and corresponding prediction configurations for each river flowing into the target plateau lake within the target historical time period; Based on the monitoring data and the prediction configuration, the prediction accuracy information corresponding to the monitoring point is determined through multiple logical judgment conditions, wherein the logical judgment conditions include judgments based on the stability, completeness and accuracy of the monitoring data; Determine the prediction level corresponding to the prediction accuracy information, wherein determining the prediction level includes comparing the prediction accuracy information with multiple threshold intervals; Based on the prediction level, a target machine learning model is selected from multiple candidate machine learning models, wherein the multiple candidate machine learning models adopt the same hybrid neural network architecture, the hybrid neural network architecture includes an input layer, multiple hidden layers and an output layer, the multiple hidden layers include a long short-term memory network layer, an attention mechanism layer and a convolutional layer; Using the target machine learning model, generate the estimated water quality and quantity sequences corresponding to each water quality parameter in the target plateau lake within the target future time period; Based on the predicted level set and each estimated water quality and quantity sequence, the change rate corresponding to each water quality parameter and the monitoring plan corresponding to each monitoring point are adaptively adjusted to obtain the prediction result information. Based on the estimated water quality and quantity sequences and corresponding prediction levels of all monitoring points, a comprehensive prediction result for the target plateau lake is generated through a weighted fusion algorithm. Output the prediction result information and the comprehensive prediction result; Based on the comparison between the predicted results and the actual monitoring data, the parameters of the target machine learning model are updated.
[0005] Further, determining the prediction accuracy information corresponding to the monitoring point based on the monitoring data and the prediction configuration includes: Based on the monitoring data and the prediction configuration, at least one prediction index score corresponding to the prediction index set is determined, wherein the prediction index set includes: pH value, dissolved oxygen concentration, turbidity, total phosphorus concentration, total nitrogen concentration, water level, flow rate, water storage capacity, water temperature, and conductivity. The at least one predicted index score is input into a pre-trained scoring model to generate a prediction accuracy score, wherein the scoring model is trained based on historical monitoring data and corresponding prediction accuracy labels, and the scoring model includes an input layer, multiple hidden layers and an output layer, wherein the multiple hidden layers include fully connected layers and activation function layers. The prediction accuracy information is generated based on the prediction accuracy score.
[0006] Furthermore, after determining the prediction accuracy information corresponding to the monitoring point based on the monitoring data and the prediction configuration, the method further includes: Based on the prediction level, the number of days the target monitoring point has been running in the target historical time period, and the prediction configuration, operation adjustment information is generated; Based on the prediction level, the total number of parameters of the target monitoring point in the target historical time period, and the prediction configuration, parameter adjustment information is generated; Based on the operational adjustment information and the parameter adjustment information, target adjustment information is generated; The target adjustment information is sent to the management terminal corresponding to the monitoring point via target communication encryption.
[0007] Furthermore, the monitoring data is periodically stored in the monitoring database; and the monitoring data is generated through the following steps: In response to determining that the monitoring data acquisition time has been reached, the monitoring dataset of the monitoring point within the target historical time period is acquired; For each monitoring data point in the monitoring dataset, perform the following steps: obtain the duration between the start time of data collection and the completion time of transmission for the monitoring data point, so as to obtain the collection duration; In response to determining that the acquisition duration is less than a preset duration, the monitoring data point is identified as a high-quality data point; Based on the types of parameters in the monitoring data points, determine the number of parameter types corresponding to the monitoring data points; obtain the number of data errors of the monitoring points within the target historical time period; Based on the monitoring dataset, the high-quality data point set, and the parameter type set, generate the total number of data points, the proportion of high-quality data, the number of data errors, and the daily average number of parameter types corresponding to the monitoring dataset; The monitoring data is generated by summing the total number of data points, the proportion of high-quality data, the daily average number of parameter types, and the number of data errors. Before generating the monitoring data, outlier detection and processing are performed on the monitoring dataset, including removing monitoring data points that exceed a preset range.
[0008] Further, determining the prediction level corresponding to the prediction accuracy information includes: In response to determining that the prediction accuracy information is in a first score interval, the prediction level corresponding to the monitoring point is configured as the first level; In response to determining that the prediction accuracy information is in the second score interval, the prediction level corresponding to the monitoring point is configured as the second level; In response to determining that the prediction accuracy information is in the third score interval, the prediction level corresponding to the monitoring point is configured as the third level; In response to determining that the prediction accuracy information is in the fourth score interval, the prediction level corresponding to the monitoring point is configured as the fourth level; In response to determining that the prediction accuracy information is in the fifth score interval, the prediction level corresponding to the monitoring point is configured as the fifth level, wherein the first score interval, the second score interval, the third score interval, the fourth score interval and the fifth score interval are sequentially consecutive score intervals.
[0009] Furthermore, before performing the following steps for each monitoring point corresponding to the target plateau lake, the method further includes: Based on the geographical and hydrological characteristics of the target plateau lake, the lake is divided into multiple monitoring areas; For each monitoring area, the monitoring task type is determined, which includes fixed monitoring tasks and mobile monitoring tasks; Based on the importance of the monitoring area and the quality of historical data, a monitoring priority is assigned to each monitoring area; Based on the monitoring task type and monitoring priority, a monitoring task set is generated, which includes the monitoring frequency, monitoring parameter list and monitoring duration for each monitoring area. The monitoring task set is allocated to each monitoring point according to a dynamic allocation algorithm, which optimizes the allocation based on the device status, geographical location and current workload of the monitoring point. Generate a monitoring task execution plan, which includes the task sequence and time schedule for each monitoring point; Send the monitoring task execution plan to the mobile devices corresponding to each monitoring point; The monitoring point is instructed to confirm receipt of the monitoring task on the mobile device and execute the monitoring task according to the task sequence; Monitor the progress of tasks in real time and dynamically adjust task allocation based on actual performance.
[0010] Furthermore, the method also includes: Based on the temporal characteristics of the monitoring data, time-related labels are generated, including data collection continuity labels, data timeliness labels, and seasonality characteristic labels. Based on the quality indicators of the monitoring data, quality labels are generated, including data integrity labels, data accuracy labels, and data consistency labels. Based on the prediction accuracy score and target adjustment information, a comprehensive label is generated, which includes a model suitability label, a prediction reliability label, and a system stability label. Based on time-based labels, an adaptive algorithm is used to adjust the data collection frequency. The formula for adjusting the collection frequency is as follows:
[0011] in This indicates the adjusted sampling frequency. Indicates the basic sampling frequency. The score represents the continuity of data collection. The score represents the timeliness of the data label. The score represents the seasonality characteristic label. , , For adjustment coefficients; Based on quality-related labels, quality control algorithms are used to adjust the quality control parameters of monitoring data, including adjusting data verification thresholds, anomaly detection sensitivity, and data completion strategies. Based on comprehensive class labels, a hyperparameter optimization algorithm is used to adjust the hyperparameters of the target machine learning model, including the learning rate, batch size, and regularization coefficient.
[0012] Furthermore, the plurality of candidate machine learning models include: For the first prediction level, the corresponding candidate machine learning model includes a 2-layer LSTM network with 128 units per layer, 4 attention heads, and a 3-layer convolutional network. For the second prediction level, the corresponding candidate machine learning model includes a 3-layer LSTM network with 256 units per layer, 8 attention heads, and a 4-layer convolutional network. For the third prediction level, the corresponding candidate machine learning model includes a 4-layer LSTM network with 512 units per layer, 16 attention heads, and a 5-layer convolutional network. For the fourth prediction level, the corresponding candidate machine learning model includes a 5-layer LSTM network with 1024 units per layer, 32 attention heads, and a 6-layer convolutional network. For the fifth prediction level, the corresponding candidate machine learning model includes a 6-layer LSTM network with 2048 units per layer, 64 attention heads, and a 7-layer convolutional network.
[0013] Furthermore, the generated predicted water quality and quantity sequences corresponding to various water quality parameters in the target plateau lake within the target future time period include: Obtain the historical sequence data corresponding to the monitoring data; The historical sequence data is preprocessed to obtain standardized sequence data, wherein the preprocessing includes missing value imputation, smoothing and feature extraction; The standardized sequence data is input into the target machine learning model to obtain the predicted sequence data for the target future time period. The predicted sequence data is post-processed to generate the estimated water quality and quantity sequence, wherein the post-processing includes inverse standardization and uncertainty quantification.
[0014] Furthermore, the target machine learning model is trained through the following steps: Obtain a training dataset, which includes multiple historical monitoring sequences and corresponding real future sequences; The training dataset is divided into a training set and a validation set; The initial machine learning model is trained using the training set, and the model parameters are updated by minimizing the loss function, wherein the loss function is a combined loss function, including a mean squared error term, a mean absolute error term, and a regularization term. The model during training is evaluated using the validation set to select the optimal model parameters; The model corresponding to the optimal model parameters is taken as the target machine learning model; The formula for the combined loss function is as follows:
[0015] in, This represents the value of the loss function. Indicates the number of samples. Indicates the first A true value, Indicates the first One predicted value, , , These are weighting coefficients. Represents the model parameter vector. Model parameter vector The square of the L2 norm.
[0016] The embodiments of the present invention have at least the following beneficial effects: 1. By introducing a machine learning-based hybrid neural network architecture, combining long short-term memory networks, attention mechanisms, and convolutional layers, the complex temporal characteristics and spatial correlations of water quality and quantity data from plateau lakes can be effectively processed. This architecture can capture long-term dependencies, local features, and key information in the data, improving prediction accuracy and solving the problem of insufficient accuracy of traditional methods when processing complex data.
[0017] 2. A dynamic adjustment mechanism is adopted to adjust the monitoring plan and model parameters in real time based on the stability of the monitoring data and the accuracy of the prediction. This method can adapt to the complex conditions of different monitoring points, flexibly respond to changes in data quality and dynamic environmental changes, improve the adaptability and reliability of the prediction model, and solve the problems of model fixation and difficulty in dynamic optimization in existing technologies. Attached Figure Description
[0018] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. The following drawings are not intentionally drawn to scale to actual size; their focus is on illustrating the main points of this disclosure.
[0019] Figure 1 This is a flowchart illustrating a machine learning-based method for predicting water quality and quantity in plateau lakes, as provided in an embodiment of the present invention. Detailed Implementation
[0020] The following is for reference. Figure 1 , Figure 1 This is a flowchart illustrating a machine learning-based method for predicting water quality and quantity in plateau lakes, as provided in an embodiment of the present invention. Figure 1As shown, a machine learning-based method for predicting the water quality and quantity of plateau lakes includes: S1. Obtain the monitoring data and corresponding prediction configuration of each river flowing into the target plateau lake within the target historical time period; S2. Based on the monitoring data and the prediction configuration, determine the prediction accuracy information corresponding to the monitoring point through multiple logical judgment conditions, wherein the logical judgment conditions include judgments based on the stability, completeness and accuracy of the monitoring data; S3. Determine the prediction level corresponding to the prediction accuracy information, wherein determining the prediction level includes comparing the prediction accuracy information with multiple threshold intervals; S4. Based on the prediction level, select a target machine learning model from multiple candidate machine learning models, wherein the multiple candidate machine learning models adopt the same hybrid neural network architecture, the hybrid neural network architecture includes an input layer, multiple hidden layers and an output layer, the multiple hidden layers include a long short-term memory network layer, an attention mechanism layer and a convolutional layer; S5. Using the target machine learning model, generate the estimated water quality and quantity sequence corresponding to each water quality parameter in the target plateau lake within the target future time period; S6. Based on the predicted level set and each estimated water quality and quantity sequence, adaptive adjustments are made to the change rate corresponding to each water quality parameter and the monitoring plan corresponding to each monitoring point to obtain the prediction result information. S7. Based on the estimated water quality and quantity sequences and corresponding prediction levels of all monitoring points, a comprehensive prediction result for the target plateau lake is generated through a weighted fusion algorithm. S8. Based on the comprehensive prediction results and combined with the preset management objectives, generate a management and control plan for the rivers flowing into the lake or the lake itself; output the prediction result information, the comprehensive prediction results, and the management and control plan; S9. Based on the comparison between the predicted results and the actual monitoring data, update the parameters of the target machine learning model.
[0021] Monitoring data is acquired periodically and stored in a monitoring database. The monitoring data includes indicators such as the total number of data points, the percentage of high-quality data, the daily average number of parameter types, and the number of data errors. High-quality data points refer to those collected for less than a preset time; these data points are considered to have high reliability. The number of parameter types refers to the number of water quality and quantity parameters included in each monitoring data point. When determining prediction accuracy, a prediction accuracy score is generated based on these indicators of the monitoring data, thereby determining the prediction level. Prediction levels are divided into multiple levels, such as Level 1 to Level 5, each corresponding to different prediction accuracy ranges. Each prediction level corresponds to a different candidate machine learning model configuration. For example, Level 1 corresponds to a simpler model architecture, such as a 2-layer LSTM network with 128 units per layer, 4 attention heads, and 3 convolutional layers; while Level 5 corresponds to a more complex model architecture, such as a 6-layer LSTM network with 2048 units per layer, 64 attention heads, and 7 convolutional layers. These model configurations reflect the differences in model complexity under different prediction accuracy requirements.
[0022] As a relatively closed body of water, the final water quality and quantity of a plateau lake are the result of the combined input fluxes from all rivers flowing into it. The absence of any single river will prevent the model from fully and accurately establishing the causal relationship between water and lake changes. Only by ensuring that every river flowing into the lake has a monitoring point can complete material and energy input information for the entire watershed be collected, thereby guaranteeing the accuracy and reliability of the final lake-specific predictions.
[0023] When generating the estimated water quality and quantity sequence, historical sequence data corresponding to the monitoring data are acquired and preprocessed, including missing value imputation, smoothing, and feature extraction, to obtain standardized sequence data. This standardized sequence data is then input into the target machine learning model to obtain predicted sequence data for the target future time period. Finally, post-processing, such as inverse standardization and uncertainty quantification, is performed on the predicted sequence data to generate the final estimated water quality and quantity sequence. During model training, a combined loss function is used to optimize the model. This loss function includes a mean squared error term, a mean absolute error term, and a regularization term. Minimizing this loss function updates the model parameters, thereby improving the model's predictive performance.
[0024] By comparing the predicted results with preset management targets, when water quality or quantity indicators are found to be about to exceed standards, the system first identifies the main rivers flowing into the lake that are causing the risk. Subsequently, the system retrieves treatment measures for this type of problem from a pre-stored knowledge base. These measures, compiled in advance based on expert experience, include various types such as flow regulation and pollution source control. The system simulates the implementation effects of each measure, comprehensively considering the degree of improvement and implementation costs, and selects the optimal set of measures to form the final plan. This plan clearly specifies the specific rivers requiring regulation, the specific measures to be taken, the intensity of implementation, and the time frame for implementation, providing managers with directly actionable instructions.
[0025] In some embodiments, determining the prediction accuracy information corresponding to the monitoring point based on the monitoring data and the prediction configuration includes: Based on the monitoring data and the prediction configuration, at least one prediction index score corresponding to the prediction index set is determined, wherein the prediction index set includes: pH value, dissolved oxygen concentration, turbidity, total phosphorus concentration, total nitrogen concentration, water level, flow rate, water storage capacity, water temperature, and conductivity. The at least one predicted index score is input into a pre-trained scoring model to generate a prediction accuracy score, wherein the scoring model is trained based on historical monitoring data and corresponding prediction accuracy labels, and the scoring model includes an input layer, multiple hidden layers and an output layer, wherein the multiple hidden layers include fully connected layers and activation function layers. The prediction accuracy information is generated based on the prediction accuracy score.
[0026] Monitoring data and predictive configurations are used to determine at least one predictive indicator score corresponding to a set of predictive indicators. The predictive indicator set includes several key water quality parameters, such as pH, dissolved oxygen concentration, turbidity, and total phosphorus concentration, which are important indicators for measuring water quality and quantity. The predictive indicator scores reflect the performance of these parameters during the prediction process. By inputting these scores into a pre-trained scoring model, a prediction accuracy score can be generated. The scoring model is trained based on historical monitoring data and corresponding prediction accuracy labels. Its structure includes an input layer, multiple hidden layers, and an output layer. The hidden layers consist of fully connected layers and activation function layers. The prediction accuracy score is ultimately used to generate prediction accuracy information, thus providing a basis for subsequent prediction level classification. This method can effectively evaluate the predictive capability of monitoring points and provide a scientific basis for selecting appropriate machine learning models.
[0027] The predicted index score is obtained through the analysis of various water quality parameters in the monitoring data. For example, pH reflects the acidity or alkalinity of the water, dissolved oxygen concentration indicates the amount of dissolved oxygen in the water, and turbidity measures the transparency of the water. The numerical range and variation of these parameters directly affect the predicted index score. The scoring model is a deep learning model. Its input layer receives the predicted index score, the hidden layer extracts and transforms features from the input data through fully connected layers and activation function layers, and the output layer generates the prediction accuracy score. The prediction accuracy score is a quantitative indicator used to evaluate the reliability and accuracy of monitoring point data in prediction tasks. The calculation of this score is based on the model's learning from historical data. By comparing the difference between predicted and actual values, the model parameters are optimized, thereby improving the accuracy of the prediction accuracy score. The prediction accuracy information is generated based on the prediction accuracy score. It describes the predictive capability of the monitoring point in a more intuitive way and provides an important reference for subsequent prediction level classification and model selection.
[0028] In some embodiments, after determining the prediction accuracy information corresponding to the monitoring point based on the monitoring data and the prediction configuration, the method further includes: Based on the prediction level, the number of days the target monitoring point has been running in the target historical time period, and the prediction configuration, operation adjustment information is generated; Based on the prediction level, the total number of parameters of the target monitoring point in the target historical time period, and the prediction configuration, parameter adjustment information is generated; Based on the operational adjustment information and the parameter adjustment information, target adjustment information is generated; The target adjustment information is sent to the management terminal corresponding to the monitoring point via target communication encryption.
[0029] Operational adjustment information is generated based on the forecast level, the number of days the monitoring point has been operating within the target historical time period, and the forecast configuration. It aims to adjust the monitoring point's operational strategy to meet the requirements of forecast accuracy. Parameter adjustment information is generated based on the forecast level, the total number of parameters at the monitoring point within the target historical time period, and the forecast configuration. It is used to optimize the parameter settings at the monitoring point. Finally, the operational adjustment information and parameter adjustment information are integrated into target adjustment information and sent to the corresponding management terminal at the monitoring point via target communication encryption to ensure information security and accuracy. This process effectively improves the dynamic adaptability and forecast accuracy of the monitoring system.
[0030] The generation of operational adjustment information and parameter adjustment information is based on the prediction level and historical operational data of the monitoring points. The number of operational days refers to the actual number of days the monitoring point operated within the target historical time period, reflecting its operational efficiency and stability. The total number of parameters refers to the number of water quality parameter types collected by the monitoring point during that time period, reflecting its monitoring capability. Predictive configuration includes information such as monitoring frequency, a list of monitoring parameters, and monitoring duration, used to guide the operation and parameter adjustments of the monitoring points. Operational adjustment information involves adjusting the monitoring frequency or optimizing the monitoring duration to improve monitoring efficiency; while parameter adjustment information includes increasing or decreasing the types of monitoring parameters to meet the needs of different prediction levels. Target adjustment information is the result of combining operational adjustment information and parameter adjustment information, used to guide the specific operations of the monitoring points. Target communication encryption refers to using encryption algorithms to encrypt the adjustment information, ensuring the security of information during transmission and preventing data leakage or tampering.
[0031] If the prediction level is low, the monitoring frequency needs to be increased to improve the timeliness and accuracy of the data. The generation of parameter adjustment information can be determined by evaluating the contribution of each parameter to the prediction accuracy, deciding whether to add or remove certain parameters. For example, for parameters with a small impact on prediction accuracy, the number of monitoring sessions can be appropriately reduced to save resources. The generation of target adjustment information requires a comprehensive consideration of the needs of operational and parameter adjustments, optimizing the operational status of monitoring points through the development of detailed adjustment plans. During encrypted transmission, encryption algorithms such as AES (Advanced Encryption Standard) can be used to encrypt the adjustment information, ensuring secure transmission. In this way, monitoring points can dynamically adjust their operational strategies and parameter configurations according to the required prediction accuracy, thereby improving the overall performance and prediction accuracy of the monitoring system.
[0032] In some embodiments, the monitoring data is periodically stored in a monitoring database; and the monitoring data is generated through the following steps: In response to determining that the monitoring data acquisition time has been reached, the monitoring dataset of the monitoring point within the target historical time period is acquired; For each monitoring data point in the monitoring dataset, perform the following steps: obtain the duration between the start time of data collection and the completion time of transmission for the monitoring data point, so as to obtain the collection duration; In response to determining that the acquisition duration is less than a preset duration, the monitoring data point is identified as a high-quality data point; Based on the types of parameters in the monitoring data points, determine the number of parameter types corresponding to the monitoring data points; obtain the number of data errors of the monitoring points within the target historical time period; Based on the monitoring dataset, the high-quality data point set, and the parameter type set, generate the total number of data points, the proportion of high-quality data, the number of data errors, and the daily average number of parameter types corresponding to the monitoring dataset; The monitoring data is generated by summing the total number of data points, the proportion of high-quality data, the daily average number of parameter types, and the number of data errors. Before generating the monitoring data, outlier detection and processing are performed on the monitoring dataset, including removing monitoring data points that exceed a preset range.
[0033] The monitoring data is stored periodically in the monitoring database. Its generation process includes determining the data acquisition time and obtaining the monitoring dataset for the corresponding time period. For each monitoring data point, the duration between the start time of acquisition and the completion time of transmission is calculated; this is the acquisition duration. If the acquisition duration is less than a preset duration, the data point is considered a high-quality data point.
[0034] The system counts the number of parameter types and data errors for each monitoring data point, ultimately generating monitoring data that includes the total number of data points, the percentage of high-quality data, the daily average number of parameter types, and the number of data errors. Before generating the monitoring data, outlier detection and processing are performed on the monitoring dataset to remove data points that exceed preset ranges, ensuring the accuracy and reliability of the data.
[0035] The monitoring data acquisition time refers to the periodic data collection time points set by the system to ensure the timeliness of the data. The collection duration refers to the time interval from the start of data collection to the completion of transmission, used to evaluate the quality of the data points. The preset duration is a threshold used to determine whether a data point is a high-quality data point. The number of parameter types refers to the number of water quality parameter types contained in each monitoring data point, reflecting the richness of the data. The number of data errors refers to the number of erroneous data points in the monitoring data within the target historical time period, used to evaluate the accuracy of the data. The total number of data points refers to the total number of data points collected within the target historical time period; the percentage of high-quality data is the ratio of the number of high-quality data points to the total number of data points, reflecting the overall quality of the data. The average number of parameter types per day refers to the average number of parameter types collected each day, used to evaluate the diversity of the data. These parameters collectively constitute the core content of the monitoring data, providing an important basis for subsequent data processing and prediction.
[0036] Outlier detection and handling require setting reasonable threshold ranges. For example, for pH values, the normal range is set to 6.5 to 8.5, and data points exceeding this range will be removed. Through these specific operational steps, the quality of monitoring data is ensured, providing high-quality data support for subsequent predictive model training and prediction result generation.
[0037] In some embodiments, determining the prediction level corresponding to the prediction accuracy information includes: In response to determining that the prediction accuracy information is in a first score interval, the prediction level corresponding to the monitoring point is configured as the first level; In response to determining that the prediction accuracy information is in the second score interval, the prediction level corresponding to the monitoring point is configured as the second level; In response to determining that the prediction accuracy information is in the third score interval, the prediction level corresponding to the monitoring point is configured as the third level; In response to determining that the prediction accuracy information is in the fourth score interval, the prediction level corresponding to the monitoring point is configured as the fourth level; In response to determining that the prediction accuracy information is in the fifth score interval, the prediction level corresponding to the monitoring point is configured as the fifth level, wherein the first score interval, the second score interval, the third score interval, the fourth score interval and the fifth score interval are sequentially consecutive score intervals.
[0038] Prediction accuracy information is a quantitative indicator generated based on the quality of monitoring data and the performance of the prediction model, used to evaluate the predictive capability of monitoring points. The score intervals are divided according to the numerical range of the prediction accuracy information, for example, from the first score interval to the fifth score interval, with each interval corresponding to a prediction level. In this way, the predictive capability of monitoring points can be divided into different levels; for example, the first level represents the highest accuracy, and the fifth level represents the lowest accuracy. This grading method provides a basis for subsequently selecting appropriate machine learning models, ensuring that the model's complexity matches the prediction accuracy requirements, thereby improving the accuracy and efficiency of predictions.
[0039] Prediction accuracy is a quantitative indicator reflecting the performance of a monitoring point in a prediction task, typically calculated using a scoring model. A score interval is a continuous range divided according to the numerical range of the prediction accuracy information; for example, the first score interval represents the highest range of prediction accuracy, while the fifth score interval represents the lowest. The prediction level is determined based on the score interval in which the prediction accuracy information falls, and is used to describe the predictive capability of the monitoring point. For example, the first level corresponds to the highest prediction accuracy and is suitable for scenarios with extremely high accuracy requirements; while the fifth level corresponds to the lowest prediction accuracy and is suitable for scenarios with relatively low accuracy requirements. This tiered approach allows for the selection of appropriate machine learning models based on different prediction accuracy needs, ensuring that the model's complexity matches the required accuracy. For example, for the first level, a more complex model is chosen to achieve high-precision predictions; while for the fifth level, a relatively simple model is chosen to save computational resources.
[0040] The division of score intervals can be adjusted according to the needs of the actual application scenario. For example, the first score interval can be set as the top 10% of the prediction accuracy score, representing the highest accuracy; while the fifth score interval can be set as the bottom 10%, representing the lowest accuracy. In practice, the criteria for dividing score intervals can be flexibly set according to the water quality characteristics of the lake and the importance of the monitoring task. For example, for lakes with drastic water quality changes, the range of score intervals can be appropriately narrowed to improve the accuracy of the classification. In this way, the division of prediction levels can more accurately reflect the predictive ability of the monitoring points, providing a scientific basis for subsequent model selection and prediction tasks.
[0041] In some embodiments, before performing the following steps for each monitoring point corresponding to the target plateau lake, the method further includes: Based on the geographical and hydrological characteristics of the target plateau lake, the lake is divided into multiple monitoring areas; For each monitoring area, the monitoring task type is determined, which includes fixed monitoring tasks and mobile monitoring tasks; Based on the importance of the monitoring area and the quality of historical data, a monitoring priority is assigned to each monitoring area; Based on the monitoring task type and monitoring priority, a monitoring task set is generated, which includes the monitoring frequency, monitoring parameter list and monitoring duration for each monitoring area. The monitoring task set is allocated to each monitoring point according to a dynamic allocation algorithm, which optimizes the allocation based on the device status, geographical location and current workload of the monitoring point. Generate a monitoring task execution plan, which includes the task sequence and time schedule for each monitoring point; Send the monitoring task execution plan to the mobile devices corresponding to each monitoring point; The monitoring point is instructed to confirm receipt of the monitoring task on the mobile device and execute the monitoring task according to the task sequence; Monitor the progress of tasks in real time and dynamically adjust task allocation based on actual performance.
[0042] Geographical features refer to the natural attributes of a lake, such as its topography, area, and depth, while hydrological features include dynamic attributes such as water flow velocity and water level changes. Monitoring tasks are categorized into fixed and mobile monitoring tasks. Fixed monitoring tasks are typically used for long-term, stable monitoring points, while mobile monitoring tasks are suitable for scenarios requiring flexible adjustments to monitoring locations. Monitoring priority is determined based on the importance of the monitoring area and the quality of historical data. Areas with high importance or poor data quality are given higher priority to allocate more monitoring resources. The monitoring task set includes monitoring frequency (e.g., how many times per day), a list of monitoring parameters (e.g., pH value, dissolved oxygen concentration), and monitoring duration (e.g., the duration of each monitoring session). The dynamic allocation algorithm is an optimization algorithm that allocates tasks based on the equipment status of the monitoring points (e.g., whether the equipment is operating normally), geographical location (e.g., whether the monitoring point is located in a critical area), and current workload (e.g., whether the equipment is overloaded), ensuring reasonable and efficient task allocation. The monitoring task execution plan is a detailed execution scheme, including the task sequence and time arrangement for each monitoring point, ensuring that monitoring tasks can be carried out as planned.
[0043] For a large plateau lake, it can be divided into several sub-regions based on its topography. Each sub-region's monitoring task type is determined according to its hydrological and geographical characteristics. Fixed monitoring tasks can be set up at key locations in the lake, such as the inlet and outlet, for long-term monitoring of water quality changes; mobile monitoring tasks can be used to monitor the lake center or areas with frequent water quality changes. Monitoring priorities can be quantified based on historical data quality; for example, areas with poor data quality can be assigned higher priority to improve data quality by increasing monitoring frequency. The dynamic allocation algorithm can be implemented using heuristic algorithms, such as genetic algorithms or simulated annealing algorithms. These algorithms can dynamically adjust task allocation based on the real-time status of monitoring points, ensuring efficient execution of monitoring tasks. For example, if equipment at a monitoring point malfunctions, the algorithm can automatically reassign tasks to other available monitoring points. In this way, the execution of monitoring tasks is more flexible and efficient, better adapting to the complex monitoring needs of plateau lakes.
[0044] In some embodiments, the method further includes: Based on the temporal characteristics of the monitoring data, time-related labels are generated, including data collection continuity labels, data timeliness labels, and seasonality characteristic labels. Based on the quality indicators of the monitoring data, quality labels are generated, including data integrity labels, data accuracy labels, and data consistency labels. Based on the prediction accuracy score and target adjustment information, a comprehensive label is generated, which includes a model suitability label, a prediction reliability label, and a system stability label. Based on time-based labels, an adaptive algorithm is used to adjust the data collection frequency. The formula for adjusting the collection frequency is as follows:
[0045] in This indicates the adjusted sampling frequency. Indicates the basic sampling frequency. The score represents the continuity of data collection. The score represents the timeliness of the data label. The score represents the seasonality characteristic label. , , For adjustment coefficients; Based on quality-related labels, quality control algorithms are used to adjust the quality control parameters of monitoring data, including adjusting data verification thresholds, anomaly detection sensitivity, and data completion strategies. Based on comprehensive class labels, a hyperparameter optimization algorithm is used to adjust the hyperparameters of the target machine learning model, including the learning rate, batch size, and regularization coefficient.
[0046] Time-related labels describe the characteristics of monitoring data over time. Data collection continuity labels reflect whether there are interruptions or missing data during collection; data timeliness labels indicate the timeliness of the data, i.e., whether the data was collected and transmitted within the specified time; seasonality labels reflect the pattern of data variation with the seasons. Quality-related labels assess data quality: data integrity labels indicate whether the data is complete; data accuracy labels measure the closeness of the data to the true value; data consistency labels assess the consistency of data across different times or monitoring points. Comprehensive labels evaluate the overall performance of the model and system: model suitability labels reflect whether the model is suitable for the current prediction task; prediction reliability labels measure the credibility of the prediction results; system stability labels assess the stability of the entire monitoring system. In the formula for adjusting the collection frequency, the base collection frequency... It is the preset initial acquisition frequency and adjustment coefficient. , , The weights are set according to actual needs and are used to balance the impact of different labels on the collection frequency.
[0047] In areas with poor data collection continuity, the value of α can be increased to improve the collection frequency. Quality control algorithms can adjust data validation thresholds, anomaly detection sensitivity, and data completion strategies. For example, for monitoring points with low data accuracy, the data validation threshold can be lowered to reduce false positives. Hyperparameter optimization algorithms are used to adjust the hyperparameters of the target machine learning model, such as the learning rate, batch size, and regularization coefficient. For example, if the model suitability label indicates that the current model's prediction performance for some monitoring points is poor, the training process can be optimized by adjusting the learning rate. Through these specific operational steps, dynamic optimization of monitoring data collection and processing can be achieved, improving the performance of the prediction model and the overall stability of the system.
[0048] In some embodiments, the plurality of candidate machine learning models include: For the first prediction level, the corresponding candidate machine learning model includes a 2-layer LSTM network with 128 units per layer, 4 attention heads, and a 3-layer convolutional network. For the second prediction level, the corresponding candidate machine learning model includes a 3-layer LSTM network with 256 units per layer, 8 attention heads, and a 4-layer convolutional network. For the third prediction level, the corresponding candidate machine learning model includes a 4-layer LSTM network with 512 units per layer, 16 attention heads, and a 5-layer convolutional network. For the fourth prediction level, the corresponding candidate machine learning model includes a 5-layer LSTM network with 1024 units per layer, 32 attention heads, and a 6-layer convolutional network. For the fifth prediction level, the corresponding candidate machine learning model includes a 6-layer LSTM network with 2048 units per layer, 64 attention heads, and a 7-layer convolutional network.
[0049] The model comprises an input layer, multiple hidden layers, and an output layer. The hidden layers consist of a Long Short-Term Memory (LSTM) layer, an attention mechanism layer, and convolutional layers. The LSTM layer handles long-term dependencies in the time-series data, the attention mechanism layer highlights important features, and the convolutional layers extract local features. The number of layers and units is adjusted according to the prediction level, achieving adaptive optimization for different accuracy requirements. This design ensures that, with limited resources, the model's performance matches the required prediction accuracy, improving both the efficiency and accuracy of predictions.
[0050] For the first prediction level, the model is configured with a 2-layer LSTM network (128 units per layer), 4 attention heads, and 3 convolutional layers. For the fifth prediction level, the model is configured with a 6-layer LSTM network (2048 units per layer), 64 attention heads, and 7 convolutional layers. The number of layers and units in the LSTM network reflects the depth and complexity of the model's processing of time-series data; more layers and units result in a stronger ability to capture long-term dependencies. The number of attention heads indicates the number of features the model can simultaneously focus on; more heads mean a stronger ability to identify important features. The number of layers in the convolutional network reflects the depth of the model's local feature extraction; more layers result in more refined local feature extraction. These parameters can be adjusted according to actual prediction accuracy requirements to achieve accurate predictions for different monitoring points.
[0051] The basic architecture of the model is determined based on the prediction level, such as selecting an appropriate number of LSTM layers and units. Then, an attention mechanism layer is constructed, setting the number of attention heads so that the model can effectively identify and process key features. During model training, input parameters include the time-series and spatial features of historical monitoring data, and model parameters are optimized by minimizing the loss function. For example, for a model at the first prediction level, the input parameters could be a time series of water quality parameters from the past 24 hours; the model learns patterns from this data to predict water quality changes in future time periods.
[0052] In some embodiments, generating the estimated water quality and quantity sequences corresponding to various water quality parameters in the target plateau lake within a target future time period includes: Obtain the historical sequence data corresponding to the monitoring data; The historical sequence data is preprocessed to obtain standardized sequence data, wherein the preprocessing includes missing value imputation, smoothing and feature extraction; The standardized sequence data is input into the target machine learning model to obtain the predicted sequence data for the target future time period. The predicted sequence data is post-processed to generate the estimated water quality and quantity sequence, wherein the post-processing includes inverse standardization and uncertainty quantification.
[0053] The division into training and validation sets ensures that the model learns patterns in the data during training, while the validation set prevents overfitting. The mean squared error (MSE) term measures the squared difference between the predicted and true values, the mean absolute error (MAE) term measures the absolute difference between the predicted and true values, and the regularization term limits the model's complexity and prevents overfitting. Weight coefficients. , , These coefficients, used to balance the importance of different error terms, can be adjusted according to actual needs. The model parameter vector w represents parameters such as weights and biases in the model. By optimizing these parameters, the loss function is minimized, thereby improving the model's predictive performance.
[0054] Preferably, the model training process can be further refined into the following steps: First, the training dataset is preprocessed, including data cleaning, standardization, and feature extraction, to improve data quality. Then, the preprocessed data is divided into a training set and a validation set, for example, a 70% training set and a 30% validation set. During training, the initial model is trained using the training set, and the model parameters are updated through backpropagation to minimize the combined loss function. In the combined loss function, the mean squared error term penalizes larger errors more severely, the mean absolute error term penalizes all errors evenly, and the regularization term prevents overfitting by limiting the size of the model parameters. The weight coefficients can be dynamically adjusted based on the model's performance on the validation set; for example, if the mean squared error of the model on the validation set is high, the value of α can be appropriately increased. Finally, the model performance is evaluated using the validation set, and the model parameters that perform best on the validation set are selected as the parameters of the target machine learning model. In this way, the model can be continuously optimized during training, improving the accuracy and reliability of predictions.
[0055] In some embodiments, the target machine learning model is trained through the following steps: Obtain a training dataset, which includes multiple historical monitoring sequences and corresponding real future sequences; The training dataset is divided into a training set and a validation set; The initial machine learning model is trained using the training set, and the model parameters are updated by minimizing the loss function, wherein the loss function is a combined loss function, including a mean squared error term, a mean absolute error term, and a regularization term. The model during training is evaluated using the validation set to select the optimal model parameters; The model corresponding to the optimal model parameters is taken as the target machine learning model; The formula for the combined loss function is as follows:
[0056] in, This represents the value of the loss function. Indicates the number of samples. Indicates the first A true value, Indicates the first One predicted value, , , These are weighting coefficients. Represents the model parameter vector. Model parameter vector The square of the L2 norm.
[0057] Historical data refers to the time series of water quality and quantity parameters recorded by monitoring points over a past period. This data forms the basis for model learning. Preprocessing steps include missing value imputation, which involves filling in missing parts of the data using appropriate methods, such as interpolation or time-series-based prediction methods; smoothing, which removes noise from the data through filtering and other methods to make the data smoother; and feature extraction, which extracts useful features for prediction from the raw data, such as calculating moving averages or extracting seasonal features.
[0058] A target machine learning model is a trained predictive model that generates future predictions based on historical input data. Predicted sequence data consists of the model's output predictions of water quality and quantity for a future time period. Inverse standardization in post-processing restores the standardized data to its original dimensions, while uncertainty quantification assesses the uncertainty of the prediction results, for example, by calculating confidence intervals to represent the reliability of the prediction.
[0059] For time-series data of water quality parameters, LSTM networks can be used to capture time dependencies, while convolutional layers can be combined to extract local features.
[0060] After model training is complete, preprocessed data is input into the model to generate predicted sequence data. In the post-processing stage, inverse standardization is used to restore the predicted values to actual water quality and quantity values, and statistical methods, such as calculating the standard deviation, are used to quantify the uncertainty of the prediction. For example, for the prediction of dissolved oxygen concentration, the predicted value output by the model is inversely standardized to obtain the specific concentration value, and the confidence interval of the predicted value is given through uncertainty quantification, thus providing more comprehensive reference information for water quality management.
[0061] The foregoing description is illustrative of the invention and should not be construed as limiting it. Although several exemplary embodiments of the invention have been described, those skilled in the art will readily understand that many modifications can be made to the exemplary embodiments without departing from the novel teachings and advantages of the invention. Therefore, all such modifications are intended to be included within the scope of the invention as defined in the claims. It should be understood that the foregoing description is illustrative of the invention and should not be construed as limiting it to the specific embodiments disclosed, and modifications to the disclosed embodiments and other embodiments are intended to be included within the scope of the appended claims. The invention is defined by the claims and their equivalents.
Claims
1. A method for predicting water quality and quantity in plateau lakes based on machine learning, characterized in that, include: Obtain monitoring data and corresponding prediction configurations for each river flowing into the target plateau lake within the target historical time period; Based on the monitoring data and the prediction configuration, the prediction accuracy information corresponding to the monitoring point is determined through multiple logical judgment conditions, wherein the logical judgment conditions include judgments based on the stability, completeness and accuracy of the monitoring data; Determine the prediction level corresponding to the prediction accuracy information, wherein determining the prediction level includes comparing the prediction accuracy information with multiple threshold intervals; Based on the prediction level, a target machine learning model is selected from multiple candidate machine learning models, wherein the multiple candidate machine learning models adopt the same hybrid neural network architecture, the hybrid neural network architecture includes an input layer, multiple hidden layers and an output layer, the multiple hidden layers include a long short-term memory network layer, an attention mechanism layer and a convolutional layer; Using the target machine learning model, generate the estimated water quality and quantity sequences corresponding to each water quality parameter in the target plateau lake within the target future time period; Based on the predicted level set and each estimated water quality and quantity sequence, the change rate corresponding to each water quality parameter and the monitoring plan corresponding to each monitoring point are adaptively adjusted to obtain the prediction result information. Based on the estimated water quality and quantity sequences and corresponding prediction levels of all monitoring points, a comprehensive prediction result for the target plateau lake is generated through a weighted fusion algorithm. Based on the comprehensive prediction results and combined with the preset management objectives, a management and control plan for the rivers flowing into the lake or the lake itself is generated; the prediction result information, the comprehensive prediction results, and the management and control plan are output. Based on the comparison between the predicted results and the actual monitoring data, the parameters of the target machine learning model are updated.
2. The method according to claim 1, characterized in that, The step of determining the prediction accuracy information corresponding to the monitoring point based on the monitoring data and the prediction configuration includes: Based on the monitoring data and the prediction configuration, at least one prediction index score corresponding to the prediction index set is determined, wherein the prediction index set includes: pH value, dissolved oxygen concentration, turbidity, total phosphorus concentration, total nitrogen concentration, water level, flow rate, water storage capacity, water temperature, and conductivity. The at least one predicted index score is input into a pre-trained scoring model to generate a prediction accuracy score, wherein the scoring model is trained based on historical monitoring data and corresponding prediction accuracy labels, and the scoring model includes an input layer, multiple hidden layers and an output layer, wherein the multiple hidden layers include fully connected layers and activation function layers. The prediction accuracy information is generated based on the prediction accuracy score.
3. The method according to claim 2, characterized in that, After determining the prediction accuracy information corresponding to the monitoring point based on the monitoring data and the prediction configuration, the method further includes: Based on the prediction level, the number of days the target monitoring point has been running in the target historical time period, and the prediction configuration, operation adjustment information is generated; Based on the prediction level, the total number of parameters of the target monitoring point in the target historical time period, and the prediction configuration, parameter adjustment information is generated; Based on the operational adjustment information and the parameter adjustment information, target adjustment information is generated; The target adjustment information is sent to the management terminal corresponding to the monitoring point via target communication encryption.
4. The method according to claim 1, characterized in that, The monitoring data is periodically stored in the monitoring database; and the monitoring data is generated through the following steps: In response to determining that the monitoring data acquisition time has been reached, the monitoring dataset of the monitoring point within the target historical time period is acquired; For each monitoring data point in the monitoring dataset, perform the following steps: obtain the duration between the start time of data collection and the completion time of transmission for the monitoring data point, so as to obtain the collection duration; In response to determining that the acquisition duration is less than a preset duration, the monitoring data point is identified as a high-quality data point; Based on the types of parameters in the monitoring data points, determine the number of parameter types corresponding to the monitoring data points; obtain the number of data errors of the monitoring points within the target historical time period; Based on the monitoring dataset, the high-quality data point set, and the parameter type set, generate the total number of data points, the proportion of high-quality data, the number of data errors, and the daily average number of parameter types corresponding to the monitoring dataset; The monitoring data is generated by summing the total number of data points, the proportion of high-quality data, the daily average number of parameter types, and the number of data errors. Before generating the monitoring data, outlier detection and processing are performed on the monitoring dataset, including removing monitoring data points that exceed a preset range.
5. The method according to claim 1, characterized in that, Determining the prediction level corresponding to the prediction accuracy information includes: In response to determining that the prediction accuracy information is in a first score interval, the prediction level corresponding to the monitoring point is configured as the first level; In response to determining that the prediction accuracy information is in the second score interval, the prediction level corresponding to the monitoring point is configured as the second level; In response to determining that the prediction accuracy information is in the third score interval, the prediction level corresponding to the monitoring point is configured as the third level; In response to determining that the prediction accuracy information is in the fourth score interval, the prediction level corresponding to the monitoring point is configured as the fourth level; In response to determining that the prediction accuracy information is in the fifth score interval, the prediction level corresponding to the monitoring point is configured as the fifth level, wherein the first score interval, the second score interval, the third score interval, the fourth score interval and the fifth score interval are sequentially consecutive score intervals.
6. The method according to claim 1, characterized in that, Before performing the following steps for each monitoring point corresponding to the target plateau lake, the method further includes: Based on the geographical and hydrological characteristics of the target plateau lake, the lake is divided into multiple monitoring areas; For each monitoring area, the monitoring task type is determined, which includes fixed monitoring tasks and mobile monitoring tasks; Based on the importance of the monitoring area and the quality of historical data, a monitoring priority is assigned to each monitoring area; Based on the monitoring task type and monitoring priority, a monitoring task set is generated, which includes the monitoring frequency, monitoring parameter list and monitoring duration for each monitoring area. The monitoring task set is allocated to each monitoring point according to a dynamic allocation algorithm, which optimizes the allocation based on the device status, geographical location and current workload of the monitoring point. Generate a monitoring task execution plan, which includes the task sequence and time schedule for each monitoring point; Send the monitoring task execution plan to the mobile devices corresponding to each monitoring point; The monitoring point is instructed to confirm receipt of the monitoring task on the mobile device and execute the monitoring task according to the task sequence; Monitor the progress of tasks in real time and dynamically adjust task allocation based on actual performance.
7. The method according to claim 3, characterized in that, The method further includes: Based on the temporal characteristics of the monitoring data, time-related labels are generated, including data collection continuity labels, data timeliness labels, and seasonality characteristic labels. Based on the quality indicators of the monitoring data, quality labels are generated, including data integrity labels, data accuracy labels, and data consistency labels. Based on the prediction accuracy score and target adjustment information, a comprehensive label is generated, which includes a model suitability label, a prediction reliability label, and a system stability label. Based on time-based tags, an adaptive algorithm is used to adjust the collection frequency of monitoring data. The formula for adjusting the collection frequency is as follows: in This indicates the adjusted sampling frequency. Indicates the basic sampling frequency. The score represents the continuity of data collection. The score represents the timeliness of the data label. The score represents the seasonality characteristic label. , , For adjustment coefficients; Based on quality-related labels, quality control algorithms are used to adjust the quality control parameters of monitoring data, including adjusting data verification thresholds, anomaly detection sensitivity, and data completion strategies. Based on comprehensive class labels, a hyperparameter optimization algorithm is used to adjust the hyperparameters of the target machine learning model, including the learning rate, batch size, and regularization coefficient.
8. The method according to claim 1, characterized in that, The plurality of candidate machine learning models include: For the first prediction level, the corresponding candidate machine learning model includes a 2-layer LSTM network with 128 units per layer, 4 attention heads, and a 3-layer convolutional network. For the second prediction level, the corresponding candidate machine learning model includes a 3-layer LSTM network with 256 units per layer, 8 attention heads, and a 4-layer convolutional network. For the third prediction level, the corresponding candidate machine learning model includes a 4-layer LSTM network with 512 units per layer, 16 attention heads, and a 5-layer convolutional network. For the fourth prediction level, the corresponding candidate machine learning model includes a 5-layer LSTM network with 1024 units per layer, 32 attention heads, and a 6-layer convolutional network. For the fifth prediction level, the corresponding candidate machine learning model includes a 6-layer LSTM network with 2048 units per layer, 64 attention heads, and a 7-layer convolutional network.
9. The method according to claim 1, characterized in that, The estimated water quality and quantity sequences corresponding to various water quality parameters in the target plateau lake within the target future time period include: Obtain the historical sequence data corresponding to the monitoring data; The historical sequence data is preprocessed to obtain standardized sequence data, wherein the preprocessing includes missing value imputation, smoothing and feature extraction; The standardized sequence data is input into the target machine learning model to obtain the predicted sequence data for the target future time period. The predicted sequence data is post-processed to generate the estimated water quality and quantity sequence, wherein the post-processing includes inverse standardization and uncertainty quantification.
10. The method according to claim 1, characterized in that, The target machine learning model is trained through the following steps: Obtain a training dataset, which includes multiple historical monitoring sequences and corresponding real future sequences; The training dataset is divided into a training set and a validation set; The initial machine learning model is trained using the training set, and the model parameters are updated by minimizing the loss function, wherein the loss function is a combined loss function, including a mean squared error term, a mean absolute error term, and a regularization term. The model during training is evaluated using the validation set to select the optimal model parameters; The model corresponding to the optimal model parameters is taken as the target machine learning model; The formula for the combined loss function is as follows: in, This represents the value of the loss function. Indicates the number of samples. Indicates the first A true value, Indicates the first One predicted value, , , These are weighting coefficients. Represents the model parameter vector. Model parameter vector The square of the L2 norm.