Step-by-step land-river-lake water quality prediction method, model and computer equipment
Through the step-by-step ‘ground-river-lake’ water quality prediction method and LSTM model, and the prediction results are explained in combination with the SHAP model, the problem of high cost, complex operations and unexplained prediction of existing water quality monitoring methods is solved, and more accurate and explainable water quality prediction is achieved.
Patent Information
- Application Number
- CN202510114446.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-06-24
AI Technical Summary
The existing water quality monitoring methods have problems such as high equipment costs, complex operations and limited coverage. The ‘black box’ characteristics of machine learning and deep learning models lead to a lack of interpretability in predicted results, making it difficult to accurately capture the complex water quality dynamics in large-scale watersheds.
The step-by-step ‘ground-river-lake’ water quality prediction method is used to predict and parameter estimate the water quality characteristics of secondary river basins through multiple LSTM feature prediction models, and the contribution of the prediction results is explained in combination with the SHAP model.
The accuracy of water quality prediction results is improved, and by interpreting the contribution of model results, it provides a clear explanation of the impact characteristics of water quality changes, supporting more effective water quality protection measures.
Smart Images

Figure CN120197803A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of water quality detection, and particularly to the fields of a distributed "ground-river-lake" water quality prediction method, model, and computer equipment. Background Art
[0002] Currently, the water quality monitoring methods for large lakes and river basins mainly rely on high-precision real-time monitoring technologies such as chemical analysis and spectral monitoring. These technologies can provide accurate water quality data, but there are problems such as high equipment costs, complex operations, and limited coverage. Especially in the long-term monitoring of large-scale river basins, these methods are difficult to promote.
[0003] In the long-term monitoring of large-scale river basins, the existing technologies usually use machine learning and deep learning models to complete the long-term detection tasks. These two technologies have advantages in dealing with complex non-linear relationships, but their "black box" characteristics lead to the lack of interpretability of the model prediction results, making it difficult for users to understand the specific impact of each input variable on the prediction results, thus restricting the application of the model in actual decision-making; in addition, traditional water quality monitoring models usually only consider the water quality prediction of a single stage, ignoring the complex dynamic relationship between river inputs and lake water quality, resulting in the existing models often being unable to accurately capture these complex system dynamics, leading to inaccurate or unexplainable prediction results. Summary of the Invention
[0004] Therefore, the present invention proposes to capture the impact of secondary river basins on the water quality of the main river basin based on a distributed framework, and perform scale alignment and normalization processing on different data during the data transfer process between different hierarchical models, which can successfully capture the complex dynamic associations between the characteristics of the primary and secondary river basins, thereby improving the accuracy of the prediction results.
[0005] Specifically, the present invention is implemented through the following technical solutions:
[0006] On the one hand, the present invention provides a distributed "ground-river-lake" water quality prediction method, which includes the following steps:
[0007] S10: Fit according to the social and economic characteristics and meteorological characteristics of the secondary river basin through a first LSTM feature prediction model to obtain the first predicted water quality feature data of the secondary river basin;
[0008] S20: Perform parameter estimation on the first predicted water quality feature data of the secondary river basin and the flow characteristics of the secondary river basin according to a water quality feature transformation function to obtain a set of predicted water quality feature data, and accumulate the set of predicted water quality feature data to obtain the second predicted water quality feature data of the secondary river basin;
[0009] S30: Fit the second predicted water quality characteristic data of the secondary watershed, the flow characteristics of the secondary watershed, the meteorological characteristics of the mainstream watershed, and the surface social and economic characteristics through the second LSTM characteristic prediction model to obtain the predicted water quality characteristic data of the mainstream watershed.
[0010] Further, the training process of the first LSTM model is as follows:
[0011] Fit the historical social and economic characteristics and historical meteorological characteristics of the input secondary watershed to obtain a predicted water quality characteristic data;
[0012] Calculate the absolute loss between the predicted water quality characteristic data and the historical water quality characteristic data of the secondary watershed;
[0013] Update the parameters of the first LSTM model according to the absolute loss until the absolute error is less than a preset threshold.
[0014] Further, the water quality characteristic transformation function is obtained through the following steps:
[0015] Calculate the Akaike information criterion values of multiple regression models based on the historical water quality characteristic data of the secondary watershed, and select the regression model with the smallest Akaike information criterion value as the optimal regression model;
[0016] Input the historical instantaneous water quality characteristics of the secondary watershed and the historical instantaneous flow characteristics of the secondary watershed corresponding to its time into the optimal regression model for parameter fitting to obtain the specified parameters of the optimal regression model, and generate a water quality characteristic transformation function according to the specified parameters of the optimal model.
[0017] Further, after step S30, it also includes:
[0018] S40: Obtain the relevant characteristics participating in the water quality characteristic prediction of the mainstream watershed and the predicted water quality characteristic data of the mainstream watershed, calculate the marginal contribution of each relevant characteristic to the predicted water quality characteristic data of the mainstream watershed respectively, and complete the ranking according to the magnitude of the marginal contribution value.
[0019] Further, the calculation formula of the marginal contribution is:
[0020]
[0021] Among them, Φ j represents the contribution of feature j to the model, Z j indicates whether the feature exists, Φ0 is a constant term, S represents all subsets used in the model, {x1,…,x n}\{x j} represents all features excluding feature j, and F(n)-F(n\j) represents the marginal contribution of including feature j to S.
[0022] Further, before step S10, it also includes data preprocessing of cleaning, processing, complementing, and standardizing the socio-economic characteristics and meteorological characteristics of the secondary basin.
[0023] On the other hand, the present invention also provides a distributed "ground-river-lake" water quality prediction model, which includes:
[0024] The first LSTM feature prediction model: used to fit the socio-economic characteristics and meteorological characteristics of the secondary basin to obtain the first predicted water quality feature data of the secondary basin;
[0025] The feature transformation model: used to perform parameter estimation on the first predicted water quality feature data of the secondary basin and the flow characteristics of the secondary basin according to a water quality feature transformation function to obtain a set of predicted water quality feature data, and accumulate the set of predicted water quality feature data to obtain the second predicted water quality feature data of the secondary basin;
[0026] The second LSTM feature prediction model: used to fit the second predicted water quality feature data of the secondary basin, the flow characteristics of the secondary basin, the meteorological characteristics of the main basin, and the surface socio-economic characteristics to obtain the predicted water quality feature data of the main basin.
[0027] Further, the distributed "ground-river-lake" water quality prediction model further includes:
[0028] The SHAP model: used to obtain the relevant features participating in the water quality feature prediction of the main basin and the predicted water quality feature data of the main basin, calculate the marginal contribution of each relevant feature to the predicted water quality feature data of the main basin respectively, and complete the ranking according to the magnitude of the marginal contribution value.
[0029] Further, the distributed "ground-river-lake" water quality prediction model further includes:
[0030] The data preprocessing module: used to perform data preprocessing of cleaning, processing, complementing, and standardizing the socio-economic characteristics and meteorological characteristics of the secondary basin.
[0031] On the other hand, the present invention also provides a computer device, which includes at least one memory and at least one processor;
[0032] The memory is used to store one or more programs; when the one or more programs are executed by the at least one processor, the at least one processor realizes the steps of a distributed "ground-river-lake" water quality prediction method as described in any one of the above.
[0033] Based on the above methods and models, it can be seen that the present invention creatively combines a distributed model architecture with a long short-term memory model (LSTM) for water quality prediction of large lakes or rivers. The first prediction water quality characteristics of the secondary domain are predicted once by multiple first LSTM feature prediction models, and then multiple second prediction water quality characteristics are obtained by estimating based on the multiple first prediction water quality characteristics output by the multiple first LSTM feature prediction models and the corresponding flow characteristics of each secondary basin. Finally, the second LSTM feature prediction model is used to perform a secondary prediction in combination with other characteristics that may affect the predicted water quality characteristics to obtain the prediction result of the mainstream domain water quality characteristics. A more accurate prediction result is obtained, and combined with the SHAP model, the influence weight result of each secondary basin on the mainstream domain and the marginal contribution of the relevant characteristics participating in the water quality characteristics prediction of the mainstream domain to the water quality characteristic data can also be obtained, which is more conducive to reasonably allocating and protecting water quality resources to protect the water quality environment of the mainstream domain in practical applications.
[0034] For better understanding and implementation, the present invention will be described in detail below with reference to the accompanying drawings. Description of the Drawings
[0035] Figure 1 It is a flowchart of a distributed water quality prediction method provided by the present invention;
[0036] Figure 2 For execution Figure 1 It is a structural block diagram of a distributed water quality prediction model for the prediction method shown;
[0037] Figure 3 It is a structural block diagram of an exemplary LSTM model provided by the present invention;
[0038] Figure 4 It is a flowchart of the transformation function generation process of the present invention;
[0039] Figure 5 It is a flowchart of another exemplary distributed water quality prediction method of the present invention;
[0040] Figure 6 It is a SHAP result diagram of the S1 site in the Dianchi Lake Basin in the third-stage model in the embodiment of the present invention. Detailed Embodiments
[0041] It should be clear that the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the embodiments of the present invention without creative efforts belong to the scope protected by the embodiments of the present invention.
[0042] The terms used in the embodiments of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the embodiments of the present invention. The singular forms "a", "the", and "said" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0043] It should be understood that the embodiments of the present invention are not limited to the exact structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the embodiments of the present invention is limited only by the appended claims.
[0044] The distributed model refers to an entity composed of multiple interconnected computers that cooperate to execute a common or different task under a set of system software (such as a distributed operating system or middleware). The distributed model is usually adopted to distribute the computing work among multiple computers, reduce the load and possible risks of centralized operation on a single computer, and provide high scalability, reliability, manageability, and flexibility. The core idea is to split complex tasks to improve the efficiency, security, and manageability of task execution.
[0045] The Long Short-Term Memory model (LSTM) is a neural network specifically designed to process time series data and can capture long-term dependencies in the data. Based on the RNN, LSTM adds a memory cell belt and three gate structures (input gate, output gate, forget gate) to overcome the problems of gradient vanishing and gradient explosion. LSTM controls the transmission and storage of information through three gate structures (forget gate, input gate, output gate), enabling it to effectively remember important information over a long time span.
[0046] Water quality data are mainly the detection data of detection sites, including physical, chemical, and biological parameters. Physical parameters include water temperature, suspended solids (SS), turbidity, transparency, conductivity, etc.; chemical parameters include ammonia nitrogen, nitrite nitrogen, nitrate nitrogen, total nitrogen, phosphate, and total phosphorus, etc., which are used to characterize the content of plant nutrient elements in water and reflect the organic pollution degree of water; and biological parameters include total bacteria count, coliform group, etc. In the embodiments of the present invention, the concentration of total phosphorus (TP) in chemical parameters is used as the core target variable for prediction to exemplify the present invention.
[0047] Meteorological characteristics obtain daily-scale data from the National Meteorological Information Center, including data such as daily precipitation, temperature, maximum wind speed, and average wind speed.
[0048] Socio - economic characteristics refer to data such as rural population, sewage treatment volume, chemical fertilizer application amount, total agricultural output value, and livestock production within the basin. These factors affecting the water quality of the basin are used as influencing factors and participate in the water quality prediction process.
[0049] The black - box model is established based on the input - output relationship, reflecting a general direct causal relationship among relevant factors.
[0050] The traditional single - stage LSTM model ignores the intermediate process and directly inputs all features for prediction. This makes the model unable to comprehensively capture the complex material migration process among land, sub - basins, and main basins, weakening the understanding and interpretation of the dynamic changes in water quality. The "black - box" characteristic of the single - stage model makes it difficult to explain the specific process of model prediction. Due to the neglect of the staged analysis of material migration, this limits the application value of the model in actual water quality management.
[0051] The present invention creatively combines a distributed model architecture with a long short - term memory model (LSTM) for water quality prediction of large lakes or rivers. Through multiple first LSTM feature prediction models, the predicted water quality characteristics of the sub - domain are predicted once. Then, based on the predicted water quality characteristics of the multiple first LSTM feature prediction models, combined with the remaining features that may affect the predicted water quality characteristics, a second LSTM feature prediction model is used for secondary prediction to obtain the predicted result of the water quality characteristics of the main basin, thus obtaining a more accurate prediction result.
[0052] Please refer to Figure 1 and Figure 2 , Figure 1 which is a flowchart of a distributed water quality prediction method provided by the present invention. Figure 2 For executing Figure 1 The structural block diagram of the distributed water quality prediction model for the prediction method shown; wherein the distributed water quality prediction model includes a first LSTM feature prediction model 10, a feature transformation model 20, and a second LSTM feature prediction model 30.
[0053] Among them, the first LSTM feature prediction model 10 is used to execute step S10: fitting the socio - economic characteristics and meteorological characteristics of the sub - basin to obtain the first predicted water quality characteristic data of the sub - basin.
[0054] The feature transformation model 20 is used to execute step S20: performing parameter estimation on the first predicted water quality characteristic data of the sub - basin and the flow characteristics of the sub - basin according to a water quality characteristic transformation function to obtain a set of predicted water quality characteristic data, and accumulating the set of predicted water quality characteristic data to obtain the second predicted water quality characteristic data of the sub - basin.
[0055] The second LSTM feature prediction model 30 is used to execute step S30: fitting the second predicted water quality feature data of the secondary basin, the flow characteristics of the secondary basin, the meteorological characteristics of the mainstream basin, and the surface socio-economic characteristics to obtain the predicted water quality feature data of the mainstream basin.
[0057] Specifically, the present invention will be described in conjunction with the following embodiments, and the application in the Dianchi Lake Basin will be used as an embodiment to explain the present technology.
[0058] The first stage (the first LSTM feature prediction model - prediction of the total phosphorus load of each secondary river flowing into Dianchi Lake)
[0059] The total phosphorus load of several rivers flowing into Dianchi Lake is used as the predicted water quality feature of the secondary basin; in this embodiment, the first LSTM feature prediction model aims to predict the total phosphorus load of the river, and the input features are the historical socio-economic characteristics, historical meteorological characteristics (such as precipitation, temperature, total agricultural output value, etc.) of the Dianchi Lake Basin, and the total phosphorus load of the historical river. In this stage, the long short-term memory model (LSTM) is used, and the structure of the LSTM model is as Figure 3 shown. It can be seen that it includes a memory cell belt, an input gate, an output gate, and a forget gate.
[0060] Forget Gate: Determines whether the information from the previous time step needs to be retained or forgotten. The output of the forget gate is calculated through the sigmoid activation function, and the value range is between 0 and 1. The formula is:
[0061] f(t) = σ(W f [h t-1 , X t )
[0062] Among them, W f is the weight of f(t) in the LSTM model, h t-1 is the state variable of the LSTM model at the previous time step, and X t is the input of the LSTM model at the current moment.
[0063] Input Gate: The input gate determines whether new information should be written into the current memory cell. The formula of the input gate is as follows:
[0064] i(t) = σ(W i [h t-1 , X t )
[0065] g(t) = tanh(W g [h t-1 , X t )
[0066] where, i(t) is the output of the input gate, representing the degree to which the current input is written into the memory cell, and W i is the weight matrix of the input gate, g(t) is the candidate information feature, and the tanh activation function is used to generate the feature map of the current input. W g is the weight matrix of the candidate memory.
[0067] Memory Cell Update: The state of the memory cell at the current time step is jointly updated by the state of the memory cell at the previous time step and the current input information. The formula is as follows:
[0068] C(t) = f(t)C(t - 1) + i(t)g(t)
[0069] C(t) is the state of the memory cell of the LSTM model at the current moment.
[0070] Output Gate: The output gate determines the output of the current hidden state information to the next time step. The formula of the output gate is as follows:
[0071] o(t) = σ(W o [h t-1 , X t )
[0072] h(t) = tanh(C(t))o(t)
[0073] Through the above process, the LSTM model can realize the prediction of the total phosphorus load output from the river to the Dianchi Lake Basin according to the preset corresponding social and economic characteristics, meteorological characteristics (such as precipitation, temperature, total agricultural output value, etc.) of the river.
[0074] The Second Stage (Feature Transformation Model - Parameter Estimation and Time Scale Transformation for Predicting Water Quality Characteristics)
[0075] In practical applications, only the monitoring data of the river pollutant concentration at a certain time interval can be obtained, that is, the time interval of the total phosphorus load of the historical river collected is relatively long. Then, the water quality characteristic data of the secondary basin obtained by the first LSTM feature prediction model may represent the total phosphorus load of the river flowing into the main basin after a relatively long time interval, and the long-time interval measurement will affect the water quality prediction result. Therefore, it is necessary to perform processing on the obtained data in terms of time scale. In the embodiment of the present invention, based on the monthly total phosphorus load data of 14 rivers flowing into the Dianchi Lake, in the embodiment of the present invention, the feature transformation model is the LOADEST model. According to the LOADEST model, the daily total phosphorus load of 14 rivers is estimated through a water quality characteristic transformation function, and then the estimation results are accumulated to obtain the monthly total phosphorus load data, and more accurate total phosphorus load data is obtained by accumulating the more accurate data with a smaller granularity. Specifically, please refer toFigure 4 , first, generate a water quality characteristic transformation function through the following steps:
[0076] According to the historical water quality characteristic data of the secondary watershed, calculate the Akaike information criterion values of multiple regression models, and select the regression model with the smallest Akaike information criterion value as the optimal regression model;
[0077] By establishing a multivariate non-linear relationship between the total phosphorus load and its corresponding flow rate in the secondary watershed, estimate the total phosphorus load and fill in the missing data to generate more accurate total phosphorus load data, so as to capture the dynamic changes of the total phosphorus load and improve the prediction accuracy of the model. Specifically: select the optimal model from 11 regression models, and the formulas of these models are as follows:
[0078]
[0079] Among them, is the load estimator; Q is the flow rate at the monitoring section; dtime is the decimal time; a1 is the first regression coefficient, a2 is the second regression coefficient, a3 is the third regression coefficient, a4 is the fourth regression coefficient, a5 is the fifth regression coefficient, a6 is the sixth regression coefficient; per is the research time period, that is, a total of 33 regression models can be obtained by three parameter estimation methods, and calculate the Akaike information criterion (AIC) of each regression model. AIC is an index to measure the fitting effect of a statistical model. The lower the AIC, the better the model fitting effect. LOADEST selects the model with the lowest AIC as the optimal model for this group of data sets.
[0080] Input the historical instantaneous water quality characteristics of the secondary watershed and the historical instantaneous flow characteristics of the secondary watershed corresponding to its time into the optimal regression model for parameter fitting, obtain the specified parameters of the optimal regression model, and generate a water quality characteristic transformation function according to the specified parameters of the optimal model.
[0081] Substitute the historical instantaneous flow characteristics into the optimal regression model to calculate the load estimator Compare the load estimator with the historical instantaneous water quality characteristics (phosphorus load) of the secondary watershed corresponding to this time, and continuously correct the coefficients of the regression model until the total error between the load estimator at each collection point and the historical instantaneous water quality characteristics of each secondary watershed is the lowest, and the specified parameters of the obtained optimal regression model are the water quality characteristic transformation function.
[0082] This water quality characteristic transformation function can estimate parameters based on the first predicted water quality characteristic data of the secondary watershed and the flow characteristics of the secondary watershed, obtain a set of predicted water quality characteristic data, and accumulate the set of predicted water quality characteristic data to obtain the second predicted water quality characteristic data of the secondary watershed.
[0083] Based on this, the total phosphorus load is summarized monthly to complete the data and convert the monthly total phosphorus load data into more continuous total phosphorus load data, thereby improving the prediction accuracy of the model. Through the above processing, the time-scale conversion and time-delay elimination of the data are completed, ensuring the effective connection between the first LSTM feature prediction model and the second LSTM feature prediction model.
[0084] The third stage (the second LSTM feature prediction model - prediction of the total phosphorus load in Dianchi Lake)
[0085] The second LSTM feature prediction model establishes the relationship between the secondary watersheds flowing into the mainstream and the mainstream. Its input features include the physicochemical characteristics of Dianchi Lake (such as water temperature, transparency, dissolved oxygen, permanganate index, and previous total phosphorus load), meteorological characteristics (such as precipitation, air temperature, and wind speed), the total phosphorus load of the river, and the river flow into the lake. Using the river flow as an input feature is to capture the stirring effect of the river inflow on the lake water and promote the potential process of endogenous release to improve the prediction accuracy of the model. Finally, the output feature of the second LSTM feature prediction model is the monthly monitored total phosphorus load of 10 monitoring stations (S1 - S10) in Dianchi Lake. Then, performance evaluation indicators are calculated to evaluate the model performance: The following three indicators are used to measure the total phosphorus prediction output by the second-stage model:
[0086] Mean Absolute Percentage Error (MAPE): To evaluate the relative error, the formula is:
[0087]
[0088] Mean Absolute Error (MAE): To evaluate the absolute deviation between the predicted value and the true value, the formula is:
[0089]
[0090] Root Mean Square Error (RMSE): To evaluate the square root of the error, the formula is:
[0091]
[0092] In the calculation of the error between the final prediction result and the actual value, the average MAPE of the second-stage model in the test dataset is 27.00%, the average MAE is 0.052 mg / L, and the RMSE is 0.075 mg / L. Compared with the traditional model at the corresponding monitoring stations, the second-stage model shows an average reduction of 3.16% in MAPE, an average reduction of 6.81% in MAE, and an average reduction of 8.86% in RMSE in the test dataset. This indicates that the overall error rate of the second-stage model is lower than that of the traditional model, and the prediction performance has been improved in terms of overall accuracy.
[0093] In summary, the proposed step-by-step water quality prediction method and model of the present invention, through the step-by-step processing method and model architecture, first perform a primary prediction on the water quality characteristics of the secondary watershed, then perform parameter estimation and time scale transformation on the predicted water quality characteristic data to ensure effective connection of the characteristics between the secondary watershed and the mainstream watershed, and then perform a secondary prediction based on the water quality characteristics after parameter estimation and the characteristics related to water quality, and finally obtain a more accurate prediction result.
[0094] However, according to the above prediction model, the uncertainty of the black box model cannot be solved. Only the water quality prediction result can be simply obtained, but the primary and secondary of the influencing characteristics and the influence weights when it changes cannot be known, which causes great trouble in the selection of water quality protection methods. Therefore, the present invention also adds a SHAP model 40 on the basis of the foregoing model to explain the contribution of each input characteristic to the prediction result. Please refer to Figure 4 , Figure 4 which is a flowchart of another exemplary step-by-step water quality prediction method of the present invention.
[0095] The SHAP model 40 is used to execute step S40: obtain the relevant characteristics participating in the water quality characteristic prediction of the mainstream watershed and the predicted water quality characteristic data of the mainstream watershed, calculate the marginal contribution of each relevant characteristic to the predicted water quality characteristic data of the mainstream watershed respectively, and complete the sorting according to the magnitude of the marginal contribution value.
[0096] Its calculation formula is as follows:
[0097]
[0098] Among them, Φ0 is a constant term, and Φ j represents the contribution of feature j to the model. Z j indicates whether the feature exists, S represents all subsets used in the model, {x1,…,x n}\{x j} represents all features excluding feature j, and F(n)-F(n\j) represents the marginal contribution of including feature j to S.
[0099] Specifically, please refer to Figure 6 , Figure 6 which is the SHAP result diagram of the S1 site in the third-stage model of the Dianchi Lake watershed in the embodiment of the present invention; to show the influence of relevant characteristics on the total phosphorus load prediction of this site, it can be seen that for the S1 site, transparency is the main contributing feature, and the contribution rates of wind speed and dissolved oxygen follow closely, and they are also important contributing features. Therefore, it can be clearly known that for the total phosphorus load, the main influencing items of the S1 site are these parameters such as transparency, wind speed, and dissolved oxygen. When the detected total phosphorus load is abnormal, the main influencing items can be focused on to effectively implement water quality protection.
[0100] In summary, the present invention creatively combines a distributed model architecture with a long short-term memory model (LSTM) for predicting the water quality of large lakes or rivers. Through multiple first LSTM feature prediction models, a primary prediction of the first predicted water quality features in the secondary domain is performed. Then, based on the multiple first predicted water quality features output by the multiple first LSTM feature prediction models and the corresponding flow features of each secondary basin, an estimation is made to obtain multiple second predicted water quality features. Finally, in combination with the remaining features that may affect the predicted water quality features, a second LSTM feature prediction model is used for a secondary prediction to obtain the prediction result of the water quality features in the main domain, resulting in a more accurate prediction result. In combination with the SHAP model, it is also possible to obtain the influence weight results of each secondary basin on the main domain, as well as the marginal contributions of the relevant features participating in the water quality feature prediction of the main domain to the water quality feature data; this is more conducive to reasonably allocating and protecting water quality resources in practical applications to protect the water quality environment of the main domain.
[0101] Based on the same inventive concept as described above, the present invention also provides an electronic device, which can be a server, a desktop computing device, or a mobile computing device (such as a laptop computing device, a handheld computing device, a tablet computer, a netbook, etc.) and other terminal devices. The device includes one or more processors and a memory, where the processor is used to execute a program to implement the above-described distributed water quality prediction method; the memory is used to store a computer program executable by the processor.
[0102] Based on the same inventive concept, the present invention also provides a computer-readable storage medium, corresponding to the embodiment of the above-described distributed water quality prediction method. The computer-readable storage medium stores a computer program thereon, and when the program is executed by a processor, it implements the steps recorded in any of the above embodiments.
[0103] The present invention may be embodied in the form of a computer program product implemented on one or more storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) that contain program code. Computer-usable storage media include permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information may be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include but are not limited to: phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device.
[0104] The above-described embodiments merely represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and the present invention also intends to include these changes and modifications.
Claims
1. A step-by-step "land-river-lake" water quality prediction method, characterized in that: include: S10: fitting the socio-economic characteristics and meteorological characteristics of the secondary watershed through the first LSTM feature prediction model to obtain the first predicted water quality characteristic data of the secondary watershed; S20: performing parameter estimation on the first predicted water quality characteristic data of the secondary watershed and the flow characteristics of the secondary watershed according to a water quality characteristic transformation function to obtain a set of predicted water quality characteristic data, and accumulating the set of predicted water quality characteristic data to obtain second predicted water quality characteristic data of the secondary watershed; S30: The second predicted water quality characteristic data of the secondary basin, the flow characteristics of the secondary basin, the meteorological characteristics of the mainstream basin, and the surface socio-economic characteristics are fitted through the second LSTM feature prediction model to obtain the predicted water quality characteristic data of the mainstream basin.
2. The step-by-step "land-river-lake" water quality prediction method according to claim 1 is characterized in that: The training process of the first LSTM model is as follows: According to the historical socio-economic characteristics and historical meteorological characteristics of the input sub-basin, a predicted water quality characteristic data is obtained by fitting; Calculate the absolute loss of the predicted water quality characteristic data compared to the historical water quality characteristic data for the said secondary watershed; Parameters of the first LSTM model are updated according to the absolute loss until the absolute error is less than a preset threshold.
3. The step-by-step "land-river-lake" water quality prediction method according to claim 1 is characterized in that: The water quality characteristic transformation function is obtained by the following steps: According to the historical water quality characteristic data of the secondary watershed, the Akaike information criterion values of multiple regression models are calculated, and the regression model with the smallest Akaike information criterion value is selected as the optimal regression model; The historical instantaneous water quality characteristics of the secondary basin and the historical instantaneous flow characteristics of the secondary basin corresponding to the time are input into the optimal regression model for parameter fitting to obtain the specified parameters of the optimal regression model, and a water quality characteristic transformation function is generated according to the specified parameters of the optimal model.
4. The step-by-step "land-river-lake" water quality prediction method according to claim 1 is characterized in that: After step S30, the following steps are also included: S40: Obtain relevant features involved in the prediction of water quality characteristics of the mainstream domain and the predicted water quality characteristic data of the mainstream domain, calculate the marginal contribution of each relevant feature to the predicted water quality characteristic data of the mainstream domain respectively, and perform sorting according to the size of the marginal contribution value.
5. A step-by-step "land-river-lake" water quality prediction method according to claim 4, characterized in that: The calculation formula of the edge contribution is: Among them, Φ j represents the contribution of feature j to the model, Z j indicates whether the feature exists, Φ0 is a constant term, S represents all subsets used in the model, {x1,…,x n }\{x j } represents all features excluding feature j, and F(n)-F(n\j) represents the marginal contribution of S including feature j.
6. A step-by-step "land-river-lake" water quality prediction method according to any one of claims 1 to 5, characterized in that: Before step S10, the process also includes data preprocessing for cleaning, processing, completing and standardizing the socio-economic characteristics and meteorological characteristics of the secondary watershed.
7. A step-by-step "land-river-lake" water quality prediction model, characterized in that: include: The first LSTM feature prediction model: used to fit the socio-economic characteristics and meteorological characteristics of the secondary watershed to obtain the first predicted water quality characteristic data of the secondary watershed; Characteristic transformation model: used for performing parameter estimation on the first predicted water quality characteristic data of the secondary watershed and the flow characteristics of the secondary watershed according to a water quality characteristic transformation function to obtain a set of predicted water quality characteristic data, and accumulating the set of predicted water quality characteristic data to obtain second predicted water quality characteristic data of the secondary watershed; The second LSTM feature prediction model is used to fit the second predicted water quality characteristic data of the secondary basin, the flow characteristics of the secondary basin, the meteorological characteristics of the mainstream basin, and the surface socio-economic characteristics to obtain the predicted water quality characteristic data of the mainstream basin.
8. The step-by-step "land-river-lake" water quality prediction model according to claim 7 is characterized in that: Also includes: SHAP model: used to obtain relevant features and water quality characteristic data involved in the prediction of water quality characteristics in the mainstream domain, calculate the marginal contribution of each relevant feature to the water quality characteristic data, and sort them according to the size of the marginal contribution value.
9. The step-by-step "land-river-lake" water quality prediction model according to claim 8 is characterized in that: Also includes: Data preprocessing module: used for data cleaning, processing, completion and standardization of socio-economic and meteorological characteristics of secondary watersheds.
10. A computer device, characterized in that: include: at least one memory and at least one processor; The memory is used to store one or more programs; When the one or more programs are executed by the at least one processor, the at least one processor implements the steps of the step-by-step "land-river-lake" water quality prediction method as described in any one of claims 1 to 6.