Water quality index prediction method and system based on deep learning model interpretability framework
By introducing an interpretability framework into the deep learning model, using large language models and interpretive algorithms, the problem of complex decision-making process in water quality prediction is solved, and accurate and robust interpretation of water quality indicator prediction and pollution source monitoring are achieved, ensuring the stability of water quality.
Patent Information
- Application Number
- CN202510229229.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-20
AI Technical Summary
The existing deep learning models have complex and opaque decision-making processes in water quality prediction, making them difficult to interpret, making it difficult to judge whether model suggestions should be adopted, which may lead to property losses and safety hazards.
Adopting an interpretability framework based on deep learning models, explanatory tasks are disassembled through large language models, allocate weights, select and call the algorithm of the explanatory framework pool, obtain explanatory results, and conduct comprehensive analysis to provide explanatory reasons with high confidence.
Accurate and robust interpretation of water quality indicator predictions is achieved, pollution sources can be identified and monitored, and measures are taken in a timely manner to reduce nitrogen emissions and prevent water quality from deteriorating.
Smart Images

Figure CN120181606A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of water environment prediction, and more specifically, to a water quality index prediction method and system based on an interpretability framework of a deep learning model. Background Art
[0002] Water quality prediction methods can provide short-term and long-term water quality conditions and future change trends, provide a scientific basis for water pollution prevention and control, provide important guidance for public health, and provide technical support for water environment governance. Data-driven water quality prediction methods, although they may be superior to traditional mechanism models in terms of performance, with the fusion of multi-modal data, the increase in the number of model layers, and the improvement of the model's non-linear expression ability, the decision-making process of the model becomes increasingly complex and opaque. This complexity makes it difficult to interpret the decision-making logic of the model, and thus difficult to determine whether the model's suggestions should be adopted. If the decision of the model is incorrect, it may lead to serious property losses and safety hazards. Conducting safety assessments on deep learning models, improving the interpretability of the models, and ensuring the transparency and understandability of their decision-making processes are crucial for the practical application of the models and for safeguarding people's property and safety.
[0003] However, existing interpretability frameworks or methods simply use one or more different interpretability methods to obtain different interpretation data, and different interpretability methods often give different interpretations due to different algorithm principles, and some are even contradictory, which is likely to mislead decision-makers and thus lead to incorrect decisions.
[0004] Therefore, how to provide a water quality index prediction method and system based on an interpretability framework of a deep learning model, screen out highly confident interpretability results, provide accurate and robust interpretability reasons, accurately predict the total nitrogen concentration, identify and monitor pollution sources, take timely measures to reduce nitrogen emissions to prevent water quality deterioration is an urgent problem to be solved by those skilled in the art. Summary of the Invention
[0005] In view of this, the present invention provides a water quality index prediction method and system based on an interpretability framework of a deep learning model. Based on a deep learning algorithm and combined with a large language model, it screens out highly confident interpretability results, provides accurate and robust interpretability reasons, accurately predicts the total nitrogen concentration, can identify and monitor pollution sources, and take timely measures to reduce nitrogen emissions to prevent water quality deterioration.
[0006] To achieve the above object, the present invention adopts the following technical scheme: A water quality index prediction method based on an interpretability framework of a deep learning model, comprising:
[0007] Collect a dataset of water quality parameters and geographical parameters, preprocess the dataset to obtain a training set and a test set;
[0008] Construct a water quality prediction model based on deep learning algorithms, and train the water quality prediction model with the training set;
[0009] Input the interpretability task into the large language model; the large language model disassembles the interpretability task and pre-assigns a weight to the result of each interpretability algorithm;
[0010] Select and call the algorithm in the interpretability framework pool according to the analysis, and access the water quality prediction model to obtain the interpretability result;
[0011] Return the interpretability result to the large language model for comprehensive analysis;
[0012] If the large language model deems the interpretability result unreasonable, ask the decision maker again whether to call other algorithms for re-interpretability analysis.
[0013] Preferably, it further includes: performing real-time interpretability monitoring on the prediction result of the water quality prediction model and providing report information.
[0014] Preferably, before giving the final interpretability result, make two judgments:
[0015] If there are no repeated important variables identified by the interpretability algorithm, ask the decision maker to call a third interpretability algorithm to strengthen the interpretation and repeat the process;
[0016] If there are repeated important variables identified by the interpretability algorithm, calculate its confidence score, combine the initial weight with the normalized importance score to obtain the final score, calculate the interpretability information entropy, and compare the interpretability information entropy with the threshold. If it does not exceed the threshold, return the final interpretability result; otherwise, call a third interpretability algorithm to repeat the process.
[0017] Preferably, returning the interpretability result to the large language model for comprehensive analysis includes:
[0018] Submit the interpretability result to the confidence score module of the large language model, and define the confidence score as the number of times the variable is identified by the interpretability algorithm and ranked in the top m;
[0019] Introduce interpretability information entropy to quantify interpretability. When the interpretability information entropy exceeds the threshold, it can be interpreted, and the monitoring stations and internal water quality variables affecting the water quality index prediction can be identified.
[0020] Preferably, the deep learning algorithm uses a graph convolutional neural network, which updates the representation of each node in the next layer by aggregating the features of neighboring nodes and iterates multiple times over the entire graph, enabling the representation of each node to capture information from more distant nodes.
[0021] Preferably, the GNNExplainer algorithm is used to interpret the prediction results of the graph convolutional neural network to obtain the node features and edges that affect the prediction of specific nodes.
[0022] The GraphLIME algorithm is used to interpret the prediction results of the graph convolutional neural network. The GraphLIME algorithm is based on a method of local non-linear feature selection to learn a non-linear interpretable surrogate model from the subgraph and obtain the features that affect the prediction of specific nodes.
[0023] Preferably, the GNNExplainer algorithm based on the perturbation strategy can determine a key subgraph structure and a small subset of node features with key roles for the prediction task; for the trained water quality prediction model, the GNNExplainer algorithm can find the subset of other monitoring stations that affect the total nitrogen concentration prediction of the current monitoring station.
[0024] Preferably, the GraphLIME algorithm based on local non-linear feature selection learns a non-linear interpretable surrogate model from the subgraph.
[0025] Preferably, a water quality index prediction system based on the interpretability framework of the deep learning model includes:
[0026] A data collection and preprocessing module for collecting a data set of water quality parameters and geographical parameters, preprocessing the data set to obtain a training set and a test set;
[0027] A model training module for constructing a water quality prediction model based on a deep learning algorithm and training the water quality prediction model through the training set;
[0028] A weight assignment module for inputting an interpretability task into a large language model; the large language model disassembles the interpretability task and pre-assigns a weight to the result of each interpretability algorithm;
[0029] An algorithm call module for calling the algorithms in the interpretability framework pool according to the analysis and accessing the water quality prediction model to obtain interpretability results;
[0030] An analysis module for returning the interpretability results to the large language model for comprehensive analysis;
[0031] A judgment module for, if the large language model deems the interpretability results unreasonable, asking the decision maker again whether to call other algorithms for re-interpretability analysis.
[0032] As can be seen from the above technical solutions, compared with the prior art, the present invention discloses a water quality index prediction method based on an interpretability framework of a deep learning model, including: collecting a data set of water quality parameters and geographical parameters, preprocessing the data set to obtain a training set and a test set; constructing a water quality prediction model based on a deep learning algorithm, and training the water quality prediction model through the training set; inputting an interpretability task into a large language model; the large language model disassembles the interpretability task and pre-assigns a weight to the result of each interpretability algorithm; selecting and calling an algorithm from an interpretability framework pool according to the analysis, and accessing the water quality prediction model to obtain an interpretability result; returning the interpretability result to the large language model for comprehensive analysis; if the large language model deems the interpretability result unreasonable, it will ask the decision maker again whether to call other algorithms for re-interpretability analysis. The present invention can identify and monitor pollution sources, take timely measures to reduce nitrogen emissions to prevent water quality deterioration. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the provided drawings.
[0034] Figure 1 It is a schematic flowchart of an interpretability framework based on a large language model provided by an embodiment of the present invention.
[0035] Figure 2 It is a schematic diagram of an interpretability framework based on algorithm principles and classification of interpretability types provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0036] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0037] An embodiment of the present invention discloses a water quality index prediction method based on an interpretability framework of a deep learning model, including:
[0038] Collecting a data set of water quality parameters and geographical parameters, preprocessing the data set to obtain a training set and a test set; obtaining data and preprocessing the data, including steps such as filling missing values and removing noise.
[0039] Build a water quality prediction model based on deep learning algorithms, and train the water quality prediction model with the training set; train a water quality prediction model based on target feature data.
[0040] Input the interpretability task into the large language model; the large language model analyzes and organizes the semantic information of the interpretability task, disassembles the interpretability task, and invokes the interpretability algorithm, and assigns a weight to the result of each interpretability algorithm in advance; among them, the decision maker inputs a query statement or an interpretability task to the large language model, the large language model analyzes and organizes the semantic information, disassembles the interpretability task, decides the invocation of the interpretability method, and assigns a weight to the result of each interpretability method in advance.
[0041] Select and invoke the algorithm of the interpretability framework pool according to the analysis, and access the water quality prediction model to obtain the interpretability result;
[0042] Return the interpretability result to the large language model for comprehensive analysis; the large language model returns graphic and text information to the decision maker according to the assigned weight and the rationality and confidence score of the interpretability result.
[0043] If the large language model believes that the interpretability result is unreasonable or cannot be understood by humans, it can ask the decision maker again whether to invoke other algorithms for re-interpretability analysis.
[0044] Through the above steps, comprehensively adopt the interpretability results of multiple algorithms, which not only enables the algorithms to verify each other but also can analyze the data from different perspectives, can identify the importance of topological structures or features to the prediction results, enhance the transparency of the deep model decision-making and improve the reliability and comprehensiveness of the interpretability results. Through this framework, it can provide strong decision-making support for managers and help formulate more effective water environment management strategies.
[0045] Specifically, it further includes: performing real-time interpretability monitoring on the prediction results of the water quality prediction model and providing report information.
[0046] In a specific embodiment of the present invention, an interpretability framework for deep learning of water environment based on a large language model, as Figure 1 shown, includes a large language model analysis module, an interpretability framework, a deep learning model to be accessed, a confidence score module, a graphic and text visualization module, a real-time monitoring module, etc.
[0047] Among them, the interpretive framework is applied to a water quality prediction model based on a graph convolutional neural network, and the predicted index focuses on the total nitrogen concentration. This is because total nitrogen is an important indicator of water eutrophication. Excessive total nitrogen concentration will lead to overgrowth of algae, which will in turn cause algal blooms and affect the health of the water ecosystem. By accurately predicting the total nitrogen concentration, pollution sources can be identified and monitored, and measures can be taken in a timely manner to reduce nitrogen emissions to prevent water quality deterioration. Therefore, the embodiments of the present invention select total nitrogen as the target feature, but this embodiment is not limited to the application of interpretation in the total nitrogen variable, nor is it limited to interpreting only deep learning models for predicting river network water quality.
[0048] Specifically, before giving the final interpretive result, two judgments are made to provide an interpretive result with high confidence and interpretable by humans;
[0049] If there are no repeated important variables identified by the interpretive algorithm, it indicates that the interpretive results vary greatly. The decision maker is asked to call the third interpretive algorithm to strengthen the interpretation and repeat the process;
[0050] If there are repeated important variables identified by the interpretive algorithm, for example, if both algorithms identify total phosphorus as one of the top m features, then calculate its confidence score, combine the initial weight with the normalized importance score to obtain the final score, calculate the interpretive information entropy, and compare the interpretive information entropy with the threshold. If it does not exceed the threshold, return the final interpretive result; otherwise, call the third interpretive algorithm to repeat the process.
[0051] Specifically, the interpretive result is returned to the large language model for comprehensive analysis, including:
[0052] The interpretive result is the node label N of the subgraph identified by the interpretive framework h (monitoring station number) and the dth feature in the hth node (the dth water quality variable at the hth monitoring station);
[0053] The interpretive result is handed over to the confidence scoring module of the large language model, and the confidence score is defined as the number of times this variable is identified by the interpretive algorithm and ranked among the top m (m = 5);
[0054] The interpretive information entropy is introduced to quantify interpretability. When the interpretive information entropy exceeds the threshold (such as -0.4), it can be interpreted, and the monitoring stations and internal water quality variables that affect the prediction of water quality indicators (such as total nitrogen concentration) can be identified.
[0055] Specifically, the deep learning algorithm uses a graph convolutional neural network. This type of graph convolutional neural network is a model with the ability to extract spatial structures. It updates the representation of each node in the next layer by aggregating the features of neighboring nodes and iterates multiple times over the entire graph, enabling the representation of each node to capture information from more distant nodes.
[0056] Specifically, the GNNExplainer algorithm is used to interpret the prediction results of the graph convolutional neural network, obtaining the node features and edges that affect the prediction of specific nodes.
[0057] The GraphLIME algorithm is used to interpret the prediction results of the graph convolutional neural network. The GraphLIME algorithm is based on the method of local non-linear feature selection, learning a non-linear interpretable surrogate model from the subgraph to obtain the features that affect the prediction of specific nodes.
[0058] Specifically, the GNNExplainer algorithm based on the perturbation strategy can determine a key subgraph structure and a small subset of node features with key roles for the prediction task. It aims to maximize the mutual information between the prediction of the graph convolutional neural network and the distribution of possible subgraph structures. In this embodiment, for the trained water quality prediction model, the GNNExplainer algorithm can find the subset of other monitoring stations that affect the total nitrogen concentration prediction of the current monitoring station.
[0059] GNNExplainer identifies other water quality variables that affect the total nitrogen concentration prediction by learning a binary node feature filter. By marginalizing all feature subsets, and then learning this binary node feature filter through Monte Carlo sampling and gradient backpropagation. The features with a filter value of 1 are selected through this filter and their feature weights are obtained, recording this numerical feature and normalizing it to the interval [0, 1].
[0060] Specifically, the GraphLIME algorithm based on local non-linear feature selection learns a non-linear interpretable surrogate model from the subgraph. Since the graph convolutional neural network model is usually non-linear, and the deeper the graph convolutional neural network model performs better, using a linear-based surrogate model is not sufficient to approximate the non-linear representation of the graph convolutional neural network. Therefore, the non-linear interpretable feature selection algorithm HSIC LASSO based on the kernel method is adopted.
[0061] In a specific embodiment of the present invention, a water quality index prediction system based on the interpretability framework of the deep learning model includes:
[0062] A data collection and preprocessing module for collecting a dataset of water quality parameters and geographical parameters, preprocessing the dataset to obtain a training set and a test set;
[0063] A model training module for constructing a water quality prediction model based on a deep learning algorithm and training the water quality prediction model with the training set;
[0064] A weight allocation module for inputting an interpretive task into a large language model; the large language model disassembles the interpretive task and pre-assigns a weight to the result of each interpretive algorithm.
[0065] An algorithm call module for selecting and calling an algorithm from an interpretive framework pool according to the analysis and accessing a water quality prediction model to obtain an interpretive result.
[0066] An analysis module for returning the interpretive result to the large language model for comprehensive analysis.
[0067] A judgment module for, if the large language model deems the interpretive result unreasonable, asking the decision maker again whether to call other algorithms for re-interpretive analysis.
[0068] In a specific embodiment of the present invention, a water quality index prediction method based on an interpretability framework of a deep learning model includes:
[0069] Step 1: Obtain an input data table from a pre-constructed deep learning module, where the input data table includes information such as river network topology, monitored water quality variables of monitoring stations, flow velocity of river channels, cross-sectional area, etc. This example takes the Yangtze River Basin as an example.
[0070] Among them, the river network topology is obtained by obtaining the longitude and latitude coordinates of each station and constructing an adjacency matrix, and calculating the distance between stations through the Euclidean distance formula. Information such as the monitored water quality variables of monitoring stations, flow velocity of river channels, cross-sectional area, etc. are self-collected data, and the time resolution of the data is monthly. The specific water quality parameters are shown in Table 1, with the unit of mg / L. The collection time is from January 2010 to December 2023. Before feature embedding, the interpolation function of pandas is used to process the missing values. The water quality variables are regarded as the features of the nodes, and the vector shape of each node is 23×168 dimensions. The flow velocity and cross-sectional area are regarded as the features of the edges, and their shapes are 2×168.
[0071] Table 1 Self-collected water quality parameters and geographical parameter types in the Yangtze River Basin
[0072]
[0073] Based on the above input data, a deep learning model for predicting river network water quality can be established. The available black-box models include deep learning models such as graph convolutional neural networks (GCN, GAT, GraphSage), recurrent neural networks (RNN, LSTM, GRU), and fully connected network MLP. One of these models can be selected, or two of them can be combined to achieve feature extraction in both space and time simultaneously. The model used in the embodiments of the present invention is a graph convolutional neural network. The graph convolutional neural network category is a type of model with the ability to extract spatial structures. It updates the representation of each node in the next layer by aggregating the features of neighboring nodes. The above process is iterated multiple times over the entire graph, enabling the representation of each node to capture information of nodes at greater distances. Then, 70% of the time series data in the data is used as the training set, and the remaining 30% is used as the test set. Three statistical metrics are used, namely the commonly used R 2 , RMSE, and Nash-Sutcliffe Efficiency coefficient (NSE). The structure and initialization parameters of the GCN model are shown in Table 2. By using the activation function relu and the Dropout layer, the learned features are non-linearly transformed through embedding, which helps the model capture more complex patterns and relationships between water quality variables and variables at other stations, while preventing overfitting of the model.
[0074] Table 2 Construction parameters of the GCN model
[0075] Model structure Default parameters The first layer of GCN (23 * 168, 128) Learning rate: 1e-4 The second layer of GCN (128, 1) Loss function: mean squared error loss nn.MSELoss() Activation function F.relu() Optimizer: Adam optimizer Dropout(p = 0.5) Number of training epochs: 200 epochs
[0076] Step 2: Edit the query language for "the most important other water quality variables and river network topological substructures affecting the total nitrogen fluctuation" based on the data table, including the importance explanation of the internal features of the obtained nodes, the importance of the network topology structure, and model explanations based on surrogate models, etc. Input in the form of speech or text to the front end of the large language model with the Transformer as the basic architecture. The language model disassembles the task and generates the code for calling the GNNExplainer algorithm embedded in PyG for the importance explanation of the internal features of the nodes and the importance of the network topology structure:
[0077] explainer = GNNExplainer(model, epochs = 200)
[0078] node_feat_mask, edge_mask = explainer.explain_node(node_idx, data.x, data.edge_index)
[0079] Among them, GNNExplainer is a class for explaining the prediction results of graph convolutional neural networks.
[0080] The model is a pre-trained graph convolutional neural network model, and GNNExplainer will explain the predictions of this model.
[0081] epochs = 200 represents the number of training epochs for the explainer. That is, GNNExplainer will perform 200 rounds of training to learn which features and edges are important for the prediction results.
[0082] explainer.explain_node is a method in the GNNExplainer class used to explain the prediction results of a specific node.
[0083] node_idx is the index of the node to be explained, indicating which node's prediction we want to understand how it is obtained.
[0084] data.x is the feature matrix of all nodes in the graph, and each row represents the feature vector of a node.
[0085] data.edge_index is the edge index matrix of the graph, which defines the connection relationships between nodes in the graph.
[0086] node_feat_mask is a mask vector whose length is equal to the number of node features. Each element in the mask vector represents the importance of the corresponding feature for the prediction of the current node. Usually, the closer the value is to 1, the more important the feature is.
[0087] edge_mask is a mask vector whose length is equal to the number of edges in the graph. Each element in the mask vector represents the importance of the corresponding edge for the prediction of the current node.
[0088] For generating code that calls the GraphLIME algorithm in the graphlime package based on the surrogate model:
[0089] explainer = GraphLIME(model, hop = 2, rho = 0.1)
[0090] coefs = explainer.explain_node(node_idx, data.x, data.edge_index)
[0091] Among them, GraphLIME is a class in the graphlime package used to explain the prediction results of graph convolutional neural networks.
[0092] The model is a pre-trained graph convolutional neural network model, and GraphLIME will explain the predictions of this model.
[0093] hop = 2 represents the number of hops when selecting neighbor nodes, that is, neighbor nodes within at most 2 hops from the target node are considered to capture the locality information of the node to be explained.
[0094] rho = 0.1 is a regularization parameter in the GraphLIME algorithm, which is used to control the complexity of the surrogate model and avoid overfitting.
[0095] explainer.explain_node is a method in the GraphLIME class, which is used to explain the prediction result of a specific node.
[0096] node_idx is the index of the node to be explained.
[0097] data.x is the feature matrix of all nodes in the graph.
[0098] data.edge_index is the edge index matrix of the graph.
[0099] coefs is a coefficient vector, which represents the importance of each feature for the prediction of the current node. The larger the absolute value of the coefficient, the greater the impact of the feature on the prediction result.
[0100] Specifically, assign weights [1, 0.5, 0.5] and return them to the decision maker for modification and review.
[0101] It should be noted that there are various types of explanations involved here. For the explanation of the topological structure, only GNNExplainer can give it, so the weight given to this result is 1. For the explanation of the importance of node internal features, both GNNExplainer and GraphLIME can provide explanations, so the results given each account for a weight of 0.5.
[0102] Specifically, GNNExplainer based on the perturbation strategy can determine a key subgraph structure and a small subset of node features with key roles for the prediction task. It aims to maximize the mutual information between the prediction of the graph convolutional neural network and the distribution of possible subgraph structures. In this embodiment, for the trained water quality prediction model, GNNExplainer attempts to find other subsets of monitoring stations that have a significant impact on the total nitrogen concentration prediction of the current monitoring station, obtained by optimizing formula (1):
[0103]
[0104] Among them, the H function represents the information entropy, G s represents the subgraph, and X s is a subset of d-dimensional node features. The MI function measures the change in the impact on the total nitrogen concentration prediction under the full graph and under the computational graph with only the node subset.
[0105] Furthermore, GNNExplainer identifies other water quality variables that affect the total nitrogen concentration prediction by learning a binary node feature filter. By marginalizing all feature subsets and then learning the binary node feature filter through Monte Carlo sampling and gradient backpropagation. The features with a filter value of 1 are selected through this filter, and their feature weights are obtained. The numerical features are recorded and normalized to the interval [0, 1].
[0106] Specifically, the interpretive model GraphLIME based on local non-linear feature selection aims to learn a non-linear interpretable surrogate model from subgraphs. Since graph convolutional neural network models are usually non-linear and the deeper the graph convolutional neural network model performs better, using a linear-based surrogate model is not sufficient to approximate the non-linear representation of the graph convolutional neural network. Therefore, the non-linear interpretable feature selection algorithm HSIC LASSO based on the kernel method is adopted. The learning objective is shown in Equation (2):
[0107]
[0108] where f represents the deep learning model to be explained, X n represents the sampling information matrix that can capture the locality of the node to be explained, g is the interpretable model to be learned, and S(v) is the set of features of the node v to be explained. In addition to considering the node itself, GraphLIME uses an N-hop sampling strategy to select the set of neighbor nodes that affect the current node, selects the Gaussian kernel function as the input and the prediction of X n from the given GNN model. The interpretation of the HSIC Lasso model is obtained through Equation (3)
[0109]
[0110] where f k is the eigenvector corresponding to the k-th feature, tr() is the trace operator, K (K) is the normalized central Gram matrix of the k-th feature, and L is the normalized central Gram matrix. Through the above method, the larger the NHSIC value, the stronger the correlation between variables. Similarly, the numerical features are recorded and normalized to the interval [0, 1].
[0111] It should be noted that, to save the energy consumption of model training, for the first model interpretability analysis, the framework will only call two or three interpretability methods. If the returned interpretability results are generally the same, the result will be used as the final interpretability result. Otherwise, when the interpretability results are inconsistent or cannot be interpreted by humans, the algorithm selects the third or fourth interpretability algorithm from the interpretability framework pool. When traversing to the 10th model, if no valuable interpretation can be obtained, the model output is "uninterpretable". As Figure 2 shown.
[0112] Steps three and four: Through the two interpretability results returned in step two, in this example, the node label N h (monitoring station number) of the sub-graph identified by the interpretability framework and the d-th feature (the d-th water quality variable of the h-th monitoring station) at the h-th node are handed over to the confidence scoring module of the large language model. The confidence score is defined as the number of times this variable is identified by the interpretability algorithm and ranked in the top 5 (m = 5 in this example). And by introducing interpretability information entropy to quantify the interpretability that can be interpreted by humans, the embodiment of the present invention defines the interpretability information entropy as formula (4):
[0113]
[0114] where where f i represents the value of the important feature obtained in step two. Further, similar to the self-information in information theory, the negative logarithm of p k can be defined as the self-interpretability penalty of this feature, and formula (4) can be written as formula (5):
[0115]
[0116] Given a threshold Threshold (which is -0.4 in this example), when the interpretability information entropy exceeds this value, it indicates that there are significant differences between the feature importance variables and can be interpreted. Through the above steps, the important monitoring stations and their internal water quality variables that affect the prediction of the total nitrogen concentration of the current monitoring station are identified.
[0117] Step five: Before giving the final interpretability result, the program will make two judgments to provide an interpretability result with high confidence and can be interpreted by humans. The following is described in two cases:
[0118] The first case: When the important variables identified by the interpretability algorithm do not repeat. In this example, when the features identified by the two algorithms do not repeat, that is, the 10 variables only appear once, indicating that the interpretability results given by the two algorithms are quite different, then the program will ask the decision maker to apply to call the third interpretability algorithm to strengthen the interpretability, and then repeat the above process.
[0119] The second case: When there are duplicate importance variables identified by the interpretive algorithm, in this example, both algorithms identify the total phosphorus concentration as one of the top 5 importance features, then the confidence score of total phosphorus is 2. The program will then enter the judgment of interpretive information entropy, multiply the initial weight by the normalized total phosphorus importance score to obtain the final total phosphorus importance score. Calculate the importance scores of other features in this logic in turn. Finally, by calculating the interpretive information entropy, when it exceeds the given threshold Threshold = 0.4, it is considered that the feature has significant interpretability, and the final interpretive result is returned. Otherwise, if it does not exceed, the program will ask the decision maker again to apply to call the third interpretive algorithm to strengthen the interpretability, and then repeat the above process.
[0120] Step 6: To prevent the deep learning model from making predictions relying on incorrect information during operation, it can be randomly checked on a weekly basis and recorded in the form of a log for the decision maker to view.
[0121] In a specific embodiment of the present invention, the overall algorithm logic is as follows:
[0122] Input:
[0123] Query statement q: The statement input by the user for requesting an explanation of the deep learning model, such as "the most important other water quality variables and river network topological substructures affecting the total nitrogen fluctuation".
[0124] Large language model LLM: A model with natural language processing and logical reasoning capabilities, responsible for disassembling the query statement, calling appropriate interpretive algorithms, and comprehensively analyzing and processing the interpretive results.
[0125] Set of interpretive algorithms F: Includes a variety of different interpretive algorithms, such as the aforementioned GNNExplainer and GraphLIME, etc., for explaining the prediction results of the deep learning model.
[0126] Deep learning model ξ: The target deep learning model to be explained, such as the graph convolutional neural network model for river network water quality prediction.
[0127] Stop condition control parameter T: Its usage method is not directly reflected in the code, but it may be a parameter that controls the number of algorithm loops or other termination conditions.
[0128] Define the top m important features: Specify the top m most important features to focus on in the interpretive result, for screening and evaluating the interpretive result.
[0129] Define the interpretive information entropy threshold Threshold: It is used to judge the interpretability degree of the interpretation result. When the interpretive information entropy exceeds this threshold, it is considered that there are significant differences among the feature importance variables and they can be interpreted.
[0130] Process:
[0131] 1: while i <= 10 do
[0132] 2: fi, Winit = LLM(q), i = 2 or 3
[0133] 3: explain = [[]] * i
[0134] 4: for fi in F:
[0135] 5: explanation = fi(ξ)
[0136] 6: Nm = choice(explanation, m)
[0137] 7: explain[i].append(Nm)
[0138] 8:
[0139] 9: i = i + 1
[0140] 10: continue
[0141] 11: else:
[0142] 12: confidence = count(Nm)
[0143] 13: Nm *= ΣWinitNm
[0144] 14: S = IE(Nm*)
[0145] 15: if S > Threshold:
[0146] 16: i = i + 1
[0147] 17: continue
[0148] 18: else:
[0149] 19: return confidence and the interpretive result LLM(confidence, Nm*)
[0150] Specifically, the explanation of the above algorithm is as follows:
[0151] 1. Outer loop (line 1)
[0152] plaintext
[0153] while i <= 10 do
[0154] This is a loop structure. i is the loop counter. Starting from the initial value, as long as i is less than or equal to 10, the operations in the loop body will continue to be executed. The purpose of this loop is to stop trying after multiple attempts with different combinations of interpretability algorithms if a satisfactory interpretation result cannot be obtained.
[0155] 2. Large Language Model Processing (Line 2)
[0156] plaintext
[0157] fi, Winit = LLM(q), i = 2 or 3
[0158] After receiving the query statement q, the large language model LLM processes it and outputs two parts of results:
[0159] fi: Represents the selected i interpretability algorithms. The value of i is 2 or 3, that is, usually 2 to 3 algorithms are selected for the initial model interpretability analysis.
[0160] Winit: The initial weights pre-allocated for the interpretation results of these algorithms.
[0161] 3. Initialize the Explanation Result List (Line 3)
[0162] plaintext
[0163] explain = [[]] * i
[0164] Create a list explain with a length of i, where each element is an empty list, used to store the interpretation results of each interpretability algorithm.
[0165] 4. Traverse the Interpretability Algorithms (Lines 4 - 7)
[0166] plaintext
[0167] for fi in F:
[0168] explanation = fi(ξ)
[0169] Nm = choice(explanation, m)
[0170] explain[i].append(Nm)
[0171] for fi in F:: Traverse the i interpretability algorithms selected by the large language model.
[0172] explanation = fi(ξ): Call the current interpretability algorithm fi to interpret the deep learning model ξ, and obtain the interpretation result explanation.
[0173] Nm = choice(explanation, m): Select the top m important features from the interpretation result explanation and store them in Nm.
[0174] explain[i].append(Nm): Add the selected top m important features to the corresponding position of the algorithm in the explain list.
[0175] 5. Check the repeatability of the interpretation results (lines 8 - 10)
[0176] plaintext
[0177]
[0178] i = i + 1
[0179] continue
[0180] Determine whether the sets of the top m important features obtained by different interpretability algorithms have no intersection (i.e., no duplicate features). If there is no intersection, it means that the interpretation results given by these algorithms are quite different.
[0181] i = i + 1: Increase the value of the loop counter i and try to select more interpretability algorithms.
[0182] continue: Skip the subsequent operations of this loop and directly enter the next loop.
[0183] 6. Calculate the confidence and comprehensive importance score (lines 12 - 13)
[0184] plaintext
[0185] confidence = count(Nm)
[0186] Nm *= ΣWinitNm
[0187] confidence = count(Nm): Calculate the number of times each feature is recognized as one of the top m important features as the confidence score of this feature.
[0188] Nm *= ΣWinitNm: Multiply the importance score Nm of each normalized feature by the initial weight Winit, and then sum to obtain the comprehensive importance score Nm*.
[0189] 7. Calculate the explanatory information entropy and make a judgment (lines 14 - 18)
[0190] plaintext
[0191] explanation
[0192] S = IE(Nm*)
[0193] if S > Threshold:
[0194] i = i + 1
[0195] continue
[0196] else:
[0197] return confidence and interpretability result LLM(confidence, Nm*)
[0198] S = IE(Nm*): Calculate the explanatory information entropy S of the comprehensive importance score Nm*.
[0199] if S > Threshold:: If the explanatory information entropy S exceeds the threshold Threshold, it indicates that the differences between the feature importance variables are not significant and the interpretability is insufficient.
[0200] i = i + 1: Increase the value of the loop counter i and try to select more interpretability algorithms.
[0201] continue: Skip the subsequent operations of this loop and directly enter the next loop.
[0202] else:: If the explanatory information entropy S does not exceed the threshold Threshold, it indicates that there are significant differences between the feature importance variables and it can be interpreted.
[0203] return confidence and interpretability result LLM(confidence, Nm*): Return the final confidence score and interpretability result, which are output by the large language model LLM.
[0204] In summary, the core idea of this algorithm is to interpret the prediction results of the deep learning model by repeatedly trying different combinations of interpretability algorithms, and judge whether a satisfactory interpretability result is obtained based on the repeatability of the interpretability results and the explanatory information entropy. If a satisfactory result is not obtained, continue to try more algorithms until the maximum number of attempts (i reaches 10).
[0205] In the present specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments, and the same or similar parts among the embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description in the method section.
[0206] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.
Claims
1. A water quality index prediction method based on the deep learning model interpretability framework, characterized in that: include: Collecting data sets of water quality parameters and geographical parameters, and preprocessing the data sets to obtain training sets and test sets; Constructing a water quality prediction model based on a deep learning algorithm, and training the water quality prediction model using the training set; Inputting the explanation task into the large language model; the large language model decomposes the explanation task and pre-assigns a weight to the result of each explanation algorithm; The algorithm of the interpretive framework pool is called according to the analysis selection, and the water quality prediction model is accessed to obtain the interpretive results; Return the interpretation results to the large language model for comprehensive analysis; If the large language model considers the explanatory results to be unreasonable, the decision maker is asked again whether to call other algorithms for further explanatory analysis.
2. The method for predicting water quality indicators based on the deep learning model interpretability framework according to claim 1 is characterized in that: Also includes: Provide real-time interpretative monitoring of water quality prediction model prediction results and report information.
3. The water quality index prediction method based on the deep learning model interpretability framework according to claim 1 is characterized in that: Before giving the final explanatory results, two judgments are made: If there is no duplication of the important variables identified by the explanatory algorithm, ask the decision maker to call a third explanatory algorithm to strengthen the explanation and repeat the process; If the importance variables identified by the explanatory algorithm are repeated, their confidence scores are calculated, and the final scores are obtained by combining the initial weights with the normalized importance scores. The explanatory information entropy is calculated and compared with the threshold. If it does not exceed the threshold, the final explanatory result is returned, otherwise the third explanatory algorithm is called to repeat the process.
4. The method for predicting water quality indicators based on the deep learning model interpretability framework according to claim 1, characterized in that: The explanatory results are returned to the large language model for comprehensive analysis, including: Submit the explanatory results to the confidence scoring module of the large language model, and define the confidence score as the number of times the variable is identified by the explanatory algorithm and ranked in the top m; The explained information entropy is introduced to quantify the interpretability. When the explained information entropy exceeds the threshold, it can be explained and the monitoring stations and internal water quality variables that affect the prediction of water quality indicators can be identified.
5. The method for predicting water quality indicators based on the deep learning model interpretability framework according to claim 1, characterized in that: The deep learning algorithm adopts a graph convolutional neural network, which updates the next layer representation of each node by aggregating the features of neighboring nodes and iterates multiple times on the entire graph, so that the representation of each node can capture node information at a longer distance.
6. The method for predicting water quality indicators based on the deep learning model interpretability framework according to claim 5 is characterized in that: Use the GNNExplainer algorithm to interpret the prediction results of the graph convolutional neural network to obtain the node features and edges that affect the prediction of specific nodes; The GraphLIME algorithm is used to explain the prediction results of the graph convolutional neural network. The GraphLIME algorithm is based on the method of local nonlinear feature selection, learns a nonlinear interpretable proxy model from the subgraph, and obtains the features that affect the prediction of specific nodes.
7. The method for predicting water quality indicators based on the deep learning model interpretability framework according to claim 6 is characterized in that: The GNNExplainer algorithm based on the perturbation strategy can determine a subgraph structure and a small subset of node features for the prediction task; for the trained water quality prediction model, the GNNExplainer algorithm can find other subsets of monitoring stations that affect the total nitrogen concentration prediction of the current monitoring station.
8. The method for predicting water quality indicators based on the deep learning model interpretability framework according to claim 1, characterized in that: The GraphLIME algorithm based on local nonlinear feature selection learns a nonlinear interpretable proxy model from a subgraph.
9. A water quality index prediction system based on a deep learning model interpretability framework, characterized in that: include: A data collection and preprocessing module is used to collect data sets of water quality parameters and geographical parameters, and preprocess the data sets to obtain training sets and test sets; A model training module, used to construct a water quality prediction model based on a deep learning algorithm, and train the water quality prediction model through the training set; A weight assignment module is used to input the explanatory task into the large language model; the large language model decomposes the explanatory task and pre-assigns a weight to the result of each explanatory algorithm; An algorithm calling module is used to call the algorithm of the explanatory framework pool according to the analysis selection and access the water quality prediction model to obtain the explanatory results; The analysis module is used to return the explanatory results to the large language model for comprehensive analysis; The judgment module is used to ask the decision maker whether to call other algorithms for further interpretation analysis if the large language model considers the interpretation result to be unreasonable.