Dissolved oxygen prediction method fusing ChatGPT field expert knowledge and deep learning
By introducing ChatGPT-EK-TabNet into the dissolved oxygen prediction model, and combining domain expert knowledge to perform feature normalization and weighted feature constraint matrix construction, the problem of insufficient prediction accuracy caused by the lack of domain knowledge in the existing methods is solved, and higher prediction accuracy and stability are achieved.
Patent Information
- Application Number
- CN202510282077.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-27
AI Technical Summary
The existing dissolved oxygen prediction methods lack domain expert knowledge, which leads to insufficient feature selection and variable relationship mining, affecting the prediction accuracy and reliability.
The dissolved oxygen prediction method based on ChatGPT-EK-TabNet is adopted to acquire expert knowledge in the field of water quality, improve feature normalization treatment and build a weighted feature constraint matrix, and integrate it into the attention mechanism of the TabNet model to enhance feature expression ability and prediction accuracy.
It significantly improves the performance of the dissolved oxygen prediction model, improves the prediction accuracy and stability, enhances the accuracy of feature selection and relationship modeling, and is suitable for situations where data is scarce or variable complexity is high.
Smart Images

Figure CN120217151A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for predicting dissolved oxygen in water environment, belonging to the cross - field of water environment science and artificial intelligence, and specifically relates to a dissolved oxygen prediction method that integrates ChatGPT domain expert knowledge and deep learning. Background Art
[0002] Dissolved oxygen (DO) is a key factor in maintaining the health of the water environment and the balance of the aquatic ecosystem. Accurate prediction of dissolved oxygen has important research significance for the healthy management of the aquatic ecosystem. Dissolved oxygen is an indispensable condition for the survival of aquatic organisms and is crucial for maintaining their respiration. On the one hand, when the dissolved oxygen in the water body is lower than a certain value, the growth and substrate metabolism of aquatic organisms will be affected. At this time, the dissolved oxygen concentration is called the critical dissolved oxygen concentration. When the dissolved oxygen continues to decrease, it will lead to the asphyxiation death of aquatic organisms. On the other hand, when the dissolved oxygen in the water body is too high, it may cause metabolic imbalance or other physiological problems in aquatic organisms, especially greatly increasing the probability of causing "gas bubble disease" in fish, and then leading to fish death. Excessive dissolved oxygen will accelerate the oxidation and decomposition rate of organic matter, thus accelerating the eutrophication process of some water bodies. Therefore, accurate prediction of dissolved oxygen in the water environment is of great significance for analyzing the healthy state of the water environment.
[0003] Most of the current dissolved oxygen prediction models are based on statistical algorithms, machine learning, and deep learning, etc. In the early research of dissolved oxygen prediction, many methods relied on empirical statistical models. These models usually collected a series of water quality parameters (such as water temperature, pH value, ammonia nitrogen concentration, etc.) and used classical statistical algorithms such as linear regression and stepwise regression for regression analysis to predict the concentration of dissolved oxygen. These algorithms have the advantages of low computational complexity and are more suitable for small - scale data sets. However, when dealing with non - linear water quality data, these models often show insufficient stability and limited prediction accuracy.
[0004] With the rapid development of computer technology, relevant researchers have introduced machine learning algorithms into the prediction of dissolved oxygen. Compared with traditional statistical methods, support vector machine (SVM), random forest (RF), and ensemble learning can handle nonlinear and high-dimensional data more effectively. In addition, machine learning methods can extract important features under complex multi-parameter conditions, thereby increasing the accuracy of prediction. However, machine learning often relies too much on data quality and quantity. When the data is insufficient or there is noise, the prediction effect and generalization ability of the model will decrease significantly. With the further development of artificial intelligence, deep learning models have been introduced into the dissolved oxygen prediction task. In particular, models based on long short-term memory network (LSTM) and convolutional neural network (CNN) have been widely used in dissolved oxygen prediction and water environment field analysis. Deep learning models can automatically learn data features and show obvious advantages in dealing with complex nonlinear relationships. However, the high complexity and computational cost of deep learning are still challenges that need to be solved urgently. In contrast, TabNet is a deep learning-based model that selectively focuses on different features by introducing an attention mechanism. Compared with traditional deep learning models (such as fully connected neural networks), TabNet has significant advantages. Through its sparse attention mechanism, it can selectively focus on different features in each layer of the network, thus efficiently processing tabular data and time series data. Nevertheless, due to the lack of guidance from domain expert knowledge, TabNet still relies on the features of the data itself for selective attention. Therefore, it may not be as accurate as expert knowledge-driven methods in mining the relationships between variables, resulting in a decrease in the prediction accuracy of the model in some cases.
[0005] ChatGPT expert domain knowledge is a knowledge base with multiple expert domains obtained through pre-training large-scale models. ChatGPT can propose expert knowledge based on the experience and theories of domain experts, provide feature correlation analysis for variables and parameters in different professional fields, and then make more reasonable decisions in data processing, model analysis, and feature relationship calculation. However, the current dissolved oxygen prediction methods do not introduce expert domain knowledge, resulting in insufficient feature selection and variable relationship mining, and further causing the model to rely too much on data-driven, thus affecting the prediction accuracy and reliability. This knowledge enhancement method not only makes up for the randomness and blindness in the feature selection process of traditional deep learning models but also effectively addresses the model's dependence on large-scale data sets. By embedding domain expert knowledge into the attention mechanism of the model, the model can perform well in the face of data scarcity or high variable complexity, significantly improving the prediction accuracy and stability of the model. Its advantage lies in that it can not only improve the performance of the dissolved oxygen prediction model but also maintain high interpretability, making the prediction results easier to analyze and apply. Summary of the Invention
[0006] To solve the problem that the lack of domain expert knowledge in the existing technology for dissolved oxygen prediction leads to inaccurate model prediction, the present invention discloses a dissolved oxygen prediction method based on ChatGPT-EK-TabNet, which solves the problem that the traditional method ignores the combination of domain knowledge and prediction tasks, resulting in limited improvement in prediction accuracy. By using ChatGPT to obtain expert knowledge in the water quality field, improving the normalization process of the maximum and minimum values of features, constructing a feature constraint matrix based on water quality data, and combining the feature relationship weights generated by ChatGPT, a weighted feature constraint matrix is formed to improve the attention mechanism of the TabNet model, providing a new idea for the prediction of dissolved oxygen in the water environment.
[0007] A dissolved oxygen prediction method based on ChatGPT-EK-TabNet provided by the present invention includes the following five steps:
[0008] Step 1: Pretreatment of water quality data;
[0009] Collect various data related to water quality, including key parameters such as water temperature, pH value, conductivity, turbidity, permanganate index, ammonia nitrogen, total phosphorus, and total nitrogen. Check the data integrity, remove missing values or outliers to ensure the accuracy and consistency of the data and guarantee the data quality.
[0010] Step 2: Normalization method based on ChatGPT domain knowledge support;
[0011] According to the data collected in Step 1, obtain the maximum and minimum expert knowledge of ChatGPT water quality domain through the algorithm process for the normalization of each index. By traversing the water quality data features one by one, send the information of each feature to ChatGPT for query. ChatGPT returns the maximum and minimum values of this feature in the water quality domain, which are saved in JSON format. The information of these maximum and minimum values will be gradually summarized into a set, and finally a complete domain knowledge base is formed. Replace the maximum and minimum values in the maximum and minimum normalization method with the obtained maximum and minimum values in the water quality domain to map the values of each feature to between 0 and 1 to ensure the consistency of the data and the stability and effectiveness of model training.
[0012] Step 3: Construct a weighted feature constraint matrix based on ChatGPT;
[0013] Based on the data obtained in Step 1, use the interface to call ChatGPT to obtain the correlation weights between various variables in the water quality field, and then construct a weighted feature constraint matrix between variables in the water quality field. Specifically, the water quality field expert knowledge of ChatGPT is based on interacting with the model to analyze each pair of features in the water quality data, and the correlation weights between the features are returned through its expert knowledge base. These weight information will be stored in a gradually improved domain knowledge set in JSON format, and finally form a complete domain knowledge base. This knowledge base is used to construct a weighted feature constraint matrix, and based on this knowledge, the scaled dot product attention mechanism is improved to ensure that the correlation between data features can be effectively utilized in the model;
[0014] Step 4: Integrate the improved scaled dot product attention mechanism with the weighted feature constraint matrix;
[0015] Based on the weighted feature constraint matrix constructed in Step 3, integrate the weighted feature constraint matrix with the scaled dot product attention mechanism. The present invention introduces a constraint mechanism enhanced by domain knowledge on the basis of the traditional attention mechanism, and constructs a scaled dot product attention mechanism integrating the feature constraint matrix. In the scaled dot product attention mechanism, calculate the dot product similarity score of the query vector and the key vector of the input feature. Multiply the calculated similarity score by the corresponding element in the weighted feature constraint matrix to obtain the weighted attention score. By integrating the weighted feature constraint matrix, the interaction of each pair of features is weighted and adjusted, so that in the attention mechanism, the influence of different features on the final result varies according to the relationship weight of the weighted features, and only makes up for the randomness and blindness in the feature selection process of the traditional deep learning model in a knowledge-enhanced way to ensure the accuracy of the model in feature interaction and make the correlation between features be fully reflected in the model;
[0016] Step 5: Construct a dissolved oxygen prediction model of ChatGPT-EK-TabNet;
[0017] This model consists of three main components. ChatGPT represents an artificial intelligence model used to obtain domain knowledge; EK represents the integration of expert knowledge (Expert Knowledge); TabNet represents a deep learning prediction model based on tabular data;
[0018] First, the expert knowledge obtained through Step 2 is used to normalize the input data, standardizing each eigenvalue to the same range to ensure the consistency of the data in the model and improve the stability of model training. Secondly, the weighted feature constraint matrix constructed in Step 3 is incorporated into the scaled dot-product attention mechanism to complete the weighted interaction and fusion between features, enabling the correlation between different features to be effectively reflected and adjusted in the attention mechanism of the model, and ensuring the reasonable distribution of feature weights. Finally, to further improve the accuracy of the model, in the improved scaled dot-product attention mechanism of the present invention, the Attentive Transformer module in the TabNet model is replaced. Although the original Attentive Transformer module can capture global information when dealing with the relationship between features, its ability to model complex interactions between variables is limited, and it lacks sufficient domain knowledge guidance, resulting in a decrease in accuracy in dissolved oxygen prediction. By introducing a new module, expert knowledge can be better utilized, feature selection and relationship modeling can be improved, thereby enhancing the overall prediction effect of the model. The dissolved oxygen is predicted through the constructed ChatGPT-EK-TabNet model, combined with the data normalized by expert domain knowledge.
[0019] Compared with the prior art, the advantages of the present invention are as follows:
[0020] (1) The present invention proposes a normalization method that introduces domain knowledge (domain maximum and domain minimum). Compared with ordinary normalization methods, the normalization method applicable to domain maximum and minimum has stronger dynamic adaptability and accuracy. Ordinary methods usually rely on static global maximum and minimum values, and are unable to effectively cope with the dynamic changes in streaming data or real-time data, and are easily affected by outliers or changes in data distribution. The normalization method based on domain knowledge can dynamically adjust the upper and lower bounds of normalization according to the characteristics of the data in the domain, making data processing more flexible and laying a data foundation for the subsequent construction of the model.
[0021] (2) The present invention proposes a weighted feature constraint matrix based on domain expert knowledge. By interacting with ChatGPT, the correlation weights between various features in the water quality domain are obtained, and these weight information are integrated into the feature constraint matrix. The feature constraint matrix is used to weight the water quality features, enabling the model to more accurately capture the interaction relationships between features. Compared with traditional feature processing methods, the present invention significantly improves the model's ability to represent complex relationships between features by incorporating domain expert knowledge, thereby enhancing the accuracy and stability of prediction.
[0022] (3) The present invention proposes an improved scaled dot - product attention mechanism. By introducing a weighted feature constraint matrix in the process of dot - product attention calculation, the similarity score between feature vectors is multiplied by the corresponding weight of the feature constraint matrix to dynamically adjust the importance of feature interaction. The improved scaled dot - product attention mechanism can effectively reduce redundant information and noise interference between features, ensure that the model pays attention to key features, and thus improve the overall performance.
[0023] (4) The present invention proposes a dissolved oxygen prediction method (ChatGPT - EK - TabNet) that integrates ChatGPT domain expert knowledge and deep learning. By combining domain expert knowledge for feature normalization processing and feature interaction weighting, it enhances the feature expression ability and prediction accuracy. The feature weighting matrix improves the scaled dot - product attention mechanism and replaces the original Attentive Transformer to more effectively capture and process complex interaction relationships between features. The model can still maintain high prediction performance and generalization ability in the case of scarce and highly heterogeneous data. Brief Description of the Drawings
[0024] Figure 1 is a schematic diagram of the prediction method process;
[0025] Figure 2 is a schematic diagram of obtaining the maximum and minimum values of expert knowledge;
[0026] Figure 3 is a schematic diagram of obtaining the relationship between domain expert knowledge variables;
[0027] Figure 4 is a structural diagram of the improved dot - product attention mechanism;
[0028] Figure 5 is a structural diagram of ChatGPT - EK - TabNet;
[0029] Figure 6 is a comparison chart of ablation experiment evaluation indicators;
[0030] Figure 7 is a comparison chart of the prediction result and the actual value of the prediction method of the present invention. Detailed Embodiment
[0031] The following will further elaborate on the present invention in conjunction with the drawings and embodiments.
[0032] The present invention proposes a dissolved oxygen prediction method (ChatGPT-EK-TabNet) that integrates the expert knowledge in the ChatGPT field and deep learning, and constructs a dissolved oxygen prediction model based on this model. First, the present invention proposes a normalization method that introduces domain knowledge (domain maximum value and domain minimum value) to enhance the robustness in data normalization, laying a foundation for better model construction. Secondly, in the process of constructing the model, the present invention proposes a weighted feature constraint matrix based on domain expert knowledge to improve the scaled dot-product attention mechanism, thereby ensuring the model's attention to key features. Finally, on this basis, the improved scaled dot-product attention mechanism is integrated into the TabNet model, replacing the original AttentiveTransformer, to more effectively capture and process the complex interaction relationships between features, providing new ideas for the river and lake dissolved oxygen prediction model.
[0033] As Figure 1 shown, it is a flowchart of the dissolved oxygen prediction method based on ChatGPT-EK-TabNet. The specific implementation manner will be described in detail with reference to this figure:
[0034] Step 1: Pretreatment of water quality data.
[0035] Collect the corresponding water quality data, including indicators such as water temperature, pH value, conductivity, turbidity, permanganate index, ammonia nitrogen, total phosphorus, and total nitrogen, and preprocess the data to remove missing values or outliers to ensure the reliability and consistency of data quality. These data will be used as the input of the model, providing a basis for dissolved oxygen prediction.
[0036] Step 2: Normalization method supported by ChatGPT domain knowledge.
[0037] Use ChatGPT to obtain the maximum and minimum values in the water quality domain. By querying each indicator one by one, ChatGPT returns the maximum and minimum values of this feature in the water quality domain and stores them in JSON format. Based on the obtained maximum and minimum values of the water quality domain information, a normalization processing knowledge base is constructed to normalize the original data, that is, map the values of each feature to between [0, 1] to ensure the consistency of the data and the stability and effectiveness of model training. The algorithm flow is shown in Appendix 1 and Table 2; the algorithms for obtaining the maximum and minimum values of the expert knowledge domain include:
[0038] (2.1) Initialize an empty set S as shown in formula (1) to store the domain knowledge of each water quality feature, that is, the minimum and maximum values.
[0039]
[0040] where S is the set of domain knowledge, is an empty set.
[0041] (2.2) For the water quality set, the feature set is W, and each feature is ω. The following operations are performed on each element ω in the water quality feature set W:
[0042] (2.3) Send a request to ChatGPT, which contains information about a specific water quality feature ω, in order to obtain relevant data of this feature in the field of water quality, as specifically shown in formulas (2) and (3):
[0043] Response = ChatGPT(ω) (2)
[0044] response1 = JSON{min(w), max(w)} (3)
[0045] Among them, ω represents the feature in the feature set, and JSON represents the return in JSON format.
[0046] Specifically, traverse all the features in W one by one, and send the specific information of each feature, such as name, data features, etc., to ChatGPT for analysis. Using the existing domain knowledge and a large amount of historical data, analyze the input water quality features and return the maximum and minimum values of this feature in the field of water quality. After successfully sending the request, ChatGPT will return a response, which contains the maximum and minimum values of the queried feature in the field of water quality. The returned response is stored in JSON format. Through this method, each water quality feature will obtain its corresponding upper and lower bounds, providing a reliable basis for subsequent data standardization and normalization processing.
[0047] (2.4) Construct a set S, and add the corresponding result to the set S, as specifically shown in formula (4):
[0048] S ← S ∪ {response1} (4)
[0049] Among them, S represents the response set, and response1 represents the response result received from ChatGPT (the minimum and maximum values in the field).
[0050] Update S to the union of itself and the new response (response1), that is, add the new response1 to the existing set S.
[0051] Specifically, add the response result received from ChatGPT (that is, the maximum and minimum values of feature ω) to the set S to form a complete knowledge base of feature upper and lower limits. Each received response is regarded as a new information unit, and it is merged with the existing set S until all the information of all features is obtained and stored in the set S.
[0052] (2.5) Repeat the operations of formulas (2), (3), and (4) until all features are traversed, and return the complete set of domain knowledge S. Traverse each feature in the water quality feature set W, and perform the operations of steps (2), (3), and (4) for each feature in turn. Gradually collect the maximum and minimum values of all features, and add this information to the set S until the entire feature set W is traversed. The finally obtained set S contains the maximum and minimum values of the expert knowledge domain of all water quality features, constituting a complete domain knowledge base.
[0053] (2.6) Initialize an empty set E to store the normalized water quality data points. The purpose of the normalization process is to standardize water quality features with different dimensions to the same range, thereby eliminating the differences in data dimensions. The set E will contain the normalized water quality data points to ensure that all input data is processed within a unified range, as shown in formula (5):
[0054]
[0055] where E is the set of sequences of normalized water quality data points, is an empty set.
[0056] (2.7) Obtain the maximum value max_val and minimum value min_val of the water quality features from the domain knowledge S. Before performing the normalization operation, extract the upper and lower limits of each water quality feature in the domain knowledge set S, as shown in formulas (6) and (7). These maximum values max_val and minimum values min_val replace the maximum and minimum values in the maximum-minimum normalization, and scale the feature values to the interval [0, 1].
[0057] max_val ← S(6)
[0058] min_val ← S(7)
[0059] where max_val is the maximum value of the expert knowledge domain, and min_val is the minimum value of the expert knowledge domain.
[0060] Based on the obtained maximum and minimum values of the water quality domain, perform the following operations for each element ω in the water quality feature set W:
[0061] (2.8) For each water quality feature ω in the feature set W, perform normalization on it. The specific formula is:
[0062]
[0063] where e is the normalized value of ω, max_val is the maximum value of the expert knowledge domain, and min_val is the minimum value of the expert knowledge domain.
[0064] (2.9) After completing the normalization operation, add the normalization result e of each feature to the set E, and gradually construct the normalized water quality data point sequence, as shown in formula (9) specifically.
[0065] E ← E ∪ {e} (9)
[0066] (2.10) Repeat the steps shown in formulas (6), (7), (8) and (9) until all data points are processed, and return the normalized water quality data point sequence E.
[0067] Step 3: Construct a weighted feature constraint matrix based on ChatGPT.
[0068] By sending a request containing queries about the relationships between various variables in the water quality field to ChatGPT, obtain the correlation weights between each water quality index. These weights are determined through the analysis of a large amount of water quality data, and a weighted feature constraint matrix between water quality features is constructed. Send a query request to ChatGPT, and the content of the request includes the relationships between each pair of variables in the water quality dataset. ChatGPT processes it using the expert knowledge in the field and returns the correlation weights between each pair of features. Based on the obtained feature correlation weights, construct a weighted feature constraint matrix. The algorithm flow is shown in Appendix 3 and Table 4.
[0069] The construction of the weighted feature constraint matrix is completed through the following steps:
[0070] (3.1) Initialize an empty set N to store the correlation weight information between features, as shown in formula (10):
[0071]
[0072] where N is the set of normalized water quality data point sequences, is an empty set.
[0073] (3.2) For each pair of features (ω1, ω2) in the water quality feature set W, perform the following operations:
[0074] For all feature pairs (ω1, ω2) in the water quality feature set W, process them pair by pair. Specifically, if the two features ω1 and ω1 are not the same feature (i.e., ω1 ≠ ω2), that is, if ω1 ≠ ω2 then, send a request to ChatGPT. The content of the request includes the specific information of these two features to obtain the correlation weight, as shown in formula (11):
[0075] Response = ChatGPT(ω1, ω2) (11)
[0076] (3.3) Receive the response from ChatGPT. The response is returned in JSON format and contains the correlation weights of the feature pair (ω1, ω2). After ChatGPT receives the request, it will analyze its knowledge base and domain data to generate a response and return the correlation weights between the two features. The returned response is stored in JSON format, specifically as shown in formula (12):
[0077] response2 = {ω correlation weight (ω1, ω2)} (12)
[0078] Among them, response2 refers to the response result received from ChatGPT (the correlation information of each pair of features), and ω correlation weight is the correlation weight between the two features.
[0079] (3.4) Add the correlation weight information in the response to the domain knowledge set N. Add the correlation weight information in the response obtained from ChatGPT to the previously constructed domain knowledge set N. In this way, gradually expand the set N to make it contain the correlation weight information between each pair of features, specifically as shown in formula (13):
[0080] N ← N ∪ {response2} (13)
[0081] (3.5) Repeat the steps of formulas (11) to (13) until all feature pairs are traversed, and return the final feature constraint matrix N, where each element contains the correlation weight information of the feature pair. Obtain the correlation weights between features pair by pair until all feature pairs are processed. Each time the correlation weights of the feature pair are received from ChatGPT, add them to the feature constraint set N, and finally form the complete feature constraint matrix N, specifically as shown in formula (14):
[0082]
[0083] Among them, the horizontal and vertical directions of the matrix are matrices composed of eight water quality variables, and the diagonal line is the correlation N ii = 1, N ij = * is the correlation between different variables.
[0084] Step Four: Improve the scaled dot - product attention mechanism by fusing the weighted feature constraint matrix.
[0085] Improve the scaled dot - product attention mechanism. Combine the weighted feature constraint matrix described in step three to adjust the result of the scaled dot - product attention mechanism. Multiply the dot - product similarity score between each pair of features by the corresponding element in the weighted feature constraint matrix to form the weighted attention score. The specific steps to improve the scaled dot - product attention mechanism are as follows:
[0086] (4.1) Convert the input data into a feature vector matrix. Calculate the query, key, and value matrices, and convert the input data into three different representations so that when calculating the attention weights, the importance can be effectively evaluated based on the interaction of different features, as shown in formulas (15), (16), and (17) specifically:
[0087]
[0088] Among them, Q is the query matrix, n is the length of the input sequence, and d q is the dimension of the query vector; K is the key matrix, and d k is the dimension of the key vector; V is the value matrix, and d v is the dimension of the value vector.
[0089] (4.2) Calculate the similarity score between the query and the key. Calculate the similarity score between the query Q and the key K. The dot - product operation is used to calculate the similarity between each query and each key vector. The similarity between the query vector and the key vector is calculated through the dot - product operation, as shown in formula (18) specifically:
[0090]
[0091] Among them, Attention Score is the similarity score between the query vector Q i and the key vector K j , Q i is the i - th query vector in the query matrix, K j is the j - th key vector in the key matrix, and · is the vector dot - product operation.
[0092] (4.3) Scaling operation. After obtaining the similarity score between the query and the key, perform the scaling operation. The purpose of scaling is to avoid the dot - product similarity score being too large when the dimension is high, resulting in too small or too large gradients of the softmax function, which affects the training effect of the model. The scaling formula is as (19):
[0093]
[0094] Among them, Scaled Score is the scaled score, and d k is the dimension of the key vector. The scaling operation is performed by dividing by To maintain the stability of the numerical value and ensure that the similarity score does not become too large due to the increase in dimensions.
[0095] (4.4) Fusion feature constraint matrix. Using the weighted feature constraint matrix in domain knowledge, multiply the scaled score element-wise with matrix N. By multiplying the scaled attention score element-wise with matrix N, adjust the attention scores between each pair of features to make them more in line with the feature relationships in a specific domain, so as to better capture the important interactions between various features during model training. The calculation formula is as in (20):
[0096]
[0097] Among them, Weighted Score is the result obtained by multiplying the scaled attention score element-wise with the weighted feature constraint matrix, ⊙ is the element-wise product, and N is the feature constraint matrix.
[0098] (4.5) Calculate the attention weights. Convert the weighted similarity score into attention weights and apply the softmax function. Through exponentiation, convert the scores into a probability distribution. The softmax function converts all attention scores into a probability distribution, and the output values are between [0, 1], and the sum of all weights is 1. The processed attention weights can represent the relative importance of each feature in the current context, specifically as in formula (21):
[0099] Attention Weights = softmax(Weighted Score)(21)
[0100] Among them, Attention Weights are the attention weights after the softmax operation (the values of the vector are converted into a probability distribution).
[0101] (4.6) Calculate the final attention output. Multiply the attention weights with the value matrix V to get the final attention output. The attention weights Attention Weights are applied to the value information of the input features through matrix multiplication with the value matrix V, so as to generate a weighted feature representation, and finally obtain the weighted output of each feature, specifically as in formula (22):
[0102] Attention Output = Attention Weights · V(22)
[0103] Among them, Attention Output represents the value matrix V adjusted by the weighted feature constraint matrix N.
[0104] Step 5: Build a dissolved oxygen prediction model for ChatGPT-EK-TabNet.
[0105] Based on the original TabNet framework, introduce the domain expert knowledge obtained through ChatGPT, and improve the feature selection layer and feature interaction module. Replace the original AttentiveTransformer with an improved scaled dot-product attention mechanism to more effectively capture and process the complex interaction relationships between features.
[0106] Specifically, it includes the following steps:
[0107] (5.1) Combine expert domain knowledge (domain maximum and minimum values) to perform feature normalization at the input layer. At the input stage of the model, first use the expert domain knowledge obtained from ChatGPT to perform feature normalization on the input data.
[0108] (5.2) During the feature interaction process, combine the scaled dot-product attention mechanism. Through the calculation of the query matrix, key matrix, and value matrix, and combine the weighted feature constraint matrix to weight the interaction between each feature, generating a weighted attention score. Multiply the scaled similarity score element-wise with the weighted feature constraint matrix N obtained from ChatGPT. By multiplying the similarity score with the weighted feature matrix, the attention mechanism can more accurately capture the feature interaction relationships that have an important impact on the final prediction, thereby generating a weighted attention score (WeightedScore).
[0109] (5.3) The model applies the weighted attention weights to the feature interaction module to accurately describe the influence of different features in complex water quality data, and replaces the original Attentive Transformer module. And finally build a ChatGPT-EK-TabNet dissolved oxygen prediction model.
[0110] Finally, according to the constructed dissolved oxygen prediction model of ChatGPT-EK-TabNet, use the model evaluation functions RMSE, MAE, and R 2 , specifically as shown in formulas (23), (24), and (25):
[0111]
[0112] where, y i is the i-th true value, is the i-th predicted value, is the average value of the true values, that is n is the number of samples.
[0113] Example 1:
[0114] Taking the daily water quality data of a reservoir in the North China Plain from 2022 to 2023 as an example, which includes several key water quality parameters such as water temperature, pH value, conductivity, turbidity, permanganate index, ammonia nitrogen, total phosphorus, and total nitrogen to predict dissolved oxygen. The data features are 8. The experiment uses the five-fold cross-validation method to verify the model performance. The specific obtained dataset is shown in Table 5, and Steps 1 to 5 of the present invention are used to build a dissolved oxygen prediction model.
[0115] Table 5 Dataset
[0116]
[0117] Step 1: Remove missing values and outliers.
[0118] Clean the original dataset to remove missing values and outliers to ensure the quality and reliability of the model training data. Specifically, missing values refer to the situation where the values of some features in the dataset are empty or not recorded, which may be caused by errors or omissions in the data collection process.
[0119] Step 2: Obtain the maximum and minimum values in the field for normalization.
[0120] First, combine ChatGPT with domain expert knowledge and a large amount of historical water quality data to determine the normal value range of the features. Specifically, for each water quality feature, such as "conductivity" or "permanganate index", the model first generates a request describing the statistical information requirements of the feature, including the maximum and minimum values of the feature under normal circumstances. This request will be sent to ChatGPT, and the detailed statistical information of the feature will be returned and output in JSON format. The process is as Figure 2 shown, including the following content:
[0121] Table 6 Content of the Returned JSON File
[0122]
[0123] Secondly, build a response set S according to the obtained maximum and minimum values in the field, that is, the result response1 after each response is successively added to the response set S. The finally constructed corresponding set S is:
[0124]
[0125] Horizontally starting from the corresponding variables are: water temperature, pH value, conductivity, turbidity, permanganate index, ammonia nitrogen, total phosphorus, total nitrogen.
[0126] Finally, after obtaining these maximum and minimum values, the model normalizes each feature based on these values, scaling the feature values to the range of [0, 1]. The steps of the normalization process are as follows: For each feature, normalization is performed using the following formula:
[0127] where min_val and max_val are the minimum and maximum values of the feature respectively, and the calculation formulas are:
[0128]
[0129] All feature values are scaled to the same range, ensuring that the model can better learn and capture the complex relationships between different features. After completing the normalization operation, the normalized result e of each feature is added to the set E, gradually constructing the normalized water quality data point sequence, specifically E ← E ∪ {e}.
[0130] Step 3: Construct a weighted feature constraint matrix.
[0131] In this embodiment, ChatGPT is combined with domain expert knowledge and a large amount of historical water quality data to obtain the correlation weights between each pair of water quality features, and a weighted feature constraint matrix is constructed based on these weights.
[0132] First, as shown in formula (10): A null set N is constructed to store the correlation weight information between features.
[0133] Secondly, for each pair of features (ω1, ω2) in the water quality feature set W, the following operations are performed: For all feature pairs (ω1, ω2) in the water quality feature set W, they are processed pair by pair. Specifically, if the two features ω1 and ω1 are not the same feature (i.e., ω1 ≠ ω2), that is, if ω1 ≠ ω2 then, a request is sent to ChatGPT. The content of the request includes the specific information of these two features to obtain the correlation weight. After receiving the response from ChatGPT, the response is returned in JSON format, containing the correlation weight of the feature pair (ω1, ω2). When ChatGPT receives the request, it will generate a response by analyzing its knowledge base and domain data, returning the correlation weight between the two features. The returned response is stored in JSON format, and the returned weight content is shown in Table 7.
[0134] Table 7 Content of the returned JSON file
[0135]
[0136]
[0137]
[0138] Finally, add the correlation weight information in the response to the domain knowledge set N, and add the correlation weight information in the response obtained from ChatGPT to the previously constructed domain knowledge set N. In this way, gradually expand the set N to include the correlation weight information between each pair of features, specifically as shown in formula (13): Receiveresponse2 in JSON format containing the correlation weight for (ω1, ω2) and formula (14): N ← N ∪ {response2}. Here, response2 refers to the response result received from ChatGPT (the correlation information for each pair of features), and finally construct the feature constraint matrix N as:
[0139]
[0140] Among them, the horizontal and vertical directions of the matrix are matrices composed of eight water quality variables, and the diagonal line is the correlation N between the same variables ii = 1, N ij is the correlation between different variables.
[0141] Step 4: Improve the scaled dot - product attention mechanism.
[0142] Adjust the result of the scaled dot - product attention mechanism by combining the weighted feature constraint matrix N. Multiply the dot - product similarity score between each pair of features by the corresponding element in the weighted feature constraint matrix to form the weighted attention score.
[0143] First, convert the input data into a feature vector matrix. Calculate the query, key, and value matrices, and convert the input data into three different representations, specifically as shown in formulas (14), (15), and (16):
[0144]
[0145] Secondly, calculate the similarity score between the query and the key. Calculate the similarity score between the query Q and the key K, specifically as shown in formula (17):
[0146]
[0147] Substitute formulas (14) and (15) into formula (17) to obtain the similarity score.
[0148] After obtaining the similarity score between the query and the key, perform a scaling operation. Substitute the Attention Score calculated by formula (17) into formula (18) to calculate the scaled score Scaled Score.
[0149]
[0150] The calculated Scaled Score is multiplied element-wise with matrix N to perform the fused feature weighted matrix operation, adjusting the attention scores between each pair of features to make them more in line with the feature relationships in a specific domain, thereby obtaining the WeightedScore.
[0151]
[0152] Again, the weighted similarity scores are converted into attention weights by applying the softmax function, which converts the scores into a probability distribution through exponentiation, as shown in formula (19):
[0153] Attention Weights = softmax(Weighted Score) (19)
[0154] Finally, the final attention output is calculated. The attention weights are multiplied with the value matrix V to obtain the final attention output, as shown in formula (20):
[0155] Attention Output = Attention Weights · V (20)
[0156] Steps Five and Six: Construct a dissolved oxygen prediction model for ChatGPT-EK-TabNet.
[0157] First, the expert knowledge in the water quality field obtained through ChatGPT, including the maximum and minimum values of key features such as water temperature, pH value, and conductivity, is used for preprocessing the input data. Secondly, the improved scaled dot-product attention mechanism replaces the original Attentive Transformer. By combining the weighted feature constraint matrix, the interaction between different water quality features can be weighted according to their actual domain relevance. Finally, the ChatGPT-EK-TabNet model optimized with expert knowledge is used to predict the dissolved oxygen based on the daily water quality data of a reservoir in the North China Plain from 2022 to 2023. The five-fold cross-validation method is used to verify the model performance, and RMSE, MAE, and R 2 are used as evaluation indicators.
[0158] The evaluation indicators for the ablation experiment are shown in Table 8 and Figure 6As shown (Model 1 to Model 4 represent: ChatGPT expert knowledge normalization + TabNet + scaled dot - product attention mechanism + weighted feature constraint matrix, ChatGPT expert knowledge normalization + TabNet + scaled dot - product attention mechanism, TabNet + scaled dot - product attention mechanism, and TabNet respectively), it can be seen that ChatGPT expert knowledge normalization + TabNet + scaled dot - product attention mechanism + weighted feature constraint matrix exhibits the best prediction performance, significantly improving the accuracy of dissolved oxygen concentration prediction compared to other methods.
[0159] Table 8 Comparison of evaluation metrics for ablation experiments
[0160] Experimental model RMSE MAE <![CDATA[R 2 > Model 1 0.3349 0.2101 0.9388 Model 2 0.4139 0.2686 0.9094 Model 3 0.4571 0.3203 0.8900 Model 4 0.4571 0.4288 0.8268
[0161] The comparison of evaluation metrics for different models is shown in Table 9. The ChatGPT - EK - TabNet model performs optimally among all models. The results show that the ChatGPT - EK - TabNet model has significant advantages in all evaluation metrics compared to other models.
[0162] Table 9 Comparison of evaluation metrics for different models
[0163] Experimental model RMSE MAE <![CDATA[R 2 > ChatGPT-EK-TabNet 0.3349 0.2101 0.9388 LSTM 0.5799 0.4362 0.8241 CNN 0.6763 0.5346 0.7629 BiLSTM 0.7257 0.6048 0.7257 SVR 0.6351 0.4499 0.7910 XGboost 0.5647 0.4172 0.8346 BP 0.6918 0.5552 0.7523
[0164] The comparison of evaluation metrics between different normalization methods and the normalization method proposed in the present invention is shown in Table 10. Compared with traditional normalization methods, the normalization method proposed in the present invention can more accurately match the characteristics of actual water quality data, significantly improving the accuracy of dissolved oxygen concentration prediction. The experimental results show that the introduction of domain expert knowledge plays a crucial role in improving the prediction performance of the model in the data pre - processing stage.
[0165] Table 10 Comparison of evaluation metrics for different normalization methods
[0166] Normalization method RMSE MAE <![CDATA[R 2 > Logarithmic normalization 0.4333 0.3056 0.9015 Quantile normalization 0.4376 0.2990 0.8995 Min-max normalization 0.5534 0.4101 0.8402 ChatGPT expert knowledge normalization 0.3349 0.2101 0.9388
[0167] Figure 7 It is a comparison curve of the true value (Actual) of the dissolved oxygen concentration and the predicted value (Predicted) of the model proposed in the present invention. The data comes from the results of five - fold cross - validation. It can be seen from the figure that the changing trends of the predicted value and the true value are very close. The model can better capture the fluctuations of the dissolved oxygen concentration. As can be seen from the evaluation metrics in Table 8, Table 9, and Table 10, the error between the predicted value and the true value using the method of the present invention is the smallest and the accuracy is the highest.
[0168] Appendix
[0169] Table 1 Algorithm: Using ChatGPT to obtain domain knowledge
[0170]
[0171] Table 2 Algorithm: ChatGPT Expert Knowledge Water Quality Normalization Method
[0172]
[0173] Table 3 Algorithm: Using ChatGPT to Obtain Relevant Knowledge for Constructing Feature Constraint Matrix
[0174]
[0175] Table 4 Algorithm: Constructing Feature Constraint Matrix
[0176]
Claims
1. A dissolved oxygen prediction method integrating ChatGPT domain expert knowledge and deep learning (ChatGPT-EK-TabNet), characterized in that: The following steps are involved: Step 1: Collect corresponding water quality data, including water temperature, pH, conductivity, turbidity, permanganate index, ammonia nitrogen, total phosphorus and total nitrogen, and remove missing values or outliers to ensure data quality; Step 2: Based on the data collected in step 1, the maximum and minimum expert knowledge in the field of ChatGPT water quality is obtained through the algorithm process for the normalization of each indicator; The algorithm process is as follows: by traversing the water quality data features one by one, the information of each feature is sent to ChatGPT for query; ChatGPT returns the maximum and minimum values of the feature in the water quality field and saves them in JSON format; these maximum and minimum value information will be gradually aggregated into a set, and finally form a complete domain knowledge base; This knowledge base is used to normalize the original data, that is, to map the value of each feature to between 0 and 1 to ensure the consistency of the data and the stability and effectiveness of the model training; Step 3: Based on the data obtained in step 1, the expert knowledge of variable relationships in the ChatGPT water quality field is obtained through the algorithm process to construct the correlation weights between various variables in the water quality field, and then construct a weighted feature constraint matrix; The ChatGPT water quality domain variable relationship expert knowledge is as follows: for each pair of features in the water quality data feature set, the correlation weights between the features are obtained through interaction with ChatGPT, and the obtained information is stored in the domain knowledge set in JSON format, and finally summarized into a complete domain knowledge base for constructing a weighted feature constraint matrix; Step 4: Improve the scaled dot product attention mechanism; Using the weighted feature constraint matrix described in step 3, in the scaled dot product attention mechanism, first calculate the dot product similarity scores of the query and key vectors of the input features, then multiply these similarity scores with the corresponding weighted feature constraint matrix elements to obtain the weighted attention scores, and then complete the weighted adjustment of the attention scores, so that each pair of feature interactions reflects different weights in the attention mechanism; Step 5: Construct the dissolved oxygen prediction model of ChatGPT-EK-TabNet; ChatGPT represents an artificial intelligence model for acquiring domain knowledge; EK represents the fusion of expert knowledge; TabNet represents a deep learning prediction model based on tabular data; The ChatGPT-EK-TabNet model is: the expert knowledge described in step 2 is used for normalization of input data, and each eigenvalue is standardized to a unified range to ensure data consistency and stability of model training; The weighted feature constraint matrix described in step 3 is integrated into the scaled dot product attention mechanism to complete the weighted adjustment and fusion of feature interactions. The prediction of dissolved oxygen is achieved through the constructed ChatGPT-EK-TabNet model.
2. The method according to claim 1, characterized in that The step 2 of obtaining the ChatGPT water quality field maximum and minimum expert knowledge through the interface refers to obtaining the maximum and minimum values of each water quality indicator in the water quality field by sending a request for specific water quality field data to ChatGPT. These maximum and minimum values are determined by combining expert knowledge and used for normalization processing. The algorithm flow is shown in Appendix 1 and Table 2; The algorithms for obtaining the maximum and minimum values in the expert knowledge domain include: (2.1) Initialize an empty set S to store the domain knowledge of each water quality feature, i.e., the minimum and maximum values: Where S is the domain knowledge set, is an empty set; (2.2) For each element ω in the water quality feature set W, the following operations are performed: Response=ChatGPT(ω) (2.3) Receive a response from ChatGPT, which is returned in JSON format and contains the minimum and maximum values of the feature ω: response1=JSON{min(w),max(w)} (2.4) Add the response result to the set S: S←S∪{response1}, which means updating S to the union of itself and the new response (response1), that is, adding the new response1 to the existing set S; Where response1 refers to the response result received from ChatGPT (domain minimum and maximum values); (2.5) Repeat steps (2.2) to (2.4) until all features are traversed and the complete domain knowledge set S is returned; (2.6) Initialize an empty set E to store the normalized water quality data points: Where E is the set of normalized water quality data point sequences; (2.7) Obtain the maximum value max_val and minimum value min_val of water quality characteristics from domain knowledge S: max_val←S min_val←S (2.8) For each element ω in the water quality feature set W, the following operations are performed: Where e is the normalized value of ω; (2.9) Add the normalized data point e to the set E: E←E∪{e}; (2.10) Repeat steps (2.7) to (2.8) until all data points are processed and the normalized water quality data point sequence E is returned.
3. The method according to claim 1, characterized in that The step 3 of obtaining the expert knowledge of variable relationships in the water quality field of ChatGPT through the interface refers to obtaining the correlation weights between various water quality indicators by sending a request containing a query on the relationship between various variables in the water quality field to ChatGPT. These weights are determined by analyzing a large amount of water quality data and are used to construct a weighted feature constraint matrix. The matrix incorporates domain knowledge and is used to scale the dot product attention mechanism, so that the model can more accurately capture the interaction relationship between features when calculating the attention score. The algorithm flow is shown in Appendix 3 and Table 4; The algorithms for acquiring expert knowledge between variables in the water quality domain include: (3.1) Initialize an empty set N to store the correlation weight information between features: Where N is the set of correlations between features; (3.2) For each pair of features (ω1, ω2) in the water quality feature set W, perform the following operations: If ω1≠ω2, that is, ifω1≠ω2then, a request is sent to ChatGPT, containing information about the feature pair (ω1,ω2), requesting the correlation weight between the two features: Response=ChatGPT(ω1,ω2); (3.3) Receive a response from ChatGPT, which is returned in JSON format and contains the relevance weights of the feature pair (ω1, ω2): response2={ω correlationweight (ω1,ω2)} (3.4) Add the relevance weight information in the response to the domain knowledge set N: N←N∪{response2}, which means updating S to the union of itself and the new response (response2), that is, adding the new response2 to the existing set N; Where response2 refers to the response result received from ChatGPT (correlation information of each pair of features); (3.5) Repeat steps (3.2) to (3.4) until all feature pairs are traversed, and return the final feature constraint matrix N, in which each element contains the correlation weight information of the feature pair; The constructed feature constraint matrix N is characterized by: The horizontal and vertical axes of the matrix are composed of eight water quality variables, and the diagonal is the correlation N between the same variables. ii =1,N ij =* represents the correlation between different variables.
4. The method according to claim 1, characterized in that: The improved scaled dot product attention mechanism in step 4 refers to adjusting the result of the scaled dot product attention mechanism using the weighted feature constraint matrix in step 3, multiplying the dot product similarity score between each pair of features by the corresponding element in the weighted feature constraint matrix, thereby forming a weighted attention score; Improved scaling dot product attention mechanism, characterized by: (4.1) The input data is converted into a feature vector matrix, which is characterized by: Where Q is the query matrix, n is the length of the input sequence, and d q is the dimension of the query vector; Where K is the bond matrix, d k is the dimension of the key vector; Where V is the value matrix, d v is the dimension of the value vector; (4.2) Calculate the similarity score between the query and the key. Calculate the similarity score between the query Q and the key K. The dot product operation is used to calculate the similarity between each query and each key vector: Where Attention Score is the query vector Q i With the key vector K j The similarity score between i is the i-th query vector in the query matrix, K j is the jth key vector in the key matrix, · is the vector dot product operation; (4.3) Scaling operation, the calculation formula is as follows: Where Scaled Score is the scaled score, d k is the dimension of the key vector, and the scaling operation is done by dividing by To maintain numerical stability and ensure that the similarity score does not become too large due to the increase in dimension; (4.4) Fusion feature constraint matrix, using the weighted feature constraint matrix in domain knowledge, multiply the scaled score by the matrix N element by element. The calculation formula is as follows: Where Weighted Score is the result of element-by-element multiplication of the scaled attention score and the weighted feature constraint matrix, ⊙ is the element-by-element product, and N is the feature constraint matrix; (4.5) Calculate the attention weight, convert the weighted similarity score into the attention weight, apply the softmax function, and convert the score into a probability distribution through an exponential operation. The calculation formula is as follows: Attention Weights=softmax(Weighted Score) Where Attention Weights is the attention weight after the softmax operation (the value of the vector is converted into a probability distribution); (4.6) Calculate the final attention output. The attention weight is multiplied by the value matrix V to obtain the final attention output. The output formula is as follows: Attention Output=Attention Weights·V Where Attention Output represents the value matrix V adjusted by the weighted feature constraint matrix N.
5. The method according to claim 1, characterized in that The construction of the ChatGPT-EK-TabNet dissolved oxygen prediction model in step 5 refers to introducing the domain expert knowledge obtained through ChatGPT on the basis of the original TabNet framework, and improving the feature selection layer and feature interaction module; it is characterized in that the original Attentive Transformer is replaced by the improved scaled dot product attention mechanism to more effectively capture and process the complex interaction relationship between features. Specifically include: (5.1) Combine expert domain knowledge (domain maximum and minimum values) to perform feature normalization at the input layer; (5.2) In the process of feature interaction, combined with the scaled dot product attention mechanism, the interaction between features is weighted by calculating the query matrix, key matrix and value matrix, and the weighted feature constraint matrix is used. The Attention Score is multiplied element by element with the weighted feature constraint matrix to generate a weighted attention score. (5.3) The model achieves an accurate description of the impact of different features in complex water quality data by applying weighted attention weights to the feature interaction module.
6. The method according to claim 5, characterized in that The ChatGPT-EK-TabNet dissolved oxygen prediction model in step 5 uses the model evaluation functions of RMSE, MAE and R 2 ; Where: y i is the ith true value, is the ith predicted value, is the average of the true values, that is n is the sample size.