Water plant equipment working condition prediction method based on multi-factor coupling
The water plant equipment operating condition prediction method, which utilizes GloVe, LSTM, PNA, and ALN to extract triplet data features and combines them with the Transformer architecture for feature fusion and prediction, solves the problem of coupling relationship between water plant equipment operating condition parameters and improves the accuracy and robustness of prediction.
Patent Information
- Application Number
- CN202511491708.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-20
- Publication Date
- 2025-11-21
AI Technical Summary
Existing technologies struggle to effectively decouple the dynamic coupling relationships between parameters of water plant equipment under different operating conditions, leading to a chain propagation of prediction errors and impacting prediction accuracy and robustness.
A water plant equipment condition prediction method based on multi-factor coupling is adopted. The GloVe algorithm, Long Short-Term Memory Network (LSTM), Primary Neighborhood Aggregation Network (PNA), and Association Link Network (ALN) are used to extract features from triple data. The Transformer architecture is combined to perform feature fusion and prediction, and candidate events are selected through cross-attention calculation.
It improves the accuracy and robustness of water plant equipment condition prediction, reduces model complexity, and solves the shortcomings of traditional methods in handling multi-dimensional nonlinear coupled data, thus achieving more accurate equipment condition prediction.
Smart Images

Figure CN120995026A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method for predicting the operating conditions of water plant equipment based on multi-factor coupling, belonging to the fields of intelligent prediction and deep learning technology. Background Technology
[0002] Predicting the operating conditions of water treatment plant equipment is crucial for ensuring the stable operation of waterworks, improving water supply quality, and reducing operating costs. However, due to the complex operating environment and numerous influencing factors, predicting equipment operating events faces many challenges.
[0003] Water plant equipment operating data is characterized by complexity and uncertainty. On the one hand, it is influenced by a combination of factors, such as the equipment's aging level, water quality conditions, and operating load. These factors are intertwined, exhibiting complex nonlinear relationships. On the other hand, equipment operation may be affected by random events such as sudden failures and abnormal operations, leading to randomness and unpredictability in the operating data. This complexity and uncertainty make equipment operating script events more difficult to predict than traditional industrial equipment data, significantly increasing the complexity when predicting long-term equipment operating conditions.
[0004] Traditional methods for predicting water plant equipment operating condition script events primarily rely on manual experience or simple statistical analysis models. Manual experience is often limited by individual knowledge and experience, making it difficult to comprehensively and accurately capture changing trends in equipment operating conditions, and prone to misjudgments and omissions. Statistical analysis methods, such as regression analysis and time series analysis, while capable of modeling and analyzing historical data to some extent, struggle to effectively handle multi-dimensional, nonlinearly coupled data and fail to fully consider the various complex factors and interactions during equipment operation. Therefore, a novel prediction framework is urgently needed that can decouple the dynamic coupling relationships between different operating parameters of water plant equipment and achieve error compensation. This framework should ensure lightweight and real-time performance while improving the overall prediction accuracy and robustness of the system, providing a reliable technology for the safe and stable operation of water plant equipment. Summary of the Invention
[0005] The purpose of this invention is to overcome the above-mentioned shortcomings and provide a water plant equipment operating condition prediction method based on multi-factor coupling. This method can decouple the dynamic coupling relationship between different operating condition parameters of the equipment to predict operating condition script events, effectively reduce the chain propagation of prediction errors, and at the same time reduce the complexity of the algorithm model and improve the robustness of the model.
[0006] The technical solution adopted in this invention is as follows: The water plant equipment operating condition prediction method based on multi-factor coupling includes the following steps: S1. Obtain historical data on the operating conditions of water plant equipment, preprocess it into a time series based on the time sequence, and construct "subject, verb, object" triplet data to form historical events. Process the operating condition data to be measured into time series and triplet data to form candidate events. S2. Input the triplet data into the word embedding module. The word embedding module uses the GloVe algorithm to convert the triplet data into vector representations. Then, the vector representations are sent to the feature extraction module. The feature extraction module uses the Long Short-Term Memory Network (LSTM) and the Main Neighborhood Aggregation Network (PNA) to extract the relationship features between historical events and candidate events. It uses the Association Link Network (ALN) to extract the contextual features of historical event triples. In the relationship features, LSTM is first used to extract the semantic features of the event chain based on the time series, and then the Main Neighborhood Aggregation Network (PNA) is used to extract the local relationship features of the events. S3. The extracted features are de-noiseed by the feature selector and then input into the improved Transformer. The attention fusion mechanism in the decoder stage is used to perform feature fusion, and the fused event representation is output. S4. Input the fused event representation into the cross-attention calculation and prediction module. The cross-attention calculation and prediction module obtains the association scores between each parameter of the candidate event and the historical event based on the calculation of local association degree. It obtains the weight relationship between each triple parameter of the candidate event and the historical event chain through the calculation of overall association degree. Then, it calculates the final similarity score between the historical event chain and each candidate event, and selects the candidate event with the high similarity score as the most likely event to occur in the future as the final prediction result. S5. Optimize the model parameters by calculating the prediction loss function to obtain the final prediction model.
[0007] In the above method, in step S2, the word embedding module, each triplet event is embedded using GloVe pre-trained word vectors and continuously updated during training. i (i∈{1,2,...,n}, where n represents the number of its historical events), its embedding representation e i It is by using the verb v i and its related parameters (subject, object). a i,0 , a i,l The formula for calculating the result when mapped to the same vector space is as follows: , in, , , These represent the weight matrices for the verb, subject, and object, respectively. b eThis represents the bias vector. For words not in the vocabulary, a zero vector is used. For events with fewer than three elements, the corresponding positional component is set to "NaN".
[0008] In step S2, the feature extraction module, the Main Neighborhood Aggregation Network (PNA) uses multiple aggregators—mean aggregator, maximum aggregator, minimum aggregator, and standard deviation aggregator—and a scaling factor based on node degree to aggregate features. PNA combines a scaling factor based on node degree with these multiple aggregators. The scaling factor based on node degree is used to determine the amount of information to be aggregated. Multiplying this scaling factor by the aggregated value increases or decreases the squareness of the incoming message features, thus mitigating the impact of differences in node degree. The specific calculation process is as follows: (1) First, perform scaling based on node degree. The scaling formula based on node degree is as follows: , in d The degree of a node. γ It is a variable parameter. Represents a value vector. Is injective functions under certain conditions; (2) Then, the neighbor node information is aggregated. The mathematical formula for PNA to aggregate neighbor node information is as follows: , in I s This indicates that scaling should not be performed. Tensor multiplication is defined, and different scalers are used to scale the features of neighbor nodes. Then, different aggregators are used to aggregate the scaled features of neighbor nodes. ⊕ is used to integrate the results obtained by combining different scalers and aggregators. Different scalers and aggregators are used to capture different aspects of the information of neighbor nodes, so as to comprehensively aggregate the information of neighbor nodes. (3) The aggregated neighbor node information is concatenated and integrated with the current node's own features. Then, through the processing of the neural network layer, the features are further transformed and extracted to obtain a feature representation that can be used for subsequent tasks. The process of extracting feature information from PNA can be defined by the following formula: , In the formula l It is the current layer of PNA. φ It is a single-layer linear neural network. X i (l) Indicates the current nodei In the l The feature vector of the layer. X j (l) Represents a node i neighboring nodes j In the l The feature vector of the layer, Edge is defined as the event node in the event graph. e i and nodes e j There are edges connecting them.
[0009] The process by which the Association Link Network (ALN) extracts contextual features from historical event triples is as follows: The event parameter is represented as e. i = {w1, w2, w3}, corresponding to the verb, subject, and object in the triple, respectively. The bag-of-words table Count = w1, w2, ..., w3 is used to determine the event chain. n In this process, the co-occurrence relationships of all words in historical event links are used to organize event contextual information, and an ALN model is constructed to extract contextual features to provide prior knowledge for event prediction. The basic components of ALN are semantic nodes and their associated connections, and the word w is calculated. i with w j Correlation weights between The formula is as follows: , in Co ( w i , w j ) indicates words w i and w j The frequency of simultaneous occurrence in the event chain DF ( w i ) indicates words w i Frequency of occurrence in the event chain, related link weight score ∈[0,1] indicates the strength of the association between words; By using the GCN algorithm to enhance the differentiation of node features, GCN automatically encodes the current node through its neighbors, thereby extracting features that enrich the graph structure of its neighbors: , in H (l) It is the first l The feature matrix of the layer, where σ is the ReLU activation function. Let I be the identity matrix and A be the adjacency matrix. A is obtained by using... Fill in the correlation score between parameters in the event chain, W (l) It is the first l The weight matrix of the layer, yes The degree matrix obtained through standardization; After updating all node vectors in the ALN to convergence, the semantic embedding vector S is computed. i as follows: , in h i It is the vector representation of node i output by GCN. N G The number of nodes in ALN, for each event. e i Semantic information is used to learn new embedding representations by integrating the contextual embeddings of corresponding words, and events. e i Embedded representation after learning contextual information The calculation process is as follows: , in Represents the parameters of a triple. , and Weight matrices corresponding to different triplet parameters. This indicates the bias term.
[0010] In step S3, the feature selector uses the following formula: , Among them W and b i These are the weight matrix and the bias term, respectively. Indicates contextual features, P ( e i ) represents the probability of semantic features passing, when P ( e i If the feature value is ≥0.5, the feature will be used for fusion; otherwise, it will be discarded.
[0011] In step S3, in the improved transformer, contextual features are input into the attention-based feature fusion layer of the decoder after being encoded. Relationship features (including semantic features of event chains and local event relationship features) are also input into the attention-based feature fusion layer after passing through a multi-head self-attention layer, an addition and normalization layer. In the attention-based feature fusion layer, event chain features and local relationship features are split into historical event features and candidate event features according to event type. The attention distribution of the two features, contextual features and historical event features, is calculated using the softmax function. This attention distribution is used as the weight of the corresponding feature state and a weighted average is calculated. The weighted average is then input into the fully connected network MLP along with the split candidate event features to output the fused feature vector representation. Finally, after normalization and feedforward processing, the fused event representation is obtained.
[0012] In step S4, the fused event representation is split into historical events, candidate events, and historical event chains by event type. In the cross-attention calculation module, this includes the calculation of overall correlation and the calculation of local correlation.
[0013] The calculation of local correlation involves first calculating the similarity between each triplet parameter of the candidate event and each historical event through MLP neural network layers, then normalizing the similarity features using softmax to obtain the weight coefficients of each triplet parameter of the candidate event and each historical event, and finally summing the weight coefficients of each triplet parameter of the candidate event to obtain the correlation score between each parameter of the candidate event and the historical event. The calculation of overall correlation involves first compressing the historical event chain features using average pooling to facilitate subsequent calculations, then using MLP and softmax to calculate the similarity between each triplet parameter of the candidate event and the historical event chain, and normalizing the similarity to obtain the weight relationship between each triplet parameter of the candidate event and the historical event chain, finally using the local calculated score and the overall correlation calculated coefficient to calculate the individual score of each triplet parameter of the candidate event, and obtaining the final score of the entire candidate event by summing the scores of each triplet parameter.
[0014] Step S5 learns the event embedding by minimizing the cross-entropy loss function L(θ): , Where n is the length of the historical event chain, S ( z , c () represents the prediction score of the candidate event c corresponding to the historical event chain, log is the logarithmic function, and L2 is the L2 regularization parameter. λ Here, θ is the regularization coefficient, and θ is the model parameter.
[0015] Compared with the prior art, the beneficial effects of the present invention are: (1) This invention proposes a novel end-to-end scene event prediction model architecture based on multi-feature fusion. It utilizes the respective advantages of the Association Linking Network (ALN), the Long Short-Term Memory Network (LSTM), and the Main Neighborhood Aggregation Network (PNA) to extract different features of the triplet time, so as to learn better features and improve the prediction accuracy of the entire architecture. It overcomes the problem that existing technologies are difficult to effectively handle the complex coupling relationship between multiple operating parameters of water plant equipment, which easily leads to the continuous accumulation of prediction errors among variables and affects the prediction accuracy.
[0016] (2) For different features, this invention uses a feature selector to filter out noisy data. At the same time, it introduces the Transformer architecture and improves the decoding end by using an attention-based multi-feature fusion method to further integrate and enhance the semantic features, contextual features, and local contextual features of the triples. This solves the problem that traditional attention at the decoding end will forcibly aggregate the features of the encoding and decoding ends, treating the event chain features and local features of the decoding end as a whole, and thus cannot extract the features of the candidate events separately.
[0017] (3) In the prediction module, a cross-selection calculation method based on attention mechanism is used. By jointly calculating the scores between the historical event chain and the candidate events and the scores of the candidate events on the historical event chain, the two different calculation methods are statistically integrated to jointly select subsequent events, so as to solve the problem that traditional event prediction only considers the influence of historical events on the candidate events. Attached Figure Description
[0018] Figure 1 This is a schematic diagram of a triplet dataset in an embodiment of the present invention; Figure 2 This is a schematic diagram of the architecture of the method model of the present invention; Figure 3 This is a schematic diagram of the aggregation process of the Primary Neighborhood Aggregation Network (PNA) in an embodiment of the present invention; Figure 4 This is a diagram of the improved transformer architecture in an embodiment of the present invention; Figure 5 This is a schematic diagram of the structure of the attention-based feature fusion layer in an embodiment of the present invention; Figure 6 This is a schematic diagram of the cross-attention calculation and prediction module in an embodiment of the present invention; Figure 7 This is a flowchart of the cross-attention calculation and prediction module in an embodiment of the present invention. Detailed Implementation
[0019] The present invention will be further described below with reference to specific embodiments.
[0020] Example 1: A method for predicting the operating conditions of water plant equipment based on multi-factor coupling, comprising the following steps: S1. Acquire historical data on the operating conditions of water plant equipment, preprocess it into time series based on chronological order, and construct "subject-verb-object" triplet data to form historical events. Process the operating condition data to be measured into time series and triplet data to form candidate events: Based on the timeline, data from each water treatment process in the water plant is preprocessed, including data stabilization and missing data completion. Ultimately, data representative of events is processed into triples (e.g., ...). Figure 1 For example, "emit (motor, abnormal noise)" is used to indicate the event that "the motor emits abnormal noise".
[0021] (1) Data cleaning For a triplet dataset, duplicate data is removed from the dataset.
[0022] (2) Missing value imputation For triples in the triple dataset that contain only a subject and a predicate, the object is represented by "NaN".
[0023] A dataset is constructed based on the optimally preprocessed data, comprising a training set, a validation set, and a test set in a ratio of 8:1:1.
[0024] S2. Input the triplet data into the word embedding module. The word embedding module uses the GloVe algorithm to convert the triplet data into vector representations, and then sends the vector representations into the feature extraction module. The feature extraction module uses a Long Short-Term Memory (LSTM) network and a Principal Neighborhood Aggregation Network (PNA) to extract the relational features between historical events and candidate events, and uses an Association Link Network (ALN) to extract the contextual features of historical event triples. In the relational features, LSTM is first used to extract the semantic features of the event chain based on the time series, and then the PNA is used to extract the local relational features of the events. (1) Embedding layer: Events are represented as triples containing three internal parameters, namely <verb, subject, object>. Each triple component is embedded using GloVe pre-trained word vectors and continuously updated during training. For event i (i∈{1,2,...,n}, where n represents the number of its historical events), its embedding representation is... e i It is by using verbs v i and its related parameters a i,0 , a i,1 It is calculated by mapping to the same vector space. The calculation formula is as follows: , in, , , These represent the weight matrices for the verb, subject, and object, respectively. b e This represents the bias vector. For words not in the vocabulary, a zero vector is used. For events with fewer than three elements, the corresponding positional component is set to "NaN".
[0025] (2) Feature extraction module: We introduce a Long Short-Term Memory (LSTM) network to extract semantic features of time-series-based event chains, a Principal Neighborhood Aggregation (PNA) network to extract local relational features of events, and an Association Link Network (ALN) to extract contextual features of triple events. (See...) Figure 2 .
[0026] ①LSTM: To explore the contextual relationships within an event chain, we introduce the Sequence Model (LSTM) to compute and learn the sequence representations of events, and extract the features contained between events by learning the semantic information passed between them. For an event chain {e1, e2, ..., e...},... n−1 We use a standard LSTM to model the narrative event sequence of the event chain. Events are recursively input into the LSTM, and their hidden state vectors are computed. The formula is as follows: , Where the hidden state vector Includes e1 to e i The event sequence contains relevant information, while the initial hidden state vector h0 is randomly initialized.
[0027] ②PNA: Based on event representation, an event evolution graph is constructed, and a graph neural network is used to learn denser relationship information between event nodes in local neighborhoods, thereby enriching and enhancing the feature relationships between local events. Traditional graph neural networks use a single aggregator, which cannot fully extract neighborhood information of single-layer nodes, easily leading to injection problems, resulting in the loss of specific neighborhood information and over-smoothing of node features. At the same time, when traditional graph neural networks use summation-type aggregators, small changes in the aggregator may cause gradients and aggregated information to be exponentially amplified or decayed, thereby reducing the model's generalization ability. To address these issues, we introduce the Main Neighborhood Aggregation Network (PNA) in the event local context feature extraction process. PNA uses multiple aggregators, such as mean aggregator, maximum aggregator, minimum aggregator, and standard deviation aggregator, and performs feature aggregation based on node degree, strengthening the differentiated expression of node neighborhood information, obtaining better node representations, effectively solving the problem of over-smoothing of node features, and improving the model's generalization ability. Typically, in graph networks, nodes with high degree receive information from more neighbors, while nodes with low degree receive less information. PNA combines a node-degree-based scaler with various aggregators. The node-degree-based scaler, which represents the amount of information to be aggregated, amplifies or attenuates the features of the incoming message by multiplying it by the aggregated value, thus mitigating the impact of differences in node degree. In this process, the outputs of different aggregators are combined with the scaling factor calculated based on node degree. This allows information from nodes with high node degrees to be appropriately attenuated, while information from nodes with low node degrees is amplified. This ensures that the model can process the information from each node in the graph more balancedly and effectively, improving the expressive and generalization capabilities of the graph neural network. Figure 3 This demonstrates the specific implementation of the main neighborhood aggregation process: First, perform scaling based on node degree. The scaling formula based on node degree is as follows: , in d The degree of a node. γ It is a variable parameter. Represents a value vector. Is injective functions under certain conditions; Next, the neighbor node information is aggregated. The mathematical formula for PNA to aggregate neighbor node information is as follows: , in I s This indicates that scaling should not be performed. Tensor multiplication is defined, and different scalers are used to scale the features of neighbor nodes. Then, different aggregators are used to aggregate the scaled features of neighbor nodes. ⊕ is used to integrate the results obtained by combining different scalers and aggregators. Different scalers and aggregators are used to capture different aspects of the information of neighbor nodes, so as to comprehensively aggregate the information of neighbor nodes. The aggregated neighbor node information is concatenated and integrated with the current node's own features. Then, through processing by neural network layers, the features are further transformed and extracted to obtain a feature representation that can be used for subsequent tasks. The process of extracting feature information from PNA can be defined by the following formula: , In the formula l This is the current layer of PNA, where φ is a linear neural network layer. X i (l) Indicates the current node i In the l The feature vector of the layer. X j (l) Represents a node i neighboring nodes j In the l The feature vector of the layer, Edge is defined as the event node in the event graph. e i and nodes e j There are edges connecting them.
[0028] In the middle of the formula l It is the current layer of PNA. φ It is a single-layer linear neural network. X i (l) Indicates the current node i In the l The feature vector of the layer. X j (l) Represents a node i neighboring nodes j In the l The feature vector of the layer, Edge is defined as the event node in the event graph. e i and nodes e j There are edges connecting them.
[0029] ③ ALN (Association Link Network): An event chain refers to a series of related events occurring in a specific order within a short period of time. Since the model input is a ternary event description with a limited vocabulary, it is difficult to fully express the semantic information of the events and the semantic background information that triggers them. Therefore, we introduce an Association Link Network (ALN) in the multi-feature extraction layer. By calculating the association relationships between words in different events, we extract the contextual information of the current event chain and use it as prior knowledge for event prediction. Candidate events usually have similar background information to historical events and contain similar historical event context information. Therefore, when the words in the triples are the same, they can be considered to have some kind of association, and the correspondence between words in the event chain will reflect the semantic connection between events. For existing event chains (e1, e2, ..., e...),... n For ease of description, we will represent the event parameter as e. i = {w1, w2, w3}, corresponding to the verb, subject, and object in the triplet, respectively. In determining the event chain, the bag-of-words list Count = w1, w2, ..., w n In this approach, we organize event contextual information by analyzing the co-occurrence relationships of all words in historical event links, and construct an ALN model to extract contextual features, providing prior knowledge for event prediction. The basic components of ALN are semantic nodes and their associated connections. We calculate the word w... i with w j Correlation weights between The formula is as follows: , in Co ( w i , w j ) indicates the word w i and w j The frequency of simultaneous occurrence in the event chain DF ( w i ) indicates words w i Frequency of occurrence in the event chain. Related link weight score. ∈[0,1] indicates the strength of the association between words.
[0030] Since ALN is a graph-based structure based on word association strength, it cannot be directly used for computation; therefore, it needs to be encoded as a high-dimensional vector. To achieve network encoding and alleviate the over-smoothing problem of graph neural networks, we enhance the differentiation of node features using the GCN algorithm. GCN automatically encodes the current node through its neighbors, thereby extracting features from its neighbors that enrich the graph structure. The specific calculation process is as follows: , in H (l) It is the first l The feature matrix of the layer, where σ is the ReLU activation function. yes The degree matrix obtained by standardization, Let I be the identity matrix and A be the adjacency matrix. A is obtained by using... Fill in the correlation score between parameters in the event chain, W (l) It is the first l Layer weight matrix.
[0031] After updating all node vectors in the ALN to convergence, we compute the semantic embedding vector. S i as follows: , Where h i It is the vector representation of node i output by GCN, N G This represents the number of nodes in ALN. For event e... i Based on semantic information, we integrate the contextual embeddings of corresponding words to learn new embedding representations for events. e i Embedded representation after learning contextual information The calculation process is as follows: , in Represents the parameters of a triple. , and Weight matrices corresponding to different triplet parameters. This indicates the bias term.
[0032] S3. After the extracted features are denoised by the feature selector, they are input into the improved Transformer. The attention fusion mechanism in the decoder stage is used to perform feature fusion, and the fused event representation is output: In the feature extraction process described above, we obtained event chain features, local contextual semantic features, and contextual features, respectively. Considering that direct fusion operations might introduce significant noise and negatively impact model performance, we dynamically select semantic features to optimize event representation during training and design a corresponding feature selector to fuse the selected semantic features with other features. The feature selection formula is as follows: , Among them W and b i These are the weight matrix and the bias term, respectively. Indicates contextual features, P ( ei () represents the probability of semantic features being passed. When P ( e i When the semantic selection criterion is ≥0.5, the features are used for fusion; otherwise, they are discarded. Through this semantic selection mechanism, our model can automatically filter semantic noise and improve the accuracy of event representation after fusing semantic features.
[0033] We achieved comprehensive extraction of multi-dimensional features by introducing the Transformer model and designing a feature fusion scheme based on an attention mechanism, such as... Figure 4 As shown. We also designed a novel feature fusion module that integrates event chain features, local contextual semantic features, and event contextual features through an attention fusion mechanism. Its calculation formula is as follows: , in η , α These are the weight scores after softmax normalization. It is a vector representation of contextual features. x i and x c These are the historical event vector and candidate event vector output by PNA, respectively. In the improved transformer, contextual features are input into the attention-based feature fusion layer of the decoder after being encoded, such as... Figure 5 The relational features (including the semantic features of the event chain and the local relational features of the event) are input into the attention-based feature fusion layer after passing through the multi-head self-attention layer, the addition and normalization layer. In the attention-based feature fusion layer, the event chain features and the local relational features are split into historical event features and candidate event features according to the event type. The attention distribution of the contextual features and historical event features is calculated by the softmax function. The attention distribution is used as the weight of the corresponding feature state and a weighted average is calculated. The weighted average is then input into the fully connected network MLP along with the split candidate event features to output the fused feature vector representation. Finally, after normalization and feedforward processing, the fused event representation can be obtained.
[0034] The entire fusion process can be defined by the following formula: , in H join This represents the event vector representation output by the attention fusion layer. f i It is the output of the hidden layer. W T b and b are the learnable weight matrix and bias vector, respectively.
[0035] S4. The fused event representation is input into the cross-attention calculation and prediction module. The cross-attention calculation and prediction module obtains the association scores between each parameter of the candidate event and the historical events based on the calculation of local association. It obtains the weight relationship between each triple parameter of the candidate event and the historical event chain through the calculation of overall association. Then, it calculates the final similarity score between the historical event chain and each candidate event, and selects the candidate event with the highest similarity score as the most likely event to occur in the future as the final prediction result. In this invention, we employ an attention-based cross-computation mechanism. This mechanism dynamically analyzes the relationship between each candidate event and historical time, revealing the strength of the correlation between historical events and candidate times. The attention-based cross-computation dynamically learns how different candidate events represent historical events, while also considering their interrelationships. Specifically, information from different candidate events focuses on different parameter dimensions of historical events, thereby determining the representation of candidate events. Subsequently, historical events give differentiated attention to each candidate event to determine its weight allocation. This attention-based cross-computation mechanism includes the attention component (CH) of candidate events on historical events and the attention component (HC) of historical events on candidate events, such as... Figure 6 As shown.
[0036] When evaluating a candidate event, we first examine its parameter information, then reread the historical event chain to identify the key aspects. We then move on to the next aspect, rereading the historical event chain again, until all aspects have been utilized. After reviewing all aspects of the candidate events and obtaining all scores, the final similarity score between the historical and candidate events should be a weighted sum of all scores. This mechanism helps to understand historical events using candidate events and may yield better results.
[0037] ①C-H Attention: Different historical events should be evaluated using a differentiated approach based on their three-dimensional representations. Attention can be drawn through historical events. f j and candidate events fc The correlation between them is used to measure the calculation process as follows: , , in w ij Let represent the similarity score between the triplet parameter i of the candidate event and the historical event j, tanh is the nonlinear activation function, i.e., hyperbolic tangent transform, n is the length of the historical event chain, and W and b are the weight matrix and bias vector, respectively. α ij Represents candidate event elements fci The attention weights for the j-th historical event ∈{v,a0,a1}.
[0038] Subsequently, based on the specific candidate event, we calculate a weighted sum of the hidden representations using attention weights to obtain a score representing the association between the candidate event triple parameters and all historical events: .
[0039] ②H-C Attention Each candidate event should focus on the same event chain. For a historical event chain h, the similarity score of each candidate event c is calculated as follows: , , , in f ck and f ci They are respectively taken from the set {v, a0, a1}. β fci The attention weights of the historical event chain on candidate events are used to determine which candidate events should be focused on the (h, c) pair. W and b represent the weight matrix and bias vector, respectively. It calculates the average of all hidden state sequences to represent a chain of historical events, which is used to determine which candidate event should receive more attention.
[0040] The merged event representation is broken down into historical events, candidate events, and historical event chains based on event type. The cross-attention calculation module includes the calculation of overall correlation and local correlation, such as... Figure 7 .
[0041] The calculation of local correlation, i.e., CH Attention, first uses an MLP neural network layer to calculate the similarity between each triplet parameter of the candidate event and each historical event. Then, the similarity features are normalized using softmax to obtain the weight coefficients of each triplet parameter of the candidate event and each historical event. Finally, the correlation score between each parameter of the candidate event and the historical event is obtained by summing the weight coefficients of each triplet parameter of the candidate event. The calculation of overall correlation, i.e., H-CA Attention, first uses average pooling to compress the features of the historical event chain for easier subsequent calculation. Then, MLP and softmax are used sequentially to calculate the similarity between each triplet parameter of the candidate event and the historical event chain. The similarity is normalized to obtain the weight relationship between each triplet parameter of the candidate event and the historical event chain. Finally, the individual score of each triplet parameter of the candidate event is calculated by jointly using the local calculated score and the overall correlation calculated coefficient. The final score of the entire candidate event is obtained by summing the scores of each triplet parameter.
[0042] S5. Optimize the model parameters by calculating the prediction loss function to obtain the final prediction model: Each event chain contains a set of candidate subsequent events, and we learn the event embeddings by minimizing the cross-entropy loss function L(θ): , Where n is the length of the historical event chain, S ( z , c () represents the prediction score of the candidate event c corresponding to the historical event chain, log is the logarithmic function, and L2 is the L2 regularization parameter. λ Here, θ is the regularization coefficient, and θ is the model parameter.
[0043] The above description is merely a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A water plant equipment operating condition prediction method based on multi-factor coupling, characterized by: The steps include the following: S1. Obtain historical data on the operating conditions of water plant equipment, preprocess it into a time series based on the time sequence, and construct "subject, verb, object" triplet data to form historical events. Process the operating condition data to be measured into time series and triplet data to form candidate events. S2. Input the triplet data into the word embedding module. The word embedding module uses the GloVe algorithm to convert the triplet data into vector representations. Then, the vector representations are sent to the feature extraction module. The feature extraction module uses the Long Short-Term Memory Network (LSTM) and the Main Neighborhood Aggregation Network (PNA) to extract the relationship features between historical events and candidate events. It uses the Association Link Network (ALN) to extract the contextual features of historical event triples. In the relationship features, LSTM is first used to extract the semantic features of the event chain based on the time series, and then the Main Neighborhood Aggregation Network (PNA) is used to extract the local relationship features of the events. S3. The extracted features are de-noiseed by the feature selector and then input into the improved Transformer. The attention fusion mechanism in the decoder stage is used to perform feature fusion, and the fused event representation is output. S4. Input the fused event representation into the cross-attention calculation and prediction module. The cross-attention calculation and prediction module obtains the association score between the triple parameters of each candidate event and the historical event based on the calculation of local association degree. It obtains the weight relationship between the triple parameters of each candidate event and the historical event chain through the calculation of overall association degree. Then, it calculates the final similarity score between the historical event chain and each candidate event, and selects the candidate event with the high similarity score as the most likely event to occur in the future as the final prediction result. S5. Optimize the model parameters by calculating the prediction loss function to obtain the final prediction model.
2. The water plant equipment operating condition prediction method based on multi-factor coupling according to claim 1, characterized in that, In step S2, the word embedding module, each triplet event is embedded using GloVe pre-trained word vectors and continuously updated during training. i Its embedded representation e i It is by using the verb v i and its related parameters a i,0 , a i,l The formula for calculating the result when mapped to the same vector space is as follows: , in, , , These represent the weight matrices for the verb, subject, and object, respectively. b e This represents the bias vector. For words not in the vocabulary, a zero vector is used. For events with fewer than three elements, the corresponding positional component is set to "NaN".
3. The water plant equipment operating condition prediction method based on multi-factor coupling according to claim 1, characterized in that, In step S2, the feature extraction module uses the main neighborhood aggregation network (PNA) to aggregate features by employing multiple aggregators, including mean aggregator, maximum aggregator, minimum aggregator, and standard deviation aggregator, as well as a scaling based on node degree. PNA combines a scaling based on node degree with multiple aggregators, where the scaling based on node degree is used to determine the amount of information to be aggregated. The scaling based on node degree is multiplied by the aggregated value to increase or decrease the size of the features of the incoming message.
4. The water plant equipment operating condition prediction method based on multi-factor coupling according to claim 1, characterized in that, The process by which the Association Link Network (ALN) extracts contextual features from historical event triples is as follows: The event parameter is represented as e. i = {w1, w2, w3}, corresponding to the verb, subject, and object in the triple, respectively. The bag-of-words table Count = w1, w2, ..., w3 is used to determine the event chain. n In this process, the co-occurrence relationships of all words in historical event links are used to organize event contextual information, and an ALN model is constructed to extract contextual features to provide prior knowledge for event prediction. The basic components of ALN are semantic nodes and their associated connections, and the word w is calculated. i with w j Correlation weights between The formula is as follows: , in Co ( w i , w j ) indicates words w i and w j The frequency of simultaneous occurrence in the event chain DF ( w i ) indicates words w i Frequency of occurrence in the event chain, related link weight score ∈[0,1] indicates the strength of the association between words; By using the GCN algorithm to enhance the differentiation of node features, GCN automatically encodes the current node through its neighbors, thereby extracting features that enrich the graph structure of its neighbors: , in H (l) It is the first l The feature matrix of the layer, where σ is the ReLU activation function. Let I be the identity matrix and A be the adjacency matrix. A is obtained by using... Fill in the correlation scores between parameters in the event chain. yes The standardized degree matrix, W (l) It is the first l Layer weight matrix; After updating all node vectors in the ALN to convergence, the semantic embedding vector is computed. S i as follows: , in h i It is the vector representation of node i output by GCN. N G The number of nodes in ALN, for each event. e i Semantic information is used to learn new embedding representations by integrating the contextual embeddings of corresponding words, and events. e i Embedded representation after learning contextual information The calculation process is as follows: , in Represents the parameters of a triple. , and Weight matrices corresponding to different triplet parameters. This indicates the bias term.
5. The water plant equipment operating condition prediction method based on multi-factor coupling according to claim 1, characterized in that, In step S3, the feature selector uses the following formula: , Among them W and b i These are the weight matrix and the bias term, respectively. Indicates contextual features, P ( e i ) represents the probability of semantic features passing, when P ( e i If the feature value is ≥0.5, the feature will be used for fusion; otherwise, it will be discarded.
6. The water plant equipment operating condition prediction method based on multi-factor coupling according to claim 1, characterized in that, In step S3, in the improved transformer, contextual features are input into the attention-based feature fusion layer of the decoder after being encoded. Relationship features are also input into the attention-based feature fusion layer after passing through a multi-head self-attention layer, an addition and normalization layer. In the attention-based feature fusion layer, relationship features are split into historical event features and candidate event features according to event type. The attention distribution of the two features, contextual features and historical event features, is calculated using the softmax function. This attention distribution is used as the weight of the corresponding feature state, and a weighted average is calculated. This average is then input into the fully connected network MLP along with the split candidate event features to output the fused feature vector representation. Finally, after normalization and feedforward processing, the fused event representation is obtained.
7. The water plant equipment operating condition prediction method based on multi-factor coupling according to claim 1, characterized in that, In step S4, the fused event representation is split into historical events, candidate events, and historical event chains by event type. In the cross-attention calculation module, this includes the calculation of overall correlation and the calculation of local correlation. The calculation of local correlation involves first calculating the similarity between each triplet parameter of the candidate event and each historical event through MLP neural network layers, then normalizing the similarity features using softmax to obtain the weight coefficients of each triplet parameter of the candidate event and each historical event, and finally summing the weight coefficients of each triplet parameter of the candidate event to obtain the correlation score between each parameter of the candidate event and the historical event. The calculation of overall correlation involves first compressing the historical event chain features using average pooling to facilitate subsequent calculations, then using MLP and softmax to calculate the similarity between each triplet parameter of the candidate event and the historical event chain, and normalizing the similarity to obtain the weight relationship between each triplet parameter of the candidate event and the historical event chain, finally using the local calculated score and the overall correlation calculated coefficient to calculate the individual score of each triplet parameter of the candidate event, and obtaining the final score of the entire candidate event by summing the scores of each triplet parameter.
8. The water plant equipment operating condition prediction method based on multi-factor coupling according to claim 1, characterized in that, Step S5 learns the event embedding by minimizing the cross-entropy loss function L(θ): , Where n is the length of the historical event chain, S ( z , c ) represents the prediction score of the candidate event c corresponding to the historical event chain, log is the logarithmic function, and L2 is the L2 regularization parameter. λ Here, θ is the regularization coefficient, and θ is the model parameter.
Citation Information
Patent Citations
Narrative event prediction method fusing event environment information
CN113887836A
Candidate event prediction method based on gated graph neural network
CN117521810A