Water pump pan-mapping fault diagnosis method based on penalty gating neural network

Through the water pump pan-mapping fault diagnosis method based on penalty-gated neural network, the problems of low accuracy and poor adaptability of traditional water pump fault diagnosis under complex working conditions are solved, high-precision fault detection and model robustness are achieved, and the level of condition monitoring and intelligent maintenance of industrial equipment is improved.

CN120804815APending Publication Date: 2025-10-17BEIJING UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510882620.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Traditional water pump fault diagnosis methods have low diagnostic accuracy under complex working conditions, lack adaptability to diverse operating modes, and are unable to cope with data distribution drift caused by changes in working conditions and load fluctuations.

Method used

A pan-mapping fault diagnosis method for water pumps based on a penalty-gated neural network is adopted. By constructing a causal graph of water pump operation characteristics, combining a gating optimization mechanism, dynamically extracting key information, and establishing a long-short-term memory neural network, a complex nonlinear mapping relationship between characteristic variables and water pump fault status is realized.

Benefits of technology

The accuracy and robustness of water pump fault detection have been significantly improved, and it can maintain high fault diagnosis accuracy under various complex working conditions, ensuring the safe and efficient operation of industrial process equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120804815A_ABST
    Figure CN120804815A_ABST
Patent Text Reader

Abstract

The invention discloses a water pump flooding mapping fault diagnosis method based on a penalty gating neural network, and belongs to the field of industrial equipment fault detection and intelligent maintenance. In order to solve the problem that the diagnosis accuracy is reduced due to data distribution drift of an existing industrial water pump under complex working conditions, the invention designs a water pump fault diagnosis model combining causal feature processing and a penalty gating long-short-term memory network, and accurate judgment of multi-state and multi-distribution faults of the water pump is realized. Experimental results show that the method can effectively improve the accuracy of water pump fault recognition in the industrial process, and long-term stable operation of equipment is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application designs a water pump general mapping fault diagnosis method based on a penalty gate neural network, which realizes accurate identification of faults of the water pump in various operating states of industrial process equipment. As an important power equipment in industrial production process, the running state of the water pump directly affects the continuity and system stability of the process flow. Accurate diagnosis of water pump faults is of great significance for optimizing equipment operation and maintenance strategies and ensuring safe and efficient operation of production lines, and belongs to the field of industrial process equipment monitoring and intelligent operation and maintenance. BACKGROUND

[0002] In modern industrial production processes, water pumps are widely used in fluid transportation, pressure regulation and cooling circulation, etc. With the expansion of the scale of industrial systems and the increase of process complexity, water pumps, as one of the core equipment, bear the important task of stable transportation and system support. Once the water pump fails, it will often lead to a decrease in production efficiency, and even cause a chain of equipment failures, causing great economic losses. Therefore, real-time monitoring and fault warning of the operating state of the industrial process water pump are of great significance.

[0003] Traditional water pump fault diagnosis methods mainly rely on single parameter monitoring or fixed threshold determination, which is difficult to cope with the problem of data distribution drift caused by changes in working conditions, load fluctuations, environmental disturbances and other factors. This method has low diagnostic accuracy in actual complex working conditions and lacks adaptability to diversified operating modes. In recent years, with the development of machine learning and deep learning technology, the use of multi-dimensional feature extraction and data modeling methods provides a new idea for improving the fault diagnosis level of industrial equipment.

[0004] The application proposes a water pump general mapping fault diagnosis method based on a penalty gate neural network, which dynamically extracts the most critical information for fault discrimination by constructing a water pump operating feature causal graph combined with a gate optimization mechanism, can effectively adapt to the complex situation of multiple data distributions corresponding to the same fault type, significantly improve the accuracy and robustness of water pump fault detection, and provide strong protection for the safe and efficient operation of industrial process equipment. SUMMARY

[0005] The application obtains a water pump general mapping fault diagnosis method based on a penalty gate neural network, takes a water pump fault state as an evaluation variable, takes water pump temperature, motor voltage, motor temperature, motor current, bearing vibration, bearing pressure as characteristic variables, establishes a causal relationship of water pump operation data through a multi-dimensional characteristic processing layer, and embeds a long short-term memory neural network optimized by a penalty gate mechanism, so that a complex nonlinear mapping relationship between the characteristic variables and the water pump fault state is realized. The method effectively solves the problem of low fault recognition accuracy of the water pump under multiple distribution data caused by changes in operating conditions in the industrial process, can maintain high fault diagnosis accuracy and model robustness under various complex working conditions, and significantly improves the level of industrial equipment state monitoring and intelligent maintenance. The application adopts the following technical solutions and implementation steps:

[0006] A water pump general mapping fault diagnosis method based on a penalty gate neural network, characterized by determining characteristic variables and evaluation variables of a water pump fault state, establishing a fault diagnosis model of the water pump, training parameters of the diagnosis model, and evaluating the fault state of the water pump, comprising the following steps:

[0007] 1. Determine the characteristic variables and evaluation variables of the water pump fault state:

[0008] Taking a sewage treatment process as the research object, y(t) is determined as the output variable of the water pump fault diagnosis model at time t; the water pump temperature x1(t) has a value interval of [80℃, 90℃], ℃ represents degrees Celsius, the motor voltage x2(t) has a value interval of [380V, 480V], V represents volts, the motor temperature x3(t) has a value interval of [30℃, 100℃], the motor current x4(t) has a value interval of [250A, 350A], A represents amperes, the bearing vibration speed x5(t) has a value interval of [1mm / s, 5mm / s], mm represents millimeters, s represents seconds, the bearing pressure x6(t) has a value interval of [1000N / m 2 ,5000N / m 2 ], N represents Newton, m 2 represents square meters, x1(t), x2(t), x3(t), x4(t), x5(t), x6(t) are characteristic variables affecting y(t), the water pump fault state is an evaluation variable, and is represented as F(t), wherein F(t)=1 represents a fault, and F(t)=0 represents normal;

[0009] 2. Establish a fault diagnosis model of the water pump:

[0010] The water pump fault diagnosis model is composed of a feature processing layer and a long short-term memory neural network, the long short-term memory neural network is composed of 80 long short-term memory units connected in series, the long short-term memory unit is divided into three parts: an input layer, a forgetting layer and an output layer, X(t)=[x1(t), x2(t), x3(t), x4(t), x5(t), x6(t)] is a model input matrix;

[0011] The feature processing layer: this layer captures the association relationship in the input data, establishes a data causal graph, and the output of the multi-dimensional feature processing layer can be expressed as:

[0012] G k (t)={V k (t),E k (t)} (1)

[0013] V k (t)={x j (t),x m (t),S(t),y(t)} (2)

[0014] E k (t)={g jy (t),g jm (t)} (3)

[0015] g jy (t)={I(x j (t);y(t))|I(x j (t);x m (t))>θ(t)} (4)

[0016] g jm (t)={I(x j (t);x m (t))|I(x j (t);x m (t))>β(t)} (5)

[0017]

[0018] Wherein, G k (t) represents a data causal graph matrix, k represents the number of causal graph matrices, k=1, 2, …, 15, V k (t) represents a set matrix of nodes in the causal graph, x j (t) represents a parent node vector directly related to the fault, j represents the number of parent node vectors, j=1, 2, …, 20, x m (t) represents a child node vector directly related to the fault, m represents the number of child node vectors, m=1, 2, …, 50, s(t) represents a spouse node vector indirectly related to the fault, and xj (t) represents x j any variable value in (t), x m (t) represents x m any variable value in (t), s(t) represents any variable value in s(t), y(t) represents a variable value in y(t), E k (t) represents a set matrix of edges in a causal graph, g jm (t) represents a set vector of edge values between input variables, I(x j (t); x m (t)) represents x j (t) and x m (t) is the mutual information value of x j (t); x m (t)) is greater than a threshold value β(t), β(t) = 0.45, then the edge is considered to exist, g jm (t) = I(x j (t); x m (t)), g jy (t) represents a set vector of edge values between input variables and fault variables, I(x j (t); y(t)) represents x j (t) and y(t) is the mutual information value of x j (t); y(t)) is greater than a threshold value θ(t), θ(t) = 0.95, then the edge is considered to exist, g jy (t) = I(x j (t); y(t)), log is a logarithm;

[0019] P(y(t) \ x j (t), s(t)) ≠ P(y(t) \ s(t)) (8)

[0020] P(x m (t) \ y(t), s(t)) ≠ P(x m (t) \ s(t)) (9)

[0021] P(x j (t) \ x m (t), y(t)) ≠ P(x j (t) \ y(t)) (10)

[0022] where, as shown in equation (8), if x j (t) affects the conditional probability of y(t), then x j (t) is considered to be a parent node, as shown in equation (9), if the change of y(t) affects the conditional probability of x m (t), then x m(t) is a child node, as shown in equation (10), if the change of y(t) cannot affect the conditional probability of s(t), then s(t) is considered as a spouse node; the expression of LSTM unit is as follows:

[0023] Input layer: This layer filters out simple data causal patterns, and the output of the input layer can be expressed as:

[0024] i(t) = σ(a(t)) (11)

[0025]

[0026]

[0027]

[0028] where i(t) represents the output vector of the input layer at time t, a(t) represents the transformation vector at time t, e is an exponential constant, x(t) represents the input vector of the input layer at time t, x(t) ∈ G k (t), h(t-1) represents the output vector of the last LSTM unit at time t-1, h(0) = 0, U i (t) represents the input layer coefficient matrix, W i (t) represents the input layer weight matrix, GC(G k (t)) represents the influence value of G k (t) on y(t), g jy (t) represents a causal edge in G k (t), represents an influence factor value,

[0029] T = 100;

[0030] Forget layer: This layer forgets low-frequency data causal patterns and remembers high-frequency data causal patterns, and the operation of forgetting low-frequency can be expressed as:

[0031] f(t) = e(x(t)U f (t) + h(t-1)W f (t) - δ(t)ASim(G k (t)) (15)

[0032]

[0033] where f(t) represents the output vector of the forget gate, U f (t) represents the forget gate coefficient matrix, W f (t) represents the forget gate weight matrix, δ(t) represents an influence factor, δ(t) = 0.75, Gk1 (t) and G k2 (t) represent two different causal matrices, ASim(G k (t)) represents the average value of all inter-causal graph similarities, K is a constant value, Sim(G k1 (t), G k2 (t)) represents the similarity value between G k1 (t) and G k2 (t), x k1 (t) represents a node value in G k1 (t), g(t) represents a causal edge weight value in G k1 (t) and G k2 (t), The operation of memorizing high frequency can be represented as:

[0034] c(t) = f(t)c(t-1) + i(t)σ(W c (t)h(t-1) + U c (t)x(t) + γ(t)Sta(G k (t))Cov(G k (t)) (18)

[0035]

[0036] where c(t) represents the output vector of the cell state at time t, c(t-1) represents the output vector of the cell state at time t-1, U c (t) represents the cell state coefficient matrix, W c (t) represents the cell state weight matrix, γ(t) represents an influence factor, γ(t) = 1, Sta(G k (t)) represents the long-term similarity value of G k (t), Cov(G k (t)) represents the global coverage value of G k (t), G k is a fixed causal matrix, φ(G k (t) = G k ) is an indicator function, which returns a value of 1 when G k (t) = G k at any time step, and returns a value of 0 when G k (t) ≠ G k ;

[0037] Output layer: The output of the output layer can be represented as:

[0038]

[0039] o(t) = σ(x(t)U o (t) + h(t-1)W o (t))(22)

[0040] where h(t) represents the output vector of the output layer at time t, o(t) represents the output vector of the output gate at time t, U o (t) represents the coefficient matrix of the output gate, W o (t) represents the weight matrix of the output gate; the final output of the network is y(t) = h 80 (t), h 80 (t) represents the output of the 80th long short-term memory unit;

[0041] 3. Training the parameters of the diagnostic model:

[0042] ① The model input is X(1), X(2), …, X(t), …, X(M), M is the number of training samples, and the weight matrix is initialized as W(1) = [W i (1), W f (1), W c (1), W o (1)], W i (1) is the set of input layer weight matrices of all units at the first time, W f (1) is the set of all unit forget gate weight matrices at the first time, W c (1) is the set of all unit cell state weight matrices at the first time, W o (1) is the set of all unit output gate weight matrices at the first time, the coefficient matrix U(1) = [U i (1), U f (1), U c (1), U o (1)], U i (1) is the set of all unit input gate coefficient matrices at the first time, U f (1) is the set of all unit forget gate coefficient matrices at the first time, U c (1) is the set of all unit cell state coefficient matrices at the first time, U o (1) is the set of all unit output gate coefficient matrices at the first time, the elements in the W(1) matrix are randomly taken from the interval [0.01, 0.18], and the elements in the U(1) matrix are randomly taken from the interval [0.01, 0.1]; let the current learning step L = 1;

[0043] ② The loss function of the model is represented as:

[0044]

[0045] where, The true value representing the fault diagnosis model output is updated by using the gradient descent method to update the weight matrix and the coefficient matrix:

[0046]

[0047] Wherein, W(t) represents the weight matrix before model adjustment, U(t) represents the coefficient matrix before model adjustment, W(t+1) represents the weight matrix after model adjustment, U(t+1) represents the coefficient matrix after model adjustment, and alpha=0.1 represents the learning rate, The partial derivative of J(t) with respect to W(t) is represented as, The partial derivative of J(t) with respect to U(t) is represented as.

[0048] ③ If the learning step L is less than M, L is increased by 1, and step ② is continued to train, if L=M, stop calculation, and step 4 is converted.

[0049] 4. Evaluate the fault state of the water pump:

[0050] The trained fault diagnosis model index y(t) is used to diagnose whether the water pump has a fault, the normal value range of y(t) is [0.3, 1.2], if y(t) belongs to this range, F(t)=0, the water pump is normal, if y(t) exceeds this range, F(t)=1, the water pump has a fault.

[0051] The creativity of the application mainly embodies in:

[0052] 1. In view of the problems of low fault diagnosis accuracy and insufficient model robustness of industrial process water pump under multiple working conditions and multiple distribution conditions, the application provides a water pump general mapping fault diagnosis method based on a penalty gate neural network. By deeply analyzing the key indicators in the water pump operation process, such as water pump temperature, motor voltage, motor temperature, motor current, bearing vibration and bearing pressure, the six variables are selected as characteristic inputs, and a fault diagnosis model combining multi-dimensional feature processing and gate optimization is established, which effectively overcomes the limitations of low accuracy and poor adaptability of traditional single monitoring method under complex industrial conditions.

[0053] 2. Model design for data complexity and dynamic change: Considering that the industrial process water pump operation data presents high nonlinearity, dynamic change and drift characteristics, the traditional method is difficult to model stably, the application proposes a long short-term memory neural network design based on causal modeling and gate mechanism optimization, which dynamically selects fault correlation features and strengthens the memory of high-frequency causal patterns, improves the model's ability to capture changes in operating conditions and adaptability to distribution drift, and significantly improves the accuracy and anti-interference ability of fault diagnosis.

[0054] Particularly, it is to be noted that the present application introduces a penalty gate neural network architecture, and designs a diagnosis model in combination with key operating characteristics of a water pump. As long as the method and technology proposed by the present application are adopted to monitor and evaluate the multiple distributed faults of an industrial process water pump, it shall belong to the protection scope of the present application. BRIEF DESCRIPTION OF DRAWINGS

[0055] Figure 1 is a water pump fault diagnosis model diagram of the present application;

[0056] Figure 2 is a prediction result diagram of the model on faults of the present application; DETAILED DESCRIPTION

[0057] The present application obtains a water pump general mapping fault diagnosis method based on a penalty gate neural network. Key characteristic variables in a water pump operating process are selected as inputs, including water pump temperature, motor voltage, motor temperature, motor current, bearing vibration and bearing pressure.

[0058] The experimental data is derived from a water pump operating data acquisition system of an industrial enterprise, and covers water pump operating periods under different working conditions. Actual detection data of water pump temperature, motor voltage, motor temperature, motor current, bearing vibration and bearing pressure are extracted as experimental samples. After data cleaning and outlier elimination, 4500 groups of effective samples are reserved. All 4500 groups of data are divided in proportion, of which 3000 groups are used as training samples to train parameters of a water pump fault diagnosis model, and 1500 groups are used as test samples to evaluate fault recognition accuracy and robustness of the model.

[0059] The present application adopts the following technical solutions and implementation steps:

[0060] A water pump general mapping fault diagnosis method based on a penalty gate neural network, characterized by determining characteristic variables and evaluation variables of a water pump fault state, establishing a water pump fault diagnosis model, training parameters of the diagnosis model, and evaluating the fault state of the water pump, comprising the following steps:

[0061] 1. Determining characteristic variables and evaluation variables of a water pump fault state:

[0062] With sewage treatment process as the research object, the output variable y(t) of the pump fault diagnosis model at time t is determined; the value interval of the pump temperature x1(t) is [80℃, 90℃], ℃ represents Celsius, the value interval of the motor voltage x2(t) is [380V, 480V], V represents volt, the value interval of the motor temperature x3(t) is [30℃, 100℃], the value interval of the motor current x4(t) is [250A, 350A], A represents ampere, the value interval of the bearing vibration speed x5(t) is [1mm / s, 5mm / s], mm represents millimeter, s represents second, the value interval of the bearing pressure x6(t) is [1000N / m 2 ,5000N / m 2 ], N represents Newton, m 2 represents square meter, x1(t), x2(t), x3(t), x4(t), x5(t), x6(t) are characteristic variables affecting y(t), the pump fault state is an evaluation variable, represented as F(t), wherein F(t) = 1 represents fault, and F(t) = 0 represents normal;

[0063] 2. Establishing a pump fault diagnosis model:

[0064] The pump fault diagnosis model is composed of a feature processing layer and a long short-term memory neural network, the long short-term memory neural network is composed of 80 long short-term memory units connected in series, the long short-term memory unit is divided into three parts: an input layer, a forgetting layer and an output layer, and X(t) = [x1(t), x2(t), x3(t), x4(t), x5(t), x6(t)] is a model input matrix;

[0065] Feature processing layer: This layer captures the association relationship in the input data, establishes a data causal diagram, and the output of the multi-dimensional feature processing layer can be represented as:

[0066] G k (t)={V k (t),E k (t)}(26)

[0067] V k (t)={x j (t),x m (t),S(t),y(t)}(27)

[0068] E k (t)={g jy (t),g jm (t)}(28)

[0069] g jy (t)={I(x j (t);y(t))|I(x j (t);xm (t))>θ(t)}(29)

[0070] g jm (t) = {I(x j (t); x m (t))|I(x j (t); x m (t))>β(t)}(30)

[0071]

[0072] Among them, G k (t) represents the data causal graph matrix, k represents the number of causal graph matrices, k = 1, 2, ..., 15, V k (t) represents the set matrix of nodes in the causal graph, x j (t) represents the parent node vector directly related to the fault, j represents the number of parent node vectors, j = 1, 2, ..., 20, x m (t) represents the child node vector directly related to the fault, m represents the number of child node vectors, m = 1, 2, ..., 50, s(t) represents the spouse node vector indirectly related to the fault, x j (t) represents x j Any variable value in (t), x m (t) represents x m (t), s(t) represents any variable value in s(t), y(t) represents any variable value in y(t), E k (t) represents the edge set matrix in the causal graph, g jm (t) represents the set vector of boundary values ​​between input variables, I(x j (t); x m (t)) represents x j (t) and x m (t), if I(x j (t); x m (t)) is greater than the threshold β(t), β(t) = 0.45, then the edge is considered to exist, g jm (t) = I(x j (t); x m (t)), g jy (t) represents the set vector of boundary values ​​between input variables and fault variables, I(x j (t); y(t)) represents x j The mutual information value of (t) and y(t) is as follows: j (t); y(t)) is greater than the threshold θ(t), θ(t) = 0.95, then the edge is considered to exist, g jy (t) = I(x j(t); y(t)), log is logarithm;

[0073] P(y(t) | x j (t), s(t)) ≠ P(y(t) | s(t)) (33)

[0074] P(x m (t) | y(t), s(t)) ≠ P(x m (t) | s(t)) (34)

[0075] P(x j (t) | x m (t), y(t)) ≠ P(x j (t) | y(t)) (35)

[0076] Where, as shown in equation (8), if x j (t) affects the conditional probability of y(t), then x j (t) is considered as a parent node, as shown in equation (9), if the change of y(t) affects the conditional probability of x m (t), then x m (t) is considered as a child node, as shown in equation (10), if the change of y(t) cannot affect the conditional probability of s(t), then s(t) is considered as a spouse node; the expression of long short-term memory unit is as follows:

[0077] Input layer: this layer filters out simple data causal patterns, and the output of the input layer can be expressed as:

[0078] i(t) = σ(a(t)) (36)

[0079]

[0080] Where, i(t) represents the output vector of the input layer at time t, a(t) represents the conversion vector at time t, e is an exponential constant, x(t) represents the input vector of the input layer at time t, x(t) ∈ G k (t), h(t-1) represents the output vector of the last long short-term memory unit at time t-1, h(0) = 0, U i (t) represents the input layer coefficient matrix, W i (t) represents the input layer weight matrix, GC(G k (t)) represents the influence value of the causal graph G k (t) on y(t) at time t, g jy (t) represents the weight value of a causal edge in G k (t), represents an influence factor value,

[0081] T = 100;

[0082] forgetting low-frequency data causal patterns and remembering high-frequency data causal patterns. The operation of forgetting low-frequency data causal patterns can be represented as:

[0083] f(t) = e(x(t)U f (t) + h(t-1)W f (t) - δ(t)ASim(G k (t)) (40)

[0084]

[0085] where f(t) represents the output vector of the forgetting gate, U f (t) represents the forgetting gate coefficient matrix, W f (t) represents the forgetting gate weight matrix, δ(t) represents an influence factor, δ(t) = 0.75, G k1 (t) and G k2 (t) represent two different causal matrices, ASim(G k (t)) represents the average value of the similarity between all causal graphs, K is a constant value, Sim(G k1 (t), G k2 (t)) represents the similarity value between G k1 (t) and G k2 (t), x k1 (t) represents a node value in G k1 (t), g(t) represents a causal edge weight value in G k1 (t) and G k2 (t), represents the intersection operation, and the operation of remembering high-frequency data causal patterns can be represented as:

[0086] c(t) = f(t)c(t-1) + i(t)σ(W c (t)h(t-1) + U c (t)x(t) + γ(t)Sta(G k (t))Cov(G k (t)) (43)

[0087]

[0088]

[0089] where c(t) represents the output vector of the cell state at time t, c(t-1) represents the output vector of the cell state at time t-1, U c (t) represents the cell state coefficient matrix, W c(t) represents a cell state weight matrix, γ(t) represents an influence factor, γ(t) = 1, Sta(G k (t) represents G k (t) represents a long-term similarity value of G k (t) represents G k (t) represents a global coverage value of G k is a fixed causal matrix, φ(G k (t) = G k is an indicator function, at any time step, when G k (t) = G k , returns a value of 1, when G k (t) ≠ G k , returns a value of 0;

[0090] Output layer: the output of the output layer can be represented as:

[0091]

[0092] o(t) = σ(x(t)U o (t) + h(t-1)W o (t)) (47)

[0093] Wherein, h(t) represents an output vector of the output layer at time t, o(t) represents an output vector of the output gate at time t, U o (t) represents a coefficient matrix of the output gate, W o (t) represents a weight matrix of the output gate; the final output of the network is y(t) = h 80 (t), h 80 (t) represents an output of the 80th long short-term memory unit;

[0094] 3. Training parameters of the diagnostic model:

[0095] ① The model input is X(1), X(2), …, X(t), …, X(M), M is the number of training samples, the weight matrix is initialized as W(1) = [W i (1), W f (1), W c (1), W o (1)], W i (1) is a set of input layer weight matrices of all units at the first time, W f (1) is a set of forget gate weight matrices of all units at the first time, W c (1) is a set of cell state weight matrices of all units at the first time, W o (1) is a set of output gate weight matrices of all units at the first time, the coefficient matrix is U(1) = [U i (1), Uf (1),U c (1),U o (1)],U i (1) is the set of all unit input gate coefficient matrices at the first moment, U f (1) is the set of all unit forget gate coefficient matrices at the first moment, U c (1) is the set of all unit cell state coefficient matrices at the first moment, U o (1) is the set of all unit output gate coefficient matrices at the first moment, W(1) matrix elements are randomly taken from the interval [0.01, 0.18], U(1) matrix elements are randomly taken from the interval [0.01, 0.1]; Let the current learning step L = 1;

[0096] ②The loss function of the model is represented as:

[0097]

[0098] Wherein, The real value output by the fault diagnosis model is represented by the gradient descent method to update the weight matrix and the coefficient matrix:

[0099]

[0100] Wherein, W(t) represents the weight matrix before the model is adjusted, U(t) represents the coefficient matrix before the model is adjusted, W(t+1) represents the weight matrix after the model is adjusted, U(t+1) represents the coefficient matrix after the model is adjusted, and α = 0.1 represents the learning rate, The partial derivative of J(t) with respect to W(t) is represented by The partial derivative of J(t) with respect to U(t) is represented by

[0101] ③If the learning step L < M, L increases by 1, and step ② is continued to train, if L = M, stop calculation, and go to step 4;

[0102] 4. Evaluate the fault state of the water pump:

[0103] Use the trained fault diagnosis model index y(t) to diagnose whether the water pump has a fault, the normal value range of y(t) is [0.3, 1.2], if y(t) belongs to this range, F(t) = 0, the water pump is normal, if y(t) exceeds this range, F(t) = 1, the water pump has a fault.

Claims

1. A water pump pan-mapping fault diagnosis method based on penalty-gated neural network, characterized in that: It includes the following steps: 1). Determine the characteristic variables and evaluation variables of the pump failure state: Taking the sewage treatment process as the research object, y(t) is determined as the output variable of the water pump fault diagnosis model at time t; the water pump temperature x1(t) has a value range of [80℃, 90℃], ℃ represents degrees Celsius, the motor voltage x2(t) has a value range of [380V, 480V], V represents volts, the motor temperature x3(t) has a value range of [30℃, 100℃], the motor current x4(t) has a value range of [250A, 350A], A represents amperes, the bearing vibration speed x5(t) has a value range of [1mm / s, 5mm / s], mm represents millimeters, s represents seconds, and the bearing pressure x6(t) has a value range of [1000N / m 2 ,5000N / m 2 ], N stands for Newton, m 2 represents square meters, x1(t), x2(t), x3(t), x4(t), x5(t), x6(t) are characteristic variables that affect y(t), and the fault state of the pump is the evaluation variable, expressed as F(t), where F(t) = 1 indicates fault and F(t) = 0 indicates normal; 2). Establish a fault diagnosis model for the pump: The pump fault diagnosis model consists of a feature processing layer and a long short-term memory neural network. The long short-term memory neural network is composed of 80 long short-term memory units connected in series. The long short-term memory units are divided into three parts: an input layer, a forgetting layer, and an output layer. X(t) = [x1(t), x2(t), x3(t), x4(t), x5(t), x6(t)] is the model input matrix; Feature processing layer: This layer captures the correlation relationships in the input data and establishes a data causal graph. The output of the multi-dimensional feature processing layer is expressed as: G k (t)={V k (t),E k (t)}(1) V k (t)={x j (t),x m (t),S(t),y(t)} (2) E k (t)={g jy (t),g jm (t)} (3) g jy (t)={I(x j (t);y(t))|I(x j (t);x m (t))>θ(t)} (4) g jm (t)={I(x j (t);x m (t))|I(x j (t);x m (t))>β(t)} (5) Among them, G k (t) represents the data causal graph matrix, k represents the number of causal graph matrices, k = 1, 2, ..., 15, V k (t) represents the set matrix of nodes in the causal graph, x j (t) represents the parent node vector directly related to the fault, j represents the number of parent node vectors, j = 1, 2, ..., 20, x m (t) represents the child node vector directly related to the fault, m represents the number of child node vectors, m = 1, 2, ..., 50, s(t) represents the spouse node vector indirectly related to the fault, x j (t) represents x j Any variable value in (t), x m (t) represents x m (t), s(t) represents any variable value in s(t), y(t) represents any variable value in y(t), E k (t) represents the edge set matrix in the causal graph, g jm (t) represents the set vector of boundary values ​​between input variables, I(x j (t); x m (t)) represents x j (t) and x m (t), if I(x j (t); x m (t)) is greater than the threshold β(t), β(t) = 0.45, then the edge is considered to exist, g jm (t) = I(x j (t); x m (t)), g jy (t) represents the set vector of boundary values ​​between input variables and fault variables, I(x j (t); y(t)) represents x j The mutual information value of (t) and y(t) is as follows: j (t); y(t)) is greater than the threshold θ(t), θ(t) = 0.95, then the edge is considered to exist, g jy (t) = I(x j (t); y(t)), log is the logarithm; P(y(t)|x j (t),s(t))≠P(y(t)|s(t)) (8) P(x m (t)|y(t),s(t))≠P(x m (t)|s(t)) (9) P(x j (t)|x m (t),y(t))≠P(x j (t)|y(t)) (10) Where, as shown in formula (8), if x j (t) affects the conditional probability of y(t), then x j (t) is the parent node, as shown in formula (9), if the change of y(t) affects x m (t), then we assume that x m (t) is a child node. As shown in formula (10), if the change of y(t) cannot affect the conditional probability of s(t), then s(t) is considered to be a spouse node. The expression of the long short-term memory unit is as follows: Input layer: This layer filters out relatively simple data causal patterns. The output of the input layer is expressed as: i(t) = σ(a(t)) (11) Among them, i(t) represents the output vector of the input layer at time t, a(t) represents the conversion vector at time t, e is the exponential constant, x(t) represents the input vector of the input layer at time t, x(t)∈G k (t), h(t-1) represents the output vector of a long short-term memory unit at time t-1, h(0)=0, U i (t) represents the input layer coefficient matrix, W i (t) represents the input layer weight matrix, GC(G k (t)) represents the causal graph G at time t k The impact of (t) on y(t), g jy (t) represents G k The weight of a causal edge in (t), Represents an impact factor value, T = 100; Forgetting layer: This layer forgets low-frequency data causal patterns and remembers high-frequency data causal patterns. The operation of forgetting low-frequency is expressed as: f(t)=e(x(t)U f (t)+h(t-1)W f (t)-δ(t)ASim(G k (t)) (15) Among them, f(t) represents the output vector of the forget gate, U f (t) represents the forget gate coefficient matrix, W f (t) represents the forget gate weight matrix, δ(t) represents an impact factor, δ(t) = 0.75, G k1 (t) and G k2 (t) represents two different causal matrices, ASim(G k (t)) represents the average similarity between all causal graphs, K is a constant value, Sim(G k1 (t),G k2 (t)) represents G k1 (t) and G k2 (t), the similarity value between k1 (t) represents G k1 (t), g(t) represents a node value in G k1 (t) and G k2 A causal edge weight value in (t), ∩ represents the intersection operation, and the high-frequency operation of memory is expressed as: c(t)=f(t)c(t-1)+i(t)σ(W c (t)h(t-1)+U c (t)x(t)+γ(t)Sta(G k (t))Cov(G k (t)))(18) Among them, c(t) represents the output vector of the cell state at time t, c(t-1) represents the output vector of the cell state at time t-1, and U c (t) represents the cell state coefficient matrix, W c (t) represents the cell state weight matrix, γ(t) represents an influence factor, γ(t) = 1, Sta(G k (t)) represents G k (t) long-term similarity value, Cov(G k (t)) represents G k The global coverage value of (t), G k is a fixed causal matrix, φ(G k (t) = G k ) is the indicator function. At any time step, when G k (t) = G k When G k (t)≠G k When , the return value is 0; Output layer: The output of the output layer is expressed as: o(t)=σ(x(t)U o (t)+h(t-1)W o (t)) (22) Among them, h(t) represents the output vector of the output layer at time t, o(t) represents the output vector of the output gate at time t, and U o (t) represents the coefficient matrix of the output gate, W o (t) represents the weight matrix of the output gate; the final output of the network is y(t) = h 80 (t), h 80 (t) represents the output of the 80th LSTM unit; 3). Train the parameters of the diagnosis model: 4). Evaluate the fault state of the pump: Use the trained fault diagnosis model index y(t) to diagnose whether there is a fault in the pump. The normal value range of y(t) is [0.3, 1.2]. If y(t) belongs to this range, then F(t) = 0, and the pump is normal. If y(t) exceeds this range, then F(t) = 1, and the pump has a fault.

2. The method according to claim 1, wherein: ① The model input is X(1), X(2),…, X(t),…, X(M), where M is the number of training samples. The weight matrix W(1) is initialized as [W i (1),W f (1),W c (1),W o (1)],W i (1) is the set of weight matrices of all unit input layers at the first moment, W f (1) is the set of forget gate weight matrices of all units at the first moment, W c (1) is the set of all unit cell state weight matrices at the first moment, W o (1) is the set of weight matrices of all unit output gates at the first moment, and the coefficient matrix U(1) = [U i (1),U f (1),U c (1),U o (1)],U i (1) is the set of all unit input gate coefficient matrices at the first moment, U f (1) is the set of forget gate coefficient matrices of all units at the first moment, U c (1) is the set of all unit cell state coefficient matrices at the first moment, U o (1) is the set of all unit output gate coefficient matrices at the first moment. The elements in the W(1) matrix are randomly selected from the interval [0.01, 0.18], and the elements in the U(1) matrix are randomly selected from the interval [0.01, 0.1]. Let the current learning step number L = 1; ② The loss function of the model is expressed as: in, Represents the true value of the fault diagnosis model output, and uses the gradient descent method to update the weight matrix and coefficient matrix: Among them, W(t) represents the weight matrix before model adjustment, U(t) represents the coefficient matrix before model adjustment, W(t+1) represents the weight matrix after model adjustment, U(t+1) represents the coefficient matrix after model adjustment, α=0.1 represents the learning rate, represents the partial derivative of J(t) with respect to W(t), represents the partial derivative of J(t) with respect to U(t); ③ If the number of learning steps L < M, L is increased by 1, and go back to step ② to continue training. If L = M, stop the calculation and go to step 4).