Air handling unit fault diagnosis method, system, equipment and medium
By using a teacher-student model architecture and a semi-supervised deep learning framework, and by selecting generalization features using Shapley values and normalized weights, the problems of high cost of labeled data and insufficient model generalization ability caused by equipment heterogeneity in the fault diagnosis of air handling units are solved, and efficient and accurate fault identification is achieved.
Patent Information
- Application Number
- CN202510773937.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-11-07
AI Technical Summary
Existing technologies for fault diagnosis of air handling units suffer from problems such as high cost of acquiring labeled data, insufficient model generalization ability due to equipment heterogeneity, and low model efficiency caused by redundancy of high-dimensional sensor data. In particular, they are not effective in diagnosing intermittent and compound faults.
A teacher-student model architecture is adopted, and the temporal features of fault data are extracted through a temporal convolutional network. The generalization features are selected by using Shapley value and normalized weights. Combined with a semi-supervised deep learning framework, the teacher network parameters are updated using the exponential moving average algorithm to achieve fault diagnosis.
It improves the accuracy and generalization ability of fault diagnosis, reduces the dependence on labeled data, enhances the efficiency and adaptability of the model, and can effectively identify early minor faults.
Smart Images

Figure CN120910677A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of HVAC system intelligent operation and maintenance, in particular to an air handling unit fault diagnosis method, system, device and medium. BACKGROUND
[0002] In recent years, although the fault diagnosis method based on machine learning has made progress in theoretical research, it still faces challenges in engineering practice: first, the multi-sensor network (including temperature, pressure, flow, etc. monitoring module) configured in the modern AHU system (i.e. air handling unit system) generates high-dimensional time series data stream. The traditional manual labeling method needs engineers to analyze data correlation combined with device logs and maintenance records, which has the defects of low efficiency and strong subjectivity. Especially for complex patterns such as intermittent faults (such as occasional leakage of valves) and compound faults (multiple components concurrent anomalies), the labeling process is prone to feature misjudgment, which directly affects the quality of model training. New production equipment is more likely to face the "cold start" dilemma due to the lack of historical fault data accumulation, resulting in insufficient reliability of the diagnosis model in the initial deployment stage. Second, AHU of different manufacturers and models has significant differences in mechanical structure, control system, etc., resulting in strong equipment dependency of the diagnosis model constructed by the traditional feature engineering method.
[0003] The current three mainstream technical solutions have certain deficiencies in dealing with the above challenges. Among them, the rule-based fault diagnosis method determines whether the current running state of the AHU exceeds the threshold by setting empirical rules. This method lacks adaptability to the dynamic changes of the equipment running state and cannot effectively identify early minor faults. The model-based fault diagnosis method compares the actual output with the model output to identify faults by constructing a mechanism model. This method relies on a large amount of prior knowledge and is difficult to generalize to AHU equipment of different structures. The deep learning-based fault diagnosis method identifies different types of faults by training a deep learning model, which shows high classification accuracy in laboratory environment, but faces problems such as insufficient labeled samples and weak generalization ability in actual deployment. SUMMARY
[0004] In order to solve the problems of high cost of labeled data acquisition, insufficient model generalization ability caused by equipment heterogeneity, and low model efficiency caused by high-dimensional sensor data redundancy in the prior art air handling unit fault diagnosis, the present application proposes an air handling unit fault diagnosis method, system, device and medium.
[0005] The present application provides an air handling unit fault diagnosis method, comprising:
[0006] From the fault data to be diagnosed, select the fault data to be diagnosed corresponding to the preselected generalization features as input data;
[0007] Inputting the input data into the trained teacher-student model obtains a fault diagnosis result;
[0008] The teacher-student model extracts time sequence features of the fault data by using a time sequence convolution network, and is trained by taking historical input data as input and corresponding fault diagnosis results as output; the generalization features are selected by calculating the Shapley values and normalized weights of the sensor features.
[0009] Preferably, the selection process of the generalization features comprises:
[0010] The Shapley values of the sensor features in the fault data are calculated, and a plurality of first candidate features of each integrated model are selected from the sensor features in descending order of the Shapley values;
[0011] The normalized weights of the first candidate features of each integrated model are calculated, and the first candidate features are ranked in descending order of the weights, and the first quantity of the first candidate features are taken as the high-weight candidate feature sets of each integrated model;
[0012] The data corresponding to the high-weight candidate feature sets are selected from the target air handling unit data set as input, and the corresponding fault type classification results and sensor feature scores are taken as output to retrain the plurality of integrated models;
[0013] The Shapley values of the sensor features in the high-weight candidate feature sets of the retrained plurality of integrated models are calculated, and a plurality of second candidate features of each integrated model are selected from the first candidate features in descending order of the Shapley values;
[0014] The normalized weights of the second candidate features of each integrated model are calculated, and the second candidate features are ranked in descending order of the weights again, and the second quantity of the second candidate features are taken as the generalization features of each integrated model;
[0015] The plurality of integrated models are trained by taking the fault data in the source air handling unit data set as input and the fault type classification results and sensor feature scores as output; the target air handling unit data set is part of the source air handling unit data set; the first quantity is greater than the second quantity; the number of the first candidate features is greater than the number of the second candidate features.
[0016] Preferably, the fault data is TxF dimension, T is the sampling time, and F is the sensor feature.
[0017] Preferably, the plurality of integrated models comprise at least one or more of the following: a random forest model, an extreme gradient boosting tree model, a light gradient boosting machine model, or a category boosting model.
[0018] Preferably, the Shapley value is determined according to the following formula:
[0019]
[0020] wherein, is the Shapley value, denotes the subset of S excluding sensor feature x j , |S| denotes the number of elements in set S, x denotes a sensor feature, j denotes a sensor feature number, n denotes the number of sensor features, f x (S j ∪{x j}) is the sensor feature score including x j , f x (S j ) is the sensor feature score excluding x j .
[0021] Preferably, the normalized weight is determined by the following formula:
[0022]
[0023] wherein, is the normalized weight of the lth sensor feature of the χth integrated model, r l denotes the ranking of the lth sensor feature based on the sensor feature contribution degree of the Shapley value, v χ denotes the score of the χth model, and λ denotes the number of models.
[0024] Preferably, the training of the teacher-student model comprises:
[0025] A1, dividing the target air handling unit data set corresponding to the generalization feature into a training data set and a validation data set;
[0026] A2, training the student network in the teacher-student model using the labeled data in the training data set, and updating the parameters of the teacher network in the teacher-student model according to the parameters of the student network using the exponential moving average update algorithm;
[0027] A3, inputting the unlabeled data in the training data set into the updated teacher network to obtain pseudo-labeled data, training the student network using the pseudo-labeled data, and updating the parameters of the teacher network according to the parameters of the trained student network using the exponential moving average update algorithm;
[0028] A4, repeating step A3 until the performance of the validation data set no longer improves;
[0029] wherein, the parameters of the student network are updated during the training process by calculating the cross-entropy loss; and the initial value of the parameters of the teacher network is obtained by random initialization.
[0030] Preferably, the parameters of the teacher network in the teacher-student model are updated by using an exponential moving average update algorithm, which is implemented according to the following formula:
[0031] θ teacher = α·θ teacher + (1-α)·θ student
[0032] wherein θ teacher is the parameter of the teacher network, θ student is the parameter of the student network, and α is a smoothing hyperparameter.
[0033] Preferably, the inputting of the input data into the trained teacher-student model to obtain the fault diagnosis result comprises:
[0034] inputting the input data into the student network in the trained teacher-student model, and extracting the time sequence features of the input data through the time sequence convolution network body in the student network;
[0035] processing the time sequence features through the residual connection and batch normalization in the student network to obtain processed time sequence features;
[0036] inputting the processed time sequence features into the teacher network in the trained teacher-student model to obtain the fault diagnosis result.
[0037] Based on the same inventive concept, the present application also provides an air handling unit fault diagnosis system, comprising: a selected input data module and a diagnosis module.
[0038] The selected input data module is used to select the fault data to be diagnosed corresponding to the preselected generalization features from the fault data to be diagnosed as input data.
[0039] The diagnosis module is used to input the input data into the trained teacher-student model to obtain the fault diagnosis result.
[0040] The teacher-student model uses the time sequence convolution network to extract the time sequence features of the fault data, and is trained by using the historical input data as input and the corresponding fault diagnosis result as output; and the generalization features are selected by calculating the Shapley values and normalized weights of the sensor features.
[0041] Preferably, the selection process of the generalization features comprises:
[0042] calculating the Shapley values of the sensor features in the fault data, and selecting a plurality of first candidate features of the respective integrated models from the sensor features in the order from large to small according to the Shapley values;
[0043] The normalized weights of the first candidate features of the plurality of ensemble models are calculated, and the first candidate features are ranked according to the weights from large to small, and the first quantity of the first candidate features are the high-weight candidate feature sets of the ensemble models;
[0044] The data corresponding to the high-weight candidate feature set is selected from the target air handling unit data set as input, and the corresponding fault type classification result and sensor feature score are retrained as output to retrain the plurality of ensemble models;
[0045] The Shapley values of the sensor features in the high-weight candidate feature set of the retrained plurality of ensemble models are calculated, and the second candidate features of the plurality of ensemble models are selected from the first candidate features according to the order of the Shapley values from large to small;
[0046] The normalized weights of the second candidate features of the plurality of ensemble models are calculated, and the second candidate features are ranked again according to the weights from large to small, and the second quantity of the second candidate features are the generalization features of the ensemble models;
[0047] The plurality of ensemble models are trained by taking the fault data in the source air handling unit data set as input, and the fault type classification result and the sensor feature score as output; the target air handling unit data set is part of the source air handling unit data set; the first quantity is greater than the second quantity; the number of the first candidate features is greater than the number of the second candidate features.
[0048] Preferably, the fault data in the input data selection module is TxF dimension, T is the sampling time, and F is the sensor feature.
[0049] Preferably, the plurality of ensemble models in the input data selection module include at least one or more of the following: random forest model, extreme gradient boosting tree model, light gradient boosting machine model, or category boosting model.
[0050] Preferably, the Shapley value in the input data selection module is determined by the following formula:
[0051]
[0052] In the formula, is the Shapley value, indicates that the subset of the set S does not contain the sensor feature x j , |S| indicates the number of elements in the set S, x is the sensor feature, j is the sensor feature number, n is the number of sensor features, and f x (S j ∪{x j}) is the sensor feature score containing x j x (Sj ) is a sensor feature score not containing x j .
[0053] Preferably, the normalized weight in the selected input data module is determined by the following formula:
[0054]
[0055] wherein, is the normalized weight of the lth sensor feature of the χth integrated model, r l represents the ranking of the sensor feature contribution degree of the lth sensor feature based on the Shapley value, v χ represents the score of the χth model, and λ represents the number of models.
[0056] Preferably, the training of the teacher-student model in the diagnosis module comprises:
[0057] A1, dividing the target air handling unit data set corresponding to the generalized feature into a training data set and a validation data set;
[0058] A2, training the student network in the teacher-student model using the labeled data in the training data set, and updating the parameters of the teacher network in the teacher-student model according to the parameters of the student network using the exponential moving average update algorithm;
[0059] A3, inputting the unlabeled data in the training data set into the updated teacher network to obtain pseudo-labeled data, training the student network using the pseudo-labeled data, and updating the parameters of the teacher network according to the parameters of the trained student network using the exponential moving average update algorithm;
[0060] A4, repeating step A3 until the performance of the validation data set no longer improves;
[0061] Wherein, the parameters of the student network are updated during the training process by calculating the cross-entropy loss; and the initial value of the parameters of the teacher network is obtained by random initialization.
[0062] Preferably, the exponential moving average update algorithm is used to update the parameters of the teacher network in the teacher-student model, and is implemented by the following formula:
[0063] θ teacher = α·θ teacher + (1-α)·θ student
[0064] Wherein, θ teacher is the parameter of the teacher network, θ student is the parameter of the student network, and α is a smoothing hyperparameter.
[0065] Preferably, the diagnosis module is used for:
[0066] inputting the input data into the student network in the trained teacher-student model, extracting time sequence features of the input data through a time sequence convolution network body in the student network;
[0067] processing the time sequence features through residual connection and batch normalization in the student network to obtain processed time sequence features;
[0068] transferring the processed time sequence features to an output layer corresponding to the number of fault categories in the student network to obtain a probability distribution of each fault category as a fault diagnosis result.
[0069] Based on the same inventive concept, the present application also provides a computer device, comprising: at least one processor and a memory; the memory and the processor are connected through a bus;
[0070] the memory is used for storing one or more programs;
[0071] when the one or more programs are executed by the at least one processor, the air handling unit fault diagnosis method described above is realized.
[0072] Based on the same inventive concept, the present application also provides a computer readable storage medium, which has an execution program stored thereon, and the execution program is executed to realize the air handling unit fault diagnosis method described above.
[0073] Compared with the prior art, the present application has the following advantages:
[0074] The present application provides an air handling unit fault diagnosis method, system, device and medium, comprising: selecting the fault data to be diagnosed corresponding to the pre-selected generalization features as input data from the fault data to be diagnosed; inputting the input data into the trained teacher-student model to obtain the fault diagnosis result; wherein the teacher-student model uses a time sequence convolution network to extract the time sequence features of the fault data, and is trained with historical input data as input and corresponding fault diagnosis results as output; the generalization features are selected by calculating the Shapley value and normalized weight of the sensor features. The present application improves the low efficiency problem of the traditional method caused by relying on full-amount labeled data, device-specific feature set and high-dimensional redundant input by using the generalization feature selection mechanism and the semi-supervised deep learning framework. BRIEF DESCRIPTION OF DRAWINGS
[0075] Figure 1 is a flowchart of the air handling unit fault diagnosis method of the present application;
[0076] Figure 2 is a flowchart of an embodiment of the air handling unit fault diagnosis method of the present application; is a flowchart of an embodiment of the air handling unit fault diagnosis method of the present application;
[0077] Figure 3 Fig. 1 is a schematic diagram of an air handling unit fault diagnosis system according to the present application;
[0078] Figure 4 Fig. 2 is a schematic diagram of an air handling unit fault diagnosis computer device according to the present application. DETAILED DESCRIPTION
[0079] The specific embodiments of the present application will be further described in detail below with reference to the accompanying drawings.
[0080] Embodiment 1
[0081] The present application provides an air handling unit fault diagnosis method, the flow chart is as shown in Figure 1 Fig. 1, comprising:
[0082] S1, from the fault data to be diagnosed, select the fault data to be diagnosed corresponding to the preselected generalized feature as input data;
[0083] S2, inputting the input data into the trained teacher-student model to obtain the fault diagnosis result;
[0084] Wherein, the teacher-student model extracts the time sequence features of the fault data to be diagnosed by using time sequence convolution network, and is trained by taking the input data as input and the fault diagnosis result as output.
[0085] Wherein, the generalized feature is selected by calculating the Shapley value and the normalized weight of the sensor feature.
[0086] The generalized feature selection process in step S1 is as follows:
[0087] 1. Candidate feature screening: On the source AHU data set, four integrated models of Random forest (RF), Extreme Gradient Boosting (XGBoost), Light Gradient Boosting Machine (LightBGM) and Categorical Boosting (CatBoost) are trained respectively. SHAP (Shapley Additive Explanations) value (i.e. Shapley value) analysis is used to quantify the contribution of each feature to the model classification result, and the feature importance ranking is generated. Through the calculation of normalized weight, the ranking results of the four models are integrated, and the top 10 high weight candidate feature set is selected.
[0088] This step aims to calculate the feature importance using Shapley Additive exPlanations (SHAP) value analysis by integrating multiple machine learning models, including Random forest (RF), Extreme Gradient Boosting (XGBoost), Light Gradient Boosting Machine (LightBGM), and Categorical Boosting (CatBoost). The top 10 high-weight candidate features are selected through the normalized feature weight calculation method.
[0089] The input of the 4 tree-based models is the fault table data, as shown in Table 1, which is TxF dimensional, where T represents the sampling time and F represents the sensor features.
[0090]
[0091] Table 1
[0092] The model output is the classification result, F1 score. The SHAP method can be used to analyze the contribution of each sensor feature to the model output using the trained model. The classification result is the fault type result. The model takes the input (time x feature) and outputs the fault type, so it is considered a multi-classification task. The score is F1-score:
[0093]
[0094] RF is an ensemble learning method based on decision trees, which is an implementation of Bagging (Bootstrap Aggregating). It improves the accuracy and stability of the model by constructing multiple independent decision trees and aggregating their prediction results (majority vote for classification tasks, average for regression tasks).
[0095] XGBoost is an optimization implementation based on gradient boosting, which trains weak learners (usually decision trees) one by one and adjusts the training direction of subsequent trees according to the error of the previous tree, and finally combines the results of all trees.
[0096] LightGBM is a high-efficiency gradient boosting framework developed by Microsoft, aiming to solve the training speed and memory occupation problem in big data scenarios. It is also based on gradient boosting of decision trees, but uses innovative optimization strategies.
[0097] CatBoost is a gradient boosting algorithm developed by Yandex, which focuses on handling categorical features and is optimized for performance and ease of use.
[0098] SHAP model (i.e., Shapley additive explanations) uses game theory methods to calculate SHAP values, which quantify the contribution of each feature to the model's prediction, taking into account all possible feature combinations and their impact on the model's contribution, ensuring fair calculation of each feature's contribution. SHAP model belongs to a class of additive feature attribution methods, where the sum of the contributions of each feature is equal to the total prediction value of the model. This is taken as a simple and transparent explanation model of the original model. If f is the original fault diagnosis model, and g is the explanation model, the SHAP model is a binary linear function, as shown in equation (1):
[0099]
[0100] In the formula, P is the number of input features, j represents the number of each input feature, from the first feature to the Pth feature, z′ j ∈{0,1} P To simplify the features, determine which features in P are on the sample decision path, i.e., whether these features are used to make predictions, their values can only be 0 or 1. n is the number of models. SHAP value, used to quantify the impact of features on the model output. represents a subset that does not contain feature x j . |S| represents the number of elements in set S, f x (S j ∪{x j}) and f x (S j ) are sensor feature scores containing x j and not containing x j .
[0101] g(z′) is the predicted output of the model under the influence of these considered features; z′ indicates which features are considered in the current sample.
[0102] Specifically: z′∈{0,1} P is a simplified feature vector. It is used to determine which of the P input features are present in the decision path of the sample. The value of each element in z′ can only be 0 or 1. 0 indicates that the feature does not exist, so it will not affect the attribution of the sample. 1 indicates that the feature exists. Where P is the number of input features.
[0103] g(z′) is a function calculated according to the formula. φ0 represents the average prediction of all samples. φ iis the Shapley value of each feature, which quantifies the impact of the feature on the model output.
[0104] Therefore, g(z') can be understood as the output value of the model when the feature presence state is determined by z'. It is calculated by multiplying the Shapley value of each existing feature (i.e., the Shapley value) with its corresponding z' value (0 or 1), and then adding the prediction mean.
[0105] Since the four models adopted by the present application are all tree-based machine learning models, the Tree SHAP method is used to calculate and identify the top 10 features ranked by the weight of each model. Next, the normalized feature weight is calculated, and the 10 features obtained by the four models are ranked, so that the features with higher importance ranking and repeated occurrence have greater weight value, and finally a set of key feature set is obtained, and the calculation method is shown in formula (2):
[0106]
[0107] In the formula, is the normalized weight of the lth (l = 1, 2,..., 10) feature of the χth (χ = 1, 2, 3, 4) set, and r l represents the importance ranking of the feature based on the SHAP value in the current set, v χ represents the score of the χth model calculated on the test set. λ represents the number of models.
[0108] The specific screening process of the candidate features is divided into the following three steps:
[0109] 1) Train four machine learning models on the source air handling unit data set;
[0110] 2) Calculate and identify the top 10 features ranked by the weight of each model using the Tree SHAP method;
[0111] 3) Rank the four sets of features obtained using the normalization method shown in formula (2) to obtain 10 candidate features.
[0112] 2. Extract generalized features: based on the candidate features in the first stage, retrain the four integrated models on the target AHU data set, and finally extract 6 device-independent generalized features through normalized weight calculation. As the input of the subsequent fault diagnosis model.
[0113] This step aims to obtain features with strong generalization ability through the idea of transfer learning. In the target AHU dataset, the 10 candidate features obtained in step one are used as input to retrain four machine learning models. Then, the Tree SHAP method is used to calculate and identify the top 6 features of each model weight. Finally, the 6 generalization features are obtained through the normalization method shown in equation (2).
[0114] The fault data to be diagnosed corresponding to the generalization features pre-selected in step S1 is: in the fault data collected from the same air handling unit, the data corresponding to the row in the table of the fault data corresponding to the generalization features selected by the above method is selected.
[0115] The training process of the teacher-student model used in step S2 is as follows:
[0116] Semi-supervised deep learning fault diagnosis: This step aims to use the obtained generalization features to use a semi-supervised deep learning framework to use a teacher-student model based on a temporal convolutional network (TCN) to carry out a fault diagnosis task.
[0117] After the generalization feature selection, we obtain a key feature set composed of 6 sensor features. Then we use the key feature set for model training. That is, only 6 rows of data of the original table dataset are used for training, i.e. the dataset dimension is (time x 6).
[0118] During training:
[0119] (1) The dataset is composed of 20% labeled data and 80% unlabeled data, and the labeled data is divided into training set and validation set according to 8:2, while ensuring that each part contains complete fault types.
[0120] (2) Prepare input data (batch_size = 128, seq_len = 96, input_dim = 6), data preprocessing has global standardization (min-max), the weights of the student network are randomly initialized, the teacher network structure and weight parameters are consistent with the student network, and the teacher network does not update the weight parameters through back propagation.
[0121] (3) During training, for each training epoch, first, the student network is trained with labeled data and the cross-entropy loss function is calculated. Then, the unlabeled data set is input into the teacher network to obtain pseudo labels (hard labels, giving clear classification results), and the pseudo labels with low confidence are discarded according to the confidence weight, and only the pseudo labels with high confidence are retained. Next, the student network is trained with the pseudo label data, and then the loss function is calculated, which is composed of the cross-entropy loss function and the KL divergence (consistency loss). The student network parameters are updated. Finally, the teacher network parameters are updated using the exponential moving average (EMA, initial a = 0.99). The EMA formula is as follows:
[0122] θ t ←αθ t +(1-α)θ s θ t ,θ s are the parameters of the teacher network and the student network, respectively, and a is a smoothing hyperparameter.
[0123] (4) When the performance of the validation set no longer improves, stop training.
[0124] During testing, only the student network is used to participate in the classification task.
[0125] TCN is a model specifically designed for time series data processing, which is developed from the best design practices of convolutional neural networks. Unlike typical recurrent neural networks such as LSTM networks, TCN does not use a gating mechanism and has a longer effective memory, is more stable during model training, and can produce comparable or more excellent performance in many different sequence processing tasks. TCN has the following two characteristics:
[0126] 1) Causal Convolution, meaning that there is no information leakage from future time to past time;
[0127] 2) Isometric Mapping, meaning that the length of the input and output sequences can remain consistent, just like a recurrent neural network.
[0128] Dilated Causal Convolution is the main component of TCN. The meaning of causal convolution is that the convolution output at time t is only related to the elements at time t or before. Dilation convolution can efficiently solve the problem of multi-scale information fusion without increasing the number of parameters and layers. Given a 1-D input sequence x ∈ R n and a discrete filter h: {0, …, k-1}, x is an n-dimensional vector, and each component (element) of the vector is a real number. R nrepresents an n-dimensional real vector space.
[0129] Dilated convolution operation on element s in sequence x may be represented as follows:
[0130]
[0131] where, d represents a convolution with dilation size d, d is the dilation factor, k is the size of the filter, s is the current computed output time step, and f is the filter kernel.
[0132] This formula differs from the standard convolution in the subscript operation s-d×e, s-d×e represents the index in the input sequence x, which points to a past time step or position, which is determined by the current output position s, the filter index e, and the dilation factor d.
[0133] For the standard convolution, the subscript operation is s-e. In the causal convolution, the filter kernel only takes the d-th signal element, so the dilated convolution with d = 1 is equivalent to the standard convolution. The dilation factor is introduced as a fixed convolution stride between adjacent filters, and in the c-th layer of the deep network, it is usually set to d(c) = 2 c-1 , d(c) represents the value of the dilation factor of the i-th layer in the deep network (specifically, the TCN network). This exponentially increasing dilation factor design enables the TCN network to rapidly expand its receptive field as the number of layers increases.
[0134] When a larger dilation factor is set, the high-layer output of the network can represent the input with a larger observation range, so the receptive field of the TCN can be flexibly expanded.
[0135] Since the effectiveness of the TCN receptive field depends on the network depth, the filter size k, and the dilation factor d, it is very important to build a larger and deeper TCN for better feature expression. However, since it is very difficult to train a deep network, directly stacking a too deep neural network will cause a decrease in expression ability. To solve this problem, a residual connection can be used to stack multiple dilated causal convolution layers to form a residual connection block. Unlike the standard residual connection, the TCN uses an additional 1×1 convolution to ensure the element-wise addition operation The dimensions of the received tensors remain the same. In the residual block, the output of a series of transformations within the branch is added to the identity mapping of the block input x. Therefore, the residual connection block R(x) of the TCN can be represented as:
[0136]
[0137] Where, θ represents a parameter set in the residual block, T(x, θ) represents a series of transformation operations on the input x, including the operations of hollow causal convolution, weight normalization, activation function, etc.
[0138] This makes the intermediate layer of the deep network adjust the series of transformations of the intermediate block by introducing the identity mapping, and a large number of practices prove that this connection mode can effectively improve the expression ability of the deep network. At this time, a deep TCN network can be constructed by stacking a series of residual connection blocks. Assuming that the input of the mth residual connection block is x m , the forward propagation of the network from the block m of the shallower layer to the block M of the deeper layer can be expressed as:
[0139]
[0140] In the formula, x m represents the input of the mth residual block, θ m is the parameter of the mth residual block.
[0141] Assuming that the overall loss function is represented as The back propagation process can be calculated as follows:
[0142]
[0143] In the formula, L is the loss function; x m represents the input of the mth residual block; x M represents the output of the Mth residual block; represents the gradient of the loss function with respect to the input of the mth layer.
[0144] Since For a small batch of training samples, it is not always equal to 1, so the gradient disappearance phenomenon is difficult to occur, which can provide stable gradient flow for network training. A small batch of training samples is a small number of samples randomly selected from the entire training set, and the small batch gradient descent can make the network obtain more robust convergence and avoid converging to a local minimum.
[0145] The semi-supervised deep learning framework SS-TCN proposed in the application includes a student network and a teacher network with consistent structure, the network main body is composed of 8 layers of TCN, adopts residual structure connection and adds batch normalization processing, the output layer is a 12-dimensional full connection layer (corresponding to the number of fault categories), and the ReLU activation function is used to prevent gradient disappearance. The parameter update adopts a differentiated strategy: the student network updates the parameters in two ways, (1) the cross-entropy loss is calculated by using the labeled data; (2) the prediction result of the teacher network on the unlabeled data is received as a pseudo label for training. The weight of the teacher network is dynamically updated by the EMA algorithm, as shown in formula (7):
[0146] θ teacher= a · theta teacher + (1-a) · theta student (7)
[0147] In the formula, theta teacher is the teacher network parameter, theta student is the student network parameter, and alpha is a smoothing hyperparameter to ensure smooth transition of network parameters.
[0148] During the training process, the loss function is composed of two parts, namely supervised loss and unsupervised loss, as shown in formula (8).
[0149]
[0150] In the formula, L is the labeled data, B is the batch training set containing labeled data and unlabeled data, b represents a single data sample in the training process, eta is the unsupervised loss weight, and C is the total number of classifications. respectively, the student network output, the teacher network output and the real sample.
[0151] Step S2 comprises:
[0152] When performing fault diagnosis, the trained student model is used to identify and classify faults. The specific diagnosis process is as follows: first, select the data to be diagnosed according to the generalization features as input data, input the fault data to be diagnosed into the student model, the model extracts the time sequence features of the data through its 8-layer TCN (time sequence convolution network) main body, and after residual connection and batch normalization processing, the features are further optimized and transmitted to the output layer. The output layer is a 12-dimensional fully connected layer corresponding to 12 different fault categories, combined with the ReLU activation function to ensure nonlinear expression capability. Finally, the student model outputs the probability distribution of each fault category according to the input data, thereby realizing accurate fault diagnosis.
[0153] As described above, the specific flowchart of the above steps is shown in Figure 2 The meaning of BP in the figure is back propagation.
[0154] Embodiment 2
[0155] Based on the same inventive concept, the present application also provides an air handling unit fault diagnosis system, as shown in Figure 3 The system comprises a selected input data module and a diagnosis module.
[0156] The selected input data module is used to select the fault data to be diagnosed corresponding to the preselected generalization features from the fault data to be diagnosed as input data.
[0157] The diagnosis module is used to input the input data into the trained teacher-student model to obtain the fault diagnosis result.
[0158] The teacher-student model extracts the time sequence characteristics of the fault data by using a time sequence convolution network, and is trained by taking the historical input data as the input and the corresponding fault diagnosis result as the output.
[0159] Preferably, the selection process of the generalization features comprises:
[0160] The Shapley values of the sensor features in the fault data are calculated, and a plurality of first candidate features of the respective integrated models are selected from the sensor features in descending order of the Shapley values;
[0161] The normalized weights of the first candidate features of the respective integrated models are calculated, and the first candidate features are ranked in descending order of the weights, and the first quantity of the first candidate features are taken as the high-weight candidate feature sets of the respective integrated models;
[0162] The data corresponding to the high-weight candidate feature sets are selected from the target air handling unit data set as the input, and the corresponding fault type classification results and sensor feature scores are taken as the output to retrain the plurality of integrated models;
[0163] The Shapley values of the sensor features in the high-weight candidate feature sets of the retrained plurality of integrated models are calculated, and a plurality of second candidate features of the respective integrated models are selected from the first candidate features in descending order of the Shapley values;
[0164] The normalized weights of the second candidate features of the respective integrated models are calculated, and the second candidate features are ranked in descending order of the weights again, and the second quantity of the second candidate features are taken as the generalization features of the respective integrated models;
[0165] The plurality of integrated models are trained by taking the fault data in the source air handling unit data set as the input and the fault type classification results and sensor feature scores as the output; the target air handling unit data set is part of the source air handling unit data set; the first quantity is greater than the second quantity; and the number of the first candidate features is greater than the number of the second candidate features.
[0166] Preferably, the fault data in the input data selection module is of TxF dimension, T is the sampling time, and F is the sensor feature.
[0167] Preferably, the plurality of integrated models in the input data selection module comprise at least one or more of the following: a random forest model, an extreme gradient boosting tree model, a light gradient boosting machine model, or a category boosting model.
[0168] Preferably, the Shapley value is determined according to the following formula:
[0169]
[0170] wherein, is the Shapley value, denotes the subset of S excluding the sensor feature x j , |S| denotes the number of elements in the set S, x denotes the sensor feature, j denotes the sensor feature number, n denotes the number of sensor features, f x (S j ∪{x j}) is the sensor feature score including x j , f x (S j ) is the sensor feature score excluding x j .
[0171] Preferably, the normalized weight in the selected input data module is determined according to the following formula:
[0172]
[0173] wherein, is the normalized weight of the lth sensor feature of the χth integrated model, r l denotes the ranking of the lth sensor feature based on the sensor feature contribution degree of the Shapley value, v χ denotes the score of the χth model, and λ denotes the number of models.
[0174] Preferably, the training of the teacher-student model in the diagnosis module comprises:
[0175] A1, dividing the target air handling unit data set corresponding to the generalization feature into a training data set and a validation data set;
[0176] A2, training the student network in the teacher-student model using the labeled data in the training data set, and updating the parameters of the teacher network in the teacher-student model according to the parameters of the student network using the exponential moving average update algorithm;
[0177] A3, inputting the unlabeled data in the training data set into the updated teacher network to obtain pseudo-labeled data, training the student network using the pseudo-labeled data, and updating the parameters of the teacher network according to the parameters of the trained student network using the exponential moving average update algorithm;
[0178] A4, repeating step A3 until the performance of the validation data set no longer improves;
[0179] wherein, the parameters of the student network are updated during the training process by calculating the cross-entropy loss; and the initial value of the parameters of the teacher network is obtained by random initialization.
[0180] Preferably, the diagnostic module uses an exponential moving average update algorithm to update the parameters of the teacher network in the teacher-student model, implemented as follows:
[0181] θ teacher =α·θ teacher +(1-α)·θ student
[0182] In the formula, θ teacher For the parameters of the teacher network, θ student Let be the parameters of the student network, and α be the smoothing hyperparameter.
[0183] Preferably, the diagnostic module is used for:
[0184] The input data is fed into the student network of the pre-trained teacher-student model, and the temporal features of the input data are extracted through the temporal convolutional network in the student network.
[0185] The temporal features are processed by residual connections and batch normalization in the student network to obtain the processed temporal features.
[0186] The processed temporal features are passed to the output layer of the student network, which corresponds to the number of fault categories in terms of dimension, to obtain the probability distribution of each fault category, which serves as the fault diagnosis result.
[0187] Example 3
[0188] like Figure 4 As shown, the present invention also provides an electronic device, which may be a computer device, a microcontroller device, a smart mobile device, etc. The electronic device in this embodiment may include a processor, a memory, a transceiver component, etc. The memory, processor, and transceiver component are connected via a bus; the memory can be used to store executable programs, and an exemplary executable program may include instructions; the processor is used to execute the instructions stored in the memory. The memory can also be used to store data, which can be accessed and / or modified when instructions are executed.
[0189] The processor can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., which are the computing core and control core of the terminal, and are suitable for implementing one or more instructions, and are specifically suitable for loading and executing one or more instructions in the storage medium to implement a corresponding method flow or a corresponding function, to implement the steps of the air handling unit fault diagnosis method in the above embodiments.
[0190] Embodiment 4
[0191] Based on the same inventive concept, the application further provides a readable storage medium, specifically an electronic device readable storage medium (Memory), which is a memory device in the electronic device, and is used for storing programs and data. It can be understood that the storage medium herein can include a built-in storage medium in the electronic device, and of course can also include an expansion storage medium supported by the electronic device. The storage medium provides a storage space, and the storage space stores an operating system of the terminal. In addition, one or more instructions suitable for being loaded and executed by the processor are also stored in the storage space, and the instructions can be one or more execution programs (including program codes). It should be noted that the storage medium herein can be a high-speed RAM memory, or a non-volatile memory such as at least one disk memory. The processor loads and executes one or more instructions stored in the storage medium, to implement the steps of the air handling unit fault diagnosis method in the above embodiments.
[0192] In summary, the present application provides an air handling unit fault diagnosis method, system, device and medium, comprising: selecting the fault data to be diagnosed corresponding to the pre-selected generalization features as input data from the fault data to be diagnosed; inputting the input data into the trained teacher-student model to obtain the fault diagnosis result; wherein the teacher-student model extracts the time sequence features of the fault data using the time sequence convolution network, and is trained by taking the historical input data as the input and the corresponding fault diagnosis result as the output; the generalization features are selected by calculating the Shapley value and the normalized weight of the sensor features. The present application improves the low efficiency problem of the traditional method caused by relying on full-amount labeled data, device-specific feature set and high-dimensional redundant input by using the generalization feature selection mechanism and the semi-supervised deep learning framework.
[0193] Those skilled in the art will appreciate that embodiments of the present application can be provided as methods, systems, or computer program products. Accordingly, the present application can be embodied in the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present application can be embodied in the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk memory, CD-ROMs, optical memory, etc.) having computer usable program code embodied thereon.
[0194] The present application is described with reference to the flowcharts and / or block diagrams of the methods, apparatus (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing apparatus to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing apparatus produce an apparatus that implements the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions specified in a flow or multiple flows and / or blocks Figure 1 The means for performing the functions specified in a flow or multiple flows and / or blocks.
[0195] These computer program instructions can also be stored in a computer-readable memory that can direct the computer or other programmable data processing apparatus to work in a specific manner, so that the instructions stored in the computer-readable memory produce a product including instruction means, which implements the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions specified in a flow or multiple flows and / or blocks Figure 1 The means for performing the functions specified in a flow or multiple flows and / or blocks.
[0196] These computer program instructions can also be loaded into a computer or other programmable data processing devices, so that a series of operational steps are generated to realize the computer-implemented processes in the computer or other programmable devices, and the instructions executed in the computer or other programmable devices provide operational steps for implementing the functions specified in the flowchart Figure 1 one flow or a plurality of flows and / or the functions specified in the block Figure 1 one flow or a plurality of flows and / or the functions specified in the block
[0197] The above merely illustrates the embodiments of the present application, and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall fall within the scope of the claims of the present application.
Claims
1. An air handling unit fault diagnosis method characterized by, The method comprises the following steps: selecting, from the fault data to be diagnosed, fault data to be diagnosed corresponding to the preselected generalization features as input data; inputting the input data into the trained teacher-student model to obtain a fault diagnosis result; wherein the teacher-student model extracts time sequence features of the fault data by using a time sequence convolution network, and is trained by taking historical input data as input and taking corresponding fault diagnosis results as output; the generalization features are selected by calculating the Shapley values and normalized weights of the sensor features.
2. The method of claim 1, wherein, The selection process of the generalization features comprises: calculating the Shapley values of the sensor features in the fault data, and selecting a plurality of first candidate features of each integrated model from the sensor features in descending order of the Shapley values; calculating the normalized weights of the first candidate features of each integrated model, and ranking the first candidate features in descending order of the weights, wherein the first quantity of the first candidate features are high-weight candidate feature sets of each integrated model; retraining the plurality of integrated models by taking the data corresponding to the high-weight candidate feature sets selected from the target air handling unit data set as input, and taking the corresponding fault type classification results and sensor feature scores as output; calculating the Shapley values of the sensor features in the high-weight candidate feature sets in the retrained plurality of integrated models, and selecting a plurality of second candidate features of each integrated model from the first candidate features in descending order of the Shapley values; calculating the normalized weights of the second candidate features of each integrated model selected again, and ranking the second candidate features again in descending order of the weights, wherein the first quantity of the second candidate features are generalization features of each integrated model; wherein the plurality of integrated models are trained by taking the fault data in the source air handling unit data set as input, and taking the fault type classification results and sensor feature scores as output; the target air handling unit data set is part of the source air handling unit data set; the first quantity is greater than the second quantity; the number of the first candidate features is greater than the number of the second candidate features.
3. The method of claim 2, wherein, The fault data is of TxF dimension, where T is the sampling time and F is the sensor feature.
4. The method of claim 2, wherein, The plurality of integrated models comprise at least one or more of the following: a random forest model, an extreme gradient boosting tree model, a light gradient boosting machine model, or a category boosting model.
5. The method of claim 2, wherein, The Shapley value is determined according to the following formula: wherein is the Shapley value, denotes the subset of S not containing the sensor feature x j , |S| denotes the number of elements in the set S, x is the sensor feature, j is the sensor feature number, n is the number of sensor features, f x (S j ∪{x j}) is the sensor feature score including x j , f x (S j ) is the sensor feature score not including x j .
6. The method of claim 2, wherein, The normalized weight is determined according to the following formula: wherein, is the normalized weight of the lth sensor feature of the χth integrated model, r l denotes the ranking of the lth sensor feature based on the sensor feature contribution degree of the Shapley value, v χ denotes the score of the χth model, and λ denotes the number of models.
7. The method of claim 1, wherein, The training of the teacher-student model comprises: A1, dividing the target air handling unit data set corresponding to the generalization features into a training data set and a validation data set; A2, training the student network in the teacher-student model by using the labeled data in the training data set, and updating the parameters of the teacher network in the teacher-student model by using the exponential moving average update algorithm according to the parameters of the student network; A3, inputting the unlabeled data in the training data set into the updated teacher network to obtain pseudo-labeled data, training the student network by using the pseudo-labeled data, and updating the parameters of the teacher network by using the exponential moving average update algorithm according to the parameters of the trained student network; A4, repeating step A3 until the performance of the validation data set no longer improves; The parameters of the student network are updated during the training process by calculating cross-entropy loss; and the initial value of the parameters of the teacher network is obtained by random initialization.
8. The method of claim 7, wherein, The parameters of the teacher network in the teacher-student model are updated by using an exponential moving average update algorithm, and the implementation is as follows: θ teacher = a · θ teacher + (1 - a) · θ student where θ teacher are the parameters of the teacher network, θ student are the parameters of the student network, and α is a smoothing hyperparameter.
9. The method of claim 1, wherein, The input data is input into the trained teacher-student model to obtain the fault diagnosis result, which comprises: The input data is input into the student network in the trained teacher-student model, and the time sequence characteristics of the input data are extracted by the time sequence convolution network in the student network; The time sequence characteristics are processed by the residual connection and batch normalization in the student network to obtain processed time sequence characteristics; The processed time sequence characteristics are transmitted to the output layer in the student network corresponding to the dimension of the fault category number to obtain the probability distribution of each fault category as the fault diagnosis result.
10. An air handling unit fault diagnostic system, comprising: It comprises: an input data module and a diagnosis module; The input data module is selected from the fault data to be diagnosed, and the fault data to be diagnosed corresponding to the pre-selected generalization features is selected as the input data; The diagnosis module is used to input the input data into the trained teacher-student model to obtain the fault diagnosis result. The teacher-student model uses the time sequence convolution network to extract the time sequence characteristics of the fault data, and is trained by taking the historical input data as the input and the corresponding fault diagnosis result as the output; and the generalization features are selected by calculating the Shapley value and the normalized weight of the sensor features.
11. The system of claim 10, wherein, The selection process of the generalization features comprises: The Shapley value of the sensor features in the fault data is calculated, and a plurality of first candidate features of each integrated model are selected from the sensor features in descending order of the Shapley value; The normalized weight of the first candidate features of each integrated model is calculated, and the first candidate features are ranked in descending order of the weight, and the first quantity of the first candidate features are used as the high-weight candidate feature set of each integrated model; The data corresponding to the high-weight candidate feature set is selected from the target air handling unit data set as the input, and the plurality of integrated models are retrained by taking the corresponding fault type classification result and sensor feature score as the output; The Shapley value of the sensor features in the high-weight candidate feature set in the retrained plurality of integrated models is calculated, and a plurality of second candidate features of each integrated model are selected from the first candidate features in descending order of the Shapley value; The normalized weight of the second candidate features of each integrated model is calculated, and the second candidate features are ranked in descending order of the weight again, and the second quantity of the second candidate features are used as the generalization features of each integrated model. The plurality of integrated models are trained by taking the fault data in the source air handling unit data set as the input and the fault type classification result and the sensor feature score as the output; the target air handling unit data set is part of the source air handling unit data set; the first quantity is greater than the second quantity; and the number of the first candidate features is greater than the number of the second candidate features.
12. The system of claim 11, wherein, The fault data in the input data module is TxF dimensional, T is the sampling time, and F is the sensor feature.
13. The system of claim 11, wherein, The multiple integrated models in the selected input data module include at least one or more of the following: a random forest model, an extreme gradient boosting tree model, a light gradient boosting machine model, or a category boosting model.
14. The system of claim 11, wherein, The Shapley value in the selected input data module is determined by the following formula: wherein is the Shapley value, denotes the subset of S not containing the sensor feature x j , |S| denotes the number of elements in the set S, x is the sensor feature, j is the sensor feature number, n is the number of sensor features, f x (S j ∪{x j}) is the sensor feature score including x j , f x (S j ) is the sensor feature score not including x j .
15. The system of claim 11, wherein, The normalized weight in the selected input data module is determined by the following formula: wherein, is the normalized weight of the lth sensor feature of the χth integrated model, r l denotes the ranking of the lth sensor feature based on the sensor feature contribution degree of the Shapley value, v χ denotes the score of the χth model, and λ denotes the number of models.
16. The system of claim 10, wherein, The training of the teacher-student model in the diagnosis module includes: A1, dividing the target air handling unit data set corresponding to the generalization feature into a training data set and a validation data set; A2, training the student network in the teacher-student model using the labeled data in the training data set, updating the parameters of the teacher network in the teacher-student model according to the parameters of the student network using the exponential moving average update algorithm; A3, inputting the unlabeled data in the training data set into the updated teacher network to obtain pseudo-labeled data, training the student network using the pseudo-labeled data, and updating the parameters of the teacher network according to the parameters of the trained student network using the exponential moving average update algorithm; A4, repeating step A3 until the performance of the validation data set no longer improves; Wherein, the parameters of the student network are updated by calculating the cross-entropy loss during the training process; the initial value of the parameters of the teacher network is obtained by random initialization.
17. The system of claim 16, wherein, The exponential moving average update algorithm is used to update the parameters of the teacher network in the teacher-student model, which is implemented by the following formula: θ teacher = a · θ teacher + (1 - a) · θ student where θ teacher are the parameters of the teacher network, θ student are the parameters of the student network, and α is a smoothing hyperparameter.
18. The system of claim 10, wherein, The diagnosis module is used to: input the input data into the student network in the trained teacher-student model, and extract the time sequence features of the input data through the time sequence convolution network body in the student network; obtain processed time sequence features after the time sequence features pass through the residual connection and batch normalization in the student network; pass the processed time sequence features to the output layer corresponding to the dimension of the fault category number in the student network to obtain the probability distribution of each fault category as the fault diagnosis result.
19. A computer device, comprising: It includes: at least one processor and a memory; The memory and the processor are connected through a bus; The memory is used to store one or more programs; When the one or more programs are executed by the at least one processor, the air handling unit fault diagnosis method of any one of claims 1-9 is implemented.
20. A computer-readable storage medium, characterized in that, It has an execution program stored thereon, and when the execution program is executed, the air handling unit fault diagnosis method of any one of claims 1-9 is implemented.
Citation Information
Cited By
Method for predicting residual life of engine based on time-frequency cooperative dual-channel TCN-Transformer
CN122310453A