A systematic feature selection method for HVAC and Refrigeration system fault diagnosis
By combining filters and encapsulation methods, using the maximum relevance minimum redundancy algorithm and mean influence value (MIV), a one-dimensional convolution-gated recurrent unit network model is constructed, which solves the scientific nature and computational cost issues of feature selection in HVAC and refrigeration system fault diagnosis, and improves the accuracy and efficiency of diagnosis.
Patent Information
- Application Number
- CN202311165815.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-11
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2043-09-11
AI Technical Summary
Existing feature selection methods lack a scientific and universal framework for HVAC and refrigeration system fault diagnosis, rely on manual experience and have high computational costs, and are difficult to adapt to the fault diagnosis needs of complex systems.
The filter method and the encapsulation method are combined, and the maximum relevance minimum redundancy algorithm and the mean influence value (MIV) are used for feature selection. A one-dimensional convolution-gated recurrent unit network model is constructed for intelligent fault diagnosis.
It achieves low-cost and efficient feature selection, improves the accuracy of fault diagnosis, reduces computational complexity and training costs, and is suitable for a wide range of HVAC and refrigeration system fault diagnosis.
Smart Images

Figure CN117194896B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent fault diagnosis, and in particular relates to a systematic feature selection method for fault diagnosis of heating, ventilation, air conditioning and refrigeration systems. Background Art
[0002] The rapid development of fields such as construction, electronics, and healthcare has placed increasingly stringent demands on HVAC and refrigeration systems. As their components become increasingly complex and diverse, HVAC and refrigeration systems are more susceptible to various types of failures. Faulty system operation results in energy waste, and in severe cases, sudden equipment shutdowns can cause significant losses. Therefore, it is necessary to monitor and diagnose faults, monitor system operating status, promptly identify and eliminate faults, and ensure the safe and efficient operation of HVAC and refrigeration systems. In recent years, with the collection of large amounts of data and the rapid development of fields such as artificial intelligence and machine learning, corresponding intelligent fault diagnosis methods are expected to better address the fault diagnosis issues of HVAC and refrigeration systems.
[0003] Selecting appropriate features is both a prerequisite and key to the application of intelligent fault diagnosis methods. Feature selection, the process of selecting a subset of relevant features when building a model, directly impacts model performance and computational complexity. However, existing feature selection methods typically rely on manual experience, domain knowledge, and specific scenarios, while ignoring measurement costs. Consequently, a scientific and universal feature selection framework applicable to general scenarios is lacking when building fault diagnosis models for HVAC and refrigeration systems.
[0004] Therefore, it is necessary to establish a comprehensive, scientific and systematic feature selection method to guide the feature selection process of intelligent fault diagnosis of HVAC and refrigeration systems. Summary of the Invention
[0005] In order to overcome the shortcomings of the above-mentioned prior art, the purpose of the present invention is to provide a systematic feature selection method for HVAC and refrigeration system fault diagnosis. The method couples the filter method and the encapsulation method, and introduces the mean influence value (MIV) into the fault diagnosis feature selection. It can systematically and comprehensively guide the feature selection process for HVAC and refrigeration system fault diagnosis to solve one or more of the above-mentioned technical problems.
[0006] In order to achieve the above object, the technical solution adopted by the present invention is:
[0007] A systematic feature selection method for HVAC and refrigeration system fault diagnosis includes the following steps:
[0008] Step 1: Obtain a dataset for intelligent fault diagnosis of HVAC and refrigeration systems, which contains the original features to be selected;
[0009] Step 2: Using the filter method, a feature subset is initially screened based on the maximum relevance and minimum redundancy algorithm;
[0010] In step 3, a coupled method is used in combination with the model for intelligent fault diagnosis to calculate the average influence value of the features, and finally the feature selection process for intelligent fault diagnosis is completed.
[0011] In one embodiment, in step 1, the operating data of the HVAC and refrigeration system units in normal states and various fault states are first obtained, and the data are pre-processed to remove outliers and standardize to obtain a data set Ω for intelligent fault diagnosis.
[0012] In one embodiment, between step 2 and step 3, the following steps are further performed:
[0013] Considering the experimental measurement conditions and measurement cost factors of the model input features, combined with the feature importance ranking in the set S, we screen again to obtain a new set of features to be selected.
[0014] The present invention also provides a corresponding HVAC and refrigeration system fault diagnosis method, which adopts the systematic feature selection method for HVAC and refrigeration system fault diagnosis to select features for intelligent fault diagnosis; then, a supervised learning fault diagnosis model is constructed based on the selected features, and the operating data of the HVAC and refrigeration system is monitored. The operating status of the system is detected through the model, and faults are discovered and diagnosed in a timely manner.
[0015] Existing intelligent fault diagnosis feature selection methods can be categorized into three types: filter, encapsulated, and embedded. Filter methods rely on evaluating the intrinsic characteristics of the data and selecting feature subsets, regardless of the selected algorithm, resulting in poor adaptability to specific problems. Encapsulated methods evaluate each candidate feature based on the performance of the selected model, requiring model training for each evaluation and resulting in high computational costs. Embedded methods are only applicable to specific algorithms. Compared to existing technologies, the feature selection method proposed in this paper combines filter and encapsulated methods, considering both the inherent mathematical connections of the data and the characteristics of the specific model used for fault diagnosis. This method offers advantages such as comprehensive considerations, a comprehensive and systematic approach, and excellent feature selection results. The proposed method is not limited to specific data or algorithms, or even to intelligent fault diagnosis for HVAC and refrigeration systems, and has a wide range of applications and excellent versatility. The proposed method innovatively introduces the mean influence value (MIV) into feature selection, requiring only a single model training, eliminating the need for multiple retraining cycles and resulting in low computational costs. The proposed method also considers actual measurement costs, making it easier to apply in the field.
[0016] The method of the present invention can arrange various fault diagnosis models, which can be used for intelligent fault diagnosis of HVAC and refrigeration systems after training, which can greatly reduce training and computing costs and improve diagnostic accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 This is the specific process of the proposed feature selection method.
[0018] Figure 2 This is the intelligent fault diagnosis process of HVAC and refrigeration systems using this feature selection method.
[0019] Figure 3 This is an example of an intelligent fault diagnosis experimental platform for HVAC and refrigeration systems using this feature selection method.
[0020] Figure 4 The results of average influence value ranking in the intelligent fault diagnosis case of HVAC and refrigeration system using this feature selection method are shown. DETAILED DESCRIPTION
[0021] To make the purpose, technical effects, and technical solutions of the embodiments of the present invention more clear, the technical solutions of the embodiments of the present invention are clearly and completely described below in conjunction with the accompanying drawings and specific fault diagnosis cases in which the present invention is implemented. Obviously, the described embodiments are only part of the embodiments of the present invention. Based on the embodiments disclosed in the present invention, other embodiments obtained by ordinary technicians in this field without making any creative efforts should fall within the scope of protection of the present invention.
[0022] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the same. Although the present invention has been described in detail with reference to the above embodiments, a person skilled in the art may still modify or make equivalent substitutions to the specific implementations of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention are within the scope of protection of the claims of the present invention to be approved.
[0023] refer to Figure 1 and Figure 2 As shown, the systematic feature selection method for HVAC and refrigeration system fault diagnosis of the present invention is specifically implemented as follows:
[0024] Step 1: Obtain a data set for intelligent fault diagnosis of HVAC and refrigeration systems, which contains the original features to be selected.
[0025] The present invention constructs an intelligent fault diagnosis database by carrying out fault diagnosis experiments on specific HVAC and refrigeration systems, measuring characteristic data of normal and fault states. Figure 3As shown, the data contains 44 features, including temperature, pressure, and flow at different locations. The specific feature data codes and definitions are shown in Table 1. The data is pre-processed to remove outliers and perform standardization to obtain the dataset Ω for intelligent fault diagnosis feature selection. In this example, refrigerant leaks of varying severity are used as the fault diagnosis targets.
[0026] Table 1 Characteristic data list
[0027]
[0028] In step 2, a forward addition method is used, and a filter method is used to preliminarily screen the feature subset S based on the maximum relevance min-redundancy algorithm (mRMR). It is planned to preliminarily screen out 20 features in this stage. The specific steps are shown in steps 21 to 26.
[0029] This step can be described in detail as follows:
[0030] Step 21: Calculate the correlation V between feature X and target output Z based on mutual information X , select the correlation V X The largest feature X is selected and the selected feature is added to the empty set S. In this embodiment, the correlation V between the 44 selected features and the refrigerant leakage fault is calculated respectively. X , correlation V X The calculation formula is shown in formula (1):
[0031] V X =I(X,Z)(1)
[0032] Where X represents a feature, Z represents the target output, and I(X, Z) represents the mutual information between variables X and Z. Its calculation formula is shown in formula (2):
[0033]
[0034] Among them, P(X) and P(Z) are the probability distributions of variables X and Z respectively, and P(X, Z) is the joint probability distribution of the two variables.
[0035] Step 22, in the complement of S C Filter out V that has non-zero correlation with the target output Z X And has zero redundancy W with other features Y in set S X features, among which the correlation V X The largest features are put into the set S in sequence, where the redundant W X The calculation formula is shown in formula (3):
[0036]
[0037] Among them, X represents a feature to be selected, Y represents the selected features in the geometry S, and |S| represents the number of features contained in the set S.
[0038] Step 23, repeat step 22 until S C The redundancy of all features in is non-zero;
[0039] Step 24, in the complement of S C The filter has a non-zero correlation V with the target output Y X , and the set S has non-zero redundancy W X The features with the maximum mutual information quotient (MIQ) are selected and added to the set S. The calculation formula of the maximum mutual information quotient MIQ is shown in formula (4):
[0040]
[0041] Step 25, repeat step 24 until the set S C The correlation of all features in V X is zero.
[0042] Step 26: Add the zero-correlation features to the set S in a random order at the end of the queue, which is the optimal feature subset S containing |S| features.
[0043] At this point, the preliminary screening based on the maximum relevance minimum redundancy algorithm (mRMR) was completed, and the ranking of 44 features in the data set was obtained. The top 20 features were selected as the candidate optimal feature subset S, including compressor inlet temperature, evaporator 1 average temperature, condenser air side temperature difference, compressor pressure difference, etc., which maximized the relevance V of the set S with respect to the target output Y. S , minimizes the redundancy W within the set S S , for further feature selection process, where the correlation V S The calculation of is shown in formula (5), the redundancy W S The calculation of is shown in formula (6):
[0044]
[0045]
[0046] After steps 1 and 2, the feature set is sorted and preliminarily screened. At this point, we can combine manual experience, consider factors such as the experimental measurement conditions and measurement costs of the model input features, and choose to retain or eliminate some features to further exclude features. Based on the perspective of field application, this embodiment considers factors such as the experimental measurement conditions and measurement costs of the model input features, combines the feature importance ranking and specific application scenarios in the set S, and excludes some features with higher measurement costs, including 3 feature data: capillary inlet pressure, condenser outlet pressure, and system mass flow, to obtain a new set S containing 17 features to be selected.
[0047] In step 3, a coupled method is used in combination with the model for intelligent fault diagnosis to calculate the average influence value of the features, and finally the feature selection process for intelligent fault diagnosis is completed.
[0048] In this step, the mean impact value (MIV) of the model features is first calculated using an encapsulated method, and the features are further filtered to obtain the final result. The specific filtering method is shown in steps 31 to 35.
[0049] Step 31 : Select the intelligent fault diagnosis supervised learning fault diagnosis model to be used, use the data set S to divide it into a training set T, a validation set V and a test set E, and train the selected intelligent fault diagnosis model.
[0050] In this embodiment, a one-dimensional convolution-gated recurrent unit network is used as an intelligent fault diagnosis supervised learning model, and a data set S containing 17 features is divided into a training set T, a validation set V, and a test set E for training.
[0051] Step 32: All sample data corresponding to feature X1 in the test set P are added or subtracted in the same proportion to construct two new test sets E1 and E2. In this embodiment, the proportion is 10%, that is, 10% is added or subtracted respectively. The calculation method is shown in formula (7) and formula (8):
[0052]
[0053]
[0054] In step 33, the trained network is tested using the new test sets E1 and E2 to obtain two new results A1 and A2. The difference between A1 and A2 is the impact of the change on the output after the independent variable is changed. The mean impact value (MIV) is obtained by averaging the values according to the number of samples. The specific calculation formula of the mean impact value is shown in formula (9):
[0055]
[0056] Among them, |P| represents the number of samples contained in set P, a1 and a2 represent the samples in sets A1 and A2 respectively.
[0057] Step 35, repeat steps 33 and 34, calculate the average impact value of all features in the data set, and perform multiple rounds of calculation to calculate the average value, obtain the calculation results and sort them, and select the required features based on the final results. In this embodiment, five features are selected for fault diagnosis, namely capillary inlet temperature, condenser air side outlet temperature, condenser air side inlet temperature, long tube temperature 3 and compressor inlet temperature, among which the final impact value is sorted as follows Figure 4 The specific codes and their meanings are shown in Table 1.
[0058] Based on the selected features for intelligent fault diagnosis, a supervised learning fault diagnosis model is constructed to monitor the operating data of the HVAC and refrigeration systems. The model detects the system's operating status and promptly detects and diagnoses faults. For example, in this embodiment, a one-dimensional convolutional-gated recurrent unit network fault diagnosis model is ultimately constructed and deployed for a specific HVAC and refrigeration system. The model monitors the system's operating data and can provide corresponding normal or faulty status judgments, enabling intelligent and timely system maintenance to ensure healthy and efficient system operation.
[0059] The method for constructing a model and monitoring the operating status of the system according to the present invention is as follows:
[0060] Step 91 : Select the intelligent fault diagnosis supervised learning fault diagnosis model to be used. In this example, a one-dimensional convolution-gated recurrent unit network fault diagnosis model is constructed.
[0061] Step 92: Based on the selected features, historical operating data of the HVAC and refrigeration systems is collected to obtain data in normal and various fault states. The data is divided into training, validation, and test sets, and the selected intelligent fault diagnosis model is trained, validated, and tested respectively to obtain the final intelligent fault diagnosis model.
[0062] In step 93, a fault diagnosis model is deployed during HVAC / Refrigeration system operation. Real-time system data is fed into the model, which then determines the system's status based on the input, reporting whether the system is normal or has a fault. Using the five features selected using the proposed feature selection method, a one-dimensional convolutional-gated recurrent unit network fault diagnosis model was constructed, achieving a final fault diagnosis accuracy of 98.1%.
Claims
1. A systematic feature selection method for HVAC and refrigeration system fault diagnosis, characterized by: The steps include: Step 1: Obtain a dataset for intelligent fault diagnosis of HVAC and refrigeration systems, which contains the original features to be selected; Step 2: Use the filter method to preliminarily screen the feature subset based on the maximum relevance and minimum redundancy algorithm. The screening method is as follows: Step 21: Based on the mutual information, calculate the correlation between the feature X and the target output Z, select the feature X with the largest correlation, and add the selected feature to the empty set S; Step 22, in the complement of S C Filter out the features that have non-zero correlation with the target output Z and zero redundancy with other features Y in the set S, select the features with the largest correlation, and put them into the set S in turn; Step 23, repeat step 22 until S C The redundancy of all features in is non-zero; Step 24, in the complement of S C Filter the features that have non-zero correlation with the target output Z and non-zero redundancy with other features Y in the set S, select the features with the largest mutual information quotient, and add the selected features to the set S; Step 25, repeat step 24 until the set S C The correlation of all features in is zero; Step 26: Add the zero-correlation features to the set S in random order at the end of the queue, which is the optimal feature subset S containing |S| features; In step 3, a coupled approach is used in combination with the model for intelligent fault diagnosis. The encapsulation approach is used to calculate the average impact value of the model features. The features are further filtered to obtain the final result, thus completing the feature selection process for intelligent fault diagnosis. The filtering method is as follows: Step 31: Select the intelligent fault diagnosis supervised learning fault diagnosis model to be used, use the data set S to divide it into a training set T, a validation set V and a test set E, and train the selected intelligent fault diagnosis model; Step 32: All sample data corresponding to feature X1 in the test set P are added and subtracted in the same proportion to construct two new test sets E1 and E2; Step 33: Use new test sets E1 and E2 to perform test calculations on the trained network, obtaining two new results A1 and A2. The difference between A1 and A2 is the impact of changing the independent variable on the output. The difference is averaged over the number of samples to obtain the average impact value. Step 35: Repeat steps 33 and 34 to calculate the average influence value of all features in the data set, perform multiple rounds of calculations to calculate the average value, obtain the calculation results and sort them, and select the required features based on the final results.
2. The method for systematic feature selection for HVAC and refrigeration system fault diagnosis according to claim 1, characterized in that: In step 1, the operating data of the HVAC and refrigeration system units in normal state and various fault states are first obtained, and the data are pre-processed to remove outliers and standardize to obtain a data set Ω for intelligent fault diagnosis.
3. The method for systematic feature selection for HVAC and refrigeration system fault diagnosis according to claim 1, characterized in that: The calculation formula of the correlation is shown in formula (1): V X =I(X,Z) (1) Where X represents a feature, Z represents the target output, and I(X, Z) represents the mutual information between variables X and Z. Its calculation formula is shown in formula (2): Where P(X) and P(Z) are the probability distributions of variables X and Z respectively, and P(X, Z) is the joint probability distribution of the two variables; The redundancy calculation formula is shown in formula (3): Among them, X represents a feature to be selected, Y represents the selected features in the geometry S, and |S| represents the number of features contained in the set S; The calculation formula of the maximum mutual information quotient is shown in formula (4):
4. The method for systematic feature selection for HVAC and refrigeration system fault diagnosis according to claim 1, characterized in that: Between step 2 and step 3, the following steps are further performed: Considering the experimental measurement conditions and measurement cost factors of the model input features, combined with the feature importance ranking in the set S, we screen again to obtain a new set of features to be selected.
5. The method for systematic feature selection for HVAC and refrigeration system fault diagnosis according to claim 1, characterized in that: The calculation formula of the average impact value is shown in formula (5): Here, |P| represents the number of samples in the set P, and a1 and a2 are the elements in A1 and A2 respectively.
6. A method for diagnosing faults in heating, ventilation, air conditioning and refrigeration systems, characterized in that: A systematic feature selection method for HVAC and refrigeration system fault diagnosis as described in any one of claims 1 to 5 is used to select features for intelligent fault diagnosis; then, a supervised learning fault diagnosis model is constructed based on the selected features, the operating data of the HVAC and refrigeration system is monitored, the operating status of the system is detected through the model, and faults are discovered and diagnosed in a timely manner.
7. The HVAC and refrigeration system fault diagnosis method according to claim 6, characterized in that: The method of constructing the model and monitoring the system operation status is as follows: Step 91, selecting the intelligent fault diagnosis supervised learning fault diagnosis model to be used; Step 92: Based on the selected features, historical operating data of the HVAC and refrigeration systems is collected to obtain data in normal and various fault states. The data is divided into training, validation, and test sets, and the selected intelligent fault diagnosis model is trained, validated, and tested respectively to obtain the final intelligent fault diagnosis model. Step 93: During the operation of the HVAC and refrigeration system, a fault diagnosis model is deployed and real-time data of the system operation is input into the fault diagnosis model. The model judges the system status based on the input and reports whether the system is in a normal state or what kind of fault has occurred.
Citation Information
Patent Citations
Feature weighted PCA face recognition method based on average influence value data transformation
CN111914718A
Sensor optimization selection method for fault diagnosis of water chilling unit
CN112990272A