Control valve fault diagnosis method based on dynamic incremental ensemble learning model
Through the dynamic incremental integrated learning model combined with multi-domain feature fusion and intelligent preprocessing, the noise and data quality problems of control valves in industrial environments are solved, efficient and reliable fault diagnosis is achieved, and the accuracy and real-time performance of fault diagnosis are improved.
Patent Information
- Application Number
- CN202510657182.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-09-05
AI Technical Summary
Existing technologies are unable to effectively deal with the degradation of data quality caused by noise, missing values and non-stationary distribution of control valves in industrial environments, as well as the problems of feature characterization and model redundancy in traditional methods, which affect the accuracy and reliability of fault diagnosis.
A method based on a dynamic incremental ensemble learning model is adopted. By fusing features in the time, frequency and spatial domains, combined with preprocessing of sliding median filtering, dynamic normalization and cubic spline interpolation, features are extracted using a one-dimensional convolutional neural network, Transformer encoder and principal component analysis, and fault classification is performed through a dynamic incremental ensemble learning model.
It significantly improves the accuracy and reliability of control valve fault diagnosis, can effectively suppress noise interference, enhance feature characterization capabilities, and balance model complexity and computational efficiency in scenarios with high real-time requirements.
Smart Images

Figure CN120597149A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of control valve fault diagnosis, and more particularly to a control valve fault diagnosis method based on a dynamic incremental ensemble learning model. Background Art
[0002] Control valves, as key components in industrial automation systems, regulate parameters such as flow, pressure, and direction of fluids and are widely used in industries such as chemical, petroleum, and electric power. As industrial production processes become increasingly complex, the requirements for control valve reliability and precision are also increasing. However, over long-term operation, control valves may be susceptible to various faults, such as a drop in air pressure and a loose valve packing cover. If these problems are not diagnosed and resolved promptly, they will seriously affect production continuity and safety.
[0003] Industrial time series data often contains noise, missing values, and non-stationary distributions. Existing preprocessing methods often use fixed strategies that struggle to cope with complex data characteristics. For example, forcing Z-score normalization on skewed data can distort the data distribution, rendering subsequent feature extraction ineffective. Traditional interpolation methods also lack the accuracy to fill non-uniform missing values, further degrading data quality.
[0004] At the same time, existing methods often rely on time-domain statistical features or single-frequency analysis, lacking the deep integration of multi-domain features. For example, using only time-domain features may overlook the frequency band energy distribution characteristics of the fault signal, while pure frequency-domain analysis struggles to capture local time series dynamics, resulting in insufficient feature representation capabilities and affecting classification accuracy.
[0005] Furthermore, traditional ensemble learning methods typically use a fixed combination of base learners, making it difficult to dynamically adjust the model structure based on data distribution. This can lead to model redundancy or insufficient generalization. Especially in scenarios with high real-time requirements, fixed models may be undeployable due to excessive inference latency. Furthermore, they lack a dynamic weighting mechanism for the contributions of base learners, making it difficult to balance accuracy and computational efficiency.
[0006] Therefore, how to design a control valve fault diagnosis method based on a dynamic incremental ensemble learning model that can solve industrial data noise interference, one-sided feature representation, and model redundancy defects, and improve the accuracy and reliability of control valve fault diagnosis is an urgent problem that technical personnel in this field need to solve. Summary of the Invention
[0007] In view of this, the present invention provides a control valve fault diagnosis method based on a dynamic incremental ensemble learning model, which improves the comprehensiveness of diagnosis by fusing time domain, frequency domain and space domain features; introduces a dynamic incremental ensemble learning model to optimize model efficiency and accuracy, and is suitable for scenarios with high real-time requirements; and adopts an intelligent preprocessing strategy based on data distribution to effectively cope with complex data distribution, thereby improving the accuracy and reliability of control valve fault diagnosis as a whole.
[0008] In order to achieve the above object, the present invention adopts the following technical solutions:
[0009] A control valve fault diagnosis method based on a dynamic incremental ensemble learning model comprises the following steps:
[0010] S1. Obtaining time series data of target valve position displacement, actual valve position displacement, and displacement deviation under various fault status types;
[0011] S2. Preprocess the time series data to obtain preprocessed time series data;
[0012] S3. Based on the preprocessed time series data, extract time domain features, frequency domain features, and spatial domain features to generate a fused feature vector;
[0013] S4. Input the fused feature vector into a dynamic incremental ensemble learning model to perform fault classification and generate a fault diagnosis result.
[0014] Furthermore, in S1, the fault state types include: gas source pressure drop, gas pipeline damage, guide sleeve mismatch, gas pipeline flattening, valve packing cover loosening, impurity precipitation in the valve core and valve seat, actuator diaphragm wear, spring aging and upper valve cover loosening.
[0015] Furthermore, the S2 includes:
[0016] S21. Use sliding median filtering to remove noise from the time series data. The sliding window length is 15, and the filtering formula is:
[0017] S'(t)=median{S(t-7),S(t-6),…,S(t),…,S(t+7)}
[0018] Where S'(t) represents the output value after filtering at time point t, and median{} represents the median operation;
[0019] S22. Normalize the median filtered data. The normalization includes dynamically selecting a normalization method based on the skewness γ1 and kurtosis γ2 of the data distribution. If |γ1|>1 or |γ2|>3, use Min-Max normalization; otherwise, use Z-score normalization.
[0020] S23. Combine the cubic spline interpolation algorithm to fill the missing values of the normalized data:
[0021] S”(t)=a(tt i ) 3 +b(tt i ) 2 +c(tt i )+d
[0022] Among them, S”(t) represents the function value after interpolation, t i represents the adjacent valid data points at time point t, and a, b, c, and d represent the polynomial coefficients.
[0023] Furthermore, in S3, performing time domain feature extraction includes:
[0024] S311, using one-dimensional convolutional neural network to extract local features of time series data;
[0025] S312, performing global temporal modeling on the local features through a Transformer encoder to generate a feature map;
[0026] S313: Perform weighted fusion on the feature maps in combination with the Attention mechanism to generate a time-domain feature vector.
[0027] Furthermore, in S3, performing frequency domain feature extraction includes:
[0028] S321, using Morlet wavelet basis function to convert time series data into frequency domain;
[0029] S322, dividing the frequency domain into a preset number of frequency bands, and calculating the energy value of each frequency band;
[0030] S323 . Based on the energy value of each frequency band, extract the energy proportion of each frequency band, the spectrum centroid, and the frequency band energy ratio to form a frequency domain feature vector.
[0031] Furthermore, in S3, performing spatial domain feature extraction includes:
[0032] S331. Perform principal component analysis on the time series data to generate a corresponding covariance matrix;
[0033] S332, performing eigenvalue decomposition on the covariance matrix, retaining the eigenvectors corresponding to the first M largest eigenvalues, and constructing a principal component space;
[0034] S333. Project the corresponding original time series data into the principal component space to generate a spatial domain feature vector.
[0035] Furthermore, the S4 includes:
[0036] S41. Based on the preset candidate base learners, the accuracy of each base learner is evaluated through cross-validation, and the base learner is selected in combination with the preset dynamic addition conditions until any termination condition is met;
[0037] S42. Generate a weighted probability prediction result based on the selected base learner, and fuse the confidence features to construct a secondary feature vector;
[0038] S43. Adopting an adaptive logistic regression model as a meta-learner and combining it with a dynamic regularization parameter adjustment strategy, the secondary feature vectors are classified and the fault diagnosis results are generated.
[0039] Furthermore, in S41, the base learner selection is performed in combination with the preset dynamic addition condition, including:
[0040] Each time a base learner is added, the accuracy improvement value Δ of the integrated model is calculated through cross-validation. If Δ ≥ 1%, it is retained, otherwise it is skipped.
[0041] Furthermore, in S41, the termination condition includes:
[0042] The number of base learners reaches 4, there are multiple base learners with accuracy ≥ 95%, and the newly added base learners cause a single inference delay ≥ 100ms.
[0043] Furthermore, in S43, the dynamic regularization parameter adjustment strategy is expressed as:
[0044]
[0045] Among them, α represents the positive influence coefficient, β represents the reverse influence coefficient, o represents the number of base learners, and p represents the number of training samples.
[0046] It can be seen from the above technical solution that compared with the prior art, the technical solution of the present invention has the following advantages:
[0047] Beneficial effects:
[0048] 1. This method dynamically selects a normalization method based on data skewness and kurtosis, and integrates sliding median filtering with cubic spline interpolation to effectively suppress noise interference and fill missing values. This improves data quality and provides more robust input for subsequent feature extraction and classification. It is particularly suitable for non-stationary, high-noise time series data commonly found in industrial environments.
[0049] 2. It integrates time, frequency, and spatial domain features, comprehensively capturing control valve fault information through multi-level feature extraction methods such as one-dimensional convolutional neural networks, Transformer encoders, and principal component analysis. The combination of local time domain features and global time series modeling, supplemented by frequency domain energy distribution and spatial domain dimensionality reduction projection, significantly enhances feature representation capabilities, effectively addressing the issues of missed detection or misjudgment caused by traditional methods due to the single nature of the features.
[0050] 3. A dynamic incremental ensemble learning model is further employed for fault classification. This model introduces a dynamic incremental mechanism, evaluates the accuracy improvement of base learners through cross-validation, and adaptively adjusts the model structure based on multiple termination criteria, such as the number of base learners and inference latency. Furthermore, a dynamic regularization strategy using weighted probability prediction and adaptive logistic regression is employed to balance model complexity and computational efficiency while ensuring high accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0052] Figure 1 A flow chart of a control valve fault diagnosis method based on a dynamic incremental ensemble learning model provided by an embodiment of the present invention;
[0053] Figure 2 A schematic diagram of a control valve fault diagnosis process for a high-temperature and high-pressure steam system in a power plant according to an embodiment of the present invention;
[0054] Figure 3 A schematic diagram of a fusion feature vector processing process based on a selected base learner provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0055] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0056] Example 1;
[0057] like Figure 1 As shown, this embodiment provides a control valve fault diagnosis method based on a dynamic incremental ensemble learning model, comprising the following steps:
[0058] S1. Obtaining time series data of target valve position displacement, actual valve position displacement, and displacement deviation under various fault status types;
[0059] S2. Preprocess the time series data to obtain preprocessed time series data;
[0060] S3. Based on the preprocessed time series data, extract time domain features, frequency domain features, and spatial domain features to generate a fused feature vector;
[0061] S4. Input the fused feature vector into a dynamic incremental ensemble learning model to perform fault classification and generate a fault diagnosis result.
[0062] This approach effectively addresses industrial data noise, incomplete feature representation, and model redundancy through intelligent adaptive preprocessing, deep fusion of time-frequency-space multi-domain features, and a dynamic incremental ensemble learning model. It significantly improves fault classification accuracy, enhances the algorithm's robustness to non-stationary data, and achieves highly reliable diagnosis under real-time constraints, providing an efficient and lightweight intelligent fault diagnosis solution for complex industrial scenarios.
[0063] In this embodiment S1, time series data of target valve position displacement, actual valve position displacement and displacement deviation under various fault state types are obtained; the fault state types include: gas source pressure drop, gas pipeline damage, guide sleeve mismatch, gas pipeline flattening, valve packing cover loosening, impurity precipitation of valve core and valve seat, actuator diaphragm wear, spring aging and upper valve cover loosening.
[0064] Specifically, at a sampling frequency of 100 Hz, the target valve position displacement Starget(t), actual valve position displacement Sactual(t) and displacement deviation ΔS(t) = Starget(t) - Sactual(t) under each fault state are collected; wherein, the single collection time is 2.1 seconds, forming a time series data with a length of 210.
[0065] The various fault status types here cover the typical failure modes of the control valve mechanical structure, pneumatic system and sealing components. Their selection is based on industrial field fault statistical reports and failure mode analysis. For example, the flattening of the gas pipeline will cause a delay in the transmission of air pressure, which manifests as an abnormal step response of the actual valve position displacement; while the wear of the actuator diaphragm may cause high-frequency jitter of the displacement deviation, etc.; different fault types have different representations in the feature space: the drop in gas source pressure mainly affects the time domain mean, while the precipitation of impurities in the valve core will cause the frequency domain energy to concentrate in the low-frequency band. In this step, by clarifying the physical meaning and data characteristics of the fault type, a clearly labeled sample set is provided for model training.
[0066] In this embodiment S2, the time series data is preprocessed to obtain preprocessed time series data; specifically, the following steps are performed:
[0067] S21. Use sliding median filtering to remove noise from the time series data. The sliding window length is 15, and the filtering formula is:
[0068] S'(t)=median{S(t-7),S(t-6),…,S(t),…,S(t+7)}
[0069] Where S'(t) represents the output value after filtering at time point t, and median{} represents the median operation;
[0070] S22. Normalize the median filtered data. The normalization includes dynamically selecting a normalization method based on the skewness γ1 and kurtosis γ2 of the data distribution. If |γ1|>1 or |γ2|>3, use Min-Max normalization; otherwise, use Z-score normalization.
[0071] S23. Combine the cubic spline interpolation algorithm to fill the missing values of the normalized data:
[0072] S”(t)=a(tt i ) 3 +b(tt i ) 2 +c(tt i )+d
[0073] Among them, S”(t) represents the function value after interpolation, t i represents the adjacent valid data points at time point t, and a, b, c, and d represent the polynomial coefficients.
[0074] The intelligent design of the pre-processing process in this embodiment is reflected in two aspects: first, the window length of the sliding median filter is determined based on the spectrum analysis of the control valve signal sampling frequency and the typical noise period, ensuring that high-frequency noise is filtered out while retaining the real signal mutation; second, the dynamic selection strategy of the normalization method distinguishes the data distribution type through the threshold judgment of skewness and kurtosis. For example, for right-skewed displacement deviation data, the use of Min-Max normalization can avoid the negative value problem caused by mean shift in Z-score normalization; and for data close to a normal distribution, Z-score can better preserve the distribution characteristics. In addition, the polynomial coefficients of the cubic spline interpolation are solved by continuously constraining the derivatives of adjacent valid data points. Compared with linear interpolation, it can more smoothly fill the long-term missing segments caused by sensor disconnection and reduce the impact of interpolation errors on feature extraction.
[0075] In this embodiment, S3, based on the preprocessed time series data, time domain features, frequency domain features, and spatial domain features are extracted to generate a fused feature vector;
[0076] The time domain feature extraction includes:
[0077] S311, using one-dimensional convolutional neural network to extract local features of time series data;
[0078] S312, performing global temporal modeling on the local features through a Transformer encoder to generate a feature map;
[0079] S313: Perform weighted fusion on the feature maps in combination with the Attention mechanism to generate a time-domain feature vector.
[0080] The network parameters of the one-dimensional convolutional neural network here are: convolution kernel size = 5, stride = 1, number of output channels = 32; the activation function is ReLU, and the output feature map size is 210×32; it can capture the short-term fluctuation pattern of valve position displacement; the number of encoding layers of the Transformer encoder is = 4, each layer contains 8 attention heads, and long-term dependencies are modeled through the self-attention mechanism; the weighted fusion strategy of the Attention mechanism on the feature map can assign higher weights to fault-sensitive periods and enhance the significance of key features.
[0081] Furthermore, frequency domain feature extraction includes:
[0082] S321, using Morlet wavelet basis function to convert time series data into frequency domain;
[0083] S322, dividing the frequency domain into a preset number of frequency bands, and calculating the energy value of each frequency band;
[0084] S323 . Based on the energy value of each frequency band, extract the energy proportion of each frequency band, the spectrum centroid, and the frequency band energy ratio to form a frequency domain feature vector.
[0085] The center frequency of the Morlet wavelet basis function is set to 0.5Hz-10Hz, covering the main frequency band of the control valve action; the frequency domain division adopts a logarithmic scale (5 frequency bands: 0-1Hz, 1-2Hz, 2-4Hz, 4-8Hz, 8-10Hz) to match the frequency band characteristics of the fault signal.
[0086] Furthermore, spatial domain feature extraction includes:
[0087] S331. Perform principal component analysis on the time series data to generate a corresponding covariance matrix;
[0088] S332, performing eigenvalue decomposition on the covariance matrix, retaining the eigenvectors corresponding to the first M largest eigenvalues, and constructing a principal component space;
[0089] S333. Project the corresponding original time series data into the principal component space to generate a spatial domain feature vector.
[0090] In principal component analysis, the basis for retaining the first M principal components is that the cumulative variance contribution rate is ≥85%, which reduces the data dimension while retaining the main variation information; after projecting into the principal component space, the clustering of samples of different fault categories is significantly improved, which is conducive to subsequent classification.
[0091] In this embodiment, S4, the fused feature vector is input into a dynamic incremental ensemble learning model to perform fault classification and generate a fault diagnosis result. Specifically, the following steps are performed:
[0092] S41. Based on the preset candidate base learners, the accuracy of each base learner is evaluated through cross-validation, and the base learner is selected in combination with the preset dynamic addition conditions until any termination condition is met;
[0093] Furthermore, the system selects base learners based on preset dynamic addition conditions. For each new base learner, the accuracy improvement Δ of the ensemble model is calculated through cross-validation. If Δ ≥ 1%, the ensemble is retained; otherwise, it is skipped. Specific termination conditions include: the number of base learners reaches 4, the accuracy of multiple base learners is ≥ 95%, and the addition of a new base learner causes a single inference delay of ≥ 100ms.
[0094] S42. Generate a weighted probability prediction result based on the selected base learner, and fuse the confidence features to construct a secondary feature vector;
[0095] S43. Adopting an adaptive logistic regression model as a meta-learner and combining it with a dynamic regularization parameter adjustment strategy, the secondary feature vectors are classified and the fault diagnosis results are generated.
[0096] Among them, the dynamic regularization parameter adjustment strategy is expressed as:
[0097]
[0098] Among them, α represents the positive influence coefficient, β represents the reverse influence coefficient, o represents the number of base learners, and p represents the number of training samples.
[0099] Here, α and β are optimized to α = 0.02 and β = 0.5 through grid search, so that as the number of base learners increases, the regularization strength increases linearly to prevent overfitting; at the same time, as the sample size increases, the reverse adjustment term gradually weakens to avoid underfitting of the model.
[0100] This embodiment proposes a control valve fault diagnosis method based on a dynamic incremental ensemble learning model. By acquiring time series data under various fault conditions and undergoing intelligent adaptive preprocessing, deep fusion of time-frequency-space multi-domain features, and fault classification based on a dynamic incremental ensemble learning model, it effectively solves the problems of industrial data noise interference, one-sided feature representation, and model redundancy.
[0101] Preprocessing methods such as sliding median filtering, dynamic normalization, and cubic spline interpolation were further adopted to improve data quality. Multi-domain features were extracted using one-dimensional convolutional neural networks, Transformer encoders, and principal component analysis techniques to enhance feature representation capabilities. Finally, through a dynamic incremental ensemble learning model, combined with a heterogeneous basis learner and an adaptive logistic regression meta-learner, the model complexity and computational efficiency were balanced while ensuring high accuracy, providing an efficient and reliable intelligent fault diagnosis solution for complex industrial scenarios.
[0102] Example 2;
[0103] In high-temperature, high-pressure steam systems in power plants, impurity deposits on control valve cores, wear on actuator diaphragms, and concurrent faults can easily lead to cascading equipment shutdowns. Traditional threshold alarm mechanisms, unable to decouple complex fault characteristics, have a high false alarm rate. This embodiment addresses this scenario by capturing transient anomalies through high-frequency data acquisition. It then combines multi-dimensional feature extraction with a dynamic incremental ensemble learning model to achieve precise isolation and real-time diagnosis of concurrent faults.
[0104] like Figure 2 As shown, the specific implementation steps include:
[0105] 1) Acquisition of time series data under fault conditions;
[0106] The sampling frequency is set to 200H, and the acquisition time is 3 seconds, which can cover the valve action cycle; the data obtained include: target valve position displacement: collected through the PLC control system, with a measurement range of 0-100% (corresponding to valve opening 0-50mm), and a resolution of ±0.1%; actual valve position displacement: using LVDT displacement sensor, with a range of 0-50mm, an accuracy of ±0.05mm, and a sampling frequency of 200Hz; displacement deviation: calculated in real time as (target valve position - actual valve position), in mm, with an accuracy of ±0.05mm.
[0107] Each sampling generates 600 data points, which are stored in CSV format and contain three columns: timestamp, target displacement, and actual displacement. The data validity is verified through offline calibration experiments.
[0108] 2) Intelligent data preprocessing;
[0109] A sliding median filter with a 25-point window length was used to remove high-frequency vibration noise while preserving the gradual trend of valve core sedimentation. A dynamic normalization strategy selected either Min-Max or Z-score processing based on data skewness and kurtosis to avoid distribution distortion. For example, right-skewed sedimentation data was normalized to [0, 1], while near-normal wear data was normalized to a mean of 0 and a variance of 1.
[0110] Furthermore, for the missing segments caused by sensor disconnection, cubic spline interpolation is used to force the interpolation curve to have continuous second-order derivatives in the jitter area, and the interpolation error MSE ≤ 0.05, ensuring the reliability of subsequent feature extraction.
[0111] 3) Multi-domain feature fusion and enhancement;
[0112] First, a coordinated extraction process across the time, frequency, and spatial domains is performed: a one-dimensional CNN extracts microsecond-level jitter associated with diaphragm wear, a Transformer encoder models the long-term trend of valve core sedimentation, and an attention mechanism assigns a double weight to periods of sudden changes. The Morlet wavelet is then divided into seven frequency bands to extract the frequency band energy ratios associated with concurrent faults. In the principal component analysis, PCA retains the first four principal components.
[0113] Furthermore, the time domain, frequency domain, and spatial domain features are spliced into a multi-dimensional fusion vector, and the redundant information is compressed through L2 regularization and input into the dynamic incremental integration model.
[0114] 4) Dynamic incremental integrated model construction and diagnosis;
[0115] In this embodiment, the candidate base learners in the dynamic incremental ensemble model include logistic regression, decision tree, support vector machine, K-nearest neighbor, random forest, XGBoost, and improved Transformer classifier, a total of seven heterogeneous models; their heterogeneity is intended to improve the generalization ability and robustness of the ensemble model through diversity. The base learner screening and optimization based on these seven heterogeneous models includes the following specific steps:
[0116] First, in the first round of cross-validation, the random forest model improved its accuracy for compound faults by Δ = 2.5%, meeting the Δ ≥ 1% threshold and thus being retained. Through feature importance analysis, the random forest model effectively identified the low-frequency energy surges caused by impurity deposition while also capturing the high-frequency jitter characteristics of diaphragm wear.
[0117] Secondly, the accuracy increased by Δ=1.8% after adding XGBoost. Its gradient boosting mechanism optimized the interactive modeling of frequency domain and spatial domain features, especially in the case of diaphragm wear fault alone, the recall rate was increased to 93%.
[0118] When the improved Transformer was further added, Δ=1.2%. Its causal convolutional layer accurately captured the slow temporal drift trend caused by impurity deposits on the valve core, while increasing the single inference latency by only 15ms. At this point, the number of base learners reached three. The system detected that the individual accuracies of Random Forest and XGBoost were 96% and 95%, respectively, meeting the high-precision termination criteria. Therefore, the screening process was terminated early, ultimately focusing on Random Forest, XGBoost, and the improved Transformer.
[0119] In addition, the decision tree and K-nearest neighbor were skipped because their Δ was only 0.6% and 0.4% respectively. Although the support vector machine (SVM) had Δ = 1.1%, the inference delay increased to 102ms after adding it, triggering the delay termination condition.
[0120] like Figure 3 As shown in the figure, the final integration of three base learners is random forest, XGBoost and improved Transformer classifier. Among them, the improved Transformer classifier includes an input layer, a multi-head self-attention layer, a feedforward neural network layer, layer normalization, gated residual connections and an output layer. Based on the output results of different models, weighted probability prediction results are further generated, and confidence features are integrated to construct secondary feature vectors.
[0121] Furthermore, as shown in Table 1 below, based on the secondary eigenvectors, an adaptive logistic regression model is used as a meta-learner to generate fault diagnosis results. The entropy weight method is used to assign weights: random forest: 0.4, XGBoost: 0.35, Transformer: 0.25; and dynamic regularization parameters are combined to prevent overfitting, and the fault type and confidence level are finally output.
[0122] Table 1
[0123]
[0124] In this example, a 200Hz sampling frequency is used to acquire time series data within the valve's operating cycle, and intelligent preprocessing techniques such as sliding median filtering and dynamic normalization are used to remove noise. Subsequently, multi-domain feature extraction techniques such as a one-dimensional CNN and a Transformer encoder are used to enhance feature representation. These features are then fed into an integrated model consisting of a random forest, XGBoost, and an improved Transformer for fault diagnosis. Ultimately, high-precision identification of valve core impurity deposits, diaphragm wear, and their associated faults is achieved, with an accuracy rate of 97.1% and a single inference latency of 83ms.
[0125] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. References to the same or similar parts between the various embodiments are sufficient. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For relevant parts, refer to the method description.
[0126] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A control valve fault diagnosis method based on a dynamic incremental ensemble learning model, characterized in that: The following steps are involved: S1. Obtaining time series data of target valve position displacement, actual valve position displacement, and displacement deviation under various fault status types; S2. Preprocess the time series data to obtain preprocessed time series data; S3. Based on the preprocessed time series data, extract time domain features, frequency domain features, and spatial domain features to generate a fused feature vector; S4. Input the fused feature vector into a dynamic incremental ensemble learning model to perform fault classification and generate a fault diagnosis result.
2. A control valve fault diagnosis method based on a dynamic incremental ensemble learning model according to claim 1, characterized in that: In S1, the fault status types include: gas source pressure drop, gas pipeline damage, guide sleeve mismatch, gas pipeline flattening, valve packing cover loosening, impurity precipitation of valve core and valve seat, actuator diaphragm wear, spring aging and upper valve cover loosening.
3. The control valve fault diagnosis method based on dynamic incremental ensemble learning model according to claim 1 is characterized in that: The S2 includes: S21. Use sliding median filtering to remove noise from the time series data. The sliding window length is 15, and the filtering formula is: S'(t)=median{S(t-7),S(t-6),…,S(t),…,S(t+7)} Where S'(t) represents the output value after filtering at time point t, and median{} represents the median operation; S22. Normalize the median filtered data. The normalization includes dynamically selecting a normalization method based on the skewness γ1 and kurtosis γ2 of the data distribution. If |γ1|>1 or |γ2|>3, use Min-Max normalization; otherwise, use Z-score normalization. S23. Combine the cubic spline interpolation algorithm to fill the missing values of the normalized data: S”(t)=a(t-t i ) 3 +b(t-t i ) 2 +c(t-t i )+d Among them, S”(t) represents the function value after interpolation, t i represents the adjacent valid data points at time point t, and a, b, c, and d represent the polynomial coefficients.
4. The control valve fault diagnosis method based on dynamic incremental ensemble learning model according to claim 1 is characterized in that: In S3, performing time domain feature extraction includes: S311, using one-dimensional convolutional neural network to extract local features of time series data; S312, performing global temporal modeling on the local features through a Transformer encoder to generate a feature map; S313: Perform weighted fusion on the feature maps in combination with the Attention mechanism to generate a time-domain feature vector.
5. The control valve fault diagnosis method based on dynamic incremental ensemble learning model according to claim 1 is characterized in that: In S3, frequency domain feature extraction includes: S321, using Morlet wavelet basis function to convert time series data into frequency domain; S322, dividing the frequency domain into a preset number of frequency bands, and calculating the energy value of each frequency band; S323 . Based on the energy value of each frequency band, extract the energy proportion of each frequency band, the spectrum centroid, and the frequency band energy ratio to form a frequency domain feature vector.
6. The control valve fault diagnosis method based on dynamic incremental ensemble learning model according to claim 1 is characterized in that: In S3, performing spatial domain feature extraction includes: S331. Perform principal component analysis on the time series data to generate a corresponding covariance matrix; S332, performing eigenvalue decomposition on the covariance matrix, retaining the eigenvectors corresponding to the first M largest eigenvalues, and constructing a principal component space; S333. Project the corresponding original time series data into the principal component space to generate a spatial domain feature vector.
7. The control valve fault diagnosis method based on dynamic incremental ensemble learning model according to claim 1 is characterized in that: Said S4 comprises: S41. Based on the preset candidate base learners, the accuracy of each base learner is evaluated through cross-validation, and the base learner is selected in combination with the preset dynamic addition conditions until any termination condition is met; S42. Generate a weighted probability prediction result based on the selected base learner, and fuse the confidence features to construct a secondary feature vector; S43. Adopting an adaptive logistic regression model as a meta-learner and combining it with a dynamic regularization parameter adjustment strategy, the secondary feature vectors are classified and the fault diagnosis results are generated.
8. The control valve fault diagnosis method based on dynamic incremental ensemble learning model according to claim 7 is characterized in that: In S41, the base learner selection is performed in combination with the preset dynamic addition condition, including: Each time a base learner is added, the accuracy improvement value Δ of the integrated model is calculated through cross-validation. If Δ ≥ 1%, it is retained, otherwise it is skipped.
9. The control valve fault diagnosis method based on dynamic incremental ensemble learning model according to claim 7 is characterized in that: In S41, the termination conditions include: The number of base learners reaches 4, there are multiple base learners with accuracy ≥ 95%, and the newly added base learners cause a single inference delay ≥ 100ms.
10. The control valve fault diagnosis method based on dynamic incremental ensemble learning model according to claim 7, characterized in that: In S43, the dynamic regularization parameter adjustment strategy is expressed as: Among them, α represents the positive influence coefficient, β represents the reverse influence coefficient, o represents the number of base learners, and p represents the number of training samples.
Citation Information
Cited By
Airborne radio frequency system filter circuit fault diagnosis method based on confidence fusion
CN120804900A