A thermal power plant primary air fan health degree evaluation method based on data fusion

CN119740183BActive Publication Date: 2026-09-29NANJING INST OF MECHATRONIC TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411555436.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-04
Publication Date
2026-09-29
Estimated Expiration
2044-11-04

AI Technical Summary

Technical Problem

传统上,仅仅依靠历史数据建立的健康度评价与故障预警模型难以适应实际的工业应用场景,需要解决模型在线自适应更新的问题,进一步提高健康度评估与故障预警方法的实时性、准确性与可靠性

Benefits of technology

[0080]本发明提供的一种基于改进多元状态估计的大型火力发电厂燃煤机组一次风机健康度评估方法及系统,针对一次风机需要大量的故障数据进行训练以及需要将故障数据划分为不同的故障、依靠单一变量阈值进行故障评估的局限性以及故障诊断模型的在线更新问题,通过对历史数据进行数据清洗后和稳态筛选后,采用概率密度等间隔抽样方法对总体进行抽样,采用模糊聚类确定每个区间间隔上的状态向量,选取典型状态向量构成记忆矩阵,并采用k最近邻算法进行更新记忆矩阵。使用设备较新的数据替代历史训练集合中质量较差的数据,提高了模型的在线自适应能力。采用基于滑动窗口状态相似度代替传统残差阈值法作为预警的评价标准。本发明采用改进的多元状态估计方法,对噪声具有较强的鲁棒性,能够有效提高一次风机状态监测的精度和可靠性。通过综合分析多种运行参数,建立多元状态估计模型,改进的多元状态估计技术能够可以实时估计和预测一次风机的状态,更准确地反映一次风机的实际健康状况,降低了误判和漏判的风险,为运行人员提供及时准确的设备状态信息。通过设定健康度评分阈值,系统能够在故障发生前发出预警,保障发电厂的安全运行和经济效益。该系统具有良好的适应性和可扩展性,可以根据不同类型的一次风机和运行环境进行调整和优化,适用于各种大型火力发电厂燃煤机组。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119740183B_ABST
    Figure CN119740183B_ABST
Patent Text Reader

Abstract

The application provides a health degree evaluation method for a thermal power plant primary air fan based on data fusion, samples the population by adopting a probability density equal-interval sampling method after data cleaning and steady-state screening of historical data, determines a state vector on each interval interval by adopting fuzzy clustering, selects a typical state vector to form a memory matrix to effectively cover a common working space, and realizes dynamic updating of the memory matrix by adopting a k nearest neighbor algorithm, so that the model has improved online self-adapting capability by replacing poor-quality data in a historical training set with new data of the equipment. The system has good adaptability and scalability, and has a high application and promotion prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of health assessment technology for primary air fans in thermal power plants, specifically a data fusion-based method for assessing the health of primary air fans in thermal power plants. Background Technology

[0002] Primary air fans are crucial large auxiliary machines in the flue gas systems of traditional coal-fired power plant boilers, primarily providing primary air for combustion. In pulverized coal boilers, primary air fans provide the power source for primary air, regulate combustion conditions, and maintain stable operation of the pulverizing system. The function of primary air is to carry pulverized coal into the furnace for combustion. It delivers air to the pulverizing system, allowing the pulverized coal to be carried into the furnace through pipelines and other conveying devices. By adjusting the airflow and pressure of the primary air fan, the amount of primary air entering the furnace can be controlled. Appropriate primary airflow is essential for stable combustion of pulverized coal. In the pulverizing system, primary air also plays a role in drying and conveying pulverized coal. Primary air fans operate in harsh environments and have a high failure rate. As a critical auxiliary machine, the operating status of the primary air fan directly affects the safety and economy of power plant unit operation. Early health assessments and fault warning systems can help operators detect faults early and prevent further escalation.

[0003] Traditionally, data-driven health assessment and fault early warning identification methods require extensive fault data for training, sometimes even necessitating the classification of fault data into different categories. However, in real-world primary wind turbine fault diagnosis scenarios, collecting and classifying large amounts of fault data is challenging. Firstly, obtaining a sufficient number of fault samples is often difficult; in actual industrial settings, primary wind turbines mostly operate in a healthy state, with very low operating frequency in fault states, resulting in limited fault sample data. Secondly, the required prior knowledge of fault mechanisms and calculations of equipment state transitions are overly complex. Most existing methods in the industry require separate modeling of multiple fault-related variables for primary wind turbines, which is too complex for industrial applications. Over time, primary wind turbine equipment gradually ages, and the initially constructed fault diagnosis model is based on the equipment's initial state or data from a certain period. As the equipment ages, newly emerging fault characteristics may not be included in the original model. The operating environment of primary wind turbines is complex and dynamically changing; changes in these environmental factors lead to changes in fault types and probabilities, necessitating updates to the fault diagnosis model to adapt to the new operating environment. Furthermore, power plants may adopt new technologies or upgrade equipment. The data characteristics and operating modes brought about by these new technologies and equipment differ from the data and modes on which the original model was based. Updating the fault diagnosis model allows it to fully utilize the new data information. Traditionally, health assessment and fault early warning models built solely on historical data are difficult to adapt to actual industrial application scenarios. It is necessary to solve the problem of online adaptive updating of the model to further improve the real-time performance, accuracy, and reliability of health assessment and fault early warning methods.

[0004] The following is a comparison with existing technologies:

[0005] Technical comparison with patent CN112067335A "A method for early warning of power plant blower faults based on multivariate state estimation"

[0006] The evaluation metrics of patent CN112067335A mainly focus on the deviation of historical data, using a deviation function to determine fault warnings. In contrast, this patent's evaluation metrics not only include health scores but also employ a sliding window state similarity method instead of the traditional residual threshold method, considering multiple aspects of working conditions and environmental influences. The two patents differ fundamentally in their approach to setting metrics.

[0007] Patent CN112067335A uses a memory matrix built based on historical data, relies on deviation thresholds to achieve early warning, and emphasizes quantitative calculation; while this patent uses fuzzy clustering and the k-nearest neighbor algorithm to achieve dynamic updates, combining quantitative and qualitative evaluations to form a comprehensive health assessment standard. The two are fundamentally different in their overall evaluation approach and implementation method.

[0008] Patent CN112067335A is mainly applied to power plant blowers, focusing on fault early warning of blowers; while this patent is applied to the health assessment and fault early warning of primary air blowers in coal-fired power plants, focusing on the overall operating status of the equipment. The two have fundamental differences in the objects of evaluation.

[0009] Comparison with patent CN118410365A "Wind turbine fault monitoring and diagnosis method based on online learning and multivariate state estimation"

[0010] The evaluation indicators of patent CN118410365A mainly focus on the preprocessing and monitoring of fault samples, emphasizing the handling of outliers in the data; while the evaluation indicators of this patent cover a wider range of fields, including equipment health scores, state similarity, etc., forming a more comprehensive evaluation system. The two are fundamentally different in terms of the scope of their indicators.

[0011] Patent CN118410365A employs online learning and the DBSCAN algorithm for data processing, and utilizes an adaptive similarity threshold for fault warning, emphasizing real-time performance and data accuracy. In contrast, this patent achieves dynamic updates through an improved memory matrix and the k-nearest neighbor algorithm, combined with a sliding window method for state similarity evaluation, focusing on online adaptive capabilities. The two patents differ fundamentally in their technical implementation.

[0012] Patent CN118410365A is mainly used for wind turbine fault monitoring, focusing on the diagnosis and monitoring of wind turbines; while this patent is applicable to the primary air fans of coal-fired power plants, focusing on overall health assessment and fault early warning. The two are fundamentally different in their application areas and target objects. Summary of the Invention

[0013] To address the aforementioned technical problems, this invention proposes a data fusion-based method for assessing the health of primary air turbines in thermal power plants, aiming to overcome the shortcomings of existing methods and improve the operational safety and reliability of coal-fired power plants.

[0014] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0015] Multivariate state estimation (MSET) does not require an explicit model training process. It only uses historical normal data to construct a memory matrix and then calculates the similarity between the current state and the normal state. It has advantages such as low training cost, simple procedure, and strong generalization ability. This invention uses MSET for the health assessment and fault early warning of a primary wind turbine.

[0016] The estimation accuracy and computational efficiency of the MSET model depend on the memory matrix, which should consist of a sufficient number of historical vectors that are similar in state to the observed vectors. This goal is usually achieved by increasing the size of the memory matrix, but a large memory matrix can negatively impact the computational speed of matrix multiplication and nonlinear operations. Furthermore, an excessive number of samples with low similarity to the observed vectors leads to increased information redundancy, which in turn reduces the accuracy of the estimation. Theoretically, the ideal matrix consists of all historical data from normal operation, but in practical applications, the amount of historical data from normal operation is very large. An excessively large matrix will result in enormous computational demands, high computer performance requirements, and excessively long computation time, failing to meet the real-time requirements of industry. However, in most MSET applications, the process memory matrix is ​​constructed from historical observation vectors without a specific sampling method or the use of equally spaced sampling, which cannot guarantee sufficient sampling density for typical operating states. Furthermore, random sampling techniques involve randomly selecting some data from historical data to construct a memory matrix, but this method is highly random and cannot guarantee coverage of all states. In addition, due to changes in operating conditions and frequent maintenance that cause changes in objects, the accuracy of MSET estimation using a fixed memory matrix may be reduced, making it difficult to estimate equipment under various operating conditions.

[0017] Therefore, the memory matrix should cover a typical normal operating state space and minimize its size. Probability density reflects the probability of a variable value within a certain region; selecting the state vectors of the process memory matrix based on probability density can effectively cover a common working space. MSET prediction accuracy essentially depends on historical data similar to the observed data; therefore, effectively constructing the process memory matrix based on probabilistic features is beneficial for improving prediction accuracy. The unit load can represent the equipment's operating state. Probability density sampling at equal intervals based on the load is used to form the improved process memory matrix D. When selecting the state vectors, the aim is to remove redundant or similar state vectors. Probability density sampling at equal intervals is used to sample the population, and fuzzy clustering is employed to determine the state vectors at each interval. Typical state vectors are then selected to form the memory matrix.

[0018] This invention provides a data fusion-based method for assessing the health of primary air turbines in thermal power plants. After cleaning and steady-state screening of historical data, a probability density equal-interval sampling method is used to sample the population. State vectors of the process memory matrix are selected to effectively cover the common working space. Fuzzy clustering is used to determine the state vectors at each interval. Typical state vectors are selected to construct the memory matrix. Newer data from the equipment replaces lower-quality data in the historical training set, improving the model's online adaptive capability. The method includes the following steps:

[0019] S1: After performing dimensionality reduction, noise reduction, and data cleaning on historical data, steady-state data is selected to provide a dataset for the model;

[0020] S2: Using the maximum and minimum values ​​of the unit load as upper and lower limits, divide the filtered data into M equal intervals, denoted as [L1, L2, ..., L...]. i , ...L M ];

[0021] S3: Sampling is performed on the population using the probability density interval sampling method. The probability density over each interval is integrated to obtain the loading value of the state vector at the loading interval L. i The probability P within i Multiply by the total number K of state vectors in the process memory matrix to obtain the number K of state vectors in each load interval. i K i =P i ×K;

[0022] S4: In each load interval L i Within this range, fuzzy clustering is used to determine the state vector X for each interval. i Set the load interval L i The optimal number of clusters is set to K. i ;

[0023] S5: Each load interval L i The selected state vector X will be internally determined. i Add to the process memory matrix D until the number of Xi reaches Ki;

[0024] S6: Real-time acquisition of primary wind turbine data; calculation of the estimated vector X based on the acquired memory matrix using the observed vector. est The sliding window method is used to calculate the overall similarity of vectors. The overall similarity is the weighted average of the basic similarity and the state similarity. Considering the stability of the primary wind turbine equipment operation and the frequency of similarity changes, the current health of the primary wind turbine is calculated. The health is the weighted average of the overall similarity and the degree of similarity fluctuation.

[0025] S7: The memory matrix is ​​updated periodically using a K-nearest neighbor dynamic approach.

[0026] As a further improvement of the present invention, the steady-state data screening in step S1 is based on the autocorrelation function, and the calculation formula is as follows:

[0027]

[0028] Among them, X t For time series observations, Let be the mean of the observations, and k be the lag order.

[0029] As a further improvement of this invention, the dimensionality reduction process in step S1 is implemented based on the incremental PCA algorithm. When dealing with large-scale datasets, traditional PCA is computationally intensive because it requires processing the entire dataset at once. Incremental PCA, on the other hand, is a method that can handle streaming or large-scale data. It processes only one or a small batch of data samples at a time, gradually updating the principal components. For example, when processing real-time primary wind turbine sensor data, the data is continuously generated, and incremental PCA can perform dimensionality reduction on this new data in real time.

[0030] Includes the following steps:

[0031] S101: Initialization

[0032] Standardize the first batch of data so that the mean of each feature is 0 and the variance is 1;

[0033] Calculate the covariance matrix C of the first batch of data. Let the first batch of data matrix be X1, the sample size be n1, and the number of features be d. Then the covariance matrix C is...

[0034] Eigenvalue decomposition of the covariance matrix C yields eigenvalues ​​λ1, λ2, ..., λ d and the corresponding feature vectors v1, v2, ..., v d Sort the eigenvalues ​​from largest to smallest, and the corresponding eigenvectors will also be sorted accordingly;

[0035] Select the eigenvectors corresponding to the k largest eigenvalues ​​as the initial principal components, where k is the preset dimension after dimensionality reduction;

[0036] S102: Incremental Update

[0037] Let the new sample be x. When the new sample arrives, first calculate the new sample mean. in It is the mean of the previous n-1 samples, updating the covariance moments. Where C old It is the previous covariance matrix estimate;

[0038] S103: Update principal components

[0039] The updated covariance matrix is ​​decomposed into eigenvalues. The new eigenvalues ​​are sorted from largest to smallest, and the corresponding eigenvectors are also sorted accordingly. The eigenvectors corresponding to the k largest eigenvalues ​​are selected as the updated principal components.

[0040] S104: Repeated Update

[0041] As new data continues to arrive, it is continuously processed. Steps S102 and S103 are repeated to continuously update the covariance matrix and principal components, thereby achieving incremental principal component analysis of the data.

[0042] As a further improvement of the present invention, the probability density in step S3 adopts the formula of the normal distribution:

[0043]

[0044] Where μ is the mean and σ is the standard deviation.

[0045] The method for dynamically updating the memory matrix based on K-nearest neighbors is as follows: For the selected observation vectors as the distances to the sample to be classified, calculate the Euclidean distance between the selected observation vectors and all vectors in the training set. Sort the calculated Euclidean distances, find the k nearest training samples to the sample to be classified, count the categories of these k neighbors, and determine the category of the sample to be classified, i.e., belonging to [L1, L2, ..., L...]. i , ...L M The interval L i .

[0046] As a further improvement to the present invention, the detailed algorithm steps of fuzzy clustering in step S4 are as follows:

[0047] S401: Initialization: Select the number of clusters C, and initialize the fuzzy membership matrix U, where U ij Represents data point x j For the membership degree of cluster i, the membership value of each data point should be between [0,1], and the sum of all membership degrees for each data point j should be 1.

[0048] S402: Calculate cluster centers: Calculate the center V of each cluster based on the current membership matrix U. i :

[0049]

[0050] In the formula, m is the fuzzy factor, which takes values ​​between (1,2);

[0051] S403: Update the membership matrix: Calculate x for each data point. j Membership degree U for each cluster i ij ;

[0052]

[0053] S404: Determine convergence and calculate the change in the membership matrix. If the stopping condition is met, such as the change being less than a set threshold or the maximum number of iterations being reached, then stop the iteration; otherwise, return to step S402.

[0054] S405: Output the final cluster centers V and fuzzy membership matrix U.

[0055] As a further improvement of the present invention, the estimated vector X output by the model in step S6 est The weight vector W is obtained from the weight vector W, which represents the similarity between the current observation vector and the historical data memory matrix.

[0056]

[0057] Finally, estimate vector X est The expression is as follows:

[0058]

[0059] As a further improvement of the present invention, the overall similarity in step S6 is the weighted average of the basic similarity and the state similarity; the basic similarity refers to the observed vector X. obs With the estimated vector X est The degree of proximity is as follows:

[0060]

[0061] S(X o ,X e )∈(0,1]

[0062] In the formula α k This represents the weight of the k-th variable in the fault warning function. α k >0, S(X) o ,X e )∈(0,1], the larger S is, the closer the two are;

[0063] State similarity refers to the similarity of the observation vector X obs The degree of similarity to the memory matrix is ​​as follows:

[0064]

[0065] S(X o The larger S is, the closer the two are;

[0066] Overall similarity is the weighted average of basic similarity and state similarity:

[0067] S=β1S(X o ,X e )+β2S(X o D)

[0068] β1∈(0,1),β2∈(0,1)

[0069] S∈(0,1],

[0070] The larger the value of S, the higher the similarity between the primary wind turbine and its historical normal operating conditions.

[0071] Considering the stability of primary air turbine operation and the frequency of similarity changes, the health design is as follows:

[0072]

[0073] α1∈(0,1),α2∈(0,1)

[0074] H∈(0,1]

[0075] and Represent the change in overall similarity at time t and the total overall similarity, respectively.

[0076] For real-time acquired primary wind turbine data, a sliding window method is used to eliminate uncertainties and random interference during equipment operation. The window width is set to N, and the overall similarity is calculated once per cycle to obtain the overall health data [H1, H2, L, H] within the specified window width. N Calculate the moving average of N consecutive health scores within the window:

[0077] The fault warning threshold Ht is determined based on the minimum average health value within the sliding window. During normal equipment operation, the minimum health value H is calculated based on the minimum average similarity between the normal observation vector and the MSET model estimated vector. N The fault warning threshold is: H AN =kH N k is less than 1 and is the alarm threshold coefficient, which is determined by the operator based on their operating experience.

[0078] As a further improvement of the present invention, step S6 determines the current health status of the primary wind turbine. If the early warning model determines that the primary wind turbine is faulty, it will promptly notify the power plant staff of the early warning signal so that corresponding preventive measures can be taken. At the same time, the early warning information will also be uploaded to the power plant's monitoring system to realize the integration of the early warning system and the power plant's monitoring system. The data upload methods include, but are not limited to, HttpAPI interface, WebSocket interface, and IEC104 interface.

[0079] Beneficial effects: The present invention has the following advantages:

[0080] This invention provides a method and system for assessing the health of primary air turbines in large-scale coal-fired power plants based on improved multivariate state estimation. It addresses the limitations of requiring extensive fault data for training, classifying fault data into different fault types, relying on single-variable thresholds for fault assessment, and the online updating of fault diagnosis models. After cleaning and steady-state screening of historical data, a probability density equal-interval sampling method is used to sample the population. Fuzzy clustering is used to determine the state vector at each interval, and typical state vectors are selected to form a memory matrix, which is then updated using the k-nearest neighbor algorithm. Newer equipment data replaces lower-quality data in the historical training set, improving the model's online adaptability. A sliding window-based state similarity method is used instead of the traditional residual threshold method as the evaluation criterion for early warning. This invention employs an improved multivariate state estimation method, which is robust to noise and effectively improves the accuracy and reliability of primary air turbine condition monitoring. By comprehensively analyzing various operating parameters and establishing a multivariate state estimation model, the improved multivariate state estimation technology can estimate and predict the state of primary air turbines in real time, more accurately reflecting the actual health status of the primary air turbines, reducing the risk of misjudgment and omission, and providing operators with timely and accurate equipment status information. By setting health score thresholds, the system can issue early warnings before faults occur, ensuring the safe operation and economic benefits of the power plant. The system has good adaptability and scalability, and can be adjusted and optimized according to different types of primary air turbines and operating environments, making it suitable for various large-scale coal-fired power plant units. Attached Figure Description

[0081] Figure 1 This is a schematic diagram of the primary wind turbine fault early warning method in this invention;

[0082] Figure 2 This is a schematic diagram of the health assessment model construction in this invention. Detailed Implementation

[0083] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments:

[0084] The present invention provides a method for assessing the health of primary air turbines in thermal power plants based on data fusion, comprising the following steps:

[0085] The memory matrix should cover a typical normal operating state space and minimize its size. Probability density reflects the probability of a variable value within a certain region; selecting the state vectors of the process memory matrix based on probability density can effectively cover a common working space. MSET prediction accuracy essentially depends on historical data similar to the observed data; therefore, effectively constructing the process memory matrix based on probabilistic features is beneficial for improving prediction accuracy. The unit load can represent the equipment's operating state. Probability density sampling at equal intervals is performed based on the load to form the improved process memory matrix D. When selecting the state vectors, the aim is to remove redundant or similar state vectors. The population is sampled using probability density sampling at equal intervals, and fuzzy clustering is used to determine the state vectors at each interval. Typical state vectors are selected to form the memory matrix. A flowchart is shown below. Figure 1 As shown, the specific steps are as follows:

[0086] S1: After performing dimensionality reduction, noise reduction, and data cleaning on historical data, steady-state data is selected to provide a high-quality dataset for the model;

[0087] S2: Using the maximum and minimum values ​​of the unit load as upper and lower limits, divide the filtered data into M equal intervals, denoted as [L1, L2, ..., L...]. i , ...L M ];

[0088] S3: Sampling is performed on the population using the probability density interval sampling method. The probability density over each interval is integrated to obtain the loading value of the state vector at the loading interval L. i The probability P within i Multiply by the total number K of state vectors in the process memory matrix to obtain the number K of state vectors in each load interval. i K i =P i ×K;

[0089] S4: In each load interval L i Within this range, fuzzy clustering is used to determine the state vector X for each interval. i Set the load interval L i The optimal number of clusters is set to K. i ;

[0090] S5: Each load interval L i The selected state vector X will be internally determined. i Add to the process memory matrix D until the number of Xi reaches Ki;

[0091] S6: Real-time acquisition of primary wind turbine data; calculation of the estimated vector X based on the acquired memory matrix using the observed vector. estThe sliding window method is used to calculate the overall similarity of vectors. The overall similarity is the weighted average of the basic similarity and the state similarity. Considering the stability of the primary wind turbine equipment operation and the frequency of similarity changes, the current health of the primary wind turbine is calculated. The health is the weighted average of the overall similarity and the degree of similarity fluctuation.

[0092] S7: The memory matrix is ​​updated periodically using a K-nearest neighbor dynamic approach.

[0093] The steady-state data screening is based on the autocorrelation function, and the calculation formula is as follows:

[0094]

[0095] Among them, X t For time series observations, Let be the mean of the observations, and k be the lag order.

[0096] The dimensionality reduction process is based on the PCA algorithm. Traditional PCA is computationally intensive when dealing with large-scale datasets because it processes the entire dataset at once. Incremental PCA, on the other hand, is a method that can handle streaming or large-scale data. It processes only one or a small batch of data samples at a time, gradually updating the principal components. For example, when processing real-time wind turbine sensor data, the data is continuously generated, and incremental PCA can perform dimensionality reduction on this new data in real time.

[0097] Includes the following steps:

[0098] S101: Initialization

[0099] Standardize the first batch of data so that the mean of each feature is 0 and the variance is 1;

[0100] Calculate the covariance matrix C of the first batch of data. Let the first batch of data matrix be X1, the sample size be n1, and the number of features be d. Then the covariance matrix C is...

[0101] Eigenvalue decomposition of the covariance matrix C yields eigenvalues ​​λ1, λ2, ..., λ d and the corresponding feature vectors v1, v2, ..., v d Sort the eigenvalues ​​from largest to smallest, and the corresponding eigenvectors will also be sorted accordingly;

[0102] Select the eigenvectors corresponding to the k largest eigenvalues ​​as the initial principal components, where k is the preset dimension after dimensionality reduction;

[0103] S102: Incremental Update

[0104] Let the new sample be x. When the new sample arrives, first calculate the new sample mean. in It is the mean of the previous n-1 samples, updating the covariance moments. Where C old It is the previous covariance matrix estimate;

[0105] S103: Update principal components

[0106] The updated covariance matrix is ​​decomposed into eigenvalues. The new eigenvalues ​​are sorted from largest to smallest, and the corresponding eigenvectors are also sorted accordingly. The eigenvectors corresponding to the k largest eigenvalues ​​are selected as the updated principal components.

[0107] S104: Repeated Update

[0108] As new data continues to arrive, it is continuously processed. Steps S102 and S103 are repeated to continuously update the covariance matrix and principal components, thereby achieving incremental principal component analysis of the data.

[0109] The probability density in the multivariate state estimation step adopts the formula of the normal distribution:

[0110]

[0111] Where μ is the mean and σ is the standard deviation.

[0112] Health assessment models such as Figure 2 As shown,

[0113] The steps for dynamically updating the memory matrix based on K-nearest neighbors are as follows: For the actual collected primary wind turbine parameter data, after dimensionality reduction, noise reduction, and data cleaning, and after filtering steady-state data, the selected observation vectors are used as the distances to the samples to be classified. The Euclidean distances between the selected observation vectors and all vectors in the training set are calculated. The calculated Euclidean distances are sorted, and the k nearest training samples to the sample to be classified are found. The categories of these k neighbors are counted, and the category of the sample to be classified is determined by voting, i.e., belonging to [L1, L2, ..., L...]. i , ...L M The interval L i Furthermore, the population is sampled using probability density equal-interval sampling, and fuzzy clustering is used to determine the state vector at each interval. Typical state vectors are selected to update the memory matrix.

[0114] Furthermore, an estimated vector is calculated based on the memory matrix and the observation vector, and the overall similarity is calculated based on the estimated vector. The overall similarity is the weighted average of the basic similarity and the state similarity. Considering the stability of the primary wind turbine equipment operation and the frequency of similarity changes, the current health of the primary wind turbine is calculated and determined. The health is the weighted average of the overall similarity and the degree of similarity fluctuation. When the health exceeds the threshold, an early warning is issued.

[0115] The detailed algorithm steps of the fuzzy clustering are as follows:

[0116] S401: Initialization: Select the number of clusters C, and initialize the fuzzy membership matrix U, where U ij Represents data point x j Membership degree of cluster i. The membership degree value of each data point should be between [0,1], and the sum of all membership degrees for each data point j should be 1:

[0117] S402: Calculate cluster centers: Calculate the center V of each cluster based on the current membership matrix U. i :

[0118]

[0119] In the formula, m is the fuzzy factor, which usually takes a value between (1,2).

[0120] S403: Update the membership matrix: Calculate x for each data point. j Membership degree U for each cluster i ij

[0121]

[0122] S404: Determine convergence, calculate the change in the membership matrix, and stop iteration if the stopping condition is met (e.g., the change is less than a set threshold, or the maximum number of iterations is reached); otherwise, return to S402.

[0123] S405: Output the final cluster centers V and fuzzy membership matrix U.

[0124] The MSET algorithm first needs to construct a health status historical data memory matrix. This involves selecting several typical and effective observation vectors from historical data to form a historical matrix. This matrix should contain as few vectors as possible while still representing the normal operating status of the equipment under almost all operating conditions. The health status historical data memory matrix needs to collect as much historical system health status data as possible to cover all dynamic operating conditions of the system. The specific implementation process of MSET is as follows:

[0125] The memory matrix D is an n×m matrix, where n represents the number of n features of the primary wind turbine equipment status, and m represents the number of observation vectors. X i Let x represent the parameter of the i-th primary fan. ij This represents the value of the i-th parameter at time j.

[0126] X j =[x 1j ,x2j ,L,x nj ] T

[0127]

[0128] The observation data vector is represented by X. obs The expression is: X obs =[x1,x2,L,x n ] T

[0129] For each observation vector X obs The MEST model generates a weight vector W based on the memory matrix D, where W = [w1, w2, ..., w...]. m ] T

[0130] The weight vector W represents the similarity between the current observation vector and the historical data memory matrix. From this, the estimated vector X of the model output can be obtained. est .

[0131]

[0132] That is, the observation vector X obs The prediction output of the MSET algorithm is a linear combination of historical observation vectors. The weight vector W directly affects the model's estimation performance. W is determined by the minimum residual ε between the observed vector and the estimated vector. The smaller the absolute value of the residual, the higher the prediction accuracy of the MSET algorithm model.

[0133] ε=|X est -X obs |=|DW-X obs |

[0134] ε 2 =(DW-X) obs ) T (DW-X obs )

[0135] minε 2 =min[(DW-X obs ) T (DW-X obs )]

[0136] When ε satisfies the minimization constraint When, the weight vector W is

[0137]

[0138] For non-linear operators, the Euclidean distance is used for calculation as follows:

[0139]

[0140] Finally, estimate vector X est The expression is as follows:

[0141]

[0142] Most fault early warning systems determine fault status solely based on the residual between the normal estimate and the actual observed value of a single variable. However, this residual-based method has limitations. Sometimes, fault information is not contained in a single variable; trends from multiple variables or a series of behavioral states are needed to confirm the fault information. Furthermore, the residual value of a single parameter is easily affected by environmental noise, which can negatively impact the accuracy of fault early warnings. A judgment criterion describing the similarity between the observed state and the normal estimated state can more fully uncover anomalies in multiple variables. A similarity metric based on inter-state fault assessment can address the limitations of simultaneous monitoring thresholds for multiple variables in multivariate state estimation. Using a sliding window-based state similarity method instead of the traditional residual threshold method as the evaluation criterion for early warning can better assess multiple fault-related variables within a given state.

[0143] Similarity is a measure of how similar two samples are; it represents the differences between them. Euclidean distance is the most common measure of distance in vector spaces.

[0144]

[0145] When anomalies occur, the improved MSET model's estimates of other variables may also exceed the normal range. Therefore, using residuals as the sole criterion for fault warning will not only miss important anomaly information but may also lead to misjudgments. Therefore, by comprehensively considering the similarity between the observed vector and the estimated vector, and the similarity between the observed vector and the historical health memory matrix, the difference between the two can be measured to characterize the equipment's fault information. Furthermore, since the amount of fault information contained in each input variable of the primary turbine in the model is different, and the measurement accuracy and data reliability of each variable vary, their contributions to fault warning also differ. Based on this, weights are assigned to the contributions of each variable to fault warning.

[0146] In engineering applications, the design formula limits the similarity range measured by Euclidean distance to the range of [0,1].

[0147] Basic similarity refers to the similarity between the observed vectors X and X. obs With the estimated vector X est The degree of proximity is as follows:

[0148]

[0149] S(Xo ,X e )∈(0,1]

[0150] In the formula α k This represents the weight of the k-th variable in the fault warning function. α k >0, S(X) o ,X e )∈(0,1], the larger S is, the closer the two are.

[0151] State similarity refers to the similarity of the observation vector X obs The degree of similarity to the memory matrix is ​​as follows:

[0152]

[0153] S(X o ,D)∈(0,1], the larger S is, the closer the two are.

[0154] Overall similarity is the weighted average of basic similarity and state similarity:

[0155] S=β1S(X o ,X e )+β2S(X o D)

[0156] β1∈(0,1),β2∈(0,1)

[0157] S∈(0,1], the larger S is, the higher the similarity between the primary wind turbine and its historical normal operating state.

[0158] Considering the stability of primary air turbine operation and the frequency of similarity changes, the health design is as follows:

[0159]

[0160] α1∈(0,1),α2∈(0,1)

[0161] H∈(0,1]

[0162] and Represent the change in overall similarity at time t and the total overall similarity, respectively.

[0163] For real-time acquired primary wind turbine data, a sliding window method is used to eliminate uncertainties and random interference during equipment operation. The window width is set to N, and the overall similarity is calculated once per cycle to obtain the overall health data [H1, H2, L, H] within the specified window width. N Calculate the moving average of N consecutive health scores within the window:

[0164] The fault warning threshold Ht is determined based on the minimum average health value within the sliding window. During normal equipment operation, the minimum health value H is calculated based on the minimum average similarity between the normal observation vector and the MSET model estimated vector. N The fault warning threshold is: H AN =kH N k is less than 1 and is the alarm threshold coefficient, which is determined by the operator based on their operating experience.

[0165] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any other way. Any modifications or equivalent changes made based on the technical essence of the present invention shall still fall within the scope of protection claimed by the present invention.

Claims

1. A method for assessing the health of primary air turbines in thermal power plants based on data fusion, characterized in that: Includes the following steps: S1: After performing dimensionality reduction, noise reduction, and data cleaning on historical data, steady-state data is selected to provide a dataset for the model; S2: Using the maximum and minimum values ​​of the unit load as upper and lower limits, divide the filtered data into M equal intervals, denoted as [L1, L2, ..., L...]. i , ...L M ]; S3: Sampling is performed on the population using the probability density interval sampling method. The probability density over each interval is integrated to obtain the loading value of the state vector at the loading interval L. i The probability P within i Multiply by the total number K of state vectors in the process memory matrix to obtain the number K of state vectors in each load interval. i , ; S4: In each load interval L i Within this range, fuzzy clustering is used to determine the state vector X for each interval. i Set the load interval L i The optimal number of clusters is set to K. i ; S5: Each load interval L i The selected state vector X will be internally determined. i Add to the process memory matrix D until the number of Xi reaches Ki; S6: Real-time acquisition of primary wind turbine data; calculation of the estimated vector based on the acquired process memory matrix using the observed vector. The sliding window method is used to calculate the overall similarity of vectors. Considering the stability of the primary wind turbine equipment operation and the frequency of similarity changes, the current health of the primary wind turbine is calculated. The health is the weighted average of the overall similarity and the degree of similarity fluctuation. The overall similarity in step S6 is the weighted average of the basic similarity and the state similarity; Basic similarity refers to the similarity between observation vectors With the estimated vector The degree of proximity is as follows: ; ; In the formula This represents the weight of the i-th feature in the fault warning. n represents the number of features of the primary air turbine equipment status, and the larger S is, the closer the two are; State similarity is the observation vector The degree of similarity to the process memory matrix is ​​as follows: ; The larger the value of S, the closer the two are; Overall similarity is the weighted average of basic similarity and state similarity: ; ; , The larger the value of S, the higher the similarity between the primary air turbine and its historical normal operating conditions; Considering the stability of primary air turbine operation and the frequency of similarity changes, the health design is as follows: ; ; ; and Let represent the change in overall similarity at time t and the total overall similarity, respectively; S7: The process memory matrix is ​​updated periodically based on K-nearest neighbor dynamics.

2. The method for assessing the health of primary air turbines in thermal power plants based on data fusion as described in claim 1, characterized in that: In step S1, the steady-state data screening is based on the autocorrelation function, and the calculation formula is as follows: ; in, For time series observations, The mean of the observed values, This is the lag order.

3. The method for assessing the health of primary air turbines in thermal power plants based on data fusion as described in claim 1, characterized in that: The dimensionality reduction process in step S1 is implemented based on the incremental PCA algorithm, and includes the following steps: S101: Initialization Standardize the first batch of data so that the mean of each feature is 0 and the variance is 1; Calculate the covariance matrix of the first batch of data Let the first batch of data matrices be... The sample size is The number of features is Then the covariance matrix ; For covariance matrix Perform eigenvalue decomposition to obtain eigenvalues. and the corresponding feature vector The eigenvalues ​​are sorted from largest to smallest, and the corresponding eigenvectors are also sorted accordingly. Select the eigenvectors corresponding to the first r largest eigenvalues ​​as the initial principal components, where r is the preset dimension after dimensionality reduction; S102: Incremental Update Let the new sample be x. When the new sample arrives, first calculate the new sample mean. ,in It is the mean of the previous n-1 samples, updating the covariance moments. ,in It is the previous covariance matrix estimate; S103: Update principal components The updated covariance matrix is ​​decomposed into eigenvalues. The new eigenvalues ​​are sorted from largest to smallest, and the corresponding eigenvectors are also sorted accordingly. The eigenvectors corresponding to the first r largest eigenvalues ​​are selected as the updated principal components. S104: Repeated Update As new data continues to arrive, it is continuously processed. Steps S102 and S103 are repeated to continuously update the covariance matrix and principal components, thereby achieving incremental principal component analysis of the data.

4. The method for assessing the health of primary air turbines in thermal power plants based on data fusion as described in claim 1, characterized in that: In step S3, the probability density adopts the formula of the normal distribution: ; Where μ is the mean and σ is the standard deviation.

5. The method for assessing the health of primary air turbines in thermal power plants based on data fusion according to claim 1, characterized in that: The detailed algorithm steps for fuzzy clustering in step S4 are as follows: S401: Initialization: Select the number of clusters C', initialize the fuzzy membership matrix U, where Representing data points For the membership degree of cluster i, the membership degree value of each data point should be between [0, 1], and the sum of all membership degrees for each data point j should be 1. ; S402: Calculate cluster centers: Calculate the center of each cluster based on the current membership matrix U. : ; In the formula, m is a fuzzy factor, with a value between (1, 2); S403: Update the membership matrix: Calculate the membership matrix for each data point. Membership degree of each cluster i ; ; S404: Determine convergence and calculate the change in the membership matrix. If the following stopping condition is met, stop the iteration when the change is less than the set threshold or the maximum number of iterations is reached; otherwise, return to step S402. S405: Output the final cluster centers V and fuzzy membership matrix U.

6. The method for assessing the health of primary air turbines in thermal power plants based on data fusion according to claim 1, characterized in that: The estimated vector output by the model in step S6 The weight vector W is obtained from the weight vector W, which is the similarity between the current observation vector and the process memory matrix. ; Finally, the estimated vector The expression is as follows: 。 7. The method for assessing the health of primary air turbines in thermal power plants based on data fusion according to claim 1, characterized in that: The primary wind turbine data collected in real time in step S6 is processed using a sliding window method to eliminate uncertainties and random interference during equipment operation. The window width is set to N, and the overall similarity is calculated once per cycle to obtain the overall health data under the specified window width. Calculate the moving average of N consecutive health scores within the window: ; The fault warning threshold Ht is determined based on the average health value within the sliding window. When the equipment is operating normally, the minimum health value is calculated based on the minimum average similarity between the normal observation vector and the MSET model estimated vector. The fault warning threshold is: , A value less than 1 represents the alarm threshold coefficient, which is determined by the operator based on their experience.

8. The method for assessing the health of primary air turbines in thermal power plants based on data fusion according to claim 1, characterized in that: Step S6 determines the current health status of the primary wind turbine. If the early warning model determines that the primary wind turbine is faulty, it will promptly notify the power plant staff of the early warning signal so that corresponding preventive measures can be taken. At the same time, the early warning information will also be uploaded to the power plant's monitoring system to achieve integration between the early warning system and the power plant's monitoring system. Data upload methods include, but are not limited to, Http API interface, WebSocket interface, and IEC104 interface.

9. The method for assessing the health of primary air turbines in thermal power plants based on data fusion according to claim 1, characterized in that: Step S7 uses the K-nearest neighbor dynamic periodic update process memory matrix. The selected observation vectors are used as the distances to the sample to be classified. The Euclidean distance between the selected observation vectors and all vectors in the training set is calculated. The calculated Euclidean distances are sorted, and the k nearest training samples to the sample to be classified are found. The categories of these k neighbors are counted, and the category of the sample to be classified is determined by voting, i.e., belonging to [L1, L2, ..., L...]. i , ...L M The interval L i .

Citation Information

Patent Citations

  • Power plant blower fault early warning method based on multivariate state estimation

    CN112067335A

  • Power transformer fault early warning system based on data mining

    CN112884089A

  • Bearing fault early warning and diagnosis method based on improved MSET and spectrum characteristics

    CN113834657A