A data cleaning method applied in thermal power production
The improved exponential weighted moving average filter with trust factor calculation enhances thermal power data cleaning by dynamically adjusting weights and identifying anomalies, ensuring data integrity and accuracy for real-time monitoring and optimization.
Patent Information
- Application Number
- CN202510156775.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-02-13
AI Technical Summary
When facing a complex and changing operating environment, existing thermal production data cleaning methods have problems with insufficient integrity and accuracy of data cleaning, especially due to equipment measurement errors, sensor aging and data transmission interference, noise, missing values and outliers are difficult to effectively identify and process.
The improved exponential weighted moving average filtering algorithm is used, combined with the confidence factor and time series model, and the boiler thermal parameters timing is cleaned. By dynamically adjusting the weight coefficient and credibility evaluation, abnormal data points are identified and suppressed, and data standardization and missing completion are carried out.
It significantly improves the stability and accuracy of data, ensures the integrity and consistency of data quality, and supports real-time monitoring and optimization control of boiler operating status.
Smart Images

Figure CN119622220B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electrical digital data. More specifically, the present invention relates to a data cleaning method applied in thermal power production. Background Art
[0002] During the thermal power production process, a large amount of real-time data generated by boilers and other equipment is crucial for the operation monitoring and optimization of the system. However, due to equipment measurement errors, sensor aging, environmental changes, and interference during data transmission, the data in the thermal power system often contains noise, missing values, or outliers. Traditional data cleaning methods usually rely on simple filtering techniques or experience-based rules to remove noise and fill missing values. However, these methods often rely too much on manually set parameters and have limited capabilities for processing complex multi-dimensional data, making it difficult to cope with dynamically changing system states. In addition, although existing data analysis methods based on mechanism models can theoretically simulate the operation of the thermal power system, due to the setting of model parameters often being based on simplified assumptions, their applicability and accuracy have certain limitations. Especially when facing a complex and changeable operating environment, the accuracy of the model is insufficient to effectively identify anomalies and missing problems in the data, thus affecting the integrity and accuracy of data cleaning.
[0003] The patent application document with the publication number CN112328590A discloses a method for deeply cleaning the operation data of thermal equipment. Based on the equipment mechanism model, this patent application document combines the mechanism model with the machine learning adjoint simulation model of healthy operation data and the mass and heat balance equation of the thermal power system, and realizes the deep cleaning of data in the intelligent operation and maintenance of thermal equipment through methods such as key parameter similarity analysis and different parameter correlation analysis.
[0004] However, the above technical solution only realizes the deep cleaning of data based on the existing model, without fully considering the applicability and limitations of the existing model in specific application scenarios, and is prone to errors when processing complex data, resulting in the problem of low integrity and accuracy of data cleaning. Summary of the Invention
[0005] To solve the problem of low integrity and accuracy of data cleaning proposed in the above background art, the present invention provides the following solutions.
[0006] In this solution, the present invention provides a data cleaning method applied to thermal power production, including: obtaining the time series of thermal parameters of a boiler, where the time series of thermal parameters includes multiple-dimensional parameter time series at different times, and taking multiple-dimensional parameters at the same time as a data point; using an improved exponentially weighted moving average filtering algorithm to clean the data points to obtain the cleaned data points; where the improved exponentially weighted moving average filtering algorithm includes a weighting coefficient, and the weighting coefficient is the product of a preset value and an adjustment factor. The adjustment factor of the th dimension parameter in the th data point , where is the credibility of the th dimension parameter in the th data point, is the confidence factor of the th data point, and is the natural constant; the confidence factor represents the degree of abnormality of the data point; the credibility , where is the value of the th dimension parameter in the th data point, is the predicted value of the th dimension parameter in the th data point, is the average value of the absolute values of the differences between the values of each dimension parameter and the predicted value in the th data point, is the value of the th dimension parameter in the th neighborhood data point of the th data point, is the standard deviation of the th data point, and is the total number of neighborhood data points.
[0007] Through the improved exponentially weighted moving average filtering algorithm, the above technical solution introduces a weighting coefficient, combines the dynamic calculation of the preset value and the adjustment factor, and can effectively identify and suppress the influence of abnormal data points. The calculation of the adjustment factor fully considers the credibility and confidence factor of the data point, ensures that the influence of the abnormal point on the filtering result is reasonably reduced, and thus improves the stability and reliability of the data. The calculation of the credibility is based on the deviation between the parameter value and the predicted value, as well as the distribution characteristics of the neighborhood data, and can accurately capture the abnormal fluctuations of the data point, further enhancing the accuracy of data cleaning. At the same time, the confidence factor reflects the degree of abnormality of the data point, enabling the algorithm to more accurately identify potential deviations or noises and perform effective correction, thereby solving the problems of low data quality integrity and analysis accuracy.
[0008] Further, the trust factor is as follows:
[0009] ;
[0010] In the formula, is the trust factor of the th data point, is the credibility of the th dimension parameter in the th data point, is the credibility of the th dimension parameter in the th data point, is the average value of the credibility of each dimension parameter in the th data point, is the standard deviation of the credibility of each dimension parameter in the th data point, is the total number of items of dimension parameters, is an empirical constant, and ≠0.
[0011] Through the definition and calculation of the trust factor, the above technical solution further enhances the accurate recognition and correction ability of the abnormality degree of data points. By calculating the deviation of the credibility of each dimension parameter from its average value and combining with the standard deviation for normalization processing, the trust factor can reflect the deviation degree and its consistency of data points among various dimensions. In particular, by comparing with the credibility differences of other dimension parameters, the abnormal fluctuations of a certain dimension in the data set can be effectively identified, so as to provide accurate recognition and control of abnormal data. Generally speaking, this technical solution enables the data cleaning process to not only suppress noise but also retain the key information of the data, improves the accuracy and reliability of data processing, and provides more accurate and stable data support for subsequent analysis, modeling and optimization.
[0012] Further, the trust factor is as follows:
[0013] ;
[0014] In the formula, is the trust factor of the th data point, is the credibility of the th dimension parameter in the th data point, is the credibility of the th dimension parameter in the th data point, is the median of the credibility of each dimension parameter in the th data point, is the The maximum credibility of each dimension parameter among the data points, is the total number of items of the dimension parameter, is an empirical constant, and ≠0.
[0015] The above technical solution is different from the mean method. The use of the median can effectively avoid the interference of extreme values and enhance the stability of the algorithm when facing asymmetric distributions or the existence of abnormal data. At the same time, the introduction of the maximum value helps to identify the upper bound deviation of the data points, further improving the sensitivity to abnormal points. By standardizing the difference between the credibility of each dimension parameter and the median and combining the empirical constant to avoid numerical instability, the algorithm can more effectively identify the abnormal fluctuations in the data points, ensure the retention of normal information during the data cleaning process, and reduce the impact of noise.
[0016] Furthermore, a time series model is used to obtain the predicted values of each dimension parameter in each data point.
[0017] Furthermore, the time series model is an autoregressive moving average model or a long short-term memory network model.
[0018] In the above technical solution, the autoregressive moving average model utilizes the linear relationship of historical data and can provide relatively accurate predictions when the data fluctuations are small and the linear characteristics are obvious; while the long short-term memory network model can capture the long-term and short-term dependence relationships in the data through its deep learning mechanism, and is particularly suitable for processing complex, non-linear time series data with long-term memory, and can better cope with the complex dynamic changes in the boiler system. By combining the advantages of these two models, not only the modeling ability of the boiler thermal parameter time series is enhanced, but also the stability and accuracy of the prediction are effectively improved, providing more reliable technical support for the real-time monitoring, fault prediction and optimal control of the boiler operation state.
[0019] Furthermore, it also includes performing data standardization and missing value filling on the thermal parameter time series.
[0020] The above technical solution significantly improves the data quality and the robustness of the model by performing data standardization and missing value filling on the thermal parameter time series data. Data standardization ensures that parameters in different dimensions are processed on the same scale, eliminates the interference that may be caused by different dimensions, and makes subsequent analysis more accurate and consistent. The missing value filling process solves the problem of vacancies caused by incomplete data collection, and ensures the continuity and integrity of the time series data through reasonable filling methods, avoiding the negative impact of missing values on the model training and prediction results.
[0021] Furthermore, the thermal parameter time series includes: boiler temperature time series, humidity time series, pressure time series, and water flow rate time series.
[0022] Furthermore, a temperature sensor is used to collect the temperature time series of the boiler, a humidity sensor is used to collect the humidity time series of the boiler, a pressure sensor is used to collect the pressure time series of the boiler, and a flow sensor is used to collect the water flow rate time series of the boiler.
[0023] The beneficial effects of the present invention are as follows:
[0024] By introducing an improved exponentially weighted moving average filtering algorithm and a multi-dimensional credibility evaluation mechanism, the present invention effectively improves the accuracy and robustness of data cleaning in the boiler system during thermal production. By cleaning and standardizing multiple thermal parameter time series such as the temperature, humidity, pressure, and water flow rate of the boiler, noise and abnormal data can be effectively removed, ensuring the quality and continuity of the data. The introduction of the credibility factor enables the algorithm to accurately capture abnormal fluctuations in data points and make intelligent adjustments according to the deviation degree of each dimension parameter, thereby improving the accuracy and stability of data cleaning. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present invention will become readily understood. In the drawings, several embodiments of the present invention are shown by way of illustration and not limitation, and the same or corresponding reference numerals denote the same or corresponding parts, wherein:
[0026] Figure 1 is a flowchart of a data cleaning method applied to thermal production according to an embodiment of the present invention, schematically shown. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] Hereinafter, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0028] The following will describe the specific embodiments of the present invention in detail with reference to the accompanying drawings.
[0029] An embodiment of a data cleaning method applied to thermal production.
[0030] As Figure 1 shown, a flowchart of a data cleaning method applied to thermal production according to an embodiment of the present invention includes the following steps:
[0031] S1: Obtain the time series of thermal parameters in multiple dimensions of the boiler at different times, and take the parameters in multiple dimensions at the same time as a data point.
[0032] In one embodiment, the time series of thermal parameters includes: the time series of boiler temperature, humidity, pressure, and water flow rate. These parameters comprehensively cover the key thermal characteristics during the operation of the boiler. By using a high-precision temperature sensor to collect the temperature data of the boiler in real time, the thermal state of the boiler can be accurately reflected; by obtaining the humidity time series data through a humidity sensor, it helps to monitor the steam and humidity changes of the boiler and provides support for thermal performance analysis; by collecting the pressure time series data through a pressure sensor, the pressure dynamics inside the boiler can be quickly captured, thereby realizing the evaluation of operation safety; at the same time, by using a flow sensor to obtain the water flow rate time series data, it can reflect the water supply status of the boiler in real time, helps to optimize the water supply efficiency and prevent abnormal working conditions.
[0033] To ensure the reliability and integrity of the data, further data standardization and missing data filling processing are performed on the time series of thermal parameters. Data standardization can map parameters in different dimensions to the same dimension, making the parameters more comparable and integrable in subsequent analysis. Especially in machine learning model training or thermal system anomaly detection, standardized data can significantly improve the performance and stability of the algorithm. For possible missing data, repair processing is carried out through missing data filling techniques, such as methods based on interpolation algorithms, historical data fitting, or model prediction filling, which can effectively restore the complete time series of thermal parameters and avoid analysis biases caused by data missing.
[0034] This series of data processing measures not only improve the integrity and consistency of the time series of thermal parameters but also enhance the accuracy and reliability of subsequent analysis. By efficiently collecting and processing core parameters such as boiler temperature, humidity, pressure, and water flow rate, the operation state of the boiler can be more deeply understood, potential risks can be identified, and a scientific basis can be provided for operation optimization, thereby realizing the intelligentization and high efficiency of equipment management. This comprehensive and refined technical solution helps to significantly improve the operation efficiency, reliability, and safety of the boiler system, and provides strong technical support for the intelligent monitoring and decision-making of industrial boiler operation.
[0035] S2: Use the improved exponentially weighted moving average filtering algorithm to clean each data point to obtain the cleaned data points.
[0036] In one embodiment, since the time series of boiler thermal parameters usually has high dimensionality and dynamics, traditional static filtering methods often struggle to handle the complex relationships of multi-dimensional data. However, the improved exponentially weighted moving average filtering algorithm can effectively adapt to such complex data characteristics, significantly enhancing the efficiency and effectiveness of data cleaning. Ultimately, by generating high-quality data points after cleaning, the accuracy of monitoring the operating state of the thermal system can be further improved, significantly reducing decision-making biases caused by data anomalies, and providing strong guarantees for the efficiency, safety, and intelligence of boiler operation.
[0037] Compared with traditional filtering algorithms, the improved algorithm dynamically adjusts the weight coefficients, fully considering the credibility and anomaly degree of data points, and can more effectively eliminate noise, smooth data, while retaining important trends and characteristic changes. This process not only improves the overall consistency of the data but also lays a solid foundation for subsequent analysis and modeling.
[0038] S3: Among them, the improved exponentially weighted moving average filtering algorithm includes a weighting coefficient, and the weighting coefficient is the product of a preset value and an adjustment factor.
[0039] In one embodiment, the preset value is:
[0040] ;
[0041] In the formula, is the preset value of the th data point, is the standard deviation of the th data point. This calculation method is based on the standard deviation as the core measure of data fluctuation degree, and can dynamically reflect the stability of data points. For data points with larger fluctuations, their standard deviation values are higher, and the corresponding preset values will be lower, thus reducing their influence weights in weighted calculations; while for data points with smaller fluctuations and higher stability, their preset values are higher and can play a more important role in the filtering process. In addition, the preset value can also be set manually to more precisely meet personalized needs in specific scenarios. This diverse way of generating preset values greatly enhances the adaptability and practicality of the algorithm.
[0042] For the adjustment factor, the adjustment factor of the th dimension parameter in the th data point, in the formula, is the credibility of the th dimension parameter in the th data point, is the credibility factor of the th data point, is the natural constant.
[0043] By weighting different dimensional parameters through adjustment factors, the influence weight of abnormal parameters can be adaptively reduced, thereby further improving the accuracy of data cleaning. After introducing the above improvement strategy, the technical effect is particularly significant. First, the dynamic adjustment mechanism of the preset value enables the filtering algorithm to adaptively adjust the weight distribution according to the volatility of data points, enhancing the ability to suppress noise data. Second, by comprehensively evaluating the credibility and trust factor, the adjustment factor incorporates the complex characteristics of multi-dimensional data points into the weight calculation process, achieving precise detection and dynamic correction of abnormal data. In addition, while ensuring the filtering effect, this improved weight mechanism fully preserves the true characteristics of data points, avoiding the loss of key information due to excessive smoothing.
[0044] S4: The adjustment factor is calculated based on the parameter credibility and the data point trust factor.
[0045] In one embodiment, the credibility , where is the value of the th dimension parameter in the th data point, is the predicted value of the th dimension parameter in the th data point, is the average value of the absolute value of the difference between the values of each dimension parameter and the predicted value in the th data point, is the value of the th dimension parameter in the th neighborhood data point of the th data point, is the standard deviation of the th data point, is the total number of neighborhood data points.
[0046] The calculation of credibility comprehensively considers the deviation between the parameter value and the predicted value, the distribution characteristics of neighborhood data, and the overall volatility of data points, making the credibility evaluation of each dimension parameter more targeted and robust. By normalizing the ratio of the deviation to the mean difference and combining the sum of squared relative deviations of neighborhood data distribution, credibility can not only reflect the degree of difference between the parameter value and the predicted value, but also reveal the consistency and rationality of data points in their surrounding environment. This method effectively retains the key features and overall trends of data while smoothing outliers and noise, thereby improving the reliability of data processing and the accuracy of subsequent analysis, providing a scientific basis for anomaly identification and performance optimization of multi-dimensional dynamic systems.
[0047] The trust factor is:
[0048] ;
[0049] In the formula, is the confidence factor of the th data point, is the credibility of the th dimension parameter among the th data points, is the credibility of the th dimension parameter among the th data points, is the average value of the credibility of each dimension parameter among the th data points, is the standard deviation of the credibility of each dimension parameter among the th data points, is the total number of items of dimension parameters, is an empirical constant, and ≠0.
[0050] The calculation of the confidence factor comprehensively considers the deviation degree of the credibility of each dimension parameter from its average value, balances the data differences between different dimensions through standardization processing, and further reveals the balance and consistency of the internal credibility distribution of data points through deviation comparison with other dimension parameters. An empirical constant is introduced to avoid the division-by-zero problem and enhance the robustness of the algorithm to adapt to different data scenarios. This scheme can accurately capture potential local anomalies in data points, ensure that the weights of anomaly points are reasonably reduced, thus better protecting the characteristics of normal data during the filtering process, while eliminating the impact of noise and anomalies on the overall data quality. This quantitative design of the confidence factor provides strong technical support for the accurate processing and reliability analysis of complex multi-dimensional data, and helps to achieve more efficient anomaly detection and intelligent optimization in dynamic systems.
[0051] In another embodiment, the confidence factor is:
[0052] ;
[0053] In the formula, is the confidence factor of the th data point, is the credibility of the th dimension parameter among the th data points, is the credibility of the th dimension parameter among the th data points, is the median of the credibility of each dimension parameter among the th data points, is the maximum value of the credibility of each dimension parameter among the th data points, is the total number of dimension parameters, is an empirical constant, and ≠0.
[0054] Compared with the calculation method based on the mean and standard deviation, the introduction of the median can more effectively reduce the interference of extreme values on the calculation of the credibility factor and improve the robustness to asymmetrically distributed data. The use of the maximum value provides an upper bound reference for the parameter credibility and enhances the sensitivity of anomaly detection. At the same time, through the normalization processing of the difference and the median in the numerator part and the consistency comparison of global parameters, the relative deviation degree between dimension parameters can be more accurately reflected. This design significantly improves the anomaly detection ability of the algorithm in complex data environments, ensures that the filtering process can more accurately identify and suppress the influence of anomaly points, while retaining the overall characteristics and key trends of multi-dimensional data, and provides more efficient and reliable technical support for data cleaning and analysis in dynamic systems.
[0055] The solution of the present invention can effectively clean the time-series data of boiler thermal parameters and remove abnormal data points by introducing an improved exponentially weighted moving average filtering algorithm and combining the dynamic calculation of credibility factors and credibility, thereby improving the accuracy and stability of the data. By weighted adjustment of data points, the algorithm can automatically identify and correct potential noise and deviation according to the credibility and anomaly degree of each data point, ensuring high-quality data. Using the time-series model for data prediction further enhances the accurate capture of future trends. Especially when dealing with multi-dimensional data such as temperature, humidity, pressure, and water flow rate during boiler operation, it can effectively cope with data fluctuations and missing problems, and improve the reliability of data analysis and decision support. In addition, data standardization and missing value filling effectively solve the problems of data inconsistency and missing, ensure the stability of model training, and provide a more accurate and comprehensive data basis for subsequent boiler condition monitoring, fault warning, and optimization control, with significant technical advantages and application values.
[0056] In the description of this specification, the meanings of "a plurality of" and "several" are at least two, such as two, three or more, etc., unless otherwise clearly and specifically defined.
[0057] Although this specification has shown and described multiple embodiments of the present invention, it is obvious to those skilled in the art that such embodiments are provided only by way of example. Those skilled in the art will think of many changes, alterations, and alternative ways without departing from the spirit and concept of the present invention. It should be understood that various alternative solutions to the embodiments of the present invention described herein can be adopted in the practice of the present invention.
Claims
1. A data cleaning method applied in thermal power production, characterized in that, Including: Obtain the time series of the thermal parameters of the boiler, where the time series of the thermal parameters includes multiple-dimensional parameter time series at different times, and use multiple-dimensional parameters at the same time as a data point; Perform data standardization and missing value filling on the time series of the thermal parameters; Use an improved exponentially weighted moving average filtering algorithm to clean each data point to obtain the cleaned data points; Among them, the improved exponentially weighted moving average filtering algorithm includes a weighting coefficient, and the weighting coefficient is the product of a preset value and an adjustment factor. The adjustment factor of the th dimension parameter in the th data point. In the formula, , is the credibility of the th dimension parameter in the th data point, is the credibility factor of the th data point, and is the natural constant. The credibility factor characterizes the degree of abnormality of the data point; The credibility , where is the value of the -th dimension parameter in the -th data point, is the predicted value of the -th dimension parameter in the -th data point, is the average value of the absolute value of the difference between the values of each dimension parameter and the predicted value in the -th data point, is the value of the -th dimension parameter in the -th neighborhood data point of the -th data point, is the standard deviation of the -th data point, is the total number of neighborhood data points.
2. The data cleaning method applied in thermal production according to claim 1, characterized in that The credibility factor is: ; In the formula, is the confidence factor of the th data point, is the confidence level of the th data point for the th dimensional parameter, is the confidence level of the th data point for the th dimensional parameter, is the average value of the confidence levels of the dimensional parameters in the th data point, is the standard deviation of the confidence levels of the dimensional parameters in the th data point, is the total number of items of the dimensional parameters, is an empirical constant, and ≠0.
3. A data cleaning method applied in thermal production according to claim 1, characterized in that, The credibility factor is: ; In the formula, is the credibility factor of the th data point, is the credibility of the th data point for the th dimension parameter, is the credibility of the th data point for the th dimension parameter, is the median of the credibility of each dimension parameter in the th data point, is the maximum value of the credibility of each dimension parameter in the th data point, is the total number of items of the dimension parameter, is an empirical constant, and ≠0.
4. A data cleaning method applied in thermal power production according to claim 1, characterized in that Use a time series model to obtain the predicted values of each dimensional parameter in each data point.
5. A data cleaning method applied in thermal power production according to claim 4, characterized in that The time series model is an autoregressive moving average model or a long short-term memory network model.
6. A data cleaning method applied in thermal production according to claim 1, characterized in that The time series of the thermal parameters includes: the time series of the boiler temperature, the time series of the humidity, the time series of the pressure, and the time series of the water flow.
7. A data cleaning method applied in thermal power production according to claim 1, characterized in that, Use a temperature sensor to collect the time series of the boiler temperature, use a humidity sensor to collect the time series of the boiler humidity, use a pressure sensor to collect the time series of the boiler pressure, and use a flow sensor to collect the time series of the boiler water flow.
Citation Information
Patent Citations
Thermal equipment operation data deep cleaning method
CN112328590A
High-dimensional time series data cleaning method based on correlation assistance
CN115022911A
Data cleaning method for flight time sequence parameters of spacecraft
CN117807069A