Desulfurization prediction data preprocessing method based on box plot-knowledge experience double drive

Through the data preprocessing method of box chart-knowledge and experience dual drive, the problem of low automation level of desulfurization system in coal-fired power plants is solved, the data is high accuracy and stability is achieved, and the applicability and intelligence of the desulfurization prediction model are improved.

CN120337014AActive Publication Date: 2025-07-18ANHUI MAANSHAN WANNENGDA POWER GENERATION CO LTD

Patent Information

Application Number
CN202510822721.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-07-18
Estimated Expiration
2045-06-19

AI Technical Summary

Technical Problem

The automation level of desulfurization systems in coal-fired power plants is low, making it difficult to cope with complex working conditions, and insufficient sensor data quality, which affects the stability and generalization capabilities of the prediction model.

Method used

The data preprocessing method based on box graph-knowledge and experience dual drive is adopted, including data acquisition and cleaning, multi-source data structure, box graph inspection, knowledge experience inspection and data noise reduction and normalization, identify and eliminate outliers, and combine expert experience to perform data correction and noise suppression.

Benefits of technology

It improves the reliability and applicability of the desulfurization prediction model, ensures the accuracy and consistency of input data, enhances the stability and robustness of data, and supports the intelligent operation and green transformation of coal-fired power plants.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337014A_ABST
    Figure CN120337014A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, in particular to a desulfurization prediction data preprocessing method based on box plot-knowledge experience double drive, comprising the following steps: S1, collecting and cleaning key parameters such as flue gas flow, temperature, pH value, current and the like, and verifying data type and numerical reasonability; s2, constructing derived features, and fusing laboratory and inspection data to form multi-source data; s3, setting a multi-class anomaly counter, and detecting an outlier based on a box plot; s4, marking semantic tags by utilizing expert rules, and identifying the consistency of working conditions; s5, counting missing, mutation and multiplexing times and updating an abnormal state; and S6, Fourier noise reduction and Z-score standardization processing are carried out. According to the invention, high-precision, strong-robustness and highly-intelligent desulfurization data preprocessing is realized, and the data quality, the system stability and the engineering applicability are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a desulfurization prediction data preprocessing method driven by box plot - knowledge and experience. Background Art

[0002] With the wide application of desulfurization prediction models in coal - fired power plants, data preprocessing, as a fundamental link in model construction and operation, has become increasingly important. China's energy structure and power production methods are undergoing a rapid transformation. Coal - fired power units need to balance tasks such as power supply guarantee and peak regulation, and gradually adapt to complex operating conditions such as deep peak regulation and low - load operation. Under this background, the operation management of desulfurization systems faces higher requirements for intelligence and automation.

[0003] Currently, the desulfurization control of coal - fired power plants mainly relies on manual intervention and empirical judgment, with a low level of automation, making it difficult to effectively cope with the complexity and dynamic changes of operating conditions. Taking the limestone - gypsum wet desulfurization system as an example, its key parameters such as the pH value of gypsum slurry and the SO2 concentration in flue gas have significant characteristics of large time delay, non - linearity, and strong coupling, resulting in difficulties in accurately modeling and dynamically regulating traditional control strategies. At the same time, with the increase in the density of sensor deployment, a large amount of multi - source heterogeneous data generated by the system also has obvious deficiencies in terms of data quality, structure, and temporal consistency, severely restricting the stability and generalization ability of prediction models. Summary of the Invention

[0004] The present invention provides a desulfurization prediction data preprocessing method driven by box plot - knowledge and experience. By constructing a preprocessing process with the capabilities of anomaly detection, data repair, noise suppression, and semantic recognition, it ensures that the input data of the model has a good foundation in terms of accuracy, stability, and expression consistency, thereby improving the reliability and applicability of the desulfurization prediction model and providing key data support for the intelligent operation and green transformation of coal - fired power units.

[0005] The desulfurization prediction data preprocessing method driven by box plot - knowledge and experience includes the following steps: S1, data collection and cleaning: Real - time collection of various key parameters during the operation of the desulfurization system. The key parameters include flue gas flow, flue gas temperature, pH value of gypsum slurry, and circulating pump current. At the same time, perform cleaning operations on the collected key parameters based on data integrity and logical consistency, including verification of the legality of data types (such as whether it is floating - point data) and verification of numerical rationality (such as whether non - negative eigenvalue appears less than zero); S2, multi - source data construction: Process and reconstruct the structure of the key parameters after the cleaning operation, construct derived features through calculation fusion, logical combination, or statistical derivation, and at the same time, integrate the auxiliary information of laboratory test data and manual inspection data to supplement external characteristic parameters closely related to the desulfurization process; S3, Box plot test: Set up missing value counters, reused value counters, mutation value counters, and abnormal status indicators for the constructed multi-source data respectively to comprehensively record and identify various abnormal statuses. At the same time, based on the statistical principle of box plots, construct the quartile intervals of features for each working condition category, and use the upper and lower limit ranges to identify potential outliers, so as to realize the detection of outliers and abnormal identification of multi-source data; S4, Knowledge and experience test: Integrate expert experience and prior knowledge to conduct multi-dimensional abnormal identification and working condition consistency verification on multi-source data, structurally express the relationships between key operating parameters, equipment status, and typical working conditions, and use a rule engine or feature combination logic to perform semantic label annotation on real-time multi-source data streams, so as to realize the identification and early warning of abnormal working conditions, boundary behaviors, or potential failure modes; S5, Abnormal statistic checker verification: Dynamically monitor and identify the status of statistical indicators of missing values, mutation values, and reused values in multi-source data; S6, Data denoising and normalization: Perform Fourier transform denoising on multi-source data. At the same time, adopt the Z-score standardization strategy to unify the scale and regularize the distribution of the denoised multi-source data.

[0006] Optionally, the data collection and cleaning in S1 include: S11, Sensor deployment and data collection: Deploy multiple types of industrial-grade sensors at the desulfurization tower process section, including flow sensors, temperature sensors, pH meters, pressure sensors, and current transformers. The sensors are connected to the data collection terminal through a signal acquisition card to achieve continuous collection of real-time data; S12, Data type verification mechanism: Use a regular expression matching strategy to automatically verify the collected key parameters. If the collected key parameters conform to the floating-point format, they are converted and stored in the database. If not, they are marked as abnormal, triggering an alarm and recording; S13, Data physical constraint verification and anomaly correction: Combine the engineering attributes of the desulfurization process and the physical laws of variables to construct a feature dimension rule library, and judge the range rationality of key parameter values. For key parameters that should be non-negative, if an abnormal value less than zero appears, it is corrected to the minimum value (such as 0).

[0007] Optionally, the construction of multi-source data in S2 includes: S21, Multi-measurement point data fusion processing: Use a dynamic fusion algorithm based on variance discrimination to process the same key parameter. When the variance of the measured value of the key parameter is less than the preset threshold, take its average value as the fusion result. When the variance exceeds the preset threshold, select the two groups of data with the smallest difference for weighted average; S22, Derived feature construction based on physical mechanisms: Combining the desulfurization process flow, calculate various derived variables based on the collected key parameters. Among them, the total sulfur content in the flue gas is calculated by multiplying the total air volume by the SO2 concentration in the original flue gas, and the sulfur content ratio is calculated by dividing the total sulfur content by the total coal quantity. S23, Heterogeneous data integration and structured conversion: Introduce external data from laboratory tests and manual inspections, and perform unified format conversion and structured coding processing on text, enumeration, and numerical information.

[0008] Optionally, the box plot test in S3 includes: S31, Feature value stability detection: For the case where multi-source data remains unchanged within a continuous time window, determine whether there is a collection failure or malfunction. S32, Missing value identification and filling: For the case where multi-source data has missing values, consider it as a missing event to trigger the compensation mechanism. S33, Time series mutation detection: Use the first-order difference jump detection method to judge the change trend of continuous time series data in multi-source data. If the continuous difference value exceeds the mutation threshold, it is regarded as a time series anomaly, increment the mutation value counter by 1, replace it with the previous moment value, and increment the reuse value counter by 1. S34, Box plot outlier detection: Use the box plot method to identify abnormal data points. For multi-source data, calculate its first quartile Q1, third quartile Q3, and interquartile range IQR, and construct the normal value interval according to the box plot principle, that is, [Q1 - k×IQR, Q3 + k×IQR], where k is the outlier sensitivity coefficient. If a data point in the multi-source data exceeds this interval range, it is determined as an outlier.

[0009] Optionally, the feature value stability detection in S31 includes: S311, Set the monitoring time window: Set the time window threshold And extract the continuous observation value sequence of multi-source data within this window. S312, Judge whether the feature value is constant: Judge whether the multi-source data remains unchanged within the time window; if the values at all times are equal, it is regarded as a constant state. S313, Mark the abnormal state: When the multi-source data remains constant within the window, set the first bit of its abnormal state register from 0 to 1, and mark it as having a sensor jam or data freeze anomaly.

[0010] Optionally, the missing value identification and filling in S32 includes: S321, Record the missing event: When it is detected that the observed value of multi-source data at the current moment is empty, increment the corresponding missing value counter by 1 and record the frequency of the occurrence of the missing value. S322. Perform missing value filling: Automatically fill the current missing value with the valid observed value at the previous moment. S323. Update the reuse counter: After completing the filling operation, increment the reuse value counter by 1 to record the number of times the compensation behavior is executed.

[0011] Optionally, the knowledge and experience verification in S4 includes: S41. Abnormality identification based on empirical rules: According to the operation mechanism and engineering practice of the desulfurization system, set the physical range for key parameters. If the collected key parameter exceeds this range, it is regarded as an abnormal value, and the key parameter is replaced with the valid value at the previous moment. At the same time, increment the reuse value counter by 1 to record the correction behavior. S42. Working condition semantic label annotation: Use the predefined working condition identification rule set, combine with the rule engine to perform real-time matching on multi-source data, and automatically assign corresponding working condition semantic labels according to the combination logic of multi-source data to realize the identification of the operating state.

[0012] Optionally, the abnormality statistic checker in S5 includes: S51. Missing value statistics and abnormality detection: Set the sliding time window length threshold and the upper limit threshold of the missing times. Real-time statistics the number of missing times of each multi-source data within the window. If the cumulative number exceeds the upper limit threshold of the missing times, it is regarded as a missing abnormality. Set the second bit of the abnormality status indicator from 0 to 1 and trigger an alarm. S52. Mutation value statistics and abnormality detection: Based on the result of the first-order difference, count the number of mutation events. If the number of mutations exceeds the sliding time window length threshold within the set window, it is marked as a fluctuation abnormality, update the third bit of the abnormality status indicator to 1, and start the alarm process. S53. Reuse value statistics and abnormality detection: Monitor the frequency of multi-source data being reused by the data at the previous moment. If the number of reuse times within the window exceeds the time window and number threshold of the reuse behavior, it is determined as a reuse abnormality, update the corresponding bit of the abnormality status indicator and issue an alarm, reflecting the existence of equipment lag or data abnormality handling problems.

[0013] Optionally, the data noise reduction and normalization in S6 includes: S61. Fourier transform noise reduction: For the high-frequency interference in multi-source data, use the discrete Fourier transform (DFT) to convert the time series to the frequency domain. According to the amplitude spectrum distribution, set the cut-off frequency threshold , filter the high-frequency noise components, and restore them to the noise-reduced time-domain signal through the inverse transform. S62. Z-score normalization: Perform Z-score normalization on the noise-reduced multi-source data.

[0014] Advantages of the present invention: In the present invention, outlier values are identified and removed through the box plot method, and the specific working condition data is compensated and corrected in combination with expert experience rules, so as to improve the authenticity and consistency of the data. At the same time, the Fourier transform noise reduction technology is introduced to effectively suppress measurement noise and short-term fluctuations, enhance the smoothness and reliability of the data, and provide high-quality input for subsequent processing.

[0015] In the present invention, by considering the coupling characteristics between variables and the characteristics of working condition switching, combining historical data and real-time monitoring results, the data quality is dynamically evaluated and abnormal repair and sequence smoothing processing are implemented, so as to improve the stability and adaptability of the preprocessing process under variable operating conditions, and ensure the accuracy of control strategies and adjustment responses.

[0016] In the present invention, by constructing a standardized and automated data preprocessing process, manual intervention is reduced, and the intelligent level and execution efficiency of desulfurization data processing are improved; the method has good versatility and scalability, is applicable to various operating loads and working conditions, and provides effective support for the refined management and low-carbon operation of the desulfurization system in coal-fired power plants. Description of the Drawings

[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1 It is a schematic flow chart of the preprocessing method according to an embodiment of the present invention. Detailed Embodiments

[0019] The present invention will be described in detail below in combination with the drawings and specific embodiments. At the same time, it should be noted here that in order to make the embodiments more detailed, the following embodiments are the best and preferred embodiments. For some well-known technologies, those skilled in the art can also adopt other alternative methods for implementation; moreover, the drawing part is only for more specifically describing the embodiments, and is not intended to specifically limit the present invention.

[0020] It should be pointed out that in the specification, it is mentioned that "an embodiment", "embodiment", "exemplary embodiment", "some embodiments", etc. indicate that the described embodiments may include specific features, structures or characteristics, but not necessarily every embodiment includes the specific features, structures or characteristics. In addition, when combining embodiments to describe specific features, structures or characteristics, implementing such features, structures or characteristics in combination with other embodiments (whether explicitly described or not) should be within the knowledge of those skilled in the relevant art.

[0021] Generally, terms can be understood, at least in part, from their use in context. For example, depending at least in part on the context, the term "one or more" as used herein can be used to describe any feature, structure, or property in the singular sense, or can be used to describe a combination of features, structures, or properties in the plural sense. Additionally, the term "based on" can be understood as not necessarily intended to convey an exclusive set of factors, but rather can alternatively, at least in part depending on the context, allow for the existence of other factors that are not necessarily explicitly described.

[0022] As Figure 1 shown, the desulfurization prediction data preprocessing method based on box plot - knowledge experience dual drive includes the following steps: Step 1, data collection and cleaning: Key parameters during the operation of the desulfurization system are collected in real - time, and preliminary cleaning is performed based on data integrity and logical consistency rules to ensure the availability and accuracy of the original data.

[0023] Specifically, it includes: 1. Sensor deployment and data collection: Multiple types of industrial - grade sensors are deployed at key process links of the desulfurization tower, including but not limited to flow sensors, temperature sensors, pH meters, pressure sensors, and current transformers, etc. These sensors are connected to the data collection terminal through a signal acquisition card to achieve continuous collection of real - time operating condition data. The collected data is transmitted to the data collection module via an industrial communication protocol and synchronously written into the database for storage as the input source for subsequent data processing.

[0024] 2. Data type verification mechanism: For the problem of type consistency of the collected data, a regular expression matching strategy is used to automatically verify the data format of each field. The specific process includes: constructing a regular expression template for matching the floating - point format, comparing each input data item one by one; if the match is successful, type conversion is performed and stored in the database; if the match fails, it is regarded as abnormal data, and the system will trigger an alarm mechanism and record the data entry for manual verification to prevent illegal formats from interfering with the calculation model.

[0025] 3. Data physical constraint verification and anomaly correction: Combining the engineering attributes of the desulfurization process and the physical laws of variables, a feature dimension rule library is constructed to determine the rationality of the range of key parameter values. For example, variables such as slurry pH value and flue gas flow theoretically should not be negative. The system automatically detects parameter values according to the preset physical rules. If the detected value is less than the theoretical lower limit, it is automatically corrected to the set minimum value (such as 0) to enhance the physical consistency of the data and the input stability of the model; data that conforms to the physical constraints is retained to ensure that valid information is not lost during the data cleaning process.

[0026] Step 2, Multi-source data construction: Derived features are constructed from the original features through calculation fusion, logical combination, etc. to meet the input requirements of the model. At the same time, external data such as laboratory tests and manual inspections are introduced to enhance the integrity of feature dimensions and the semantic relevance of desulfurization.

[0027] Specifically include: 1. Multi-measurement point data fusion processing: For multiple sensor data of the same physical quantity set at different measurement points, a dynamic fusion algorithm is used to improve the robustness of feature values and reduce the impact of single-point measurement errors on the overall model. For example, when processing data from multiple temperature sensors, first calculate the real-time variance σ² of the three sensors to evaluate the consistency between measurement values. When σ² is less than or equal to the preset threshold, the data is considered relatively consistent, and the arithmetic mean of the three is directly used as the fusion output. When σ² is greater than this threshold, it indicates that there are large deviations. At this time, further calculate the absolute differences of the three pairwise combinations, select the pair with the smallest difference as the reliable data source, and use the mean of this pair of data as the final fusion result.

[0028] 2. Construction of derivative features based on physical mechanisms: Combining the desulfurization process flow and its physical logical relationships, use the collected basic features to calculate derivative variables to enhance the engineering rationality of features. For example: The total coal quantity can be obtained by accumulating the rotational speed feedback values of each coal feeder; the total sulfur content in the flue gas can be calculated by multiplying the total air volume by the SO2 concentration in the original flue gas; the sulfur content ratio is expressed as the ratio of the total sulfur amount to the total coal amount. Such features not only enrich the model input but also improve the model's interpretability of key desulfurization indicators.

[0029] 3. Heterogeneous data integration and structured transformation: To supplement important operating condition information that cannot be covered by the real-time monitoring system, this module further introduces external data sources such as laboratory test data and manual inspection data. Perform unified format conversion and structured processing on the obtained heterogeneous information (including text type, enumeration type, and numerical type). For example: Text fields are converted into standard labels by using rule extraction and mapping methods; enumeration data is converted into numerical vectors that can participate in modeling through one-hot encoding. The above processing ensures that various types of information can be effectively integrated into the main data framework, providing a more comprehensive description of the operating conditions for the model.

[0030] Step 3, Box plot test: Set up abnormal counters such as missing values, reused values, and mutation values for each feature, use the box plot method to establish the interquartile range, and identify outliers according to the operating condition type to achieve multi-dimensional detection and marking of data anomalies.

[0031] Specifically include: 1. For the case where the feature value remains unchanged within a continuous time window, judge whether there is a collection failure or malfunction. The specific method is: Set the time window threshold For each feature value within the window monitor the changes within; if a certain eigenvalue remains constant within the system automatically sets the first digit of the anomaly marker corresponding to this feature from 0 to 1, indicating that there may be an anomaly in this feature (such as sensor jamming or data freezing).

[0032] 2. For the case where eigenvalue is missing, the system monitors it in real time through this module, and when detecting that the data is empty, immediately regards it as a missing event to trigger the compensation mechanism. The compensation process includes the following steps: First, increment the missing value counter corresponding to the feature by 1 to record the frequency of missing occurrences; Second, to ensure the integrity and continuity of the data sequence, the system automatically fills and replaces it with the valid data of this feature at the previous time point; Finally, after completing the filling operation, synchronously increment the reused value counter by 1 to count the number of compensation actions and assist in subsequent data quality assessment and model adjustment.

[0033] 3. Use the first-order difference jump detection method to judge the change trend of continuous time-series data. Specifically, by calculating the difference sequence between the eigenvalue of adjacent time steps, monitor whether there is an instantaneous drastic change in the data. When the difference value of a certain feature always exceeds the mutation threshold within a continuous period, the system determines that there is a potential time-series anomaly in this feature. For the detected mutation situation, the system increments the corresponding mutation value counter by 1, replaces and fills it with the valid data at the previous moment, and at the same time increments the reused value counter by 1 to record this compensation behavior and ensure the smoothness and continuity of the data sequence.

[0034] 4. Introduce a statistical test method based on box plots to identify anomalies in feature data. For any feature data set, first calculate its first quartile (Q1), third quartile (Q3), and interquartile range (IQR = Q3 - Q1), and construct the normal value interval of this feature according to the principle of box plots, that is, [Q1 - k×IQR, Q3 + k×IQR], where k is the anomaly sensitivity coefficient, and the common value is 1.5 or 3. If a certain data point exceeds this interval range, it is determined as an outlier. For the detected outliers, the system replaces them with the valid values at the previous moment to maintain data continuity, and at the same time increments the reused value counter by 1 for subsequent anomaly frequency statistics and model reliability analysis.

[0035] Step 4, Knowledge and Experience Verification: Construct a knowledge graph of desulfurization working conditions based on expert knowledge, structurally express the semantic relationship between variables and working conditions, and use a rule engine to label the real-time data stream to achieve the identification of complex working conditions, the auxiliary determination of abnormal behaviors, and early warning support.

[0036] Specifically include: 1. Based on the operation mechanism and engineering practice of the desulfurization system, the reasonable physical ranges of key parameters are set by pooling expert knowledge. For each target variable, an upper limit value and a lower limit value are set. If the collected data exceeds this range, it is determined to be physically unreasonable and marked as an outlier. For such data, the system replaces it with the valid data at the previous moment to ensure the continuity and credibility of the data. At the same time, the reuse value counter is incremented by 1 to record this correction operation.

[0037] 2. Using a predefined set of working condition identification rules and combining with a rule engine, logical matching and feature combination judgment are performed on the input data to achieve automatic annotation of working condition semantic labels. This process is achieved by establishing a rule base composed of multiple Boolean logic expressions, and each rule is composed of a combination logic of a set of process variables and expert-set thresholds. The system performs real-time matching on each rule and maps the input n-dimensional feature vector to the corresponding semantic label set through a label mapping function, so as to accurately describe the current operating state and provide a data basis and decision support for subsequent control strategy adjustment and anomaly warning.

[0038] Step 5, Anomaly statistic checker: Dynamically monitor the missing, mutation, reuse, etc. counters within a set time window. When the number of times exceeds the limit, automatically update the anomaly status flag and issue a warning to achieve closed-loop management and multi-source collaborative identification of abnormal behaviors.

[0039] Specifically, it includes: 1. Missing value statistics and anomaly detection: Preset the sliding time window length and the upper limit threshold of the number of missing times. During the real-time data inflow process, the system continuously counts the number of missing times of each feature within this window. If a certain feature value is missing at a certain moment, the system records the status as 1 through an indicator function, otherwise it is 0. At the end of the window period, the system determines whether this count value exceeds the preset threshold. If it exceeds, it is regarded as a missing anomaly. At this time, the system sets the second bit of the anomaly status register of the corresponding feature from 0 to 1, indicating that there may be problems such as sensor damage, data acquisition interruption, or communication anomaly, and outputs a warning signal.

[0040] 2. Mutation value statistics and anomaly detection: The system continuously tracks and monitors the mutation behavior of feature data. By setting the time window length and the mutation number threshold, the number of jump points identified by the first-order difference is counted. When it is detected that a feature value mutates at a certain moment and is replaced with the data at the previous moment, the system marks it as 1, otherwise it is 0. If the cumulative number of mutations within the window period exceeds the preset threshold, it is determined that there is a fluctuation anomaly for this feature, automatically update the third bit of the anomaly status register to 1, and start the corresponding alarm process.

[0041] 3. Reuse Value Statistics and Anomaly Detection: Set the time window and the threshold of the number of times for the feature reuse behavior, which is specifically used to judge the frequency of a certain feature value being replaced by the previous moment's data due to anomalies. Within the sliding window, the system counts the cumulative number of times by marking whether the current value is generated by reusing the previous moment's data. When the number of times of this reuse behavior exceeds the set threshold, the system determines it as a reuse anomaly, which may reflect equipment response lag, abnormal data processing logic, or external environmental interference. At this time, the third bit flag of the anomaly status register is also updated to 1, and a warning prompt is issued.

[0042] Step 6, Data Denoising and Normalization: Use Fourier transform to perform frequency-domain denoising on the original data to suppress random perturbations; at the same time, apply the Z-score normalization method to unify the feature scale and distribution, improving the stability of model training and the standardization of input data.

[0043] Specifically include: 1. Fourier Transform Denoising: For the high-frequency perturbations that may exist in the original time series, use the discrete Fourier transform to convert it to the frequency-domain space. Let the original time series signal be: , convert the signal from the time domain to the frequency domain through the discrete Fourier transform: , where, represents the th frequency component, is the total number of data points, is the imaginary unit. In the frequency domain, according to the amplitude spectrum distribution characteristics of the signal, set the cut-off frequency threshold , and filter out the high-frequency noise components (i.e., ), that is: ; Then restore it to the denoised time-domain signal through the inverse transform: , where, is the frequency index, representing the position of a certain frequency component, represents the absolute value of the frequency index, used to uniformly measure the distance from the center frequency to eliminate the negative frequency in the symmetric frequency, represents the frequency component after removing the high frequency, represents the denoised time-domain signal. This method can effectively weaken the high-frequency noise caused by equipment jitter, sampling error, or environmental interference on the basis of retaining the key structural features of the signal, which is beneficial to improving the stability of the data and the robustness of subsequent model analysis.

[0044] Z-score Normalization: To eliminate the scale differences between different feature variables and improve the training efficiency and convergence stability of the model, the system performs Z-score normalization processing on each feature variable. Let the sample sequence of a certain feature variable be , where n is the total number of samples, and its mean and standard deviation are calculated as follows: ; . Then the Z-score normalization formula is expressed as: , where represents the normalization result of sample . After normalization, the mean of the feature data is 0 and the standard deviation is 1, which helps to improve the robustness and convergence speed of the model when processing data with different scales.

[0045] The present invention covers any alternatives, modifications, equivalent methods, and solutions made to the essence and scope of the present invention. To enable the public to have a thorough understanding of the present invention, specific details are described in detail in the following preferred embodiments of the present invention, and those skilled in the art can fully understand the present invention without these detailed descriptions. In addition, well-known methods, processes, procedures, components, and circuits are not described in detail to avoid unnecessary confusion to the essence of the present invention.

[0046] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A data preprocessing method for desulfurization prediction based on box plot - knowledge and experience dual drive, characterized in that It includes the following steps: S1, Data collection and cleaning: Real-time collection of various key parameters during the operation of the desulfurization system. The key parameters include flue gas flow, flue gas temperature, pH value of gypsum slurry, and current of the circulation pump. At the same time, perform cleaning operations on the collected key parameters based on data integrity and logical consistency, including verification of the legality of data types and verification of numerical rationality; S2, Multi-source data construction: Process and reconstruct the structure of the key parameters after the cleaning operation, construct derived features through calculation fusion, logical combination, or statistical derivation. At the same time, integrate the auxiliary information of laboratory test data and manual inspection data, and supplement external characteristic parameters closely related to the desulfurization process; S3, Box plot inspection: Set missing value counters, reused value counters, mutation value counters, and abnormal status indicators for the constructed multi-source data respectively to comprehensively record and identify various abnormal statuses. At the same time, based on the statistical principle of the box plot, construct the quartile intervals of the features for each working condition category respectively, and use the upper and lower limit ranges to identify potential outliers to achieve outlier detection and anomaly recognition of multi-source data; S4, Knowledge and experience inspection: Integrate expert experience and prior knowledge to perform multi-dimensional anomaly recognition and working condition consistency verification on multi-source data, structurally express the relationships between key operating parameters, equipment status, and typical working conditions, and use a rule engine or feature combination logic to perform semantic label annotation on the real-time multi-source data stream to achieve the recognition and early warning of abnormal working conditions, boundary behaviors, or potential failure modes; S5, Abnormal statistic checker verification: Dynamically monitor and identify the status of statistical indicators of missing values, mutation values, and reused values in multi-source data; S6, Data noise reduction and normalization: Perform Fourier transform noise reduction processing on multi-source data. At the same time, adopt the Z-score standardization strategy to unify the scale and regularize the distribution of the multi-source data after noise reduction processing.

2. The preprocessing method for desulfurization prediction data based on the double drive of box plot and knowledge experience according to claim 1, wherein The data collection and cleaning in S1 include: S11, Sensor deployment and data collection: Deploy multiple types of industrial-grade sensors at the technological links of the desulfurization tower, including flow sensors, temperature sensors, pH meters, pressure sensors, and current transformers. The sensors are connected to the data collection terminal through a signal acquisition card to achieve continuous collection of real-time data; S12, Data type verification mechanism: Automatically verify the collected key parameters using a regular expression matching strategy. If the collected key parameters conform to the floating-point format, they are converted and stored in the database. If not, they are marked as abnormal, triggering an alarm and recording; S13, Data physical constraint verification and anomaly correction: Combine the engineering attributes of the desulfurization process and the physical laws of variables to construct a feature dimension rule library, and judge the range rationality of the key parameter values. For key parameters that should be non-negative, if an abnormal value less than zero appears, it is corrected to the minimum value.

3. The preprocessing method for desulfurization prediction data based on box plot-knowledge experience dual drive according to claim 2, wherein The multi-source data construction in S2 includes: S21, Multi - point measurement data fusion processing: Use a dynamic fusion algorithm based on variance discrimination to process the same key parameter. When the variance of the measured value of the key parameter is less than the preset threshold, take its average value as the fusion result. When the variance exceeds the preset threshold, select two groups of data with the smallest difference for weighted averaging; S22, Derived feature construction based on physical mechanism: Combine the desulfurization process flow and calculate various derived variables based on the collected key parameters. Among them, the total sulfur content in the flue gas is calculated by multiplying the total air volume by the SO2 concentration of the raw flue gas, and the sulfur content ratio is calculated by dividing the total sulfur content by the total coal consumption; S23, Heterogeneous data integration and structured conversion: Introduce external data from laboratory tests and manual inspections, and perform unified format conversion and structured coding processing on text - type, enumeration - type, and numerical - type information.

4. The preprocessing method for desulfurization prediction data based on the double drive of box plot and knowledge experience according to claim 3, wherein The box - plot test in S3 includes: S31, Feature value stability detection: For the case where multi - source data remains unchanged within a continuous time window, judge whether there is a collection failure or malfunction; S32, Missing value identification and filling: For the case where there are missing values in multi - source data, consider it as a missing event to trigger the compensation mechanism; S33, Time - series mutation detection: Use the first - order difference jump detection method to judge the change trend of continuous time - series data in multi - source data. If the continuous difference value exceeds the mutation threshold, it is regarded as a time - series anomaly, increment the mutation value counter by 1, and replace it with the value of the previous moment. At the same time, increment the reuse value counter by 1; S34, Box - plot outlier detection: Use the box - plot method to identify outlier data points. For multi - source data, calculate its first quartile Q1, third quartile Q3, and inter - quartile range IQR, and construct a normal value interval according to the box - plot principle, that is, [Q1 - k×IQR, Q3 + k×IQR], where k is the anomaly sensitivity coefficient. If a data point in the multi - source data exceeds this interval range, it is determined as an outlier.

5. The preprocessing method for desulfurization prediction data based on box plot - knowledge experience dual drive according to claim 4, wherein The feature value stability detection in S31 includes: S311, Set the monitoring time window: Set the time window threshold , and extract the continuous observation value sequence of the multi-source data within this window; S312, Judge whether the feature value is constant: Judge whether the multi - source data remains unchanged within the time window; if the values at all times are equal, it is regarded as a constant state; S313, Mark the abnormal state: When the multi - source data remains constant within the window, set the first bit of its abnormal state register from 0 to 1 and mark it as an abnormal state.

6. The desulfurization prediction data preprocessing method based on box plot-knowledge experience dual drive according to claim 5, characterized in that The abnormal state includes sensor jamming or data freezing anomalies.

7. The preprocessing method for desulfurization prediction data based on box plot - knowledge experience dual drive according to claim 6, wherein The missing value identification and filling in S32 includes: S321, Record the missing event: When it is detected that the observed value of the multi - source data at the current moment is empty, increment the corresponding missing value counter by 1 and record the frequency of the missing occurrence; S322, Perform missing value filling: Automatically fill the current missing value with the valid observed value at the previous moment; S323, Update the reuse counter: After completing the filling operation, increment the reuse value counter by 1 to record the number of times the compensation behavior is executed.

8. The preprocessing method for desulfurization prediction data based on box plot-knowledge experience dual drive according to claim 7, characterized in that, The knowledge and experience test in S4 includes: S41. Anomaly recognition based on empirical rules: According to the operation mechanism of the desulfurization system and engineering practice, set the physical range for key parameters. If the collected key parameters exceed this range, they are regarded as outliers, and the valid values of the key parameters at the previous moment are used for replacement. At the same time, increment the reuse value counter by 1 to record the correction behavior. S42. Working condition semantic label annotation: Use the predefined working condition recognition rule set, combine the rule engine to perform real-time matching on multi-source data, and automatically assign corresponding working condition semantic labels according to the combination logic of multi-source data to realize the recognition of the operating state.

9. The preprocessing method for desulfurization prediction data based on box plot-knowledge experience dual drive according to claim 8, wherein The anomaly statistic checker in S5 includes: S51. Missing value statistics and anomaly detection: Set the sliding time window length threshold and the upper limit threshold of the number of missing times. Real-time statistics of the number of missing times of each multi-source data within the window. If the cumulative number exceeds the upper limit threshold of the number of missing times, it is regarded as a missing anomaly, set the second bit of the anomaly status indicator from 0 to 1, and trigger an alarm. S52. Mutation value statistics and anomaly detection: Based on the result of the first-order difference, count the number of mutation events. If the number of mutations exceeds the sliding time window length threshold within the set window, it is marked as a fluctuation anomaly, update the third bit of the anomaly status indicator to 1, and start the alarm process. S53. Reuse value statistics and anomaly detection: Monitor the frequency of multi-source data reused by the data at the previous moment. If the number of reuse times exceeds the time window and number threshold of the reuse behavior within the window, it is determined as a reuse anomaly, update the corresponding bit of the anomaly status indicator and issue an alarm, reflecting the existence of equipment lag or data anomaly handling problems.

10. The preprocessing method for desulfurization prediction data based on box plot - knowledge experience dual drive according to claim 9, characterized in that The data noise reduction and normalization in S6 include: S61, Fourier transform noise reduction: For high-frequency interference in multi-source data, the discrete Fourier transform is used to convert the time series to the frequency domain. According to the amplitude spectrum distribution, a cut-off frequency threshold is set , the high-frequency noise components are filtered, and the denoised time-domain signal is restored through inverse transformation; S62. Z-score standardization: Perform Z-score standardization processing on the noise-reduced multi-source data.

Citation Information

Patent Citations

  • Method suitable for power data quality assessment and rule check

    CN106649840A

  • Long short-term memory network industrial soft measurement method for desulfurization process flue gas

    CN115688865A

  • Bed temperature prediction and control optimization method for 660MW ultra-supercritical circulating fluidized bed boiler

    CN116910472A

  • Data acquisition and verification method of extruder energy consumption online monitoring system

    CN118093564A

  • Data acquisition fusion analysis system

    CN118227826A

Cited By

  • Equipment state monitoring method and system based on automatic working condition calibration and multi-dimensional signal analysis

    CN120668219A