Data Analysis-Based Fault Early Warning Method for Activated Carbon Regeneration Equipment
By constructing a state space point cloud and a multi-level point connection network for activated carbon regeneration equipment, and combining topology quantification of deviation and aggregation of dynamic margin, the problem of missing equipment status monitoring information in existing technologies is solved, enabling earlier fault warning and risk identification.
Patent Information
- Application Number
- CN202511403771.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-29
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2045-09-29
AI Technical Summary
In existing technologies for industrial equipment condition monitoring, static feature modeling of multidimensional sensor data results in a lack of state space characterization information, making it difficult to reflect changes in equipment operating status in the early stages and making it difficult to identify and warn of potential risks in a timely manner.
By synchronously collecting real-time values such as furnace temperature, steam flow rate, and flue gas oxygen content of activated carbon regeneration equipment within a sliding time window, a state space point cloud of activated carbon regeneration is constructed, a multi-level point connection network is established, a life cycle distribution map of loop structure is identified, the topology quantization deviation and aggregation dynamic margin are calculated, and a two-dimensional state feature vector is generated for fault early warning.
It improves the completeness and accuracy of equipment status characterization, enhances the ability to detect early anomalies, identifies potential risks earlier, and reduces production losses caused by sudden shutdowns.
Smart Images

Figure CN120892894B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of pattern recognition technology, and in particular to a fault early warning method for activated carbon regeneration equipment based on data analysis. Background Technology
[0002] In industrial equipment condition monitoring applications, it is necessary to collect multi-dimensional sensor data during equipment operation, analyze the characteristics of the operating status, and establish corresponding discrimination models to identify abnormal equipment conditions.
[0003] Current technologies for industrial equipment condition monitoring rely on multi-dimensional sensor data and discriminant models for equipment condition identification. However, multi-dimensional sensor data is typically modeled using only static features, resulting in a lack of information in the characterization of the state space. This makes it difficult to promptly reflect potential risks in the early stages of changes in equipment operating status, and the reliance on simple threshold comparisons across a single dimension can easily lead to delayed early warnings. Therefore, improvements are needed. Summary of the Invention
[0004] The purpose of this invention is to address the shortcomings of existing technologies by proposing a fault early warning method for activated carbon regeneration equipment based on data analysis.
[0005] To achieve the above objectives, the present invention adopts the following technical solution: a fault early warning method for activated carbon regeneration equipment based on data analysis, comprising the following steps:
[0006] Within a preset sliding time window, real-time values of furnace temperature, steam flow rate, flue gas oxygen content, and outlet temperature of the activated carbon regeneration equipment are collected simultaneously to establish a spatial point cloud of activated carbon regeneration status.
[0007] Based on the activated carbon regeneration state spatial point cloud, a multi-level point connection network is constructed by successively increasing the scale of the connection distance between points. Loop structures formed at different connection distance scales and the distance scales of their appearance and disappearance are identified to obtain a loop structure life cycle distribution map. Then, the difference between the current loop structure life cycle distribution map and the baseline distribution map corresponding to the normal operation of the equipment is calculated to generate the topology quantification deviation.
[0008] Within a preset sliding time window, the control power of the furnace temperature is selected as the input and the actual value of the furnace temperature is selected as the output. A difference equation describing the relationship between the two is established, the coefficients of the difference equation are solved, and the aggregate dynamic margin is calculated.
[0009] The topology quantization deviation is used as the first dimension, and the aggregated dynamic margin is used as the second dimension. Together, they form a two-dimensional state feature vector. It is determined whether the two-dimensional state feature vector falls within a preset two-dimensional warning area, and a Boolean value is output according to the determination result to generate a device operating status signal.
[0010] Preferably, the step of obtaining the activated carbon regeneration state spatial point cloud is as follows:
[0011] Within a preset sliding time window, real-time values of furnace temperature, steam flow rate, flue gas oxygen content, and outlet temperature are recorded synchronously with a unified timestamp. Any missing or duplicate records are removed to form a synchronous sequence of real-time values of furnace temperature, steam flow rate, flue gas oxygen content, and outlet temperature.
[0012] Based on the synchronization sequence of the real-time values of furnace temperature, steam flow rate, flue gas oxygen content, and outlet temperature, the corresponding four values are extracted one by one in chronological order and spliced into coordinate quadruplets in a fixed order. Identical coordinate quadruplets are removed and unique identifiers are retained to generate a high-dimensional coordinate point sequence.
[0013] Based on the high-dimensional coordinate point sequence, the time sequence label is removed and the storage buckets are divided according to the adjacent distance relationship. All coordinate quadruples are collected in no particular order and mapped to the same coordinate system to form a spatial point cloud of activated carbon regeneration state.
[0014] Preferably, the steps for obtaining the loop structure lifecycle distribution map are as follows:
[0015] Based on the activated carbon regeneration state spatial point cloud, a progressively increasing point connection distance scale is set. For each distance scale, coordinate point pairs whose Euclidean distance is not greater than the distance scale are calculated one by one and connected. All the connections and nodes are organized into a structural graph to form a multi-level point connection network.
[0016] Based on the multi-level point connection network, loop structures are identified scale by scale, and the birth and death scales of each loop structure are recorded. The life cycle length of each loop structure is calculated and mapped to the interval sequence. The distance scale range is divided into multiple equal-width intervals. For each interval, the number of loops in all life cycle coverage intervals is counted, and the normalized probability value is calculated. At the same time, the average life cycle length of the loops in the coverage intervals is calculated to obtain the average life cycle length, thus forming a loop structure life cycle distribution map.
[0017] Preferably, the step of obtaining the topology quantization deviation is as follows:
[0018] Based on the loop structure lifecycle distribution diagram, the topology quantization deviation is calculated.
[0019] Preferably, the step of obtaining the aggregation dynamic margin is as follows:
[0020] Based on the preset sliding time window, the control power and actual value of furnace temperature are extracted at each moment, a one-to-one corresponding input and output pairing sequence is established, abnormal pairing points are removed and the sampling interval is kept consistent, and an input-output synchronization sequence of furnace temperature control power and actual value of furnace temperature is generated.
[0021] Based on the input-output synchronization sequence of the control power and the actual value of the furnace temperature, the order and lag order of the difference equation are set, the input and output quantities are expanded into a matrix expression for each sample, the coefficients of the difference equation are calculated, and all eigenvalues are obtained to obtain the set of eigenvalues.
[0022] Calculate the aggregate dynamic margin based on the set of eigenvalues.
[0023] Preferably, the steps for obtaining the two-dimensional state feature vector are as follows:
[0024] Based on the topological structure quantization deviation and the aggregation dynamic margin, a binary sequence is assembled in a fixed order of the first and second dimensions, the decimal places are unified and the original units are maintained, the value is checked for emptiness and the null value handling method and boundary inclusion rules are recorded to form a two-dimensional state feature vector.
[0025] Preferably, the step of acquiring the device operating status signal is as follows:
[0026] Based on the two-dimensional state feature vector, the minimum and maximum coordinates and coordinate axis directions of the preset two-dimensional warning area are read, and the size relationship between the first dimension and the second dimension and the corresponding boundary is compared item by item. The boundary inclusion rule is applied to obtain the landing area judgment mark of the two-dimensional state feature vector.
[0027] Preferably, the step of acquiring the device operating status signal further includes: generating a Boolean value based on the landing area determination mark of the two-dimensional status feature vector, according to the rule that falling into the preset two-dimensional warning area is recorded as true and not falling into the preset two-dimensional warning area is recorded as false, attaching a timestamp and source identifier and writing it into the output channel to generate the device operating status signal.
[0028] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0029] This invention synchronously collects multi-dimensional real-time values such as furnace temperature, steam flow rate, flue gas oxygen content, and outlet temperature within a preset sliding time window. It treats the multi-source data at each sampling moment as high-dimensional coordinate points and constructs a state space point cloud, enabling the equipment's operating state to be described in a multi-dimensional space. This improves the completeness and accuracy of state characterization. By successively increasing the scale of the connection distance between points on the point cloud data, a multi-level point connection network is established, and the formation and disappearance characteristics of loop structures at different scales are extracted. The deviation is calculated by combining the difference between the loop structure lifecycle distribution map and the normal operation baseline distribution, which can sensitively reflect the evolution of the topology. The subtle changes enhance the ability to detect early anomalies. By selecting the furnace temperature control power and the actual furnace temperature to establish a difference equation and solve for the characteristic roots, the aggregate dynamic margin is calculated based on the overall distribution of all characteristic roots. This allows for a comprehensive quantification of the stability of the equipment's dynamic characteristics, avoiding the limitations of relying on a single characteristic index. The topology quantification deviation and the aggregate dynamic margin are used to construct a two-dimensional state feature vector. By judging its relative relationship with the warning area in the two-dimensional space, state classification is achieved, improving the accuracy of fault risk identification. This enables earlier identification of potential risks and reduces production losses caused by sudden shutdowns. Attached Figure Description
[0030] Figure 1 This is a schematic diagram of the steps of the present invention. Detailed Implementation
[0031] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0032] Please see Figure 1 This invention provides a technical solution: a fault early warning method for activated carbon regeneration equipment based on data analysis, comprising the following steps:
[0033] Within a preset sliding time window, real-time values of furnace temperature, steam flow rate, flue gas oxygen content, and outlet temperature of the activated carbon regeneration equipment are collected simultaneously to establish a spatial point cloud of activated carbon regeneration status.
[0034] Based on the point cloud of the activated carbon regeneration state space, a multi-level point connection network is constructed by successively increasing the scale of the connection distance between points. The loop structure formed at different connection distance scales and the distance scales of appearance and disappearance are identified to obtain the loop structure life cycle distribution map. Then, the difference between the current loop structure life cycle distribution map and the baseline distribution map corresponding to the normal operation of the equipment is calculated to generate the topology quantitative deviation.
[0035] Within a preset sliding time window, the control power of the furnace temperature is selected as the input and the actual value of the furnace temperature is selected as the output. A difference equation describing the relationship between the two is established, the coefficients of the difference equation are solved, and the aggregate dynamic margin is calculated.
[0036] The topology deviation is used as the first dimension, and the dynamic margin is used as the second dimension. Together they form a two-dimensional state feature vector. It is determined whether the two-dimensional state feature vector falls within the preset two-dimensional warning area, and a Boolean value is output according to the determination result to generate the device operation status signal.
[0037] The steps for obtaining the spatial point cloud of activated carbon regeneration state are as follows:
[0038] Within a preset sliding time window, real-time values of furnace temperature, steam flow rate, flue gas oxygen content, and outlet temperature are recorded synchronously with a unified timestamp. Any missing or duplicate records are removed to form a synchronous sequence of real-time values of furnace temperature, steam flow rate, flue gas oxygen content, and outlet temperature.
[0039] Based on the synchronous sequence of real-time furnace temperature, real-time steam flow, real-time flue gas oxygen content, and real-time outlet temperature, the corresponding four values are extracted one by one in chronological order and spliced into coordinate quadruplets in a fixed order. Identical coordinate quadruplets are removed and unique identifiers are retained to generate a high-dimensional coordinate point sequence.
[0040] Based on the high-dimensional coordinate point sequence, the time sequence label is removed and the storage buckets are divided according to the adjacent distance relationship. All coordinate quadruples are collected in no particular order and mapped to the same coordinate system to form a spatial point cloud of activated carbon regeneration state.
[0041] Specifically, within a preset 10-minute sliding time window, the window slides forward in 1-minute increments. Data is acquired via the programmable logic controller (PLC) data interface connected to the equipment, sampling at a frequency of 1 Hz. Real-time values of four key process parameters—furnace temperature, steam flow rate, flue gas oxygen content, and outlet temperature—are recorded simultaneously. The acquisition time of all data points is calibrated using the Network Time Protocol (NTP) to ensure timestamp consistency. First, the acquired raw data undergoes preliminary verification, and the effective operating range for each parameter is set. These ranges are determined based on the equipment design manual and statistical analysis of historical data from at least three months of normal operation. For example, taking furnace temperature as an example, the historical data shows a mean of 900℃ and a standard deviation of 15℃. Therefore, the effective range is set to the mean plus or minus three times the standard deviation, i.e., [855℃, 945℃]. Similarly, the effective range for steam flow rate is set to [1.2 t / h, 2.8 t / h], and the effective range for flue gas oxygen content is set to [2.0%, ]. [5.0%], the effective range of the outlet temperature is set to [150℃, 250℃]. Iterate through each record within the sliding time window and check whether all four values in the record fall within their respective effective ranges. If any value exceeds the preset range, the entire record corresponding to that timestamp is marked as invalid and removed directly. Then, handle the issues of missing and duplicate data. Check whether each timestamp completely contains all four parameter values. If a record at a certain timestamp is missing any parameter, the record is considered incomplete and removed. Then, identify duplicate records by comparing the timestamps of adjacent records. If two or more records are found to have the same timestamp, only the first record is kept and the rest of the duplicate records are deleted. After the above data cleaning process, a synchronous sequence of real-time furnace temperature, real-time steam flow, real-time flue gas oxygen content, and real-time outlet temperature values is finally obtained with complete data, no duplication, and all values within the normal fluctuation range.
[0042] Based on the synchronized sequence of real-time furnace temperature, steam flow rate, flue gas oxygen content, and outlet temperature generated in the previous step, the data records in the sequence are processed one by one according to their timestamps. For each record, the values of the four parameters are extracted strictly in the fixed order of "furnace temperature, steam flow rate, flue gas oxygen content, and outlet temperature," and these four values are concatenated into a four-dimensional vector, i.e., a coordinate quadruple. For example, if the four values at a certain moment are 900.5℃, 2.15 t / h, 3.5%, and 180.2℃, then the generated coordinate quadruple is (900.5, 2.15, 3.5, ...). 180.2) After generating all coordinate quadruples, a deduplication operation is performed to identify and retain all unique states that the equipment has operated in. This deduplication is not based on timestamps, but on the numerical content of the coordinate quadruples themselves. Considering that sensor measurement data are floating-point numbers, direct equivalent comparisons may lead to misjudgments due to accuracy issues. Therefore, a tolerance threshold based on the measurement accuracy of each parameter is introduced for comparison. This tolerance threshold is set with reference to the sensor's technical specifications. For example, the accuracy of the furnace temperature sensor is ±0.5℃, the steam flow meter is ±0.01 t / h, the oxygen analyzer is ±0.1%, and the outlet temperature sensor is ±0.2℃. Therefore, when comparing two coordinate quadruples A=(T1, F1, O1, Tout1) and B=(T2, F2, O2, Tout1),... When Tout2), two coordinate quadruples are considered identical only if all four conditions are met simultaneously: |T1-T2|<0.5, |F1-F2|<0.01, |O1-O2|<0.1, and |Tout1-Tout2|<0.2. The deduplication process is implemented using a hash set data structure. It iterates through all generated coordinate quadruples. For each quadruple, it checks whether it exists in the hash set within a tolerance range. If it does not exist, it is added to the set and assigned a unique integer identifier starting from 1, while retaining the time order of its first appearance. If it already exists, the current quadruple is discarded. After the traversal is completed, the set stores all unique coordinate quadruples and their unique identifiers, thus generating a high-dimensional coordinate point sequence.
[0043] Based on the high-dimensional coordinate point sequence generated in the previous process, all time-related information is first removed, including the original timestamps and the time sequence labels retained during deduplication. This completely removes the temporal relationships between data points, treating each coordinate quadruple as an independent geometric point for subsequent spatial structure analysis. Before mapping these points to a unified four-dimensional coordinate system, to eliminate the influence of differences in physical dimensions and numerical ranges, the data for each dimension needs to be normalized using the min-max normalization method. The normalization formula is as follows: ,in It is the value of the original data point in the i-th dimension, while and These are the minimum and maximum values recorded for this dimension under long-term normal operation of the equipment. For example, based on historical data, the furnace temperature... and The reference temperatures are 850℃ and 950℃, steam flow rates are 1.0 t / h and 3.0 t / h, flue gas oxygen content is 1.5% and 5.5%, and outlet temperatures are 140℃ and 260℃, respectively. These baseline values are pre-calculated and stored to perform a uniform scale transformation on all points within the current sliding window, ensuring that the values of all dimensions fall within the [0, 1] interval. After normalization, to improve the efficiency of subsequent calculations of adjacent point pairs, a spatial grid partitioning method is used to divide the four-dimensional space into multiple hypercube storage buckets. The side length of the storage bucket, i.e., the grid size, is a key parameter, and its setting needs to balance computational efficiency and analytical accuracy. It is set to a fixed value of 0.05, which is obtained based on empirical analysis as half of the minimum attention distance scale in subsequent topology analysis. For each normalized four-dimensional point P(x, y, z, w), its storage bucket index (floor(x / 0.05), floor(y / 0.05)) is calculated. The coordinates are assigned to the corresponding storage buckets, and finally, all the normalized coordinate quadruples are stored in the storage buckets in this form, which has no order and is only organized according to the spatial position relationship. Together, they form a geometric snapshot representing the current operating state of the equipment, that is, the spatial point cloud of the activated carbon regeneration state.
[0044] The steps to obtain the life cycle distribution map of the loop structure are as follows:
[0045] Based on the spatial point cloud of activated carbon regeneration state, the distance scale between points is set up in progressively increasing order. For each distance scale, the coordinate pairs of points whose Euclidean distance is not greater than the distance scale are calculated one by one and connected. All the connections and nodes are organized into a structural graph to form a multi-level point connection network.
[0046] Based on a multi-level point connection network, loop structures are identified scale by scale, and the birth and death scales of each loop structure are recorded. The life cycle length of each loop structure is calculated and mapped to an interval sequence. The distance scale range is divided into multiple equal-width intervals. For each interval, the number of loops in all life cycle coverage intervals is counted, and the normalized probability value is calculated. At the same time, the average life cycle length of the loops in the coverage intervals is calculated to obtain the average life cycle length, thus forming a loop structure life cycle distribution map.
[0047] Specifically, based on the activated carbon regeneration state space point cloud obtained in the aforementioned steps, a sequence of point connection distance scales is set, starting from 0.01 and ending at 0.5, with a step size of 0.01. The range and step size of this distance scale are determined after analyzing the point cloud data generated during normal equipment operation. The lower limit of 0.01 ensures that the densest local data structure can be captured, while the upper limit of 0.5 is sufficient to cover the vast majority of point pairs in the point cloud, allowing the point cloud to evolve into a fully connected graph. The step size of 0.01 strikes a balance between computational accuracy and efficiency. For each distance scale in this sequence, such as 0.01, 0.02, up to 0.5, a complete network construction process is performed. Specifically, for the current distance scale, all coordinate point pairs in the activated carbon regeneration state space point cloud are traversed, and the Euclidean distance of each pair of coordinate points in the four-dimensional normalized space is calculated, that is, the coordinate components corresponding to the two quadruplets are calculated. If the square root of the sum of squares of the differences is less than or equal to the current distance scale, the two coordinate points are considered connected, and an edge is established between them. To optimize computational efficiency, the storage buckets partitioned in the previous steps are used to calculate the distance only between adjacent or same bucket pairs of points, instead of performing a global brute-force pairing. After completing the distance judgment and edge connection operation for all point pairs, all coordinate points in the point cloud are treated as nodes, and all established edges are organized to form an undirected graph structure. This structure graph is stored in the form of an adjacency list, where each node is associated with a list recording all other nodes directly connected to it. As the distance scale increases progressively from 0.01 to 0.5, the number of edges in the point connection network increases monotonically, and the network structure changes from sparse to dense. This series of structure graphs generated at different distance scales together constitute a multi-level point connection network.
[0048] Based on a multi-level point-connected network, starting from the smallest distance scale, the structural graph at each scale is analyzed one by one to identify loop structures. A depth-first search algorithm is used for loop identification. When traversing the nodes of the graph, if a previously visited node is encountered that is not its parent node, a loop is identified. Whenever a new loop structure is identified, the distance scale that initially formed the loop is recorded as the birth scale of the loop structure. Simultaneously, the evolution of this loop structure in subsequent larger distance scales is continuously tracked. When, at a certain distance scale, all nodes constituting the loop form one or more smaller loops due to the appearance of new chords (i.e., edges connecting non-adjacent nodes), or when the entire loop is filled by a higher-dimensional simplex (e.g., a tetrahedron), the loop structure is considered to have disappeared. The smallest distance scale that caused its disappearance is recorded as the death scale of the loop structure. The scale is then used to calculate the lifespan of each loop structure by subtracting the birth scale from the death scale. Next, the entire range of distance scales, from 0.01 to 0.5, is divided into 50 equally wide intervals, each with a width of (0.5-0.01) / 50=0.0098. For any of these 50 intervals, such as the u-th interval, all identified loop structures are traversed, and the number of loops whose lifespans (i.e., the closed interval from the birth scale to the death scale) intersect with the current u-th distance scale interval is counted. This number is divided by the total number of all identified loops to obtain the normalized probability value of the interval. At the same time, for all loops whose lifespans intersect with the interval, the average of their lifespan lengths is calculated to obtain the average lifespan length of the interval. The normalized probability values and average lifespan lengths of these 50 intervals are organized in order to form a loop structure lifespan distribution map.
[0049] The steps for obtaining the topology quantization deviation are as follows:
[0050] Based on the loop structure lifecycle distribution diagram, the topology quantization deviation is calculated using the following formula:
[0051] ;
[0052] in, This represents the probability value of the loop structure's lifetime within the u-th distance scale interval of the current sliding time window. This represents the probability value of the loop structure's lifecycle for the u-th distance scale interval during normal equipment operation. The average lifetime length of the u-th distance scale interval within the current sliding time window. The average lifespan length of the u-th distance scale interval during normal equipment operation. The reference lifetime length is used to achieve dimensionless lifetime measurement. The number of distance scale intervals. Quantify the deviation of the topology.
[0053] Specifically, the formula: The above formula quantifies the difference between the current device state's topology and the normal state baseline using a weighted Hellinger distance. The core of the formula is... It measures the difference in the probability distribution of loop structure occurrence at each distance scale interval, while the exponential weight term The design assigns greater influence to loop structures with longer lifecycles because long-lifecycle loops typically correspond to more stable and representative system structural characteristics. When the distribution of these important characteristics deviates, the weighting terms amplify this deviation, resulting in a more significant final outcome. It is more sensitive to deep structural changes in the state of the equipment.
[0054] The steps for obtaining this parameter are as follows: This parameter represents the probability value of the loop structure lifecycle in the u-th distance scale interval within the current sliding time window. It is directly derived from the loop structure lifecycle distribution map generated in the previous step. Specifically, in the loop structure lifecycle distribution map, the normalized probability value corresponding to the interval with index u is found. This value is obtained by counting the number of loops whose lifecycles cover the u-th interval and then dividing it by the total number of loops identified in the current time window. For example, in the previous step, the distance scale range was divided into 50 intervals. If 1500 loop structures are identified in the current 10-minute time window, and the lifecycles of 180 loops intersect with the 10th interval [0.0992, 0.1090), then the probability value of the loop structure lifecycle in the 10th interval is... The calculation result is 180 divided by 1500, which is 0.12.
[0055] The acquisition steps are as follows: This parameter represents the loop structure lifecycle probability value of the u-th distance scale interval when the equipment is operating normally. It is a preset value used as a comparison benchmark. Its acquisition requires processing the data collected during the long-term operation phase when the equipment is confirmed to be fault-free (e.g., continuous operation for one month). This month's data is divided into multiple sliding time windows of the same size as the window used during early warning (e.g., thousands of 10-minute windows). The aforementioned complete point cloud construction and loop analysis process is performed for each window to obtain a loop structure lifecycle distribution map. Finally, the probability values of the u-th interval in all these distribution maps are averaged to obtain a stable and representative benchmark probability value. For example, by analyzing 3000 normal operating windows, the average probability value of the 10th interval is calculated to be 0.11, then set... .
[0056] The steps to obtain this parameter are as follows: this parameter represents the average lifetime length of the u-th distance scale interval within the current sliding time window, and... Similarly, it is directly obtained from the lifecycle distribution map of the loop structure in the current time window. The calculation process involves finding all loops whose lifecycles intersect with the u-th interval, summing the lifecycle lengths (death scale minus birth scale) of these loops, and then dividing by the number of these loops to obtain the average value. For example, for the 180 loops intersecting with the 10th interval, their lifecycle lengths are calculated, resulting in a numerical sequence {0.05, 0.07, ..., 0.12}. The average of these values is 0.088. .
[0057] The steps for obtaining this parameter are as follows: This parameter represents the average lifespan length of the u-th distance scale interval during normal device operation, and is part of the baseline distribution map. Its acquisition method is the same as... Similarly, in generating the baseline distribution map, in addition to averaging the probability values, the average lifetime length of the u-th interval calculated for each window is also averaged to obtain a stable baseline average lifetime length. For example, through the analysis of 3000 normal operation windows, the average lifespan length of the 10th interval was calculated to be 0.092. Therefore, the following setting is made: .
[0058] The steps for obtaining this parameter are as follows: This parameter is the reference lifetime length, used to perform dimensionless processing on the lifetime length to avoid its numerical value having a disproportionate impact on the exponential weight. Its value is set to the theoretically possible maximum lifetime length, that is, the difference between the set maximum and minimum distance scales. According to the aforementioned settings, the distance scale range is [0.01, 0.5]. Therefore, The value is fixed at 0.5 minus 0.01, which is 0.49.
[0059] The steps to obtain this parameter are as follows: the parameter is the total number of distance scale intervals, which is a parameter preset during loop analysis and determines the accuracy of the loop structure life cycle distribution map. In the aforementioned steps, it has been clearly stated that the entire distance scale range is divided into 50 equal-width intervals. Therefore, the parameter... The value is fixed at 50.
[0060] Calculations based on parameters:
[0061] Taking the calculation of the contribution of the 10th interval (u=10) to the total deviation as an example, let's substitute the previously obtained parameter values:
[0062] , , , , .
[0063] Calculate the index weighting terms:
[0064] ;
[0065] Calculate the probability difference term:
[0066] ;
[0067] Calculate the contribution value of the 10th interval:
[0068] ;
[0069] Repeat the above calculation and summation for all 50 intervals (u=1 to 50). Here are the results for two additional intervals: for example, the contribution value of interval 25 is 0.0004120, and the contribution value of interval 40 is 0.0001588. Summate the contribution values of all 50 intervals to obtain the total. .
[0070] Finally, the topology quantization deviation is calculated. :
[0071] ;
[0072] The results indicate that the device's operating state within the current sliding time window deviates from the baseline of the normal state in terms of topology by 0.27. This value is a dimensionless scalar, and its magnitude reflects the degree of deviation. A value close to 0 means that the current state is highly similar to the normal state in terms of topology, while a larger value indicates a difference. This value of 0.27 will be used as a dimension for subsequent state evaluation, comparing it with a preset warning threshold to determine whether there is a potential risk of device failure. For example, if the warning threshold is set to 0.25 (this threshold can be obtained by analyzing a large number of normal and critical failure states...), then... If the value distribution is selected to effectively distinguish between the two (such as the 99th percentile of the normal value distribution), then the current result of 0.27 will trigger a warning signal.
[0073] The steps to obtain the aggregate dynamic margin are as follows:
[0074] Based on the preset sliding time window, the control power and actual value of furnace temperature are extracted at each moment, a one-to-one corresponding input and output pairing sequence is established, abnormal pairing points are removed and the sampling interval is kept consistent, and an input-output synchronization sequence of furnace temperature control power and actual value of furnace temperature is generated.
[0075] Based on the input-output synchronization sequence of the control power of the furnace temperature and the actual value of the furnace temperature, the order and lag order of the difference equation are set, the input and output quantities are expanded into a matrix expression for each sample, the coefficients of the difference equation are calculated, and all eigenvalues are obtained to obtain the set of eigenvalues.
[0076] Based on the set of eigenvalues, the aggregate dynamic margin is calculated using the following formula:
[0077] ;
[0078] in, Let r be the r-th eigenvalue. Let the modulus of the r-th eigenvalue be . The total number of characteristic roots. This is for aggregate dynamic margin.
[0079] Specifically, within a preset 10-minute sliding time window, similar to the previous steps, the control power (expressed as a percentage) and the actual measured value of the furnace temperature (expressed in degrees Celsius) are synchronously extracted hourly from the control system of the activated carbon regeneration equipment at a frequency of 1 Hz. By comparing the timestamps of each data point, the control power and actual temperature value at the same moment are paired to establish an initial input-output pairing sequence. Then, anomaly removal is performed on this sequence. The criteria for anomaly removal include two aspects: first, numerical range verification. According to the equipment operating procedures, the reasonable range for control power is 0% to 100%, and the reasonable range for furnace temperature is 850℃ to 950℃. Any pairing point outside this range is considered invalid and removed. Second, rate of change verification... By analyzing long-term normal operating data of the equipment, the maximum normal rate of change of furnace temperature was statistically determined to be 1.5℃ per second. Therefore, a rate of change threshold of 2.0℃ / second was set. The first-order difference between adjacent temperature values in the sequence was checked. If the absolute value of the difference exceeds the threshold when divided by the sampling interval (1 second), the point is considered an abrupt change and is removed. After removing all aberrant pairs, the continuity of the timestamps in the sequence is checked to maintain a consistent 1-second sampling interval. If a jump or missing timestamp is found, resulting in an interval that is not 1 second, the data segments before and after the breakpoint are split, and only the longest continuous data segment is retained for subsequent analysis. Through the above processing, a clean, time-series-continuous, and uniformly sampled input-output synchronization sequence of furnace temperature control power and actual furnace temperature value is finally generated.
[0080] Based on the input-output synchronization sequence of the furnace temperature control power and the actual furnace temperature generated in the previous step, an autoregressive (ARX) model with external input is used to establish a difference equation describing the dynamic relationship between the two. This model assumes that the current actual furnace temperature is a linear combination of its past actual values and control power values from several past times. First, the model's order (i.e., how many historical temperature values need to be traced) and lag order (i.e., how many time intervals the control power's influence on temperature is delayed) need to be determined. The order and lag order are optimized using the Akaike Information Criterion (AIC). Specifically, a candidate order range is set, for example, output order from 1 to 5, input order from 1 to 5, and lag order from 1 to 3 sampling periods. All possible combinations of order and lag are iterated. For each combination, the ARX model is fitted using the input-output synchronization sequence data, and its corresponding AIC value is calculated. The model with the highest AIC value is selected. The smaller set of orders and lag orders is used as the final model parameters. For example, calculations show that the AIC value is minimized when the output order is 2, the input order is 2, and the lag order is 1. This determines the form of the difference equation. Then, based on the determined orders and lag orders, the data in the input-output synchronization sequence are expanded sample by sample to construct a matrix expression for an overdetermined system of equations. In this expression, the left side of the equation system is a one-dimensional column vector composed of actual furnace temperature values, and the right side is the product of a design matrix composed of historical temperature and historical control power values and a column vector of coefficients of the difference equation to be solved. The matrix expression is solved using the standard least squares method to calculate the coefficients of the difference equations. Finally, using the coefficients related to historical temperature values, the characteristic polynomial of the system is constructed, and all roots of this polynomial are solved using numerical methods. These roots are the characteristic roots of the system. All the solved characteristic roots are organized together to obtain the characteristic root set.
[0081] formula: The above formula provides a single index for quantifying the dynamic stability margin of a system, namely the aggregate dynamic margin. This originates from discrete-time control theory, which states that all eigenvalues (poles) of a stable system must lie within the unit circle of the complex plane, and the magnitude of the eigenvalues... The value represents the distance from the origin. The farther away from the unit circle boundary (i.e., the smaller the modulus), the faster the response and the better the stability. This formula calculates the root mean square (RMS) value of all eigenvalues, aggregating the combined impact of all dynamic modes on stability. This RMS value can be seen as the "average equivalent radius" of the eigenvalues from the origin in the complex plane. When the system is very stable, all roots are close to the origin, and this value is close to 0. When the system tends to be unstable, at least one root is close to the unit circle, and this value is close to 1. Subtracting this RMS value from 1 yields an intuitive margin indicator. Its range is [0, 1]. The larger the value, the better the dynamic performance and the greater the margin. Conversely, the smaller the value, the worse the dynamic performance of the system.
[0082] The steps to obtain this parameter are as follows: The parameter represents the r-th eigenvalue in the eigenvalue set. It is a complex number and is directly derived from the calculation result of the previous step. In the previous step, after identifying the coefficients of the difference equation using the least squares method, the characteristic polynomial was constructed. For example, if the determined output order is 2, the obtained characteristic polynomial is... , where the coefficient and It was obtained through data fitting, for example and Then the polynomial is Using the quadratic formula, we can find the two characteristic roots as follows: and These two complex numbers are the elements in the set of characteristic roots.
[0083] The steps to obtain the characteristic polynomial are as follows: the parameter is the total number of eigenvalues, and its value is uniquely determined by the output order of the established difference equation (i.e., the order of the autoregressive part). In the previous step, the AIC criterion was used for optimization, and the final selected output order was 2. Therefore, the characteristic polynomial is a second-order polynomial, which must have two roots (possibly real numbers or conjugate complex numbers). Thus, the total number of eigenvalues is... In this example, it is determined to be 2.
[0084] The steps for obtaining this parameter are as follows: the parameter is the modulus of the r-th eigenvalue, that is, the distance from the corresponding point of the complex number in the complex plane to the origin. It is calculated as the square root of the sum of the squares of the real and imaginary parts. For the first eigenvalue obtained above... The calculation process of its modulus is as follows: For the second eigenvalue Its mold is also .
[0085] Calculations based on parameters:
[0086] Substitute the previously obtained parameter values into the formula for calculating the aggregate dynamic margin:
[0087] , , .
[0088] First, calculate the sum of squares of the moduli of all eigenvalues:
[0089] ;
[0090] Then, calculate the mean of the modulus squared:
[0091] ;
[0092] Next, calculate the square root of the mean:
[0093] ;
[0094] Finally, calculate the aggregate dynamic margin. :
[0095] ;
[0096] The results indicate that within the current sliding time window, the aggregate dynamic margin of the furnace temperature control system is 0.2254. This value of 0.2254 indicates that the system currently possesses a certain degree of dynamic stability, but is not in an optimal state. This value will be used as the second dimension of the two-dimensional state feature vector to comprehensively evaluate the equipment's operating status through continuous monitoring. The changing trend of the value can reveal the degradation process of the equipment's dynamic performance. For example, if the value is stable at around 0.5 during normal operation, but the current calculated value is 0.2254, it indicates that the system response is slowing down or the oscillation is increasing. If the value drops further and falls below the preset warning threshold (e.g., 0.2), it indicates potential faults such as aging of the heating element, sluggish sensor response, or mismatch of controller parameters.
[0097] The steps for obtaining the two-dimensional state feature vector are as follows:
[0098] Based on the topological structure, the deviation and aggregation dynamic margin are quantized, and binary sequences are assembled in a fixed order of the first and second dimensions. The decimal places are unified and the original units are maintained. The numerical values are checked for emptiness and the null value handling method and boundary inclusion rules are recorded to form a two-dimensional state feature vector.
[0099] Specifically, based on the topology quantization deviation and aggregation dynamic margin calculated in the preceding steps, these two values are assembled in a fixed order, with the topology quantization deviation as the first dimension and the aggregation dynamic margin as the second dimension, forming an initial binary sequence. For example, if the calculated topology quantization deviation is 0.27 and the aggregation dynamic margin is 0.2254, then the initial binary sequence is (0.27, 0.2254). To maintain data format consistency, the decimal places of these two values are uniformly processed, setting all values to retain four decimal places and rounding them off. Therefore, the above sequence is processed into (0.2700, ...). During the assembly process, the original units of these two values (0.2254) are maintained, meaning they are dimensionless. Next, the assembled binary sequence is validated to check if any dimension is null or invalid (e.g., NaN). This may be due to upstream calculation steps failing to produce valid output due to insufficient data or other anomalies. The null value handling method is set to "previous value filling," meaning that if the calculation result of the current sliding window is null, the two-dimensional state feature vector calculated by the previous valid window is used instead, and this filling event is recorded. At the same time, the boundary inclusion rule is defined. This rule defines how to determine when a certain dimension value of the feature vector is exactly equal to the region boundary value in the subsequent region determination. Here, all boundaries are set to be closed intervals, i.e., including the boundary value. For example, if the lower boundary of the warning region is 0.25, when the calculated value is 0.25, it is determined to fall into the warning region. All this information, including the uniformly formatted values, null value handling records, and boundary inclusion rule settings, is integrated to form the final two-dimensional state feature vector.
[0100] The steps for acquiring equipment operating status signals are as follows:
[0101] Based on the two-dimensional state feature vector, the minimum and maximum coordinates and coordinate axis directions of the preset two-dimensional warning area are read, the size relationship between the first dimension and the second dimension and the corresponding boundary is compared item by item, and the boundary inclusion rule is applied to obtain the landing area judgment mark of the two-dimensional state feature vector.
[0102] Based on the landing zone determination mark of the two-dimensional state feature vector, a Boolean value is generated according to the rule that falling into the preset two-dimensional warning area is recorded as true, and not falling into the preset two-dimensional warning area is recorded as false. A timestamp and source identifier are attached and written to the output channel to generate the device operation status signal.
[0103] Specifically, based on the two-dimensional state feature vector generated in the previous step, the definition parameters of the preset two-dimensional warning region are read from the system configuration file. This two-dimensional warning region is a rectangular area, determined by its minimum and maximum coordinates in the two-dimensional state space. These coordinates define the boundary of the warning region. The minimum coordinate represents the lower left corner of the region, and the maximum coordinate represents the upper right corner. Simultaneously, the direction definition of each coordinate axis is read. The first dimension (X-axis) represents the topology quantization deviation; a larger value indicates a more abnormal state. The second dimension (Y-axis) represents the aggregation dynamic margin; a smaller value indicates a more abnormal state. Therefore, the warning region is usually located in an area with a larger X-axis value and a smaller Y-axis value. The boundary of the two-dimensional warning region... The boundary is obtained through statistical analysis of a large amount of historical data. Specifically, the method involves collecting two-dimensional state feature vectors generated by the equipment under different states such as normal operation, minor faults, moderate faults, and severe faults. These points are then plotted as a scatter plot on a two-dimensional plane. Using a support vector machine (SVM) or similar classification algorithm, the decision boundary that optimally distinguishes the normal state point set from all fault state point sets is found. Based on this boundary, a certain safety margin is set to define the boundary of the warning area. For example, analysis shows that when the topology quantization deviation is greater than 0.25 or the aggregation dynamic margin is less than 0.2, the equipment has a high risk of failure. Therefore, the minimum coordinates of the preset two-dimensional warning area are (0.25, ...). The maximum coordinate is (+∞, 0.2). Then, the two dimensions of the currently calculated two-dimensional state feature vector are compared with the corresponding boundaries of the warning area item by item. For example, for the two-dimensional state feature vector (0.2700, 0.2254), its first dimension 0.2700 is greater than the minimum X coordinate of the warning area 0.25, and its second dimension 0.2254 is greater than the maximum Y coordinate of the warning area 0.2. When comparing, the previously set boundary inclusion rule is applied. According to the comparison result, it is determined whether the feature vector falls into the warning area. In this example, since the first dimension meets the condition but the second dimension does not, the point does not fall into the rectangular warning area defined by the two conditions. Finally, a Boolean judgment mark is obtained, such as "not fallen into".
[0104] Based on the landing zone determination mark of the two-dimensional state feature vector obtained in the previous step, the final signal generation logic is executed. This logic follows a simple binarization rule: if the landing zone determination mark is "falling into the preset two-dimensional warning area", a Boolean value of "True" is generated; if the landing zone determination mark is "not falling into the preset two-dimensional warning area", a Boolean value of "False" is generated. This Boolean value directly reflects whether the equipment is currently in a warning state that requires attention. To ensure the traceability and integrity of the generated signal, additional information is added to the Boolean value. First, the precise timestamp of the end time of the current sliding time window is added, which is consistent with the time base used when the input data was collected. Second, a source identifier is added. This identifier is a unique string used to indicate that the signal was generated by this "Data Analysis-Based Fault Early Warning Method for Activated Carbon Regeneration Equipment", to distinguish it from other alarm or status signals in the equipment control system. For example, if the determination result is "False", the current time is "2023-10-27". If the source identifier is "TPL-ADM_Warning_v1.0", then the final packaged information is {value: False, timestamp: "2023-10-27 10:30:00", source: "TPL-ADM_Warning_v1.0"}. Finally, this structured information containing a boolean value, timestamp, and source identifier is written to a predefined output channel. This output channel can be a message queue (such as MQTT or Kafka), a database table, or a dedicated log file, which can be subscribed to or queried by upper-level monitoring systems (such as SCADA or DCS) to realize the real-time release and integration of fault warning information. This final output structured information is the device operating status signal.
[0105] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.
Claims
1. A fault early warning method for activated carbon regeneration equipment based on data analysis, characterized in that, Includes the following steps: Within a preset sliding time window, real-time values of furnace temperature, steam flow rate, flue gas oxygen content, and outlet temperature of the activated carbon regeneration equipment are collected simultaneously to establish a spatial point cloud of activated carbon regeneration status. Based on the activated carbon regeneration state spatial point cloud, a multi-level point connection network is constructed by successively increasing the scale of the connection distance between points. Loop structures formed at different connection distance scales and the distance scales of their appearance and disappearance are identified to obtain a loop structure life cycle distribution map. Then, the difference between the current loop structure life cycle distribution map and the baseline distribution map corresponding to the normal operation of the equipment is calculated to generate the topology quantification deviation. Within a preset sliding time window, the control power of the furnace temperature is selected as the input and the actual value of the furnace temperature is selected as the output. A difference equation describing the relationship between the two is established, the coefficients of the difference equation are solved, and the aggregate dynamic margin is calculated. The topology quantization deviation is used as the first dimension, and the aggregated dynamic margin is used as the second dimension. Together they form a two-dimensional state feature vector. It is determined whether the two-dimensional state feature vector falls within a preset two-dimensional warning area. Based on the determination result, a Boolean value is output to generate a device operating status signal. The steps for obtaining the lifecycle distribution map of the loop structure are as follows: Based on the activated carbon regeneration state spatial point cloud, a progressively increasing point connection distance scale is set. For each distance scale, coordinate point pairs whose Euclidean distance is not greater than the distance scale are calculated one by one and connected. All the connections and nodes are organized into a structural graph to form a multi-level point connection network. Based on the multi-level point connection network, loop structures are identified on a scale-by-scale basis, and the birth and death scales of each loop structure are recorded. The life cycle length of each loop structure is calculated and mapped to the interval sequence. The distance scale range is divided into multiple equal-width intervals. For each interval, the number of loops in all life cycle coverage intervals is counted, and the normalized probability value is calculated. At the same time, the average life cycle length of the loops in the coverage intervals is calculated to obtain the average life cycle length, thus forming a loop structure life cycle distribution map. The steps for obtaining the aggregation dynamic margin are as follows: Based on the preset sliding time window, the control power and actual value of furnace temperature are extracted at each moment, a one-to-one corresponding input and output pairing sequence is established, abnormal pairing points are removed and the sampling interval is kept consistent, and an input-output synchronization sequence of furnace temperature control power and actual value of furnace temperature is generated. Based on the input-output synchronization sequence of the control power and the actual value of the furnace temperature, the order and lag order of the difference equation are set, the input and output quantities are expanded into a matrix expression for each sample, the coefficients of the difference equation are calculated, and all eigenvalues are obtained to obtain the set of eigenvalues. Calculate the aggregate dynamic margin based on the set of eigenvalues.
2. The method for early warning of activated carbon regeneration equipment based on data analysis according to claim 1, characterized in that, The steps for obtaining the spatial point cloud of the activated carbon regeneration state are as follows: Within a preset sliding time window, real-time values of furnace temperature, steam flow rate, flue gas oxygen content, and outlet temperature are recorded synchronously with a unified timestamp. Any missing or duplicate records are removed to form a synchronous sequence of real-time values of furnace temperature, steam flow rate, flue gas oxygen content, and outlet temperature. Based on the synchronization sequence of the real-time values of furnace temperature, steam flow rate, flue gas oxygen content, and outlet temperature, the corresponding four values are extracted one by one in chronological order and spliced into coordinate quadruplets in a fixed order. Identical coordinate quadruplets are removed and unique identifiers are retained to generate a high-dimensional coordinate point sequence. Based on the high-dimensional coordinate point sequence, the time sequence label is removed and the storage buckets are divided according to the adjacent distance relationship. All coordinate quadruples are collected in no particular order and mapped to the same coordinate system to form a spatial point cloud of activated carbon regeneration state.
3. The method for early warning of activated carbon regeneration equipment based on data analysis according to claim 1, characterized in that, The steps for obtaining the topology quantization deviation are as follows: Based on the loop structure lifecycle distribution diagram, the topology quantization deviation is calculated.
4. The method for early warning of activated carbon regeneration equipment based on data analysis according to claim 1, characterized in that, The steps for obtaining the two-dimensional state feature vector are as follows: Based on the topological structure quantization deviation and the aggregation dynamic margin, a binary sequence is assembled in a fixed order of the first and second dimensions, the decimal places are unified and the original units are maintained, the value is checked for emptiness and the null value handling method and boundary inclusion rules are recorded to form a two-dimensional state feature vector.
5. The method for early warning of activated carbon regeneration equipment based on data analysis according to claim 1, characterized in that, The steps for obtaining the device operating status signal are as follows: Based on the two-dimensional state feature vector, the minimum and maximum coordinates and coordinate axis directions of the preset two-dimensional warning area are read, and the size relationship between the first dimension and the second dimension and the corresponding boundary is compared item by item. The boundary inclusion rule is applied to obtain the landing area judgment mark of the two-dimensional state feature vector.
6. The method for early warning of activated carbon regeneration equipment based on data analysis according to claim 5, characterized in that, The step of obtaining the device operation status signal further includes: generating a Boolean value based on the landing area determination mark of the two-dimensional status feature vector, according to the rule that falling into the preset two-dimensional warning area is recorded as true and not falling into the preset two-dimensional warning area is recorded as false, attaching a timestamp and source identifier and writing it into the output channel to generate the device operation status signal.
Citation Information
Patent Citations
Method for comprehensively evaluating intersection design schemes in variable-weight manner on basis of Monte-Carlo simulation
CN105760634A
Roller kiln system anomaly detection method based on exergy model
CN111059896A