A digital-twin-based electric dust removal power supply fault early warning method and system
By identifying voltage drops and current distortions in electrostatic precipitator power supplies using sliding windows and local statistical features, and confirming abnormal characteristics by combining voltage and current correlation rules, this technology solves the problem of insufficient timeliness and accuracy in fault warnings for electrostatic precipitator power supplies in existing technologies, and achieves high-precision fault warnings.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 阳城国际发电有限责任公司
- Filing Date
- 2026-04-15
- Publication Date
- 2026-06-19
AI Technical Summary
Existing technologies struggle to efficiently identify fault symptoms such as voltage drops and current distortions from the real-time data stream of electrostatic precipitators, resulting in insufficient timeliness and accuracy of early warnings and an inability to achieve early fault warnings for electrostatic precipitators.
Data segments are extracted by sliding windows, local statistical features are calculated, abrupt change points are identified, candidate abnormal segments are screened, and valid abnormal feature segments are confirmed according to the correlation rules between voltage and current. These segments are then packaged into standard data packets and pushed to the fault diagnosis and early warning analysis queue.
It achieves high-precision automatic conversion from real-time data streams to standardized fault characteristics, improving the timeliness and accuracy of electrostatic precipitator power supply fault early warning.
Smart Images

Figure CN122238931A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of power supply failure early warning, and in particular relates to a method and system for early warning of power supply failures in electrostatic precipitators based on digital twins. Background Technology
[0002] In the field of industrial environmental protection equipment maintenance, the stable operation of electrostatic precipitator power supplies is crucial for ensuring continuous production and compliance with pollutant emission standards. Sudden failures often lead to production interruptions and environmental risks. Therefore, accurate fault early warning is a key aspect of improving operation and maintenance efficiency. Current mainstream early warning methods mostly rely on offline analysis of historical data or monitoring with fixed thresholds. These methods are ill-suited to the dynamic changes and complex relationships in equipment operating conditions, and their timeliness and accuracy are limited. Alarms are often issued only after fault characteristics have fully manifested, thus missing the opportunity for early warning.
[0003] The root of this limitation lies in the fact that fault symptoms in electrostatic precipitator power supplies are often hidden in the real-time dynamic sequence of key parameters such as secondary voltage and current, manifesting as brief voltage drops or transient current distortions. The core technical challenge in capturing these symptoms is the need to efficiently and accurately identify these characteristic segments containing fault information from the continuously flowing, high-speed real-time data stream synchronously generated by the digital twin. This first requires establishing a real-time data extraction mechanism that can closely coordinate with fault screening logic to ensure the targeted and timely supply of data. However, simply achieving data extraction is insufficient. The deeper technical challenge lies in designing an algorithm that can automatically identify and capture these non-stationary, sporadic abnormal features within a continuous time-series data stream, because fault characteristics are not continuous but rather brief pulses or pattern deviations mixed within a large amount of normal data.
[0004] Therefore, how to combine data extraction for fault screening with automatic capture of abnormal features in real-time data streams, so as to extract standardized potential fault data packets from massive twin data in a timely and accurate manner, has become a key issue in realizing early warning and precise intervention of electrostatic precipitator power supplies. Summary of the Invention
[0005] To address the aforementioned technical problems, this invention provides a method for early warning of electrostatic precipitator power supply failures based on digital twins, comprising: The secondary voltage and secondary current signals of the electrostatic precipitator power supply are continuously acquired from the digital twin to form a real-time data stream; The real-time data stream is continuously segmented using a sliding window of fixed time length to obtain data segments to be analyzed arranged in chronological order. For each of the data segments to be analyzed, calculate its local statistical characteristics; Based on the local statistical characteristics, identify the mutation points in the data segment to be analyzed; Based on the number of mutation points, candidate abnormal segments are selected from the data segments to be analyzed; Based on the preset voltage and current correlation rules, the candidate abnormal segments are reviewed to obtain valid abnormal feature segments; Extract the original time-series data, timestamps, and local statistical features corresponding to the effective abnormal feature segments, and encapsulate them into a data packet with a unified structure; The data packet is pushed to the fault diagnosis and early warning analysis queue.
[0006] Optionally, the step of continuously acquiring the secondary voltage and secondary current signals of the electrostatic precipitator power supply from the digital twin to form a real-time data stream includes: Acquire high-frequency real-time data streams of the secondary voltage and secondary current of the electrostatic precipitator power supply; The high-frequency real-time data stream is segmented using a sliding time window to obtain voltage and current sequences within the window; The instantaneous power sequence is calculated based on the voltage and current sequences. If the variance of the instantaneous power sequence exceeds a preset threshold, the power supply operating state is determined to have entered the fluctuation range. Within the fluctuation range, abnormal data points in the instantaneous power sequence are detected; Based on the distribution density and consecutive occurrence frequency of the abnormal data points, the potential fault level of the power supply equipment is determined.
[0007] Optionally, the step of continuously capturing the real-time data stream using a sliding window of fixed time length to obtain data segments to be analyzed arranged in chronological order includes: Acquire the real-time data stream; The real-time data stream is continuously segmented using a sliding window of fixed time length to obtain data segments; If the time series of the data segment meets the preset requirements, it is taken as the data segment to be analyzed; Extract the change characteristics of the data segment to be analyzed within the time window; Identify anomalous changes in the data segment to be analyzed; Based on the correspondence between the abnormal change characteristics and the sliding window, the time range of the abnormality is determined.
[0008] Optionally, calculating the local statistical characteristics for each of the data segments to be analyzed includes: The data segment to be analyzed is divided into multiple time slices using a sliding window method, and the center value of the data in each time slice is calculated. Calculate the dispersion of data within the same time slice, and combine the center value with the dispersion to form the feature set of the time slice; If the dispersion of the feature set exceeds a preset threshold, the corresponding time slice is marked as a high-fluctuation segment; Based on the labeling results of the high-fluctuation segments, abnormal segments are identified.
[0009] Optionally, determining the mutation points in the data segment to be analyzed based on the local statistical characteristics includes: Calculate the local mean and local standard deviation of the data segment to be analyzed; The dynamic threshold is obtained based on the preset multiple and the local standard deviation; Calculate the difference between the instantaneous value at each sampling point and the local mean; If the difference is greater than the dynamic threshold, the sampling point is determined to be a mutation point; Based on the location of all the mutation points, the abnormal intervals in the data segment to be analyzed are determined.
[0010] Optionally, the step of filtering candidate anomalous segments from the data segment to be analyzed based on the number of mutation points includes: The sliding window method is used to divide the data sequence into multiple data segments; Calculate the number of mutation points within each of the data segments; The number of mutation points in each data segment is compared with a preset density threshold; If the number of mutation points exceeds the preset density threshold, the corresponding data segment is marked as a candidate anomalous segment; Based on the location information of all the candidate abnormal segments, the distribution range of the abnormal event in the original sequence is determined.
[0011] Optionally, the step of reviewing the candidate abnormal segments according to preset voltage and current correlation rules to obtain valid abnormal feature segments includes: Extract the occurrence times of voltage drop points and current distortion points from the time series data of the candidate abnormal segments; A preset time window is used to match the voltage drop point with the current distortion point; If a voltage drop point and a current distortion point are located within the same time window, then it is determined that there is time coupling between the two. Count the number of temporally coupled drop points and distortion points within a single candidate anomalous segment; If the number exceeds a preset threshold, the candidate abnormal segment is determined to be a valid abnormal feature segment.
[0012] On the other hand, the present invention also provides a digital twin-based electrostatic precipitator power supply fault early warning system, comprising: The data stream acquisition module is used to continuously acquire the secondary voltage and secondary current signals of the electrostatic precipitator power supply from the digital twin, forming a real-time data stream; The sliding window segmentation module is used to continuously segment the real-time data stream using a sliding window of fixed time length to obtain data segments to be analyzed arranged in chronological order. The local feature calculation module is used to calculate the local statistical features of each of the data segments to be analyzed; The mutation point determination module is used to determine the mutation points in the data segment to be analyzed based on the local statistical characteristics. The candidate fragment determination module is used to filter out candidate abnormal fragments from the data fragments to be analyzed based on the number of mutation points. The association rule verification module is used to verify the candidate abnormal segments according to the preset voltage and current association rules to obtain valid abnormal feature segments; The feature data encapsulation module is used to extract the original time-series data, timestamps and local statistical features corresponding to the effective abnormal feature fragments, and encapsulate them into a data packet with a unified structure. The data packet push module is used to push the data packet to the fault diagnosis and early warning analysis queue.
[0013] On the other hand, the present invention also provides an electronic device including a memory, a processor, and a computing program stored in the memory and executable on the processor, wherein the processor implements the method when executing the computing program.
[0014] On the other hand, the present invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method.
[0015] Compared with the prior art, the present invention has the following advantages and technical effects: This invention extracts data segments by using a sliding window and calculates local statistical features. It identifies abrupt change points based on dynamic thresholds to screen candidate abnormal segments. Then, it performs coupling verification according to preset voltage and current correlation rules. After confirming the valid abnormal feature segments, it encapsulates them with timestamps and statistical features into a standard data package and pushes it to the analysis queue. This achieves high-precision automatic conversion and extraction of real-time data into standardized fault features. Attached Figure Description
[0016] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a schematic diagram of the method flow according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the system structure according to an embodiment of the present invention. Detailed Implementation
[0017] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.
[0018] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.
[0019] Example 1 This embodiment provides a method for early warning of electrostatic precipitator power supply failure based on digital twins, including: like Figure 1 As shown, the secondary voltage and secondary current signals of the electrostatic precipitator power supply are continuously acquired from the digital twin to form a real-time data stream with high-frequency sampling. A sliding window of fixed time length is used to continuously capture the real-time data stream, resulting in a series of data segments to be analyzed arranged in chronological order. For each data segment to be analyzed, its mean and standard deviation within a short time window are calculated as local statistical features of that segment; The instantaneous value of each sampling point in the data segment to be analyzed is compared with the local statistical characteristics of the segment. If the instantaneous value deviates from the mean by more than a preset multiple of the standard deviation, the sampling point is determined to be a mutation point. If the number of mutation points in a data segment to be analyzed exceeds a preset density threshold, the segment is determined to be a candidate abnormal segment. According to the preset voltage and current correlation rules, the candidate abnormal segments are reviewed. If there is a temporal coupling relationship between voltage drop and current distortion within the segment, the segment is confirmed as a valid abnormal feature segment. Extract the original time-series data corresponding to the effective abnormal feature segments, their timestamps, and the calculated local statistical features, and combine and encapsulate them into a data packet with a unified structure; The encapsulated data packet is pushed to the fault diagnosis and early warning analysis queue to complete the transformation from real-time data stream to standardized potential fault data packet.
[0020] Furthermore, the step of continuously acquiring the secondary voltage and secondary current signals of the electrostatic precipitator power supply from the digital twin to form a high-frequency sampled real-time data stream includes: High-frequency real-time data streams of secondary voltage and secondary current of the electrostatic precipitator power supply are obtained from the digital twin; A sliding time window is used to segment the data stream, resulting in voltage and current sequences within the window; Calculate the product of the voltage sequence and the current sequence within each time window to obtain the instantaneous power sequence; If the variance of the instantaneous power sequence exceeds a preset threshold, the power supply is judged to have entered the fluctuation range. Within the fluctuation range, the isolated forest algorithm is used to detect outlier data points in the instantaneous power sequence; The potential fault level of the power supply equipment is determined based on the distribution density and consecutive occurrence frequency of abnormal data points.
[0021] Furthermore, a sliding window of fixed time length is used to continuously capture the real-time data stream, resulting in a series of data segments to be analyzed arranged in chronological order, including: Obtain real-time data stream; The data stream is continuously captured using a sliding window of fixed time length; Obtain data segments arranged in chronological order; If the time sequence of a data segment meets the preset requirements, then the segment is determined to be the data segment to be analyzed. Based on the data segment to be analyzed, extract the change characteristics of the data stream within the time window; The isolated forest algorithm is used to identify anomalous changes in data segments. Based on the correspondence between abnormal features and sliding windows, the time range of the abnormality is determined.
[0022] Furthermore, for each of the data segments to be analyzed, calculating its mean and standard deviation within a short time window as local statistical characteristics of that segment includes: Acquire time series data stream; The data stream is divided using a sliding window method to obtain multiple time slices; For each time slice, calculate the center value of the data within it; Calculate its dispersion based on data from the same time slice; The center value and dispersion of each time slice are combined to form the feature set of that time slice; If the dispersion of the feature set exceeds a preset threshold, the corresponding time slice is marked as a high-fluctuation segment; Based on the labeling results of high-fluctuation segments, the isolated forest algorithm is used to identify anomalous segments.
[0023] Furthermore, the step of comparing the instantaneous value of each sampling point in the data segment to be analyzed with the local statistical characteristics of the segment, and determining that the sampling point is a mutation point if the instantaneous value deviates from the mean by more than a preset multiple of the standard deviation, includes: Obtain the time series data segment to be analyzed; Calculate the local mean and local standard deviation of the data segment; The dynamic threshold is obtained by multiplying a preset factor by the local standard deviation; For each sampling point, calculate the difference between its instantaneous value and the local mean; If the difference is greater than the dynamic threshold, the sampling point is determined to be a mutation point; Based on the location of all mutation points, identify the anomalous intervals in the data segment.
[0024] Furthermore, if the number of mutation points in a data segment to be analyzed exceeds a preset density threshold, then the segment is determined to be a candidate anomalous segment, including: Obtain the data sequence to be analyzed; The data sequence is segmented using a sliding window method to obtain multiple data segments; For each data segment, calculate the number of internal mutation points; The number of mutation points in each data segment is compared with a preset density threshold; If the number of mutation points exceeds a preset density threshold, the data segment is marked as a candidate anomalous segment; Based on the location information of all candidate abnormal segments, the distribution range of the abnormal event in the original sequence is determined.
[0025] Furthermore, the candidate abnormal segments are reviewed according to preset voltage and current correlation rules. If a voltage drop and current distortion within a segment are coupled in time, the segment is confirmed as a valid abnormal feature segment, including: Obtain time-series data of candidate anomaly segments; Extract the occurrence times of voltage sag points and current distortion points from time series data; A preset time window is used to match the drop point with the distortion point; If a voltage drop point and a current distortion point are located within the same time window, then it is determined that there is time coupling between the two. Count the number of sudden drop points and distortion points with temporal coupling within a single candidate segment; If the number exceeds a preset threshold, the candidate segment is determined to be a valid abnormal feature fragment.
[0026] Furthermore, the extracted original time-series data corresponding to the effective abnormal feature fragments, their timestamps, and the calculated local statistical features are combined and encapsulated into a data packet with a unified structure, including: Obtain raw time-series stream data; Identify the occurrence points of characteristic segments in a time-series stream; Extract the original data corresponding to the occurrence points of the feature segments; Record the timestamp of the occurrence point of the feature segment; Calculate the local statistical values of the original data within the feature segment; The raw data, timestamps, and statistical values are encapsulated into a structured data package.
[0027] Furthermore, the step of pushing the encapsulated data packet to the fault diagnosis and early warning analysis queue to complete the transformation from real-time data stream to standardized potential fault data packet includes: Acquire standardized fault data packets from the fault diagnosis and early warning analysis queue; Using preset parsing rules, the device identifier and timestamp sequence are separated from the data packet; Based on the device identification code, match the normal operating parameter range of the device from the historical database; If the monitored values in the timestamp sequence continuously exceed the normal operating parameter range, the data packet is determined to contain an abnormal event; For abnormal events, the duration and fluctuation range of the monitored values exceeding the threshold are extracted as feature vectors; The feature vector is input into a pre-defined isolated forest model to obtain the anomaly score of the anomalous event. If the abnormal score exceeds the preset isolation threshold, an early warning record containing the device identifier, timestamp, and abnormal type will be generated.
[0028] Example 2 This embodiment provides a method for early warning of electrostatic precipitator power supply failure based on digital twins, including: Step S101: Continuously acquire the secondary voltage and secondary current signals of the electrostatic precipitator power supply from the digital twin to form a real-time data stream with high-frequency sampling.
[0029] High-frequency real-time data streams of the secondary voltage and secondary current of the electrostatic precipitator power supply are acquired from a digital twin. A sliding time window is used to segment the data stream, obtaining voltage and current sequences within each window. The product of the voltage and current sequences within each time window is calculated to obtain the instantaneous power sequence. If the variance of the instantaneous power sequence exceeds a preset threshold, the power supply is considered to have entered a fluctuation range. Within this fluctuation range, an isolated forest algorithm is used to detect abnormal data points in the instantaneous power sequence. Based on the distribution density and frequency of consecutive occurrences of these abnormal data points, the potential fault level of the power supply equipment is determined.
[0030] The secondary voltage and secondary current signals of the electrostatic precipitator power supply are continuously acquired from the digital twin and synchronously sampled at a frequency of 1,000 times per second by a data acquisition agent deployed at the edge, forming a high-frequency real-time data stream.
[0031] For example, 1000 voltage data points (unit: kV) and 1000 current data points (unit: mA) can be obtained within 1 second. These data points are timestamped in milliseconds and then pushed to the stream processing platform in real-time via the MQTT protocol in JSON format (e.g., {"timestamp": 1720000000123", "U2": 72.5", "I2": 1200.8}). Upon receiving the data stream, the platform first performs data quality verification, for example, applying a threshold-based outlier detection algorithm. If the voltage values of five consecutive sampling points exceed a preset safety range (e.g., greater than 80.0 kV or less than 30.0 kV), the data for that period is marked as invalid and an alarm is triggered. Subsequently, a sliding window calculation is performed on the valid data stream. For example, every 200 milliseconds (containing 200 original sampling points) serves as an analysis window. Within the window, the average value of the secondary voltage, the effective value of the secondary current, and the instantaneous power of both (P=U) are calculated. (I) To further analyze the operating status, the system calculates the difference between peak and valley voltage values in real time and analyzes the frequency of spark discharges in conjunction with the current waveform. For example, if 15 pulses with an instantaneous current value exceeding 1500.0 mA and a duration of less than 2 milliseconds are detected within a 10-second observation period, it is inferred that dense spark discharges may have occurred in the electric field. These characteristic indicators, processed and analyzed in real time, along with the raw data stream, are continuously written into the time-series database, providing input for subsequent energy efficiency optimization and fault prediction models.
[0032] Step S102: A sliding window of fixed time length is used to continuously capture the real-time data stream to obtain a series of data segments to be analyzed arranged in chronological order.
[0033] Acquire real-time data stream. Continuously capture segments of the data stream using a sliding window of fixed time length. Obtain data segments arranged in chronological order. If the time sequence of a data segment meets preset requirements, that segment is identified as the data segment to be analyzed. Based on the data segment to be analyzed, extract the change characteristics of the data stream within the time window. Use the isolated forest algorithm to identify abnormal change characteristics in the data segment. Determine the time range of anomaly occurrence based on the correspondence between abnormal characteristics and the sliding window.
[0034] In real-time data stream processing, a sliding window with a fixed duration of 5 seconds is used to continuously capture data. The window sliding step size is set to 1 second, which means that the window slides forward once every 1 second, and each time the data within the most recent 5 seconds is captured as a segment to be analyzed.
[0035] For example, for a continuously generated temperature sensor data stream in the format "timestamp:temperature value", at system time T=10 seconds, the current window will cover all data points from T=6 seconds to T=10 seconds (inclusive), forming a data segment to be analyzed. Then, at T=11 seconds, the window slides, and a new segment covers the data from T=7 seconds to T=11 seconds, and so on, resulting in a series of data segments arranged chronologically with a 1-second overlap. For each captured 5-second data segment, a specific analysis algorithm is immediately applied for processing, such as calculating the average temperature within that time period to monitor for anomalies. Assuming a segment contains 5 data points with temperature values of [22.1, 22.3, 35.6, 22.2, 22.0] degrees Celsius, the analysis process first calls a simple statistical function to calculate the average, resulting in (22.1 + 22.3 + 35.6 + 22.2 + 22.0) / 5 = 24.84 degrees Celsius. Next, the result is compared with a preset threshold (e.g., 26 degrees Celsius). Since 24.84 is less than 26, the system determines that the data for this window period is normal. If the average temperature calculated for the next segment exceeds the threshold, an alarm mechanism is triggered.
[0036] Step S103: For each data segment to be analyzed, calculate its mean and standard deviation within a short time window as local statistical features of the segment.
[0037] Acquire the time-series data stream. Segment the data stream using a sliding window method to obtain multiple time slices. For each time slice, calculate the median value of the data within it. Calculate the dispersion of the data within the same time slice. Combine the median value and dispersion of each time slice to form a feature set for that time slice. If the dispersion in the feature set exceeds a preset threshold, the corresponding time slice is marked as a high-volatility segment. Based on the marking results of high-volatility segments, use the Isolation Forest algorithm to identify anomalous segments.
[0038] For each data segment to be analyzed, such as a sensor time-series signal with 1024 sampling points, we first define a short time window with a length of 256 sampling points, sliding in steps of 128 sampling points. For the first window, i.e., the data points with indices from 1 to 256, we apply the mean algorithm, i.e., calculate the arithmetic mean of all data points within the window, specifically the formula μ=(Σx_i) / n, where n=256. Assuming the sum of the data points in this window is 5120.0, then its local mean μ_1=5120.0 / 256=20.0. Next, we calculate the standard deviation of this window to measure the dispersion of the data, using the algorithm σ=sqrt((Σ(x_i-μ)^2) / (n-1)), where n-1 is used for unbiased estimation. We first calculate the sum of squares of the deviations of each data point from the mean of 20.0. Assuming this sum is 2048.0, the standard deviation σ_1 = sqrt(2048.0 / 255) ≈ sqrt(8.0314) ≈ 2.834. This gives us the local statistical characteristics of the first data segment (μ_1 = 20.0, σ_1 ≈ 2.834). Then, we slide the window to indices 129 to 384 and repeat the calculation to obtain the second set of characteristic values, for example (μ_2 = 22.5, σ_2 ≈ 3.1). By continuously sliding the window and calculating, we generate a series of local statistical characteristic sequences for the entire data segment being analyzed. These sequences effectively reflect the central trend and fluctuations of the signal within a short time window.
[0039] Step S104: Compare the instantaneous value of each sampling point in the data segment to be analyzed with the local statistical characteristics of the segment. If the instantaneous value deviates from the mean by more than a preset multiple of the standard deviation, the sampling point is determined to be a mutation point.
[0040] Obtain the time series data segment to be analyzed. Calculate the local mean and local standard deviation of the data segment. Multiply the local standard deviation by a preset factor to obtain a dynamic threshold. For each sampling point, calculate the difference between its instantaneous value and the local mean. If the difference is greater than the dynamic threshold, the sampling point is determined to be a mutation point. Based on the locations of all mutation points, determine the abnormal intervals in the data segment.
[0041] First, information technology is used to acquire the data segment to be analyzed, such as a voltage time series containing 1000 sampling points at a sampling frequency of 1kHz and a data segment duration of 1 second. Next, the local statistical characteristics of this segment are calculated. The specific algorithm is as follows: using a sliding window method with a window size of 101 points (corresponding to approximately 0.1 seconds), the local mean and local standard deviation are calculated by taking the 50 neighboring points before and after each sampling point as the center.
[0042] For example, for the 300th sampling point, take 101 data points from index 250 to 350, calculate the local mean μ_local=(Σx_i) / 101 using the formula, and calculate the local standard deviation σ_local=sqrt(Σ(x_i-μ_local)^2 / 100). Then, compare the instantaneous value of each sampling point with its corresponding local statistical feature, with a preset multiple threshold of 3.
[0043] For example, the instantaneous value of the 300th sampling point is x_300 = 5.2V, while its local mean μ_local is 3.1V and its local standard deviation σ_local is 0.6V. The calculated deviation is |5.2 - 3.1| = 2.1V, which is approximately 3.5 times the standard deviation (2.1 / 0.6 ≈ 3.5 times), exceeding the preset threshold of 3 times. Therefore, the system automatically identifies this sampling point as a mutation point. To ensure the robustness of the analysis, the influence of potential mutation points is excluded when calculating local statistical characteristics. For example, when calculating the local mean and standard deviation of the 300th point, this point's value can be temporarily excluded, and only its neighboring 100 points are used for calculation. This prevents mutation points from artificially inflating the standard deviation and causing missed detections.
[0044] Step S105: If the number of mutation points in a data segment to be analyzed exceeds a preset density threshold, the segment is determined to be a candidate abnormal segment.
[0045] Obtain the data sequence to be analyzed. Segment the data sequence using a sliding window method to obtain multiple data segments. For each data segment, calculate the number of mutation points within it. Compare the number of mutation points in each data segment with a preset density threshold. If the number of mutation points exceeds the preset density threshold, mark the data segment as a candidate anomalous segment. Based on the location information of all candidate anomalous segments, determine the distribution range of the anomalous event in the original sequence.
[0046] In the data processing workflow, the genome sequence is first divided into continuous data fragments using a sliding window algorithm. For example, each fragment is set to be 1000 base pairs in length, with a 200-base-pair overlap between adjacent fragments to ensure that boundary mutations are not missed. Next, a mutation detection algorithm based on statistical hypothesis testing is applied, such as using the Bayesian Information Criterion (BIC) combined with a piecewise linear model to analyze sequence fluctuations within each fragment. When the change in the BIC value exceeds a preset ΔBIC = 10, a mutation is identified at that location. Subsequently, the system automatically calculates the mutation density within each fragment, which is the number of mutations divided by the fragment length. If a fragment (e.g., a fragment from position 5000 to 6000) detects 7 mutations, its density is 7 / 1000 = 0.007. This density value is compared with a preset density threshold, which can be obtained through training with historical normal data, for example, set to 0.005. Since 0.007 is greater than 0.005, the system automatically identifies the fragment as a candidate anomalous fragment and stores its identifier in the database for subsequent in-depth analysis.
[0047] Step S106: According to the preset voltage and current correlation rules, the candidate abnormal segments are reviewed. If the voltage drop and current distortion within the segment are coupled in time, the segment is confirmed as a valid abnormal feature segment.
[0048] Obtain time-series data of candidate anomaly segments. Extract the occurrence times of voltage sags and current distortions from the time-series data. Match sags and distortions using a preset time window. If a voltage sag and a current distortion occur within the same time window, they are considered to be temporally coupled. Count the number of temporally coupled sags and distortions within a single candidate segment. If this number exceeds a preset threshold, the candidate segment is determined to be a valid anomaly feature segment.
[0049] In the preset voltage and current correlation rules, voltage sag is defined as a drop in the effective voltage value exceeding 15% of the rated value (e.g., 220V) within 0.1 seconds and lasting for at least 3 cycles. Current distortion is defined as an instantaneous total harmonic distortion (THD) value exceeding 20% or a sudden increase in the content of a specific harmonic (e.g., the 5th) exceeding 8%. When reviewing candidate abnormal segments, a sliding time window (0.2 seconds window length, 0.01 seconds step size) is first used to simultaneously extract the effective voltage value sequence and the current THD sequence within the segment. The time coupling relationship is quantified by calculating the dynamic time warping (DTW) distance between the two or the time difference correlation coefficient at a set threshold (e.g., ±20 milliseconds).
[0050] For example, in a candidate segment starting at 10:05:23.450, the algorithm detected a voltage drop to 185V (a 15.9% decrease) at 10:05:23.452, while the current THD jumped to 28.5% at 10:05:23.455. Calculations showed that the time difference between the voltage drop and the start of the current distortion was 3 milliseconds, less than the preset 20-millisecond coupling threshold, and their trends were highly synchronized within the following 0.15 seconds (DTW distance less than 0.1). Further analysis of harmonic components revealed a sudden increase in the 5th harmonic content from 3% to 12%, consistent with the phase characteristics of the voltage dip. This confirmed a clear causal relationship between the electrical disturbances within the segment, classifying it as a valid anomalous segment, and recording its start and end times, minimum voltage value, maximum current THD, and coupling strength coefficient (e.g., 0.92).
[0051] Step S107: Extract the original time-series data corresponding to the effective abnormal feature fragment, the timestamp of its occurrence, and the calculated local statistical features, and combine and encapsulate them into a data packet with a unified structure.
[0052] Acquire raw time-series stream data. Identify the occurrence points of feature segments in the time-series stream. Extract the raw data corresponding to the occurrence points of feature segments. Record the timestamps of the occurrence points of feature segments. Calculate the local value statistics of the raw data within the feature segments. Encapsulate the raw data, timestamps, and statistics into a structured data packet.
[0053] First, the raw time-series data is segmented using a sliding window algorithm. For example, vibration signals are collected at a rate of 100 sampling points per second, with a window length of 500 points (5 seconds) and a step size of 100 points (1 second). When data within a window exceeds a preset threshold (e.g., the root mean square value is greater than 2.5 for three consecutive windows), it is identified as a valid anomalous feature segment. Next, the corresponding raw data points are extracted, for example, 320 raw amplitude values lasting 3.2 seconds from the timestamp 2023-10-27 14:05:23.450, with a numerical sequence such as [0.12, 0.85, 1.34, -0.92]. Simultaneously, the precise start and end timestamps of this anomalous segment are recorded. Then, local statistical features are calculated for this segment data using a specific algorithm: calculating the segment's mean (e.g., 0.68), standard deviation (e.g., 1.23), peak value (e.g., 3.45), and zero-crossing rate (e.g., 45 times per second).
[0054] Step S108: Push the encapsulated data packet to the fault diagnosis and early warning analysis queue to complete the transformation from real-time data stream to standardized potential fault data packet.
[0055] Acquire standardized fault data packets from the fault diagnosis and early warning analysis queue. Using preset parsing rules, separate the device identifier and timestamp sequence from the data packets. Based on the device identifier, match the device's normal operating parameter range from the historical database. If the monitored values in the timestamp sequence consistently exceed the normal operating parameter range, the data packet is determined to contain an abnormal event. For abnormal events, extract the duration and fluctuation amplitude of the monitored values exceeding the threshold as feature vectors. Input the feature vectors into a preset isolated forest model to obtain the abnormality score for the abnormal event. If the abnormal score exceeds a preset isolated threshold, generate an early warning record containing the device identifier, timestamp, and abnormality type.
[0056] After the data packet is encapsulated, the system pushes it to the fault diagnosis and early warning analysis queue through a RabbitMQ-based message queue service. Specifically, the system calls the AMQP protocol interface to send the encapsulated JSON format data packet, such as the device ID "CNC-001", timestamp "2023-10-27 14:30:25.550", vibration feature vector [0.12, 0.85, 1.34, 0.07] and temperature value 45.6, to the queue named "fault_diagnosis_queue" in persistent mode. At the same time, the message priority is set to 5, and the production end release confirmation mechanism is enabled, thereby completing the transformation from real-time data stream to standardized potential fault data packet. Subsequently, the diagnostic analysis service, acting as a consumer, pulls data packets from the queue. It first performs JSON parsing and data integrity verification, then calls a pre-trained Isolation Forest algorithm model for anomaly detection. This model has been trained with 100 decision trees using historical normal data. It calculates anomaly scores for the input vibration feature vector. If the score exceeds a threshold of 0.65, an early warning process is triggered. The system performs similarity matching between the data packet and a historical fault case database, using a cosine similarity algorithm. If the feature similarity with the case "bearing wear" reaches 0.8 or higher, an early warning work order with an 85% confidence level is automatically generated and stored in the early warning database. Simultaneously, the early warning information is pushed to the monitoring dashboard in real time via a WebSocket service. like Figure 2 As shown, on the other hand, this embodiment also provides a digital twin-based electrostatic precipitator power supply fault early warning system, which mainly includes: The data stream acquisition module is used to continuously acquire the secondary voltage and secondary current signals of the electrostatic precipitator power supply from the digital twin, forming a high-frequency sampled real-time data stream; The sliding window segmentation module is used to continuously segment the real-time data stream using a sliding window of fixed time length to obtain a series of data segments to be analyzed arranged in chronological order. The local feature calculation module is used to calculate the mean and standard deviation of each data segment to be analyzed within a short time window, as the local statistical features of the segment. The mutation point determination module is used to compare the instantaneous value of each sampling point in the data segment to be analyzed with the local statistical characteristics of the segment. If the instantaneous value deviates from the mean by more than a preset multiple of the standard deviation, the sampling point is determined to be a mutation point. The candidate segment determination module is used to determine that a segment is a candidate abnormal segment if the number of mutation points in a segment to be analyzed exceeds a preset density threshold. The association rule verification module is used to verify the candidate abnormal segments according to the preset voltage and current association rules. If there is a temporal coupling relationship between voltage drop and current distortion in the segment, the segment is confirmed as a valid abnormal feature segment. The feature data encapsulation module is used to extract the original time-series data corresponding to the effective abnormal feature fragment, the timestamp of its occurrence, and the calculated local statistical features, and combine and encapsulate them into a data packet with a unified structure. The data packet push module is used to push the encapsulated data packet to the fault diagnosis and early warning analysis queue, completing the transformation from real-time data stream to standardized potential fault data packet.
[0057] On the other hand, this embodiment also provides an electronic device, including a memory, a processor, and a computing program stored in the memory and executable on the processor, wherein the processor implements the method when executing the computing program.
[0058] On the other hand, this embodiment also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method.
[0059] The above are merely preferred embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for early warning of power supply failure in electrostatic precipitators based on digital twins, characterized in that, include: The secondary voltage and secondary current signals of the electrostatic precipitator power supply are continuously acquired from the digital twin to form a real-time data stream; The real-time data stream is continuously segmented using a sliding window of fixed time length to obtain data segments to be analyzed arranged in chronological order. For each of the data segments to be analyzed, calculate its local statistical characteristics; Based on the local statistical characteristics, identify the mutation points in the data segment to be analyzed; Based on the number of mutation points, candidate abnormal segments are selected from the data segments to be analyzed; Based on the preset voltage and current correlation rules, the candidate abnormal segments are reviewed to obtain valid abnormal feature segments; Extract the original time-series data, timestamps, and local statistical features corresponding to the effective abnormal feature segments, and encapsulate them into a data packet with a unified structure; The data packet is pushed to the fault diagnosis and early warning analysis queue.
2. The method according to claim 1, characterized in that, The process of continuously acquiring secondary voltage and secondary current signals from the electrostatic precipitator power supply from the digital twin to form a real-time data stream includes: Acquire high-frequency real-time data streams of the secondary voltage and secondary current of the electrostatic precipitator power supply; The high-frequency real-time data stream is segmented using a sliding time window to obtain voltage and current sequences within the window; The instantaneous power sequence is calculated based on the voltage and current sequences. If the variance of the instantaneous power sequence exceeds a preset threshold, the power supply operating state is determined to have entered the fluctuation range. Within the fluctuation range, abnormal data points in the instantaneous power sequence are detected; Based on the distribution density and consecutive occurrence frequency of the abnormal data points, the potential fault level of the power supply equipment is determined.
3. The method according to claim 1, characterized in that, The method of continuously capturing the real-time data stream using a sliding window of fixed time length to obtain data segments to be analyzed arranged in chronological order includes: Obtain the real-time data stream; The real-time data stream is continuously segmented using a sliding window of fixed time length to obtain data segments; If the time sequence of the data segment meets the preset requirements, it is taken as the data segment to be analyzed; Extract the change characteristics of the data segment to be analyzed within the time window; Identify anomalous changes in the data segment to be analyzed; Based on the correspondence between the abnormal change characteristics and the sliding window, the time range of the abnormality is determined.
4. The method according to claim 1, characterized in that, The step of calculating the local statistical characteristics for each of the data segments to be analyzed includes: The data segment to be analyzed is divided into multiple time slices using a sliding window method, and the center value of the data in each time slice is calculated. Calculate the dispersion of data within the same time slice, and combine the center value with the dispersion to form the feature set of the time slice; If the dispersion of the feature set exceeds a preset threshold, the corresponding time slice is marked as a high-fluctuation segment; Based on the labeling results of the high-fluctuation segments, abnormal segments are identified.
5. The method according to claim 1, characterized in that, The step of determining the mutation points in the data segment to be analyzed based on the local statistical characteristics includes: Calculate the local mean and local standard deviation of the data segment to be analyzed; The dynamic threshold is obtained based on the preset multiple and the local standard deviation; Calculate the difference between the instantaneous value at each sampling point and the local mean; If the difference is greater than the dynamic threshold, the sampling point is determined to be a mutation point; Based on the location of all the mutation points, the abnormal intervals in the data segment to be analyzed are determined.
6. The method according to claim 1, characterized in that, The step of selecting candidate anomalous segments from the data segment to be analyzed based on the number of mutation points includes: The sliding window method is used to divide the data sequence into multiple data segments; Calculate the number of mutation points within each of the data segments; The number of mutation points in each data segment is compared with a preset density threshold; If the number of mutation points exceeds the preset density threshold, the corresponding data segment is marked as a candidate anomalous segment; Based on the location information of all the candidate abnormal segments, the distribution range of the abnormal event in the original sequence is determined.
7. The method according to claim 1, characterized in that, The process involves reviewing the candidate abnormal segments according to preset voltage and current correlation rules to obtain valid abnormal feature segments, including: Extract the occurrence times of voltage drop points and current distortion points from the time series data of the candidate abnormal segments; A preset time window is used to match the voltage drop point with the current distortion point; If a voltage drop point and a current distortion point are located within the same time window, then it is determined that there is time coupling between the two. Count the number of temporally coupled drop points and distortion points within a single candidate anomalous segment; If the number exceeds a preset threshold, the candidate abnormal segment is determined to be a valid abnormal feature segment.
8. A fault early warning system for electrostatic precipitator power supply based on digital twin, characterized in that, include: The data stream acquisition module is used to continuously acquire the secondary voltage and secondary current signals of the electrostatic precipitator power supply from the digital twin, forming a real-time data stream; The sliding window segmentation module is used to continuously segment the real-time data stream using a sliding window of fixed time length to obtain data segments to be analyzed arranged in chronological order. The local feature calculation module is used to calculate the local statistical features of each of the data segments to be analyzed; The mutation point determination module is used to determine the mutation points in the data segment to be analyzed based on the local statistical characteristics. The candidate fragment determination module is used to filter out candidate abnormal fragments from the data fragments to be analyzed based on the number of mutation points. The association rule verification module is used to verify the candidate abnormal segments according to the preset voltage and current association rules to obtain valid abnormal feature segments; The feature data encapsulation module is used to extract the original time-series data, timestamps and local statistical features corresponding to the effective abnormal feature fragments, and encapsulate them into a data packet with a unified structure. The data packet push module is used to push the data packet to the fault diagnosis and early warning analysis queue.
9. An electronic device comprising a memory, a processor, and a computing program stored in the memory and executable on the processor, characterized in that, When the processor executes the computing program, it implements the method of any one of claims 1-7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1-7.