Pre-cluster data election method based on stacked sliding time window
By employing stacked sliding time window technology and a steady-state threshold dynamic adjustment algorithm, the problems of data asynchrony and surge in communication volume during multi-channel data acquisition were solved, achieving efficient and reliable data processing and accurate judgment, thereby improving the stability and real-time performance of the power monitoring system.
Patent Information
- Application Number
- CN202511234624.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-01
- Publication Date
- 2025-11-11
AI Technical Summary
Existing technologies lack multi-source data fusion and intelligent judgment mechanisms in multi-channel data acquisition, leading to data asynchrony and a surge in communication volume. They cannot effectively handle the problem of asynchronous multi-channel data frequencies, affecting the accuracy of judgments by the dispatch center.
A front-end cluster data election method based on a stacked sliding time window is adopted. After generating trusted data by performing single-machine data election locally, it is sent to other front-end servers for multi-machine verification, which reduces the amount of communication and realizes the distributed sharing of computing tasks. The sliding time window mechanism and steady-state threshold dynamic adjustment algorithm are used for data alignment and processing.
It increases the probability of successful data convergence and overall processing efficiency, improves data response speed and processing timeliness, ensures efficient and orderly data transmission at different processing stages, and enhances the accuracy and reliability of data judgment.
Smart Images

Figure CN120934976A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to data election, and more specifically, to a front-end cluster data election method based on a stacked sliding time window. Background Technology
[0002] In power monitoring systems, the front-end system undertakes core functions such as multi-channel data access, protocol parsing, and data uploading to SCADA (Supervisory Control and Data Acquisition) systems, making it a crucial link in ensuring stable system operation. To improve system reliability and stability, the power industry widely adopts various data acquisition modes, including channel cold standby acquisition, primary / standby front-end cold standby acquisition, and primary / standby front-end dual-plane acquisition. Among these, primary / standby front-end dual-plane acquisition has become one of the mainstream acquisition architectures due to its higher fault tolerance and data redundancy.
[0003] However, current systems mostly employ a "single-channel upload" strategy, which selects data from one channel out of multiple channels and uploads it to the SCADA system. This approach faces the following challenges in actual operation: due to channel acquisition delays, data loss, or inconsistent sampling frequencies, data from multiple channels may become out of sync. If the selected channel exhibits an anomaly, it will directly affect the accuracy of the dispatch center's judgment and may even lead to erroneous operations.
[0004] In existing technologies, although patent CN115471211A demonstrates how establishing a telemetry data verification mechanism can improve the reliability of switch telemetry data in distribution automation systems, it primarily targets the validity detection of data from devices such as FTUs, TTUs, and DTUs, lacking multi-source data fusion and intelligent judgment mechanisms. Patent CN115396752B proposes a Redis-based dual-plane data acquisition method, which can avoid data loss during faults, but its "centralized judgment" mechanism requires all data to be transmitted to the main front-end server for unified analysis and processing, resulting in a surge in communication volume and increased load on the main front-end server, easily creating bottlenecks. It also fails to address the alignment difficulties caused by different frequencies of multi-channel data. Furthermore, while patent CN112446801B describes a data quality improvement system that introduces multi-dimensional verification modules such as models, graphics, and links, it still has shortcomings in handling concurrent multi-channel acquisition and data redundancy judgment, and cannot identify the optimal channel data in real time. While the master-slave interaction method proposed in CN115622251A can improve the flexibility and operational efficiency of the control system, it still does not involve the management and optimization strategy of front-end data redundancy.
[0005] In summary, existing technologies still have many shortcomings in processing multi-channel acquired data. They lack a comprehensive judgment mechanism for the consistency, validity, and accuracy of multi-source data, fail to make full use of the resources in the acquisition cluster, and do not combine mathematical methods such as sliding time windows to perform dynamic analysis and selection of data.
[0006] No effective solutions have yet been proposed to address the problems in the relevant technologies. Summary of the Invention
[0007] To address the problems in related technologies, this invention proposes a front-end cluster data election method based on a stacked sliding time window, in order to overcome the aforementioned technical problems existing in the current related technologies.
[0008] Therefore, the specific technical solution adopted by the present invention is as follows:
[0009] A pre-cluster data election method based on stacked sliding time windows, the method comprising:
[0010] S1. Based on the physical connection between the front-end server and the gateway, configure the corresponding communication parameters and collect data from all channels on the same cross-section through the front-end server;
[0011] S2. Use the stacked sliding time window technology to align the data collected from each channel, analyze the data deviation and generate an initial deviation threshold, and synchronize the initialization progress to all front-end servers.
[0012] S3. Each front-end server performs a single-machine election based on its local channel data, determines the target channel data, and uploads it to the remaining front-end servers in the front-end cluster in a preset format. A multi-machine comparison mechanism is then executed to generate a comparison sequence.
[0013] S4. Based on the generated comparison sequence, combined with the election result deviation and the local device management authority, the front-end server executes the corresponding data processing strategy.
[0014] Furthermore, a layered sliding time window technique is used to align the data acquired from each channel, analyze data deviations, generate an initial deviation threshold, and synchronously initialize the progress to all front-end servers, including:
[0015] S21. In the same acquisition section, when the number of target data is not less than the preset channel number threshold, the data acquired by each channel is aligned using the stacked sliding time window technology.
[0016] S22. Analyze the dispersion of data between channels and generate the initial deviation threshold for the current section by calculating the data deviation;
[0017] S23. The initial deviation threshold is processed using a steady-state threshold dynamic adjustment algorithm, and the processed data is uploaded to the data acquisition and monitoring control system, and the progress is synchronously initialized to all front-end servers.
[0018] Furthermore, within the same acquisition section, when the amount of target data is not less than a preset channel number threshold, the data acquired from each channel is aligned using a stacked sliding time window technique, including:
[0019] S211. Based on the time difference of data transmission from each channel and the real-time requirements of the data acquisition and monitoring control system, set the initial length of the sliding time window.
[0020] S212. When the target channel is refreshed, the time of data transmission is used as the timing origin. Combined with the set initial length and preset step size, the data of the remaining channels within the corresponding time range is searched and matched.
[0021] S213. Based on the search and matching results, dynamically adjust the initial length of the sliding time window, generate an optimized time window configuration, and perform alignment processing on the data of each channel in the current section according to the time window configuration.
[0022] Furthermore, the search matching results include:
[0023] If the time for a successful data match is greater than the first target threshold, then the length of the current time window is maintained.
[0024] If the time for a successful data match is less than the second target threshold, the time window will be reduced to the third target threshold in the next cycle.
[0025] If there is no matching data in the current time window, the target channel data will be used as the local matching result and sent to the remaining front-end servers for judgment, and the matching time window length will be increased in the next cycle.
[0026] Furthermore, the dispersion of data between channels is analyzed, and an initial deviation threshold for the current section is generated by calculating the data deviation, including:
[0027] S221. Based on the aligned channel data, extract the target data points within the same time period, and calculate the average value of the target data points as the reference data for the corresponding acquisition section.
[0028] S222. Using the reference data as a reference, select the data point closest to the reference data from all channels as the selected point, and calculate the absolute deviation between the remaining data points and the selected point.
[0029] S223. Sum the absolute deviation values of all channels, obtain the average value, and get the initial deviation threshold within the current time window.
[0030] Furthermore, each front-end server performs a single-machine election based on its local channel data to determine the target channel data, and uploads it to the remaining front-end servers in the front-end cluster according to a preset format. A multi-machine comparison mechanism is then executed to generate a comparison sequence, including:
[0031] S31. Based on the initialization progress synchronization results, each front-end server performs a single-machine election on the channel data connected to its local machine to obtain the target channel data;
[0032] S32. The acquired target channel data is sent to the remaining front-end servers in the front-end cluster according to the preset format, and then enters the delay mode.
[0033] S33. When all front-end servers in the cluster complete data reception confirmation or reach the preset synchronization timeout threshold, execute the multi-machine comparison mechanism to generate a comparison sequence.
[0034] Furthermore, delay modes include:
[0035] If the processing results of the remaining front-end servers are received within the preset waiting time window, the processing results of multiple front-end servers will be analyzed.
[0036] If no processing result is received from the remaining front-end servers within the preset waiting time window, the waiting time window will be extended in the next processing cycle.
[0037] Further analysis of the processing results from multiple front-end servers includes:
[0038] If the difference between the timestamp of the remaining front-end server data received and the time base of the local data is within the preset consistency threshold, the data reporting mechanism is triggered, and the local machine or the corresponding front-end server is notified to send the data.
[0039] If the difference between the timestamp of the remaining front-end server data and the time base of the local data exceeds the preset consistency threshold, the current reporting operation is stopped, and a sliding time window waiting process is entered. After all target channel data is received, a multi-front-end server data comparison is performed, and a judgment is made based on the comparison results.
[0040] Furthermore, extending the waiting time window includes:
[0041] Determine whether the extended waiting time window length exceeds the difference between the arrival time of the remaining front-end server data and the local data time; if it exceeds, shorten the waiting time window for the next cycle according to the preset configuration; if it does not exceed, extend the waiting time window by the preset step size.
[0042] Further data processing strategies include:
[0043] If the election result deviation is within the preset threshold range and the local machine has device management authority, the data will be sent to the data acquisition and monitoring control system and the remaining front-end servers will be notified; if the election result deviation is within the preset threshold range and the local machine does not have device management authority, the data confirmation will be sent to the remaining front-end servers and the front-end servers with management authority will execute the data upload.
[0044] If the election result deviates beyond a preset threshold, a negative confirmation is sent, and the front-end server with jurisdiction will send the data back and mark the data as abnormal.
[0045] The beneficial effects of this invention are as follows:
[0046] 1. This invention performs single-machine data election locally, generates a reliable data according to predetermined verification rules, and then sends it to other front-end servers for multi-machine verification. This reduces the amount of communication between multiple front-end servers and realizes the distributed sharing of computing tasks, thereby reducing the processing pressure on the main server caused by high concurrency load and improving the probability of successful data convergence and overall processing efficiency.
[0047] 2. This invention achieves efficient data processing by employing a layered sliding time window mechanism, thereby improving data response speed and processing timeliness. Through continuous sliding and dynamic switching of the time window, it ensures that data can be efficiently and orderly transmitted to the data acquisition and monitoring control system at different processing stages, achieving synergistic optimization of timeliness and reliability.
[0048] 3. This invention uses a dynamic adjustment algorithm based on a steady-state threshold to utilize reliable data deviations during historical steady-state periods as a reference for consistency judgment. As the amount of historical data increases, the system can gradually and dynamically converge the deviation threshold to a smaller range, thereby improving the accuracy and reliability of data judgment in election decisions. Attached Figure Description
[0049] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0050] Figure 1 This is a flowchart of a pre-cluster data election method based on a stacked sliding time window according to an embodiment of the present invention;
[0051] Figure 2 This is a duplex mode diagram in a front-end cluster data election method based on a stacked sliding time window according to an embodiment of the present invention;
[0052] Figure 3 This is a half-duplex mode diagram in a front-end cluster data election method based on a stacked sliding time window according to an embodiment of the present invention;
[0053] Figure 4 This is a simplex mode diagram in a front-end cluster data election method based on a stacked sliding time window according to an embodiment of the present invention;
[0054] Figure 5 This is a single-machine data selection logic diagram in a front-end cluster data election method based on a stacked sliding time window according to an embodiment of the present invention;
[0055] Figure 6 This is a flowchart of a pre-cluster data election method based on a stacked sliding time window according to an embodiment of the present invention. Detailed Implementation
[0056] To further illustrate the various embodiments, the present invention provides accompanying drawings, which are part of the disclosure of the present invention. These drawings are mainly used to illustrate the embodiments and can be used in conjunction with the relevant descriptions in the specification to explain the operating principles of the embodiments. With reference to these drawings, those skilled in the art should be able to understand other possible implementation methods and the advantages of the present invention.
[0057] According to an embodiment of the present invention, a method for selecting front-end cluster data based on a stacked sliding time window is provided.
[0058] The present invention will now be further described in conjunction with the accompanying drawings and specific embodiments, such as... Figure 1 As shown, according to an embodiment of the present invention, a pre-cluster data election method based on a stacked sliding time window includes:
[0059] S1. Based on the physical connection between the front-end server and the gateway, configure the corresponding communication parameters and collect data from all channels on the same cross-section through the front-end server.
[0060] It should be noted that, depending on the different on-site deployments, the front-end data access includes the following three modes:
[0061] 1. Full-duplex mode: such as Figure 2 As shown, the primary and backup front-end servers establish links with the primary and backup gateways respectively. Each front-end server establishes 4 links with the primary and backup gateways, and each gateway accesses data from two planes respectively.
[0062] 2. Half-duplex mode: such as Figure 3 As shown, this addresses the situation where some gateway devices on-site only support one link connection per network port.
[0063] 3. Simplex mode: such as Figure 4As shown, this addresses the situation where some gateway devices on-site only accept one connection.
[0064] S2. Use the stacked sliding time window technology to align the data collected from each channel, analyze the data deviation and generate an initial deviation threshold, and synchronously initialize the progress to all front-end servers.
[0065] In this optional embodiment, the data acquired from each channel is aligned using a stacked sliding time window technique, data deviations are analyzed and an initial deviation threshold is generated, and the progress is synchronously initialized to all front-end servers, including:
[0066] S21. In the same acquisition section, when the number of target data is not less than the preset channel number threshold, the data acquired by each channel is aligned using the stacked sliding time window technique.
[0067] In this optional embodiment, when the number of target data is not less than a preset channel number threshold in the same acquisition section, the data acquired by each channel is aligned using the stacked sliding time window technique, including:
[0068] S211. Based on the time difference of data transmission from each channel and the real-time requirements of the data acquisition and monitoring control system, set the initial length of the sliding time window.
[0069] S212. When the target channel is refreshed, the time of data transmission is used as the timing origin. Combined with the set initial length and preset step size, the data of the remaining channels within the corresponding time range is searched and matched.
[0070] In this optional embodiment, the search matching results include:
[0071] If the time for a successful data match is greater than the first target threshold, then the length of the current time window is maintained.
[0072] If the time for a successful data match is less than the second target threshold, the time window will be reduced to the third target threshold in the next cycle.
[0073] If there is no matching data in the current time window, the target channel data will be used as the local matching result and sent to the remaining front-end servers for judgment, and the matching time window length will be increased in the next cycle.
[0074] S213. Based on the search and matching results, dynamically adjust the initial length of the sliding time window, generate an optimized time window configuration, and perform alignment processing on the data of each channel in the current section according to the time window configuration.
[0075] S22. Analyze the dispersion of data between channels and generate the initial deviation threshold for the current section by calculating the data deviation.
[0076] In this optional embodiment, analyzing the dispersion of data between channels and generating an initial deviation threshold for the current cross-section by calculating data deviation includes:
[0077] S221. Based on the aligned channel data, extract the target data points within the same time period, and calculate the average value of the target data points as the reference data for the corresponding acquisition section.
[0078] S222. Using the reference data as a reference, select the data point closest to the reference data from all channels as the selected point, and calculate the absolute deviation between the remaining data points and the selected point.
[0079] S223. Sum the absolute deviation values of all channels, obtain the average value, and get the initial deviation threshold within the current time window.
[0080] S23. The initial deviation threshold is processed using a steady-state threshold dynamic adjustment algorithm, and the processed data is uploaded to the data acquisition and monitoring control system, and the progress is synchronously initialized to all front-end servers.
[0081] It's worth noting the use of cascaded sliding time windows: Sliding time windows are a commonly used method in data quality processing, particularly suitable for cleaning, denoising, anomaly detection, and missing value imputation of time-series data. The core idea is to use dynamic windows to calculate and correct local data, ensuring data quality remains stable over time. The specific implementation methods and steps are as follows:
[0082] 1. Set window time:
[0083] Based on the characteristics of power systems, data is highly time-sensitive, and the window time should not be set too long. In this invention, a window time of 1 to 10 seconds is set, and the sliding step is set to 1 second.
[0084] The window size is dynamically adjusted based on the data matching degree. If data is refreshed in channel A, the upload time of this data is used as the timing origin. Data from other channels is searched and matched in steps of 1 / 10 of the time window (when the time window is 1 second, the step size should be 100 milliseconds). If the successful matching time is greater than 60% of the time window, the current time window size is maintained; if it is less than 30% of the current time window, the next matching time window is reduced by 50% (not less than the minimum time window of 1 second). The window size should be reduced as much as possible when the data matching degree is high to reduce the resource consumption of data processing. If there is no matching data in the current time window, the data in channel A is used as the matching data of this front-end machine and sent to other front-end machines for further judgment. The next calculation time window is increased by 1 second (i.e., one sliding step).
[0085] 2. Dynamic calculation of windows:
[0086] As shown in Table 1, considering the high real-time requirements of power system monitoring and the fact that preprocessing is only for qualitative analysis of data, "median filtering" is adopted. The following issues must also be considered when using median filtering:
[0087] (1) Cold start problem: When there is insufficient data in the initial stage, a historical global threshold can be used or the window can be gradually widened. Each subsequent window should contain at least one historical data point. Keep as much data as possible within the time window to ensure the accuracy of "median filtering".
[0088] (2) Overfitting risk: Avoid threshold jitter caused by excessively small window size or overly sensitive feedback.
[0089] (3) Concept drift: Periodically recalibrate the model or statistics (e.g., reset the window daily / weekly).
[0090] Table 1: Preprocessing Algorithm Table
[0091]
[0092] Because the data may not be completely consistent, a method based on "dynamic thresholds for statistics" is used to process the data consistency of the analog quantities. The principle is as follows:
[0093] (1) Calculate the threshold in real time based on the standard deviation of the statistical indicators within the sliding window:
[0094] Select data points with good data quality from all channels within the same time period (time deviation set to 100 milliseconds, configurable), and take the average value as the baseline data; select the data point closest to the baseline data as the selected point, calculate the deviation between other data points and the selected point, sum the absolute values of all deviations, and then average them as the threshold; record all thresholds within the time window to form a threshold sequence, which is used to prepare for the dynamic selection of subsequent thresholds.
[0095] (2) Calculate the mean (μ) and standard deviation (σ) within the window:
[0096] The expression for the mean is: avg_value=(value1+value2+....+valuen) / n; where avg_value represents the average value of the cross-sectional data; valuen represents the value of a data point in channel n; and n represents the number of channels.
[0097] Standard deviation: At time t1, the average of the tolerances of channel 1 and other channels at the same time is denoted as difft1; the standard deviation σ = (difft1 + ... + difftn) / n, and 3σ is chosen as the dynamic threshold.
[0098] S3. Each front-end server performs a single-machine election based on its local channel data, determines the target channel data, and sends it to the remaining front-end servers in the front-end cluster in a preset format. A multi-machine comparison mechanism is then executed to generate a comparison sequence.
[0099] In this optional embodiment, each front-end server performs a single-machine election based on its local channel data to determine the target channel data, and uploads it to the remaining front-end servers in the front-end cluster according to a preset format. A multi-machine comparison mechanism is then executed to generate a comparison sequence, including:
[0100] S31. Based on the initialization progress synchronization results, each front-end server performs a single-machine election on the channel data connected to its local machine to obtain the target channel data.
[0101] S32. The acquired target channel data is sent to the remaining front-end servers in the front-end cluster according to the preset format, and then enters the delay mode.
[0102] In this optional embodiment, the delay mode includes:
[0103] If the processing results of the remaining front-end servers are received within the preset waiting time window, the processing results of multiple front-end servers will be analyzed.
[0104] In this optional embodiment, the analysis of the processing results of multiple front-end servers includes:
[0105] If the difference between the timestamp of the remaining front-end server data received and the time base of the local data is within the preset consistency threshold, the data reporting mechanism is triggered, and the local machine or the corresponding front-end server is notified to send the data.
[0106] If the difference between the timestamp of the remaining front-end server data and the time base of the local data exceeds the preset consistency threshold, the current reporting operation is stopped, and a sliding time window waiting process is entered. After all target channel data is received, a multi-front-end server data comparison is performed, and a judgment is made based on the comparison results.
[0107] If no processing result is received from the remaining front-end servers within the preset waiting time window, the waiting time window will be extended in the next processing cycle.
[0108] In this optional embodiment, extending the waiting time window includes:
[0109] Determine whether the extended waiting time window length exceeds the difference between the arrival time of the remaining front-end server data and the local data time; if it exceeds, shorten the waiting time window for the next cycle according to the preset configuration; if it does not exceed, extend the waiting time window by the preset step size.
[0110] It should be noted that the process of uploading the acquired target channel data to the remaining front-end servers in the front-end cluster according to a preset format and entering a delay mode includes:
[0111] Step 1: Based on incremental coding pipeline technology, the acquired target channel data is split into logical blocks according to a preset format and a differential fingerprint sequence is generated, which is then sent to the remaining front-end servers in the front-end cluster in parallel.
[0112] 1. The acquired target channel data is split into several logical data blocks according to preset rules.
[0113] 2. Use the differential fingerprint algorithm to process each logic block and generate the corresponding differential fingerprint sequence.
[0114] (1) Extract feature parameters from each logical data block and input them into a preset hash function to generate the corresponding hash value;
[0115] (2) Match the corresponding weights according to the obtained hash values, construct the weighted fingerprint vector of each logical block, and sum all the weighted fingerprint vectors to form a global fingerprint vector;
[0116] (3) The global fingerprint vector is processed by bit symbolization and converted into a binary sequence to generate the corresponding differential fingerprint sequence.
[0117] 3. Calculate the difference between the differential fingerprint sequence generated in the current period and the differential fingerprint sequence saved in the previous period, and identify the differential data blocks.
[0118] 4. Incrementally encode the identified differential data blocks, encapsulate them into data packets in a unified format, and send the data packets in parallel to the remaining front-end servers in the front-end cluster using a multi-channel mechanism.
[0119] Step 2: Activate the chain confirmation clock mechanism to trigger the remaining front-end servers to verify the differential fingerprint sequence step by step in topological order, and synchronously enter the delay mode.
[0120] 1. After the front-end server completes the parallel data upload, the chained confirmation clock mechanism is started to generate a unified verification timing benchmark and send a verification start signal to the remaining front-end servers in the front-end cluster.
[0121] 2. Upon receiving the start signal, each front-end server receives the differential fingerprint sequence in sequence according to the preset network topology and performs consistency verification locally.
[0122] 3. Once the front-end server completes local consistency verification or receives a verification pass signal, it synchronously enters the delay mode.
[0123] S33. When all front-end servers in the cluster complete data reception confirmation or reach the preset synchronization timeout threshold, execute the multi-machine comparison mechanism to generate a comparison sequence.
[0124] S4. Based on the generated comparison sequence, combined with the election result deviation and the local device management authority, the front-end server executes the corresponding data processing strategy.
[0125] In this optional embodiment, the data processing strategy includes:
[0126] If the election result deviation is within the preset threshold range and the local machine has device management authority, the data will be sent to the data acquisition and monitoring control system and the remaining front-end servers will be notified; if the election result deviation is within the preset threshold range and the local machine does not have device management authority, the data confirmation will be sent to the remaining front-end servers and the front-end servers with management authority will execute the data upload.
[0127] If the election result deviates beyond a preset threshold, a negative confirmation is sent, and the front-end server with jurisdiction will send the data back and mark the data as abnormal.
[0128] It should be added that a front-end cluster data election method based on stacked sliding time windows solves the data preprocessing problem in the case of multi-machine and multi-channel acquisition of the front-end cluster, including the preprocessing of a single front-end server and the data comparison mechanism of multiple machines in the front-end cluster.
[0129] 1. Processing a single data acquisition server:
[0130] The acquisition frequency of each channel is not exactly the same. A sliding time window method is used to align the data sections. Considering the time difference between the data transmission of each channel, and at the same time ensuring the real-time transmission of SCADA (Data Acquisition and Monitoring Control System) data, a short delay processing mode is adopted, with a delay value of ±100 to ±500 milliseconds and a default value of ±100 milliseconds. When there are multiple data within the time window, "delay / 10" is used as the step size until the multi-channel data is aligned.
[0131] First, we process the data from two channels on the same plane. Under normal circumstances, since the data from these two channels comes from the same data source, the data has high data and time consistency. The following situations exist:
[0132] (1) For analog data:
[0133] If the data deviation is within ±3σ and the time difference is within ±100 milliseconds, it indicates that the data quality is consistent and the data is reliable.
[0134] If the data deviation is outside the range of ±3σ and the time difference exceeds ±100 milliseconds, it indicates that the data quality of a certain channel is abnormal.
[0135] (2) For numerical data:
[0136] If the data is consistent and the time difference is within ±100 milliseconds, it indicates that the data quality is consistent and the data is reliable.
[0137] If the data is inconsistent and the time difference exceeds ±100 milliseconds, it indicates that the data quality of a certain channel is abnormal.
[0138] (3) For event (SOE) data:
[0139] Since SOEs are sent to the front-end server with timestamps and allow for delayed transmission, a maximum time window of 60 seconds is used for SOE judgment. At the same time, a 1-hour or 1024-entry SOE buffer is set for each front-end server. If the SOE values of the two channels are consistent with the timestamps within the time window, the data is normal; if the timestamps are the same but the values are different, the SOE is abnormal; if when channel A sends an SOE, channel B does not have a SOE with a consistent timestamp within the sliding window, and there is no such record in the SOE buffer, it indicates that the SOE transmission of channel B is abnormal.
[0140] After data election, each front-end server selects one channel as its valid channel, based on the criteria shown in Table 2.
[0141] Table 2: Selection Criteria Table
[0142]
[0143] The decision results are communicated to other data acquisition servers in the cluster via a message bus. The messages use a small binary format, and the message body format is shown in Table 3.
[0144] Table 3: Message Body Format Table
[0145]
[0146]
[0147] The response message format is shown in Table 4:
[0148] Table 4: Response Message Format Table
[0149]
[0150]
[0151] 2. Data decision-making mechanism for multiple front-end servers within the cluster:
[0152] After single-machine processing is completed, the process enters the waiting phase for processing results from other front-end servers. If processing results from other front-end servers are received within the waiting time window (default ±100 milliseconds), the processing results from multiple front-end servers are analyzed. This paper only focuses on two dimensions: data consistency and time consistency. Based on the multi-machine processing results, taking two front-end servers as an example, as shown in Table 5:
[0153] Table 5: Examples of Front-End Servers
[0154]
[0155]
[0156] The use of the stacked time window is as follows: if the consistency between the received data from other front-end servers and the local data is within the threshold range, the data is immediately uploaded or other front-end servers are notified to upload the data without waiting for all channels in the same section to finish receiving the data. If the data consistency does not meet the threshold requirement, the sliding time window duration must be waited for all valid channel data to be received before multi-machine data comparison is performed. Based on the comparison results, a judgment is made on whether to upload the data locally, notify other machines to upload the data, or determine if the data is abnormal.
[0157] If no processing result is received from any other front-end server within the waiting time, the waiting time window needs to be extended for the next processing. Generally, it is increased in steps of 100 milliseconds. Considering the real-time requirements of power system data, it is not recommended that the waiting time window exceed 500 milliseconds (except for SOE). Conversely, if the length of the waiting time window exceeds the time difference between the arrival time of data from other front-end servers and the local data, the waiting time window can be reduced according to the configuration. It can be by doubling or doubling the step size. Each of these two methods has its own advantages and disadvantages.
[0158] One step size: Reduces the delay in sending data to SCADA (Data Acquisition and Monitoring Control System). In cases of unstable data arrival, less data may fall within the waiting time window, requiring readjustment of the waiting time window.
[0159] Double step size: The data transmission delay to SCADA is slightly longer, and more data from the front-end servers in the cluster may be obtained during the waiting time.
[0160] Through data interaction between multiple front-end servers, each front-end server can dynamically know information such as the number of front-end servers in the cluster. If information from all front-end servers in the cluster is received before the waiting window expires, the data can be quickly sent to SCADA (Supervisory Control and Data Acquisition), further improving the real-time performance of power monitoring system data.
[0161] This method is not suitable for half-duplex and simplex modes due to the limited number of data sources.
[0162] 3. Other optimizations:
[0163] Based on the above processing logic, when the data from front-end server A is normal, the data transmission to SCADA is completed by front-end server A. During the configuration phase, a front-end machine to which a certain "communication unit" belongs can be defined. Under normal circumstances, the data transmission to SCADA is completed by the designated front-end server. If the designated front-end server is abnormal, the logic of automatically identifying the transmitting machine will be restored. The load can be further evenly distributed to different front-end servers, thereby increasing the stability of the acquisition system.
[0164] 4. Exception handling:
[0165] This invention categorizes and manages anomalies encountered during the above processing into three levels: Level I (lowest), Level II, and Level III (highest).
[0166] Level I: Data at a certain cross-section is abnormal, and there is no periodic anomaly. This type of anomaly is only serialized locally.
[0167] Level II: Multiple consecutive cross-sections show abnormal data. This type of data is serialized locally and then sent to SCADA (Supervisory Control and Data Acquisition) to alert the dispatcher for manual intervention.
[0168] Level III: Data is consistently abnormal. Based on the local serialization of this type of data, the scheduler is reminded to intervene manually and the data in this channel is suspended for a period of time. The suspension time also uses a sliding time window. This article uses a cyclical pattern of (1 minute, 1 hour, 1 day).
[0169] As shown in Table 6, the configuration of the communication unit should include the following fields:
[0170] Table 6: Field Table
[0171]
[0172] The single-machine data selection logic is as follows: Figure 5 As shown, the data from the four channels are labeled ①②③④, where data ① and ③ are valid, data ② is discarded because it exceeds the threshold, and data ④ is discarded because it exceeds the time window.
[0173] In a specific embodiment, such as Figure 6 As shown, taking the dual-front (A, B) duplex mode as an example:
[0174] Step 1: Configure the gateway according to the actual physical connection of the gateway on site, as shown in Table 7.
[0175] Gateway A's IP addresses: 172.20.99.100, 172.21.99.100;
[0176] The IP addresses of gateway B are 172.20.99.200 and 172.21.99.200.
[0177] Table 7: Gateway Configuration Table
[0178] property value Ip1 172.20.99.100 Port1 2404 Flag1 Plane A IP2 172.21.99.100 Port2 2404 Flag2 Plane B IP3 172.20.99.200 Port3 2404 Flag3 Plane A IP4 172.21.99.200 Port4 2404 Flag4 Plane B BelongMachine mac1
[0179] Step 2: Configure the cluster identifier of each front-end server, and configure the initial time window length, initial delay processing, and the number of historical records or duration to be saved for each data point.
[0180] Step 3: Program execution and data collection. When the number of sampling points on the same acquisition section is greater than or equal to half the number of channels, these data deviations are analyzed to form an initial value for the data deviation threshold. During this stage, the data sent to SCADA (Data Acquisition and Monitoring Control System) is sent by the initially configured machine, and the local initialization progress is synchronized among multiple machines.
[0181] Step 4: After initialization, the data election process begins. Each front-end machine uses the four channels connected to its local machine to elect a trusted channel, which is then transmitted over the network using a message in Table 3 format and enters a delay processing flow.
[0182] Step 5: Once all front-end servers receive data simultaneously, or the timeout period is reached, analyze the received data from other front-end servers and the data generated locally to form a comparison sequence.
[0183] Step 6: If the local election result deviates from the election results of other front-end servers within the threshold range, and the local machine has jurisdiction over this device, then immediately send data to SCADA and send data to SCADA to notify other nodes.
[0184] Step 7: When the election results of this machine deviate from the election results of other front-end servers within the threshold range, the jurisdiction of this device is not on this machine. In this case, data confirmation is sent to other nodes, and the front-end machine with jurisdiction is responsible for sending the data.
[0185] Step 8: When the local election result deviates from the election results of other front-end servers by more than a threshold, a negative acknowledgment is sent to other nodes. The front-end server with jurisdiction then sends the data and marks the data as suspicious.
[0186] Step 9: Handle all anomalies during the data election process according to the rules above.
[0187] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A pre-cluster data election method based on a stacked sliding time window, characterized in that, The method includes: S1. Based on the physical connection between the front-end server and the gateway, configure the corresponding communication parameters and collect data from all channels on the same cross-section through the front-end server; S2. Use the stacked sliding time window technology to align the data collected from each channel, analyze the data deviation and generate an initial deviation threshold, and synchronize the initialization progress to all front-end servers. S3. Each front-end server performs a single-machine election based on its local channel data, determines the target channel data, and uploads it to the remaining front-end servers in the front-end cluster in a preset format. A multi-machine comparison mechanism is then executed to generate a comparison sequence. S4. Based on the generated comparison sequence, combined with the election result deviation and the local device management authority, the front-end server executes the corresponding data processing strategy.
2. The method for selecting front-end cluster data based on a stacked sliding time window according to claim 1, characterized in that, The process of aligning data collected from each channel using a stacked sliding time window technique, analyzing data deviations and generating an initial deviation threshold, and synchronously initializing the progress to all front-end servers includes: S21. In the same acquisition section, when the number of target data is not less than the preset channel number threshold, the data acquired by each channel is aligned using the stacked sliding time window technology. S22. Analyze the dispersion of data between channels and generate the initial deviation threshold for the current section by calculating the data deviation; S23. The initial deviation threshold is processed using a steady-state threshold dynamic adjustment algorithm, and the processed data is uploaded to the data acquisition and monitoring control system, and the progress is synchronously initialized to all front-end servers.
3. The method for selecting front-end cluster data based on a stacked sliding time window according to claim 2, characterized in that, In the same acquisition section, when the number of target data is not less than a preset channel number threshold, the alignment processing of the data acquired by each channel using the stacked sliding time window technique includes: S211. Based on the time difference of data transmission from each channel and the real-time requirements of the data acquisition and monitoring control system, set the initial length of the sliding time window. S212. When the target channel is refreshed, the time of data transmission is used as the timing origin. Combined with the set initial length and preset step size, the data of the remaining channels within the corresponding time range is searched and matched. S213. Based on the search and matching results, dynamically adjust the initial length of the sliding time window, generate an optimized time window configuration, and perform alignment processing on the data of each channel in the current section according to the time window configuration.
4. The method for selecting front-end cluster data based on a stacked sliding time window according to claim 3, characterized in that, The search matching results include: If the time for a successful data match is greater than the first target threshold, then the length of the current time window is maintained. If the time for a successful data match is less than the second target threshold, the time window will be reduced to the third target threshold in the next cycle. If there is no matching data in the current time window, the target channel data will be used as the local matching result and sent to the remaining front-end servers for judgment, and the matching time window length will be increased in the next cycle.
5. The method for selecting front-end cluster data based on a stacked sliding time window according to claim 4, characterized in that, The analysis of the dispersion of data between channels, and the generation of the initial deviation threshold for the current section by calculating the data deviation, includes: S221. Based on the aligned channel data, extract the target data points within the same time period, and calculate the average value of the target data points as the reference data for the corresponding acquisition section. S222. Using the reference data as a reference, select the data point closest to the reference data from all channels as the selected point, and calculate the absolute deviation between the remaining data points and the selected point. S223. Sum the absolute deviation values of all channels, obtain the average value, and get the initial deviation threshold within the current time window.
6. The method for selecting a front-end cluster data based on a stacked sliding time window according to claim 1, characterized in that, Each front-end server performs a single-machine election based on its local channel data to determine the target channel data, and uploads it to the remaining front-end servers in the front-end cluster according to a preset format. A multi-machine comparison mechanism is then executed to generate a comparison sequence, including: S31. Based on the initialization progress synchronization results, each front-end server performs a single-machine election on the channel data connected to its local machine to obtain the target channel data; S32. The acquired target channel data is sent to the remaining front-end servers in the front-end cluster according to the preset format, and then enters the delay mode. S33. When all front-end servers in the cluster complete data reception confirmation or reach the preset synchronization timeout threshold, execute the multi-machine comparison mechanism to generate a comparison sequence.
7. The method for selecting a front-end cluster data based on a stacked sliding time window according to claim 6, characterized in that, The delay modes include: If the processing results of the remaining front-end servers are received within the preset waiting time window, the processing results of multiple front-end servers will be analyzed. If no processing result is received from the remaining front-end servers within the preset waiting time window, the waiting time window will be extended in the next processing cycle.
8. The method for selecting front-end cluster data based on a stacked sliding time window according to claim 7, characterized in that, The analysis of the processing results from multiple front-end servers includes: If the difference between the timestamp of the remaining front-end server data received and the time base of the local data is within the preset consistency threshold, the data reporting mechanism is triggered, and the local machine or the corresponding front-end server is notified to send the data. If the difference between the timestamp of the remaining front-end server data and the time base of the local data exceeds the preset consistency threshold, the current reporting operation is stopped, and a sliding time window waiting process is entered. After all target channel data is received, a multi-front-end server data comparison is performed, and a judgment is made based on the comparison results.
9. The method for selecting front-end cluster data based on a stacked sliding time window according to claim 8, characterized in that, The extended waiting time window includes: Determine whether the extended waiting time window length exceeds the difference between the arrival time of the remaining front-end server data and the local data time; if it exceeds, shorten the waiting time window for the next cycle according to the preset configuration; if it does not exceed, extend the waiting time window by the preset step size.
10. The method for selecting a pre-cluster data based on a stacked sliding time window according to claim 1, characterized in that, The data processing strategy includes: If the election result deviation is within the preset threshold range and the local machine has device management authority, the data will be sent to the data acquisition and monitoring control system and the remaining front-end servers will be notified; if the election result deviation is within the preset threshold range and the local machine does not have device management authority, the data confirmation will be sent to the remaining front-end servers and the front-end servers with management authority will execute the data upload. If the election result deviates beyond a preset threshold, a negative confirmation is sent, and the front-end server with jurisdiction will send the data back and mark the data as abnormal.
Citation Information
Patent Citations
A Redis-based dual-plane data acquisition method and system
CN115396752B
Information interaction system and method for power grid regulation and control master station and substation
CN115622251A