An edge-computing-based power distribution energy consumption data preprocessing method
By acquiring current waveform data at edge nodes and utilizing dynamic window sliding segmentation and synchronous offset verification, the real load signal can be accurately identified, solving the problem of misjudgment of interference signals in edge-side data preprocessing. This enables efficient load characteristic analysis and fault diagnosis, while reducing communication latency and costs.
Patent Information
- Application Number
- CN202511913422.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-18
- Publication Date
- 2026-03-17
- Estimated Expiration
- 2045-12-18
AI Technical Summary
Existing edge-side data preprocessing schemes have significant limitations in accurately distinguishing between real load characteristics and transient interference signals. Especially in power distribution monitoring scenarios with complex electromagnetic environments, it is difficult to effectively distinguish between real load signals and interference signals, leading to misjudgments and data transmission delays, and failing to meet the real-time intelligent decision-making requirements of power distribution systems.
By acquiring current waveform data at edge nodes, and utilizing dynamic window sliding segmentation and adaptive anchoring of local amplitude fluctuation rate, the continuity deviation of pulse propagation path is tracked, interference traces are isolated, and combined with synchronous offset verification of adjacent segments, the true load pulsation subsequence is identified, and an edge-cleaned waveform stream is constructed.
It improves the accuracy of load forecasting and fault diagnosis under edge node conditions, reduces data retransmission and cloud processing delays caused by misjudgments, reduces communication costs, and meets the real-time and autonomous requirements of power distribution systems.
Smart Images

Figure CN121350422B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and more specifically, to a method for preprocessing power distribution energy consumption data based on edge computing. Background Technology
[0002] With the in-depth development of smart grid and energy internet technologies, the refined management and intelligent operation and maintenance of power distribution systems have become the core support for achieving dual-carbon goals and ensuring power supply security. In the scenario of power distribution network energy consumption monitoring and optimization, real-time collection and analysis of energy consumption data from massive electrical equipment is of key significance for load forecasting, fault diagnosis, demand response, and energy efficiency assessment. In order to reduce data transmission costs and meet the requirements of millisecond-level response, edge computing architecture has been widely deployed in power distribution energy consumption monitoring systems in recent years, effectively alleviating the bandwidth pressure and latency bottleneck of cloud computing centers and providing a technical foundation for real-time intelligent decision-making in power distribution systems.
[0003] Existing edge-side data preprocessing schemes have significant limitations in accurately distinguishing between real load characteristics and transient interference signals. This problem is particularly prominent in power distribution monitoring scenarios with complex electromagnetic environments. Specifically, the raw energy consumption data collected by power distribution systems generally contains various transient interference components such as motor start-up and shutdown impacts, switching operation transients, and electromagnetic coupling crosstalk. These interference signals are highly similar to real load changes and power fluctuations in both time-domain waveforms and frequency-domain distributions, with blurred boundaries and intertwined characteristics in the feature space. For example, in power distribution monitoring in industrial parks, the current step characteristics generated during the normal startup of high-power equipment significantly overlap with the spike pulses caused by line faults or electromagnetic interference in terms of amplitude change rate, duration, and spectral composition. This similarity makes it difficult to establish clear discrimination criteria in the preprocessing stage. This difficulty in identification not only leads to the effective load signal being misjudged as interference and filtered out, but also causes interference components to be retained as real data and transmitted to subsequent analysis stages, resulting in erroneous conclusions in equipment anomaly detection and load characteristic analysis. The current mainstream approach to solving this problem is to deploy complex deep learning classification models using high-performance edge servers, or to upload the raw data completely to the cloud and use centralized computing power for fine-grained processing. However, the former requires a significant increase in the hardware configuration and power consumption budget of edge nodes, making the deployment cost for massive distributed monitoring points prohibitive, and the model inference latency may still exceed real-time requirements. The latter, on the other hand, violates the original design intent of edge computing, not only consuming a large amount of communication bandwidth resources, but also causing data backlog and loss of timeliness during network fluctuations, failing to meet the core requirements of power distribution systems for local autonomy and rapid response, ultimately restricting the deep application and large-scale promotion of edge computing architecture in the field of smart power distribution.
[0004] In view of this, the present invention proposes a power distribution energy consumption data preprocessing method based on edge computing to solve the above problems. Summary of the Invention
[0005] To overcome the aforementioned deficiencies of the prior art and to achieve the above objectives, the present invention provides the following technical solution: a power distribution energy consumption data preprocessing method based on edge computing, comprising:
[0006] Step S1: Collect real-time current waveform data of the power distribution line through edge nodes, and extract the time series label and amplitude peak value of each sampling point to obtain the original waveform attribute set;
[0007] Step S2: Perform dynamic window sliding segmentation on the original waveform attribute set, where the window boundary is adaptively anchored by the local amplitude fluctuation rate to obtain a transient sensitive segment sequence;
[0008] Step S3: For each segment in the transient sensitive segment sequence, track the continuity deviation of its internal pulse propagation path and isolate the discontinuous pulse subset to obtain the interference trace isolation set;
[0009] Step S4: Based on the interference trace isolation set and the remaining part of the transient sensitive segment sequence, perform synchronization offset verification between adjacent segments, identify and mark the real load pulsation subsequence;
[0010] Step S5: Separate and reassemble the real load pulsation subsequence with the interference trace isolation set to construct the edge-cleaned waveform stream, and record the purity imprint of each subsequence to obtain the preprocessed data stream set.
[0011] The technical effects and advantages of the edge computing-based power distribution energy consumption data preprocessing method of this invention are as follows:
[0012] This invention obtains a transient-sensitive segment sequence by adaptively anchoring the dynamic window boundary using local amplitude fluctuation rate, accurately locating the time domain intervals requiring key identification. The pulse propagation path of the real load signal exhibits temporal continuity and a gradual increase in peak amplitude, while transient interference signals, due to electromagnetic coupling crosstalk and the randomness of switching transients, exhibit path discontinuities and time-sequence jumps. The continuity deviation of the pulse propagation path in the transient-sensitive segment reflects the degree of isolation of the interference signal. Combined with the isolation of discontinuous pulse subsets, an interference trace isolation set is obtained to reduce the dependence of real load feature analysis on transient interference signals. Based on the remaining portion of the transient-sensitive segment sequence between adjacent segments... Synchronous offset verification results, as well as the amplitude gradient smoothness of the remaining part of the transient sensitive segment sequence, adaptively identify the real load pulsation subsequence. This avoids the misjudgment of large-scale load periodic fluctuations that can easily occur when analyzing only the continuity characteristics of a single segment. It also reduces the probability of transient interference signals being misidentified as real load characteristics. As a result, the edge-cleaned waveform stream and its purity imprint constructed based on separation and recombination accurately reflect the real energy consumption characteristics of the power distribution system. This improves the accuracy of load forecasting, fault diagnosis, and energy efficiency assessment based on the pre-processed data stream set under the limited computing power of edge nodes. At the same time, it reduces the communication costs and latency overhead of data retransmission and secondary processing in the cloud due to misjudgment. Attached Figure Description
[0013] Figure 1 This is a schematic diagram of a power distribution energy consumption data preprocessing method based on edge computing according to the present invention;
[0014] Figure 2 This is a schematic diagram of a power distribution energy consumption data preprocessing system based on edge computing according to the present invention. Detailed Implementation
[0015] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0016] Example 1
[0017] Please see Figure 1 As shown in this embodiment, a power distribution energy consumption data preprocessing method based on edge computing includes:
[0018] Step S1: Collect real-time current waveform data of the power distribution line through edge nodes, and extract the time label and amplitude peak value of each sampling point to obtain the original waveform attribute set.
[0019] In actual operation, power distribution networks exhibit complex time-varying characteristics in line currents. These include both regular current changes caused by normal load fluctuations and abnormal fluctuations introduced by factors such as line interference, electromagnetic pulses, and switching transients. Traditional centralized data acquisition systems directly upload raw waveform data to cloud servers for processing, resulting in significant data transmission latency, high network bandwidth consumption, and poor real-time performance. Edge computing architecture pushes data preprocessing capabilities down to edge nodes closer to the data source, enabling preliminary cleaning and feature extraction as soon as data is generated. This effectively reduces the amount of invalid data transmitted and improves the real-time response capability of energy consumption monitoring.
[0020] Preferably, in some possible implementations of the embodiments of the present invention, step S1 includes:
[0021] Step S11: Use the current sensor built into the edge node to sample the power distribution line at the millisecond level and capture continuous current waveform data.
[0022] This embodiment uses a Hall effect current sensor as the core acquisition component of the edge node. This sensor features non-contact measurement, fast response speed, and high linearity. The sensor's sampling frequency is set to 2kHz, meaning it acquires an instantaneous current value every 0.5 milliseconds, effectively capturing the current components in the power distribution line. During actual acquisition, the analog current signal output by the sensor is converted into a digital signal by a 16-bit analog-to-digital converter, with a quantization accuracy of 0.01 amperes. The edge node's local buffer stores the most recent 30 seconds of continuous current waveform data in a circular queue, occupying approximately 120KB of memory.
[0023] It should be noted that in other embodiments of the present invention, electromagnetic current transformers, Rogowski coils and other sensor types can also be used. The sampling frequency can be adjusted from 1kHz to 10kHz according to the frequency characteristics of the monitored object. The number of bits of the analog-to-digital converter can be selected from 12 bits to 24 bits. The sampling accuracy and buffer size can be flexibly configured according to the hardware resources of the edge node and the application scenario, and are not limited here.
[0024] Step S12: Mark the acquisition time point of the continuous current waveform data as a time sequence label, calculate the peak amplitude of each point, and construct the original waveform attribute set.
[0025] While continuous current waveform data contains rich time-domain information, the original sampled value sequence lacks structured spatiotemporal correlation identifiers, which is not conducive to subsequent segmentation processing and feature extraction. To establish a spatiotemporal indexing system for waveform data, it is necessary to assign precise time-series labels to each sampling point and extract its local peak features.
[0026] Preferably, in some possible implementations of the embodiments of the present invention, step S12 includes:
[0027] The absolute time position of each sampling point is extracted from the continuous current waveform data to generate a time-series label sequence. Specifically, when the edge node starts the acquisition task, it synchronizes its local real-time clock with the network time protocol server to obtain the current absolute timestamp (with microsecond precision), which is then recorded as the acquisition start time. For the first sampling points ( , The absolute temporal position of a sample point (the total number of sampling points in the buffer) is calculated as follows: In the formula, For the first Time-series labels for each sampling point; The sampling start time is Δt; Δt is the sampling period, for a sampling frequency of 2kHz. It is 500 microseconds.
[0028] For each position in the time-series label sequence, the extreme values of the amplitude within the microsecond interval before and after it are scanned to determine the peak amplitude. The current waveform of a power distribution line exhibits a sawtooth pattern at a microscale, and the instantaneous value of a single sampling point may be affected by quantization noise and high-frequency interference, failing to accurately represent the true current level at that moment. To improve the robustness of the amplitude characteristics, this embodiment employs a local peak extraction strategy: for the... A sampling point is defined, and a time window centered on it is defined. ,in The window width is half a window, and its value is set to two sampling periods, or 1000 microseconds. The maximum current amplitude is searched within this window and denoted as the [value]. Peak amplitude of each sampling point ;
[0029] The time-series label sequences are paired one-to-one with peak amplitudes and fused into a structured set of original waveform attributes. Specifically:
[0030] The original waveform attribute set is as follows: ;
[0031] Each element This is called a waveform attribute point, which fully describes the temporal position and amplitude characteristics of the i-th sampling point. This set is stored in the memory of the edge node in the form of a two-dimensional table, which facilitates subsequent streaming processing in temporal order.
[0032] It should be noted that half the window width The value of needs to be determined comprehensively based on the sampling frequency and line characteristics, and is generally set to 2 to 5 sampling periods. In other embodiments, median filtering, mean filtering and other methods can be used to smooth the instantaneous current value and then directly use it as the amplitude feature, or derivative indicators such as effective value (RMS) and crest factor can be extracted, which are not limited here.
[0033] Step S2: Perform dynamic window sliding segmentation on the original waveform attribute set, where the window boundary is adaptively anchored by the local amplitude fluctuation rate to obtain a transient sensitive segment sequence.
[0034] The current waveform of power distribution lines exhibits typical non-stationary characteristics. Events such as load switching, fault transients, and electromagnetic interference introduce localized, severe fluctuations into the waveform. These fluctuating segments carry crucial abnormal information and need to be separated from the stable background for focused analysis. Traditional fixed-window segmentation methods uniformly divide the waveform according to a preset time length, which cannot adapt to the time-varying nature of fluctuation characteristics. They easily fragment the waveform of the same event into different segments or mix unrelated stable and fluctuating segments together, reducing the targeted nature of subsequent processing.
[0035] The dynamic window sliding segmentation method adaptively adjusts the segmentation boundary based on the local characteristics of the waveform, enabling precise location of transient regions with drastic fluctuations. Local amplitude volatility reflects the degree of drastic change in the peak amplitude of the waveform within a short time window. When the volatility exceeds a preset threshold, it indicates that the location may correspond to the start or end of a load surge or disturbance event, making it suitable as a segmentation boundary. Through this adaptive anchoring mechanism, the original waveform attribute set can be divided into several transient-sensitive segments, each corresponding to a relatively independent current fluctuation process.
[0036] Preferably, in some possible implementations of the embodiments of the present invention, step S2 includes:
[0037] Step S21: Initialize the starting width of the sliding window to a preset fixed millisecond pulse length.
[0038] Transient events commonly encountered in power distribution lines (such as switch operations, capacitor switching, and motor starting) typically last between 10 and 100 milliseconds, while interference pulses (such as lightning induction and arc discharge) usually last between 1 and 10 milliseconds. To balance the ability to capture events at different time scales, this embodiment sets the initial width of the sliding window. It is 20 milliseconds, corresponding to 40 sampling points at a sampling frequency of 2kHz.
[0039] During initialization, the left boundary of the sliding window is located at the first waveform attribute point in the original waveform attribute set, and the right boundary is located at the 40th waveform attribute point. The window contains 40 waveform attribute points.
[0040] It should be noted that the starting width The time scale can be adjusted according to the typical events of the monitored object. For high-voltage transmission lines, it can be set to 50 to 200 milliseconds, and for low-voltage user side, it can be set to 5 to 20 milliseconds. In other embodiments of the present invention, the starting width can also be set to a fixed number of sampling points rather than a fixed time length, which is not limited here.
[0041] Step S22: Advance the window point by point along the original waveform attribute set, and calculate the rolling variance of the peak amplitude within the window as the local amplitude volatility.
[0042] The sliding window advances point by point along the time label direction, with each sampling point as a step size. At the... After the next advancement, the left boundary of the window is located at the [number]th [position]. The right boundary is located at the waveform attribute point. The window contains the waveform attribute point. To the There are a total of 40 waveform attribute points. For the... For each window location, extract all peak amplitudes within the window and calculate the variance of this data set as the local amplitude volatility.
[0043] Variance measures the degree of dispersion of data from the mean. If the peak amplitude changes drastically within a window, the variance value is large; if the peak amplitude is relatively stable within a window, the variance value is small. By calculating the variance at each window position in a rolling manner, a local amplitude volatility sequence is obtained, which characterizes the distribution of the volatility intensity of the original waveform attribute set along the time axis.
[0044] Step S23: When the local amplitude volatility exceeds the preset anchoring threshold, anchor the end of the current window as the segmentation boundary and generate a transient sensitive segment.
[0045] Preset anchoring threshold This is used to determine whether the local amplitude fluctuation rate reaches the significance level of a transient event. This embodiment is set based on the statistical characteristics of historical operating data. The current waveform data of the power distribution line under normal operating conditions for 7 consecutive days were selected, and the 95th percentile of the global local amplitude fluctuation rate was calculated as the anchoring threshold.
[0046] During the sliding process, the local amplitude volatility of the current window is compared with the anchoring threshold in real time: if the local amplitude volatility is less than or equal to the anchoring threshold, the current window is considered to be in a stable segment, and the window continues to move forward; if the local amplitude volatility is greater than the anchoring threshold, the current window is considered to have entered a transient volatility region, and the right boundary of the anchoring window (i.e., the first...) is... Timing label of each waveform attribute point () represents the dividing boundary.
[0047] After anchoring the segmentation boundary, all waveform attribute points from the previous segmentation boundary (or the starting point of the set) to the current segmentation boundary constitute a transient sensitive segment. For example, if the first segmentation boundary is anchored at... (Sampling point 150), the second segmentation boundary is anchored at... (At the 380th sampling point), two transient sensitive segments are generated:
[0048] Paragraph 1: Contains waveform attribute points The corresponding time span is 0 to 74.5 milliseconds, with a total of 150 sampling points;
[0049] Paragraph 2: Contains waveform attribute points The corresponding time span is 75.0 to 189.5 milliseconds, with a total of 230 sampling points.
[0050] It should be noted that the setting of the anchoring threshold should take into account factors such as line voltage level, load type, and environmental interference level. The threshold can be appropriately increased for high-voltage lines and areas with concentrated industrial loads, and appropriately decreased for low-voltage residential power consumption areas. In other embodiments of the present invention, a dynamic threshold strategy can also be adopted, which adjusts the threshold in real time according to the volatility statistical characteristics of the previous time window, and is not limited here.
[0051] Step S24: Repeat steps S22 to S23 until the entire original waveform attribute set is covered to obtain the transient sensitive segment sequence.
[0052] The sliding window starts from the first waveform attribute point and progresses point by point to the (N-39)th waveform attribute point (N is the total number of sampling points in the original waveform attribute set), completely scanning the entire waveform data. Whenever a local amplitude fluctuation rate exceeds the anchoring threshold, a new segmentation boundary and transient-sensitive segment are generated. Finally, M transient-sensitive segments are obtained (the value of M depends on the number of transient events in the waveform), denoted as the transient-sensitive segment sequence. Each segment contains a start time series label, an end time series label, and the corresponding set of waveform attribute points.
[0053] For example, a 30-second waveform data set containing 60,000 sampling points may result in 15 transient-sensitive segments after dynamic window sliding segmentation. Each segment's duration ranges from 10 to 200 milliseconds, covering all significant fluctuation regions. The volatility of the stable operating region remains below a threshold and is not segmented separately, but rather serves as a transition region between adjacent segments.
[0054] It should be noted that in other embodiments of the present invention, the starting boundary of a segment can be anchored when the rising edge of volatility is detected (from below the threshold to above the threshold), and the ending boundary of a segment can be anchored when the falling edge of volatility is detected (from above the threshold to below the threshold), so that each segment is more accurately aligned with the actual duration of the transient event; a hysteresis comparison mechanism can also be introduced, setting rising and falling thresholds (the falling threshold is lower than the rising threshold), to avoid excessive segmentation caused by repeated jumps in volatility near the threshold, which will not be elaborated here.
[0055] Step S3: For each segment in the transient sensitive segment sequence, track the continuity deviation of its internal pulse propagation path and isolate the discontinuous pulse subset to obtain the interference trace isolation set.
[0056] Although the transient sensitive segment locates the time region of current fluctuation, the segment may contain a mixture of fluctuation components of various natures: one type is a continuous pulse sequence caused by actual load changes, whose peak amplitude shows a smooth gradual change or periodic regularity along the time axis, and the time interval is uniform; the other type is an isolated pulse jump caused by external interference or measurement anomalies, whose peak amplitude suddenly appears and disappears quickly, lacks continuity with the preceding and following waveforms, and the time interval is irregular.
[0057] The physical propagation of current in a power distribution line follows the principle of continuity: the establishment and decay of load current are constrained by parameters such as line inductance and load impedance, and there is an upper limit to the rate of change. There is a smooth causal relationship between current values at adjacent moments. However, interference signals (such as electromagnetic pulses, sensor glitches, and data packet loss and retransmission) are not subject to these physical constraints and can produce large jumps in a very short time, which are manifested as discontinuous isolated peaks in the waveform.
[0058] By tracking the temporal continuity of peak amplitudes within transient sensitive segments, these discontinuous pulses can be identified and isolated, classified as interference traces, and prevented from contaminating subsequent load characteristic analysis.
[0059] Preferably, in some possible implementations of the embodiments of the present invention, step S3 includes:
[0060] Step S31: For each segment in the transient-sensitive segment sequence, concatenate the peak amplitudes sequentially along the time label to obtain the simulated pulse propagation path.
[0061] For the first transient-sensitive segment sequence Each segment is processed, and all waveform attribute points contained therein are extracted. The peak amplitudes of all waveform attribute points are arranged in chronological order according to their time labels to form a peak amplitude sequence. This sequence forms a broken line on the time-amplitude two-dimensional plane, simulating the propagation trajectory of the current pulse within the time range of that segment, and is called the simulated pulse propagation path.
[0062] For example, a certain segment contains 50 sampling points, corresponding to a time span of 20 to 44.5 milliseconds. The peak amplitude gradually increases from 12.0 amperes to 18.0 amperes and then falls back to 10.0 amperes. The simulated pulse propagation path shows an envelope shape that first rises and then falls, reflecting the process of load input and stabilization.
[0063] Step S32: Check the timing gap between adjacent peak amplitudes in the simulated pulse propagation path. If the gap exceeds the preset continuity threshold, mark it as a deviation point.
[0064] In the pulse propagation path of the actual load current, the time interval between adjacent sampling points should be equal to the sampling period. (For 2kHz sampling, Δt = 0.5 milliseconds), allowing for minor clock jitter (typically no more than ±10 microseconds). If the time interval between adjacent sampling points deviates significantly from the sampling period, it indicates that data loss, transmission delay, or clock anomaly may have occurred at that location, disrupting the timing continuity of the pulse propagation path.
[0065] Preferably, in some possible implementations of the embodiments of the present invention, step S32 includes:
[0066] Select each pair of adjacent peak amplitudes along the simulated pulse propagation path, calculate their timing tag difference as the timing interval, and for each segment... Inner The and the first Waveform attribute points , Represents the number of attribute points within a paragraph, used to calculate time intervals. : In the formula, For paragraph Inner The time interval between adjacent peak amplitudes; and The first The and the first The timing labels for each waveform attribute point are then compared one by one with a continuity threshold preset based on line impedance. The continuity threshold reflects the maximum allowable time interruption of current changes in the distribution line. In this embodiment, the continuity threshold is set based on the line's impedance characteristics and load response speed: for industrial distribution lines with a high proportion of resistive-inductive loads, the time constant of current changes is generally in the range of 1 to 5 milliseconds, so the continuity threshold is set to 1.5 milliseconds, meaning the timing gap is allowed to deviate by no more than ±1000 microseconds (corresponding to ±2 sampling periods) from the 500 microsecond sampling period; for residential distribution lines with predominantly resistive loads, the time constant is smaller, so the continuity threshold can be set to 1.0 millisecond. If the timing gap... satisfy: ,in, If a preset continuity threshold is used, an anomaly in the temporal continuity between adjacent peak amplitudes is determined, and the first peak is immediately marked on the simulated pulse propagation path. Each waveform attribute point is a deviation point.
[0067] For example, if the continuity threshold The sampling period is 1500 microseconds. If it is 500 microseconds, then the timing gap is allowed to be... to That is, within the range of (-500 microseconds, 1500 microseconds), a negative time interval is unreasonable in practice, so the effective range is (0 microseconds, 1500 microseconds). If the time labels of a pair of adjacent sampling points are 10000 microseconds and 12000 microseconds respectively, the time interval is 2000 microseconds > 1500 microseconds, which exceeds the continuity threshold and is marked as a deviation point.
[0068] Step S33: Expand the micro-window around each deviation point and extract the amplitude of its isolated peaks as a subset of discontinuous pulses.
[0069] Identifying deviation points pinpoints the locations of temporal continuity anomalies, but a single deviation point may only serve as a boundary marker for transient interference, as the interference signal itself may span multiple consecutive sampling points. To fully extract the interference waveform segment, a micro-window needs to be extended around the deviation point to capture all peak amplitudes within a certain range before and after the deviation point.
[0070] For each waveform attribute point marked as a deviation point within a paragraph, define a micro-window centered on that point. ,in This is half the width of a micro-window. This embodiment sets... With a duration of 5 milliseconds, corresponding to ±10 sampling points, it can cover the duration of most short-term interference pulses.
[0071] Within the micro-window, extract all waveform attribute points whose timing labels fall into the micro-window to form a subset of discontinuous pulses with the deviation point as the core.
[0072] For example, the 120th sampling point (time sequence label) within a certain paragraph A deviation point (59,500 microseconds) is marked, with a micro-window range of [54,500 microseconds, 64,500 microseconds], corresponding to sampling points 110 to 130. These 21 sampling points and their peak amplitudes constitute a discontinuous pulse subset. If the peak amplitude within this subset exhibits a sharp jump (e.g., a sudden change from 10.0 amperes to 28.0 amperes and then falling back to 9.5 amperes), lasting approximately 10 milliseconds, it conforms to the typical characteristics of an electromagnetic interference pulse.
[0073] It should be noted that the half-width of the micro-window is determined based on the typical duration of the interference pulse, and can be set to 1 to 3 milliseconds for high-frequency electromagnetic interference. If the micro-windows of multiple deviation points overlap, they can be merged to remove duplicates and avoid the same interference segment being extracted repeatedly. In other embodiments of the present invention, an adaptive window width can also be used to dynamically adjust the size of the micro-window according to the jump amplitude of the peak amplitude before and after the deviation point, which is not limited here.
[0074] Step S34: Remove the discontinuous pulse subset from the transient sensitive segment sequence to obtain the interference trace isolation set.
[0075] The discontinuous pulse subsets identified in all transient-sensitive segments are aggregated to form an interference trace isolation set. Simultaneously, waveform attribute points contained in all discontinuous pulse subsets are deleted from the original transient-sensitive segment sequence, retaining the remaining portions as the continuous pulse part of the segment for subsequent real load characteristic analysis. For example, if segment 1 originally had 200 sampling points, after removing two discontinuous pulse subsets (a total of 42 sampling points), the remaining 158 sampling points constitute the continuous pulse part of segment 1.
[0076] For example, for a 30-second waveform data containing 60,000 sampling points, after pulse continuity tracking, 25 discontinuous pulse subsets may be identified, involving approximately 600 sampling points (accounting for 1% of the total). These isolated pulses are isolated to the interference trace isolation set and do not participate in subsequent load feature extraction and energy consumption calculation, effectively avoiding the impact of interference signals on the accuracy of energy consumption statistics.
[0077] Step S4: Based on the interference trace isolation set and the remaining part of the transient sensitive segment sequence, perform synchronous offset verification between adjacent segments, identify and mark the real load pulsation subsequence.
[0078] After processing in step S3, the transient sensitive segment sequence has eliminated discontinuous interference pulses, but the temporal correlation between segments has not been fully utilized. The actual load changes of power distribution lines often exhibit periodic or quasi-periodic characteristics. For example, industrial production equipment starts and stops at fixed intervals, and air conditioning compressors operate according to temperature feedback cycles. These load fluctuations are manifested in the waveform as a temporal synchronization relationship between multiple segments: the start time, duration, or peak amplitude of adjacent segments show regular intervals.
[0079] In contrast, random interference (such as lightning strikes, radio frequency interference, etc.) does not exhibit this kind of temporal synchronicity in its occurrence time and intensity, and the time offset between adjacent segments is randomly distributed. By checking whether the time offset between adjacent segments falls within the load cycle synchronization zone, it is possible to further distinguish between real load pulsations and residual interference signals, thereby improving the accuracy of data preprocessing.
[0080] Meanwhile, the actual load current variation is constrained by physical inertia, and the gradual change of peak amplitude along the time axis should present a smooth and continuous envelope, while the amplitude change of interference signals often manifests as abrupt changes and oscillations. By evaluating the smoothness of the gradual change of peak amplitude within a segment, the authenticity of the segment can be verified from the amplitude characteristic dimension.
[0081] Preferably, in some possible implementations of the embodiments of the present invention, step S4 includes:
[0082] Step S41: Select the remaining part of every two adjacent segments in the transient sensitive segment sequence, and extract the peak amplitude of the end and the beginning of the segment as anchor points.
[0083] For the transient sensitive segment sequence, the first The paragraph and the first Each segment, after removing the discontinuous pulse subset, is denoted as a segment. The rest of the paragraphs The remaining part.
[0084] Extract paragraphs The last waveform attribute point of the remaining part is used as the tail anchor point to extract the paragraph. The first waveform attribute point of the remaining portion serves as the beginning anchor point. The anchor point marks the boundary between adjacent segments on the time axis, and its peak amplitude reflects the current level at the segment boundary.
[0085] Step S42: Calculate the time tag offset between anchor points and verify whether the offset conforms to the preset load cycle synchronization band.
[0086] The time tag offset between anchor points is defined as the difference between the time tags of the first anchor points of two adjacent paragraphs. The load cycle synchronization band reflects the time cycle range of typical periodic loads of the power distribution line. This embodiment sets the synchronization band based on the load composition characteristics of the monitored line: For commercial power distribution lines containing a large amount of air conditioning load, the start-stop cycle of air conditioning compressors is generally 5 to 15 minutes, and the load cycle synchronization band is set to [280,000 microseconds, 920,000 microseconds], allowing a cycle fluctuation of ±20,000 microseconds; For industrial production lines, the equipment cycle may be 10 seconds to 5 minutes, and the synchronization band range needs to be adjusted according to the actual process parameters.
[0087] Verify offset Does it fall within the load cycle synchronization zone? :like Then determine the paragraph. and paragraphs It has periodic synchronization and may correspond to repetitive pulsations of the same type of load;
[0088] like or If the two segments do not have periodic synchronization, they may belong to different types of loads or random interference.
[0089] It should be noted that the setting of the load cycle synchronization band needs to be based on the load type statistics and historical operation data of the line. For mixed load lines, multiple synchronization bands can be set to correspond to loads of different cycles. In other embodiments of the present invention, autocorrelation analysis, spectrum analysis and other methods can also be used to automatically extract the dominant cycle from historical waveform data and dynamically update the synchronization band range, which is not limited here.
[0090] Step S43: If the offset falls into the synchronization band, then merge the remaining parts of the two adjacent segments into a continuous pulse unit.
[0091] If the paragraph is determined and paragraphs It features periodic synchronization, connecting the remaining parts of two segments end-to-end on the timeline to merge them into a single continuous pulse unit. The merged pulse unit contains the segments. and paragraphs All waveform attribute points for the remaining portion. If the offsets of multiple consecutive segments all fall within the synchronization band, they are sequentially merged into a longer continuous pulse unit. For example, if the offsets of segments 1, 2, and 3 fall within the synchronization band for each pair, the three segments are merged into one continuous pulse unit, covering the complete time span from the beginning of segment 1 to the end of segment 3.
[0092] If the offset does not fall within the synchronization band, it is determined that the two segments do not have periodic synchronization, and the two segments remain independent and are not merged. After pairwise verification and merging of all adjacent segments, the transient sensitive segment sequence is reorganized into several continuous pulsating units, each unit corresponding to a possible periodic load pulsation process.
[0093] For example, there were originally 15 transient sensitive segments. After synchronization offset verification, segments {1, 2, 3} were merged into unit 1, segments {5, 6} were merged into unit 2, segments {8, 9, 10, 11} were merged into unit 3, and the remaining segments {4, 7, 12, 13, 14, 15} remained independent because their offsets did not conform to the synchronization band. Finally, 10 continuous pulsating units were obtained.
[0094] Step S44: For each continuous pulsating unit, identify its amplitude gradient smoothness. If the amplitude gradient smoothness is higher than the preset interference threshold, mark it as a real load pulsating subsequence.
[0095] Preferably, in some possible implementations of the embodiments of the present invention, step S44 includes:
[0096] For each continuous pulsation unit, the cumulative coefficient of variation of the difference between adjacent amplitudes is calculated along its peak amplitude sequence, serving as the amplitude gradient smoothness. Specifically, for units containing... Extracting peak amplitude sequence from continuous pulsating units of waveform attribute points. Calculate the adjacent amplitude difference sequence ,in .
[0097] The adjacent amplitude difference reflects the point-to-point variation of the peak amplitude. The actual load current is constrained by line impedance and load time constant, and the amplitude change rate has a physical upper limit; therefore, the adjacent amplitude difference should be within a reasonable range. However, the amplitude of interference signals can change instantaneously, and the adjacent amplitude difference may exhibit abnormally large values. The coefficient of variation (COP) of the adjacent amplitude difference sequence is calculated and defined as the ratio of the standard deviation to the mean of the adjacent amplitude difference sequence. The COP, being the ratio of the standard deviation to the mean, eliminates the influence of dimensions and reflects the relative dispersion of the data. If the adjacent amplitude difference changes smoothly, the COP is small; if there are abrupt changes in the adjacent amplitude difference, the COP is large.
[0098] The smoothness of the amplitude gradient is compared with a preset interference threshold. This embodiment sets the interference threshold based on the variation pattern of the actual load current: for normal resistive-inductive loads, the amplitude change rate during current rise and fall is relatively uniform, and the coefficient of variation of adjacent amplitude differences generally does not exceed 0.3; for interference signals, instantaneous jumps cause extreme values in the difference between adjacent amplitudes, and the coefficient of variation is usually greater than 0.8. This embodiment sets the preset interference threshold to 0.5 as the dividing line between the actual load and interference signals.
[0099] If the amplitude gradient smoothness is less than the preset interference threshold, the continuous pulsating unit is immediately marked as a real load pulsating subsequence, and its peak amplitude change is considered to conform to physical laws and belong to the real load current waveform; if the amplitude gradient smoothness is greater than or equal to the preset interference threshold, the unit is determined to contain residual interference or measurement abnormality and is not marked as a real load pulsating subsequence.
[0100] After amplitude gradient smoothness screening, the set of continuous pulsating units is divided into two parts: units marked as true load pulsating subsequences will be used for subsequent energy consumption calculations and load characteristic analysis; unmarked units, together with the interference trace isolation set, are classified as anomalous data and will not participate in energy consumption statistics.
[0101] It should be noted that the setting of the interference threshold needs to be combined with the load type and line parameters. For drastically changing impact loads (such as welding machines and elevators), the threshold can be appropriately increased to 0.6 to 0.8, while for stable loads (such as lighting and heating), the threshold can be decreased to 0.3 to 0.4. In other embodiments of the present invention, other forms of smoothness index can also be used, such as the ratio of the range to the mean of adjacent amplitude differences, which are not limited here.
[0102] Step S5: Separate and reassemble the real load pulsation subsequence with the interference trace isolation set to construct the edge-cleaned waveform stream, and record the purity imprint of each subsequence to obtain the preprocessed data stream set.
[0103] After the aforementioned multi-level screening steps, the original waveform attribute set has been divided into two categories: the real load pulsation subsequence represents the actual power consumption characteristics of the distribution line and has high data quality and reliability; the interference trace isolation set contains various abnormal fluctuations and measurement errors, and has lower data reliability. To facilitate subsequent energy consumption analysis and anomaly diagnosis, the two types of data need to be physically separated and structurally reorganized, and quality evaluation indicators need to be added to each subsequence to form a complete preprocessed data stream set.
[0104] The edge-cleaned waveform stream, as the main output of the preprocessing, carries the real load current waveform after deep cleaning and can be directly used for applications such as energy consumption calculation, load prediction, and equipment condition assessment. The purity imprint, as a quantitative identifier of data quality, provides a reliability reference for downstream analysis algorithms and supports adaptive weighted processing based on data quality.
[0105] Preferably, in some possible implementations of the embodiments of the present invention, step S5 includes:
[0106] Step S51: Connect the real load pulsation subsequences in series according to the original time sequence label order to generate the main waveform chain.
[0107] All subsequences in the set of real load pulsation subsequences are sorted in ascending order according to the sequence of their first anchor point time labels. The sorted subsequences are then concatenated to form the main waveform chain.
[0108] Preferably, in some possible implementations of the embodiments of the present invention, the generation of the main waveform chain includes:
[0109] For each subsequence in the real load pulsation subsequence, its first and last peak amplitudes are selected as bridging anchor points, and the potential jump risk is estimated based on the amplitude difference between anchor points, generating a bridging anchor point risk set. Specifically, for the... A subsequence, denoted by the amplitude of its first peak. The peak amplitude at the tail end is ;No. The peak amplitude at the beginning of each subsequence is Define bridging transition risk. In the formula, For the first The and the first The risk of bridging transitions between subsequences.
[0110] If bridging jump risk A large value (e.g., exceeding 3.0) indicates a significant difference in current levels between adjacent subsequences, and direct series connection may produce unrealistic amplitude abrupt changes at the bridging point. [All] Bridge points with a risk value greater than 3.0 and their risk values are recorded as the bridge anchor risk set.
[0111] Based on the bridging anchor risk set, miniature gradient transition segments are generated by sequentially bridging the tail and head anchors of adjacent subsequences along the time-series labels. For bridging points with a risk value greater than 3.0, at the... The tail anchor point of each subsequence ( , ) and the The first anchor point of each subsequence ( , Insert a linear gradient transition segment between ) .
[0112] The time span of the gradual transition section is [ , ], inserted uniformly within this time range One virtual sampling point (in this embodiment, we take...) For 5), for the first virtual sampling points ( Time tag and peak amplitude The calculation is as follows: and
[0113] For bridging points with a risk value less than or equal to 3.0, the first... The tail anchor point of the subsequence and the first The first anchor points of each subsequence are connected on the time axis, without inserting virtual sampling points.
[0114] Based on the micro-gradient transition segment, the continuity weights of the micro-gradient transition segment are adjusted node by node to form a preliminary draft of a seamless ring link, which is temporarily fixed in the memory cache of the edge nodes. The continuity weights reflect the confidence level of each waveform attribute point; the weight of the real sampling point is set to 1.0, and the weight of the virtual interpolation point decreases according to its distance from the real anchor point. For the th wavelet in the micro-gradient transition segment... The continuity weight of each virtual sampling point is calculated as follows: In the formula, For the first The continuous weights of each virtual sampling point are distributed in a triangular pattern, with the weights being the largest at the midpoint of the transition section and the smallest at both ends.
[0115] All the real sampling points, virtual interpolation points and their continuity weights of all subsequences are assembled in time-series label order to form a preliminary seamless ring link draft, which is stored in the memory buffer of the edge node.
[0116] Based on the initial draft of a seamless ring link, a lightweight hash check is performed on each serial node to verify the uniqueness of the time sequence label and lock the link integrity level by level, thereby iteratively optimizing and generating the final main waveform chain. The hash check method is as follows: for the time sequence label of each waveform attribute point, a hash value is calculated through hash operation, and the presence of duplicate values in the hash value sequence is checked. If a duplicate hash value is found, it indicates a time sequence label conflict, and the time sequence label of the corresponding node needs to be checked and fine-tuned: an offset of 1 microsecond is added, and the process is repeated until the hash value is unique.
[0117] After hash verification, a unique link index ID is assigned to each node of the main waveform chain, and a bidirectional mapping table between time series labels and link indices is established to lock the integrity and order of the links. Finally, the main waveform chain is stored as a structured array, with each element containing four fields: {link index ID, time series label, peak amplitude, and continuity weight}.
[0118] It should be noted that the number of virtual sampling points inserted can be adaptively adjusted according to the bridging transition risk value. The greater the risk, the more points are inserted to enhance the smoothing effect. The allocation strategy of the continuity weight can also adopt other forms such as Gaussian distribution and exponential decay. The modulus of hash verification can be adjusted according to the total number of nodes, and is not limited here.
[0119] Step S52: Place the interference trace isolation set in a separate buffer, keeping it physically separated from the main waveform chain.
[0120] The interference trace isolation set data is completely separated from the main waveform chain and stored in an independent buffer at the edge node, allocated a different memory address space, to avoid confusion with the cleaned main waveform chain.
[0121] The interference trace isolation set is organized in list form, with each element recording the {start time label, end time label, peak amplitude sequence, and interference type label} of an interference segment. The interference type label is divided into two categories based on the identified source: temporal discontinuity and amplitude abrupt change.
[0122] Step S53: Based on the peak amplitude density of each real load pulsation subsequence in the main waveform chain, calculate the proportion of non-deviation points within it as a purity imprint.
[0123] Purity imprinting quantifies the data quality of a true load pulsation subsequence, reflecting the proportion of valid sampling points in the subsequence. Peak amplitude density is defined as the number of valid sampling points per unit time in the subsequence, with non-biased points referring to sampling points not marked as biased points.
[0124] For a given real load pulsation subsequence in the main waveform chain, the total number of sampling points (including real sampling points and virtual interpolation points) and the number of sampling points marked as deviation points are counted. The purity imprint of this real load pulsation subsequence is defined as the ratio of the number of sampling points not marked as deviation points to the total number of sampling points. The closer the purity imprint is to 1, the higher the proportion of valid data in the subsequence and the better the data quality.
[0125] Step S54: Add a purity imprint to each real load pulsation subsequence, merge them into an edge-cleaned waveform stream, and obtain a preprocessed data stream set.
[0126] Each real load pulsation subsequence in the main waveform chain is bound to its corresponding purity imprint, forming the basic unit of the edge-cleaned waveform stream. The edge-cleaned waveform stream is organized in the form of a structured data stream, and each unit contains five fields: {subsequence ID, start time series label, end time series label, waveform attribute point sequence, purity imprint}.
[0127] The preprocessed data stream consists of two parts:
[0128] 1. Edge-cleaned waveform stream: Contains all real load pulsation subsequences and their purity imprints, serving as the main output stream for subsequent applications such as energy consumption analysis, load forecasting, and equipment monitoring;
[0129] 2. Interference Trace Isolation Set: Contains all abnormal fluctuation segments and their interference type labels, serving as an auxiliary output stream for diagnostic applications such as data quality auditing and interference source localization.
[0130] The two sets of data are archived separately in the storage system of the edge node. The edge cleansing waveform stream is stored in the high-speed cache area, which supports real-time reading and streaming to the upper layer application; the interference trace isolation set is stored in the low-speed storage area, which can be queried on demand and analyzed offline.
[0131] It should be noted that the calculation method of purity imprint can be adjusted according to application requirements, such as introducing a weighted average with continuity weights, or considering the contribution of amplitude gradient smoothness; the storage format of the edge cleansing waveform stream can be further reduced by compression coding (such as Huffman coding); in other embodiments of the present invention, extended information such as anomaly severity score and possible interference source type inference can also be added to the interference trace isolation set to support more in-depth data quality analysis, which is not limited here.
[0132] This embodiment obtains a transient-sensitive segment sequence by adaptively anchoring the dynamic window boundary using local amplitude fluctuation rate, accurately locating the time domain intervals that need to be identified. The pulse propagation path of the real load signal exhibits temporal continuity and a gradual peak amplitude, while the transient interference signal, due to electromagnetic coupling crosstalk and the randomness of switching operation transients, exhibits path discontinuity and time interval jumps. The continuity deviation of the pulse propagation path of the transient-sensitive segment reflects the degree of isolation of the interference signal. Combined with the isolation of discontinuous pulse subsets, an interference trace isolation set is obtained to reduce the dependence of real load feature analysis on transient interference signals. Based on the remaining part of the transient-sensitive segment sequence between adjacent segments... Synchronous offset verification results, as well as the amplitude gradient smoothness of the remaining part of the transient sensitive segment sequence, adaptively identify the real load pulsation subsequence. This avoids the misjudgment of large-scale load periodic fluctuations that can easily occur when analyzing only the continuity characteristics of a single segment. It also reduces the probability of transient interference signals being misidentified as real load characteristics. As a result, the edge-cleaned waveform stream and its purity imprint constructed based on separation and recombination accurately reflect the real energy consumption characteristics of the power distribution system. This improves the accuracy of load forecasting, fault diagnosis, and energy efficiency assessment based on the pre-processed data stream set under the limited computing power of edge nodes. At the same time, it reduces the communication costs and latency overhead of data retransmission and secondary processing in the cloud due to misjudgment.
[0133] Example 2
[0134] Please see Figure 2 As shown, parts not described in detail in this embodiment are described in Embodiment 1. A power distribution energy consumption data preprocessing system based on edge computing is provided, including:
[0135] Data acquisition module: Collects real-time current waveform data of power distribution lines through edge nodes, and extracts the time label and amplitude peak value of each sampling point to obtain the original waveform attribute set;
[0136] Segmentation module: Performs dynamic window sliding segmentation on the original waveform attribute set, where the window boundary is adaptively anchored by the local amplitude fluctuation rate to obtain a transient sensitive segment sequence;
[0137] Interference filtering module: For each segment in the transient sensitive segment sequence, it tracks the continuity deviation of the internal pulse propagation path and isolates the discontinuous pulse subset to obtain the interference trace isolation set;
[0138] The labeling module performs synchronization offset verification between adjacent segments based on the interference trace isolation set and the remaining part of the transient sensitive segment sequence, and identifies and labels the real load pulsation subsequence;
[0139] Reassembly module: Separates and reassembles the real load pulsation subsequence with the interference trace isolation set, constructs an edge-cleaned waveform stream, and records the purity imprint of each subsequence to obtain a preprocessed data stream set.
[0140] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. An edge computing-based power consumption data preprocessing method, characterized in that, Comprise: Step S1: Collect real-time current waveform data of power distribution lines through edge nodes, and extract time sequence label and amplitude peak value of each sampling point to obtain original waveform attribute set; Step S2: Perform dynamic window sliding segmentation on the original waveform attribute set, wherein the window boundary is anchored by the local amplitude fluctuation rate to obtain a sequence of transient sensitive paragraphs; Step S3: For each paragraph in the sequence of transient sensitive paragraphs, track the continuity deviation of the internal pulse propagation path, and isolate the discontinuous pulse subset to obtain the interference trace isolation set; Step S4: Based on the interference trace isolation set and the remaining part of the sequence of transient sensitive paragraphs, perform synchronization offset checking between adjacent paragraphs, identify and mark the real load pulsation subsequence; Step S5: Separate and recombine the real load pulsation subsequence and the interference trace isolation set to construct an edge purification waveform stream, and record the purity mark of each subsequence to obtain the preprocessed data stream set, comprising: Step S51: concatenate the real load pulsation subsequence in the original time sequence label order to generate a main waveform chain; Step S52: Place the interference trace isolation set in an independent buffer area, which is physically separated from the main waveform chain; Step S53: Based on the peak amplitude density of each real load pulsation subsequence in the main waveform chain, calculate the proportion of internal non-deviation points as the purity mark; Step S54: Attach the purity mark to each real load pulsation subsequence to fuse into an edge purification waveform stream to obtain the preprocessed data stream set.
2. The edge computing based power consumption data preprocessing method of claim 1, wherein, Step S1 includes: Step S11: Use the current sensor built-in the edge node to millisecond-level sample the power distribution line to capture continuous current waveform data; Step S12: Label each point of the continuous current waveform data with the collection time as the time sequence label, and calculate the peak amplitude of each point to construct the original waveform attribute set.
3. The edge computing based power consumption data preprocessing method of claim 2, wherein, Label each point of the continuous current waveform data with the collection time as the time sequence label, and calculate the peak amplitude of each point to construct the original waveform attribute set, comprising: Extract the absolute time sequence position of each sampling point from the continuous current waveform data to generate a time sequence label sequence; For each position in the time sequence label sequence, scan the amplitude extreme value in the microsecond interval before and after it to determine the peak amplitude; Pair the time sequence label sequence with the peak amplitude one by one to fuse into a structured original waveform attribute set.
4. The edge computing based power consumption data preprocessing method of claim 1, wherein, Step S2 includes: Step S21: Initialize the starting width of the sliding window to a preset fixed millisecond pulse length; Step S22: Push the window along the original waveform attribute set point by point, and calculate the rolling variance of the peak amplitude in the window as the local amplitude fluctuation rate; Step S23: When the local amplitude fluctuation rate exceeds the preset anchor threshold, anchor the current window tail as the segmentation boundary to generate a transient sensitive paragraph; Step S24: Repeat steps S22 to S23 until the entire original waveform attribute set is covered to obtain the sequence of transient sensitive paragraphs.
5. The edge computing based power consumption data pre-processing method of claim 1, wherein, Step S3 includes: Step S31: For each paragraph in the sequence of transient sensitive paragraphs, concatenate the peak amplitude in time sequence label order to obtain a simulated pulse propagation path; Step S32: Check the time interval between adjacent peak amplitudes in the simulated pulse propagation path. If the interval exceeds the preset continuity threshold, mark it as a deviation point. Step S33: Expand the micro window around each deviation point and extract the isolated peak amplitude inside as a discontinuous pulse subset. Step S34: Remove the discontinuous pulse subset from the transient sensitive paragraph sequence to obtain the interference trace isolation set.
6. The edge computing based power consumption data preprocessing method according to claim 5, characterized in that, Checking the time interval between adjacent peak amplitudes in the simulated pulse propagation path, if the interval exceeds the preset continuity threshold, then mark it as a deviation point, including: Selecting each pair of adjacent peak amplitudes on the simulated pulse propagation path, calculating the time label difference as the time interval; Compare the time interval with the preset continuity threshold based on the line impedance; If the time interval exceeds the range of the continuity threshold, mark the pair as a deviation point on the simulated pulse propagation path. 7.The edge computing based power consumption data preprocessing method of claim 1, wherein, Step S4 includes: Step S41: Select the remaining part of each two adjacent paragraphs in the transient sensitive paragraph sequence, and extract the peak amplitudes at the tail and head as anchor points; Step S42: Calculate the time label offset between the anchor points, and check whether the offset meets the preset load cycle synchronization band; Step S43: If the offset falls within the synchronization band, merge the remaining parts of the two adjacent paragraphs into a continuous pulsation unit; Step S44: Discriminate the amplitude gradient smoothness of all continuous pulsation units one by one, and if the amplitude gradient smoothness is higher than the preset interference threshold, mark it as a real load pulsation subsequence.
8. The edge computing based power consumption data preprocessing method according to claim 7, characterized in that, Discriminating the amplitude gradient smoothness of all continuous pulsation units one by one, and if the amplitude gradient smoothness is higher than the preset interference threshold, mark it as a real load pulsation subsequence, including: For each continuous pulsation unit, calculate the cumulative coefficient of variation of adjacent amplitude difference along its peak amplitude sequence as the amplitude gradient smoothness; Compare the amplitude gradient smoothness with the preset interference threshold; If the amplitude gradient smoothness exceeds the preset interference threshold, mark the continuous pulsation unit as a real load pulsation subsequence. 9.The edge computing based power consumption data preprocessing method of claim 1, wherein, Concatenate the real load pulsation subsequences in the original time label order to generate the main waveform chain, including: For each subsequence in the real load pulsation subsequence, select its head and tail peak amplitudes as bridge anchor points, and estimate the potential jump risk based on the amplitude difference between the anchor points to generate a bridge anchor point risk set; Based on the bridge anchor point risk set, pair the tail anchor points and head anchor points of adjacent subsequences in the time label order to generate micro gradient transition segments; Based on the micro gradient transition segment, adjust the continuity weight of the micro gradient transition segment node by node to form a preliminary gapless ring link draft, and temporarily solidify it in the memory cache of the edge node; Based on the preliminary gapless ring link draft, perform lightweight hash verification on each concatenated node to verify the uniqueness of the time label and lock the link integrity level by level, thereby iteratively optimizing to generate the final main waveform chain.
Citation Information
Patent Citations
High-frequency discharge signal identification method
CN120103088A
AI intelligent communication data processing method and system based on edge computing
CN120475383A