Power distribution energy consumption data preprocessing method based on edge calculation
By collecting and processing current waveform data of power distribution lines at edge nodes, and using dynamic window sliding segmentation and synchronous offset verification, the real load signal is accurately identified, solving the problem of misjudgment of interference signals in edge side data preprocessing, realizing efficient load forecasting and fault diagnosis, and reducing latency and communication costs.
Patent Information
- Application Number
- CN202511913422.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-18
- Publication Date
- 2026-01-16
- Estimated Expiration
- 2045-12-18
Smart Images

Figure CN121350422A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, more particularly, the present application relates to a power distribution energy consumption data preprocessing method based on edge computing. BACKGROUND
[0002] With the in-depth development of smart grid and energy internet technology, fine management and intelligent operation of power distribution system have become the core support to achieve the double carbon goal and ensure the safety of power supply; in the power distribution network energy consumption monitoring and optimization scene, real-time collection and analysis of massive energy consumption data of electrical equipment is of great significance for load forecasting, fault diagnosis, demand response and energy efficiency evaluation; in order to reduce data transmission cost and meet the demand of millisecond level response, in recent years, edge computing architecture has been widely deployed in power distribution energy consumption monitoring system, effectively relieving the bandwidth pressure and delay bottleneck of cloud computing center, and providing a technical basis for real-time intelligent decision of power distribution system.
[0003] The existing edge side data preprocessing scheme has significant limitations in accurately distinguishing real load characteristics and transient interference signals, which is particularly prominent in the complex electromagnetic environment of power distribution monitoring scene; specifically, the original energy consumption data collected by the power distribution system is generally mixed with motor start-stop impact, switch operation transient, electromagnetic coupling crosstalk and other transient interference components; these interference signals and real load mutation, power fluctuation have high similarity in time domain waveform and frequency domain distribution, and the boundary of the feature space is blurred and interwoven. For example, in the industrial park power distribution monitoring, the current step feature generated by the normal starting process of high-power equipment has significant overlapping area with the sharp pulse caused by line fault or electromagnetic interference in the amplitude change rate, duration and frequency spectrum component, etc. This feature similarity makes it difficult to establish a clear discrimination criterion in the preprocessing link. This discrimination difficulty not only leads to the effective load signal being misjudged as interference and filtered out, but also makes the interference components be retained as real data and transmitted to the subsequent analysis link, which causes false conclusions in equipment anomaly detection and load characteristic analysis; the current mainstream method to solve this problem is to deploy complex deep learning classification models on high-performance edge servers, or to upload the original data to the cloud for fine processing using centralized computing power; however, the former needs to significantly improve the hardware configuration and power consumption budget of the edge node, which is difficult to bear the deployment cost of massive distributed monitoring points, and the model inference delay may still exceed the real-time requirement; the latter violates the original intention of edge computing, not only consumes a large amount of communication bandwidth resources, but also causes data backlog and timeliness loss when the network fluctuates, which cannot meet the core needs of local autonomy and rapid response of power distribution system, and ultimately restricts the deep application and large-scale promotion of edge computing architecture in the field of intelligent power distribution.
[0004] In view of this, the application provides a power distribution energy consumption data preprocessing method based on edge computing to solve the above problems. SUMMARY
[0005] In order to overcome the above-mentioned defects of the prior art, in order to achieve the above-mentioned purpose, the application provides the following technical scheme: a power distribution energy consumption data preprocessing method based on edge computing, comprising: Step S1: collecting real-time current waveform data of the power distribution line through the edge node, and extracting the time sequence label and amplitude peak value of each sampling point to obtain an original waveform attribute set; Step S2: performing dynamic window sliding segmentation on the original waveform attribute set, wherein the window boundary is adaptively anchored by the local amplitude fluctuation rate to obtain a transient sensitive paragraph sequence; Step S3: for each paragraph in the transient sensitive paragraph sequence, tracking the continuity deviation of the internal pulse propagation path, and isolating the discontinuous pulse subset to obtain an interference trace isolation set; Step S4: based on the interference trace isolation set and the remaining part of the transient sensitive paragraph sequence, performing synchronization offset checking between adjacent paragraphs, discriminating and marking a real load pulsation sub-sequence; Step S5: separating and recombining the real load pulsation sub-sequence and the interference trace isolation set, constructing an edge purification waveform stream, and recording the purity mark of each sub-sequence to obtain a data stream set after preprocessing.
[0006] The power distribution energy consumption data preprocessing method based on edge computing has the following technical effects and advantages: The application obtains a transient sensitive paragraph sequence by locally adapting the amplitude fluctuation rate to anchor the dynamic window boundary, accurately positioning the time domain interval that needs to be distinguished; the pulse propagation path of the real load signal presents the characteristics of time sequence continuity and peak amplitude gradualness, while the transient interference signal presents the characteristics of path discontinuity and time sequence gap jumping due to the randomness of electromagnetic coupling crosstalk and switching operation transient, the continuity deviation of the pulse propagation path of the transient sensitive paragraph presents the degree of isolation of the interference signal, and the interference trace isolation set is obtained by combining the isolation of the discontinuous pulse subset, so as to reduce the dependence of real load feature analysis on transient interference signals; according to the synchronization offset check result of the remaining part of the transient sensitive paragraph sequence between adjacent paragraphs, and the amplitude gradualness smoothness of the remaining part of the transient sensitive paragraph sequence, the real load pulsation sub-sequence is adaptively distinguished, avoiding the misjudgment of large-scale load periodic fluctuations caused by analyzing the continuity characteristics of a single paragraph, reducing the probability of misidentifying transient interference signals as real load features, so that the edge purification waveform flow and its purity mark constructed based on separation and recombination accurately reflect the real energy consumption characteristics of the power distribution system, improve the accuracy of load prediction, fault diagnosis and energy efficiency evaluation based on the data flow set completed by preprocessing under the condition of limited computing power of edge nodes, and reduce the communication cost and delay overhead of data retransmission and cloud secondary processing caused by misjudgment. BRIEF DESCRIPTION OF DRAWINGS
[0007] Figure 1 FIG. 1 is a schematic diagram of a power distribution energy consumption data preprocessing method based on edge computing according to the present application; Figure 2 FIG. 2 is a schematic diagram of a power distribution energy consumption data preprocessing system based on edge computing according to the present application. DETAILED DESCRIPTION
[0008] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0009] Embodiment 1 Please refer to Figure 1 The present embodiment is a power distribution energy consumption data preprocessing method based on edge computing, which comprises: Step S1: Collecting real-time current waveform data of the power distribution line through the edge node, and extracting the time sequence label and amplitude peak value of each sampling point to obtain an original waveform attribute set.
[0010] In actual operation, the line current of the power distribution network presents complex time-varying characteristics, which not only contains regular current changes caused by normal load fluctuations, but also contains abnormal fluctuations introduced by non-normal factors such as line interference, electromagnetic pulses, and switching transients. The traditional centralized data acquisition system directly uploads the original waveform data to the cloud server for processing, which has problems such as large data transmission delay, high network bandwidth occupation, and poor real-time performance. The edge computing architecture can complete preliminary cleaning and feature extraction at the first time of data generation by sinking the data preprocessing capability to the edge node close to the data source, effectively reducing the transmission amount of invalid data and improving the real-time response capability of energy consumption monitoring.
[0011] Preferably, in some possible implementation manners of the embodiment of the present application, step S1 comprises: Step S11: millisecond-level sampling of the power distribution line is performed by using the current sensor built-in the edge node to capture continuous current waveform data.
[0012] In this embodiment, a Hall effect current sensor is selected as the core acquisition component of the edge node, which has the characteristics of non-contact measurement, fast response speed, and high linearity. The sampling frequency of the sensor is set to 2 kHz, that is, an instantaneous current value is collected every 0.5 milliseconds, which can effectively capture the current components in the power distribution line. In the actual collection process, the analog current signal output by the sensor is converted into a digital signal by a 16-bit analog-to-digital converter, and the quantization accuracy is 0.01 ampere. The local buffer area of the edge node stores the continuous current waveform data of the last 30 seconds in a circular queue manner, and occupies about 120 KB of memory.
[0013] It should be noted that in other embodiments of the present application, sensor types such as electromagnetic current transformers and Rogowski coils can also be used, the sampling frequency can be adjusted to 1 kHz to 10 kHz according to the frequency characteristics of the monitored object, the number of bits of the analog-to-digital converter can be selected to be 12 to 24, and the sampling accuracy and buffer size can be flexibly configured according to the hardware resources of the edge node and the application scenario, which is not limited herein.
[0014] Step S12: each point of the continuous current waveform data is labeled with the collection time as a time sequence label, and the peak amplitude of each point is calculated to construct an original waveform attribute set.
[0015] Although the continuous current waveform data contains rich time domain information, the original sampling value sequence lacks structured spatio-temporal correlation identification, which is not conducive to subsequent segmentation processing and feature extraction. In order to establish a spatio-temporal index system of the waveform data, an accurate time sequence label needs to be assigned to each sampling point, and the local peak value feature thereof needs to be extracted.
[0016] Preferably, in some possible implementation manners of the embodiment of the present application, step S12 comprises: The absolute time sequence position of each sampling point is extracted from the continuous current waveform data to generate a time sequence label sequence. Specifically, when starting the collection task, the edge node synchronizes the local real-time clock with the network time protocol server to obtain the current absolute timestamp (with a precision of microseconds), denoted as the collection start time . For the th sampling point (N , is the total number of sampling points in the buffer), the absolute time sequence position is calculated as: ; in the formula, is the time sequence label of the th sampling point; is the collection start time; Δt is the sampling period, and for a 2 kHz sampling frequency, is 500 microseconds.
[0017] For each position in the time sequence label sequence, the amplitude extreme value in the microsecond interval before and after the position is scanned to determine the peak amplitude. The power distribution line current waveform presents a sawtooth characteristic on a microscopic scale, and the instantaneous value of a single sampling point may be affected by quantization noise and high-frequency interference, and thus cannot accurately represent the true current level at the time. To improve the robustness of the amplitude feature, the local peak extraction strategy is adopted in this embodiment: for the th sampling point, a time window is defined with the sampling point as the center, where is the half window width, which is 2 sampling periods, i.e., 1000 microseconds. The maximum current amplitude in the window is searched and denoted as the peak amplitude of the th sampling point. The time sequence label sequence and the peak amplitude are paired one by one to be fused into a structured original waveform attribute set. The specific form is: The original waveform attribute set is: ; Each element is called a waveform attribute point, which completely describes the time sequence position and amplitude feature of the i th sampling point. The set is stored in the form of a two-dimensional table in the memory of the edge node, which is convenient for subsequent streaming processing in time sequence order.
[0018] It should be noted that the value of the half window width needs to be determined comprehensively according to the sampling frequency and the line characteristics, and is generally set to 2 to 5 sampling periods; in other embodiments, the current instantaneous value can be smoothed by using methods such as median filtering and mean filtering, and then directly used as the amplitude feature, or the effective value (RMS), the crest factor, and other derived indicators can be extracted, which are not limited herein.
[0019] Step S2: performing dynamic window sliding segmentation on the original waveform attribute set, wherein a window boundary is adaptively anchored by a local amplitude fluctuation rate, to obtain a sequence of transient sensitive paragraphs.
[0020] The current waveform of a power distribution line has typical non-stationary characteristics. Load switching, fault transients, electromagnetic interference and other events can introduce local severe fluctuations in the waveform. These fluctuation segments carry key abnormal information and need to be separated from the stationary background for focused analysis. Traditional fixed window segmentation methods uniformly divide the waveform according to a preset time length, which cannot adapt to the time-varying nature of fluctuation characteristics, and easily splits the waveform of the same event into different segments or mixes unrelated stationary segments and fluctuation segments together, reducing the relevance of subsequent processing.
[0021] The dynamic window sliding segmentation method adaptively adjusts the segmentation boundary according to the local characteristics of the waveform, and can accurately locate the transient region with severe fluctuations. The local amplitude fluctuation rate reflects the degree of change in the peak amplitude of the waveform within a short time window. When the fluctuation rate exceeds a preset threshold, it indicates that the position may correspond to the start or end time of a load mutation or an interference event, and is suitable as a segmentation boundary. Through this adaptive anchoring mechanism, the original waveform attribute set can be divided into several transient sensitive paragraphs, each corresponding to a relatively independent current fluctuation process.
[0022] Preferably, in some possible implementation manners of the embodiment of the present application, step S2 comprises:
[0023] Step S21: initializing the starting width of the sliding window as a preset fixed millisecond pulse length.
[0024] Common transient events in power distribution lines (such as switch operation, capacitor switching, motor starting, etc.) generally have a duration in the range of 10 to 100 milliseconds, and interference pulses (such as lightning induction, arc discharge, etc.) generally have a duration in the range of 1 to 10 milliseconds. In order to balance the capture ability for events of different time scales, the starting width of the sliding window is set to 20 milliseconds in this embodiment, corresponding to 40 sampling points under a 2 kHz sampling frequency.
[0025] When initialized, the left boundary of the sliding window is located at the 1st waveform attribute point of the original waveform attribute set, and the right boundary is located at the 40th waveform attribute point, and the window contains 40 waveform attribute points.
[0026] It should be noted that the starting width may be adjusted according to the typical event time scale of the monitored object. For high-voltage transmission lines, it can be set to 50 to 200 milliseconds, and for low-voltage user side, it can be set to 5 to 20 milliseconds. In other embodiments of the present application, the starting width can also be set to a fixed number of sampling points rather than a fixed time length, which is not limited here.
[0027] Step S22: The window is pushed forward point by point along the original waveform attribute set, and the rolling variance of the peak amplitude in the window is calculated as the local amplitude fluctuation rate.
[0028] The sliding window is pushed forward point by point in the time tag direction with a single sample point as a step. After the first push, the left boundary of the window is located at the 1st waveform attribute point, the right boundary is located at the 40th waveform attribute point, and the window contains 40 waveform attribute points from the 1st to the 40th. After the second push, the left boundary of the window is located at the 2nd waveform attribute point, the right boundary is located at the 80th waveform attribute point, and the window contains 40 waveform attribute points from the 2nd to the 81st. After the third push, the left boundary of the window is located at the 3rd waveform attribute point, the right boundary is located at the 120th waveform attribute point, and the window contains 40 waveform attribute points from the 3rd to the 122nd. After the fourth push, the left boundary of the window is located at the 4th waveform attribute point, the right boundary is located at the 160th waveform attribute point, and the window contains 40 waveform attribute points from the 4th to the 163rd.
[0029] The variance measures the dispersion degree of data deviating from the mean value. If the peak amplitude in the window changes dramatically, the variance value is large; if the peak amplitude in the window is relatively stable, the variance value is small. By rolling calculation of the variance of each window position, the local amplitude fluctuation rate sequence is obtained, which describes the fluctuation intensity distribution of the original waveform attribute set along the time axis.
[0030] Step S23: When the local amplitude fluctuation rate exceeds the preset anchoring threshold, the tail end of the current window is anchored as a segmentation boundary, and a transient sensitive paragraph is generated.
[0031] The preset anchoring threshold is used to determine whether the local amplitude fluctuation rate reaches the significance level of the transient event. In this embodiment, the anchoring threshold is set based on the statistical characteristics of historical operation data: : The current waveform data of the power distribution line under normal operating conditions for 7 consecutive days is selected, and the 95th percentile of the global local amplitude fluctuation rate is calculated as the anchoring threshold.
[0032] During the sliding process, the local amplitude fluctuation rate of the current window is compared with the anchoring threshold in real time: if the local amplitude fluctuation rate is less than or equal to the anchoring threshold, it is considered that the current window is in a stable paragraph, and the window continues to be pushed forward; if the local amplitude fluctuation rate is greater than the anchoring threshold, it is considered that the current window enters a transient fluctuation region, and the right boundary of the window (i.e., the time tag of the 40th waveform attribute point ) is the segmentation boundary.
[0033] After the segmentation boundary is anchored, all waveform attribute points between the last segmentation boundary (or the starting point of the set) and the current segmentation boundary form a transient sensitive paragraph. For example, if the 1st segmentation boundary is anchored at (the 150th sample point), and the 2nd segmentation boundary is anchored at (the 380th sample point), two transient sensitive paragraphs are generated: Paragraph 1: contains waveform attribute points Corresponding time span 0 to 74.5 milliseconds, a total of 150 sampling points; Paragraph 2: contains waveform attribute points Corresponding time span 75.0 to 189.5 milliseconds, a total of 230 sampling points.
[0034] It should be noted that the setting of the anchor threshold should comprehensively consider factors such as line voltage level, load type, and environmental interference level. The threshold value can be appropriately increased for high-voltage lines and industrial load concentration areas, and the threshold value can be appropriately reduced for low-voltage residential power areas. In other embodiments of the present application, a dynamic threshold strategy can also be used to adjust the threshold value in real time according to the fluctuation rate statistical characteristics of the previous time window, which is not limited here.
[0035] Step S24: Repeat steps S22 to S23 until the entire original waveform attribute set is covered, and obtain a sequence of transient sensitive paragraphs.
[0036] The sliding window starts from the 1st waveform attribute point and advances point by point to the (N-39)th waveform attribute point (N is the total number of sampling points of the original waveform attribute set), and scans the entire waveform data completely. Whenever a local amplitude fluctuation rate exceeds the anchor threshold, a new segmentation boundary and a transient sensitive paragraph are generated. Finally, M transient sensitive paragraphs (the value of M depends on the number of transient events in the waveform) are obtained, which are recorded as a sequence of transient sensitive paragraphs. Each paragraph contains a starting time label, an ending time label, and a corresponding set of waveform attribute points.
[0037] For example, for 30 seconds of waveform data containing 60000 sampling points, after dynamic window sliding segmentation, 15 transient sensitive paragraphs may be obtained, each with a duration of 10 to 200 milliseconds, covering all significant fluctuation regions. The fluctuation rate of the smooth running region is always below the threshold value and will not be segmented into paragraphs alone, but as a transition region between adjacent paragraphs.
[0038] It should be noted that in other embodiments of the present application, the paragraph starting boundary can be anchored when the fluctuation rate rising edge (from below the threshold to above the threshold) is detected, and the paragraph ending boundary can be anchored when the fluctuation rate falling edge (from above the threshold to below the threshold) is detected, so that each paragraph is more accurately aligned with the actual duration of the transient event; A hysteresis comparison mechanism can also be introduced, with a rising threshold and a falling threshold (the falling threshold is lower than the rising threshold), to avoid excessive segmentation caused by repeated fluctuations of the fluctuation rate around the threshold. This will not be described again.
[0039] Step S3: For each paragraph in the sequence of transient sensitive paragraphs, track the continuity deviation of the internal pulse propagation path, and isolate the discontinuous pulse subset to obtain the interference trace isolation set.
[0040] Although the time region of current fluctuation is located in the transient sensitive paragraph, the paragraph may mix fluctuation components of multiple natures: one type is a continuous pulse sequence caused by real load change, whose peak amplitude presents a smooth gradient or periodic law along the time axis, and the time interval is uniform; another type is an isolated pulse jump caused by external interference or measurement anomaly, whose peak amplitude suddenly appears and disappears quickly, and the waveform lacks continuity before and after, and the time interval is irregular.
[0041] The physical propagation process of current in distribution lines follows the continuity principle: the establishment and decay of load current are constrained by line inductance, load impedance and other parameters, the change rate has an upper limit, and there is a smooth causal relationship between current values at adjacent time instants. Interference signals (such as electromagnetic pulses, sensor glitches, data packet retransmission, etc.) are not subject to these physical constraints and can produce large amplitude jumps in a very short time, appearing as discontinuous isolated peaks in the waveform.
[0042] By tracking the time continuity of peak amplitudes in the transient sensitive paragraph, these discontinuous pulses can be identified and isolated, classified as interference traces, and prevented from polluting subsequent load feature analysis.
[0043] Preferably, in some possible implementation manners of the embodiments of the present application, step S3 comprises: Step S31: for each paragraph in the sequence of transient sensitive paragraphs, the peak amplitudes are sequentially connected in the order of time labels to obtain a simulated pulse propagation path.
[0044] For the first paragraph in the sequence of transient sensitive paragraphs, all waveform attribute points contained therein are extracted, the peak amplitudes of all waveform attribute points are arranged in the order of time labels, and a peak amplitude sequence is formed. The sequence forms a polyline in the time-amplitude two-dimensional plane, simulating the propagation trajectory of the current pulse in the time range of the paragraph, which is called a simulated pulse propagation path.
[0045] For example, a paragraph contains 50 sampling points, corresponding to a time span of 20 to 44.5 milliseconds, and the peak amplitude gradually rises from 12.0 amperes to 18.0 amperes and then falls to 10.0 amperes, and the simulated pulse propagation path presents an envelope line form of first rising and then falling, reflecting the process of load input and stabilization.
[0046] Step S32: check the time interval between adjacent peak amplitudes in the simulated pulse propagation path, and if the interval exceeds a preset continuity threshold, mark it as a deviation point.
[0047] In the pulse propagation path of real load current, the time interval between adjacent sampling points should be equal to the sampling period (For 2kHz sampling, Δt = 0.5ms), allowing the presence of small clock jitter (usually no more than ±10μs). If the time interval between adjacent sampling points deviates significantly from the sampling period, it indicates that data loss, transmission delay or clock anomaly may have occurred at this position, resulting in the timing continuity of the pulse propagation path being broken.
[0048] Preferably, in some possible implementations of the embodiments of the present application, step S32 comprises: Selecting each pair of adjacent peak amplitudes on the simulated pulse propagation path, calculating the difference of their timing labels as the timing gap, for paragraph the th waveform attribute point , representing the number of attribute points in the paragraph, calculating the timing gap : ; in which, is the timing gap of the th pair of adjacent peak amplitudes in the paragraph; and are the timing labels of the th and the th waveform attribute point, respectively. Compare the timing gap with the preset continuity threshold based on the line impedance. The continuity threshold reflects the maximum allowable time interval of current change of the distribution line. This embodiment sets the continuity threshold based on the impedance characteristics and load response speed of the line: for industrial distribution lines with a high proportion of resistive-inductive loads, the time constant of current change is generally in the range of 1 to 5ms, and the continuity threshold is set to 1.5ms, i.e., the timing gap is allowed to deviate by no more than ±1000μs (corresponding to ±2 sampling periods) based on the sampling period of 500μs; for residential distribution lines dominated by resistive loads, the time constant is smaller, and the continuity threshold can be set to 1.0ms. If the timing gap satisfies: wherein, represents the preset continuity threshold, it is determined that the timing continuity between the pair of adjacent peak amplitudes is abnormal, and the th waveform attribute point on the simulated pulse propagation path is marked as a deviation point in real time.
[0049] For example, if the continuity threshold is 1500μs and the sampling period is 500μs, the timing gap is allowed to be in the range of to In practice, it is unreasonable for the timing gap to be negative, so the effective range is (0, 1500) microseconds. If the timing labels of two adjacent sampling points are 10000 microseconds and 12000 microseconds respectively, the timing gap is 2000 microseconds > 1500 microseconds, exceeding the continuity threshold, and the point is marked as a deviation point.
[0050] Step S33: Expand the micro window around each deviation point, and extract the isolated peak amplitudes inside it as a discontinuous pulse subset.
[0051] The identification of the deviation point locates the position of the timing continuity anomaly, but a single deviation point may only be a boundary marker of transient interference, and the interference signal itself may span multiple consecutive sampling points. To completely extract the interference waveform segment, a micro window needs to be expanded around the deviation point to capture all peak amplitudes within a certain range before and after the deviation point.
[0052] For each waveform attribute point marked as a deviation point in the paragraph, define a micro window centered on it , where is the micro window half-width. In this embodiment, the micro window half-width is set to 5 milliseconds, corresponding to ±10 sampling points, which can cover the duration of most short-time interference pulses.
[0053] Within the micro window range, extract all waveform attribute points whose timing labels fall within the micro window to form a discontinuous pulse subset centered on the deviation point.
[0054] For example, the 120th sampling point (timing label 59500 microseconds) in a paragraph is marked as a deviation point, and the micro window range is [54500 microseconds, 64500 microseconds], corresponding to the 110th to 130th sampling points. These 21 sampling points and their peak amplitudes form a discontinuous pulse subset. If the peak amplitudes in this subset exhibit a sharp spike (e.g., from 10.0 amperes to 28.0 amperes and back to 9.5 amperes), with a duration of about 10 milliseconds, it meets the typical characteristics of an electromagnetic interference pulse.
[0055] It should be noted that the micro window half-width is determined according to the typical duration of the interference pulse, and for high-frequency electromagnetic interference, it can be set to 1 to 3 milliseconds. If the micro windows of multiple deviation points overlap, they can be merged and de-duplicated to avoid the same interference segment being extracted repeatedly. In other embodiments of the present application, an adaptive window width can also be used to dynamically adjust the micro window size according to the jump amplitude of the peak amplitudes before and after the deviation point, which is not limited here.
[0056] Step S34: Remove the discontinuous pulse subsets from the transient sensitive paragraph sequence to obtain the interference trace isolation set.
[0057] The discontinuous pulse subsets identified in all transient sensitive paragraphs are aggregated to form an interference trace isolation set; at the same time, all waveform attribute points contained in the discontinuous pulse subsets are deleted from the original transient sensitive paragraph sequence, and the remaining part is retained as the continuous pulse part of the paragraph for subsequent real load feature analysis. For example, if paragraph 1 originally has 200 sampling points, after removing 2 discontinuous pulse subsets (a total of 42 sampling points), the remaining 158 sampling points constitute the continuous pulse part of paragraph 1.
[0058] For example: for 30 seconds of waveform data containing 60000 sampling points, after pulse continuity tracking, 25 discontinuous pulse subsets may be identified, involving about 600 sampling points (1% of the total), which are isolated to the interference trace isolation set and do not participate in subsequent load feature extraction and energy consumption calculation, effectively avoiding the influence of interference signals on the accuracy of energy consumption statistics.
[0059] Step S4: Based on the interference trace isolation set and the remaining part of the transient sensitive paragraph sequence, the synchronization offset between adjacent paragraphs is checked, and the real load pulsation subsequence is identified and marked.
[0060] After the processing of step S3, the transient sensitive paragraph sequence has removed the internal discontinuous interference pulses, but the time sequence correlation between paragraphs has not been fully utilized. The real load change of the power distribution line often presents periodic or quasi-periodic characteristics, for example, industrial production equipment starts and stops according to a fixed beat, and air conditioner compressors operate according to a temperature feedback period. These load pulsations exhibit time synchronization between multiple paragraphs: the starting time, duration or peak amplitude of adjacent paragraphs present regular intervals.
[0061] On the contrary, the occurrence time and intensity of random interference (such as lightning induction, radio frequency interference, etc.) do not have this time synchronization, and the time offset between adjacent paragraphs presents a random distribution. By checking whether the time offset between adjacent paragraphs falls within the load cycle synchronization band, real load pulsation and residual interference signals can be further distinguished, improving the accuracy of data preprocessing.
[0062] At the same time, the change process of real load current is constrained by physical inertia, and the gradual change of peak amplitude along the time axis should present a smooth envelope, while the amplitude change of interference signals often presents sudden changes and oscillations. By evaluating the gradual smoothness of the peak amplitude within the paragraph, the authenticity of the paragraph can be verified from the amplitude feature dimension.
[0063] Preferably, in some possible implementation manners of the embodiments of the present application, step S4 includes: Step S41: Selecting the remaining part of each two adjacent paragraphs in the transient sensitive paragraph sequence, extracting the peak amplitudes at the tail end and the head end as anchor points.
[0064] For the transient sensitive segment sequence, the first The paragraph and the first Each segment, after removing the discontinuous pulse subset, is denoted as a segment. The rest of the paragraphs The remaining part.
[0065] Extract paragraphs The last waveform attribute point of the remaining part is used as the tail anchor point to extract the paragraph. The first waveform attribute point of the remaining portion serves as the beginning anchor point. The anchor point marks the boundary between adjacent segments on the time axis, and its peak amplitude reflects the current level at the segment boundary.
[0066] Step S42: Calculate the time tag offset between anchor points and verify whether the offset conforms to the preset load cycle synchronization band.
[0067] The time tag offset between anchor points is defined as the difference between the time tags of the first anchor points of two adjacent paragraphs. The load cycle synchronization band reflects the time cycle range of typical periodic loads of the power distribution line. This embodiment sets the synchronization band based on the load composition characteristics of the monitored line: For commercial power distribution lines containing a large amount of air conditioning load, the start-stop cycle of air conditioning compressors is generally 5 to 15 minutes, and the load cycle synchronization band is set to [280,000 microseconds, 920,000 microseconds], allowing a cycle fluctuation of ±20,000 microseconds; For industrial production lines, the equipment cycle may be 10 seconds to 5 minutes, and the synchronization band range needs to be adjusted according to the actual process parameters.
[0068] Verify offset Does it fall within the load cycle synchronization zone? :like Then determine the paragraph. and paragraphs It has periodic synchronization and may correspond to repetitive pulsations of the same type of load; like or If the two segments do not have periodic synchronization, they may belong to different types of loads or random interference.
[0069] It should be noted that the setting of the load cycle synchronization band needs to be based on the load type statistics and historical operation data of the line. For mixed load lines, multiple synchronization bands can be set to correspond to loads of different cycles. In other embodiments of the present invention, autocorrelation analysis, spectrum analysis and other methods can also be used to automatically extract the dominant cycle from historical waveform data and dynamically update the synchronization band range, which is not limited here.
[0070] Step S43: If the offset falls into the synchronization band, then merge the remaining parts of the two adjacent segments into a continuous pulse unit.
[0071] If the determination paragraph and paragraph has periodic synchronism, the remaining part of the two paragraphs is connected head to tail on the time axis, and merged into a continuous pulsation unit. The merged pulsation unit contains all the waveform attribute points of the remaining part of the paragraph and paragraph . If the offset of multiple consecutive paragraphs all falls into the synchronization band, they are sequentially merged into longer continuous pulsation units. For example, if the offset of paragraphs 1, 2, and 3 all falls into the synchronization band, the three paragraphs are merged into a continuous pulsation unit, covering the complete time span from the start of paragraph 1 to the end of paragraph 3.
[0072] If the offset does not fall into the synchronization band, i.e., the two paragraphs are determined not to have periodic synchronism, the two paragraphs remain independent and are not merged. After the pairwise checking and merging of all adjacent paragraphs, the transient sensitive paragraph sequence is reorganized into several continuous pulsation units, each corresponding to a possible periodic load pulsation process.
[0073] For example, there are originally 15 transient sensitive paragraphs. After synchronization offset checking, paragraphs {1, 2, 3} are merged into unit 1, paragraphs {5, 6} are merged into unit 2, and paragraphs {8, 9, 10, 11} are merged into unit 3. The remaining paragraphs {4, 7, 12, 13, 14, 15} remain independent because their offsets do not conform to the synchronization band. Finally, 10 continuous pulsation units are obtained.
[0074] Step S44: Gradual amplitude variation smoothness of all continuous pulsation units is discriminated one by one. If the gradual amplitude variation smoothness is higher than a preset disturbance threshold, it is marked as a real load pulsation subsequence.
[0075] Preferably, in some possible implementation manners of the embodiments of the present application, step S44 comprises: For each continuous pulsation unit, the cumulative coefficient of variation of adjacent amplitude differences along its peak amplitude sequence is calculated as the gradual amplitude variation smoothness. The specific method is as follows: for a continuous pulsation unit containing waveform attribute points, the peak amplitude sequence is extracted, and the adjacent amplitude difference sequence is calculated, where .
[0076] The adjacent amplitude difference reflects the point-by-point change amount of the peak amplitude. The real load current is constrained by the line impedance and the load time constant. The amplitude change rate has a physical upper limit, and the adjacent amplitude difference should be within a reasonable range. The amplitude of the interference signal can jump instantaneously, and the adjacent amplitude difference can have an abnormally large value. The coefficient of variation of the adjacent amplitude difference sequence is calculated, which is defined as the ratio of the standard deviation to the mean of the adjacent amplitude difference sequence. The coefficient of variation is the ratio of the standard deviation to the mean, which eliminates the dimensional influence and reflects the relative dispersion degree of the data. If the adjacent amplitude difference changes smoothly, the coefficient of variation is small. If the adjacent amplitude difference has a mutation, the coefficient of variation is large.
[0077] The amplitude gradual change smoothness is compared with the preset interference threshold value. In this embodiment, the interference threshold value is set based on the change rule of the real load current: for normal resistive and inductive loads, the amplitude change rate during the current rising and falling processes is relatively uniform, and the coefficient of variation of the adjacent amplitude difference is generally not more than 0.3; for interference signals, the instantaneous jump causes the adjacent amplitude difference to have an extreme value, and the coefficient of variation is usually greater than 0.8. In this embodiment, the preset interference threshold value is set to 0.5, which is used as a dividing line to distinguish real loads and interference signals.
[0078] If the amplitude gradual change smoothness is less than the preset interference threshold value, the continuous pulsation unit is immediately marked as a real load pulsation sub-sequence, and it is considered that the change of the peak amplitude conforms to the physical law and belongs to the real load current waveform. If the amplitude gradual change smoothness is greater than or equal to the preset interference threshold value, it is determined that the unit may contain residual interference or measurement abnormalities, and is not marked as a real load pulsation sub-sequence.
[0079] After the amplitude gradual change smoothness is discriminated, the continuous pulsation unit set is divided into two parts: the units marked as real load pulsation sub-sequences will be used for subsequent energy consumption calculation and load characteristic analysis; the unmarked units are classified as abnormal data together with the interference trace isolation set, and do not participate in energy consumption statistics.
[0080] It should be noted that the setting of the interference threshold value needs to be combined with the load type and the line parameter. For impact loads with rapid changes (such as electric welders and elevators), the threshold value can be appropriately increased to 0.6 to 0.8. For stable loads (such as lighting and heating), the threshold value can be reduced to 0.3 to 0.4. In other embodiments of the present application, other forms of smoothness indicators can also be used, such as the ratio of the range to the mean of the adjacent amplitude difference, which is not limited here.
[0081] Step S5: The real load pulsation sub-sequences and the interference trace isolation set are separated and recombined to construct an edge purification waveform stream, and the purity marks of each sub-sequence are recorded to obtain a preprocessed data stream set.
[0082] After the multi-level screening of the preceding steps, the original waveform attribute set has been divided into two categories: the real load fluctuation sub-sequence represents the actual power consumption characteristics of the power distribution line, with high data quality and reliability; the interference trace isolation set contains various abnormal fluctuations and measurement errors, with low data reliability. In order to facilitate subsequent energy consumption analysis and abnormal diagnosis, it is necessary to physically separate and structurally reorganize the two types of data, and add quality evaluation indicators to each sub-sequence to form a complete preprocessed data stream set.
[0083] The edge purification waveform stream, as the main output of preprocessing, carries the real load current waveform after deep cleaning, and can be directly used for energy consumption calculation, load prediction, equipment state evaluation and other applications; the purity mark, as a quantitative indicator of data quality, provides a reliability reference for downstream analysis algorithms, and supports adaptive weighted processing based on data quality.
[0084] Preferably, in some possible implementation manners of the embodiment of the application, step S5 comprises: Step S51: concatenating the real load fluctuation sub-sequences in the original time sequence label order to generate a main waveform chain.
[0085] All sub-sequences in the real load fluctuation sub-sequence set are arranged in ascending order according to the order of their head anchor point time sequence labels. The arranged sub-sequences are concatenated in turn to form a main waveform chain.
[0086] Preferably, in some possible implementation manners of the embodiment of the application, the generation of the main waveform chain comprises: For each sub-sequence in the real load fluctuation sub-sequence, the head and tail peak amplitudes thereof are selected as bridge anchor points, and the potential jump risk is estimated based on the amplitude difference between the anchor points to generate a bridge anchor point risk set. The specific method is as follows: for the i-th sub-sequence, let the head peak amplitude be , and the tail peak amplitude be ; the head peak amplitude of the j-th sub-sequence is . Define the bridge jump risk ; in the formula, is the bridge jump risk between the i-th and j-th sub-sequences. If the bridge jump risk is large (such as more than 3.0), it indicates that there is a significant difference in current level between adjacent sub-sequences, and direct concatenation may cause an unrealistic amplitude mutation at the bridge point. All bridge points with greater than 3.0 and their risk values are recorded as the bridge anchor point risk set.
[0087] If the bridge jump risk is large (such as more than 3.0), it indicates that there is a significant difference in current level between adjacent sub-sequences, and direct concatenation may cause an unrealistic amplitude mutation at the bridge point. All bridge points with greater than 3.0 and their risk values are recorded as the bridge anchor point risk set.
[0088] Based on the bridge anchor point risk set, the tail end anchor point and the head end anchor point of the adjacent subsequence are paired in the order of the time sequence label to generate a micro-gradual transition section. For the bridge point with a risk value greater than 3.0, a linear gradual transition section is inserted between the tail end anchor point of the first subsequence and the head end anchor point of the second subsequence.
[0089]
[0090] For the bridge point with a risk value less than or equal to 3.0, the tail end anchor point of the first subsequence and the head end anchor point of the second subsequence are directly connected on the time axis without inserting a virtual sampling point.
[0091]
[0092] The real sampling points, the virtual interpolation points and the continuity weights of all subsequences are assembled in the order of the time sequence label to form a preliminary gapless annular link draft, which is stored in the memory buffer of the edge node.
[0093] Based on the preliminary gapless ring link draft, a lightweight hash check is performed on each serial node to verify the uniqueness of the timing label and lock the link integrity level by level, thereby iteratively optimizing the final main waveform chain. The hash check method is as follows: for the timing label of each waveform attribute point, the hash value is calculated through hash operation, and it is checked whether there is a repeated value in the hash value sequence. If a repeated hash value is found, it means that there is a timing label conflict, and the timing label of the corresponding node needs to be checked and fine-tuned: increase the offset by 1 microsecond, repeat until the hash value is unique.
[0094] After completing the hash check, a unique link index ID is assigned to each node of the main waveform chain, a bidirectional mapping table of timing label and link index is established, and the integrity and order of the link are locked. The final main waveform chain is stored as a structured array, each element containing four fields: {link index ID, timing label, peak amplitude, continuity weight}.
[0095] It should be noted that the number of inserted virtual sampling points can be adjusted adaptively according to the bridge jump risk value, and the more points are inserted to enhance the smoothing effect when the risk is greater; the allocation strategy of continuity weight can also adopt other forms such as Gaussian distribution and exponential decay; the modulus of hash check can be adjusted according to the total number of nodes, which is not limited here.
[0096] Step S52: Place the interference trace isolation set in an independent buffer area, which is physically separated from the main waveform chain.
[0097] The interference trace isolation set data is completely separated from the main waveform chain and stored in an independent buffer area of the edge node, which is allocated a different memory address space to avoid confusion with the purified main waveform chain.
[0098] The interference trace isolation set is organized in the form of a list, and each element records a {start timing label, end timing label, peak amplitude sequence, interference type label} of an interference segment. The interference type label is divided into two categories: timing discontinuity and amplitude mutation according to the identification source.
[0099] Step S53: Based on the peak amplitude density of each real load pulsation subsequence in the main waveform chain, calculate the proportion of internal unbiased points as the purity mark.
[0100] The purity mark quantifies the data quality of the real load pulsation subsequence, reflecting the proportion of effective sampling points in the subsequence. The peak amplitude density is defined as the number of effective sampling points in a unit of time in the subsequence, and the unbiased point refers to the sampling point that is not marked as a biased point.
[0101] For a real load fluctuation sub-sequence in the main waveform chain, the total number of sampling points (including real sampling points and virtual interpolation points) contained in the sub-sequence is counted, and the number of sampling points marked as deviation points is counted. The purity mark of the real load fluctuation sub-sequence is defined as the ratio of the number of sampling points not marked as deviation points to the total number of sampling points. The closer the purity mark is to 1, the higher the proportion of effective data in the sub-sequence, and the better the data quality.
[0102] Step S54: Attach a purity mark to each real load fluctuation sub-sequence, fuse into an edge purification waveform stream, and obtain a pre-processed data stream set.
[0103] Bind each real load fluctuation sub-sequence in the main waveform chain with its corresponding purity mark to form a basic unit of the edge purification waveform stream. The edge purification waveform stream is organized in a structured data stream form, and each unit contains five fields: {sub-sequence ID, start time sequence label, end time sequence label, waveform attribute point sequence, and purity mark}.
[0104] The pre-processed data stream set consists of two parts: 1. Edge purification waveform stream: contains all real load fluctuation sub-sequences and their purity marks, as the main output stream, for subsequent energy consumption analysis, load prediction, device monitoring, and other applications; 2. Interference trace isolation set: contains all abnormal fluctuation segments and their interference type labels, as an auxiliary output stream, for data quality audit, interference source localization, and other diagnostic applications.
[0105] The two parts of data are archived in the storage system of the edge node, the edge purification waveform stream is stored in the cache area, supporting real-time reading and streaming to upper-layer applications; the interference trace isolation set is stored in the low-speed storage area, for on-demand query and offline analysis.
[0106] It should be noted that the calculation method of the purity mark can be adjusted according to application requirements, such as introducing a weighted average of continuity weight, or considering the contribution of amplitude gradual smoothing degree; the storage format of the edge purification waveform stream can be compressed and encoded (such as Huffman coding) to further reduce the data volume; in other embodiments of the present application, the interference trace isolation set can also be attached with extended information such as abnormal severity score and possible interference source type inference to support more in-depth data quality analysis, which is not limited herein.
[0107] The embodiment obtains a transient sensitive paragraph sequence by locally adaptive anchoring the dynamic window boundary with the amplitude fluctuation rate, accurately locates the time domain interval that needs to be distinguished, and the pulse propagation path of the real load signal presents the characteristics of time sequence continuity and peak amplitude gradualness, while the transient interference signal presents the characteristics of path discontinuity and time sequence gap jumping due to the randomness of electromagnetic coupling crosstalk and switching operation transient, the continuity deviation of the pulse propagation path of the transient sensitive paragraph presents the degree of isolation of the interference signal, and the interference trace isolation set is obtained by isolating the discontinuous pulse subset, so as to reduce the dependence of real load feature analysis on the transient interference signal; according to the synchronization offset check result of the remaining part of the transient sensitive paragraph sequence between adjacent paragraphs and the amplitude gradualness smoothness of the remaining part of the transient sensitive paragraph sequence, the real load pulsation subsequence is adaptively distinguished, the continuity characteristics of a single paragraph are analyzed, the large range load periodic fluctuation is avoided, the probability that the transient interference signal is misidentified as the real load feature is reduced, so that the edge purification waveform flow based on separation and recombination and the purity mark thereof accurately reflect the real energy consumption characteristics of the power distribution system, improve the accuracy of load prediction, fault diagnosis and energy efficiency evaluation according to the data flow set completed by preprocessing under the condition of limited computing power of the edge node, and reduce the communication cost and delay overhead of data retransmission and cloud secondary processing caused by misjudgment.
[0108] Embodiment 2 Please refer to Figure 2 As shown in the figure, some parts of the embodiment are not described in detail, see the description of embodiment 1, and a power distribution energy consumption data preprocessing system based on edge computing is provided, comprising: A data acquisition module: collecting real-time current waveform data of a power distribution line through an edge node, and extracting the time sequence label and amplitude peak value of each sampling point to obtain an original waveform attribute set; A segmentation module: performing dynamic window sliding segmentation on the original waveform attribute set, wherein the window boundary is adaptively anchored by the local amplitude fluctuation rate to obtain a transient sensitive paragraph sequence; An interference screening module: for each paragraph in the transient sensitive paragraph sequence, tracking the continuity deviation of the internal pulse propagation path, and isolating the discontinuous pulse subset to obtain an interference trace isolation set; A marking module: based on the interference trace isolation set and the remaining part of the transient sensitive paragraph sequence, performing synchronization offset check between adjacent paragraphs, distinguishing and marking the real load pulsation subsequence; A recombination module: separating and recombining the real load pulsation subsequence and the interference trace isolation set to construct an edge purification waveform flow, and recording the purity mark of each subsequence to obtain a data flow set completed by preprocessing.
[0109] The above merely describes the preferred embodiments of the present application and is not used to limit the present application, and although the present application is described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions recorded in the foregoing embodiments, or equivalently replace some technical features thereof. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. An edge computing-based power consumption data preprocessing method, characterized in that, Comprise: Step S1: Collect real-time current waveform data of power distribution lines through edge nodes, and extract time sequence label and amplitude peak value of each sampling point to obtain original waveform attribute set; Step S2: Perform dynamic window sliding segmentation on the original waveform attribute set, wherein the window boundary is anchored by the local amplitude fluctuation rate to obtain a sequence of transient sensitive paragraphs; Step S3: For each paragraph in the sequence of transient sensitive paragraphs, track the continuity deviation of the internal pulse propagation path, and isolate the discontinuous pulse subset to obtain the interference trace isolation set; Step S4: Based on the interference trace isolation set and the remaining part of the sequence of transient sensitive paragraphs, perform synchronization offset checking between adjacent paragraphs, identify and mark the real load pulsation subsequence; Step S5: Separate and reorganize the real load pulsation subsequence and the interference trace isolation set to construct an edge purification waveform stream, and record the purity mark of each subsequence to obtain the preprocessed data stream set.
2. The edge computing based power consumption data preprocessing method of claim 1, wherein, Step S1 includes: Step S11: Use the current sensor built-in in the edge node to millisecond-level sample the power distribution line to capture continuous current waveform data; Step S12: Label the collection time of each point of the continuous current waveform data as a time sequence label, and calculate the peak amplitude of each point to construct an original waveform attribute set. 3.The edge computing based power consumption data preprocessing method of claim 2, wherein, Label the collection time of each point of the continuous current waveform data as a time sequence label, and calculate the peak amplitude of each point to construct an original waveform attribute set, comprising: Extract the absolute time sequence position of each sampling point from the continuous current waveform data to generate a time sequence label sequence; for each position in the time sequence label sequence, scan the amplitude extreme value in the microsecond interval before and after it to determine the peak amplitude; pair the time sequence label sequence with the peak amplitude one by one to fuse into a structured original waveform attribute set.
4. The edge computing based power consumption data preprocessing method of claim 1, wherein, Step S2 includes: Step S21: Initialize the starting width of the sliding window to a preset fixed millisecond pulse length; Step S22: Push the window along the original waveform attribute set point by point, and calculate the rolling variance of the peak amplitude in the window as the local amplitude fluctuation rate; Step S23: When the local amplitude fluctuation rate exceeds the preset anchor threshold, anchor the tail end of the current window as the segmentation boundary to generate a transient sensitive paragraph; Step S24: Repeat steps S22 to S23 until the entire original waveform attribute set is covered to obtain a sequence of transient sensitive paragraphs.
5. The edge computing based power consumption data pre-processing method of claim 1, wherein, Step S3 includes: Step S31: For each paragraph in the sequence of transient sensitive paragraphs, concatenate the peak amplitude along the time sequence label to obtain a simulated pulse propagation path; Step S32: Check the time interval between adjacent peak amplitudes in the simulated pulse propagation path, and if the interval exceeds the preset continuity threshold, mark it as a deviation point; Step S33: Expand a micro window around each deviation point to extract the internal isolated peak amplitude as a discontinuous pulse subset; Step S34: Remove the discontinuous pulse subset from the sequence of transient sensitive paragraphs to obtain the interference trace isolation set.
6. The edge computing based power consumption data preprocessing method according to claim 5, characterized in that, Check the time interval between adjacent peak amplitudes in the simulated pulse propagation path, and if the interval exceeds the preset continuity threshold, mark it as a deviation point, comprising: The time sequence gap is calculated by selecting each pair of adjacent peak amplitudes on the simulated pulse propagation path, and the time sequence gap is compared with the continuity threshold preset based on the line impedance. If the time sequence gap exceeds the range of the continuity threshold, the pair is marked as a deviation point on the simulated pulse propagation path. 7.The edge computing based power consumption data preprocessing method of claim 1, wherein, Step S4 includes: Step S41: selecting the remaining part of each two adjacent paragraphs in the transient sensitive paragraph sequence, and extracting the peak amplitudes at the tail end and the head end as anchor points; Step S42: calculating the time sequence label offset between the anchor points, and checking whether the offset conforms to the preset load cycle synchronization band; Step S43: if the offset falls within the synchronization band, the remaining parts of the two adjacent paragraphs are combined into a continuous pulsation unit; Step S44: the amplitude gradual smoothness of all continuous pulsation units is discriminated one by one, and if the amplitude gradual smoothness is higher than the preset interference threshold, the real load pulsation subsequence is marked.
8. The edge computing based power consumption data preprocessing method according to claim 7, characterized in that, The amplitude gradual smoothness of all continuous pulsation units is discriminated one by one, and if the amplitude gradual smoothness is higher than the preset interference threshold, the real load pulsation subsequence is marked, including: For each continuous pulsation unit, the cumulative coefficient of variation of adjacent amplitude difference is calculated along the peak amplitude sequence as the amplitude gradual smoothness; the amplitude gradual smoothness is compared with the preset interference threshold; if the amplitude gradual smoothness exceeds the preset interference threshold, the continuous pulsation unit is immediately marked as a real load pulsation subsequence. 9.The edge computing based power consumption data preprocessing method of claim 1, wherein, Step S5 includes: Step S51: the real load pulsation subsequences are concatenated in the original time sequence label order to generate a main waveform chain; Step S52: the interference trace isolation set is placed in an independent buffer area, which is physically separated from the main waveform chain; Step S53: based on the peak amplitude density of each real load pulsation subsequence in the main waveform chain, the proportion of internal non-deviation points is calculated as a purity mark; Step S54: a purity mark is attached to each real load pulsation subsequence to fuse into an edge purification waveform stream, and a preprocessed data stream set is obtained. 10.The edge computing based power consumption data preprocessing method of claim 8, wherein, The real load pulsation subsequences are concatenated in the original time sequence label order to generate a main waveform chain, including: For each subsequence in the real load pulsation subsequence, the head and tail peak amplitudes are selected as bridge anchor points, and the potential jump risk is estimated based on the amplitude difference between the anchor points to generate a bridge anchor point risk set; based on the bridge anchor point risk set, the tail end anchor point and the head end anchor point of adjacent subsequences are bridged in the time sequence label order to generate a micro gradual transition segment; based on the micro gradual transition segment, the continuity weight of the micro gradual transition segment is adjusted node by node to form a preliminary gap-free ring link draft, which is temporarily solidified in the memory buffer of the edge node; based on the preliminary gap-free ring link draft, a lightweight hash check is performed on each concatenated node to verify the uniqueness of the time sequence label and lock the link integrity level by level, thereby iteratively optimizing the generation of the final main waveform chain.
Citation Information
Patent Citations
High-frequency discharge signal identification method
CN120103088A
AI intelligent communication data processing method and system based on edge computing
CN120475383A
Electricity utilization information acquisition equipment of electric energy meter metering pulse simulation system
CN120633110A
Power distribution station overload risk alarm method and device based on multi-source sensing data fusion
CN120810951A
Abnormal mode data processing system driven by power marketing big data
CN120930032A