Power grid data acquisition and analysis system based on big data
Through technical means such as standardized filling, time label alignment, dynamic feature screening, spatio-temporal correlation weights and core density estimation, the problem of cross-level semantic distortion in the power grid data acquisition system is solved, and the accuracy of grid fault positioning and scheduling reliability are achieved.
Patent Information
- Application Number
- CN202510694603.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-05-28
AI Technical Summary
Due to the lack of cross-level data semantic correlation processing of existing power grid data, key operating features are irreversiblely lost in step-by-step compression, affecting the accuracy of fault location and reliability of scheduling decisions.
Standardized filling and hierarchical alignment of time tags are performed through the device acquisition module, the feature extraction module filters dynamic feature sets, the semantic association module generates spatiotemporal association weights based on regional physical topological relationships and reactive circulation path sensitivity, the feature clustering module combines kernel density estimation and spatial weighting, and the cross-layer analysis module corrects dynamic aggregate weights, and finally generates device positioning instructions and scheduling strategies.
It realizes the deep coupling of multi-level data fusion and the physical characteristics of the power grid, improves the accuracy of fault positioning and scheduling reliability, ensures the consistency of data characteristics and electrical laws, reduces the dependence of manual intervention, and realizes the automation and accuracy of fault positioning and scheduling decisions.
Smart Images

Figure CN120498050A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power network monitoring, and more specifically, to a power grid data acquisition and analysis system based on big data. Background Art
[0002] Power grid operations require data collaboration at the device, regional, and network levels to enable condition monitoring and dispatch decisions. Existing power grid data collection systems typically employ a hierarchical aggregation mechanism, compressing raw data collected by underlying devices at each level before transmitting it to upper-level analysis platforms to reduce transmission and storage costs. This model relies on pre-set statistical rules (such as mean calculation and threshold filtering) to reduce the dimensionality of multi-source data, generating aggregated datasets at different levels for system access.
[0003] Existing hierarchical aggregation mechanisms, during the cross-level data fusion process, lack adaptive processing of the semantic associations of raw data, leading to the irreversible loss of key operational characteristics during the progressive compression process. For example, high-frequency transient signals at the device level are masked by averaging during regional aggregation, while regional load trend statistics cannot be traced back to specific device anomalies. This prevents upper-level analytical models from accurately linking causal logic between cross-level data, directly impacting fault location accuracy and the reliability of scheduling decisions. Summary of the Invention
[0004] In order to overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides a power grid data acquisition and analysis system based on big data to solve the problems raised in the above-mentioned background technology.
[0005] To achieve the above object, the present invention provides the following technical solutions: A power grid data acquisition and analysis system based on big data, comprising: Equipment acquisition module: obtains raw data from power grid equipment and performs standardized filling and time label layer alignment processing to generate standardized equipment-level operating data; Feature extraction module: Filters dynamic feature sets from standardized device-level operating data. The dynamic feature sets contain abnormal waveform segments and steady-state parameter offsets within a preset period. Semantic association module: determines whether the dynamic feature set triggers a cross-device association event. If so, it generates the spatiotemporal association weight of the dynamic feature set based on the regional physical topology and reactive circulation path sensitivity. Feature clustering module: Based on the spatiotemporal correlation weights and the device-level feature distribution density within a predefined time period, the dynamic feature sets within a preset correlation radius are combined and reconstructed into regional-level feature aggregation packages; Cross-layer parsing module: Based on the current flow direction and transient energy distribution characteristics between the devices in the regional feature aggregation package, the dynamic aggregation weight of the regional feature aggregation package is corrected; Decision generation module: Based on the corrected dynamic aggregation weight and regional feature aggregation package, it matches the preset alarm level template and outputs a combination of device positioning instructions and regional scheduling strategies.
[0006] In a preferred embodiment, the raw data of the power grid equipment side is obtained and subjected to standardization filling and time label hierarchical alignment processing to generate standardized device-level operation data, including: Based on the preset filling rules corresponding to the device type, the original data obtained from the power grid equipment end is standardized and filled with missing values; Determine the hierarchical time tag alignment strategy based on the device type and signal sampling frequency. For power frequency devices, align time tags based on the cycle period. For high-frequency sampling devices, align time tags based on the microsecond time window. The padded and aligned data are integrated into standardized device-level operating data according to a unified time axis.
[0007] In a preferred embodiment, screening the dynamic feature set from the standardized device-level operating data includes: Abnormal waveform segments are extracted based on a dynamic sliding window detection method. The sliding window period is consistent with the preset period corresponding to the device type. When the fluctuation amplitude of the data within the window exceeds the dynamic threshold, it is marked as an abnormal waveform segment. The dynamic threshold is adjusted based on the average peak-to-valley difference of the normal waveform in the device's historical operating data. The steady-state parameter offset is determined based on the steady-state baseline calculation method. The steady-state baseline value of each time stamp within a preset period is extracted from the standardized device-level operation data. The steady-state baseline value is the average of the historical operation data of the same device in the period without abnormalities. The absolute deviation between the current data and the steady-state baseline value is calculated as the offset. When the offset exceeds the preset offset threshold, it is determined to be a valid steady-state parameter offset.
[0008] In a preferred embodiment, determining whether a dynamic feature set triggers a cross-device association event, and if so, generating a spatiotemporal association weight of the dynamic feature set based on the regional physical topology and the sensitivity of the reactive circulating current path, includes: Whether the cross-device correlation event triggering conditions are met is determined based on the temporal overlap between the electrical distance between devices and the dynamic feature set. The electrical distance is calculated based on the shortest connection path length between devices in the regional physical topology. The temporal overlap is the time window overlap ratio of the abnormal waveform segment or steady-state parameter offset. If a cross-device association event is triggered, the spatiotemporal association weight is adjusted according to the reactive circulation path sensitivity. The reactive circulation path sensitivity is calculated by the ratio of the reactive circulation mutual information entropy between devices to the topological distance. The initial value of the spatiotemporal association weight is the inverse of the electrical distance between devices. The final spatiotemporal association weight is the initial value of the spatiotemporal association weight multiplied by the reactive circulation path sensitivity correction coefficient.
[0009] In a preferred embodiment, based on the spatiotemporal correlation weight and the device-level feature distribution density within a predefined time period, the dynamic feature sets within a preset correlation radius are combined and reconstructed into a regional-level feature aggregation package, including: The device-level feature distribution density is calculated based on the kernel density estimation method. The kernel function bandwidth is adjusted according to the device spatial distribution sparsity. The device spatial distribution sparsity is determined by the ratio of the average distance between devices to the preset association radius. Generate a regional feature density distribution map based on the device-level feature distribution density, and rasterize the regional feature density distribution map through a heat map; Extract high-density areas from the regional feature density distribution map based on a preset correlation radius, where the high-density areas are areas where the density values of continuous grid cells exceed a preset density threshold; The dynamic feature set in the high-density area is weightedly fused according to the spatiotemporal correlation weight and reconstructed into a regional feature aggregation package. The weighted fusion method is to multiply the dynamic feature value by the spatiotemporal correlation weight of the corresponding device, accumulate it, and then divide it by the total weight sum.
[0010] In a preferred embodiment, the grid unit density value is the weighted sum of the device-level feature distribution density within the corresponding spatial range, the weight is the spatiotemporal correlation weight of the device in the corresponding unit, and the preset density threshold is determined based on the statistical distribution of density values in historical regional-level feature aggregation events.
[0011] In a preferred embodiment, based on the current flow direction and transient energy distribution characteristics between the devices in the regional-level feature aggregation package, the dynamic aggregation weight of the regional-level feature aggregation package is corrected, including: The current flow direction is determined by the current phase difference and voltage gradient direction between devices; Extract the transient energy distribution features from the regional feature aggregation package. The transient energy distribution features are decomposed by wavelet packets to extract the frequency band energy ratio of each device's transient waveform. The frequency band energy ratio is the ratio of the energy of a specified frequency band to the total energy. Adjust the initial distribution ratio of dynamic aggregation weights between devices based on the direction of current flow. The adjustment range is a function of the flow direction and the consistency of the topological path. The adjusted weight is normalized and corrected based on the transient energy distribution characteristics. The correction coefficient is the ratio of the equipment transient energy proportion to the regional average transient energy proportion. The final dynamic aggregation weight is the adjusted weight multiplied by the correction coefficient.
[0012] In a preferred embodiment, the current phase difference is the phase offset angle between the device voltage and current waveforms at the same timestamp, and the voltage gradient direction is the changing trend of the voltage amplitude between devices along the topological path.
[0013] In a preferred embodiment, according to the modified dynamic aggregation weight and the regional feature aggregation package, a preset alarm level template is matched and a device positioning instruction and regional scheduling strategy combination is output, including: Based on the corrected dynamic aggregation weight and the spatial distribution density of devices in the regional feature aggregation package, the corresponding alarm level in the preset alarm level template is matched. The alarm level template is divided by the joint distribution threshold of the dynamic aggregation weight and the spatial density of devices in historical fault events; Generate device location instructions based on the alarm level. The device location instructions contain the spatial coordinates and associated topological paths of the device with high dynamic aggregation weight. The spatial coordinates are calculated by the weighted centroid of the geographic range in the regional feature aggregation package and the device-level feature distribution density. Generate regional scheduling strategy combinations based on alarm levels and device positioning instructions.
[0014] In a preferred embodiment, the regional scheduling strategy combination includes load reduction ratio, voltage adjustment priority and backup line switching order. The load reduction ratio is dynamically allocated according to the power proportion of high dynamic aggregation weight equipment, and the voltage adjustment priority is sorted by the equipment voltage gradient direction and the consistency of the flow path.
[0015] Compared with the prior art, the present invention has the following beneficial effects: 1. Through the deep coupling of multi-level data fusion and the physical characteristics of the power grid, this approach effectively addresses the issues of key feature loss and cross-level semantic distortion in traditional hierarchical aggregation mechanisms. The device acquisition module's standardized padding and hierarchical alignment of time tags ensures consistent time bases and format compatibility for heterogeneous data from multiple sources, providing high-precision input for subsequent feature extraction. A dynamic feature screening mechanism adaptively extracts abnormal waveform segments and steady-state offsets based on device type and real-time fluctuation amplitude, avoiding the masking of high-frequency transient signals by traditional mean or threshold filtering. The semantic association module maps discrete device features to the actual grid connection network through a joint analysis of regional physical topology and reactive circulating current path sensitivity, establishing cross-device causal relationships and addressing the inability of manual rules to capture implicit electrical interactions. The feature clustering module combines spatiotemporal weights and distribution density to preserve device-level details when reconstructing regional data. High-density region extraction technology also highlights key operating states, significantly improving the physical interpretability of data aggregation. The cross-layer parsing module further incorporates current flow direction and transient energy distribution characteristics, dynamically adjusting weights to match real-time grid conditions and ensuring consistency between data features and electrical laws. The dispatching instructions finally generated are directly linked to the equipment location and regional operating status, forming a closed-loop link from data collection to decision execution, greatly improving the fault location accuracy and dispatching reliability.
[0016] 2. While improving data semantic consistency, it also takes into account the real-time and adaptability of dynamic grid operations. Standardized device-level operational data is processed through time tag alignment and differentiated filling rules, laying the foundation for multi-source data fusion and avoiding feature correlation bias caused by time misalignment in traditional methods. Dynamic feature set screening rules dynamically adjust thresholds based on historical device operating conditions and real-time fluctuations, replacing fixed threshold filtering, making anomaly detection more tailored to actual grid operation scenarios. The generation of spatiotemporal correlation weights integrates reactive current path sensitivity with topological physical connectivity, quantifying the implicit electrical impact between devices and enabling cross-device correlation analysis to combine mathematical statistical rigor with grid physical regularity. Regional-level feature aggregation packages utilize kernel density estimation and spatial weighting to preserve device-level feature details while reflecting the overall trend of regional operating conditions, addressing the irreversible feature loss caused by traditional hierarchical compression. Dynamic aggregation weight correction incorporates transient energy distribution and flow direction to ensure that data aggregation results match grid dynamics in real time, avoiding decision lags caused by static weights. The matching mechanism between alarm level templates and scheduling strategies is based on a joint analysis of dynamic weights and spatial density, directly mapping data features to control instructions, reducing reliance on manual intervention and achieving automation and precision in fault location and scheduling decisions. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 This is a structural diagram of a power grid data acquisition and analysis system based on big data in the present invention. DETAILED DESCRIPTION
[0018] The following will provide a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0019] Example: Figure 1 The present invention provides a structural diagram of a power grid data acquisition and analysis system based on big data, which includes the following modules: Equipment acquisition module: obtains raw data from power grid equipment and performs standardized filling and time label layer alignment processing to generate standardized equipment-level operating data; Feature extraction module: Filters dynamic feature sets from standardized device-level operating data. The dynamic feature sets contain abnormal waveform segments and steady-state parameter offsets within a preset period. Semantic association module: determines whether the dynamic feature set triggers a cross-device association event. If so, it generates the spatiotemporal association weight of the dynamic feature set based on the regional physical topology and reactive circulation path sensitivity. Feature clustering module: Based on the spatiotemporal correlation weights and the device-level feature distribution density within a predefined time period, the dynamic feature sets within a preset correlation radius are combined and reconstructed into regional-level feature aggregation packages; Cross-layer parsing module: Based on the current flow direction and transient energy distribution characteristics between the devices in the regional feature aggregation package, the dynamic aggregation weight of the regional feature aggregation package is corrected; Decision generation module: Based on the corrected dynamic aggregation weight and regional feature aggregation package, it matches the preset alarm level template and outputs a combination of device positioning instructions and regional scheduling strategies.
[0020] First, the device acquisition module acquires raw data from power grid devices and generates standardized device-level operational data through standardized padding and hierarchical alignment of time tags, addressing time base and format compatibility issues for heterogeneous multi-source data. Subsequently, the feature extraction module selects dynamic feature sets containing anomalous waveform segments and steady-state parameter offsets from the standardized data based on device type, preset operating condition rules, and real-time fluctuation amplitude, ensuring that key operational features are not obscured by averaging during data dimensionality reduction. This dynamic feature set is fed into the semantic association module, which triggers cross-device association events by determining the current phase difference and voltage gradient direction between devices. This module then generates spatiotemporal association weights based on regional physical topology and reactive circulating current path sensitivity, linking discrete device features into a topological network with electrical and physical meaning. The feature clustering module uses kernel density estimation to generate regional feature aggregation packages based on the spatiotemporal association weights and device-level feature distribution density. Using heat map rasterization and high-density region extraction, the device-level dynamic features are reconstructed into structured data reflecting the regional operational status. The cross-layer analysis module further normalizes the dynamic aggregation weights in the aggregation package using weight correction coefficients based on current flow direction and transient energy distribution characteristics, achieving deep coupling between electrical parameters and data characteristics. Finally, the decision generation module matches the modified weights with the preset alarm level template in the aggregation package and outputs a combined instruction containing the spatial coordinates of the device, the topological path, and the load scheduling strategy, forming a closed-loop link from data acquisition to decision execution.
[0021] The system solves the cross-level semantic distortion problem in existing technologies through a multi-level data fusion mechanism. The standardized filling and time alignment processing of the equipment acquisition module eliminates time misalignment and format conflicts in multi-source data, laying the foundation for subsequent feature association. The feature extraction module dynamically screens abnormal waveforms and steady-state offsets to avoid the loss of high-frequency transient signals caused by traditional statistical rules. The semantic association module introduces the sensitivity of reactive circulating current paths and regional physical topology relationships to map device-level features to the real physical connection network of the power grid, solving the defect that artificial rules cannot associate cross-device causal logic. The feature clustering module uses kernel density estimation and spatial weighted fusion to retain device-level detailed features while generating regional-level aggregated data that can be reversely parsed. The cross-layer parsing module couples current flow direction and transient energy distribution, corrects weights to reflect real-time electrical status, and improves the physical consistency of aggregated data. The decision generation module generates positioning and dispatching instructions based on dynamic weights and alarm templates, directly linking data features with power grid control actions, ultimately achieving a dual improvement in fault location accuracy and dispatching reliability.
[0022] Obtain the raw data from the power grid equipment side and perform standardization padding and time label hierarchical alignment to generate standardized device-level operating data. The specific implementation is as follows: When performing standardized filling of missing values on the original data at the power grid equipment end, the operation is performed according to the preset filling rules corresponding to the equipment type. The preset filling rules are defined by a mapping relationship table between the equipment type and the missing value processing logic. For example, the missing values of the current transformer equipment are filled by linear interpolation of adjacent timestamp data, and the missing values of the voltage transformer equipment are filled by the mean of the historical data of the same equipment during the same period. The specific method of linear interpolation filling is: if the current data of a certain timestamp is missing, the current value of the valid timestamp before the timestamp and the current value of the valid timestamp after the timestamp are taken for arithmetic average, and the average value is used as the filling value for the missing timestamp. The mean filling method for the historical data of the same equipment during the same period is: extract the voltage data of all normal dates in the same period from the historical database of the voltage transformer equipment, calculate the average value of the voltage value of each timestamp, and use the average value as the filling value for the missing timestamp. The interval between adjacent timestamps in the preset filling rule is set according to the maximum allowable data interruption duration corresponding to the device type. For example, the maximum allowable data interruption duration of the current transformer is 10 seconds. If the interval between valid data before and after the missing timestamp exceeds 10 seconds, it is judged as invalid filling and marked as abnormal data.
[0023] The mapping table contains four fields: device type number, device type name, maximum allowed interruption duration, and filling method. For example, the filling method for a current transformer with device type number CT001 is linear interpolation of adjacent timestamps, with a maximum allowed interruption duration of 10 seconds. The filling method for a voltage transformer with device type number VT002 is historical average filling, with a maximum allowed interruption duration of 30 seconds.
[0024] When performing hierarchical timestamp alignment on padded data, the hierarchical timestamp alignment strategy is determined based on the device type and signal sampling frequency. Device types are categorized into two types: power frequency devices and high-frequency sampling devices. Power frequency devices include equipment with a 50 Hz power frequency signal as the reference, such as transformers and circuit breakers. High-frequency sampling devices include equipment with a sampling frequency higher than 1 kHz, such as synchronized phasor measurement devices and traveling wave ranging devices. The time tag alignment strategy for power frequency devices is based on cycle period. Specifically, the original timestamps are rounded to a 20 ms cycle, and data within the same cycle are assigned the same timestamp. For example, data with timestamps of 15 ms and 35 ms are aligned to timestamps of 0 ms and 20 ms, respectively. The time tag alignment strategy for high-frequency sampling devices is based on microsecond time windows. Specifically, the original timestamps are rounded to integer multiples of 1 μs, using a 1 μs time window as the reference. For example, data with a timestamp of 1234567.89 μs is aligned to a timestamp of 1234568 μs. In the time tag hierarchical alignment strategy, the reference time error for cycle alignment does not exceed ±0.1ms, and the reference time error for microsecond time window alignment does not exceed ±0.01μs. This error range is set according to the requirements of power system synchronous phasor measurement and will not be repeated here.
[0025] When padded and aligned data is consolidated into standardized device-level operational data along a unified timeline with a time resolution of 1μs, all device data time tags are converted to absolute timestamps within this timeline. For example, a cycle-aligned time tag of 0ms for a power-frequency device is converted to 0μs on the unified timeline, while a microsecond-aligned time tag of 1234568μs for a high-frequency sampling device is retained as 1234568μs on the unified timeline. During the consolidation process, if data from multiple devices with the same timestamp exists, they are stored in a multidimensional array in ascending order by device number. The ascending order for device numbers prioritizes sorting by device type number from smallest to largest, and for devices of the same type, by installation location number from smallest to largest. The standardized device-level operational data format consists of a structured table containing timestamps, device numbers, and data values. The timestamps have a precision of 1μs, the device numbers are globally unique identifiers, and the data values are floating-point numbers that have undergone standardized padding and alignment.
[0026] Filtering dynamic feature sets from standardized device-level operational data is implemented by the following steps: When extracting abnormal waveform segments using the sliding window dynamic detection method, the sliding window period is set based on the preset period corresponding to the device type. Device types include current transformers, voltage transformers, and traveling wave ranging devices. The sliding window period for different device types is defined using a mapping table between device type numbers and preset periods. This mapping table contains three fields: device type number, device type name, and preset period. For example, the preset period corresponding to a current transformer with device type number CT001 is 1 minute, the preset period corresponding to a traveling wave ranging device with device type number TW003 is 1 second, and the preset period corresponding to a voltage transformer with device type number VT005 is 10 seconds. The specific operation of the sliding window dynamic detection method is as follows: using the sliding window period as the time length and half the window period as the sliding step size, sequentially intercept consecutive segments of standardized device-level operating data and calculate the fluctuation amplitude of the data within each window. The fluctuation amplitude is calculated as the difference between the maximum and minimum values within the window. For example, if the maximum current within a current transformer window is 100A and the minimum current is 80A, the fluctuation amplitude is 20A. The dynamic threshold is adjusted based on the average peak-to-valley difference of the normal waveform in the device's historical operating data. If the historical data is less than 30 days, the dynamic threshold is calculated based on the actual number of available days and multiplied by 1.5. For example, if a current transformer has only 15 days of historical data and the average peak-to-valley difference is 12A, the dynamic threshold is set to 18A. When the fluctuation amplitude within a window exceeds the dynamic threshold, the data within that window is marked as an abnormal waveform segment by the abnormal marking unit. The data input to the abnormal marking unit is the window fluctuation amplitude and the dynamic threshold, and the data output is the time range record of the abnormal waveform segment.
[0027] When determining the steady-state parameter offset using the steady-state baseline calculation method, the steady-state baseline value is obtained by filtering data from anomaly-free periods from the historical operating data of the same device. The criteria for anomaly-free periods are that the historical data contains no abnormal waveform segments and the fluctuation amplitude is less than 80% of the dynamic threshold. For example, if the dynamic threshold of a current transformer is 22.5A, the fluctuation amplitude during anomaly-free periods must be less than 18A. The historical data from the anomaly-free periods is aligned to the timeline of the standardized device-level operating data based on the timestamp. The alignment rule is to ignore the year information and only retain the month, day, hour, minute, second, and microsecond information. For example, if the historical data has a timestamp of 08:00:00.000000 on October 1, 2022, it is aligned to 08:00:00.000000 on October 1, 2023. The steady-state baseline value corresponding to each timestamp is calculated. The steady-state baseline value is the average of the historical values for the same timestamp during the anomaly-free periods. For example, if the current value at the timestamp 08:00:00.000000 on October 1, 2023, during a historical period without anomalies is 95A, then the steady-state baseline value is 95A. When calculating the absolute deviation between the current data and the steady-state baseline value, the absolute deviation is calculated as the absolute value of the difference between the current data value and the steady-state baseline value. For example, if the current value at the current timestamp is 105A and the steady-state baseline value is 95A, the absolute deviation is 10A. The preset offset threshold is determined based on the statistical distribution of the device's historical offset values. Specifically, the 95th percentile of the absolute deviations at each timestamp during a historical period without anomalies is calculated and used as the preset offset threshold. For example, if the 95th percentile of the historical absolute deviations is 8A, the preset offset threshold is set to 8A. When the absolute deviation exceeds the preset offset threshold, the offset at that timestamp is considered a valid steady-state parameter offset. The valid steady-state parameter offset is recorded as the timestamp, current data value, steady-state baseline value, and offset.
[0028] Abnormal waveform segments and valid steady-state parameter offsets together constitute a dynamic feature set. The data structure of the dynamic feature set is a structured table containing timestamps, feature types, and feature values. Feature types are divided into two categories: abnormal waveform segments and steady-state parameter offsets. For example, the feature type of an abnormal waveform segment is marked as "waveform anomaly," and the feature values are the window start timestamp, end timestamp, and fluctuation amplitude; the feature type of a steady-state parameter offset is marked as "steady-state offset," and the feature values are the timestamp, current data value, steady-state baseline value, and offset. The timestamp accuracy of the dynamic feature set is consistent with standardized device-level operating data, both at 1 microsecond.
[0029] Determine whether the dynamic feature set triggers a cross-device association event. If so, generate the spatiotemporal association weight of the dynamic feature set based on the regional physical topology and reactive circulating current path sensitivity. The specific implementation is as follows: When determining whether the triggering conditions for cross-device correlation events are met based on the temporal overlap between the electrical distance between devices and the dynamic feature set, the electrical distance is calculated based on the shortest path length between the devices in the regional physical topology. The regional physical topology is defined by a grid connection topology graph. The topology graph is constructed as follows: each power device is mapped as a node in the graph, and the physical connections between devices (such as transmission lines and switches) are mapped as undirected edges between the nodes. The edge weight is the physical length of the connection path or its equivalent impedance. For example, if a transformer is mapped as node T1 and a circuit breaker is mapped as node B2, and they are connected by a transmission line, an edge with an edge weight of 1 is added between nodes T1 and B2 (indicating a single connection segment). The shortest path length is calculated using the Dijkstra algorithm. For example, if devices A and B are connected via two switches and a transmission line, the shortest path length is 3. Temporal overlap is calculated based on the overlap ratio of the time window of abnormal waveform segments or steady-state parameter offsets. The time window is the range from the start timestamp to the end timestamp of the abnormal waveform segment or steady-state parameter offset. For example, if the abnormal waveform time window for device A is 08:00:00 to 08:00:10, and the abnormal waveform time window for device B is 08:00:05 to 08:00:15, the overlapping time window is 08:00:05 to 08:00:10, and the overlap ratio is 50% (the length of the overlapping window (5 seconds) divided by the length of the shorter window (10 seconds). When the electrical distance is less than the preset electrical distance threshold and the time overlap is greater than the preset overlap ratio threshold, the trigger condition for a cross-device correlation event is determined. The preset electrical distance threshold is determined based on the statistical distribution of electrical distances between devices in historical correlation events. For example, the 75th percentile of historical electrical distances is used as the threshold. If the historical electrical distance data is [2, 3, 4, 5], the 75th percentile is 4.25, and the threshold is set to 4. The preset overlap ratio threshold is set empirically, for example, to 30%.
[0030] If a cross-device correlation event is triggered, the spatiotemporal correlation weight is adjusted based on the reactive circulation path sensitivity. Reactive circulation path sensitivity is calculated as the ratio of the mutual information entropy of the reactive circulation between devices to the topological distance. The mutual information entropy of the reactive circulation between devices is calculated by extracting the reactive power sequence between devices from historical reactive circulation data. The reactive power values are discretized into multiple intervals according to a preset bin width. The joint frequency of the reactive power values of the two devices in the same interval is counted to calculate the joint probability distribution. The independent frequency of the reactive power values of each device in each interval is counted to calculate the independent probability distribution. The expected logarithm of the ratio of the joint probability distribution to the independent probability distribution is calculated by multiplying the joint probability value in each interval by the natural logarithm of the ratio of the joint probability to the independent probability in that interval. The calculated results across all intervals are then summed to obtain the mutual information entropy. For example, the reactive power sequences of devices A and B are discretized into 10 intervals. The probability distribution is calculated after calculating the joint frequency matrix. The final entropy value is 0.8 bits. The topological distance is the length of the shortest connection path between devices in the region's physical topology. For example, the topological distance between device A and device B is 3. The reactive circulating current path sensitivity is the mutual information entropy divided by the topological distance. For example, 0.8 / 3 ≈ 0.267. The initial value of the spatiotemporal association weight is the inverse of the electrical distance between devices. For example, if the electrical distance between device A and device B is 3, the initial weight is 1 / 3 ≈ 0.333. The final spatiotemporal association weight is the initial value multiplied by the reactive circulating current path sensitivity correction factor. For example, 0.333 × 0.267 ≈ 0.089.
[0031] The spatiotemporal correlation weight calculation results are recorded in a structured table containing the device pair number, initial weight, correction factor, and final weight. For example, for device pair AB001, the initial weight for devices A and B is 0.333, the correction factor is 0.267, and the final weight is 0.089. The timestamp accuracy of the structured table is consistent with that of the dynamic feature set: 1 microsecond. When there is no physical connection between devices, the electrical distance is set to 100,000 (indicating infinity), and the initial spatiotemporal correlation weight is set to 0. If there is a physical connection between devices but the topological distance exceeds the preset electrical distance threshold (for example, 100), the initial spatiotemporal correlation weight is also set to 0. If the temporal overlap is 0, the cross-device correlation event is not triggered. The preset bin width is dynamically adjusted based on the historical reactive power range. For example, if the historical reactive power range is -10 MVar to +10 MVar, the bin width is set to 2 MVar, divided into 10 intervals.
[0032] Based on the spatiotemporal correlation weights and the device-level feature distribution density within a predefined time period, the dynamic feature sets within the preset correlation radius are combined and reconstructed into a regional-level feature aggregation package. The specific implementation is as follows: When calculating device-level feature distribution density using kernel density estimation, the kernel function bandwidth is adjusted based on the device spatial distribution sparsity. Sparsity is calculated as the ratio of the average inter-device distance to the preset association radius, where the average inter-device distance is the mean Euclidean distance between all pairs of devices in the region. For example, if the preset association radius is 500 meters and the average inter-device distance is 100 meters, the sparsity is 100 / 500 = 0.2. The kernel function bandwidth is adjusted by multiplying the inverse of the sparsity by the base bandwidth, which is determined based on the historical data distribution range for the device type. For example, if the base bandwidth of a current transformer is 50 meters and the sparsity is 0.2, the adjusted bandwidth is 50 × (1 / 0.2) = 250 meters. Kernel density estimation is calculated using a Gaussian kernel function, whose mathematical form is a negative exponential relationship between height and distance from the device to the grid center. The Gaussian kernel height is determined by the device's spatiotemporal association weight. For example, if device A is 100 meters from the grid center and the adjusted bandwidth is 250 meters, the Gaussian kernel value is the spatiotemporal correlation weight 0.3 multiplied by exp(-(100²) / (2×250²)), which is approximately 0.3×0.939=0.282.
[0033] When generating a regional feature density distribution map based on device-level feature distribution density, the map is rasterized using a heatmap. The grid cell size of the heatmap is one-tenth of the preset correlation radius. For example, if the preset correlation radius is 500 meters, the grid cell size is 50 meters x 50 meters. The data format of the regional feature density distribution map is a two-dimensional matrix. The matrix row index represents the latitudinal coordinate, and the column index represents the longitude coordinate. The origin is located at the lower left corner of the region. The grid cells are arranged in ascending longitude, and the row index is arranged in ascending latitude. For example, if the origin coordinates are (100.00° East longitude, 30.00° North latitude) and the grid cell size is 50 meters x 50 meters, an increase of 1 in the row index corresponds to an increase of 50 meters in latitude to the north, and an increase of 1 in the column index corresponds to an increase of 50 meters in longitude to the east. The density value of each grid cell is the weighted sum of the kernel density estimates of all devices within the corresponding spatial range, with the weight being the spatiotemporal correlation weight of the devices within that grid cell. For example, a grid cell covers device A (weight 0.3) and device B (weight 0.5). The kernel density estimate of device A is 0.8, and that of device B is 1.2. Then the grid density value is 0.3×0.8+0.5×1.2=0.84.
[0034] When extracting high-density areas from a regional feature density distribution map based on a preset correlation radius, high-density areas are defined as regions where the density values of consecutive grid cells exceed a preset density threshold. The preset density threshold is determined by the statistical distribution of density values in historical regional feature aggregation events. For example, the 80th percentile of the historical density values is used as the threshold. If the historical density values are [0.5, 0.8, 1.0, 1.2], the 80th percentile is 1.0, and the threshold is set to 1.0. The criterion for determining consecutive grid cells is that the density values of adjacent grid cells (including the upper, lower, left, right, and diagonal directions) all exceed the threshold. For example, if a grid cell has a density value of 1.1 and five of its eight adjacent grid cells have values exceeding 1.0, it is considered part of a high-density area. If no grid cells in the regional feature density distribution map exceed the threshold, it is determined that there is no high-density area in the current time period, and the reconstruction step is skipped.
[0035] When the dynamic feature set within a high-density area is weighted and fused based on spatiotemporal correlation weights, the dynamic feature set includes abnormal waveform segments and steady-state parameter offsets. Weighted fusion is performed by multiplying the dynamic feature values by the corresponding device's spatiotemporal correlation weights, summing the values, and then dividing by the total weighted sum. For example, if a high-density area contains device A (weight 0.3), device B (weight 0.5), and device C (weight 0.2), and the abnormal waveform amplitude for device A is 10, for device B it is 15, and for device C it is 8, then the weighted fusion value is (10 × 0.3 + 15 × 0.5 + 8 × 0.2) / (0.3 + 0.5 + 0.2) = 12.1. The reconstructed regional feature aggregation package consists of a structured record containing a timestamp, geographic range, and weighted fusion feature values. The timestamp precision is consistent with the dynamic feature set (1 microsecond), and the geographic range is the coordinates of the minimum bounding rectangle of the high-density area. The minimum bounding rectangle coordinates of a high-density area are determined by the following steps: traversing the longitude and latitude coordinates of all grid cells within the high-density area, recording the minimum latitude, maximum latitude, minimum longitude, and maximum longitude. The coordinates of the minimum bounding rectangle vertices are (minimum longitude, minimum latitude) and (maximum longitude, maximum latitude). For example, if a high-density area covers longitudes 100.00° to 100.05° and latitudes 30.00° to 30.03°, the minimum bounding rectangle vertices are (100.00°, 30.00°) and (100.05°, 30.03°). If the high-density area contains multiple discrete sub-areas, a separate feature aggregation package is generated for each sub-area.
[0036] When the device spatial distribution sparsity is 0 (i.e., all devices are concentrated at the same point), the kernel function bandwidth is automatically set to 1% of the preset correlation radius (e.g., 500 meters x 1% = 5 meters) to prevent density estimation distortion. If the preset correlation radius exceeds the actual device distribution range (e.g., if the maximum device spacing is 300 meters and the correlation radius is set to 500 meters), the correlation radius is automatically adjusted to 1.2 times the maximum device spacing (300 x 1.2 = 360 meters). During weighted fusion, if the total weight sums to 0 (e.g., if all device weights abnormally return to zero), the weighted fusion is replaced by the average of the number of devices. For example, if the eigenvalues of three devices are 10, 15, and 8, the average is 11. If the minimum bounding rectangle area of a high-density area exceeds 50% of the coverage area of the preset correlation radius, the area is split into multiple sub-areas, each with an area not exceeding 25% of the area of the correlation radius. For example, a correlation radius of 500 meters corresponds to a coverage area of 0.785 square kilometers. If the minimum bounding rectangle area is 0.5 square kilometers (exceeding 50%), the area is divided into two sub-areas separated by longitude or latitude.
[0037] Based on the current flow direction and transient energy distribution characteristics between the devices in the regional feature aggregation package, the dynamic aggregation weight of the regional feature aggregation package is modified. The specific implementation is as follows: When determining the direction of current flow based on the current phase difference and voltage gradient between devices, the current phase difference is the phase offset angle between the device voltage and current waveforms at the same timestamp. The phase offset angle is calculated by extracting the phase angle difference of the fundamental component through Fourier transform. The Fourier transform sampling frequency is 10 kHz, and the time window is one power frequency cycle (20 ms). For example, if the voltage phase angle of device A is 30 degrees and the current phase angle is 45 degrees, the phase difference is 15 degrees. The voltage gradient direction is determined by the change in the voltage amplitude between devices along the topological path. The topological path is the shortest connection path between devices in the physical topology of the region. For example, if the topological path from device A to device B passes through two switches, and the voltage amplitude of device A is 10 kV and that of device B is 9.8 kV, the voltage gradient direction is from device A to device B. The current flow direction is determined as follows: if the current phase difference from device A to device B is less than the phase difference from device B to device A, and the voltage gradient direction is consistent with the topological path, then the flow direction is from device A to device B. If the voltage gradient direction conflicts with the topological path (for example, the voltage of device A is higher than that of device B, but there is no direct connection in the topological path), the power flow direction is determined to be invalid and no weight adjustment is performed.
[0038] When extracting transient energy distribution features from regional feature aggregation packages, wavelet packet decomposition is used to extract the frequency band energy contribution of each device's transient waveform. The frequency band division rules for wavelet packet decomposition are dynamically adjusted based on the typical transient frequency range corresponding to the device type. For example, the transient frequency of transformers is concentrated in the 0-1000Hz range, while that of traveling wave ranging devices is concentrated in the 1000-3000Hz range. For transformers, the transient waveform is divided into two equal sub-bands (0-500Hz and 500-1000Hz); for traveling wave ranging devices, it is divided into three equal sub-bands (1000-1500Hz, 1500-2000Hz, and 2000-2500Hz). The frequency band boundary values are determined by the spectral peak distribution of the device's historical transient data. For example, if the historical spectrum peaks of a transformer are at 300Hz and 800Hz, the frequency bands are divided into 0-300Hz, 300-800Hz, and 800-1000Hz. The frequency band energy contribution is calculated as the ratio of the energy value of each sub-band to the total energy value. For example, if device A has an energy of 80 J in the 0-300 Hz band and a total energy of 200 J, the frequency band energy contribution is 40%. The transient energy distribution features of all devices in the regional feature aggregation package are stored by frequency band as a frequency band energy matrix. The matrix row index corresponds to the device number, the column index corresponds to the frequency band number, and the matrix element value is the frequency band energy contribution.
[0039] When adjusting the initial distribution ratio of dynamic aggregation weights between devices based on the current flow direction, if the flow direction is from device A to device B, the weight ratio of device A increases, while the weight ratio of device B decreases. The adjustment range is a function of the consistency of the flow direction and the topological path. The consistency function is calculated by the cosine of the angle between the current phase difference and the voltage gradient direction. For example, if the current phase difference between devices A and B is 15 degrees, and the angle between the voltage gradient direction and the topological path is 10 degrees, the consistency function value is cos(10 degrees) ≈ 0.985. The phase difference weight coefficient is determined by regression analysis of the phase difference and the weight adjustment range in historical power flow events (for example, a fitting coefficient of 0.1). The adjustment range is 0.985 × 0.1 = 9.85%. The initial distribution ratio is set based on the historical fault contribution of the devices. For example, if the initial weight of device A is 0.3 and that of device B is 0.5, the adjusted weight of device A is 0.3 × 1.0985 ≈ 0.329, and that of device B is 0.5 × 0.9015 ≈ 0.451. If the device has no historical fault data, the initial weight is allocated according to the device capacity ratio. For example, if the capacity of device A is 100MVA and that of device B is 150MVA, the initial weight is 100 / (100+150)=0.4, and the initial weight of device B is 0.6.
[0040] When normalizing the adjusted weights based on transient energy distribution characteristics, the correction factor is the ratio of the device's transient energy contribution to the regional average transient energy contribution. The regional average transient energy contribution is calculated by averaging the energy contributions of all devices in the same frequency band in the frequency band energy matrix. For example, in the 0-300 Hz frequency band, if device A contributes 40%, device B contributes 35%, and device C contributes 45%, then the regional average is (40 + 35 + 45) / 3 = 40%. The correction factor is calculated by dividing the device energy contribution by the regional average. For example, the correction factor for device A is 40% / 40% = 1.0, and for device B it is 35% / 40% = 0.875. The final dynamic aggregate weight is the adjusted weight multiplied by the correction factor. For example, if device A's adjusted weight is 0.329 multiplied by the correction factor of 1.0, the final weight is 0.329; if device B's adjusted weight is 0.451 multiplied by the correction factor of 0.875, the final weight is 0.394. The upper and lower limits of the correction factor are determined by the 99% confidence interval of the equipment transient energy contribution in historical data. For example, if the historical contribution ranges from 5% to 200%, the upper limit is set to 2.0 and the lower limit is set to 0.5. If the equipment contribution exceeds 200%, the correction factor is set to 2.0; if it is less than 5%, it is set to 0.5.
[0041] The sum of the normalized and corrected dynamic aggregate weights is forced to 1. For example, the weight of device A is 0.329, the weight of device B is 0.394, and the weight of other devices is 0.277, which sums to 1.0. If the energy of a frequency band is zero after wavelet packet decomposition (for example, if the device has no transient waveform), the energy percentage of that frequency band is set to 50% of the regional average to avoid abnormal correction factor calculation. When the current phase difference between devices is 0 degrees (i.e., the voltage and current are in phase), the default power flow direction is the voltage gradient direction. If the voltage gradient direction is invalid, the weights are redistributed based on the device capacity ratio.
[0042] Based on the modified dynamic aggregation weight and regional feature aggregation package, the preset alarm level template is matched and a combination of device positioning instructions and regional scheduling strategies is output. The specific implementation is as follows: When the modified dynamic aggregation weight matches the spatial distribution density of devices in the regional feature aggregation package, the alarm level template is divided based on the joint distribution threshold of the dynamic aggregation weight and device spatial density in historical fault events. The dynamic aggregation weight and device spatial density data of historical fault events are distributed in a two-dimensional scatter plot and divided into three alarm level ranges using the K-means clustering algorithm. The K-means clustering algorithm is implemented by randomly selecting three initial cluster centers from the historical data, iteratively calculating the Euclidean distance of each data point to the center, and reassigning clusters until the center stabilizes. The Level 1 alarm range corresponds to dynamic aggregation weights > 0.5 and device spatial density > 1.2; the Level 2 alarm range is dynamic aggregation weights between 0.3 and 0.5 and device spatial density between 0.8 and 1.2; and the Level 3 alarm range is dynamic aggregation weights < 0.3 or device spatial density < 0.8. For example, a historical fault event with a dynamic aggregation weight of 0.6 and a device spatial density of 1.5 would fall into the Level 1 alarm range. During real-time matching, if the corrected dynamic aggregation weight is 0.55 and the device spatial density is 1.3, a Level 1 alarm is generated. If device spatial density data is missing, it is replaced by the average density of the grid cell where the device is located and the three adjacent grid cells in the regional feature aggregation package. For example, if the density of the grid cell where the device is located is 1.0, and the density of the three adjacent grid cells is 1.1, 1.0, and 0.9, then the device spatial density is (1.0+1.1+1.0+0.9) / 4=1.0.
[0043] When generating device location instructions based on alarm levels, the device location instructions contain the spatial coordinates and associated topological paths of devices with high dynamic aggregation weights. The spatial coordinates are calculated by weighting the centroid of the geographic range in the regional feature aggregation package and the device-level feature distribution density. The weighted centroid is calculated by converting the device's geographic coordinates (longitude and latitude) to decimal degrees, multiplying them by the corrected dynamic aggregation weights, and summing them, then dividing by the total weight. For example, device A's coordinates are 100.00 degrees east longitude and 30.00 degrees north latitude, with a weight of 0.3; device B's coordinates are 100.01 degrees east longitude and 30.01 degrees north latitude, with a weight of 0.5. The total weight sum is 0.8. The longitude of the weighted centroid is (100.00 × 0.3 + 100.01 × 0.5) / 0.8 = 100.00625 degrees, and the latitude is (30.00 × 0.3 + 30.01 × 0.5) / 0.8 = 30.00625 degrees. The decimal degree conversion rule for geographic coordinates is: if the original coordinates are in degrees, minutes, and seconds format (for example, 100 degrees 15 minutes and 30 seconds east longitude), they are converted to 100 + 15 / 60 + 30 / 3600 ≈ 100.2583 degrees. The associated topological path is the shortest path from the weighted centroid to the nearest key node (such as a substation) in the regional physical topology. The shortest path is calculated using the Dijkstra algorithm, and the path length is based on the actual number of physical connections between devices. For example, the shortest path from the centroid to the substation passes through two switches and one transmission line, with a path length of 3.
[0044] When generating a regional scheduling strategy combination based on the alarm level and device location instructions, the strategy includes load reduction ratio, voltage adjustment priority, and backup line switching order. Load reduction ratios are dynamically allocated based on the power share of devices with high dynamic aggregation weights. The power share is the ratio of the device's real-time power to the region's total power. The region's total power is the sum of the real-time power of all devices with high dynamic aggregation weights. For example, if device A has a power of 10 MW and device B has a power of 15 MW, and the region's total power is 25 MW, device A accounts for 40% (10 / 25) and device B accounts for 60% (15 / 25). If the alarm level is level 1, the load reduction ratio is set to 1.5 times the load reduction ratio: device A's reduction ratio is 40% x 1.5 = 60%, and device B's is 60% x 1.5 = 90%. If the maximum allowable load reduction ratio is 80%, device B will actually receive 80%. Voltage adjustment priority is ranked based on the consistency of the device's voltage gradient direction and the power flow path. The consistency is measured by the cosine of the angle between the voltage gradient direction and the power flow path. The cosine value is calculated as the dot product of the device's gradient direction vector and the power flow path vector, divided by the product of the two vectors' moduli. For example, if device A's gradient direction vector is (1, 0) and its power flow path vector is (0.98, 0.17), with an angle of 10 degrees, the cosine value is approximately 0.985. If device B's vector is (0.87, 0.5) and its power flow path vector is (0.98, 0.17), with an angle of 30 degrees, the cosine value is approximately 0.866. Therefore, device A has higher priority than device B. The order of backup line switching is based on the electrical distance of the topological paths. Electrical distance is determined by the shortest path length in the regional physical topology. Shorter distances have higher priority. For example, if the electrical distance of path A is 3 and that of path B is 5, path A will be prioritized.
[0045] When the alarm level template has no matching interval (for example, a dynamic aggregation weight of 0.2 and a density of 1.3), the lowest alarm level (level 3) is triggered by default. If the total weight sums to 0 when calculating the weighted centroid (for example, if all device weights are abnormally reset to zero), the centroid coordinates are set to the region's geographic center. The region's geographic center is calculated using the geographic boundaries of all grid cells in the regional feature aggregation package, taking the average of the minimum longitude, maximum longitude, minimum latitude, and maximum latitude. For example, if the minimum longitude is 100.00 degrees, the maximum longitude is 100.05 degrees, the minimum latitude is 30.00 degrees, and the maximum latitude is 30.03 degrees, the geographic center is (100.00 + 100.05) / 2 = 100.025 degrees east longitude, and (30.00 + 30.03) / 2 = 30.015 degrees north latitude. If the calculated load reduction ratio exceeds the maximum curtailment capacity of the device (for example, the maximum curtailment allowed for device A is 50%), the maximum capacity is used. For example, if the calculated result is 60%, the actual load is 50%. When calculating voltage adjustment priority, if the cosine values of all devices are less than 0.5, the voltage gradient amplitude is sorted from high to low by absolute value. For example, if the gradient amplitude of device A is 2 kV / m and that of device B is 1 kV / m, device A will be prioritized. If all backup lines are unavailable (for example, if both paths A and B fail), the load is distributed to the remaining lines based on the device capacity ratio. For example, if line C has a capacity of 50 MW and line D has a capacity of 30 MW, the load distribution ratio is 50:(50 + 30) = 5:8.
[0046] This embodiment achieves cross-level semantic consistency through the deep integration of multi-level data aggregation and the physical characteristics of the power grid. Existing hierarchical aggregation mechanisms typically rely on preset statistical rules, failing to address the drawbacks of lost averaging of high-frequency transient signals and disconnected causal logic across devices. This embodiment introduces hierarchical alignment of time tags and device-type-differentiated filling rules during the data preprocessing phase to ensure a unified time base for multi-source data and compatibility with device heterogeneity, providing accurate input for subsequent feature association. During dynamic feature screening, thresholds are dynamically adjusted based on historical device operating conditions and real-time fluctuation amplitudes, replacing traditional fixed threshold filtering. This allows the extraction of abnormal waveform segments and steady-state offsets to adapt to dynamic power grid operation scenarios. The semantic association process innovatively combines reactive current path sensitivity with regional physical topology relationships, quantifying the strength of electrical interactions between devices through mutual information entropy calculations and generating spatiotemporal correlation weights. During feature clustering and cross-layer parsing, kernel density estimation is used to couple transient energy distribution characteristics with current flow direction, dynamically adjusting aggregation weights. This allows regional-level data reconstruction to preserve device-level details while reflecting the real-time operating status of the power grid. The final decision is generated based on the matching logic between dynamic weights and alarm templates, directly mapping data features to scheduling actions, forming a closed-loop link from data dimensionality reduction to control instructions. By cross-domain coupling of electrical and physical laws with data analysis models, the contradiction between semantic distortion and causal disconnection in traditional hierarchical aggregation is resolved.
[0047] The calculations involved in the embodiments are all dimensionless numerical calculations, and the preset parameters and thresholds in the calculations are set by those skilled in the art according to actual conditions.
[0048] The above embodiments may be implemented in whole or in part through software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments may be implemented in whole or in part in the form of a computer program product.
[0049] Those skilled in the art will appreciate that the modules and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application of the technical solution and the invention constraints. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0050] In addition, each functional module in each embodiment of the present application may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.
[0051] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.
[0052] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
[0053] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A power grid data acquisition and analysis system based on big data, characterized in that: include: Equipment acquisition module: obtains raw data from power grid equipment and performs standardized filling and time label layer alignment processing to generate standardized equipment-level operating data; Feature extraction module: Filters dynamic feature sets from standardized device-level operating data. The dynamic feature sets contain abnormal waveform segments and steady-state parameter offsets within a preset period. Semantic association module: determines whether the dynamic feature set triggers a cross-device association event. If so, it generates the spatiotemporal association weight of the dynamic feature set based on the regional physical topology and reactive circulation path sensitivity. Feature clustering module: Based on the spatiotemporal correlation weights and the device-level feature distribution density within a predefined time period, the dynamic feature sets within a preset correlation radius are combined and reconstructed into regional-level feature aggregation packages; Cross-layer parsing module: Based on the current flow direction and transient energy distribution characteristics between the devices in the regional feature aggregation package, the dynamic aggregation weight of the regional feature aggregation package is corrected; Decision generation module: Based on the corrected dynamic aggregation weight and regional feature aggregation package, it matches the preset alarm level template and outputs a combination of device positioning instructions and regional scheduling strategies.
2. A power grid data acquisition and analysis system based on big data according to claim 1, characterized in that: Obtain raw data from power grid devices and perform standardized padding and hierarchical alignment of time tags to generate standardized device-level operating data, including: Based on the preset filling rules corresponding to the device type, the original data obtained from the power grid equipment end is standardized and filled with missing values; Determine the hierarchical time tag alignment strategy based on the device type and signal sampling frequency. For power frequency devices, align time tags based on the cycle period. For high-frequency sampling devices, align time tags based on the microsecond time window. The padded and aligned data are integrated into standardized device-level operating data according to a unified time axis.
3. The power grid data acquisition and analysis system based on big data according to claim 1, characterized in that: Filter dynamic feature sets from standardized device-level operational data, including: Abnormal waveform segments are extracted based on a dynamic sliding window detection method. The sliding window period is consistent with the preset period corresponding to the device type. When the fluctuation amplitude of the data within the window exceeds the dynamic threshold, it is marked as an abnormal waveform segment. The dynamic threshold is adjusted based on the average peak-to-valley difference of the normal waveform in the device's historical operating data. The steady-state parameter offset is determined based on the steady-state baseline calculation method. The steady-state baseline value of each time stamp within a preset period is extracted from the standardized device-level operation data. The steady-state baseline value is the average of the historical operation data of the same device in the period without abnormalities. The absolute deviation between the current data and the steady-state baseline value is calculated as the offset. When the offset exceeds the preset offset threshold, it is determined to be a valid steady-state parameter offset.
4. The power grid data acquisition and analysis system based on big data according to claim 1, characterized in that: Determine whether the dynamic feature set triggers a cross-device association event. If so, generate the spatiotemporal association weights of the dynamic feature set based on the regional physical topology and reactive circulating current path sensitivity, including: Whether the cross-device correlation event triggering conditions are met is determined based on the temporal overlap between the electrical distance between devices and the dynamic feature set. The electrical distance is calculated based on the shortest connection path length between devices in the regional physical topology. The temporal overlap is the time window overlap ratio of the abnormal waveform segment or steady-state parameter offset. If a cross-device association event is triggered, the spatiotemporal association weight is adjusted according to the reactive circulation path sensitivity. The reactive circulation path sensitivity is calculated by the ratio of the reactive circulation mutual information entropy between devices to the topological distance. The initial value of the spatiotemporal association weight is the inverse of the electrical distance between devices. The final spatiotemporal association weight is the initial value of the spatiotemporal association weight multiplied by the reactive circulation path sensitivity correction coefficient.
5. The power grid data acquisition and analysis system based on big data according to claim 1, characterized in that: Based on the spatiotemporal correlation weights and the device-level feature distribution density within a predefined time period, the dynamic feature sets within the preset correlation radius are combined and reconstructed into a regional-level feature aggregation package, including: The device-level feature distribution density is calculated based on the kernel density estimation method. The kernel function bandwidth is adjusted according to the device spatial distribution sparsity. The device spatial distribution sparsity is determined by the ratio of the average distance between devices to the preset association radius. Generate a regional feature density distribution map based on the device-level feature distribution density, and rasterize the regional feature density distribution map through a heat map; Extract high-density areas from the regional feature density distribution map based on a preset correlation radius, where the high-density areas are areas where the density values of continuous grid cells exceed a preset density threshold; The dynamic feature set in the high-density area is weightedly fused according to the spatiotemporal correlation weight and reconstructed into a regional feature aggregation package. The weighted fusion method is to multiply the dynamic feature value by the spatiotemporal correlation weight of the corresponding device, accumulate it, and then divide it by the total weight sum.
6. The power grid data acquisition and analysis system based on big data according to claim 5, characterized in that: The grid cell density value is the weighted sum of the device-level feature distribution density within the corresponding spatial range, and the weight is the spatiotemporal correlation weight of the device in the corresponding unit. The preset density threshold is determined based on the statistical distribution of density values in historical regional-level feature aggregation events.
7. The power grid data acquisition and analysis system based on big data according to claim 1, characterized in that: Based on the current flow direction and transient energy distribution characteristics between the devices in the regional feature aggregation package, the dynamic aggregation weight of the regional feature aggregation package is modified, including: The current flow direction is determined by the current phase difference and voltage gradient direction between devices; Extract the transient energy distribution features from the regional feature aggregation package. The transient energy distribution features are decomposed by wavelet packets to extract the frequency band energy ratio of each device's transient waveform. The frequency band energy ratio is the ratio of the energy of a specified frequency band to the total energy. Adjust the initial distribution ratio of dynamic aggregation weights between devices based on the direction of current flow. The adjustment range is a function of the flow direction and the consistency of the topological path. The adjusted weight is normalized and corrected based on the transient energy distribution characteristics. The correction coefficient is the ratio of the equipment transient energy proportion to the regional average transient energy proportion. The final dynamic aggregation weight is the adjusted weight multiplied by the correction coefficient.
8. The power grid data acquisition and analysis system based on big data according to claim 7, characterized in that: The current phase difference is the phase offset angle between the device voltage and current waveforms at the same timestamp, and the voltage gradient direction is the changing trend of the voltage amplitude between devices along the topological path.
9. The power grid data acquisition and analysis system based on big data according to claim 1, characterized in that: Based on the modified dynamic aggregation weights and regional feature aggregation packages, the system matches the preset alarm level template and outputs a combination of device positioning instructions and regional scheduling strategies, including: Based on the corrected dynamic aggregation weight and the spatial distribution density of devices in the regional feature aggregation package, the corresponding alarm level in the preset alarm level template is matched. The alarm level template is divided by the joint distribution threshold of the dynamic aggregation weight and the spatial density of devices in historical fault events; Generate device location instructions based on the alarm level. The device location instructions contain the spatial coordinates and associated topological paths of the device with high dynamic aggregation weight. The spatial coordinates are calculated by the weighted centroid of the geographic range in the regional feature aggregation package and the device-level feature distribution density. Generate regional scheduling strategy combinations based on alarm levels and device positioning instructions.
10. The power grid data acquisition and analysis system based on big data according to claim 9, characterized in that: The regional scheduling strategy combination includes load reduction ratio, voltage adjustment priority and backup line switching order. The load reduction ratio is dynamically allocated according to the power proportion of high dynamic aggregation weight equipment, and the voltage adjustment priority is sorted by the consistency of the equipment voltage gradient direction and the flow path.
Citation Information
Patent Citations
Artificial intelligence data aggregation method based on big data
CN118520020A
Method and system for verifying and correcting integrity of flight data and storage medium
CN119336746A
Distributed storage node fault detection system
CN119718741A
Multi-source heterogeneous data real-time access method and system
CN119739776A
Data processing method and system for smart city and storage medium
CN119989224A
Cited By
Remote monitoring system and method for fire alarm equipment
CN120744785A
Automatic data grading protection method and system based on multi-feature fusion
CN120763957A
A data automatic grading protection method and system based on multi-feature fusion
CN120763957B
Power grid fault diagnosis method and system based on space-time correlation analysis
CN121051348A
High-precision voltage fluctuation real-time monitoring method and system based on electric energy meter
CN121476698A