A substation online monitoring system based on big data analysis
Through the online substation monitoring system analyzed by big data, the composite entropy feature vector is constructed using multi-scale entropy indexes, which solves the lag and lack of early warning problems of traditional monitoring methods, realizes early risk assessment and early warning of substation equipment and systems, and improves the detection ability of complex coupling states.
Patent Information
- Application Number
- CN202510751290.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2045-06-06
AI Technical Summary
Traditional substation monitoring methods are difficult to effectively utilize the inherent dynamic information of the power grid, lack sensitivity warnings for early functional disorders and critical system transformations, and lack online evaluation methods for dynamic interactions between devices and subsystems.
The online substation monitoring system based on big data analysis is adopted, and through data acquisition, signal preprocessing, system response entropy calculation, entropy feature baseline management, degradation evaluation and critical early warning modules, a composite entropy feature vector is constructed to achieve health status assessment and risk warning of substation equipment and systems.
It can capture the signal of functional degradation of equipment and system disorder in the early stage, provide preventive maintenance time, quantify the intensity of information flow between equipment, identify potential interaction risks, reduce interference to normal operation, and improve detection capabilities of complex coupling states.
Smart Images

Figure CN120301041B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of power system status monitoring, and in particular to an online monitoring system for substations based on big data analysis. Background Art
[0002] Substations are the core hubs of power systems, and their safe and stable operation is crucial. Currently, substation equipment condition monitoring relies primarily on threshold determination of electrical parameters, dissolved gas in oil, partial discharge, and other information acquired through online monitoring devices, as well as regular offline preventive testing and manual inspections.
[0003] These traditional methods have many limitations: 1) Alarms based on physical parameter thresholds often lag behind the actual degradation of equipment performance, making it difficult to effectively detect early or latent faults; 2) They mainly focus on physical damage to the equipment or parameter excursions, and lack effective monitoring methods for "functional disorder" or early performance degradation caused by changes in control logic, dynamic interaction, or overall system characteristics; 3) It is difficult to accurately assess and predict trends in complex system degradation processes caused by the coupling of multiple factors; 4) Existing methods lack effective early warning capabilities, especially for major systemic critical transition events such as voltage collapse and frequency instability. They can often only be detected when the event occurs or is about to occur, making it difficult to gain sufficient response time; 5) There is also a lack of effective online assessment methods for the health status of dynamic interactions between devices and subsystems. Summary of the Invention
[0004] The present invention provides an online substation monitoring system based on big data analysis to solve the technical difficulties that traditional substation monitoring methods have ineffectively utilizing inherent dynamic information of the power grid, are not sensitive enough to early functional disorders, and lack advance warning of critical transition behaviors of the system.
[0005] The present invention provides a substation online monitoring system based on big data analysis, the system comprising:
[0006] A data acquisition module configured to collect multi-source dynamic signals and auxiliary operating condition data from multiple monitoring points within the substation in real time;
[0007] A signal preprocessing and feature window extraction module, connected to the data acquisition module, configured to preprocess the collected multi-source dynamic signals and extract feature window data that can reflect the dynamic response of the system according to a preset logic;
[0008] a system response entropy calculation module, connected to the signal preprocessing and feature window extraction module, configured to calculate multiple types of system response entropy indicators for the feature window data and combine them into a composite entropy feature vector;
[0009] an entropy feature baseline management module, connected to the system response entropy calculation module and the data acquisition module, configured to store the historical composite entropy feature vector and the corresponding auxiliary operating condition data, and to construct and maintain a health status baseline of the composite entropy feature vector under different auxiliary operating condition data using a machine learning model;
[0010] The degradation assessment and critical warning module is connected to the system response entropy calculation module and the entropy feature baseline management module, and is configured to compare the composite entropy feature vector calculated in the current cycle with the health status baseline corresponding to the current auxiliary operating condition data, analyze their deviations, and evaluate the degradation status of substation equipment and the system critical transition risk based on multiple preset abnormality judgment rules, and output evaluation results and warning information.
[0011] Preferably, the system response entropy calculation module is specifically configured to calculate at least two items of multi-scale permutation entropy, multi-scale fuzzy entropy and transfer entropy for a preselected signal pair to form the composite entropy feature vector.
[0012] Preferably, the signal preprocessing and feature window extraction module is specifically configured to extract the feature window data by combining one or both of event triggering logic and sliding window logic, and when event triggering logic is adopted, the module is also configured to characterize the disturbance that triggers the event.
[0013] Preferably, the machine learning model adopted by the entropy feature baseline management module takes the auxiliary operating condition data and, optionally, the disturbance characterization description provided by the signal preprocessing and feature window extraction module as input, and outputs the expected value of the composite entropy feature vector under the predicted health state and its normal fluctuation range, thereby establishing and dynamically adjusting the health state baseline.
[0014] Preferably, the degradation assessment and critical warning module is specifically configured to calculate the deviation of the current composite entropy feature vector relative to the health status baseline, and apply at least two of the amplitude abnormality judgment rules, trend abnormality judgment rules, change rate abnormality judgment rules, critical transition precursor judgment rules and information flow interaction abnormality judgment rules to conduct a comprehensive analysis of the deviation and one or both of the composite entropy feature vectors.
[0015] Preferably, the critical transition precursor judgment rule includes monitoring the statistical characteristics of the time series of a specific entropy indicator in the composite entropy feature vector or the time series of the deviation, wherein the statistical characteristics include at least one of the variance and the first-order autocorrelation coefficient. When the statistical characteristics exhibit a specific evolutionary pattern indicating that the system is approaching an unstable state, a critical warning is triggered.
[0016] Preferably, the information flow interaction anomaly judgment rule is based on analyzing the transfer entropy value calculated for the preselected key signal pairs in the composite entropy feature vector, and evaluating the health of the dynamic interaction relationship between devices or subsystems by comparing the transfer entropy value calculated in the current cycle with its health status baseline under the corresponding working conditions, or comparing the change in the transfer entropy symmetry between signal pairs.
[0017] Preferably, the multi-source dynamic signal includes at least one selected from the instantaneous waveforms of voltage and current on the high, medium and low voltage sides of the transformer, the bus voltage and system frequency, and the input, output and internal state signal groups of the key control system; the auxiliary operating condition data includes at least one selected from the system active and reactive load levels, the grid operation topology structure code, and the environmental parameter group.
[0018] The technical solution provided by this application has at least the following technical effects or advantages:
[0019] By analyzing the entropy characteristics of the system's dynamic response, the present invention can capture the system's "functional degradation" or "disorder" signals caused by material aging, structural changes, or decreased control performance earlier than the physical parameters of the equipment deteriorate significantly, thereby buying valuable time for preventive maintenance and intervention, and realizing the transformation from "repairing it when it breaks" to "intervening when it becomes dysfunctional."
[0020] The present invention is dedicated to capturing early signals that the system is about to enter an unstable state by monitoring whether the entropy sequence and its statistical characteristics show precursory patterns such as critical slowdown, thereby gaining a decision window for operators to take emergency preventive measures.
[0021] Using the transfer entropy metric, this system quantifies the intensity and direction of information flow between key devices or control systems. By monitoring abnormal changes in transfer entropy, it can identify potential interaction risks that are typically difficult to directly observe, such as control coordination issues, increased undesirable dynamic coupling, or information blockage. This extends monitoring from focusing on "nodes" to "connections" and "networks."
[0022] The system primarily relies on analyzing responses to naturally occurring background disturbances and normal operating disturbances in the power grid. It eliminates the need to actively inject external detection signals into the system and eliminates the need for equipment shutdown, minimizing disruption to normal substation operation.
[0023] Since the entropy index measures the "disorder" or "complexity" of the overall dynamic behavior of the system rather than specific preset fault characteristics, this system also has the potential to detect abnormal conditions whose patterns are unknown or caused by the complex coupling of multiple factors and are difficult to describe with traditional single characteristics. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 This is an architecture diagram of a substation online monitoring system based on big data analysis in the present invention. DETAILED DESCRIPTION
[0025] The present invention relates to an online substation monitoring system based on big data analysis to solve the technical difficulties that traditional substation monitoring methods have ineffectively utilizing inherent dynamic information of the power grid, are not sensitive enough to early functional disorders, and lack advance warning of critical transition behaviors of the system.
[0026] The above technical solution will be described in detail below in conjunction with the accompanying drawings and specific implementation methods of the specification to better understand the above technical solution. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments of the present invention. It should be understood that the present invention is not limited to the example embodiments used only to explain the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention. In addition, it should be noted that, for the convenience of description, only the parts related to the present invention, rather than all, are shown in the drawings.
[0027] like Figure 1 The architecture diagram of a substation online monitoring system based on big data analysis is shown in the figure. The system includes the following modules:
[0028] A data acquisition module configured to collect multi-source dynamic signals and auxiliary operating condition data from multiple monitoring points within the substation in real time;
[0029] A signal preprocessing and feature window extraction module, connected to the data acquisition module, configured to preprocess the collected multi-source dynamic signals and extract feature window data that can reflect the dynamic response of the system according to a preset logic;
[0030] a system response entropy calculation module, connected to the signal preprocessing and feature window extraction module, configured to calculate multiple types of system response entropy indicators for the feature window data and combine them into a composite entropy feature vector;
[0031] an entropy feature baseline management module, connected to the system response entropy calculation module and the data acquisition module, configured to store the historical composite entropy feature vector and the corresponding auxiliary operating condition data, and to construct and maintain a health status baseline of the composite entropy feature vector under different auxiliary operating condition data using a machine learning model;
[0032] The degradation assessment and critical warning module is connected to the system response entropy calculation module and the entropy feature baseline management module, and is configured to compare the composite entropy feature vector calculated in the current cycle with the health status baseline corresponding to the current auxiliary operating condition data, analyze their deviations, and evaluate the degradation status of substation equipment and the system critical transition risk based on multiple preset abnormality judgment rules, and output evaluation results and warning information.
[0033] 1. The data acquisition module is responsible for collecting multi-source, high-precision dynamic signals from key equipment and measurement and control modules in the substation.
[0034] The selection of collection points should cover key locations that can reflect the dynamic behavior of core equipment such as transformers, circuit breakers, busbars, control systems, etc.
[0035] For transformers, the three-phase voltage and current transient waveforms on the high, medium, and low voltage sides are collected; for buses, the bus voltage and frequency are collected; and for lines, the line current and voltage are collected. These signals are crucial for analyzing electromagnetic transient response and system stability.
[0036] Collect input signals, output signals such as AVR excitation current and excitation voltage, and key internal state variables of key control systems such as automatic voltage regulators (AVRs), speed regulators (GOVs), and relay protection devices. These signals directly reflect the dynamic response characteristics and health of the control system.
[0037] Auxiliary operating conditions: Auxiliary operating conditions that reflect the system operating background are collected synchronously, such as load active / reactive power, system topology switch position, switch status, ambient temperature, humidity, etc. These are used for subsequent adaptive adjustment of the entropy baseline.
[0038] Sensors and acquisition devices must have sufficient bandwidth. For example, for transient voltage and current waveforms, the bandwidth must cover at least the DC to several megahertz (MHz) range to capture high-frequency disturbances and fast transient processes. Resolution must be at least 16 bits or higher to ensure sensitivity to subtle signal changes.
[0039] The sampling rate is set based on the specific analysis objectives. For example, if broadband dynamic impedance spectroscopy or information from high-frequency noise is required, a sampling rate of 1 MHz or higher is preferred. If the focus is only on lower-frequency control system response entropy or electromechanical oscillation modes, a sampling rate of several kHz to tens of kHz may also meet the requirements, but it must ensure that the Nyquist sampling theorem is met and sufficient harmonic and transient component information is retained.
[0040] 2. The signal preprocessing and feature window extraction module is responsible for purifying the original collected data and extracting the effective feature window for entropy calculation.
[0041] First, digital filtering is performed to eliminate strong interference in specific frequency bands, such as notch processing of the power frequency and its harmonics. If it is not the analysis object, it is also necessary to pay attention to retaining the broadband dynamic information required for analysis to avoid excessive filtering that may cause distortion of the useful signal.
[0042] The operational process involves, for example, applying a high-pass filter to remove DC offsets, and then optionally applying an adaptive notch filter targeting specific known interference frequencies, such as a 50 Hz or 60 Hz power frequency. Filter parameters, such as cutoff frequency and order, must be carefully selected to balance noise rejection and signal fidelity.
[0043] Furthermore, data alignment: using high-precision timestamps, signals from different acquisition points and different devices are accurately aligned on the time axis.
[0044] Its operation process is: using a unified time base, resample or interpolate each signal to ensure that the sampling points at the same time are strictly corresponding.
[0045] Event triggering mechanism based on feature window extraction logic: Continuously monitors preset key signals, such as bus voltage RMS, system frequency, instantaneous value change rate of key line current, or short-time energy.
[0046] When a specific indicator of the monitored signal, such as the absolute value of the rate of change or short-term energy increment, continuously exceeds a dynamic or fixed preset threshold for a certain period of time, an event recording is triggered. This threshold can be set based on the statistical characteristics of historical data, such as the mean of the normal fluctuation range plus N times the standard deviation. It can also be adaptively adjusted according to operating conditions to distinguish between normal operating fluctuations and significant disturbance events such as short circuit faults, large load switching, and system oscillation.
[0047] Taking the event time (the "event time") as the benchmark, relatively steady-state data is captured for a period of time forward, for example, a few seconds (the "pre-event duration"), to characterize the system state before the event. Dynamic response data is captured for a period of time backward, for example, a few to tens of seconds (the "post-event duration"), sufficient to cover the system response decay process and the transition to a new steady-state or sustained oscillation state. The resulting characteristic window data segment consists of all relevant signals within the time interval from the "event time" minus the "pre-event duration" to the "event time" plus the "post-event duration." The specific lengths of the "pre-event duration" and "post-event duration" can be set or adaptively adjusted based on typical fault response times and analysis requirements.
[0048] Disturbance characterization: preliminary analysis of the signal segment that triggers the event, especially the signal itself that causes the trigger.
[0049] First, the approximate type of the disturbance event is extracted, such as voltage sag, frequency swing, high-order harmonic injection, pulse disturbance, etc., which can be preliminarily judged through waveform morphology and spectrum analysis, as well as its key parameters, such as disturbance amplitude or depth (change relative to the rated value), disturbance duration, and the main frequency components of the disturbance signal (obtained through fast Fourier transform FFT or wavelet analysis).
[0050] These extracted disturbance features will be passed as a vector to the subsequent entropy feature baseline management module and degradation assessment and critical warning module for conditional analysis and more accurate baseline matching.
[0051] Sliding window mechanism (background monitoring during periods without significant events): When there are no significant external events triggering the system, the "background" response entropy of the system is continuously calculated to monitor the system's endogenous dynamics and subtle changes.
[0052] It adopts a fixed length, for example, a few seconds, namely the "total window length", which should be long enough to cover the sliding window of the characteristic time of the system dynamic mode of interest.
[0053] The sliding step size is set between 10% and 50% of the total window length. The choice of sliding step size is a trade-off between computational load and real-time monitoring. A smaller step size provides finer temporal resolution but requires more computation.
[0054] Data segments with a length of "total window length" are extracted from the aligned signal data stream in sequence for subsequent entropy calculation.
[0055] 3. The system response entropy calculation module is responsible for calculating various entropy indicators. For each feature window data, whether it is an event-triggered window or a sliding window:
[0056] The signal selection logic is as follows: Based on the current analysis objective, such as evaluating transformer winding status, controller performance, and system stability margin, one or more groups of the most relevant signals are selected from the characteristic window data. For example, when analyzing transformer dynamic characteristics, the differential current between the high and low voltage sides and the voltage and current on each side are prioritized; when analyzing AVR controller behavior, the generator terminal voltage, excitation current, excitation voltage, and the AVR reference input and control output signals are selected.
[0057] Furthermore, for slow trends that may exist in certain signals, such as slow load growth or drift caused by temperature changes, these trends are not dynamic responses of interest and can be removed by, for example, subtracting the fitting curve after polynomial fitting, or by using the first-order difference method, to avoid affecting the calculation of the entropy value.
[0058] Furthermore, the selected signal segments are normalized, for example, using Z-score normalization, where each data point is subtracted from the mean value of the signal segment and then divided by the standard deviation or the maximum and minimum values of the signal segment to normalize the data linearly to a preset range, such as 0 to 1 or between -1 and 1. The purpose is to eliminate the influence of different signal amplitude dimensions, making subsequent entropy calculations, especially the similarity tolerance parameters in fuzzy entropy, more comparable, and potentially improving the numerical stability of certain entropy algorithms.
[0059] Multi-scale permutation entropy calculation: Quickly assess the pattern complexity and randomness of a time series at different time scales. Higher values generally indicate a more complex or closer-to-random sequence.
[0060] Specific calculation steps:
[0061] 1. For a selected signal sequence of length N, set a maximum analysis scale factor. For example, select a scale factor based on the characteristic time period range of the dynamic behavior you want to analyze, such as an equivalent scale of 1 to 20 sampling periods. For each scale factor, gradually increase from 1 to the set maximum scale factor, splitting the original time series into continuous, non-overlapping segments, each containing the number of data points corresponding to the scale factor. Calculate the average value of all data points in each segment, and these average values form a new, shortened "coarse-grained" time series. When the scale factor is 1, this coarse-grained sequence is the original sequence itself.
[0062] 2. Select an embedding dimension for each coarse-grained sequence. For example, based on the signal characteristics and computing resources, it is usually selected to be 3 to 7. If the embedding dimension is too small, the dynamic information of the sequence may be lost. If it is too large, the computational complexity increases significantly and longer data is required to obtain reliable probability estimates. It may also be more sensitive to noise. A time delay parameter is also selected (usually for sequences that have been coarse-grained or original sequences that are fully sampled and not specifically downsampled, the time delay parameter is 1. If the original signal has obvious periodicity and is not completely smoothed by coarse-graining, this parameter can be adjusted to better capture its structure. For example, it can be selected by referring to the first minimum of the average mutual information function of the reference signal).
[0063] 3. For the current coarse-grained sequence, divide it into a series of overlapping vectors with a dimension equal to the selected embedding dimension. Specifically, starting from the first point in the sequence, jump to the number of points required for the embedding dimension according to the selected time delay parameter to form a vector. Then, starting from the second point in the sequence, the next vector is formed in the same manner, and so on, until the end of the sequence. For each vector thus formed, the magnitude of each element within it is compared, and based on their order from smallest to largest (or largest to smallest), a "permutation pattern" corresponding to the vector is determined. For example, if the embedding dimension is 3, and the three elements of a vector are 5.1, 2.5, and 8.0, then when they are arranged in ascending order of magnitude, the second element (2.5) in the original vector is the smallest, the first element (5.1) is the second largest, and the third element (8.0) is the largest. This represents a specific permutation pattern. For a given embedding dimension, there are a total of factorial possible permutations of the embedding dimension.
[0064] 4. Count the number of times each unique pattern appears in all the vectors obtained in step 3. Then, divide the number of occurrences of each pattern by the total number of generated vectors to get the probability of occurrence of the pattern.
[0065] 5. Calculate the permutation entropy at the current scale using the definition of Shannon entropy. The calculation method is: for all permutation patterns that actually appear in the data, multiply the probability of that pattern by the logarithm of that probability (usually the natural logarithm or the base-2 logarithm), then add up all these products and take the negative of the sum. This negative sum is the permutation entropy at the current scale.
[0066] 6. Repeat steps 2 to 5 to calculate the permutation entropy value at each scale from scale factor 1 to the set maximum scale factor. These permutation entropy values at different scales are arranged in scale order to form a numerical vector, which is the multi-scale permutation entropy (MSPE) vector.
[0067] Parameter selection basis description:
[0068] Embedding dimension: Usually based on experience or recommendations from relevant research literature, such as between 3 and 7. The choice should balance the adequacy of information extraction, computational efficiency, and the requirements for data length.
[0069] Time delay parameter: For coarse-grained sequences, this parameter is usually set to 1. If calculated directly on the original high-sampling rate sequence, its selection can be based on the first zero crossing of the signal's autocorrelation function or the first minimum point of the average mutual information, in order to reduce redundant information between data points.
[0070] The maximum scale factor depends on the time scale you wish to analyze and should cover the characteristic timescales of the system dynamics you are interested in. For example, if you are interested in second-order oscillations and the system sampling rate is in the kilohertz range, you may need to set the maximum scale factor larger to ensure that the coarse-grained time resolution can capture second-order dynamics.
[0071] Multi-scale fuzzy entropy calculation: Assess the regularity of time series at different scales. Fuzzy entropy is more robust than sample entropy, especially in the presence of noise. Higher values generally indicate more irregular and complex sequences.
[0072] Calculation steps:
[0073] 1. The first step in the calculation of multi-scale permutation entropy is to obtain coarse-grained sequences under different scale factors.
[0074] 2. For each coarse-grained sequence, select the embedding dimension (the selection principle is the same as that of the multi-scale permutation entropy), a parameter called "similarity tolerance" (usually 0.1 to 0.25 times the standard deviation of the original signal or the current coarse-grained signal; the fuzzy entropy is less sensitive to the choice of this parameter than the sample entropy), and the type of fuzzy function and its related parameters. For example, a Gaussian function is selected as the membership function, which is usually in the form of an exponential power of the natural constant e, which is related to the distance between the two vectors, the similarity tolerance, and an exponential parameter that determines the steepness of the similarity boundary.
[0075] 3. For the current coarse-grained sequence, construct a series of pattern vectors with a dimension equal to the selected embedding dimension (similar to the approach in multi-scale permutation entropy, but the time delay parameter is usually fixed to 1). Then, for any two such pattern vectors, first calculate the distance between them. A commonly used distance is the Chebyshev distance, which is the maximum absolute value of the difference between the corresponding elements of the two vectors. Then, using the selected fuzzy membership function and similarity tolerance parameters, calculate the "fuzzy similarity" between the two pattern vectors. The result is a value between 0 and 1 that indicates their degree of similarity.
[0076] 4. For each pattern vector, calculate the average (or sum, the specific definition in different literature may vary slightly, here we use average as an example) of its fuzzy similarity with all other pattern vectors of the same dimension (excluding itself, or depending on the specific definition, it can be included). Then average all these averages (or, more precisely, first calculate the similarity between all possible pairs of vectors and then take the total average of these similarities) to obtain a global similarity measure based on the current embedding dimension, which is called "m-dimensional global similarity".
[0077] 5. Increase the embedding dimension by one unit, repeat steps 3 and 4, construct a pattern vector with one dimension increased, and calculate the corresponding global similarity measure, which is called "m plus 1-dimensional global similarity".
[0078] 6. The fuzzy entropy value at the current scale is usually calculated as the negative of the natural logarithm of the ratio of "m-dimensional global similarity" to "m plus 1-dimensional global similarity". Or equivalently, it can be calculated as the natural logarithm of "m-dimensional global similarity" minus the natural logarithm of "m plus 1-dimensional global similarity". This ensures that the entropy value is positively correlated with the complexity or irregularity of the signal.
[0079] 7. MSFE vector output: Repeat steps 2 to 6 to calculate the fuzzy entropy value at each scale from scale factor 1 to the set maximum scale factor. These fuzzy entropy values at different scales are arranged in scale order to form a numerical vector, which is the multi-scale fuzzy entropy (MSFE) vector.
[0080] Parameter selection basis description:
[0081] Embedding dimension: The selection principle is the same as multi-scale permutation entropy.
[0082] Similarity tolerance: It is generally recommended to set it between 10% and 25% of the signal standard deviation. In actual applications, experimentation and optimization may be required based on signal characteristics and noise levels.
[0083] Fuzzy function type and its parameters: Gaussian function and exponential function are commonly used choices. The exponential parameter in the function affects the smoothness of the similarity boundary.
[0084] Transfer entropy calculation: quantifies the extent to which the historical information of a signal sequence called source signal X improves the ability to predict the future state of another signal sequence called target signal Y, thereby evaluating the strength of the directional information flow from source signal X to target signal Y or can be interpreted as a causal influence.
[0085] Calculation steps:
[0086] 1. Select a pair of signal sequences whose interaction relationship needs to be analyzed and label them as source signal X and target signal Y. For example, X can be the control output signal of an automatic voltage regulator (AVR) and Y can be the terminal voltage signal of a generator.
[0087] The historical embedding dimension of the target signal Y: determines how many consecutive or equally spaced sampling values of the target signal Y in the past are used to predict its future.
[0088] The historical embedding dimension of the source signal X determines how many consecutive or equally spaced past sampling values of the source signal X are used to assist in predicting the future of the target signal Y.
[0089] 2. Determine the future value of the target signal Y to be predicted. The choice of these parameters significantly impacts the calculation of the transfer entropy and requires careful consideration based on the signal's autocorrelation, cross-correlation, known system delay characteristics, and the desired timescale for analysis. This typically requires prior knowledge or optimization through methods such as traversal search and information criteria (such as the Akaike Information Criterion (AIC) and the Bayesian Information Criterion (BIC). For example, the prediction lead can be set to a typical control action delay for the system.
[0090] 3. The calculation of transfer entropy is to compare the predictability of the future value of the target signal Y under two conditions. This involves estimating multiple joint probability density functions and conditional probability density functions. Since these probability distributions are usually unknown in practice, non-parametric estimation is required from limited observation data. Common estimation methods include:
[0091] Kernel density estimation: A smooth kernel function (e.g., a Gaussian kernel) is placed at each data point. The sum of all these kernel functions forms a smooth estimate of the overall probability density function. The "bandwidth" (i.e., degree of smoothness) of the kernel function is a key parameter that needs to be carefully chosen (e.g., using empirical rules such as Silverman's rule or through cross-validation).
[0092] K-nearest neighbor estimation: By calculating the distance between each data point and its k nearest neighbors in its high-dimensional state space, the local probability density of the point is estimated. The choice of parameter k (the number of neighbors) affects the smoothness and bias of the estimate.
[0093] Data binning / discretization: The range of each variable (signal sample value) is divided into several discrete intervals ("bins"). The corresponding probabilities are then approximated by counting the frequency of data points in each state combination (i.e., a multidimensional "bin" composed of "bins" of different signals at different times). This method is relatively simple, but the accuracy of the results depends on the binning strategy and is susceptible to the so-called "curse of dimensionality" when dealing with high-dimensional data (i.e., the amount of data required grows exponentially with the dimensionality).
[0094] The logical process of transferring entropy value calculation:
[0095] a. First, consider using only the values of the target signal Y at several past moments to predict its value at a certain future moment, and calculate the uncertainty or amount of information in this prediction.
[0096] b. Then, use the values of the target signal Y at several past moments and the values of the source signal X at several past moments to predict the value of the target signal Y at the same future moment, and again calculate the uncertainty or information content of this joint prediction.
[0097] c. The value of the transfer entropy is the uncertainty (or amount of information) calculated in step a minus the uncertainty (or amount of information) calculated in step b. This difference indicates how much the uncertainty in predicting the future value of the target signal Y is reduced by the introduction of the historical information of the source signal X. If the source signal X has no predictive power for the future of the target signal Y, this difference is zero or close to zero. If the source signal X contains useful information about the future of the target signal Y (beyond the information contained in the target signal Y's own history), this difference is positive. In practical calculations, this is usually achieved by calculating a weighted sum involving the logarithm of the ratio of the joint probability and the conditional probability for all observed data points (or all possible state combinations if binning is used).
[0098] 4. Calculate the transfer entropy from source signal X to target signal Y. Similarly, by swapping the roles of source signal X and target signal Y and repeating the above steps, the transfer entropy from target signal Y to source signal X can be calculated. These two values together describe the bidirectional information flow interaction between signal X and signal Y.
[0099] Entropy feature vector construction: The multi-scale permutation entropy vector and multi-scale fuzzy entropy vector calculated for the same feature window data are collected, as well as the transfer entropy values calculated for several pre-selected key interaction signal pairs, for example, the transfer entropy value from a controller output signal to its controlled object state signal, or the transfer entropy value from an output signal of key device A to an input signal of device B.
[0100] These entropy values from different sources (which may already be vectors, such as multi-scale entropy, or single values, such as transfer entropy) are concatenated in a predetermined order to form a single, potentially high-dimensional sequence of values (the total dimensionality depends on the maximum scaling factor and the number of selected transfer entropy pairs), known as the "entropy eigenvector." This vector comprehensively characterizes the dynamic characteristics of the system response within that characteristic window from multiple perspectives.
[0101] 4. The entropy feature baseline management module is responsible for storing historical entropy features and establishing and managing the health baseline model for status assessment.
[0102] Store each entropy feature vector calculated historically. Stored along with each entropy feature vector are: an accurate timestamp; a detailed grid operating parameter vector corresponding to the time the feature window was generated (e.g., the active and reactive power output of the generator set, the load factor of key transmission lines, the voltage level of the main busbar, the numerical encoding of the key switches and circuit breakers in the system, ambient temperature, humidity, etc., which need to be quantized or encoded as numerical vectors); and, if the feature window is triggered by a specific event, the feature vector of the disturbance event extracted in the signal preprocessing and feature window extraction module.
[0103] Baseline model establishment and selection: Establish a model that can accurately predict the expected value (and the normal fluctuation range of this expected value) of the entropy eigenvector when the system is in a "healthy" state under certain conditions based on the current operating parameters of the substation and, in some cases, the characteristics of the disturbance that has occurred.
[0104] Model type selection: Considering the possible complex nonlinear relationship between operating conditions and entropy characteristics, it is preferred to use a machine learning model with nonlinear fitting capabilities.
[0105] Gaussian Process Regression (GPR): GPR is a nonparametric regression method based on Bayesian theory. Rather than directly learning a fixed functional form, it assumes that the output data points corresponding to any set of input data points follow a joint Gaussian distribution. By defining a "kernel function" to describe the similarity or correlation between different input points, GPR can learn the underlying relationship between input and output from the training data.
[0106] The input of the model is the current operating condition parameter vector and (optionally) the disturbance feature vector. The output of the model is the predicted average expected value of the entropy feature vector when the system is in a "healthy" state under the current input conditions, as well as the variance (or more complete covariance matrix) associated with this prediction. This variance information is very useful because it can naturally give the normal fluctuation range or confidence interval of the entropy characteristics in a healthy state. It can provide the uncertainty (i.e., variance) of the prediction results, which is very helpful in defining what is the "normal range"; it can also have good learning performance for small sample data sets; the choice of kernel function provides the model with great flexibility to adapt to different types of data relationships.
[0107] Neural Networks (NN): For example, a multilayer perceptron (MLP), a feedforward neural network, can be used. For complex scenarios where the temporal dependencies of the entropy feature sequence must be considered, recurrent neural network (RNN) variants (such as long short-term memory (LSTM) networks or gated recurrent modules (GRU)) can also be considered.
[0108] Typical structure (taking MLP as an example): MLP usually consists of an input layer (used to receive the current operating parameter vector and disturbance feature vector), one or more hidden layers (each hidden layer contains several neurons, using nonlinear activation functions such as ReLU function or tanh hyperbolic tangent function for nonlinear transformation and extraction of features) and an output layer (used to predict the entropy feature vector under the healthy state).
[0109] The training process involves supervised training using a backpropagation algorithm, utilizing a large number of pre-labeled "healthy" data pairs (i.e., pairs of operating condition / disturbance feature inputs and corresponding health entropy feature outputs). The goal of training is to adjust the connection weights between neurons in the network to minimize the difference (e.g., mean squared error) between the model's predicted entropy features and the actual health entropy features. This allows the model to learn and fit very complex nonlinear relationships.
[0110] Cluster analysis combined with regression:
[0111] Step 1 (Operating Condition Clustering): First, a clustering analysis algorithm (e.g., the commonly used k-means algorithm, DBSCAN density clustering algorithm, or the probability-based Gaussian mixture model (GMM)) is applied to a large number of operating condition parameter vectors in the historical data. The goal is to automatically group similar operating conditions into the same cluster (category), thereby identifying several typical and representative substation operating condition regions.
[0112] Step 2 (Partitioned Regression): Then, within each operating condition cluster (region) identified through clustering, a relatively simple regression model is established. This model can be linear regression, polynomial regression, or a simpler Gaussian process regression model or neural network model. This partitioned model is specifically designed to predict the health entropy characteristics under the conditions of this specific operating condition cluster. This approach can decompose a global, potentially very complex nonlinear modeling problem into several local, relatively simple modeling problems, which is sometimes easier to implement and interpret.
[0113] Model training and update logic: A selected baseline model is fully trained offline using a large amount of entropy feature vectors and their corresponding operating condition data collected and calculated when the system is confirmed to be in a "healthy" operating state (for example, during the initial commissioning of new equipment, after equipment overhaul and acceptance, or during long-term stable operation without any abnormal alarms). The quality of the training data and its coverage of various possible healthy operating conditions are crucial.
[0114] When the current system is confirmed to be healthy through other independent means (e.g., manual inspection confirmation or diagnostic results from other monitoring systems), and the current operating conditions or the entropy characteristics calculated based on them exhibit a certain degree of deviation from the existing baseline model's prediction, but still within the normal range, these new, verified "healthy" data points can be used to incrementally update or fine-tune the baseline model. For example, for a Gaussian process regression model, the new data points can be added to the original training dataset and the model's hyperparameters re-optimized; for a neural network model, this new data can be used for a few additional training cycles (i.e., fine-tuning). This mechanism helps the model adapt to slow drift in the health baseline over time due to factors such as normal, extremely slow aging of the equipment or minor adjustments to its operating mode. However, the frequency and triggering conditions of online updates need to be carefully set to prevent data from early failures or degradation from being mistakenly learned as the new "healthy" baseline.
[0115] 5. The degradation assessment and critical warning module compares the currently calculated entropy feature vector with the health benchmark provided by the baseline model to perform abnormal judgment and warning.
[0116] The calculation of abnormal deviation includes:
[0117] Obtain current entropy features: Obtain the latest calculated entropy feature vector, namely the "current entropy vector", from the system response entropy calculation module.
[0118] Obtain the current operating condition: obtain the current operating condition parameter vector synchronized with the "current entropy vector" from the data acquisition module or related system, that is, the "current operating condition vector".
[0119] Predicting a healthy baseline: Input the "current operating condition vector" into the pre-trained baseline model in the entropy feature baseline management module. The model outputs a prediction: the expected value of the entropy feature vector if the system is in a "healthy" state under the "current operating condition vector" (i.e., the "predicted healthy entropy baseline"), as well as (if supported by the baseline model, such as a Gaussian process regression model) uncertainty information associated with this prediction, such as the predicted variance of each entropy component or the predicted covariance matrix of the entire entropy vector.
[0120] Calculate the deviation: Calculate the deviation of the "current entropy vector" from the "predicted health entropy baseline".
[0121] Calculation method based on Mahalanobis distance (if the baseline model can provide covariance matrix information):
[0122] a. First, calculate the difference vector between the "current entropy vector" and the "predicted health entropy baseline", that is, the new vector obtained by subtracting the corresponding elements of the two one by one.
[0123] b. Then, obtain the inverse matrix of the covariance matrix of the "Predicted Health Entropy Baseline".
[0124] c. Multiply the transpose of the difference vector (that is, convert the column vector into a row vector) by the inverse matrix of the covariance matrix to obtain a resulting row vector.
[0125] d. Multiply this resulting row vector by the original difference vector (column vector) to get a single value.
[0126] Finally, take the square root of this value to get the Mahalanobis distance. The Mahalanobis distance takes into account the correlation between different components in the entropy eigenvector and their respective fluctuations. It is a statistically dimensionless distance that can more accurately reflect the degree of deviation.
[0127] Calculation method based on the normalized Euclidean distance of each dimension (if the baseline model does not provide a covariance matrix, or as a simplification):
[0128] a. For each component in the "Current Entropy Vector", calculate the difference between it and the corresponding component in the "Predicted Health Entropy Baseline".
[0129] b. Divide this difference by the standard deviation of the component in the healthy state (this standard deviation can be obtained from the training data of the baseline model or approximated by the prediction variance of the Gaussian process regression model). This gives a normalized difference.
[0130] c. Perform this normalized difference calculation on all entropy components.
[0131] d. Then, calculate the Euclidean distance of the vector of these normalized differences by squaring each normalized difference, adding up all the squared values, and taking the square root of the sum. This calculated "abnormal deviation" is a comprehensive scalar value that quantifies the overall degree to which the current system's entropy state deviates from its healthy baseline.
[0132] Examples of abnormal judgment and warning rules:
[0133] Rule 1 (Amplitude Abnormality Judgment):
[0134] Judgment logic: If the currently calculated "abnormal deviation" value is greater than a preset amplitude judgment threshold, and this state has lasted for more than a preset minimum duration (for example, the time length of 3 to 5 consecutive feature windows), then the system determines that an "amplitude abnormality" has occurred.
[0135] Threshold Setting Explanation: The amplitude judgment threshold can be set based on the statistical distribution of deviation values calculated from a large number of healthy operating samples, such as the 99th percentile of the distribution or the mean plus three standard deviations. The purpose of setting a minimum duration is to avoid false alarms caused by transient, non-continuous noise or disturbances.
[0136] Rule 2 (Trend Abnormality Judgment):
[0137] Judgment Logic: Continuously monitor the "abnormal deviation" numerical sequence within a recent period (for example, the last N feature window periods), or select the time series of one or several key entropy indicators. Using a sliding window approach, apply linear regression analysis to these numerical sequences within the window and calculate the slope of the trend line that changes over time. If the absolute value of this slope is greater than a pre-set slope judgment threshold, and this trend with a significant slope has persisted for more than a pre-set minimum trend duration, the system determines that a "trend anomaly" has occurred. This typically indicates that a degraded state of the system may be developing or accumulating.
[0138] Explanation of Threshold Setting: The slope threshold is used to define a significant rate of deterioration in deviation (or in some special cases, an abnormal "improvement" that may also indicate a problem). The minimum trend duration is used to ensure that the observed trend is stable and not a short-term fluctuation.
[0139] Rule 3 (judgment of abnormal rate or accelerated deterioration):
[0140] Judgment Logic: First, the first-order difference of the "abnormal deviation" time series is calculated, representing the rate of change of the deviation. Then, the first-order difference of this rate of change series is calculated, resulting in the second-order difference of the "abnormal deviation," which can be understood as the "acceleration" of the deviation change. If this "acceleration" exceeds a pre-set acceleration threshold and is positive (indicating an accelerating rate of increase in deviation, i.e., accelerating deterioration), the system determines that a "rate anomaly" or "accelerated deterioration" has occurred. This typically indicates that the system state may be rapidly becoming unstable or experiencing a more serious problem.
[0141] Explanation of threshold setting: The acceleration judgment threshold is used to capture the sharp increase in the deterioration trend of the system status.
[0142] Rule 4 (Criticality Warning - Based on Criticality Slowing Down Theory):
[0143] Judgment logic: Select several key entropy indicators that are particularly sensitive to the overall stability of the system or the health status of key equipment (for example, the permutation entropy or fuzzy entropy value of a specific scale, or the comprehensive "abnormal deviation" calculated previously). For the time series of these selected indicators, continuously calculate their statistical characteristics within a sliding observation window, especially their variance (variance, which measures the amplitude of data fluctuations) and their first-order autocorrelation coefficient, which measures the correlation between the current value and the value at the previous moment). If the variance value calculated within the sliding window shows a significant increasing trend and its value itself exceeds a preset variance critical threshold, and the first-order autocorrelation coefficient calculated within the sliding window shows a significant increasing trend and its value itself is very close to 1 (at least when one of these occurs), and this trend of increasing variance and increasing autocorrelation (at least one of them) persists for a certain period of time, then the system will issue a "critical transition warning."
[0144] Theoretical Background: This rule is based on the "critical slowing down" phenomenon, which is ubiquitous in complex dynamical systems. This theory states that as a system approaches a critical point (also known as a tipping point or tipping point) where its state undergoes a sudden change or instability, it takes increasingly longer to recover from small perturbations to its original equilibrium state. This weakening of resilience is typically manifested in the time series data of the system's state variables as an increase in variance and stronger autocorrelation (in particular, the time series becomes more like a random walk, resulting in a first-order autocorrelation coefficient approaching 1).
[0145] Explanation of threshold setting: The setting of the variance critical threshold and the autocorrelation critical threshold needs to be determined in combination with theoretical analysis of the system under study, numerical simulation experimental results, or statistical analysis of historical data from similar systems. The goal is to be able to effectively capture these early warning signals before the system enters a critical state.
[0146] Rule 5 (Transfer entropy anomaly judgment):
[0147] Judgment logic: For the transfer entropy value calculated in the system response entropy calculation module and used to describe the information flow between the preset key signal pairs (for example, the transfer entropy from signal X to signal Y, and the transfer entropy from signal Y to signal X):
[0148] If a calculated transfer entropy value (for example, the transfer entropy from X to Y) deviates significantly from the healthy baseline range it should have under the current operating conditions (this range is determined by the entropy feature baseline management module based on historical health data and current operating conditions, for example, it exceeds the interval consisting of the predicted mean plus or minus several standard deviations, or falls outside the extremely low or extremely high percentile of the healthy distribution),
[0149] If the asymmetry between the transfer entropy from X to Y and the transfer entropy from Y to X (for example, the difference or ratio between the two) changes significantly, deviating from the normal pattern it usually exhibits in a healthy state, the system determines that a "transfer entropy anomaly" has occurred. This may indicate a problem with the control coordination between related devices, or an undesirable increase in dynamic coupling between them, or a blockage in the transfer of information between them.
[0150] Warning level fusion logic:
[0151] Method 1 (Scorecard-based fusion): For each successfully triggered anomaly judgment rule, pre-set a weight or score representing its risk level (for example, when Rule 4, indicating a critical transition, is triggered, its score may be the highest; when Rule 1, indicating only a minor anomaly, is triggered, its score may be lower). The scores of all rules triggered during the current assessment cycle are accumulated. The final system warning level is then determined based on whether the accumulated total score falls within pre-defined score ranges (for example, a total score between 0 and 10 indicates "system status normal," 11 to 30 indicates "system status requiring attention," 31 to 60 indicates "system status alarm," and 60 or above indicates "system status severe alarm or critical warning").
[0152] Method 2 (fusion based on logical rule combination): Use a set of predefined, expert-knowledge-based logical rules in the form of IF-THEN-ELSE (if-then-else) to determine the warning level. For example, the following rules can be set: IF (Rule 4's critical warning condition is triggered) THEN the final warning level is directly determined to be "critical warning" (the highest level); ELSEIF (Rule 3's accelerated deterioration condition is triggered) THEN the final warning level is determined to be "severe alarm"; ELSEIF (Rule 2's trend anomaly condition is triggered AND Rule 1's amplitude anomaly condition is also triggered) THEN the final warning level is determined to be "alarm"; ELSEIF (Rule 1's amplitude anomaly condition is triggered OR Rule 5's transfer entropy anomaly condition is triggered) THEN the final warning level is determined to be "concern"; ELSE (i.e., if all the above conditions are not met), the final warning level is determined to be "normal."
[0153] Output: After fusion logic processing, the module ultimately outputs a comprehensive assessment of the system's overall health (e.g., a health index ranging from 0 to 100) and a clear warning level. This information also includes information about the specific anomaly detection rules that were triggered, along with the current value and evolution of the entropy indicators most relevant to these rules, for O&M personnel's reference.
[0154] Through the collaborative work of the above modules, this embodiment achieves a comprehensive analysis of the substation system response entropy, thereby enabling earlier and more accurate assessment of equipment degradation status and early warning of critical transitions, providing strong technical support for ensuring the safe and stable operation of the substation.
[0155] It should be understood that the embodiments disclosed in the present invention and the above description can enable those skilled in the art to use the present invention to implement the present invention. At the same time, the present invention is not limited to the embodiments mentioned above. It should be understood that those skilled in the art can still modify the technical solutions described in the above embodiments or replace some of the technical features therein with equivalents; and such modifications or replacements do not deviate from the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention and are all included in the scope of protection of the present invention.
Claims
1. A substation online monitoring system based on big data analysis, characterized in that: The system comprises: A data acquisition module configured to collect multi-source dynamic signals and auxiliary operating condition data from multiple monitoring points within the substation in real time; A signal preprocessing and feature window extraction module, connected to the data acquisition module, configured to preprocess the collected multi-source dynamic signals and extract feature window data that can reflect the dynamic response of the system according to a preset logic; a system response entropy calculation module, connected to the signal preprocessing and feature window extraction module, configured to calculate multiple types of system response entropy indicators for the feature window data and combine them into a composite entropy feature vector; an entropy feature baseline management module, connected to the system response entropy calculation module and the data acquisition module, configured to store the historical composite entropy feature vector and the corresponding auxiliary operating condition data, and to construct and maintain a health status baseline of the composite entropy feature vector under different auxiliary operating condition data using a machine learning model; The degradation assessment and critical warning module is connected to the system response entropy calculation module and the entropy feature baseline management module, and is configured to compare the composite entropy feature vector calculated in the current cycle with the health status baseline corresponding to the current auxiliary operating condition data, analyze their deviations, and evaluate the degradation status of substation equipment and the system critical transition risk based on multiple preset abnormality judgment rules, and output evaluation results and warning information.
2. The substation online monitoring system based on big data analysis according to claim 1, characterized in that: The system response entropy calculation module is specifically configured to calculate at least two items of multi-scale permutation entropy, multi-scale fuzzy entropy and transfer entropy for a preselected signal pair to form the composite entropy feature vector.
3. The substation online monitoring system based on big data analysis according to claim 1, characterized in that: The signal preprocessing and feature window extraction module is specifically configured to extract the feature window data by combining one or both of event triggering logic and sliding window logic, and when event triggering logic is adopted, the module is also configured to characterize the disturbance that triggers the event.
4. The substation online monitoring system based on big data analysis according to claim 3, characterized in that: The machine learning model adopted by the entropy feature baseline management module takes the auxiliary working condition data and the disturbance characterization description provided by the signal preprocessing and feature window extraction module as input, and the expected value of the composite entropy feature vector under the predicted health state and its normal fluctuation range as output, thereby establishing and dynamically adjusting the health state baseline.
5. The substation online monitoring system based on big data analysis according to claim 1, characterized in that: The degradation assessment and critical warning module is specifically configured to calculate the deviation of the current composite entropy feature vector relative to the health status baseline, and apply at least two of the amplitude abnormality judgment rules, trend abnormality judgment rules, change rate abnormality judgment rules, critical transition precursor judgment rules and information flow interaction abnormality judgment rules to conduct a comprehensive analysis of the deviation and one or both of the composite entropy feature vectors.
6. The substation online monitoring system based on big data analysis according to claim 5, characterized in that: The critical transition precursor judgment rule includes monitoring the statistical characteristics of the time series of a specific entropy indicator in the composite entropy feature vector or the time series of the deviation, wherein the statistical characteristics include at least one of the variance and the first-order autocorrelation coefficient. When the statistical characteristics show a specific evolution pattern indicating that the system is approaching an unstable state, a critical warning is triggered.
7. The substation online monitoring system based on big data analysis according to claim 5, characterized in that: The information flow interaction anomaly judgment rule is based on analyzing the transfer entropy value calculated for the preselected key signal pairs in the composite entropy feature vector, and evaluating the health of the dynamic interaction relationship between devices or subsystems by comparing the transfer entropy value calculated in the current cycle with its health status baseline under the corresponding working conditions, or comparing the change in the transfer entropy symmetry between signal pairs.
8. The substation online monitoring system based on big data analysis according to claim 1, characterized in that: The multi-source dynamic signals include at least one selected from the instantaneous waveforms of voltage and current on the high, medium and low voltage sides of the transformer, the bus voltage and system frequency, and the input, output and internal state signal groups of the key control system; the auxiliary operating condition data include at least one selected from the system active and reactive load levels, the grid operation topology structure code, and the environmental parameter group.
Citation Information
Patent Citations
Electrical operation risk assessment and early warning system based on big data analysis
CN119539496A
Ecological environment health intelligent assessment and management system based on PSR and entropy weight dynamic self-adaption
CN120088113A