Power data exception monitoring method and device and storage medium
By collecting multimodal data and using SARIMA model and residual distribution analysis, the problem of difficult to capture the dynamic correlation of spatially distributed nodes in the prior art is solved, and efficient abnormal monitoring and fault diagnosis of power systems and the Internet of Things is achieved.
Patent Information
- Application Number
- CN202411939082.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-23
AI Technical Summary
The existing anomaly monitoring methods are difficult to capture the complex dynamic correlation between spatially distributed nodes in power systems, communication networks and industrial Internet of Things, resulting in insufficient timeliness and accuracy of abnormal event detection, affecting the stability of the system and troubleshooting efficiency.
Multimodal data is collected, time series is constructed, preprocessed through edge calculation, prediction is performed using SARIMA model, residual distribution is generated, combined with spatial dependence and statistical analysis, abnormal information is extracted and classified.
It realizes comprehensive monitoring of space nodes, improves the accuracy of abnormal detection and fault diagnosis efficiency, can capture subtle exceptions and identify the correlation between space nodes, and improves the system's monitoring capabilities and response efficiency.
Smart Images

Figure CN120030486A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a method, device and storage medium for monitoring abnormal power data. Background Art
[0002] In the fields of power systems, communication networks, and industrial Internet of Things, abnormal monitoring of spatially distributed nodes is of great significance to the stable operation of the system. These systems are usually composed of multiple nodes, which are both geographically distributed and functionally related. Each node may be affected by internal or external factors, such as equipment failure, environmental changes, or network attacks, so timely and accurate detection and response to abnormal events is the key to ensure the efficient operation of the system.
[0003] Existing anomaly monitoring methods usually rely on a single data source, such as equipment operation data or environmental sensor data. However, this approach ignores the importance of multimodal data fusion, such as combining equipment operation data, environmental data, and status data to comprehensively analyze the status of nodes. In addition, many methods are based on static relational models and assume that the association between nodes is fixed, but in practical applications, the association between nodes may change dynamically in time and space, such as power grid load transfer or switching of communication links.
[0004] The limitations of existing methods are particularly evident when there are complex functional correlations or geographical distributions between spatial nodes. For example, in power systems, there is an obvious functional dependency between upstream transformers and downstream distribution equipment; in communication networks, link failures may cause chain anomalies in multiple nodes; and in industrial IoT, anomalies in sensor nodes may be amplified by anomalies in neighboring nodes. These complex dynamic correlations are often not captured by existing static models, resulting in insufficient timeliness and accuracy in detecting abnormal events.
[0005] In addition, existing methods usually lack the ability to comprehensively analyze spatial and temporal dependencies. For example, the anomalies of nodes in the same area may show a spatial clustering effect, while anomalies at different time points may show trend changes. Ignoring these dependencies will lead to misjudgment of anomaly detection results, thereby affecting the troubleshooting and recovery efficiency of the system.
[0006] In summary, the current anomaly monitoring methods have significant deficiencies in capturing the complex dynamic correlations between spatially distributed nodes, and are unable to meet the needs of modern power systems, communication networks, and industrial Internet of Things for efficient and accurate monitoring. This limitation not only affects the performance of anomaly detection, but also poses a potential threat to the reliability and stability of the system. Therefore, there is an urgent need for an innovative anomaly monitoring method that can comprehensively consider multimodal data fusion, dynamic spatial correlation, and time trend analysis to improve the system's monitoring capabilities and the efficiency of responding to anomalies. Summary of the invention
[0007] In order to solve the above technical problems, the present application provides a method, device and storage medium for monitoring abnormal power data.
[0008] The technical solution provided in this application is described below:
[0009] The first aspect of the present application provides a method for monitoring abnormal power data, the method comprising:
[0010] Collecting multimodal data of all spatial nodes, wherein the multimodal data includes at least equipment operation data, environmental data and status data;
[0011] Construct corresponding time series data from each item of the multimodal data;
[0012] Preprocessing the time series data by an edge computing device;
[0013] Detecting explicit anomalies in the time series data based on a preset anomaly threshold and outputting first anomaly information;
[0014] Determine a SARIMA model and its parameters for each spatial node, wherein the parameters include a seasonal autoregressive order p, a difference order d, a moving average order q, and a seasonal period s;
[0015] Inputting the time series data of each spatial node into the SARIMA model so that the SARIMA model generates data prediction values;
[0016] The data residual between the predicted value and the actual value of the data is statistically generated into a residual distribution;
[0017] constructing an anomaly check value according to the residual distribution;
[0018] Generate second abnormality information according to the abnormal check value and each residual in the residual distribution;
[0019] Extracting the spatial dependencies of all spatial nodes according to the residual distribution to obtain third abnormal information;
[0020] The faults of all spatial nodes are classified according to the first abnormal information, the second abnormal information and the third abnormal information detected.
[0021] Optionally, extracting the spatial dependencies of all spatial nodes according to the residual distribution to obtain the third abnormal information includes:
[0022] Constructing a residual distribution sequence according to the residual distribution;
[0023] The coefficient corr between the residual distribution sequences of any two spatial nodes is calculated by the following formula, and the coefficient corr is used to represent the spatial correlation of the two spatial nodes:
[0024]
[0025] Among them, e i and e j are the residual distribution sequences of node i and node j respectively. The numerator in the formula represents the covariance of the two spatial nodes, and the denominator represents the variance of the two spatial nodes.
[0026] Optionally, extracting the spatial dependencies of all spatial nodes according to the residual distribution to obtain the third abnormal information includes:
[0027] Constructing a residual distribution sequence according to the residual distribution;
[0028] Extract the statistical characteristics of the residual distribution sequence of all spatial nodes;
[0029] Use DBSCAN algorithm to perform cluster analysis on the statistical characteristics of all spatial nodes;
[0030] The clustering result obtained by the cluster analysis is compared with the spatial distribution of the spatial nodes, and the third abnormal information is generated according to the comparison result.
[0031] Optionally, also include:
[0032] Constructing a spatial weight matrix, wherein the spatial weight matrix is used to represent the spatial relationship between various spatial nodes;
[0033] Using the spatial weight matrix to perform weighted calculation on all residual distribution sequences, and obtain residual weighted values;
[0034] The coefficient corr is calculated using the residual weighted value.
[0035] Optionally, constructing an abnormality check value according to the residual distribution includes:
[0036] The mean of the residual distribution is calculated by the following formula:
[0037]
[0038] The standard deviation of the residual distribution is calculated by the following formula:
[0039]
[0040] The abnormal threshold is determined by the following formula:
[0041] T upper =μ+k·σ;
[0042] T lower =μ-k·σ;
[0043] Among them, k is the adjustment factor, e t represents the residual distribution sequence, n represents the number of data points contained in the residual distribution sequence, T upper Indicates the upper limit value, T lower Indicates the lower limit value;
[0044] The abnormality check value is calculated by the following formula:
[0045] S t =max(0,|e t -μ|-T);
[0046] Among them, T = k·σ.
[0047] Optionally, generating the second abnormality information according to the abnormality check value and each residual in the residual distribution includes:
[0048] The second abnormal information includes the degree of abnormality;
[0049] The degree of abnormality is calculated by the following formula:
[0050]
[0051] Where D t Indicates the degree of abnormality.
[0052] Optionally, the constructing of a spatial weight matrix, wherein the spatial weight matrix is used to represent the spatial relationship between various spatial nodes and includes:
[0053] Determine all spatial nodes;
[0054] Constructing a functional relationship tree between spatial nodes, wherein the functional relationship tree is used to represent the relationship type between two nodes;
[0055] constructing a correlation matrix according to the functional relationship tree;
[0056] According to the association matrix, a weight value is assigned to each node pair.
[0057] A second aspect of the present application provides a power data abnormality monitoring device, comprising:
[0058] A data acquisition unit, used to collect multimodal data of all spatial nodes, wherein the multimodal data includes at least equipment operation data, environmental data and status data;
[0059] A time series construction unit, used to construct corresponding time series data from each item of the multimodal data;
[0060] A preprocessing unit, configured to preprocess the time series data through an edge computing device;
[0061] A first output unit, configured to detect an explicit anomaly in the time series data based on a preset anomaly threshold, and output first anomaly information;
[0062] A model determination unit, used to determine the SARIMA model and its parameters for each spatial node, wherein the parameters include a seasonal autoregressive order p, a difference order d, a moving average order q, and a seasonal period s;
[0063] An input unit, which inputs the time series data of each spatial node into the SARIMA model so that the SARIMA model generates a data prediction value;
[0064] A residual statistical unit, used for calculating the data residual between the predicted value and the actual data value, and generating a residual distribution through statistics;
[0065] A check value construction unit, used to construct an abnormal check value according to the residual distribution;
[0066] A second output unit, configured to generate second abnormality information according to the abnormality check value and each residual in the residual distribution;
[0067] A third output unit is used to extract the spatial dependencies of all spatial nodes according to the residual distribution to obtain third abnormal information;
[0068] A classification unit is used to classify the faults of all spatial nodes according to the first abnormal information, the second abnormal information and the third abnormal information detected.
[0069] A third aspect of the present application provides a power data abnormality monitoring device, the device comprising:
[0070] Processor, memory, input-output unit, and bus;
[0071] The processor is connected to the memory, the input and output unit, and the bus;
[0072] The memory stores a program, and the processor calls the program to execute the first aspect and any optional method in the first aspect.
[0073] A fourth aspect of the present application provides a computer-readable storage medium, on which a program is stored. When the program is executed on a computer, the program executes the first aspect and any optional method in the first aspect.
[0074] It can be seen from the above technical solutions that this application has the following advantages:
[0075] 1. By collecting multimodal data including equipment operation data, environmental data and status data, this method can comprehensively reflect the operation status of spatial nodes and avoid the omission of abnormal detection that may be caused by a single data source.
[0076] 2. Using the residual distribution to generate anomaly check values can further mine potential anomaly information. At the same time, combined with statistical methods, the abnormal characteristics of the data can be extracted to improve the ability to capture subtle anomalies.
[0077] 3. Extracting spatial dependencies through residual distribution analysis, combined with the correlation and mutual influence between spatial nodes, can capture abnormal patterns of distributed nodes. This is particularly suitable for scenarios where there is functional correlation or geographical distribution between spatial nodes in distributed systems, and improves the ability to identify cross-node anomalies.
[0078] 4. The detected abnormal information is divided into dominant abnormalities, trend abnormalities and spatial dependency abnormalities, which can accurately classify the faults of spatial nodes. This comprehensive analysis of multi-source information helps to clarify the cause of the abnormality. It improves the efficiency of fault diagnosis and provides a reliable basis for the formulation of subsequent response strategies. BRIEF DESCRIPTION OF THE DRAWINGS
[0079] In order to more clearly illustrate the technical solution in the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0080] Figure 1 A schematic diagram of an embodiment of a method for monitoring abnormal power data provided in this application;
[0081] Figure 2 A flowchart of another embodiment of the power data abnormality monitoring method provided in this application;
[0082] Figure 3 A schematic diagram of the structure of an embodiment of the power data abnormality monitoring device provided in this application;
[0083] Figure 4 This is a schematic diagram of the structure of another embodiment of the power data anomaly monitoring device provided in this application. DETAILED DESCRIPTION
[0084] It should be noted that the method provided in this application can be applied to a terminal, a system, or a server. For example, the terminal can be a smart phone or a computer, a tablet computer, a smart TV, a smart watch, a portable computer terminal, or a fixed terminal such as a desktop computer. For the convenience of explanation, this application uses the terminal as an example for explanation.
[0085] See also Figure 1 The present application first provides an embodiment of a method, which comprises:
[0086] S101, collecting multimodal data of all spatial nodes, where the multimodal data includes at least equipment operation data, environmental data, and status data;
[0087] The collected data may include the device's voltage, current, temperature, power, load and other status information. Environmental data may include external environmental influencing parameters such as temperature, humidity, wind speed, air pressure and so on.
[0088] Status data may include equipment health status, alarm status, historical maintenance records, etc.
[0089] The acquisition tool can use sensors, data acquisition devices (such as PLC, SCADA system) or built-in monitoring modules for data acquisition. The acquisition method can obtain real-time data through communication protocols (such as Modbus, OPC-UA). Store in the data center or transmit directly to the edge computing device.
[0090] S102, constructing corresponding time series data from each item of the multimodal data;
[0091] The collected multimodal data are organized into a series of continuous observations in chronological order, such as [x1, x2, ..., xn][x_1, x_2, ..., x_n][x1, x2, ..., xn], where each data point contains a timestamp and a corresponding value.
[0092] The collected data can also be synchronized in time. Noise data can be filtered and missing values can be filled. Resampling can be performed according to the time granularity (such as seconds, minutes, hours) to ensure the consistency of time intervals. The output results include multimodal time series data of each spatial node, which is convenient for subsequent analysis.
[0093] S103, preprocessing the time series data through an edge computing device;
[0094] This step preprocesses the time series data through edge computing devices
[0095] Preprocessing methods may include:
[0096] Data cleaning: Remove outliers, duplicate values, or data that does not conform to the correct format.
[0097] Noise reduction: Remove signal noise through filtering (such as Kalman filtering or sliding average).
[0098] Normalization: Standardize or normalize data to unify the dimensions.
[0099] Feature extraction: Calculate feature values (such as maximum, minimum, mean, standard deviation) to prepare for subsequent analysis.
[0100] S104, detecting explicit anomalies in the time series data based on a preset anomaly threshold, and outputting first anomaly information;
[0101] In this step, the threshold range (such as the maximum allowable value and the minimum allowable value) is set by historical data or expert experience. When the value of the time series data exceeds the threshold range, it is recorded as an explicit anomaly. The first anomaly information includes the time, node and specific parameter value of the anomaly.
[0102] First, the cleaned and normalized data is obtained from the time series data preprocessed in step S103.
[0103] The data of each node is represented as {(t1, x1), (t2, x2), ..., (tn, xn)} where ti is the time and xi is the observation value.
[0104] The threshold can be set in the following ways:
[0105] Based on historical data, calculate the normal range of the parameter (such as mean ± 3 times the standard deviation).
[0106] Alternatively, set the maximum and minimum values based on equipment operating specifications or engineering experience.
[0107] Iterate over each point (ti, xi) of the time series data.
[0108] Determine whether it exceeds the threshold range: if it exceeds, mark it as an abnormal point.
[0109] For each outlier point (ti, xi)(t_i, x_i)(ti, xi), record:
[0110] Timestamp: the time ti when the exception occurred.
[0111] Node: The node_id to which the abnormal data belongs.
[0112] Parameter value: outlier value xi.
[0113] Exception type: such as exceeding the upper limit or the lower limit. It can be expressed in structured data form. For example, an example expressed in JSON format is as follows:
[0114]
[0115]
[0116] Here is a pseudocode example:
[0117]
[0118] An example format for time series data is as follows:
[0119]
[0120] In the above example, the first exception information is returned after calling the detect_explicit_anomalies method:
[0121]
[0122] S105, determining a SARIMA model and its parameters for each spatial node, wherein the parameters include a seasonal autoregressive order p, a difference order d, a moving average order q, and a seasonal period s;
[0123] SARIMA is a time series analysis model that is used to capture seasonal and trend characteristics.
[0124] The following first explains the definition of each parameter:
[0125] p: seasonal autoregressive order, indicating the number of lags in the autoregression.
[0126] d: difference order, used to eliminate non-stationarity.
[0127] q: Moving average order, representing the noise part of the model.
[0128] s: Seasonal period, which represents the period length of the seasonal pattern.
[0129] In this step, historical data can be used for parameter estimation (such as AIC criterion) and model fitting. Finally, the SARIMA model and its parameters of each node will be output.
[0130] In this embodiment, before modeling, it is necessary to verify whether the time series data is stable. If it is not stable, it is necessary to perform differential processing. It is possible to visually observe whether there is trend or periodicity in the data. The Augmented Dickey-Fuller (ADF) test can also be used to determine whether the data is stable. If the test result shows that the sequence is non-stationary (p value is greater than 0.05), the data is differentiated (such as first-order difference or seasonal difference). The unit root test can be performed again on the differentiated data to ensure stability.
[0131] In this embodiment, in order to determine the optimal combination of p, d, q, s, the following steps may be used:
[0132] Iterate over all possible combinations of parameters, including the seasonal period length s.
[0133] A model is fitted for each combination, and the quality of fit of the model (such as the AIC value) is recorded.
[0134] AIC criterion: Use the Akaike Information Criterion (AIC) to evaluate the quality of the model. The smaller the AIC value, the better the model. The AIC expression is: AIC = 2k-2ln(L), where k is the number of model parameters and L is the likelihood function value of the model.
[0135] After selecting the optimal parameters, fit the SARIMA model to the data of each spatial node:
[0136] Use the selected p, d, q, s parameters. When fitting the model, generate predicted values and residuals. Check whether the residuals of the model are white noise (no significant autocorrelation). The distribution graph of the residuals and the autocorrelation function (ACF) graph can be plotted to evaluate the fitting effect. For data of multiple spatial nodes, a model can be established independently for each node. Specifically, each node corresponds to a set of time series data. Use multi-threading or distributed computing frameworks (such as concurrent.futures in Python) to process multiple nodes simultaneously to speed up the calculation. Finally, output the model parameters and fitting results of each node.
[0137] Through the above steps, the optimal SARIMA model and parameters can be systematically determined for all spatial nodes, laying the foundation for subsequent data prediction and anomaly detection. This process takes into account the stability of the data, the quality of model fitting, and the computational efficiency, ensuring that the time series characteristics of each node can be fully captured and expressed.
[0138] S106, inputting the time series data of each spatial node into the SARIMA model, so that the SARIMA model generates a data prediction value;
[0139] After preprocessing, the time series data is used to predict the data values at future time points using the SARIMA model. The model generates a sequence of predicted values for subsequent residual calculations.
[0140] S107, generating a residual distribution by statistically calculating the data residual between the predicted value and the actual value of the data;
[0141] After preprocessing, the time series data from the spatial nodes have ensured the stability of the data. And the corresponding SARIMA model and optimal parameters p, d, q, s have been established for each spatial node.
[0142] Next, load the corresponding SARIMA model and parameters for each spatial node. The first part of the historical time series data can be used as the model training set. A small amount of data is reserved for verifying the prediction effect.
[0143] Use the preset p, d, q, s parameters to fit historical data. The SARIMA model will generate forecast values y^t for multiple future time points, where each forecast point is calculated using the following formula:
[0144] Y_t=μ+φ1y_t-1+φ2y_t-2+…+θ1∈_t-1+θ2∈t-2+…;
[0145] Among them, φ represents the autoregressive coefficient, θ represents the moving average coefficient, and ∈t is the random error term.
[0146] In this embodiment, the model first fits parameters based on historical training data, and then gradually predicts future values, and after each prediction, the results are used for the next prediction.
[0147] According to the specified prediction range (such as n time points in the future), the predicted value sequence is output, and the prediction quality is evaluated by comparing some actual values with the predicted values and calculating error indicators such as mean square error (MSE) or mean absolute error (MAE).
[0148] For multiple spatial nodes, the data of each node can be input into its corresponding SARIMA model. Use multithreading or distributed computing frameworks (such as Python's concurrent.futures or Apache Spark) to make predictions for multiple nodes simultaneously. Finally, the predicted values of all nodes are combined into a complete result set, including timestamps, node IDs, and predicted values.
[0149] Here is a code example, taking the statsmodels library in Python as an example, the sample code is as follows:
[0150] from statsmodels.tsa.statespace.sarimax import SARIMAX
[0151] import pandas as pd
[0152] #Example: Single-node time series data
[0153] data=pd.Series([value1, value2, value3, ...], index=pd.date_range(start='2023-01-01', periods=len(values), freq='D'))
[0154] #SARIMA model parameters
[0155] p,d,q=1,1,1#non-seasonal parameters
[0156] P,D,Q,s=1,1,1,12#Seasonal parameters
[0157] #1. Model training
[0158] model=SARIMAX(data,order=(p,d,q),seasonal_order=(P,D,Q,s))
[0159] model_fit=model.fit(disp=False)
[0160] #2. Future predictions
[0161] forecast_steps = 10 # Forecast the next 10 steps
[0162] forecast=model_fit.forecast(steps=forecast_steps)
[0163] #3. Output predicted value
[0164] print(forecast).
[0165] S108, constructing an abnormality check value according to the residual distribution;
[0166] In this step, a set of residual data is first obtained by calculating the residual (i.e., the difference between the actual value and the predicted value), and then the residual data is statistically analyzed to calculate its mean and standard deviation to reflect the central trend and dispersion of the prediction error. On this basis, the statistical distribution of the residual is used to determine the abnormality check value, such as constructing an abnormal detection range based on the mean and multiple standard deviations, or setting the abnormal range by the quantile method. In addition, the probability that the residual exceeds the normal range can be calculated by the probability density function as an abnormality check value to measure the severity of the abnormality. Finally, according to the abnormality detection rules, the residual points that exceed the range are marked, and the information such as the time, node and residual value of the abnormal point is recorded to provide a basis for subsequent steps. This dynamic verification process realizes adaptive anomaly detection for different nodes and data patterns by combining the statistical characteristics of the residual.
[0167] In one embodiment, constructing an anomaly check value according to the residual distribution includes:
[0168] The mean of the residual distribution is calculated by the following formula:
[0169]
[0170] The standard deviation of the residual distribution is calculated by the following formula:
[0171]
[0172] The abnormal threshold is determined by the following formula:
[0173] T upper =μ+k·σ;
[0174] T lower =μ-k·σ;
[0175] Among them, k is the adjustment factor, e t represents the residual distribution sequence, n represents the number of data points contained in the residual distribution sequence, T upper Indicates the upper limit value, T lower Indicates the lower limit value;
[0176] The abnormality check value is calculated by the following formula:
[0177] S t =max(0,|e t -μ|-T);
[0178] Among them, T = k·σ.
[0179] S109, generating second abnormality information according to the abnormal check value and each residual in the residual distribution;
[0180] In this step, based on the abnormal check value constructed in the previous step, each residual in the residual distribution is checked one by one to determine whether there is an abnormality. The specific implementation process is as follows:
[0181] First, compare the residual value of each node with the corresponding anomaly check range. If the residual value exceeds the anomaly check range, the point is considered to be abnormal. For each abnormal residual point, record the characteristic information of the anomaly, including the time of occurrence, node location, residual value, and the degree of exceeding the check range. If the probability density function (PDF) is used, the confidence of the abnormal point can be further calculated, that is, the probability value of the residual belonging to the abnormal area. Based on the direction and degree of the residual exceeding the range, classify the anomaly. For example: If the residual value is much larger than the upper limit, it may indicate an abnormality of high node data (such as equipment overload). If the residual value is much smaller than the lower limit, it may indicate an abnormality of low node data (such as equipment failure or missing data).
[0182] Summarize the information of all abnormal points and generate a second abnormal information set, including the time, node, residual value, abnormal type, etc. of the abnormal point. If necessary, nodes with abnormalities at multiple consecutive time points can be aggregated, marked as persistent abnormalities, and attached with relevant feature information. The generated second abnormal information is stored or output to provide data support for subsequent steps (such as generating third abnormal information in combination with spatial dependency). Through the above process, potential abnormal information can be efficiently extracted from the residual distribution, and the accuracy and robustness of the detection can be ensured through quantified abnormality check values.
[0183] In a specific embodiment, the second abnormal information includes the degree of abnormality;
[0184] The degree of abnormality is calculated by the following formula:
[0185]
[0186] Where D t Indicates the degree of abnormality.
[0187] S110, extracting the spatial dependencies of all spatial nodes according to the residual distribution to obtain third abnormal information;
[0188] In this step, the dependency relationship between spatial nodes is extracted according to the residual distribution of each spatial node to capture the abnormal situation caused by spatial association and generate the third abnormal information. The specific implementation process is as follows:
[0189] Define spatial dependencies between nodes based on their geographic proximity, functional relevance, or network topology.
[0190] Proximity: For example, constructing a spatial weight matrix W based on the inverse relationship of distance.
[0191] Functionality: For example, construct a functional association weight matrix based on the power transmission relationship, communication links, etc. between nodes.
[0192] Normalize the weight values in the matrix so that the sum of the weights in each row is 1.
[0193] Perform weighted averaging on the residual distribution of each spatial node, for example:
[0194] Constructing a residual distribution sequence according to the residual distribution;
[0195] The coefficient corr between the residual distribution sequences of any two spatial nodes is calculated by the following formula, and the coefficient corr is used to represent the spatial correlation of the two spatial nodes:
[0196]
[0197] Among them, e i and e j are the residual distribution sequences of node i and node j respectively. The numerator in the formula represents the covariance of the two spatial nodes, and the denominator represents the variance of the two spatial nodes.
[0198] In this calculation method, the weighted residual can also be used to perform the calculation.
[0199] In one possible implementation, the spatial weight matrix can be constructed as follows:
[0200] Step a: determine all spatial nodes;
[0201] Define all spatial nodes according to the physical or logical characteristics of the specific system. For example:
[0202] Physical space: such as sensor locations, device nodes, network nodes, etc.
[0203] Logical space: such as data flow nodes, functional modules, etc.
[0204] Each node is assigned a unique identifier N1, N2, …, Nk for subsequent matrix construction.
[0205] Collect basic information about nodes, such as geographic coordinates, topological location, functional description, or interrelated data flow characteristics.
[0206] Step b: constructing a functional relationship tree between spatial nodes, wherein the functional relationship tree is used to represent the relationship type between two nodes;
[0207] Determine the types of functional relationships that may exist between nodes, such as:
[0208] Direct dependency: The functionality of one node depends on another node.
[0209] Data transmission relationship: the direction and strength of data flow between nodes.
[0210] Physical proximity: The geographical proximity of nodes.
[0211] A tree structure can be used to represent the functional relationship between nodes, for example, it can include:
[0212] Node: The vertices of the tree represent spatial nodes.
[0213] Edge: The edge of a tree represents the functional relationship between two nodes.
[0214] Weight: In the initial stage, the relationship type can be recorded through the label on the edge.
[0215] The relationship tree can be directly constructed based on known system design documents or topology diagrams, or it can be inferred through data analysis (such as correlation analysis and dependency modeling).
[0216] Finally, the function relationship tree T = (V, E) is output, where V is the node set and E is the relationship edge set between nodes.
[0217] Step c: constructing a correlation matrix according to the functional relationship tree;
[0218] In this step, a k×k matrix A needs to be constructed, where k is the total number of nodes. The initial value of the matrix is set to 0. Traverse the edge set E of the functional relationship tree:
[0219] If there is a functional relationship between nodes Ni and Nj, A[i][j] or A[j][i] (depending on whether the relationship is bidirectional) is assigned a value of 1.
[0220] For edges with different relationship strengths, the matrix can be filled with relationship type encoding or initial weight values (such as dependency degree, transmission frequency, etc.).
[0221] Finally, the association matrix is output: A[i][j]=w, which represents the initial relationship strength w between nodes Ni and Nj.
[0222] Step d: assign a weight value to each node pair according to the association matrix.
[0223] According to the strength of the relationship between the node pairs, the weight value calculation method can be determined in the following ways:
[0224] Geographic distance: Based on the physical location of the node, the weight value is inversely proportional to the distance.
[0225] Relationship strength: Weighted using the initial weight value in the functional relationship tree.
[0226] Data correlation: Use time series or system operation data to calculate correlation coefficients (such as Pearson correlation coefficient).
[0227] Through the above steps, the constructed spatial weight matrix W comprehensively considers the functional relationship and correlation between nodes, laying the foundation for further spatial dependency analysis and anomaly detection. For example, it can be used to analyze key nodes in the network, identify high-risk areas, or optimize system resource allocation.
[0228] More specifically, cluster analysis can be performed on the abnormal conditions of multiple neighboring nodes to identify abnormalities that may be caused by regional problems (such as equipment group failure or regional external interference).
[0229] S111. Classify the faults of all spatial nodes according to the first abnormal information, the second abnormal information, and the third abnormal information detected.
[0230] First, define the fault classification criteria
[0231] Obvious abnormal fault (first abnormal information):
[0232] If the node operating parameters exceed the preset threshold range, it will be directly marked as an explicit fault.
[0233] Predict abnormal failure (second abnormal information):
[0234] Based on the abnormal check value of time series prediction and the results of residual analysis, node failures where data deviates from the normal trend are identified.
[0235] Spatial correlation abnormal failure (third abnormal information):
[0236] The abnormal spatial dependency between nodes indicates that the failure may be caused by abnormal propagation of associated nodes or regional external factors.
[0237] Merge the first, second and third abnormal information according to the node to construct the node abnormal feature set:
[0238] Fi = {Ti(1), Ti(2), Ti(3)};
[0239] Wherein, Ti(1): the explicit anomaly marker of node i; Ti(2): the predicted anomaly marker of node i; Ti(3): the spatial anomaly marker of node i.
[0240] The above information is encoded into a feature vector Xi, for example:
[0241] Xi=[xi(1),xi(2),xi(3)];
[0242] The application fault classification algorithm makes inferences based on logic:
[0243] Explicit fault: Only when Ti(1)=1, the fault is classified as a device hardware fault.
[0244] Predicted fault: Only when Ti(2)=1, the fault is classified as a trend fault.
[0245] Spatial fault: Only when Ti(3)=1, the fault is classified as a spatial correlation fault.
[0246] When multiple anomalies exist simultaneously, they are classified as comprehensive faults based on their weights (e.g., Ti(2)>Ti(3) indicates that they are mainly trend-based).
[0247] Clustering algorithms (such as K-Means) can also be used to group nodes with similar features.
[0248] Supervised learning models (such as decision trees and random forests) can also be used to predict node failure types.
[0249] Generate a classification result for each node (such as "hardware failure", "trend failure", "spatial correlation failure" or "comprehensive failure"). If the spatial correlation abnormal nodes are densely distributed, mark the area as "regional failure".
[0250] The present application also provides a specific embodiment, in which the clustering results are obtained by clustering the statistical characteristics of the residual distribution sequences of all spatial nodes, and the clustering results are compared with the spatial distribution of the spatial nodes to generate third abnormal information. This embodiment is described in detail below.
[0251] In this embodiment, in order to deeply analyze the abnormal characteristics of spatial nodes and extract their spatial dependencies, statistical features are extracted from the residual distribution sequence of all spatial nodes, and the nodes are grouped using a clustering algorithm to further explore the potential correlation of abnormal nodes. By comparing the clustering results with the spatial distribution of the nodes, the diffusion pattern and regional characteristics of abnormal events in space can be effectively identified, thereby generating third abnormal information, providing more accurate data support and scientific basis for fault location and correlation analysis of complex systems.
[0252] See also Figure 2 , the method provided in this embodiment includes:
[0253] S201, collecting multimodal data of all spatial nodes, wherein the multimodal data includes at least equipment operation data, environmental data and status data;
[0254] S202, constructing corresponding time series data from each item of the multimodal data;
[0255] S203, preprocessing the time series data through an edge computing device;
[0256] S204, detecting explicit anomalies in the time series data based on a preset anomaly threshold, and outputting first anomaly information;
[0257] S205, determining a SARIMA model and its parameters for each spatial node, wherein the parameters include a seasonal autoregressive order p, a difference order d, a moving average order q, and a seasonal period s;
[0258] S206, inputting the time series data of each spatial node into the SARIMA model, so that the SARIMA model generates a data prediction value;
[0259] S207, generating a residual distribution by statistically calculating the data residual between the predicted value and the actual value of the data;
[0260] S208, constructing an abnormality check value according to the residual distribution;
[0261] S209, generating second abnormality information according to the abnormal check value and each residual in the residual distribution;
[0262] S210, constructing a residual distribution sequence according to the residual distribution;
[0263] In this step, the residual data of each spatial node at multiple time points are arranged in time order to form a residual distribution sequence. These sequences can reflect the abnormal trends and characteristics of the node over the entire time period. For example, for the residual value [e1, e2, …, en] of a node, where et is the residual calculated at time point t, the residual distribution sequence of the node is formed by arranging these values continuously.
[0264] S211, extracting statistical features of the residual distribution sequence of all spatial nodes;
[0265] Extract features that can describe the statistical properties of each node's residual distribution sequence, including but not limited to:
[0266] Mean: reflects the central tendency of the residual distribution.
[0267] Variance: Measures the volatility of the residual distribution.
[0268] Skewness and kurtosis: describe the shape characteristics of the residual distribution (such as symmetry and peakedness).
[0269] Quantile values: used to capture the tail behavior of the residual distribution (such as outliers).
[0270] The extracted statistical features form a high-dimensional feature vector for subsequent clustering analysis.
[0271] S212, using the DBSCAN algorithm to perform cluster analysis on the statistical characteristics of all spatial nodes;
[0272] DBSCAN (Density-Based Spatial Clustering of Applications with Noise) is a density-based clustering algorithm used to discover high-density areas and identify noise points. The specific implementation steps include:
[0273] The statistical feature vectors of all spatial nodes extracted in step S211 are used as input data.
[0274] Algorithm parameter setting: Set two key parameters:
[0275] E (neighborhood radius): defines the similarity distance between nodes.
[0276] MinPts (minimum number of points): Specifies the minimum number of points that constitute a high-density area within the E radius.
[0277] The algorithm scans all nodes, forms several clusters according to density, and marks the noise points that cannot be classified. The cluster category to which each node belongs or the information marked as a noise point is obtained.
[0278] The following is an example to illustrate:
[0279] First, input data preparation is performed. From the statistical feature vector extracted in step S211, the feature of each spatial node is represented as a d-dimensional vector Xi = [xi 1, xi2, ..., xid], where d represents the number of statistical features. For example, it may include mean, variance, skewness, and kurtosis. The feature vectors of all nodes form a data set X = {X1, X2, ..., XN}, where N is the total number of spatial nodes.
[0280] Set DBSCAN parameters, E (neighborhood radius): used to determine the similarity between nodes. You can set it by:
[0281] Calculate the distribution of k-nearest neighbor distances between feature vectors and select a suitable E value by finding the "elbow point".
[0282] Set MinPts (minimum number of points) to specify the minimum number of points required to form a high-density cluster within the radius E. Generally, MinPts = d + 1 is selected or adjusted based on the size of the data set.
[0283] Use Euclidean distance or other suitable distance metrics (such as Manhattan distance, cosine similarity, etc.) to calculate the similarity between feature vectors:
[0284]
[0285] Construct a distance matrix D = [dij], where dij represents the distance between node i and node j.
[0286] Execute the DBSCAN algorithm, and for each node Xi, calculate the number of points in its E neighborhood (including itself). If the number of points in the neighborhood is ≥ MinPts, Xi is marked as a core point.
[0287] If Xi does not meet the core point condition but is within the neighborhood of a core point, it is marked as a boundary point.
[0288] If Xi does not belong to the neighborhood of any core point, it is marked as a noise point.
[0289] Starting from any unprocessed core point, recursively expand the core points and boundary points in its neighborhood to form a cluster. Repeat this process until all points are processed.
[0290] Each spatial node is assigned to a cluster category (such as cluster C1, C2, ..., Ck), or marked as a noise point (Cnoise). The clustering result includes the category label of each node for subsequent analysis and comparison.
[0291] If the data dimension d ≥ 2d, the high-dimensional feature vector can be mapped to 2D or 3D space through dimensionality reduction algorithms (such as PCA or t-SNE) to intuitively display the clustering results. You can also use colors to distinguish different clusters and mark noise points to verify the clustering effect.
[0292] Through the above steps, DBSCAN can efficiently process the statistical characteristics of spatial nodes and form clustering results. This density-driven clustering method is suitable for processing spatial data with irregular distribution and can effectively identify noise points, providing a basis for anomaly detection and further analysis.
[0293] S213, comparing the clustering result obtained by the cluster analysis with the spatial distribution of the spatial nodes, and generating third abnormal information according to the comparison result;
[0294] By mapping the clustering results in step S212 onto the geographical or topological distribution of spatial nodes, the following abnormal patterns are identified:
[0295] Regional anomalies: Nodes in a cluster are concentrated in a specific area, which may represent a regional system failure.
[0296] Random anomalies: Noise points with unconcentrated distribution may be caused by random failures or data noise.
[0297] Cross-region correlation anomaly: Nodes from multiple different regions are distributed in the same cluster, which may indicate correlation failures in different regions.
[0298] The comparison results are comprehensively analyzed to generate the third anomaly information, including anomaly categories, node sets and their spatial distribution characteristics, to provide support for further diagnosis of the system.
[0299] S214. Classify the faults of all spatial nodes according to the detected first abnormal information, second abnormal information, and third abnormal information.
[0300] The steps not described in detail in this embodiment are implemented in a similar manner to the corresponding steps in the aforementioned embodiments and will not be described again here.
[0301] The above embodiments describe the method provided in the present application in detail. The following describes an embodiment of the device provided in the present application:
[0302] See also Figure 3 The present application provides a power data abnormality monitoring device, which is characterized by comprising:
[0303] A data collection unit 301 is used to collect multimodal data of all spatial nodes, where the multimodal data includes at least equipment operation data, environmental data and status data;
[0304] A time series construction unit 302 is used to construct corresponding time series data from each item of the multimodal data;
[0305] A preprocessing unit 303, configured to preprocess the time series data through an edge computing device;
[0306] A first output unit 304 is used to detect explicit anomalies in the time series data based on a preset anomaly threshold and output first anomaly information;
[0307] A model determination unit 305 is used to determine a SARIMA model and its parameters for each spatial node, wherein the parameters include a seasonal autoregressive order p, a difference order d, a moving average order q, and a seasonal period s;
[0308] An input unit 306 is used to input the time series data of each spatial node into the SARIMA model so that the SARIMA model generates a data prediction value;
[0309] A residual statistical unit 307, used for calculating the data residual between the predicted value and the actual value of the data, and generating a residual distribution through statistics;
[0310] A check value construction unit 308, configured to construct an abnormal check value according to the residual distribution;
[0311] A second output unit 309 is used to generate second abnormality information according to the abnormality check value and each residual in the residual distribution;
[0312] A third output unit 310 is used to extract the spatial dependencies of all spatial nodes according to the residual distribution to obtain third abnormal information;
[0313] The classification unit 311 is used to classify the faults of all spatial nodes according to the first abnormal information, the second abnormal information and the third abnormal information detected.
[0314] Optionally, the third output unit 310 is specifically used for:
[0315] Constructing a residual distribution sequence according to the residual distribution;
[0316] The coefficient corr between the residual distribution sequences of any two spatial nodes is calculated by the following formula, and the coefficient corr is used to represent the spatial correlation of the two spatial nodes:
[0317]
[0318] Among them, e i and e j are the residual distribution sequences of node i and node j respectively. The numerator in the formula represents the covariance of the two spatial nodes, and the denominator represents the variance of the two spatial nodes.
[0319] Optionally, the third output unit 310 is specifically used for:
[0320] Constructing a residual distribution sequence according to the residual distribution;
[0321] Extract the statistical characteristics of the residual distribution sequence of all spatial nodes;
[0322] Use DBSCAN algorithm to perform cluster analysis on the statistical characteristics of all spatial nodes;
[0323] The clustering result obtained by the cluster analysis is compared with the spatial distribution of the spatial nodes, and the third abnormal information is generated according to the comparison result.
[0324] Optionally, a weight matrix construction unit 312 is further included, which is used to:
[0325] Constructing a spatial weight matrix, wherein the spatial weight matrix is used to represent the spatial relationship between various spatial nodes;
[0326] Using the spatial weight matrix to perform weighted calculation on all residual distribution sequences, and obtain residual weighted values;
[0327] The coefficient corr is calculated using the residual weighted value.
[0328] Optionally, the residual statistics unit 307 is specifically used for:
[0329] The mean of the residual distribution is calculated by the following formula:
[0330]
[0331] The standard deviation of the residual distribution is calculated by the following formula:
[0332]
[0333] The abnormal threshold is determined by the following formula:
[0334] T upper =μ+k·σ;
[0335] T lower =μ-k·σ;
[0336] Among them, k is the adjustment factor, e t represents the residual distribution sequence, n represents the number of data points contained in the residual distribution sequence, T upper Indicates the upper limit value, T lower Indicates the lower limit value;
[0337] The abnormality check value is calculated by the following formula:
[0338] S t =max(0,|e t -μ|-T);
[0339] Among them, T = k·σ.
[0340] Optionally, the residual statistics unit 307 is specifically used for:
[0341] The second abnormal information includes the degree of abnormality;
[0342] The degree of abnormality is calculated by the following formula:
[0343]
[0344] Where D t Indicates the degree of abnormality.
[0345] Optionally, the weight matrix construction unit 312 is specifically used for:
[0346] Determine all spatial nodes;
[0347] Constructing a functional relationship tree between spatial nodes, wherein the functional relationship tree is used to represent the relationship type between two nodes;
[0348] constructing a correlation matrix according to the functional relationship tree;
[0349] According to the association matrix, a weight value is assigned to each node pair.
[0350] See also Figure 4 The present application also provides a power data abnormality monitoring device, comprising:
[0351] Processor 401, memory 402, input and output unit 403, bus 404;
[0352] The processor 401 is connected to the memory 402, the input and output unit 403 and the bus 404;
[0353] The memory 402 stores a program, and the processor 401 calls the program to execute any of the above methods.
[0354] The present application also relates to a computer-readable storage medium on which a program is stored, wherein when the program is run on a computer, the computer is caused to execute any of the above methods.
[0355] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0356] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0357] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0358] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0359] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, read-only memory), random access memory (RAM, random access memory), disk or optical disk and other media that can store program code.
Claims
1. A method for monitoring abnormal power data, characterized in that: The method comprises: Collecting multimodal data of all spatial nodes, wherein the multimodal data includes at least equipment operation data, environmental data and status data; Construct corresponding time series data from each item of the multimodal data; Preprocessing the time series data by an edge computing device; Detecting explicit anomalies in the time series data based on a preset anomaly threshold and outputting first anomaly information; Determine a SARIMA model and its parameters for each spatial node, wherein the parameters include a seasonal autoregressive order p, a difference order d, a moving average order q, and a seasonal period s; Inputting the time series data of each spatial node into the SARIMA model so that the SARIMA model generates data prediction values; The data residual between the predicted value and the actual value of the data is statistically generated into a residual distribution; constructing an anomaly check value according to the residual distribution; Generate second abnormality information according to the abnormal check value and each residual in the residual distribution; Extracting the spatial dependencies of all spatial nodes according to the residual distribution to obtain third abnormal information; The faults of all spatial nodes are classified according to the first abnormal information, the second abnormal information and the third abnormal information detected.
2. The power data abnormality monitoring method according to claim 1, characterized in that: The step of extracting the spatial dependencies of all spatial nodes according to the residual distribution to obtain the third abnormal information comprises: Constructing a residual distribution sequence according to the residual distribution; The coefficient corr between the residual distribution sequences of any two spatial nodes is calculated by the following formula, and the coefficient corr is used to represent the spatial correlation of the two spatial nodes: Among them, e i and e j are the residual distribution sequences of node i and node j respectively. The numerator in the formula represents the covariance of the two spatial nodes, and the denominator represents the variance of the two spatial nodes.
3. The method for monitoring abnormal power data according to claim 1, characterized in that: The step of extracting the spatial dependencies of all spatial nodes according to the residual distribution to obtain the third abnormal information comprises: Constructing a residual distribution sequence according to the residual distribution; Extract the statistical characteristics of the residual distribution sequence of all spatial nodes; Use DBSCAN algorithm to perform cluster analysis on the statistical characteristics of all spatial nodes; The clustering result obtained by the cluster analysis is compared with the spatial distribution of the spatial nodes, and the third abnormal information is generated according to the comparison result.
4. The method for monitoring abnormal power data according to claim 2, characterized in that: Also includes: Constructing a spatial weight matrix, wherein the spatial weight matrix is used to represent the spatial relationship between various spatial nodes; Using the spatial weight matrix to perform weighted calculation on all residual distribution sequences, and obtain residual weighted values; The coefficient corr is calculated using the residual weighted value.
5. The method for monitoring abnormal power data according to claim 1, characterized in that: The constructing an abnormality check value according to the residual distribution comprises: The mean of the residual distribution is calculated by the following formula: The standard deviation of the residual distribution is calculated by the following formula: The abnormal threshold is determined by the following formula: T upper =μ+k·σ; T lower =μ-k·σ; Among them, k is the adjustment factor, e t represents the residual distribution sequence, n represents the number of data points contained in the residual distribution sequence, T upper Indicates the upper limit value, T lower Indicates the lower limit value; The abnormality check value is calculated by the following formula: S t =max(0, e t -μ-T); Among them, T = k·σ.
6. The method for monitoring abnormal power data according to claim 5, characterized in that: Generating second abnormality information according to the abnormality check value and each residual in the residual distribution includes: The second abnormal information includes the degree of abnormality; The degree of abnormality is calculated by the following formula: Where D t Indicates the degree of abnormality.
7. The method for monitoring abnormal power data according to claim 4, characterized in that: The constructing of a spatial weight matrix, wherein the spatial weight matrix is used to represent the spatial relationship between various spatial nodes, comprises: Determine all spatial nodes; Constructing a functional relationship tree between spatial nodes, wherein the functional relationship tree is used to represent the relationship type between two nodes; constructing a correlation matrix according to the functional relationship tree; According to the association matrix, a weight value is assigned to each node pair.
8. A power data abnormality monitoring device, characterized in that: include: A data acquisition unit, used to collect multimodal data of all spatial nodes, wherein the multimodal data includes at least equipment operation data, environmental data and status data; A time series construction unit, used to construct corresponding time series data from each item of the multimodal data; A preprocessing unit, configured to preprocess the time series data through an edge computing device; A first output unit, configured to detect an explicit anomaly in the time series data based on a preset anomaly threshold, and output first anomaly information; A model determination unit, used to determine the SARIMA model and its parameters for each spatial node, wherein the parameters include seasonal autoregressive order p, difference order d, moving average order q and seasonal period s; An input unit, which inputs the time series data of each spatial node into the SARIMA model so that the SARIMA model generates a data prediction value; A residual statistical unit, used for calculating the data residual between the predicted value and the actual data value, and generating a residual distribution through statistics; A check value construction unit, used to construct an abnormal check value according to the residual distribution; A second output unit, configured to generate second abnormality information according to the abnormality check value and each residual in the residual distribution; A third output unit is used to extract the spatial dependencies of all spatial nodes according to the residual distribution to obtain third abnormal information; A classification unit is used to classify the faults of all spatial nodes according to the first abnormal information, the second abnormal information and the third abnormal information detected.
9. A power data abnormality monitoring device, characterized in that: The device comprises: Processor, memory, input-output unit, and bus; The processor is connected to the memory, the input and output unit, and the bus; The memory stores a program, and the processor calls the program to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a program stored thereon, wherein the program, when executed on a computer, performs the method according to any one of claims 1 to 7.
Citation Information
Cited By
Multifunctional detection system for photovoltaic module direct current cable
CN120610199A
Electric energy meter abnormity identification method and device based on LSTM model
CN120744474A
Power distribution network operation planning scheme intelligent generation method, system, equipment and medium
CN121256730A