Electricity utilization safety and fault detection method and system based on big data analysis

By collecting and analyzing multi-dimensional electricity consumption data and utilizing density clustering and machine learning techniques, abnormal electricity consumption states in the power system can be identified, solving the problem that existing technologies cannot accurately identify and trace the root cause of faults, and realizing proactive predictive maintenance of the power system.

CN121542782APending Publication Date: 2026-02-17XIAMEN ZHIDIAN FUTURE INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511758966.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-27
Publication Date
2026-02-17

AI Technical Summary

Technical Problem

Existing technologies cannot effectively identify early, hidden abnormal power consumption conditions in power systems, resulting in high false alarm and missed alarm rates, difficulty in tracing the root cause of faults, and a lack of predictive maintenance, leading to passive and lagging safety management.

Method used

By collecting multi-dimensional electricity consumption data and performing digital filtering, a density clustering algorithm is used to group and construct an association strength matrix. Anomaly candidate regions are compared, and the root cause is determined by using a machine learning classification model and an electricity consumption knowledge graph. Control signals are then generated for intervention.

Benefits of technology

It enables real-time monitoring, fault diagnosis, and root cause tracing of power systems, improves the accuracy and sensitivity of anomaly identification, and realizes the transformation from passive response to proactive predictive maintenance, thereby enhancing the system's self-adaptation and self-learning capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121542782A_ABST
    Figure CN121542782A_ABST
Patent Text Reader

Abstract

The invention provides an electricity utilization safety and fault detection method and system based on big data analysis, and is applied to the technical field of fault prediction and health management of a power system. The method comprises the steps of collecting and cleaning power consumption data, constructing an incidence matrix by density clustering grouping, comparing the incidence matrix with a reference matrix to determine an abnormal candidate area, extracting a feature subset to obtain a classification label, reasoning a knowledge graph to determine a root source, and generating control signal intervention. According to the scheme, the accuracy and the sensitivity of early electricity consumption abnormity identification can be improved, so that the false alarm rate and the missing report rate are reduced, and active predictive maintenance is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of fault prediction and health management of power systems, and particularly relates to a power safety and fault detection method and system based on big data analysis. BACKGROUND

[0002] In modern industrial production and social life, the safe and stable operation of the power system is crucial. Traditional power safety management methods, such as passive protection devices relying on leakage protectors, fuses, or fixed threshold alarms based on a single parameter such as current or voltage, have many limitations.

[0003] Firstly, these methods have a single sensing dimension and cannot mine the internal correlations between multi-dimensional data. For example, the increase in temperature and current is normally correlated under different loads of healthy equipment, but this correlation relationship will change abnormally when the line is aging. Traditional methods cannot capture this change in the "relationship" level, resulting in a large number of early hidden dangers being ignored.

[0004] Secondly, the fault diagnosis capability is weak and it is difficult to trace the root cause. When an alarm such as overcurrent occurs, the traditional system cannot distinguish whether the root cause is mechanical overload of the equipment, motor start transient, or poor line contact, making it difficult for maintenance personnel to troubleshoot and inefficient.

[0005] Thirdly, there is a lack of predictability and proactivity. Traditional methods are post or in-process responses and cannot provide early warning for slow and degenerative faults such as line aging and equipment insulation degradation, missing the maintenance opportunity and often leading to unplanned downtime and causing huge economic losses.

[0006] Finally, it cannot adapt to complex operating conditions. The normal range and mutual relationship of each electrical parameter of the same equipment are completely different under different operating conditions such as start, no load, and full load. Using fixed and one-size-fits-all thresholds for judgment is extremely prone to a large number of false positives and false negatives.

[0007] It should be noted that the information disclosed in the above background section is only used to strengthen the understanding of the background of the present application, and therefore can include information that does not constitute prior art known to those of ordinary skill in the art. SUMMARY

[0008] Therefore, the present application provides a power consumption safety and fault detection method and system based on big data analysis, aiming to solve the problems of inaccurate identification of early and hidden power consumption abnormal state, high false and missed report rate in complex and variable power consumption scenarios, superficial fault diagnosis, difficulty in tracing to the physical root cause, lack of prediction ability for fault development trend, and passive lag of safety management in the prior art. By collecting multi-dimensional power consumption data and performing digital filtering, using a density clustering algorithm to group power consumption feature vectors and constructing a power consumption correlation strength matrix, comparing the power consumption correlation strength matrix with a reference matrix to determine an abnormal candidate area, extracting a power consumption feature subset from the abnormal candidate area and inputting it into a machine learning classification model to obtain a classification label, reasoning in a power consumption knowledge graph according to the classification label to determine the root cause, and generating a control signal for a controller in the system where the power monitoring device is located to intervene in the potential power consumption abnormality, the real-time monitoring, fault diagnosis, root cause tracing and predictive maintenance of the safety state of the power consumption system are realized.

[0009] The embodiment of the present application provides a power consumption safety and fault detection method based on big data analysis, comprising: Collecting multi-dimensional power consumption data output by a power monitoring device, the multi-dimensional power consumption data representing the physical operating state and containing current, voltage and temperature parameters, and performing digital filtering processing on the multi-dimensional power consumption data to obtain cleaned data; According to the cleaned data, using a density clustering algorithm to group power consumption feature vectors, and constructing a power consumption correlation strength matrix representing the correlation characteristics between parameters based on the grouping result; Comparing the power consumption correlation strength matrix with a preset reference matrix representing historical normal power consumption state to determine an abnormal candidate area indicating a potential power consumption abnormality; Extracting a power consumption feature subset from the abnormal candidate area, and inputting the power consumption feature subset into a pre-trained machine learning classification model for processing to obtain a classification label representing an abnormal type; According to the classification label, reasoning in a preset power consumption knowledge graph to determine the root cause of the potential power consumption abnormality; According to the root cause of the potential power consumption abnormality, generating a control signal for a controller in the system where the power monitoring device is located to intervene in the potential power consumption abnormality.

[0010] In some optional embodiments, the multi-dimensional power consumption data further includes at least one of power factor, total harmonic distortion and residual current.

[0011] In some optional embodiments, the digital filtering processing uses a wavelet transform algorithm or a Kalman filtering algorithm.

[0012] In some optional embodiments, a density clustering algorithm is used to group the electricity consumption feature vectors, specifically including: clustering the electricity consumption feature vectors into data clusters that correspond to different stable operating conditions of the objects monitored by the power monitoring equipment.

[0013] In some optional embodiments, constructing an electricity consumption correlation strength matrix specifically includes: for each data cluster, independently calculating the Pearson correlation coefficient between its internal multidimensional electricity consumption data parameters to construct a correlation strength matrix corresponding to the operating conditions.

[0014] In some optional embodiments, the electricity correlation strength matrix is ​​compared with a reference matrix, specifically by calculating the Frobenius norm of the difference between the electricity correlation strength matrix and the reference matrix to obtain a scalar value that quantifies the degree of difference between the electricity correlation strength matrix and the reference matrix.

[0015] In some alternative embodiments, the subset of electricity characteristics includes at least one statistical characteristic among the mean, variance, kurtosis, or kurtosis of current, voltage, or temperature parameters within the abnormal candidate region.

[0016] In some optional embodiments, the subset of electrical characteristics may further include: identifiers for characterizing operating conditions.

[0017] In some alternative embodiments, the machine learning classification model is a gradient boosting decision tree model or a deep neural network model.

[0018] In some optional embodiments, reasoning is performed in the electricity knowledge graph, specifically including mapping category tags to entry nodes in the electricity knowledge graph and performing multi-hop retrieval along paths representing physical causal relationships to locate the root cause.

[0019] In some alternative embodiments, before generating the control signal, the following steps are also included: A long short-term memory network model is used to predict the future evolution trend of key electrical parameters related to the root cause, in order to determine the remaining effective lifespan or the time when the equipment related to the root cause reaches the failure threshold.

[0020] In some optional embodiments, it also includes: In a multi-criteria decision-making model, the remaining effective life or time to reach the failure threshold, the preset severity level of the failure consequences, and the importance level of the equipment asset are integrated to calculate the maintenance response priority. The control signals are generated based on the maintenance response priority.

[0021] In some optional embodiments, the method further includes: Perform short-time Fourier transform or wavelet transform on the time-series data of the residual current to obtain the time-frequency spectrum that characterizes the time-frequency characteristics of the time-series data of the residual current. Extract a set of frequency domain features from the time-spectrum graph. The set of frequency domain features includes at least one of the fundamental amplitude, the energy proportion of a specific subharmonic, and the energy distribution characteristics of the high-frequency band. And based on the frequency domain feature set, determine whether there is hidden leakage caused by insulation deterioration or early fault caused by line aging in the electrical circuit.

[0022] In some optional embodiments, the multidimensional electricity consumption data also includes active power, and the method further includes: Multiple active power and current data pairs collected simultaneously are fitted into a real-time power-current characteristic curve that characterizes the current operating status of a single electrical device. The real-time power-current characteristic curve is compared with a pre-stored reference power-current characteristic curve that characterizes a single electrical device in a healthy state. Furthermore, based on the morphological drift of the real-time power-current characteristic curve relative to the reference power-current characteristic curve, it can be determined whether a single electrical device has an early abnormal operating state caused by mechanical wear or changes in process load.

[0023] In some optional embodiments, the method further includes: Based on the execution feedback from the actuator, obtain validated fault diagnosis and handling case data; And by using case data, incremental training can be performed on machine learning classification models or causal relationship paths in the electricity knowledge graph can be updated.

[0024] This invention provides a power safety and fault detection system based on big data analysis, comprising: The data processing unit is configured to collect multi-dimensional electricity consumption data, including parameters such as current, voltage, and temperature, which characterize the physical operating status output by the power monitoring equipment, and to perform digital filtering on the multi-dimensional electricity consumption data to obtain cleaned data. The association modeling unit is configured to group the electricity consumption feature vectors using a density clustering algorithm based on the cleaned data, and to construct an electricity consumption association strength matrix that characterizes the association properties between parameters based on the grouping results. An anomaly location unit is configured to compare the electricity consumption correlation strength matrix with a preset reference matrix that characterizes historical normal electricity consumption status in order to determine anomaly candidate regions that indicate potential electricity consumption anomalies. The type classification unit is configured to extract a subset of electricity consumption features from anomaly candidate regions and input the subset of electricity consumption features into a pre-trained machine learning classification model for processing to obtain classification labels that characterize the anomaly type. The root cause diagnosis unit is configured to perform reasoning based on classification tags in a preset electricity consumption knowledge graph to determine the root cause of potential electricity consumption anomalies. The control command generation unit is configured to generate control signals for controlling actuators in the system where the power monitoring equipment is located, based on the root cause, in order to intervene in potential power consumption anomalies.

[0025] In some optional embodiments, the association modeling unit is specifically configured to: The electricity consumption feature vectors are clustered into data clusters that correspond to different stable operating conditions of the objects monitored by the power monitoring equipment. Furthermore, for each data cluster, the correlation strength between its internal multidimensional electricity consumption data parameters is calculated independently to construct a correlation strength matrix corresponding to the operating conditions.

[0026] In some alternative embodiments, the system further includes: The line condition diagnosis module is configured to perform time-frequency analysis on the residual current time-series data in multi-dimensional power consumption data to obtain a time-frequency spectrum diagram characterizing the time-frequency characteristics of the residual current time-series data; and to extract preset frequency domain features from the time-frequency spectrum diagram to determine whether there are early faults in the electrical line caused by insulation deterioration or line aging.

[0027] In some optional embodiments, the system is also configured to: Based on the execution feedback from the actuator, obtain validated fault diagnosis and handling case data; And by using case data, incremental training can be performed on machine learning classification models or causal relationship paths in the electricity knowledge graph can be updated.

[0028] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit the invention.

[0029] The present invention provides a method and system for power safety and fault detection based on big data analysis, which has the following beneficial effects: By employing a long short-term memory network model to predict the future evolution trends of root cause-related electrical parameters, the remaining effective lifespan of electrical equipment or the time when it reaches the failure threshold can be determined, providing a quantitative time basis for predictive maintenance. By predicting potential equipment failure risks in advance, unplanned downtime and economic losses caused by sudden failures can be avoided, thereby improving the safety, reliability, and operational efficiency of the power system. Attached Figure Description

[0030] Other features, objects, and advantages of the invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings.

[0031] Figure 1This is a flowchart of an embodiment of the present invention: a method for power safety and fault detection based on big data analysis; Figure 2 This is a schematic diagram of the structure of an electrical safety and fault detection system based on big data analysis according to an embodiment of the present invention. Detailed Implementation

[0032] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided so that the invention will be more comprehensive and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0033] Furthermore, the accompanying drawings are merely illustrative of the invention and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0034] The flowchart shown in the attached diagram is merely an illustrative example and does not necessarily include all steps. For example, some steps may be broken down, while others may be combined or partially combined. Therefore, the actual execution order may change depending on the specific circumstances.

[0035] Complex nonlinear relationships exist among equipment state parameters in power systems, exhibiting different correlation patterns under varying operating conditions. Modeling these correlation patterns can reflect the health status of the equipment. During equipment operation, changes in the strength of correlations between parameters may foreshadow potential faults. Machine learning-based classification algorithms can identify different types of abnormal states, extracting and classifying abnormal features. Constructing a graph that integrates equipment operational knowledge enables the tracing of fault root causes. Based on predictive information, fault development trends can be quantified, thus providing a basis for maintenance decisions.

[0036] like Figure 1 As shown in the figure, this invention provides a method for power safety and fault detection based on big data analysis. The method includes the following steps: Step S100: Collect multi-dimensional electricity consumption data representing the physical operating status output by the power monitoring equipment, and perform digital filtering on the multi-dimensional electricity consumption data to obtain cleaned data. In this step, the power monitoring equipment is deployed at the location of the monitored electrical equipment, for example, near the power input terminal of the motor drive system, cable joint, or near the heat-generating part of the equipment. These devices can be smart meters, multi-functional power meters, temperature sensors, etc., capable of collecting parameters such as current, voltage, and temperature of the electrical equipment in real time. The sampling frequency can be selected according to the actual application scenario and equipment characteristics. For example, for rapidly changing power parameters, the sampling frequency can be set to 1kHz, while for slowly changing parameters such as temperature, the sampling frequency can be set to 1Hz. The purpose of digital filtering is to remove noise introduced during the acquisition process and improve the accuracy of the data. In one implementation, a moving average filter can be used to smooth the data and reduce the impact of random noise. In some other optional implementations, a median filter or a Gaussian filter can also be used for data smoothing.

[0037] Step S200: Based on the cleaned data, a density clustering algorithm is used to group the electricity consumption feature vectors, and an electricity consumption correlation strength matrix characterizing the correlation between parameters is constructed based on the grouping results. In this step, firstly, the cleaned multidimensional electricity consumption data obtained in step S100 is used to construct electricity consumption feature vectors. Then, a density clustering algorithm is used to group these electricity consumption feature vectors, aiming to discover the inherent structure of the data and group electricity consumption data with similar characteristics into the same group. A feasible density clustering algorithm is DBSCAN (Density-Based Spatial Clustering of Applications with Noise), which can automatically discover clusters in the data without pre-specifying the number of clusters. Based on the grouping results, an electricity consumption correlation strength matrix characterizing the correlation between parameters is constructed. Each element of this matrix represents the correlation strength between different parameters; for example, the Pearson correlation coefficient can be used to measure the linear correlation between two parameters. In some other optional implementations, mutual information or Spearman rank correlation coefficient can be used to measure the correlation strength between parameters.

[0038] Step S300: Compare the electricity consumption correlation strength matrix with a preset reference matrix representing historical normal electricity consumption states to determine anomaly candidate regions indicating potential electricity consumption anomalies. In this step, a reference matrix library is pre-established, storing electricity consumption correlation strength matrices under different normal operating conditions. The electricity consumption correlation strength matrix constructed in step S200 is compared with the matrix in the reference matrix library, and the presence of anomalies is determined by calculating the difference between the two matrices. One way to calculate the matrix difference is to calculate the Euclidean distance between the two matrices. When the difference exceeds a preset threshold, the region is considered an anomaly candidate region, indicating that there may be potential electricity consumption anomalies in the region. In some other optional implementations, Manhattan distance or Chebyshev distance can be used to measure the difference between the matrices.

[0039] Step S400: Extract a subset of electricity consumption features from the anomaly candidate region and input the subset of electricity consumption features into a pre-trained machine learning classification model for processing to obtain classification labels representing the anomaly type. In this step, a subset of electricity consumption features is extracted from the anomaly candidate region determined in step S300. This feature subset should contain features that can effectively distinguish different types of anomalies. For example, statistical features of current, voltage, and temperature, such as mean and variance, can be extracted. Then, the extracted subset of electricity consumption features is input into a pre-trained machine learning classification model for processing. The role of this model is to classify the data in the anomaly candidate region into different anomaly types. One feasible machine learning classification model is a support vector machine (SVM). In some other optional implementations, classification models such as decision trees or random forests can be used.

[0040] Step S500: Based on the classification labels, reasoning is performed in a pre-defined electricity consumption knowledge graph to determine the root cause of potential electricity consumption anomalies. In this step, an electricity consumption knowledge graph is pre-constructed, which contains knowledge about the structure, working principle, failure modes, and causes of failures of electrical equipment. Reasoning is performed in the electricity consumption knowledge graph based on the classification labels obtained in step S400, with the aim of finding the root cause of potential electricity consumption anomalies. For example, if the classification label indicates "motor overheating," the knowledge graph can infer that the cause may be motor bearing wear or poor heat dissipation. The reasoning method can employ rule-based reasoning or probability-based reasoning. In some other optional implementations, deep learning methods can also be used for knowledge graph reasoning.

[0041] Step S600: Based on the root cause of the potential power consumption anomaly, generate control signals for controlling the actuators in the system where the power monitoring equipment is located to intervene in the potential power consumption anomaly. In this step, based on the root cause of the potential power consumption anomaly determined in step S500, generate control signals for controlling the actuators in the system where the power monitoring equipment is located. The actuators can be circuit breakers, relays, fans, etc. By controlling these actuators, potential power consumption anomalies can be intervened. For example, when motor overheating is detected, the fan can be controlled to increase heat dissipation, or the circuit breaker can be controlled to cut off the power supply. The control signals can be digital signals or analog signals. In some other optional implementations, maintenance suggestions can also be generated to prompt maintenance personnel to perform inspections and repairs.

[0042] Through the above steps, this embodiment enables real-time monitoring, fault diagnosis, and root cause tracing of the safety status of the power system. Step S100 achieves the acquisition and preprocessing of multi-dimensional data, laying the foundation for subsequent analysis; Step S200 effectively extracts the operating status characteristics of the power equipment through density clustering and the construction of a correlation strength matrix; Step S300 quickly locates candidate regions for anomalies by comparing with a reference matrix; Step S400 accurately identifies anomaly types using a machine learning classification model; Step S500 deeply traces the root cause of faults using a knowledge graph; and Step S600 automatically intervenes in anomalies by generating control signals. These steps work together to solve the technical problems existing in the prior art, such as the inability to effectively utilize the complex correlations between multi-dimensional power data, the inability to adaptively identify different operating conditions, fault diagnosis remaining at the surface level, and the lack of predictive ability for fault development trends. This achieves accurate identification and proactive intervention for early and hidden power anomalies.

[0043] Through the above solution, this embodiment can improve the accuracy and sensitivity of anomaly identification, realize the intelligent leap from "phenomenon alarm" to "root cause diagnosis", and realize the mode transformation from "passive response" to "proactive predictive maintenance", thereby enhancing the system's self-adaptation and self-learning capabilities.

[0044] In one specific implementation, based on the above embodiments, in addition to current, voltage and temperature, the power monitoring equipment also simultaneously collects power factor, total harmonic distortion and residual current data.

[0045] Power factor reflects the load characteristics and energy utilization efficiency of electrical equipment. An abnormal decrease in power factor may indicate problems such as inductive or capacitive load mismatch, or malfunction of reactive power compensation devices. Specifically, incorporating power factor into multidimensional data helps to more comprehensively assess the operating status of equipment. Total harmonic distortion (THD) reflects the degree of harmonic pollution in the power grid. Excessive harmonic content can lead to equipment overheating, increased line losses, and degraded power quality. Residual current characterizes the insulation condition of electrical circuits. Any non-zero value may indicate a risk of leakage current and is an indicator for the early detection of insulation aging and safety hazards. In some alternative implementations, zero-sequence voltage can be used instead of residual current to characterize the degree of three-phase imbalance.

[0046] Through the above solution, this embodiment can improve the comprehensiveness of the electrical safety status assessment, enabling the system to not only monitor basic electrical parameters such as voltage and current, but also to capture deeper electrical problems such as reduced power factor, excessive harmonic content, and abnormal residual current, thereby discovering potential safety hazards earlier.

[0047] In one specific implementation, the digital filtering process employs either a wavelet transform algorithm or a Kalman filter algorithm.

[0048] When the power consumption data contains impulsive interference, such as the spike current generated by frequent motor starts and stops, a wavelet transform algorithm is selected. Specifically, the original current signal is decomposed into low-frequency and high-frequency coefficients using a db4 wavelet basis. Then, an adaptive threshold function is used to quantize the high-frequency coefficients to filter out noise. Next, wavelet reconstruction is performed on the processed coefficients at each level to obtain the denoised current signal.

[0049] Alternatively, when the noise in the power consumption data mainly manifests as stationary Gaussian white noise, the Kalman filter algorithm is selected. First, a state-space model of the system is established, where the state variables include the effective voltage value and frequency, and the measured variables are the acquired instantaneous voltage values. Then, using the Extended Kalman Filter (EKF) algorithm, the state variables are recursively estimated based on the statistical characteristics of the system's process noise and measurement noise. Specifically, based on the state estimate from the previous time step and the control input, the current state value and measured value are predicted. Next, the Kalman gain is calculated, and the predicted state value is corrected based on the measured value to obtain the optimal state estimate for the current time step.

[0050] In some other alternative implementations, wavelet decomposition can be performed using the sym wavelet basis or the coif wavelet basis; alternatively, the unscented Kalman filter (UKF) algorithm can be used instead of the extended Kalman filter algorithm.

[0051] Through the above solution, this embodiment can effectively suppress various noises in power consumption data, improve the signal-to-noise ratio of the data, and thus improve the accuracy of subsequent operating condition identification and anomaly detection.

[0052] In one specific implementation, based on the above embodiments, a density clustering algorithm is used to analyze historical electricity consumption data. Specifically, the collected multi-dimensional electricity consumption data (including current, voltage, temperature, power factor, etc.) is constructed into electricity consumption feature vectors. Then, the DBSCAN algorithm is used to perform cluster analysis on these vectors. There is no need to pre-specify the number of clusters, enabling the system to automatically detect different operating states of the equipment. The parameters of the clustering process, such as the neighborhood radius Eps and the minimum cluster size MinPts, are selected through hyperparameter optimization methods such as grid search or Bayesian optimization to ensure the accuracy and stability of the clustering results. After clustering, each data cluster represents a specific operating condition. For example, for a wind turbine, this might include conditions such as "start-up," "low-speed operation," "high-speed operation," and "stop." In other optional implementations, other clustering algorithms can also be used, such as the K-means algorithm or the Gaussian Mixture Model (GMM) algorithm, as long as they can divide the electricity consumption feature vectors into different clusters with clear physical meaning.

[0053] Through the above scheme, this embodiment can automatically identify different stable operating conditions of the power monitoring equipment, providing accurate operating condition information for subsequent correlation modeling and anomaly detection, thereby improving the accuracy and sensitivity of anomaly detection.

[0054] In one specific implementation, constructing an electricity consumption correlation strength matrix involves the following steps: For each data cluster obtained through density clustering, firstly, extracting the parameter columns requiring correlation analysis from the corresponding multidimensional electricity consumption data set, such as current (I), voltage (U), temperature (T), and power factor (PF). Specifically, if the data cluster contains N data points, each parameter column corresponds to a numerical vector of length N. Then, for any two parameter columns (e.g., current and temperature), calculating the correlation strength between them using the Pearson correlation coefficient formula. Next, filling a matrix with the calculated Pearson correlation coefficient values ​​between all parameter pairs into a matrix where the rows and columns represent different parameters, and each element represents the correlation coefficient between the corresponding row and column parameters. Specifically, this matrix is ​​a symmetric matrix where the diagonal elements are 1, and the off-diagonal elements are the Pearson correlation coefficients between the corresponding row and column parameters. Finally, this matrix serves as the correlation strength matrix corresponding to the operating condition, used for subsequent anomaly detection. In other alternative implementations, in addition to the Pearson correlation coefficient, other correlation measurement methods can be selected, such as Spearman rank correlation coefficient, mutual information coefficient, etc., to meet the correlation analysis needs of different types of parameters or nonlinear relationships.

[0055] Through the above scheme, this embodiment can construct a correlation strength matrix that matches the characteristics of different operating conditions, thereby more accurately capturing the changes in the correlation between parameters under different operating conditions, and thus improving the sensitivity and accuracy of anomaly detection.

[0056] In one specific implementation, when comparing the electricity consumption correlation strength matrix with a reference matrix, the real-time correlation strength matrix is ​​first obtained. The baseline correlation strength matrix for the corresponding working conditions Then, calculate the difference between the two matrices to obtain the difference matrix. Specifically, the difference matrix Each element in All are equal to the real-time correlation strength matrix. The element at the corresponding position in Subtract the baseline correlation strength matrix The element at the corresponding position in Next, the difference matrix is ​​calculated. The Frobenius norm, this norm The calculation method is as follows: ,in, Let be the dimension of the matrix. Then, let be the Frobenius norm. As a scalar value, it quantifies the overall degree of difference between the real-time electricity consumption correlation strength matrix and the baseline matrix. In some other alternative implementations, in addition to the Frobenius norm, other matrix norms can be selected to quantify matrix differences, such as the nuclear norm or the infinity norm. Each norm reflects matrix differences from a different perspective and can be selected according to the actual application scenario.

[0057] Through the above scheme, this embodiment can use the Frobenius norm to transform the degree of deviation of the multidimensional parameter correlation into a scalar value that is easy to judge, thereby simplifying the complexity of anomaly detection and improving detection efficiency.

[0058] In one specific implementation, the electricity consumption feature subset includes: for anomaly candidate regions determined by the anomaly localization unit, the type classification unit first extracts current, voltage, and temperature parameters from the time-series data corresponding to that region. Specifically, for each parameter, its mean, variance, kurtosis, and kurtosis are calculated within a sliding window. The length of the sliding window is configurable, for example, set to 10 sampling points to capture instantaneous changes. The calculation formula is as follows: Mean:

[0059] variance:

[0060] Kuroshi:

[0061] kurtosis:

[0062] in, Indicates the parameter in the first place The value of each sampling point, This is the length of the sliding window. These statistical features, along with the condition identifier and matrix variability, constitute a feature subset, which is then input into a pre-trained machine learning classification model. In some other alternative implementations, statistical features such as median, skewness, and percentiles may be included in addition to mean, variance, kurtosis, and kurtosis.

[0063] Through the above approach, this embodiment can extract more comprehensive and distinctive abnormal features, thereby improving the accuracy of machine learning classification models in identifying abnormal types.

[0064] In one specific implementation, during the extraction of a subset of power consumption features from the anomaly candidate region, the identifier of the operating condition is first added as an independent feature to the feature vector. Specifically, if One-Hot encoding is used, when the CNC machine tool is in the "mold steel machining condition," a four-dimensional sub-vector [0, 0, 0, 1] will be added to this feature vector, corresponding to the four operating conditions: standby, spindle idling, aluminum alloy machining, and mold steel machining, respectively. Then, the complete feature vector containing the operating condition identifier is input into the GBDT model for training and prediction. In other optional implementations, the operating condition identifier can also be encoded using integers, such as 1 representing standby, 2 representing spindle idling, etc., and the GBDT model can be adjusted accordingly.

[0065] Through the above solution, this embodiment can improve the machine learning model's ability to distinguish abnormal patterns under different working conditions, thereby improving the accuracy of abnormal classification.

[0066] In one specific implementation, the machine learning classification model can employ a Gradient Boosting Decision Tree (GBDT) model. First, a historical dataset containing various types of abnormal electricity usage samples is collected and organized, with each sample clearly labeled with its abnormality type (e.g., overcurrent, undervoltage, harmonic exceedance, etc.). Specifically, an efficient GBDT framework such as XGBoost or LightGBM is used, with appropriate hyperparameters configured for tree depth, learning rate, and number of iterations, and cross-validation is employed for model training and optimization. The trained GBDT model can output the probability of a sample belonging to each abnormality type based on the input feature vector, selecting the type with the highest probability as the final classification result.

[0067] In other alternative implementations, the machine learning classification model can also employ a deep neural network (DNN) model. First, a multilayer perceptron (MLP) or convolutional neural network (CNN) is constructed, with its input layer dimension matching the dimension of the feature vector and its output layer dimension matching the number of outlier types. Then, using a deep learning framework such as TensorFlow or PyTorch, the model is trained using backpropagation, and activation functions such as ReLU or Sigmoid are used to increase the model's non-linear expressiveness. To prevent overfitting, regularization techniques such as dropout or batch normalization can be introduced. The trained DNN model can also output the probability that a sample belongs to each outlier type based on the input feature vector, selecting the type with the highest probability as the final classification result.

[0068] Through the above scheme, this embodiment can utilize the powerful nonlinear fitting capability of GBDT or DNN models to more accurately identify complex types of power consumption anomalies, thereby improving the accuracy and efficiency of fault diagnosis.

[0069] In one specific implementation, the reasoning process within the electricity knowledge graph includes the following steps. First, anomaly type classification labels, such as "spindle motor overheating" or "insufficient lubrication system oil pressure," are mapped to a specific node in the knowledge graph, serving as the starting point for reasoning. Then, the system proceeds from this starting node, performing multi-hop reasoning along predefined edges that characterize the physical structure and causal relationships of the equipment. Specifically, these edges may include the following types: Component connection relationships: For example, "spindle motor" is connected to "spindle bearing", and "spindle bearing" is connected to "lubrication system".

[0070] Cause and effect relationship: For example, "insufficient oil pressure in the lubrication system" leads to "poor lubrication of the spindle bearing", "poor lubrication of the spindle bearing" leads to "increased spindle friction", and "increased spindle friction" leads to "increased spindle motor temperature".

[0071] State evolution relationship: For example, "insulation aging" evolves into "increased leakage current", and "increased leakage current" evolves into "insulation breakdown".

[0072] The inference process employs either a depth-first search or a breadth-first search algorithm, iteratively traversing the aforementioned edges. In each iteration, the system evaluates the confidence level of the current node and adjusts it based on the edge weights. Inference stops when the confidence level falls below a preset threshold or the maximum number of hops is reached. The final result of the inference is finding a series of potential root causes leading to the initial anomaly type, along with their confidence scores. The node with the highest confidence level is considered the root cause of the potential power consumption anomaly.

[0073] In other alternative implementations, the inference process can employ different graph databases and inference engines. For example, the Neo4j graph database can be used to store the knowledge graph, and its built-in Cypher query language can be used for inference; alternatively, the Apache Jena framework can be used, leveraging its rule engine for inference. Furthermore, the inference algorithm can also employ probabilistic graphical models, such as Bayesian networks or Markov random fields, to more accurately quantify dependencies between nodes and to perform uncertain inference.

[0074] Through the above solution, this embodiment can more effectively utilize expert knowledge and equipment mechanisms to delve deeper into potential fault causes from surface phenomena, providing maintenance personnel with more accurate diagnostic information and improving troubleshooting efficiency.

[0075] In one specific implementation, before generating the control signal, the system further includes mapping the "insufficient lubrication system oil pressure" information output by the root cause diagnosis unit to the input features of a Long Short-Term Memory (LSTM) network model. This model, pre-trained with a large amount of historical data, is capable of predicting future trends in electrical parameters closely related to the health of the lubrication system, such as spindle motor current and motor winding temperature. Specifically, the LSTM model receives motor current and temperature time-series data sampled at 1-minute intervals over the past hour, along with the current operating condition identifier (e.g., "mold steel processing"), using a sliding window approach. The model then predicts the motor current and temperature values ​​every 5 minutes over the next 3 hours. Next, the system compares the predicted temperature sequence with a preset alarm threshold (e.g., 90 degrees Celsius). If the predicted temperature exceeds this threshold at any point within the next 3 hours, a potential overheating risk is identified. The system also calculates the root mean square (RMS) value of the motor current sequence and compares it with the RMS value of historical normal operating data. If the predicted RMS value is significantly higher than normal, it indicates that the motor load is too high, posing an overload risk. The system then calculates the remaining effective life (RUL) or time to failure threshold (TTF) of the electrical equipment associated with the lubrication system based on the prediction results. For example, if the predicted motor temperature will exceed the alarm threshold in 2 hours, the TTF is set to 2 hours. In other alternative implementations, in addition to the LSTM model, other types of time-series prediction models, such as Kalman filtering, support vector regression (SVR), or Transformer-based prediction models, can be used to improve prediction accuracy and robustness.

[0076] Through the above scheme, this embodiment can predict the future evolution trend of electrical parameters related to the identified root cause, thereby determining the remaining effective life of the electrical equipment or the time when it reaches the failure threshold, providing a basis for subsequent predictive maintenance decisions.

[0077] In one specific implementation, based on the above embodiments, the method further includes: First, directly mapping the predicted Remaining Effective Lifetime (RUL) from the output value of the Long Short-Term Memory (LSTM) network model to a level of 0-10. For example, an RUL less than 1 month is mapped to level 10, and an RUL greater than 1 year is mapped to level 1. Then, a preset severity level of the failure consequence is extracted from a knowledge graph, which assigns a severity level to each potential root cause of failure based on historical data, expert experience, and industry standards. For example, if the root cause is "spindle bearing wear," the consequence could be spindle seizure or even breakage, causing equipment downtime and production delays, and therefore it is rated as level 8 (1-10, with 10 being the most severe). Next, the equipment asset level is obtained from the Enterprise Asset Management (EAM) system, which assigns a level to each piece of equipment based on factors such as its value, utilization rate, and extent of involvement in the production process. For example, CNC machine tools are equipment on the production line. If they stop, the entire production line will stop. Therefore, they are rated as level 9 (1-10, with 10 being the lowest).

[0078] Specifically, the multi-criteria decision-making model uses a weighted average algorithm to calculate and maintain response priorities:

[0079] in, It is the remaining effective lifespan level. It is the severity level of the consequences of the failure. It is the equipment asset rating. , , These are the weights of these three factors, and .

[0080] For example, if , , ,and , , ,but:

[0081] Then, several priority thresholds are preset. For example, a priority greater than 8 is considered "emergency," 4-8 is considered "normal," and less than 4 is considered "general." Based on the calculated priority, corresponding control signals are generated. For "emergency" situations, the control signal immediately triggers a shutdown and alarm, while simultaneously notifying maintenance personnel; for "normal" situations, the control signal generates a maintenance work order and schedules maintenance for the next planned downtime period; for "general" situations, the control signal only records the event and awaits further observation.

[0082] In other alternative implementations, the multi-criteria decision model does not employ a weighted average algorithm, but instead uses other decision-making methods, such as the Analytic Hierarchy Process (AHP), fuzzy comprehensive evaluation, or TOPSIS. These methods can more comprehensively consider the interactions and priorities among various factors. The weighting coefficients are not determined using fixed empirical values, but rather using objective weighting methods such as entropy weighting, which automatically learn and adjust weights based on historical data.

[0083] Through the above scheme, this embodiment can comprehensively assess the urgency of maintenance based on multiple factors such as remaining effective life, fault severity, and equipment type, and generate corresponding control signals, thereby achieving more reasonable maintenance decisions, avoiding over-maintenance or under-maintenance, and further improving the reliability and operation and maintenance efficiency of the power system.

[0084] In one specific implementation, based on the above embodiments, the method further includes: the system acquiring the residual current of the CNC machine tool during operation from the data acquisition layer. The system obtains time-series data, with a sampling frequency of 1 kHz. Then, it performs a Short-Time Fourier Transform (STFT) on the acquired residual current time-series data, with the following parameters: a window length of 1024 sampling points (1.024 seconds), a frame shift of 512 sampling points (0.512 seconds), and a Hamming window as the window function. Through STFT processing, the time-spectrum diagram of the residual current is obtained. ,in Indicates time, Indicates frequency.

[0085] Specifically, a set of frequency domain features is extracted from the time-spectrum graph. The extracted frequency domain features include: Fundamental amplitude : The spectral amplitude at the fundamental frequency (usually 50Hz) of the power system.

[0086] Energy proportion of specific subharmonics Calculate the energy of the 3rd harmonic (150Hz) and the 5th harmonic (250Hz). and And calculate their total harmonic energy. The proportion in, i.e. and .energy The calculation method is based on the corresponding frequency. A nearby bandwidth Within this, the amplitude of the time-frequency spectrum is integrated: Bandwidth It can be set to, for example, 10Hz.

[0087] High-frequency energy distribution characteristics Calculate the total energy of frequency components above 1kHz. And calculate its proportion in the total energy, i.e. .

[0088] Finally, the system extracts the frequency domain feature set. The data is input into a pre-defined rule base, and the system determines whether the electrical circuit has hidden leakage caused by insulation deterioration or early faults caused by line aging, based on the preset rules. For example, if... and The value of also increases significantly at the same time, and If the value remains stable, it is determined that there is early degradation in the inter-turn insulation of the motor winding; if If the value increases significantly, it is judged that there is a high-frequency discharge phenomenon in the electrical circuit, which may be caused by aging of the circuit or loose connection.

[0089] In some alternative implementations, Continuous Wavelet Transform (CWT) can be used instead of Short-Time Fourier Transform. CWT convolves the signal using mother wavelet functions of different scales, providing better time-frequency resolution, especially in the low-frequency range. In this case, the extracted frequency domain features can be the modulus maxima of the wavelet coefficients and their corresponding scales. The rules in the rule base also need to be adjusted accordingly to accommodate the characteristics of the wavelet coefficients. In addition to energy proportion, the ratio of harmonic amplitude to fundamental amplitude can also be used as a frequency domain feature of a specific harmonic.

[0090] Through the above scheme, this embodiment can extract frequency domain features that characterize the insulation status and aging degree of electrical circuits from the time-frequency characteristics of residual current, thereby achieving effective detection of hidden leakage and early faults and reducing safety risks.

[0091] In one specific implementation, the multi-dimensional power consumption data also includes active power, and the method further includes: the system first synchronously collects active power (P) and current (I) data of the CNC machine tool from a multi-functional power meter, with a time resolution of 1 second. Specifically, in each sampling period, the system pairs the collected P and I data (… , The data is stored as a single data point within a sliding window (e.g., the most recent 60 seconds of data). Then, every 5 minutes, the system uses this window's data to fit a first-order linear model using the least squares method, resulting in a real-time power-current characteristic curve describing the current operating state. The function expression for this curve is: ,in, Represents the slope. This represents the intercept. Next, the system retrieves the baseline power-current characteristic curve of the CNC machine tool under the "mold steel machining condition" from the database, obtained by fitting historical health data in the same manner. The system calculates the shape drift of the real-time curve and the baseline curve. The shape drift is determined by both the slope drift and the intercept drift. The slope drift is defined as follows: The intercept drift is defined as The total morphological drift is defined as the weighted average of the two: ,in, and These are the weights for slope drift and intercept drift, respectively, and their values ​​can be adjusted based on experience or historical data. For example, if slope changes have a greater impact on equipment status, then [the weights can be set]. When the shape drift amount If the percentage exceeds a preset threshold (e.g., 5%), the system determines that the CNC machine tool is in an early abnormal operating state.

[0092] In other alternative implementations, in addition to using a first-order linear model, higher-order polynomial models (such as quadratic or cubic functions) can be employed for fitting to more accurately describe the nonlinear characteristics of the power-current characteristic curve. Furthermore, besides using the least squares method, other curve fitting algorithms, such as the RANSAC algorithm, can be used to improve robustness to outlier data points. Moreover, the method for calculating the morphological drift can also be adjusted. For example, the average power difference between the two curves at multiple specific current values ​​can be calculated as a measure of the morphological drift.

[0093] Through the above-described scheme, this embodiment can sensitively detect early abnormal operating conditions caused by mechanical wear or changes in process load by analyzing changes in the correlation between active power and current, thus achieving early warning of hidden faults that are difficult to detect using traditional methods. Compared with existing technologies, this embodiment can detect potential mechanical faults in equipment earlier, providing users with a longer maintenance window, avoiding unplanned downtime, and reducing maintenance costs.

[0094] In one specific implementation, based on the above embodiments, the method further includes: after completing a fault diagnosis and intervention, the system first receives feedback data from the actuator, indicating whether the intervention measures (e.g., adjusting the cutting speed of the CNC machine tool or restarting the lubrication pump) have successfully resolved the problem. Specifically, if the motor temperature returns to the normal range within a predetermined time period (e.g., 30 minutes) after adjusting the cutting speed, the intervention is considered successful. Then, the system encapsulates this complete event into a "case," which includes: multi-dimensional power consumption data at the time of the anomaly, the initial diagnostic result given by the machine learning classification model (e.g., "increased mechanical friction resistance"), the root cause inferred from the knowledge graph (e.g., "insufficient oil pressure in the lubrication system"), and the actuator's intervention measures and the final feedback result. Next, the system performs the following operations based on the case's verification results: if the case confirms that the diagnosis is correct and the intervention is effective, the system uses the case for incremental training of the machine learning model, specifically using the gradient descent algorithm to fine-tune the parameters of the GBDT model to improve its accuracy in diagnosing similar faults in the future. The goal of incremental training is to minimize the cross-entropy loss between the model's predicted results and the actual fault type in the case. Simultaneously, the system updates the weights of relevant causal paths in the knowledge graph, for example, increasing the probability that "insufficient lubrication system oil pressure" leads to "increased mechanical friction resistance." In other optional implementations, case data can also be used to train or update ontology concepts in the knowledge graph, for example, by adding new equipment types or failure modes. In yet another optional implementation, the machine learning model can be a deep neural network, and incremental training can be performed using the backpropagation algorithm.

[0095] Through the above approach, this embodiment enables the system to continuously learn and evolve from actual operational experience, thereby continuously improving its diagnostic accuracy and intervention effectiveness.

[0096] like Figure 2 As shown, this embodiment of the invention provides a power safety and fault detection system based on big data analysis, used to implement the power safety and fault detection method based on big data analysis of any of the above embodiments. The system includes: Data processing unit M100 is configured to receive multi-dimensional electricity consumption data from power monitoring equipment. This multi-dimensional electricity consumption data includes multiple parameters characterizing the operating status of the power equipment. The power monitoring equipment can be a smart meter, a multi-function power meter, a temperature sensor, or a current sensor. Data processing unit M100 is also configured to perform preprocessing operations on the received multi-dimensional electricity consumption data. Preprocessing operations include, but are not limited to: data cleaning, such as removing noise through sliding window averaging or Gaussian filtering; data transformation, such as converting non-stationary signals into stationary signals for easier subsequent analysis; and data normalization, such as scaling data of different dimensions to a uniform numerical range. In one implementation, data processing unit M100 includes a field-programmable gate array (FPGA) to accelerate the data preprocessing process. In other alternative implementations, data processing unit M100 can employ a digital signal processor (DSP) or a general-purpose CPU to implement the data preprocessing function.

[0097] The association modeling unit M200, communicatively connected to the data processing unit M100, is configured to analyze the cleaned multidimensional electricity consumption data and generate an association model characterizing the correlation between parameters. The association modeling unit M200 uses a clustering algorithm to group the electricity consumption data, grouping data points with similar characteristics into the same group. Clustering is an unsupervised learning algorithm used to discover the inherent structure in data. Based on the grouping results, the association modeling unit M200 constructs an association model that quantifies the statistical dependencies between different parameters. In one embodiment, the association modeling unit M200 uses GPU-accelerated computation. In other optional embodiments, the association modeling unit M200 can employ other statistical measures, such as mutual information or distance correlation coefficients, to construct the association model.

[0098] An anomaly location unit M300, communicatively connected to the association modeling unit M200, is configured to compare the real-time generated association model with a preset reference model. The reference model represents the parameter association characteristics of the power equipment under historical normal operating conditions. The anomaly location unit M300 calculates the difference metric between the real-time association model and the reference model and compares it with a preset threshold. If the difference metric exceeds the threshold, a potential power consumption anomaly is identified, and an anomaly candidate region is determined. The anomaly location unit M300 can use methods such as Euclidean distance, Manhattan distance, or cosine similarity to calculate the difference metric. In some optional implementations, the anomaly location unit M300 can also use statistical methods such as hypothesis testing and control charts for anomaly detection.

[0099] The type classification unit M400, communicatively connected to the anomaly localization unit M300, is configured to analyze data within anomaly candidate regions and classify anomaly types. The type classification unit M400 extracts a subset of features from the anomaly candidate regions, including statistical features, spectral features, or time-domain features extracted from electricity consumption data. The extracted features are input into a pre-trained machine learning classification model, such as a support vector machine (SVM), decision tree, or neural network. Based on the input feature subset, the machine learning classification model outputs a classification label characterizing the anomaly type, such as "overload," "short circuit," or "insulation fault." The classification label can indicate the severity and potential impact of the anomaly. In some alternative implementations, the type classification unit M400 can also use an expert system or rule engine to classify anomaly types based on predefined rules.

[0100] The root cause diagnosis unit M500, communicatively connected to the type classification unit M400, is configured to determine the root cause of potential power consumption anomalies based on classification tags. The root cause diagnosis unit M500 performs reasoning within a pre-built knowledge base containing structural information, operating principles, fault modes, and diagnostic rules for power equipment. The knowledge base can be stored in the form of a relational database, graph database, or ontology library. Based on the classification tags, the root cause diagnosis unit M500 queries the knowledge base and applies reasoning rules to locate the root cause of potential power consumption anomalies. The root cause diagnosis unit M500 outputs diagnostic conclusions, such as "motor bearing wear," "poor wiring contact," or "cooling system failure." In one embodiment, the knowledge base is an equipment maintenance manual provided by the power equipment manufacturer. In other optional embodiments, the root cause diagnosis unit M500 can also use methods such as case-based reasoning or model-based reasoning for fault diagnosis.

[0101] The control command generation unit M600, communicatively connected to the root cause diagnosis unit M500, is configured to generate control signals for controlling actuators in the system where the power monitoring equipment is located, based on diagnostic results. Actuators can be circuit breakers, relays, frequency converters, or alarms. The control command generation unit M600 converts the diagnostic conclusions into specific control commands, such as "cut off power," "reduce load," or "send alarm." These control commands are sent to the corresponding actuators to intervene in potential power consumption anomalies, preventing accidents or reducing losses. The control command generation unit M600 can generate control commands according to preset strategies or user-defined rules. In some alternative implementations, the control command generation unit M600 can also generate control commands that achieve specific objectives, such as minimizing downtime or maximizing production efficiency, based on optimization algorithms.

[0102] The aforementioned data processing unit M100, correlation modeling unit M200, anomaly location unit M300, type classification unit M400, root cause diagnosis unit M500, and control command generation unit M600 work together in a coordinated manner. First, the data processing unit M100 cleans the collected raw data, providing a high-quality data foundation for subsequent analysis. Then, the correlation modeling unit M200 extracts valuable correlation information from the cleaned data to construct a correlation model that reflects the equipment's operating status. Next, the anomaly location unit M300 compares the real-time correlation model with a reference model to quickly identify potential anomaly areas. Afterward, the type classification unit M400 performs a refined analysis of the anomaly areas to determine the specific type of anomaly. Finally, the root cause diagnosis unit M500 traces the root cause of the anomaly, and the control command generation unit M600 generates control signals to drive the actuators to intervene. This close collaboration enables a closed-loop process from data acquisition to fault diagnosis and control, which can promptly detect and address potential electrical safety hazards. It solves problems in existing technologies such as the inability to effectively utilize the complex correlations between multi-dimensional electrical data, the lack of adaptive identification capabilities for different operating conditions, fault diagnosis remaining at the surface level, and the lack of predictive capabilities for fault development trends.

[0103] Through the above solution, this embodiment can achieve comprehensive monitoring, intelligent diagnosis and active control of the operating status of power equipment, timely detection and handling of potential power safety hazards, reduction of equipment failure and downtime, improvement of power system safety and reliability, and ultimately reduction of operation and maintenance costs and improvement of production efficiency.

[0104] In one specific implementation, the correlation modeling unit M200 is configured to perform the following functions: First, using a density clustering algorithm, the collected electricity consumption feature vectors are divided into multiple data clusters, each corresponding to different stable operating conditions of the monitored object by the power monitoring equipment. For example, for a wind turbine, these operating conditions may include "start-up," "low wind speed power generation," "rated wind speed power generation," "high wind speed limiting," and "shutdown." Then, for each data cluster, the correlation modeling unit M200 independently calculates the correlation strength between its internal multidimensional electricity consumption data parameters, constructing a correlation strength matrix corresponding to that operating condition. When calculating the correlation strength, methods such as Pearson correlation coefficient, mutual information, or distance correlation coefficient can be used. For example, under the "rated wind speed power generation" condition, active power and wind speed typically have a high positive correlation, while under the "shutdown" condition, this correlation is close to zero. By establishing an independent correlation strength matrix for each operating condition, the normal operating mode of the equipment under different states can be reflected more accurately. In other alternative implementations, algorithms such as fuzzy clustering or spectral clustering can be used instead of density clustering, or methods such as dynamic time warping (DTW) can be used to measure the similarity of time series data to adapt to different data characteristics and application scenarios.

[0105] Through the above scheme, this embodiment can accurately model the power consumption characteristics of the monitored objects under different stable operating conditions, thereby improving the accuracy and reliability of anomaly detection.

[0106] In one specific implementation, the system further includes a line condition diagnosis module configured to receive residual current time-series data output by the data processing unit M100. This module contains a time-frequency analysis engine that can employ either a Short-Time Fourier Transform (STFT) algorithm or a Continuous Wavelet Transform (CWT) algorithm. When STFT is selected, a Hanning window is configured as the window function, with a window length set to one power frequency cycle (20ms) and a sliding step size set to 5ms to generate a time-frequency spectrum of the residual current. This time-frequency spectrum is presented as a two-dimensional image, where the horizontal axis represents time, the vertical axis represents frequency, and the color intensity of pixels represents the energy intensity at the corresponding time and frequency. When CWT is selected, a Morlet wavelet is configured as the mother wavelet, and the scaling factor is adaptively adjusted according to the frequency range being analyzed, similarly generating a time-frequency spectrum of the residual current. Subsequently, a pre-trained feature extractor based on a convolutional neural network (CNN) is used to extract a set of frequency domain features from the time-frequency spectrum. The CNN model comprises multiple convolutional layers, pooling layers, and fully connected layers. Its training dataset consists of a large number of residual current time-frequency spectrum samples containing different types of electrical line faults (such as insulation aging, inter-turn short circuits, and poor contact). The extracted frequency domain feature set includes the fundamental amplitude, the energy proportion of specific harmonics (such as the 3rd and 5th harmonics), and high-frequency energy distribution characteristics. These features are input into a pre-defined rule base, which integrates electrical safety standards, equipment operation experience, and expert knowledge. Based on the pattern matching results of the frequency domain features, the rule base determines whether the electrical line has hidden leakage caused by insulation degradation or early faults caused by line aging. For example, if a significant increase in the energy proportion of the 3rd harmonic is detected, and the high-frequency energy distribution is relatively dispersed, it is determined that there is a risk of insulation aging, and corresponding warning information is given. In some other optional implementations, the time-frequency analysis engine can also employ the Wigner-Vell distribution (WVD) or a modified Hilbert-Huang transform (HHT) algorithm to obtain higher resolution time-frequency spectrum maps.

[0107] Through the above solution, this embodiment can accurately identify early faults in electrical lines, avoid unexpected shutdowns caused by insulation deterioration or line aging, and thus improve the safety of the power system.

[0108] In one specific implementation, the data processing unit M100 is also configured to receive execution feedback information from actuators (e.g., lubrication system control valves of a CNC machine tool). This feedback information includes the execution result (success or failure) of lubricating oil pressure adjustment, confirmation of spindle speed adjustment command execution, and inspection and maintenance records entered by manual maintenance engineers. The feedback information is associated with the abnormal event that caused the feedback, forming a complete case data set. This case data is stored in a case database and structured according to a predefined schema, for example, using JSON format, including fields such as: event ID, device ID, timestamp, anomaly type, root cause diagnosis, remedial measures, execution result, and engineer remarks.

[0109] Machine learning classification models in the Type Classification Unit M400, such as the GBDT model, are incrementally trained periodically (e.g., weekly) using newly added case data from the case database. Incremental training employs online learning algorithms, such as stochastic gradient descent, to avoid repeated training on historical data, thereby reducing computational costs. In one implementation, each incremental training iteration uses only case data added within the most recent week, and a learning rate parameter is set to control the impact of the new data on the model.

[0110] The power consumption knowledge graph in the root cause diagnostic unit M500 is also updated periodically (e.g., monthly) based on data from the case database. Updates include adding new entities (e.g., new failure modes) and relationships (e.g., causal relationships between equipment components), as well as adjusting the weights of existing entities and relationships. In one implementation, if a causal path is validated multiple times in the case database, the weight of that path increases; if a causal path is disproven in the case database, the weight of that path decreases.

[0111] In other alternative implementations, incremental training of the machine learning classification model can employ other online learning algorithms, such as the Adaptive Moment Estimation (Adam) algorithm; knowledge graph updates can use other knowledge graph learning methods, such as the TransE algorithm. In further alternative implementations, case data can be used to adaptively adjust the dynamic thresholds in the anomaly localization unit M300; for example, if false alarms frequently occur in the Frobenius norm under a certain operating condition, the threshold for that condition can be appropriately increased. The data structure used for storing case data can also be a relational database or a NoSQL database.

[0112] Through the above approach, this embodiment can utilize execution feedback information to achieve adaptive updates of the machine learning classification model and the electricity knowledge graph, thereby continuously improving the system's diagnostic accuracy and predictive capabilities.

[0113] The above description, in conjunction with specific preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art, various simple deductions or substitutions can be made without departing from the concept of the present invention, and all such modifications and substitutions should be considered within the scope of protection of the present invention.

Claims

1. A method for power safety and fault detection based on big data analysis, characterized in that, The method comprises: collecting multi-dimensional power consumption data containing current, voltage and temperature parameters representing physical operating states from power monitoring devices, and performing digital filtering on the multi-dimensional power consumption data to obtain cleaned data; grouping power consumption feature vectors using a density clustering algorithm according to the cleaned data, and constructing a power consumption correlation strength matrix representing correlation characteristics between the parameters based on the grouping result; comparing the power consumption correlation strength matrix with a preset reference matrix representing historical normal power consumption states to determine abnormal candidate regions indicating potential power consumption abnormalities; extracting a power consumption feature subset from the abnormal candidate regions, and inputting the power consumption feature subset into a pre-trained machine learning classification model for processing to obtain a classification label representing an abnormal type; reasoning in a preset power consumption knowledge graph according to the classification label to determine a root cause of the potential power consumption abnormality; generating a control signal for controlling actuators in a system where the power monitoring devices are located to intervene in the potential power consumption abnormality according to the root cause.

2. The method of claim 1, wherein, The multi-dimensional power consumption data further includes at least one of power factor, total harmonic distortion and residual current; the method further comprises: performing short-time Fourier transform or wavelet transform on time series data of the residual current to obtain a time-frequency spectrum representing time-frequency characteristics of the time series data of the residual current; extracting a frequency domain feature set from the time-frequency spectrum, the frequency domain feature set including at least one of fundamental amplitude, energy proportion of a specific harmonic and high-frequency energy distribution characteristics; determining whether an electrical line has hidden leakage caused by insulation deterioration or early failure caused by line aging according to the frequency domain feature set.

3. The method of claim 1, wherein, The digital filtering process uses a wavelet transform algorithm or a Kalman filter algorithm.

4. The method of claim 1, wherein, The grouping of power consumption feature vectors using a density clustering algorithm specifically includes clustering the power consumption feature vectors into data clusters corresponding to different stable operating conditions of objects monitored by the power monitoring devices; The construction of the power consumption correlation strength matrix specifically includes independently calculating the Pearson correlation coefficient between internal multi-dimensional power consumption data parameters for each data cluster to construct a correlation strength matrix corresponding to the operating condition.

5. The method of claim 1, wherein, The comparison of the power consumption correlation strength matrix with the reference matrix specifically includes calculating the Frobenius norm of the difference between the power consumption correlation strength matrix and the reference matrix to obtain a scalar value quantifying the difference between the power consumption correlation strength matrix and the reference matrix.

6. The method of claim 1, wherein, The power consumption feature subset includes at least one statistical feature of the mean, variance, kurtosis or skewness of the current, voltage or temperature parameters in the abnormal candidate region.

7. The method of claim 1, wherein, The machine learning classification model is a gradient boosting decision tree model or a deep neural network model.

8. The method of claim 1, wherein, Before generating the control signal, the method further comprises: using a long short-term memory network model to predict the future evolution trend of key electrical parameters related to the root cause to determine the remaining effective life or the time to reach the failure threshold of the power consumption equipment related to the root cause; In a multi-criteria decision model, the remaining useful life or the time to reach the failure threshold, a pre-set failure consequence severity level, and a device asset importance level are fused to calculate a maintenance response priority; The control signal is generated according to the maintenance response priority.

9. The method of claim 1, wherein, The method further comprises: According to the execution feedback of the executor, verified failure diagnosis and treatment case data are obtained; And the case data is used to incrementally train the machine learning classification model or update the causal relationship path in the power consumption knowledge graph.

10. A power safety and fault detection system based on big data analytics, characterized in that, Comprise: A data processing unit configured to collect multi-dimensional power consumption data containing current, voltage and temperature parameters representing physical operating states output by a power monitoring device, and to perform digital filtering processing on the multi-dimensional power consumption data to obtain cleaned data; An association modeling unit configured to group power consumption feature vectors using a density clustering algorithm based on the cleaned data, and to construct a power consumption association strength matrix representing the association characteristics between the parameters based on the grouping results; An anomaly positioning unit configured to compare the power consumption association strength matrix with a pre-set reference matrix representing historical normal power consumption states to determine an abnormal candidate area indicating a potential power consumption anomaly; A type classification unit configured to extract a power consumption feature subset from the abnormal candidate area, and to input the power consumption feature subset into a pre-trained machine learning classification model for processing to obtain a classification label representing an anomaly type; A root cause diagnosis unit configured to infer in a pre-set power consumption knowledge graph according to the classification label to determine the root cause of the potential power consumption anomaly; A control instruction generation unit configured to generate a control signal for controlling an executor in a system where the power monitoring device is located according to the root cause to intervene in the potential power consumption anomaly.