A low-voltage area fault prediction method and system based on big data
By using big data technology to process and analyze multi-source data from low-voltage distribution areas, faults can be identified and located, solving the problems of large data volume and complex data types in low-voltage distribution areas. This enables accurate identification and rapid response to early faults, thereby improving power supply reliability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGZHOU ANDIAN MEASUREMENT & CONTROL TECH CO LTD
- Filing Date
- 2026-06-05
- Publication Date
- 2026-07-31
AI Technical Summary
Low-voltage distribution areas have large amounts of data and diverse types, making it difficult for traditional methods to identify early, weak fault signals, resulting in faults not being detected in time and consequently causing power outages in the areas in use.
A fault prediction method based on big data is adopted. The multi-source data of low-voltage distribution area is standardized and stored in layers through the fault prediction terminal. A data lake is constructed by combining multi-dimensional data index. High-dimensional features are extracted by real-time stream processing framework and wavelet transform. Fault identification is performed by graph neural network combined with deep learning model. Fault location is performed by combining power grid topology.
It improves data processing efficiency, can capture early weak fault signals, improve prediction accuracy, shorten fault response time, realize the transformation from passive emergency repair to proactive prevention, and improve the operation and maintenance efficiency of distribution network and power supply reliability.
Smart Images

Figure CN122490233A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of low-voltage distribution area fault prediction technology, specifically to a low-voltage distribution area fault prediction method and system based on big data. Background Technology
[0002] A low-voltage distribution area refers to the power supply area consisting of all users and lines supplied by a distribution transformer through low-voltage lines on the low-voltage side of the transformer.
[0003] The low-voltage distribution area has a large amount of data and diverse types. If early weak fault signals are difficult to identify using traditional methods, the faults in the low-voltage distribution area cannot be detected in time. When the fault signal is collected using traditional methods, the fault in the low-voltage distribution area has become more serious. At this time, the power supply area may experience a power outage due to the fault in the low-voltage distribution area. Summary of the Invention
[0004] To address the aforementioned technical problems, this paper provides a method and system for predicting low-voltage distribution area faults based on big data. This technical solution solves the problems mentioned in the background section.
[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A method for predicting low-voltage transformer area faults based on big data includes: S100. Based on the fault prediction terminal, the low-voltage distribution area is collected and processed through the power distribution network monitoring platform to obtain multi-source data of the low-voltage distribution area. The multi-source data of the low-voltage distribution area includes distribution area operation data, user electricity consumption data and environmental data. S200. Based on the fault prediction terminal, the multi-source data of the low-voltage distribution area is standardized and filtered through a hierarchical storage strategy. A data lake is constructed by combining a multi-dimensional data index and the multi-source data of the low-voltage distribution area is classified and stored. The multi-dimensional data index is set based on time, space and equipment type. S300: Based on the fault prediction terminal, a real-time stream processing framework is used to aggregate and transform the data in the data lake, and high-dimensional features related to faults are extracted by wavelet transform combined with principal component analysis. S400: Based on the fault prediction terminal, the system uses graph neural networks combined with deep learning models to identify the operating status of low-voltage distribution areas according to high-dimensional features and outputs fault level classification results. S500, based on the fault prediction terminal, determines the faulty transformer area according to the fault level classification results and transformer area operation data, and performs fault location by combining the time difference method with graph theory algorithm in combination with the power grid topology.
[0006] Preferably, step S200, based on the fault prediction terminal, standardizes the multi-source data of the low-voltage distribution area, filters the data through a hierarchical storage strategy, and constructs a data lake by combining multi-dimensional data indexes to classify and store the multi-source data of the low-voltage distribution area, specifically including the following steps: S201. Based on the fault prediction terminal, a hierarchical storage algorithm is used to calculate and process the multi-source data of the low-voltage distribution area to obtain the priority score of the multi-source data. S202. Based on the fault prediction terminal, the multi-source data of the low-voltage distribution area is allocated to the first storage layer and the second storage layer according to the priority score of the multi-source data. S203. Based on the fault prediction terminal, the priority scores of multi-source data and the preset priority thresholds are compared and judged. S204. When the priority score of multi-source data is greater than or equal to the preset priority threshold, the multi-source data of the low-voltage distribution area is stored in the first storage layer. S205. When the priority score of the multi-source data is less than the preset priority threshold, the multi-source data of the low-voltage station area is stored in the second storage layer.
[0007] Preferably, step S201, based on the fault prediction terminal, uses a hierarchical storage algorithm to calculate and process multi-source data from the low-voltage distribution area to obtain priority scores for the multi-source data, specifically including the following steps: S2011. Based on the fault prediction terminal, the multi-source data of the low-voltage distribution area is classified according to a predefined set of categories to obtain the classified low-voltage distribution area data; wherein, the categories include voltage data, current data, power data, frequency data and environmental parameter data. S2012. Based on the fault prediction terminal, according to each category, the classified data is processed to extract key feature vectors of different categories, wherein the key feature vectors include real-time score, volatility score, outlier frequency and data source importance score. S2013. Based on the fault prediction terminal, the classification of low-voltage distribution area data is processed according to the key feature vectors of different categories and the connectivity index in the power grid topology to obtain the importance score of the classification of low-voltage distribution area data. S2014. Based on the fault prediction terminal, the DBSCAN algorithm for clustering is used to detect and process the classified low-voltage distribution area data to obtain abnormal events of different types of low-voltage distribution area data. S2015. Based on the fault prediction terminal, the importance score of the classified low-voltage distribution area data is dynamically adjusted according to the abnormal events of different types of low-voltage distribution area data.
[0008] Preferably, step S300, based on the fault prediction terminal, uses a real-time stream processing framework to aggregate and transform the data in the data lake, and extracts fault-related high-dimensional features through wavelet transform combined with principal component analysis, including the following steps: S301. Based on the fault prediction terminal, multi-source data of low-voltage distribution areas in the first and second storage layers are aggregated in real time according to three dimensions: time, space and equipment type, to obtain aggregated low-voltage distribution area data. S302. Based on the fault prediction terminal, the aggregated low-voltage distribution area data is standardized to obtain unified low-voltage distribution area data. S303. Based on the fault prediction terminal, an adaptive wavelet transform method is adopted to dynamically select the optimal wavelet basis function according to the data characteristics of each storage layer. Multi-level wavelet decomposition is performed on the unified low-voltage distribution area data of each storage layer to extract approximation coefficients and detail coefficients. The seasonal components are reconstructed by accumulating the detail coefficients. Specifically, each storage layer is the first storage layer and the second storage layer. S304. Based on the fault prediction terminal, the reconstructed high-frequency and low-frequency seasonal components are organized into high-frequency feature matrices and low-frequency feature matrices, respectively, and a comprehensive feature matrix is formed by feature splicing. S305. Based on the fault prediction terminal, the principal component analysis algorithm is applied to reduce the dimensionality of the comprehensive feature matrix to obtain key fault features.
[0009] Preferably, the S400 method, based on the fault prediction terminal, identifies the operating status of low-voltage distribution areas using a graph neural network combined with a deep learning model based on high-dimensional features, and outputs fault level classification results, specifically includes the following steps: S401. Based on the fault prediction terminal, perform data extraction and processing on the database system to obtain historical data of transformer area operation data, user electricity consumption data, and historical fault records. S402. Based on the fault prediction terminal, the historical data of user electricity consumption data is standardized to obtain standard historical data. S403. Based on the fault prediction terminal, historical fault records are manually divided into several fault levels, and a support vector machine is set between every two levels. S404. Based on the fault prediction terminal, perform feature extraction processing on standard historical data to obtain historical feature data; S405. Based on the fault prediction terminal, the support vector machine is trained using historical feature data. S406. Based on the fault prediction terminal, determine whether the level classification result output by the support vector machine is consistent with the level corresponding to the historical fault record. When the number of consistent judgments reaches the preset condition, end the model training. S407. Based on the fault prediction terminal, input the real-time feature data into the trained classification model to obtain the fault level classification results output by several support vector machines.
[0010] Preferably, the S500 method, based on the fault prediction terminal, determines the faulty transformer area according to the fault level classification results and transformer area operation data, and performs fault location using a time-difference method combined with graph theory algorithms in conjunction with the power grid topology. Specifically, this includes the following steps: S501. Based on the fault prediction terminal, within a preset time period, the trained classification model is processed by data retrieval according to a preset sampling frequency to obtain several fault level classification results output by the classification model. S502. Based on the fault prediction terminal, the fault level classification results are compared with the preset fault levels. S503. Based on the fault prediction terminal, perform data calculation and processing on the level classification results of faults greater than the preset fault level and all level classification results to obtain the weight of the fault level to be verified. S504. If the weight of the fault level to be verified is greater than the preset weight threshold, based on the fault prediction terminal, the source of the operating data of the transformer area corresponding to the level classification result greater than the preset fault level is determined. S505. Based on the fault prediction terminal, determine the color used to represent the alarm level according to the weight of the fault level to be verified, and fill the color in the table corresponding to the faulty area in the form of a table. S506. Based on the fault prediction terminal, according to the set of nodes in the power grid topology, calculate the total number of shortest paths between any two nodes and the number of paths passing through each intermediate node in the shortest path, and obtain the connectivity index. S507. Based on the fault prediction terminal, determine the initial set of nodes where the fault occurs according to the fault level classification results. S508. Based on the fault prediction terminal, the time difference method is used to calculate the time difference of the fault signals collected by each monitoring node. Combined with the connectivity index, the fault range is narrowed down through graph theory algorithm, and finally the fault point is located.
[0011] Furthermore, a low-voltage distribution area fault prediction system based on big data is proposed to implement the aforementioned low-voltage distribution area fault prediction method based on big data, including: A fault prediction terminal is used to control data transmission and information interaction between various modules. The data acquisition module is used to collect multi-source data of low-voltage distribution areas in real time through the power distribution network monitoring platform. The multi-source data includes distribution area operation data, user electricity consumption data and environmental data. The data storage module is used to filter data through a hierarchical storage strategy, construct a data lake by combining multi-dimensional data indexes, and classify and store the multi-source data. The feature extraction module is used to aggregate and transform the data in the data lake using a real-time stream processing framework, and extract fault-related high-dimensional features by combining wavelet transform with principal component analysis. The diagnostic module is used to identify the operating status of the low-voltage distribution area based on the high-dimensional features, using a graph neural network combined with a deep learning model, and output the fault level classification result. The fault location module is used to determine the faulty transformer area based on the fault level classification results and the transformer area operation data, and to perform fault location by combining the time difference method with graph theory algorithm in combination with the power grid topology.
[0012] Compared with existing technologies, this invention provides a method and system for predicting low-voltage distribution area faults based on big data, which has the following beneficial effects: This invention effectively solves the problems of large data volume, complex data types, and high real-time requirements in low-voltage distribution areas by using a multi-source data fusion and hierarchical storage strategy, thereby improving data processing efficiency. It employs wavelet transform combined with principal component analysis to extract fault features, enabling the capture of early, weak fault signals that are difficult to identify using traditional methods. It utilizes graph neural networks combined with deep learning models to achieve accurate fault level classification, improving prediction accuracy. It achieves precise fault location through time-difference method combined with graph theory algorithms, narrowing the investigation scope and shortening fault response time. Overall, it realizes a shift from passive emergency repair to proactive prevention, significantly improving the operation and maintenance efficiency and power supply reliability of the distribution network. Attached Figure Description
[0013] Figure 1 This is a flowchart illustrating steps S100-S500 in a low-voltage distribution area fault prediction method based on big data proposed in this invention. Figure 2 This is a structural block diagram of a low-voltage distribution area fault prediction system based on big data proposed in this invention. Detailed Implementation
[0014] The following description is intended to disclose the invention and enable those skilled in the art to implement it. The preferred embodiments described below are merely examples, and other obvious variations will occur to those skilled in the art.
[0015] Reference Figure 1 As shown, a low-voltage distribution area fault prediction method based on big data includes: S100. Based on the fault prediction terminal, the low-voltage distribution area is collected and processed through the power distribution network monitoring platform to obtain multi-source data of the low-voltage distribution area. The multi-source data of the low-voltage distribution area includes distribution area operation data, user electricity consumption data and environmental data. In step S100, multi-source data from low-voltage distribution areas is collected in real time through the distribution network monitoring platform. This multi-source data includes distribution area operation data, user electricity consumption data, and environmental data. Specifically, the distribution network monitoring platform accesses a first server storing equipment ledgers via a ledger interface to obtain distribution area operation data. This data includes, but is not limited to, electrical parameters such as voltage, current, active power, reactive power, power factor, frequency, and harmonics, collected once per second. Simultaneously, the distribution network monitoring platform accesses a second server storing user information via a user data acquisition interface to obtain user electricity consumption data. This data includes, but is not limited to, real-time electricity consumption, load curves, and electricity consumption behavior characteristics for each user, collected every fifteen minutes. Environmental data is acquired through meteorological sensors deployed near the distribution areas, including meteorological parameters such as temperature, humidity, air pressure, wind speed, and rainfall, collected every ten minutes. All three types of data are aggregated to the data processing center through a unified data access gateway. S200. Based on the fault prediction terminal, the multi-source data of the low-voltage distribution area is standardized and filtered through a hierarchical storage strategy. A data lake is constructed by combining a multi-dimensional data index and the multi-source data of the low-voltage distribution area is classified and stored. The multi-dimensional data index is set based on time, space and equipment type. S300: Based on the fault prediction terminal, a real-time stream processing framework is used to aggregate and transform the data in the data lake, and high-dimensional features related to faults are extracted by wavelet transform combined with principal component analysis. S400: Based on the fault prediction terminal, the system uses graph neural networks combined with deep learning models to identify the operating status of low-voltage distribution areas according to high-dimensional features and outputs fault level classification results. S500, based on the fault prediction terminal, determines the faulty transformer area according to the fault level classification results and transformer area operation data, and performs fault location by combining the time difference method with graph theory algorithm in combination with the power grid topology. Example
[0016] S200, based on the fault prediction terminal, standardizes the multi-source data of the low-voltage distribution area, filters the data through a hierarchical storage strategy, and constructs a data lake by combining multi-dimensional data indexes to classify and store the multi-source data of the low-voltage distribution area. The specific steps include the following: S201. Based on the fault prediction terminal, a hierarchical storage algorithm is used to calculate and process the multi-source data of the low-voltage distribution area to obtain the priority score of the multi-source data. S202. Based on the fault prediction terminal, the multi-source data of the low-voltage distribution area is allocated to the first storage layer and the second storage layer according to the priority score of the multi-source data. S203. Based on the fault prediction terminal, the priority scores of multi-source data and the preset priority thresholds are compared and judged. S204. When the priority score of multi-source data is greater than or equal to the preset priority threshold, the multi-source data of the low-voltage distribution area is stored in the first storage layer. S205. When the priority score of the multi-source data is less than the preset priority threshold, the multi-source data of the low-voltage station area is stored in the second storage layer.
[0017] Specifically, S201, based on the fault prediction terminal, uses a hierarchical storage algorithm to calculate and process multi-source data from low-voltage distribution areas to obtain priority scores for the multi-source data, including the following steps: S2011. Based on the fault prediction terminal, the multi-source data of the low-voltage distribution area is classified according to a predefined set of categories to obtain the classified low-voltage distribution area data; wherein, the categories include voltage data, current data, power data, frequency data and environmental parameter data. S2012. Based on the fault prediction terminal, according to each category, the classified data is processed to extract key feature vectors of different categories, wherein the key feature vectors include real-time score, volatility score, outlier frequency and data source importance score. S2013. Based on the fault prediction terminal, the classification of low-voltage distribution area data is processed according to the key feature vectors of different categories and the connectivity index in the power grid topology to obtain the importance score of the classification of low-voltage distribution area data. S2014. Based on the fault prediction terminal, the DBSCAN algorithm for clustering is used to detect and process the classified low-voltage distribution area data to obtain abnormal events of different types of low-voltage distribution area data. S2015. Based on the fault prediction terminal, the importance score of the classified low-voltage distribution area data is dynamically adjusted according to the abnormal events of different types of low-voltage distribution area data. In this embodiment, the standardization process employs the Z-score standardization method, uniformly mapping data of different dimensions to a standard normal distribution space with a mean of zero and a standard deviation of one, thus eliminating the impact of dimensional differences on subsequent analysis. The specific implementation of the hierarchical storage strategy is as follows: First, a hierarchical storage algorithm is used to calculate priority scores. This algorithm classifies multi-source data according to a predefined set of categories, including voltage data, current data, power data, frequency data, and environmental parameter data. For each category, key feature vectors are extracted from the classified data. These key feature vectors include real-time performance score, volatility score, outlier frequency, and data source importance score. Specifically, the real-time performance score is calculated based on the time difference between the data's timestamp and the current moment; a smaller time difference results in a higher real-time performance score. The volatility score is calculated based on the standard deviation of the data within the most recent hour; a larger standard deviation results in a higher volatility score. The outlier frequency is calculated based on the number of times the data exceeds three times the standard deviation within the most recent 24 hours. The data source importance score is weighted according to the importance of the corresponding equipment in the power grid topology. Then, importance scores are calculated based on key feature vectors and connectivity indices in the power grid topology. The connectivity indices are calculated by combining the degree centrality and betweenness centrality of nodes. Finally, anomalous events are detected and identified using the clustering-based DBSCAN algorithm, with the neighborhood radius of the DBSCAN algorithm being... The importance score is set to 0.5, and the minimum sample size (MinPts) is set to 10. The importance score is dynamically adjusted based on detected anomalies; when an anomaly is detected, the importance score of the corresponding category increases by 30%. The importance score is compared to a preset priority threshold. When the priority score is greater than or equal to the preset threshold, it is stored in the first storage layer, which uses a Redis in-memory database to store high-priority real-time hot data. When the priority score is less than the preset threshold, it is stored in the second storage layer, which uses the Hadoop Distributed File System (HDFS) to store low-priority historical cold data. The multi-dimensional data index is set based on three dimensions: time, space, and device type. The time dimension is indexed at the granularity of year, month, day, and hour; the space dimension is indexed at the granularity of substation number and line number; and the device type dimension is indexed at the granularity of device types such as transformers, switches, and capacitors. Example
[0018] Step S300: Based on the fault prediction terminal, the data in the data lake is aggregated and transformed using a real-time stream processing framework, and high-dimensional features related to faults are extracted by combining wavelet transform with principal component analysis, including the following steps: S301. Based on the fault prediction terminal, multi-source data of low-voltage distribution areas in the first and second storage layers are aggregated in real time according to three dimensions: time, space and equipment type, to obtain aggregated low-voltage distribution area data. S302. Based on the fault prediction terminal, the aggregated low-voltage distribution area data is standardized to obtain unified low-voltage distribution area data. S303. Based on the fault prediction terminal, an adaptive wavelet transform method is adopted to dynamically select the optimal wavelet basis function according to the data characteristics of each storage layer. Multi-level wavelet decomposition is performed on the unified low-voltage distribution area data of each storage layer to extract approximation coefficients and detail coefficients. The seasonal components are reconstructed by accumulating the detail coefficients. Specifically, each storage layer is the first storage layer and the second storage layer. S304. Based on the fault prediction terminal, the reconstructed high-frequency and low-frequency seasonal components are organized into high-frequency feature matrices and low-frequency feature matrices, respectively, and a comprehensive feature matrix is formed by feature splicing. S305. Based on the fault prediction terminal, the principal component analysis algorithm is applied to reduce the dimensionality of the comprehensive feature matrix to obtain key fault features. In this embodiment, the real-time stream processing framework uses Apache Flink, with a parallelism of 8 and a checkpoint interval of 30 seconds. The specific aggregation and transformation process is as follows: multi-source data from the first and second storage layers are aggregated in real time according to three dimensions: time, space, and device type. Aggregation is performed with a five-minute sliding window in the time dimension, by substation area in the spatial dimension, and by transformer in the device type dimension. The aggregated data is then standardized to eliminate differences in data sources and dimensions. An adaptive wavelet transform method is used to dynamically select the optimal wavelet basis function based on the characteristics of the data in each storage layer. The db4 wavelet basis is used for the high-frequency real-time data in the first storage layer, and the sym8 wavelet basis is used for the low-frequency historical data in the second storage layer. Five-level wavelet decomposition is performed on the data in each storage layer to extract approximation coefficients and detail coefficients. Seasonal components are reconstructed by accumulating the detail coefficients from the second to the fifth level. The reconstructed high-frequency and low-frequency seasonal components were organized into high-frequency feature matrices and low-frequency feature matrices, respectively. The dimension of the high-frequency feature matrix was the sampling frequency multiplied by the number of features, and the dimension of the low-frequency feature matrix was the number of aggregation windows multiplied by the number of features. A comprehensive feature matrix was formed by concatenating these features. Principal component analysis was applied to reduce the dimensionality of the comprehensive feature matrix, retaining principal components with a cumulative contribution rate of 95%, and extracting key fault features. The final feature vector had a dimension of 50. Example
[0019] S400, based on the fault prediction terminal, uses graph neural networks combined with deep learning models to identify the operating status of low-voltage distribution areas according to high-dimensional features, and outputs fault level classification results. The specific steps include the following: S401. Based on the fault prediction terminal, perform data extraction and processing on the database system to obtain historical data of transformer area operation data, user electricity consumption data, and historical fault records. S402. Based on the fault prediction terminal, the historical data of user electricity consumption data is standardized to obtain standard historical data. S403. Based on the fault prediction terminal, historical fault records are manually divided into several fault levels, and a support vector machine is set between every two levels. S404. Based on the fault prediction terminal, perform feature extraction processing on standard historical data to obtain historical feature data; S405. Based on the fault prediction terminal, the support vector machine is trained using historical feature data. S406. Based on the fault prediction terminal, determine whether the level classification result output by the support vector machine is consistent with the level corresponding to the historical fault record. When the number of consistent judgments reaches the preset condition, end the model training. S407. Based on the fault prediction terminal, input the real-time feature data into the trained classification model to obtain the fault level classification results output by several support vector machines. In this embodiment, historical data of transformer area operation and user electricity consumption are acquired, with a time span of the most recent twelve months. Historical fault records corresponding to the historical data are obtained through operation logs, totaling 3,200 records. The historical data is standardized to obtain standard historical data conforming to the IEC 61968 standard. The historical fault records are artificially divided into four fault levels: Level 1 severe fault, Level 2 moderate fault, Level 3 general fault, and Level 4 minor fault. A support vector machine (SVM) is set between every two levels, for a total of three SVMs. Feature extraction is performed on the standard historical data to obtain historical feature data. The SVMs are trained based on the historical feature data, using a radial basis function kernel with a penalty parameter C set to 10. Set the threshold to 0.1, and use five-fold cross-validation to evaluate the model. Determine if the classification results output by the support vector machine (SVM) are consistent with the corresponding levels in historical fault records. If the number of consistent judgments reaches a preset condition (consistency rate greater than or equal to 90%), model training ends. Input real-time feature data into the trained classification model to obtain fault level classification results from three SVM outputs. The majority vote among the three SVM outputs is taken as the final fault level classification result. Example
[0020] S500, based on the fault prediction terminal, determines the faulty transformer area according to the fault level classification results and transformer area operation data, and performs fault location using the time difference method combined with graph theory algorithm in conjunction with the power grid topology. Specifically, the steps include: S501. Based on the fault prediction terminal, within a preset time period, the trained classification model is processed by data retrieval according to a preset sampling frequency to obtain several fault level classification results output by the classification model. S502. Based on the fault prediction terminal, the fault level classification results are compared with the preset fault levels. S503. Based on the fault prediction terminal, perform data calculation and processing on the level classification results of faults greater than the preset fault level and all level classification results to obtain the weight of the fault level to be verified. S504. If the weight of the fault level to be verified is greater than the preset weight threshold, based on the fault prediction terminal, the source of the operating data of the transformer area corresponding to the level classification result greater than the preset fault level is determined. S505. Based on the fault prediction terminal, determine the color used to represent the alarm level according to the weight of the fault level to be verified, and fill the color in the table corresponding to the faulty area in the form of a table. S506. Based on the fault prediction terminal, according to the set of nodes in the power grid topology, calculate the total number of shortest paths between any two nodes and the number of paths passing through each intermediate node in the shortest path, and obtain the connectivity index. S507. Based on the fault prediction terminal, determine the initial set of nodes where the fault occurs according to the fault level classification results. S508. Based on the fault prediction terminal, the time difference method is used to calculate the time difference of the fault signals collected by each monitoring node. Combined with the connectivity index, the fault range is narrowed down by graph theory algorithm, and finally the fault point is located. In this embodiment, within a preset time period (the most recent hour), several fault level classification results output by the classification model are acquired based on a preset sampling frequency (every five minutes). These fault level classification results are compared with a preset fault level (Level 3 general fault), and the ratio of classification results with fault levels greater than the preset level is calculated. If the calculated ratio is greater than a preset ratio (60%), the faulty transformer area is determined based on the source of the transformer area operation data corresponding to the faulty transformer area classification results. A color is determined based on the calculated ratio to represent the alarm level: yellow for ratios between 60% and 70%, orange for ratios between 70% and 80%, and red for ratios above 80%. The colors are then filled into a table corresponding to the faulty transformer area. The specific process of fault location using the time-difference method combined with graph theory algorithms, based on the power grid topology, is as follows: Given a set of nodes in the power grid topology with a total number of N nodes, the total number of shortest paths between any two nodes and the number of paths passing through intermediate nodes in the shortest paths are calculated to obtain a connectivity index. The connectivity index is measured by the betweenness centrality of the nodes. Based on the fault level classification results, the initial set of nodes where the fault occurred is determined. The initial set of nodes consists of the transformer nodes corresponding to the distribution areas with fault levels of Level 1 or 2. The time difference method is used to calculate the time difference between the fault signals collected by each monitoring node, with an accuracy of milliseconds. Combined with connectivity indicators, the Dijkstra algorithm in graph theory is used to narrow down the fault range, successively eliminating nodes with connectivity indicators below the threshold, and finally locating the fault point with a positioning accuracy down to the pole number level.
[0021] Reference Figure 2 As shown, a low-voltage distribution area fault prediction system based on big data is used to implement the aforementioned low-voltage distribution area fault prediction method based on big data, including: A fault prediction terminal is used to control data transmission and information interaction between various modules. The data acquisition module is used to collect multi-source data of low-voltage distribution areas in real time through the power distribution network monitoring platform. The multi-source data includes distribution area operation data, user electricity consumption data and environmental data. The data storage module is used to filter data through a hierarchical storage strategy, construct a data lake by combining multi-dimensional data indexes, and classify and store the multi-source data. The feature extraction module is used to aggregate and transform the data in the data lake using a real-time stream processing framework, and extract fault-related high-dimensional features by combining wavelet transform with principal component analysis. The diagnostic module is used to identify the operating status of the low-voltage distribution area based on the high-dimensional features, using a graph neural network combined with a deep learning model, and output the fault level classification result. The fault location module is used to determine the faulty transformer area based on the fault level classification results and the transformer area operation data, and to perform fault location by combining the time difference method with graph theory algorithm in combination with the power grid topology.
[0022] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention. The scope of protection claimed by the appended claims and their equivalents is defined.
Claims
1. A method for predicting faults in low-voltage distribution areas based on big data, characterized in that, include: S100. Based on the fault prediction terminal, the low-voltage distribution area is collected and processed through the power distribution network monitoring platform to obtain multi-source data of the low-voltage distribution area. The multi-source data of the low-voltage distribution area includes distribution area operation data, user electricity consumption data and environmental data. S200. Based on the fault prediction terminal, the multi-source data of the low-voltage distribution area is standardized and filtered through a hierarchical storage strategy. A data lake is constructed by combining a multi-dimensional data index and the multi-source data of the low-voltage distribution area is classified and stored. The multi-dimensional data index is set based on time, space and equipment type. S300: Based on the fault prediction terminal, a real-time stream processing framework is used to aggregate and transform the data in the data lake, and high-dimensional features related to faults are extracted by wavelet transform combined with principal component analysis. S400: Based on the fault prediction terminal, the system uses graph neural networks combined with deep learning models to identify the operating status of low-voltage distribution areas according to high-dimensional features and outputs fault level classification results. S500, based on the fault prediction terminal, determines the faulty transformer area according to the fault level classification results and transformer area operation data, and performs fault location by combining the time difference method with graph theory algorithm in combination with the power grid topology.
2. The low-voltage distribution area fault prediction method based on big data according to claim 1, characterized in that, S200, based on the fault prediction terminal, standardizes the multi-source data of the low-voltage distribution area, filters the data through a hierarchical storage strategy, and constructs a data lake by combining multi-dimensional data indexes to classify and store the multi-source data of the low-voltage distribution area. Specifically, the steps include: S201. Based on the fault prediction terminal, a hierarchical storage algorithm is used to calculate and process the multi-source data of the low-voltage distribution area to obtain the priority score of the multi-source data. S202. Based on the fault prediction terminal, the multi-source data of the low-voltage distribution area is allocated to the first storage layer and the second storage layer according to the priority score of the multi-source data. S203. Based on the fault prediction terminal, the priority scores of multi-source data and the preset priority thresholds are compared and judged. S204. When the priority score of multi-source data is greater than or equal to the preset priority threshold, the multi-source data of the low-voltage distribution area is stored in the first storage layer. S205. When the priority score of the multi-source data is less than the preset priority threshold, the multi-source data of the low-voltage station area is stored in the second storage layer.
3. The low-voltage distribution area fault prediction method based on big data according to claim 2, characterized in that, S201, based on the fault prediction terminal, uses a hierarchical storage algorithm to calculate and process multi-source data from low-voltage distribution areas to obtain priority scores for the multi-source data. This specifically includes the following steps: S2011. Based on the fault prediction terminal, the multi-source data of the low-voltage distribution area is classified according to a predefined set of categories to obtain the classified low-voltage distribution area data; wherein, the categories include voltage data, current data, power data, frequency data and environmental parameter data. S2012. Based on the fault prediction terminal, according to each category, the classified data is processed to extract key feature vectors of different categories, wherein the key feature vectors include real-time score, volatility score, outlier frequency and data source importance score. S2013. Based on the fault prediction terminal, the classification of low-voltage distribution area data is processed according to the key feature vectors of different categories and the connectivity index in the power grid topology to obtain the importance score of the classification of low-voltage distribution area data. S2014. Based on the fault prediction terminal, the DBSCAN algorithm for clustering is used to detect and process the classified low-voltage distribution area data to obtain abnormal events of different types of low-voltage distribution area data. S2015. Based on the fault prediction terminal, the importance score of the classified low-voltage distribution area data is dynamically adjusted according to the abnormal events of different types of low-voltage distribution area data.
4. The low-voltage distribution area fault prediction method based on big data according to claim 1, characterized in that, Step S300, based on the fault prediction terminal, uses a real-time stream processing framework to aggregate and transform the data in the data lake, and extracts fault-related high-dimensional features through wavelet transform combined with principal component analysis, including the following steps: S301. Based on the fault prediction terminal, multi-source data of low-voltage distribution areas in the first and second storage layers are aggregated in real time according to three dimensions: time, space and equipment type, to obtain aggregated low-voltage distribution area data. S302. Based on the fault prediction terminal, the aggregated low-voltage distribution area data is standardized to obtain unified low-voltage distribution area data. S303. Based on the fault prediction terminal, an adaptive wavelet transform method is adopted to dynamically select the optimal wavelet basis function according to the data characteristics of each storage layer. Multi-level wavelet decomposition is performed on the unified low-voltage distribution area data of each storage layer to extract approximation coefficients and detail coefficients. The seasonal components are reconstructed by accumulating the detail coefficients. Specifically, each storage layer is the first storage layer and the second storage layer. S304. Based on the fault prediction terminal, the reconstructed high-frequency and low-frequency seasonal components are organized into high-frequency feature matrices and low-frequency feature matrices, respectively, and a comprehensive feature matrix is formed by feature splicing. S305. Based on the fault prediction terminal, the principal component analysis algorithm is applied to reduce the dimensionality of the comprehensive feature matrix to obtain key fault features.
5. The low-voltage distribution area fault prediction method based on big data according to claim 1, characterized in that, The S400, based on the fault prediction terminal, identifies the operating status of low-voltage distribution areas using graph neural networks combined with deep learning models based on high-dimensional features, and outputs fault level classification results, specifically including the following steps: S401. Based on the fault prediction terminal, perform data extraction and processing on the database system to obtain historical data of transformer area operation data, user electricity consumption data, and historical fault records. S402. Based on the fault prediction terminal, the historical data of user electricity consumption data is standardized to obtain standard historical data. S403. Based on the fault prediction terminal, historical fault records are manually divided into several fault levels, and a support vector machine is set between every two levels. S404. Based on the fault prediction terminal, perform feature extraction processing on standard historical data to obtain historical feature data; S405. Based on the fault prediction terminal, the support vector machine is trained using historical feature data. S406. Based on the fault prediction terminal, determine whether the level classification result output by the support vector machine is consistent with the level corresponding to the historical fault record. When the number of consistent judgments reaches the preset condition, end the model training. S407. Based on the fault prediction terminal, input the real-time feature data into the trained classification model to obtain the fault level classification results output by several support vector machines.
6. The low-voltage distribution area fault prediction method based on big data according to claim 1, characterized in that, The S500, based on the fault prediction terminal, determines the faulty transformer area according to the fault level classification results and transformer area operation data, and performs fault location using a time-difference method combined with graph theory algorithms in conjunction with the power grid topology. Specifically, the steps include: S501. Based on the fault prediction terminal, within a preset time period, the trained classification model is processed by data retrieval according to a preset sampling frequency to obtain several fault level classification results output by the classification model. S502. Based on the fault prediction terminal, the fault level classification results are compared with the preset fault levels. S503. Based on the fault prediction terminal, perform data calculation and processing on the level classification results of faults greater than the preset fault level and all level classification results to obtain the weight of the fault level to be verified. S504. If the weight of the fault level to be verified is greater than the preset weight threshold, based on the fault prediction terminal, the faulty area is determined according to the source of the operating data of the area corresponding to the level classification result that is greater than the preset fault level. S505. Based on the fault prediction terminal, determine the color used to represent the alarm level according to the weight of the fault level to be verified, and fill the color in the table corresponding to the faulty area in the form of a table.
7. The low-voltage distribution area fault prediction method based on big data according to claim 6, characterized in that, In S500, the fault location method combining time difference and graph theory algorithm with the power grid topology specifically includes the following steps: S506. Based on the fault prediction terminal, according to the set of nodes in the power grid topology, calculate the total number of shortest paths between any two nodes and the number of paths passing through each intermediate node in the shortest path, and obtain the connectivity index. S507. Based on the fault prediction terminal, determine the initial set of nodes where the fault occurs according to the fault level classification results. S508. Based on the fault prediction terminal, the time difference method is used to calculate the time difference of the fault signals collected by each monitoring node. Combined with the connectivity index, the fault range is narrowed down through graph theory algorithm, and finally the fault point is located.
8. A low-voltage distribution area fault prediction system based on big data, used to implement the low-voltage distribution area fault prediction method based on big data as described in any one of claims 1-7, characterized in that, include: A fault prediction terminal is used to control data transmission and information interaction between various modules. The data acquisition module is used to collect multi-source data of low-voltage distribution areas in real time through the power distribution network monitoring platform. The multi-source data includes distribution area operation data, user electricity consumption data and environmental data. The data storage module is used to filter data through a hierarchical storage strategy, construct a data lake by combining multi-dimensional data indexes, and classify and store the multi-source data. The feature extraction module is used to aggregate and transform the data in the data lake using a real-time stream processing framework, and extract fault-related high-dimensional features by combining wavelet transform with principal component analysis. The diagnostic module is used to identify the operating status of the low-voltage distribution area based on the high-dimensional features, using a graph neural network combined with a deep learning model, and output the fault level classification result. The fault location module is used to determine the faulty transformer area based on the fault level classification results and the transformer area operation data, and to perform fault location by combining the time difference method with graph theory algorithm in combination with the power grid topology.