Power distribution system fault prediction system and prediction method based on big data analysis
By designing data acquisition, preprocessing and feature extraction modules in the power distribution system fault prediction system, and combining big data analysis models for fault mode mining and prediction, the shortcomings of existing systems in data quality and feature extraction are solved, and higher fault prediction accuracy and operation and maintenance management efficiency are achieved.
Patent Information
- Application Number
- CN202510065534.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-16
AI Technical Summary
The existing power distribution system fault prediction system based on big data analysis has shortcomings in data quality and feature extraction, which has affected the accuracy and reliability of the prediction results.
A system including data collection, data storage, data preprocessing, feature extraction, big data analysis and fault prediction and notification modules is designed. By collecting and integrating data in real time, data preprocessing and feature extraction are performed, failure mode mining and prediction are used for big data analysis models, and early warning notifications are sent in a timely manner.
The quality of data acquisition and preprocessing is improved, key features related to failures are extracted, the accuracy and reliability of the big data analysis model is enhanced, and the accuracy of fault prediction and operation and maintenance management are improved.
Smart Images

Figure CN120011715A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of transmission line system fault prediction, and in particular to a distribution system fault prediction system and a prediction method based on big data analysis. Background Art
[0002] The distribution system fault prediction system based on big data analysis is an intelligent system that collects, stores, analyzes and mines a large amount of data generated in the distribution system to identify potential fault risks in advance and predict the time, location and type of fault. The system mainly relies on big data technology, artificial intelligence (AI), machine learning (ML) algorithms and data mining methods to realize fault diagnosis and prediction of the distribution system, aiming to improve the operational reliability of the power system, reduce power outage time, reduce losses, and achieve efficient operation and maintenance management.
[0003] The distribution system fault prediction system based on big data analysis uses prediction methods mainly to achieve early identification and early warning of distribution system faults, thereby improving system reliability, reducing the frequency of faults, and optimizing resource allocation.
[0004] As the requirements for power supply reliability of distribution systems increase, fault prediction methods based on big data analysis have emerged. Data collection obtains real-time operating data from distribution system sensors, but the sensors have problems with accuracy, stability and data transmission integrity, which affect the quality of collected data and are not conducive to subsequent analysis. In addition, there are deficiencies in the accuracy of outlier judgment and the adaptability of standardization and normalization, which may interfere with subsequent feature extraction and analysis. Optimization and improvement are needed to improve the accuracy and reliability of fault prediction, but there are outliers, missing values and consistency. Summary of the invention
[0005] In order to make up for the above shortcomings, the present invention provides a distribution system fault prediction system and prediction method based on big data analysis, aiming to improve the existing technology that pays more attention to the prediction model itself rather than the quality of feature data, resulting in the prediction results may be interfered by useless features. At the same time, the key features are not refined in combination with the operating characteristics of the specific distribution system, and there is a lack of in-depth exploration of the equipment operating characteristics and electricity consumption patterns.
[0006] In a first aspect, the present invention provides the following technical solution: a distribution system fault prediction system based on big data analysis, comprising:
[0007] Data acquisition module, data storage module, data preprocessing module, feature extraction module, big data analysis module and fault prediction and notification module;
[0008] The data acquisition module is used to collect real-time operating data of sensors in the power distribution system;
[0009] The data storage module is used to store and manage the collected historical data and real-time data by adopting data storage technology;
[0010] The data preprocessing module is used to perform preliminary processing on the collected data to improve the data quality and availability;
[0011] The feature extraction module is used to extract fault-related features from the preprocessed data using data analysis and machine learning algorithms.
[0012] The big data analysis module is used to analyze and model the extracted features through learning technology to explore potential failure modes and laws;
[0013] The fault prediction and notification module is used to predict possible future faults based on the results of the analysis model, send a warning notification when a fault is predicted, and send a warning notification to relevant personnel.
[0014] Preferably, the data acquisition module includes a data collection unit, a data transmission unit and a data verification unit, the data collection unit is responsible for collecting real-time operation data from the sensor, the data transmission unit is used to transmit the collected data to the data storage module, and the data verification unit is used to perform preliminary verification on the collected data, and the preliminary verification is used to detect the integrity and accuracy of the data. The sensors include smart meters and monitoring equipment. The operating data include: voltage, current, power, temperature, humidity, long-term voltage stability data of field equipment, and coordinated operation characteristics of multiple transformers between workshops and periodic data of power consumption of the command center. The coordinated operation characteristics of the multiple transformers include load distribution characteristics, correlation of operating states, fault transmission characteristics, efficiency coordination characteristics, standby and main switching characteristics, scheduling and switching characteristics and time series characteristics. The long-term voltage stability data of the field equipment includes voltage fluctuation amplitude, voltage change rate, voltage deviation, voltage abnormality rate, voltage fluctuation frequency, voltage long-term trend characteristics and voltage stability coefficient. The periodic data of power consumption of the command center includes daily cycle power consumption pattern, weekly cycle power consumption pattern, monthly cycle power consumption pattern, holiday power consumption characteristics, periodic load fluctuation characteristics, task-driven power consumption pattern and power load forecasting data.
[0015] Preferably, the data storage module includes a data storage unit, a data management unit and a data access unit. The data storage unit is used to store historical data and real-time data using data storage technology. The data storage technology includes a distributed database or a data warehouse. The data management unit is used to classify, index and backup the stored data. The data access unit is used to provide a data access interface for other modules to connect and call the stored data.
[0016] Preferably, the data preprocessing module includes a data cleaning unit, a data screening unit and a data optimization unit, the data cleaning unit is used to clean up outliers and erroneous data in the data, the data screening unit is used to filter out useful data according to set rules, and the data optimization unit is used to perform preliminary processing on the filtered data, and the preliminary processing includes cleaning, screening, denoising and normalization.
[0017] Preferably, the feature extraction module includes a feature analysis unit, a feature calculation unit and a feature evaluation unit. The feature analysis unit is used to use a data analysis method to determine the feature type related to the fault, and the features include voltage fluctuation features, current harmonic features and power factor change features. The feature calculation unit is used to extract these features through algorithm calculation, and the feature evaluation unit is used to evaluate the validity and relevance of the extracted features.
[0018] Preferably, the big data analysis module includes a model training unit, a model evaluation unit and a model adjustment unit. The model training unit is used to perform model training using learning techniques, and the learning techniques include convolutional neural networks, recurrent neural networks and random forests. The model evaluation unit is used to perform functional evaluation on the trained model, and the functionality includes the accuracy and generalization ability of the model. The model adjustment unit is used to adjust and improve the parameters of the model according to the evaluation results.
[0019] Preferably, the fault prediction and notification module includes a fault analysis unit, a fault location prediction unit, a notification generation unit and a notification sending unit. The fault analysis unit is used to analyze the type of fault that occurs based on the results of the analysis model. The fault location prediction unit is used to determine the specific location where the fault occurs and the approximate time when the fault occurs. The notification generation unit is used to generate corresponding early warning notification content when a fault is detected. The notification sending unit is used to send the prediction content to relevant personnel for early warning notification, and send the early warning notification via text messages, emails, system pop-ups, etc. The prediction content includes the fault type, fault location and fault occurrence time, and the methods of sending early warning notifications include text messages, emails, and system pop-ups.
[0020] In a second aspect, the present invention provides the following technical solution, a distribution system fault prediction method based on big data analysis, comprising the following steps:
[0021] S1. First, data collection and integration are carried out by real-time collection of parameters from field equipment in the power distribution system, including double-circuit incoming lines, 110kv step-down stations, transformers in each workshop, transformers in the headquarters and outsourcing. The parameters of field equipment include electrical parameters, equipment status and environmental data of multiple monitoring points in each distribution room and distribution cabinet, and they are integrated into a unified data set;
[0022] S2, then perform data preprocessing, first clean up outliers and missing values, standardize and normalize data from different equipment and locations to make them comparable and consistent, and reduce data errors caused by differences in workshops and equipment;
[0023] S3, feature extraction engineering, first analyze the domain knowledge and data to extract the characteristics of the entire distribution system status and fault trend. The characteristics of the distribution system status and fault trend include the long-term voltage stability characteristics of field equipment, the coordinated operation characteristics of multiple transformers between workshops, and the periodic characteristics of power consumption of the command center;
[0024] S4. After extracting features, a big data analysis model is established, and learning technology is used to train the feature data covering the entire system, and the model parameters are adjusted and improved according to the evaluation results;
[0025] S5. Model evaluation and optimization Through cross-validation, the accuracy of the model in predicting different locations of workshop, headquarters and outsourcing faults is evaluated, and the parameters of the model are adjusted according to the evaluation results to improve the model performance;
[0026] S6. Real-time monitoring and prediction uses the data collected in real time from the entire distribution system to input the trained model to perform real-time fault prediction, and continuously update the prediction results including the positions of each workshop and the headquarters;
[0027] S7. Fault warning and decision support: When it is predicted that the fault risk at a specific location exceeds the set threshold, a warning signal will be issued in a timely manner, and detailed information on the location, time and type of the fault will be provided to provide decision support for operation and maintenance personnel to formulate maintenance strategies for that location.
[0028] In the third aspect, the invention provides the following technical solution: a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned distribution system fault prediction method based on big data analysis when executing the computer program.
[0029] In a fourth aspect, the present invention provides the following technical solution: a readable storage medium having a computer program stored thereon, and the computer program, when executed by a processor, implements the above-mentioned distribution system fault prediction method based on big data analysis.
[0030] The present invention has the following beneficial effects:
[0031] 1. In the present invention, the data acquisition module collects real-time operation data from sensors in the power distribution system, and the sensors include smart meters and monitoring equipment. The operation data include: voltage, current, power, temperature, humidity, long-term voltage stability data of field equipment, and the coordinated operation characteristics of multiple transformers between workshops and the periodic data of power consumption of the headquarters. The long-term voltage stability data of the field equipment includes voltage fluctuation amplitude, voltage change rate, voltage deviation, voltage abnormality rate, voltage fluctuation frequency, voltage long-term trend characteristics and voltage stability coefficient. The periodic data of power consumption of the headquarters include daily cycle power consumption mode, weekly cycle power consumption mode, monthly cycle power consumption mode, holiday power consumption characteristics, periodic load fluctuation characteristics, task-driven power consumption mode and power load forecast data; the present invention provides the use of big data such as long-term voltage stability data of field equipment, coordinated operation data of multiple transformers between workshops and periodic data of power consumption of the headquarters, and the technical effects that can be achieved are: key features related to faults in the operation status of equipment can be accurately identified, interference of useless features on the prediction model can be avoided, and the input quality of feature data can be optimized to provide high-value data support for model training and reduce the effect of prediction error;
[0032] 2. In the present invention, the data preprocessing module cleans up outliers, missing values, and performs standardization and normalization based on the collected data, thereby avoiding errors and inconsistencies in the original data, allowing subsequent feature extraction, big data analysis and other links to be carried out smoothly based on reliable data, effectively improving the accuracy of the entire fault prediction process.
[0033] 3. In the present invention, the analysis model generated by the big data analysis module and the related results obtained are used by the fault prediction and notification module to predict possible future faults based on these results. Once the fault risk is predicted to exceed the set threshold, detailed early warning notification content will be generated in time and sent to relevant personnel through SMS, email, system pop-up windows and other methods to ensure that the operation and maintenance personnel can know the potential fault situation at the first time, make preparations in advance, ensure that the distribution system can be maintained in time, minimize the adverse effects of the fault, and ensure the stable operation of the system.
[0034] 4. In the present invention, the features related to the fault extracted by the feature extraction module are analyzed and modeled by the big data analysis module using learning technology based on these features, so that the model can focus on key fault features for learning and training, dig out potential fault modes and laws, and convert feature data into an analysis model that has practical guiding significance for fault prediction, thereby providing a scientific and reasonable basis for fault prediction and enhancing the ability of the entire system to predict faults.
[0035] 5. In the present invention, through in-depth mining and analysis of long-term voltage stability, coordinated operation characteristics and electricity consumption periodicity data, the problem of failing to effectively utilize high-quality feature data and failing to combine the specific characteristics of the distribution system in the prior art is solved, which greatly improves the accuracy, efficiency and practical application effect of fault prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 This is a system architecture diagram of the power distribution system fault prediction system based on big data analysis proposed by the present invention;
[0037] Figure 2 This is a data acquisition module architecture diagram of the power distribution system fault prediction system based on big data analysis proposed by the present invention;
[0038] Figure 3 This is a data storage module architecture diagram of the power distribution system fault prediction system based on big data analysis proposed by the present invention;
[0039] Figure 4 This is a data preprocessing module architecture diagram of the power distribution system fault prediction system based on big data analysis proposed by the present invention;
[0040] Figure 5 This is a feature extraction module architecture diagram of the power distribution system fault prediction system based on big data analysis proposed by the present invention;
[0041] Figure 6 This is a big data analysis module architecture diagram of the power distribution system fault prediction system based on big data analysis proposed by the present invention;
[0042] Figure 7 This is a diagram of the fault prediction and notification module architecture of the power distribution system fault prediction system based on big data analysis proposed by the present invention;
[0043] Figure 8 This is a flow chart of the method for power distribution system fault prediction based on big data analysis proposed in the present invention. DETAILED DESCRIPTION
[0044] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0045] Embodiment 1
[0046] Reference Figure 1-Figure 7 In a first embodiment of the present invention, the present invention provides a distribution system fault prediction system based on big data analysis, comprising:
[0047] Data acquisition module, data storage module, data preprocessing module, feature extraction module, big data analysis module and fault prediction and notification module;
[0048] The data acquisition module is used to collect real-time operating data of sensors in the power distribution system;
[0049] The data storage module is used to store and manage the collected historical data and real-time data by adopting data storage technology;
[0050] The data preprocessing module is used to perform preliminary processing on the collected data to improve data quality and availability;
[0051] The feature extraction module is used to extract fault-related features from the preprocessed data using data analysis and machine learning algorithms;
[0052] The big data analysis module is used to analyze and model the extracted features through learning technology to explore potential failure modes and laws;
[0053] The fault prediction and notification module is used to predict possible future faults based on the results of the analysis model and send warning notifications to relevant personnel when a fault is predicted.
[0054] The data acquisition module includes a data collection unit, a data transmission unit and a data verification unit. The data collection unit is responsible for collecting real-time operation data from the sensor, the data transmission unit is used to transmit the collected data to the data storage module, and the data verification unit is used to perform preliminary verification on the collected data. The preliminary verification is used to detect the integrity and accuracy of the data. The sensors include smart meters and monitoring equipment. The sensors include smart meters and monitoring equipment. The operating data include: voltage, current, power, temperature, humidity, long-term voltage stability data of field equipment, and coordinated operation characteristics of multiple transformers between workshops and periodic data of power consumption of the command center. The coordinated operation characteristics of the multiple transformers include load distribution characteristics, correlation of operating states, fault transmission characteristics, efficiency coordination characteristics, standby and main switching characteristics, scheduling and switching characteristics and time series characteristics. The long-term voltage stability data of the field equipment includes voltage fluctuation amplitude, voltage change rate, voltage deviation, voltage abnormality rate, voltage fluctuation frequency, voltage long-term trend characteristics and voltage stability coefficient. The periodic data of power consumption of the command center includes daily cycle power consumption pattern, weekly cycle power consumption pattern, monthly cycle power consumption pattern, holiday power consumption characteristics, periodic load fluctuation characteristics, task-driven power consumption pattern and power load forecast data.
[0055] Specifically, the data collection unit collects real-time operating data such as voltage, current, power, temperature and humidity from sensors such as smart meters and monitoring equipment, so as to obtain relevant information in multiple dimensions during the operation of the equipment, and provide a basic data source for subsequent data processing and analysis. It can capture the key operating parameters reflected by various sensors in a relatively comprehensive and real-time manner, ensure the timeliness of the data, provide first-hand information for understanding the actual operating status of related equipment, and facilitate subsequent accurate assessment of equipment conditions and discovery of potential problems;
[0056] The data transmission unit accurately transmits various operating data acquired by the data collection unit to the data storage module, acting as a bridge for data transmission, ensuring that the data can be smoothly transferred to the storage link for further retention and processing, so that the collected data can be properly stored to avoid data loss, and provides the prerequisite for subsequent in-depth analysis, query and other operations based on the stored data, ensuring the continuity of the entire data link;
[0057] The data verification unit conducts preliminary verification on various collected operation data, focusing on the integrity and accuracy of the data, judging whether there are missing parts in the data or whether the data values themselves are accurate and reasonable, and screening out problematic data in advance. It can effectively improve data quality, reduce the number of erroneous and incomplete data entering the subsequent process, avoid erroneous analysis and decision-making caused by data problems, lay a good data foundation for the entire data processing process, and ensure that all subsequent data-based operations are based on reliable data;
[0058] The coordinated operation characteristics of multiple transformers between workshops are specifically measured through transformer load rate, load balancing coefficient, peak load sharing rate, output voltage synchronization, reactive power distribution, harmonic component analysis, voltage regulation response time, load fluctuation synchronization, dynamic switching characteristics, unit energy consumption, operation efficiency comparison, temperature rise monitoring, transformer failure frequency, operation abnormality correlation, equipment health index HI value, short-term voltage fluctuation rate, long-term voltage stability coefficient, current distribution stability, as well as coordinated operation rules and load periodicity rules, load coordination index, and load distribution prediction accuracy. Through real-time collection, analysis and modeling of these parameters, the operating status and coordination of multiple transformers are achieved, the fault risk is effectively identified, and the stability and operation efficiency of the distribution system are improved.
[0059] The data storage module includes a data storage unit, a data management unit and a data access unit. The data storage unit is used to store historical data and real-time data using data storage technology. The data storage technology includes a distributed database or a data warehouse. The data management unit is used to classify, index and backup the stored data. The data access unit is used to provide a data access interface for other modules to connect and call the stored data.
[0060] Specifically, the data storage unit uses data storage technologies such as distributed databases or data warehouses to properly store the collected historical data and real-time data, which is equivalent to creating a safe and reliable "storage warehouse" for the data so that it can be retained for a long time to meet the subsequent data usage needs in different scenarios. It can achieve orderly storage of large amounts of data, ensure data stability and security through appropriate data storage technology, prevent data loss, and facilitate subsequent modules to quickly locate and extract corresponding data when needed, providing a stable data resource foundation for the entire data application system;
[0061] The data management unit manages the stored data in many aspects, including classification operations, dividing data of different types and sources according to specific rules; index management to facilitate quick data search; and backup management to ensure that data has redundant copies to prevent data damage or loss in case of unexpected situations so that it can be restored in time, ensuring the good management status of data in all aspects. It improves the management efficiency of data, makes data clear and easy to find and use. At the same time, backup management enhances the risk resistance of data. Even in the event of storage failures, human errors, etc., it can maximize the integrity and availability of data and escort the continuous application of data.
[0062] The data access unit provides a data access interface and builds a connection channel between other modules and stored data, so that other modules can smoothly call the stored data according to the corresponding rules and permissions, realize the circulation and sharing of data between different modules, and promote the coordinated use of data by all parts of the entire system to carry out work. It breaks the data barriers between modules, allowing data to flow flexibly within the system, facilitating different functional modules to obtain data on demand, helping to improve the coordination and operating efficiency of the entire system, and realizing diversified functions and business expansion based on a unified data foundation.
[0063] The data preprocessing module includes a data cleaning unit, a data screening unit and a data optimization unit. The data cleaning unit is used to clean up outliers and erroneous data in the data. The data screening unit is used to filter out useful data according to set rules. The data optimization unit is used to perform preliminary processing on the filtered data. The preliminary processing includes cleaning, screening, denoising and normalization.
[0064] Specifically, the data cleaning unit focuses on the abnormal values and erroneous data in the data, and identifies and cleans them through specific algorithms and rules. It removes or corrects the values that are obviously beyond the reasonable range and the data with incorrect format, ensuring that the quality of the data meets the basic requirements of subsequent analysis and application. It effectively improves the purity and accuracy of the data, reduces the misleading of data analysis results caused by abnormal and erroneous data, and enables various subsequent operations based on the data to be based on relatively reliable data, thereby enhancing the overall availability of the data.
[0065] The data screening unit identifies and extracts useful data that is valuable for current business needs, analysis goals, etc. from massive amounts of data based on pre-set rules. For example, data can be screened according to a specific time period, operating data corresponding to a specific device, or data that meets certain characteristics, eliminating irrelevant data, narrowing the scope of data processing, and improving processing efficiency. Data that meets the requirements can be accurately located, allowing subsequent data processing, analysis, and other links to focus on key data, avoiding the waste of resources on processing large amounts of useless data, and at the same time helping to obtain analysis results that are more in line with actual needs and more targeted;
[0066] The data optimization unit performs various preliminary processing on the screened data, including further data cleaning to remove a small amount of abnormal data that may remain during the screening process; screening again to ensure that the data better meets the requirements; denoising to eliminate interference factors in the data, and normalization to unify data of different dimensions into a specific range to facilitate subsequent comparisons and calculations. It optimizes the quality and characteristics of the data in all aspects, making the data more standardized, clean and more comparable, providing a high-quality data foundation for subsequent advanced applications such as in-depth data mining and modeling, and improving the scientificity and effectiveness of the overall data processing and analysis process;
[0067] The feature extraction module includes a feature analysis unit, a feature calculation unit and a feature evaluation unit. The feature analysis unit is used to determine the feature type related to the fault by using a data analysis method. The features include voltage fluctuation features, current harmonic features and power factor change features. The feature calculation unit is used to extract these features through algorithm calculation. The feature evaluation unit is used to evaluate the validity and relevance of the extracted features.
[0068] Specifically, the feature analysis unit uses data analysis methods to deeply explore the feature types related to faults, such as voltage fluctuation features, current harmonic features, and power factor change features. It accurately locates key features (see Table 2) to provide a clear direction for subsequent extraction and evaluation, which helps to focus on key points and improve the pertinence of the entire fault analysis process;
[0069] The feature calculation unit uses the algorithm to perform specific calculations on the determined features, so as to accurately extract the corresponding features from the original data. Specifically, the feature calculation unit collects the original operating data of multiple transformers (such as voltage, current, power, etc.) in real time, combines mathematical modeling and algorithm calculation, and accurately extracts the target features from them. For example, the calculation of the load balancing coefficient is realized: first, the real-time load data of each transformer is monitored by sensors and integrity verification is performed, and the load value of each transformer is normalized and then its proportion in the total load is calculated; secondly, the calculation formula of the load balancing coefficient is defined. The formula is based on the mean square error principle, and the deviation degree of the load ratio of each transformer is compared with the ideal uniform load ratio. The formula is:
[0070]
[0071] The balance index is obtained; then, in the algorithm implementation, real-time calculation is performed through an automated program, the load data is input into the calculation module, the characteristic value is extracted and output, and stored in the characteristic database; finally, the validity and relevance of the feature are verified through the feature evaluation module, and the accuracy and practicality in the operation status evaluation and system optimization of multiple transformers are verified. During use, it can dynamically reflect the load distribution balance of multiple transformers, provide key support for fault prediction and optimal scheduling, thereby improving the operation stability and efficiency of the distribution system, achieving quantitative presentation of the features, and providing intuitive and reliable data basis for fault diagnosis, equipment status monitoring, etc., which is convenient for further analysis and judgment.
[0072] The feature evaluation unit evaluates the effectiveness and relevance of the extracted features to identify the actual value of the features for fault analysis. Specifically, first, the original feature set is obtained from the feature extraction module, and a training data set containing input features and output fault categories is constructed based on historical operating data and known fault records. Then, statistical analysis and machine learning methods are used to evaluate the effectiveness and relevance of the features, respectively. The effectiveness evaluation quantifies the importance of each feature to the classification task by calculating the distribution difference of the features (such as using statistical indicators such as mean and variance) and the ability to distinguish fault categories (such as information gain, Fisher score, etc.). The relevance evaluation uses correlation coefficients (such as Pearson correlation coefficient or mutual information) to analyze the degree of correlation between the features and the fault target variables, and screen out related Then, based on the comprehensive ranking of effectiveness and relevance, the feature set is input into a variety of machine learning models (such as random forests and support vector machines) through cross-validation to train and compare the model prediction performance (such as accuracy, recall, and F1 value), so as to evaluate the contribution of each feature in the model; finally, based on the above results, invalid features or features that do not significantly improve the model performance are eliminated, and an optimized feature set is generated for subsequent analysis and modeling, so that the key features with practical value for fault analysis can be accurately identified during use (see Table 2), effectively improving the accuracy, reliability, and generalization ability of the analysis model, while reducing the occupation of computing resources by redundant features, and providing high-quality data support for distribution system fault prediction. Filter out high-quality features and avoid interference from useless features, so that subsequent work based on features is more scientific and accurate, and the quality of decision-making is improved.
[0073] The big data analysis module includes a model training unit, a model evaluation unit and a model adjustment unit. The model training unit is used to use learning techniques to train the model. The learning techniques include convolutional neural networks, recurrent neural networks and random forests. The model evaluation unit is used to evaluate the functionality of the trained model. The functionality includes the accuracy and generalization ability of the model. The model adjustment unit is used to adjust and improve the model parameters according to the evaluation results.
[0074] Specifically, the model training unit uses learning technologies such as convolutional neural networks, recurrent neural networks, and random forests to carry out model training based on relevant data, so that the model has processing and analysis capabilities. The model can learn the data rules through learning, effectively analyze and predict the corresponding problems, provide a basis for subsequent applications, and enhance the value of data utilization;
[0075] The model evaluation unit conducts a comprehensive evaluation of the trained model from functional aspects such as accuracy and generalization ability. Specifically, first, the model is verified in stages using the constructed training set and test set. The accuracy evaluation calculates the model's prediction performance indicators on the test set, including accuracy, precision, recall, and F1 value, which are used to measure the accuracy and comprehensiveness of the model for fault classification or prediction. Secondly, in order to evaluate the generalization ability of the model, the original data is divided into multiple subsets using the cross-validation method, which are alternately used as training sets and validation sets. The mean and variance of the performance indicators of each validation are calculated, and the consistency of the model on different data sets is judged by the size of the variance. Thirdly, in order to address possible overfitting or underfitting problems, the learning curve is used to analyze the model's performance in the training set and validation set. The error change trend on the power distribution system is used to determine the adaptability of the model complexity and data fitting. At the same time, regularization methods (such as L1 or L2 regularization) are used to optimize the structural parameters of the model to improve the generalization performance. Subsequently, the real-time prediction effect of the model is tested on an independent validation set or real distribution system data, and its actual performance on unseen data is recorded to further verify its practicality. Finally, through the importance scoring mechanism, the feature weights and parameter settings in the model are retrospectively analyzed, and a performance report is output and targeted optimization suggestions are provided in combination with specific business needs. This method can comprehensively and accurately evaluate the functionality and practicality of the model, achieve efficient application in distribution system fault prediction, significantly improve prediction accuracy and adaptability to complex scenarios, and determine the performance of the model in different scenarios. Clearly knowing the actual level of the model facilitates the discovery of potential problems, provides a clear basis for subsequent model adjustments, and ensures reliable model quality.
[0076] The model adjustment unit reasonably adjusts and improves the parameters of the model based on the evaluation results, optimizes the performance of the model, and makes it more suitable for actual application needs. It improves the overall performance of the model, enhances its accuracy and generalization ability, better copes with various complex situations, and improves the effectiveness and accuracy of big data analysis.
[0077] The fault prediction and notification module includes a fault analysis unit, a fault location prediction unit, a notification generation unit and a notification sending unit. The fault analysis unit is used to analyze the type of fault that occurs according to the results of the analysis model. The fault location prediction unit is used to determine the specific location of the fault and the approximate time when the fault occurs. The notification generation unit is used to generate corresponding early warning notification content when a fault is detected. The notification sending unit is used to send the prediction content to relevant personnel for early warning notification. The early warning notification is sent through SMS, email, system pop-up windows, etc. The prediction content includes the fault type, fault location and fault occurrence time. The methods for sending early warning notifications include SMS, email, and system pop-up windows.
[0078] Specifically, the fault analysis unit, based on the results output by the analysis model, deeply analyzes the specific type of the fault that has occurred, provides a key basis for subsequent precise response to the fault, can accurately identify the fault category, and help relevant personnel quickly understand the nature of the fault, so as to formulate targeted solutions and improve fault handling efficiency. Specifically: first, use voltage, current, power, temperature, etc. and historical fault records to build a training data set containing multiple fault type labels, and use the feature extraction module to extract the voltage fluctuation amplitude, short-circuit current peak, temperature rise abnormal trend, etc., and input these features into the trained random forest, support vector machine or convolutional neural network; secondly, when a fault occurs, the real-time collected system operation data is input into the model for prediction, and the model outputs specific fault type labels, such as overload fault, short-circuit fault, grounding fault or equipment aging fault; then, the fault analysis unit uses a specific rule engine to analyze the results output by the analysis model. The results are further interpreted and combined with contextual information (such as equipment operation time, environmental conditions, and system configuration) to generate multi-dimensional fault description information, including fault cause analysis (such as overcurrent caused by overload), fault impact range (such as a transformer and its downstream equipment), and possible expansion risks (such as voltage drop caused by short circuit); then, the fault analysis unit compares the analysis results with the relevant expert system database, matches the existing treatment solutions and maintenance suggestions, and generates a detailed fault diagnosis report, including fault type, cause analysis, and priority treatment measures; finally, the fault diagnosis results are sent to the operation and maintenance personnel in a timely manner through the integrated notification module, and displayed through SMS, email or system interface to assist them in quickly formulating targeted solutions. This method can effectively combine historical experience and real-time data, accurately identify fault categories, deeply analyze fault causes and impacts, improve the efficiency and accuracy of fault handling, and provide reliable protection for the stable operation of the distribution system;
[0079] The fault location prediction unit uses relevant data and analysis methods to accurately lock the specific location of the fault, and estimate the approximate time of the fault, thereby realizing the spatial and temporal location of the fault. It provides strong support for maintenance personnel to quickly arrive at the fault site and make preparations in advance, reducing the time and scope of fault investigation, and improving the speed of emergency response. Specifically: collect multi-point real-time operation data from the sensor network of the distribution system, and synchronize the data using distributed data transmission technology; then, based on the power topology map and hierarchical structure model, map the monitoring data to the physical location of the distribution system, and quickly narrow the fault scope through network analysis algorithms (such as the shortest path algorithm or the current path model), and combine fault characteristic data (such as voltage drop, abnormal current peak, rapid temperature rise, etc.) to predict the time of occurrence of the fault, ensuring that potential faults can be perceived in advance; then After that, the specific location of the fault is further accurately located through multi-point data cross-validation (such as time difference and amplitude difference analysis of data from different monitoring points), and a detailed fault location and prediction report is generated in combination with the possible types and severity of the fault model output; finally, the report information is pushed to maintenance personnel in real time through the integrated notification module, including the specific fault location, occurrence time prediction and impact range, to support rapid response and precise maintenance, so as to comprehensively consider time, space and equipment characteristics, significantly improve the accuracy of fault location and prediction, shorten fault troubleshooting time, effectively reduce the power outage scope and recovery time of the distribution system, and improve the reliability of system operation and emergency response efficiency;
[0080] Once a fault is detected, the notification generation unit will promptly generate the corresponding warning notification content based on the relevant information, covering the key elements of the fault and ensuring the completeness of the notification information. The warning notification content is guaranteed to be detailed and accurate, so that the personnel receiving the notification can fully understand the fault situation, so as to arrange the corresponding work in advance and make preventive preparations;
[0081] The notification sending unit sends warning notifications containing predicted contents such as fault type, location and occurrence time to relevant personnel through various means such as SMS, email, and system pop-up windows. This ensures that warning notifications can be conveyed to personnel who need to pay attention to fault conditions in a timely manner through multiple channels, improves the coverage and timeliness of information transmission, and helps to coordinate the response to faults.
[0082] Embodiment 2:
[0083] Reference Figure 8 In a second embodiment of the present invention, the present invention provides a distribution system fault prediction method based on big data analysis, comprising the following steps:
[0084] S1. First, data collection and integration are carried out by real-time collection of parameters from field equipment in the power distribution system, including double-circuit incoming lines, 110kv step-down stations, transformers in each workshop, transformers in the headquarters and outsourcing. The parameters of field equipment include electrical parameters, equipment status and environmental data of multiple monitoring points in each distribution room and distribution cabinet, and they are integrated into a unified data set;
[0085] S2. Then perform data preprocessing, first clean up outliers and missing values, standardize and normalize data from different equipment and locations to make them comparable and consistent, and reduce data errors caused by differences in workshops and equipment.
[0086] S3. When performing feature engineering, domain knowledge and data are first analyzed to extract the characteristics of the entire distribution system status and fault trends. The characteristics of the distribution system status and fault trends include the long-term voltage stability characteristics of field equipment, the coordinated operation characteristics of multiple transformers between workshops, and the periodic characteristics of power consumption at the command center.
[0087] S4. After extracting features, a big data analysis model is established, and learning technology is used to train the feature data covering the entire system. The model parameters are adjusted and improved based on the evaluation results.
[0088] S5. Model evaluation and optimization Through cross-validation, the accuracy of the model in predicting different locations of workshop, command center and outsourced faults is evaluated, and the parameters of the model are adjusted according to the evaluation results to improve the model performance.
[0089] S6. Real-time monitoring and prediction uses the real-time collected data from the entire distribution system to input the trained model to perform real-time fault prediction, and continuously update the prediction results including those of each workshop and the headquarters' additional coordination locations.
[0090] S7. Fault warning and decision support: When it is predicted that the fault risk at a specific location exceeds the set threshold, a warning signal will be issued in a timely manner, and detailed information on the location, time and type of the fault will be provided to provide decision support for operation and maintenance personnel to formulate maintenance strategies for that location.
[0091] Specifically, the data acquisition and integration in S1 collects the electrical, status and environmental data of multiple monitoring points of various field equipment in the distribution system (such as double-circuit incoming lines, 110kv step-down stations, etc.) in real time and integrates them into a unified data set (see Table 1). Comprehensively gather multi-source data to provide a complete and centralized data foundation for subsequent analysis and processing, facilitating unified management and utilization;
[0092] Data preprocessing in S2 cleans up outliers and missing values in the data, standardizes and normalizes data from different sources, and eliminates data errors caused by differences in equipment and workshops. It improves data quality, enhances data comparability and consistency, and allows subsequent analysis to be based on more reliable and standardized data to ensure the accuracy of the results;
[0093] In S3, feature engineering combines domain knowledge to analyze data and extracts features related to the state of the power distribution system and fault trends, such as voltage stability of field equipment. Mining out key features (see Table 2) and concentrating the effective information contained in the data will help to accurately grasp the system state in the future and provide important support for establishing an analysis model.
[0094] In S4, a big data analysis model is established to use learning technology to train the model based on the extracted feature data, and the model parameters are optimized based on the evaluation results to improve the model performance. An effective analysis tool that fits the power distribution system is built, which can better mine data patterns after adjustment and improvement, and prepare for fault prediction, etc.
[0095] In S5, model evaluation and optimization uses cross-validation to evaluate the accuracy of the model in fault prediction at different locations, and then adjusts parameters based on the results to further improve model performance. Accurately control model quality and optimize in a targeted manner, so that it can more accurately predict fault conditions at each location, enhancing practicality and reliability;
[0096] Real-time monitoring and prediction in S6 inputs the real-time collected distribution system data into the trained model to achieve real-time prediction of faults at each location and update the results. Specifically: the collected raw data is preprocessed, feature extracted and modeled, and then input into the prediction model to generate prediction results for specific fault types and fault locations. Combined with the domain knowledge of the distribution system, key features related to the fault are extracted (see Table 2). These features are used as model inputs, including but not limited to: voltage deviation, abnormal current peak, harmonic content, load distribution characteristics, temperature rise trend and historical status characteristics. The extracted features are combined into a feature vector as model input, such as X = voltage deviation, load rate, harmonic content, current peak and temperature rise trend. It can timely grasp the real-time status of the system, dynamically output prediction results, help detect fault risks in advance, and provide timely reference for operation and maintenance;
[0097] Fault warning and decision support in S7 When the fault risk exceeds the threshold, a warning is issued. The alarm threshold is the input data (real-time acquisition): voltage: 9.5kV (rated 10kV), current: 120A (rated 100A), temperature rise: 90℃ (normal <80℃), harmonic content: 12% (normal <10%), then based on the threshold setting, the model predicts the output fault type: overload fault, fault location: workshop 2 transformer, fault time: expected within the next 5 minutes, and provides detailed fault information to assist operation and maintenance personnel in formulating maintenance strategies. Timely remind operation and maintenance personnel of potential faults, provide key basis for their decision-making, facilitate rapid response, and reasonably arrange maintenance work to ensure stable operation of the system.
[0098] Embodiment 3
[0099] The third embodiment of the present invention is based on the same inventive concept. The present invention proposes a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps of the distribution system fault prediction method based on big data analysis of the above embodiment.
[0100] Embodiment 4
[0101] The fourth embodiment of the present invention is based on the same inventive concept. A computer device proposed by the present invention includes: a processor and a memory; the processor and the memory communicate with each other; the memory is used to store instructions; the processor is used to execute the instructions in the memory, and execute the distribution system fault prediction method based on big data analysis of the above embodiment.
[0102] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0103] Table 1: Unified dataset
[0104]
[0105]
[0106] In Table 1:
[0107] Voltage (kV): actual voltage value at the monitoring point;
[0108] Current (A): the current value of the corresponding monitoring point;
[0109] Power (kW): Active power calculated by voltage, current and power factor; Temperature (℃): Ambient temperature of the monitoring equipment;
[0110] Humidity (%): monitor the ambient humidity of the equipment;
[0111] Status: Whether the monitoring point is operating normally.
[0112] Table 2: Key characteristics
[0113]
[0114] In Table 2:
[0115] Voltage (kV): real-time voltage value of each monitoring point;
[0116] Voltage deviation (%): the deviation of the voltage at the monitoring point relative to the rated voltage;
[0117] Voltage fluctuation range (kV): short-term fluctuation range of voltage at the monitoring point;
[0118] Voltage change rate (kV / s): the rate at which voltage changes per unit time;
[0119] Current (A): actual current value at the monitoring point;
[0120] Power factor: The power factor at the monitoring point reflects the efficiency of electric energy utilization;
[0121] Temperature (℃): Equipment operating environment temperature;
[0122] Humidity (%): Humidity of the equipment operating environment;
[0123] Load rate (%): the percentage of equipment load to rated load;
[0124] Status: The current operating status of the device (normal / abnormal).
[0125] Finally, it should be noted that the above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, it is still possible for those skilled in the art to modify the technical solutions described in the aforementioned embodiments or to make equivalent substitutions for some of the technical features therein. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A distribution system fault prediction system based on big data analysis, characterized in that: include: Data acquisition module, data storage module, data preprocessing module, feature extraction module, big data analysis module and fault prediction and notification module; The data acquisition module is used to collect real-time operating data of the voltage change rate of the sensor in the power distribution system; The data storage module is used to store and manage the collected historical data and real-time data by adopting data storage technology; The data preprocessing module is used to perform preliminary processing on the collected data to improve the data quality and availability; The feature extraction module is used to extract features related to the fault from the preprocessed data by using data analysis and machine learning algorithms; the feature extraction module includes a feature analysis unit, a feature calculation unit and a feature evaluation unit, the feature analysis unit is used to determine the feature type related to the fault by using a data analysis method, the features include the long-term voltage stability features of the field equipment, the coordinated operation features of multiple transformers between workshops and the periodic features of the power consumption of the command center, the feature calculation unit is used to extract these features by algorithm calculation, and the feature evaluation unit is used to evaluate the validity and relevance of the extracted features; The big data analysis module is used to analyze and model the extracted features through learning technology to explore potential failure modes and laws; The fault prediction and notification module is used to predict possible future faults based on the results of the analysis model, send a warning notification when a fault is predicted, and send a warning notification to relevant personnel.
2. The power distribution system fault prediction system based on big data analysis according to claim 1, characterized in that: The data acquisition module includes a data collection unit, a data transmission unit and a data verification unit. The data collection unit is responsible for collecting real-time operation data from the sensor, the data transmission unit is used to transmit the collected data to the data storage module, and the data verification unit is used to perform preliminary verification on the collected data. The preliminary verification is used to detect the integrity and accuracy of the data.
3. The power distribution system fault prediction system based on big data analysis according to claim 2 is characterized in that: The sensors include smart meters and monitoring equipment. The operating data include: voltage, current, power, temperature, humidity, long-term voltage stability data of field equipment, and coordinated operation characteristics of multiple transformers between workshops and periodic data of power consumption of the command center. The coordinated operation characteristics of the multiple transformers include load distribution characteristics, correlation of operating states, fault transmission characteristics, efficiency coordination characteristics, standby and main switching characteristics, scheduling and switching characteristics and time series characteristics. The long-term voltage stability data of the field equipment includes voltage fluctuation amplitude, voltage change rate, voltage deviation, voltage abnormality rate, voltage fluctuation frequency, voltage long-term trend characteristics and voltage stability coefficient. The periodic data of power consumption of the command center includes daily cycle power consumption pattern, weekly cycle power consumption pattern, monthly cycle power consumption pattern, holiday power consumption characteristics, periodic load fluctuation characteristics, task-driven power consumption pattern and power load forecasting data.
4. The power distribution system fault prediction system based on big data analysis according to claim 1, characterized in that: The data storage module includes a data storage unit, a data management unit and a data access unit. The data storage unit is used to store historical data and real-time data using data storage technology. The data storage technology includes a distributed database or a data warehouse. The data management unit is used to classify, index and backup the stored data. The data access unit is used to provide a data access interface for other modules to connect and call the stored data.
5. The power distribution system fault prediction system based on big data analysis according to claim 1, characterized in that: The data preprocessing module includes a data cleaning unit, a data screening unit and a data optimization unit. The data cleaning unit is used to clean up abnormal values and erroneous data in the data, the data screening unit is used to filter out useful data according to set rules, and the data optimization unit is used to perform preliminary processing on the filtered data, and the preliminary processing includes cleaning, screening, denoising and normalization.
6. The power distribution system fault prediction system based on big data analysis according to claim 1, characterized in that: The big data analysis module includes a model training unit, a model evaluation unit and a model adjustment unit. The model training unit is used to perform model training using learning techniques, and the learning techniques include convolutional neural networks, recurrent neural networks and random forests. The model evaluation unit is used to perform functional evaluation on the trained model, and the functionality includes the accuracy and generalization ability of the model. The model adjustment unit is used to adjust and improve the parameters of the model according to the evaluation results.
7. The power distribution system fault prediction system based on big data analysis according to any one of claims 1 to 6, characterized in that: The fault prediction and notification module includes a fault analysis unit, a fault location prediction unit, a notification generation unit and a notification sending unit. The fault analysis unit is used to analyze the type of fault that occurs according to the results of the analysis model. The fault location prediction unit is used to determine the specific location of the fault and the approximate time when the fault occurs. The notification generation unit is used to generate corresponding early warning notification content when a fault is detected. The notification sending unit is used to send the prediction content to relevant personnel for early warning notification. The early warning notification is sent through text messages, emails, system pop-ups, etc. The prediction content includes the fault type, fault location and fault occurrence time. The methods for sending early warning notifications include text messages, emails, and system pop-ups.
8. A distribution system fault prediction method based on big data analysis, characterized in that: The following steps are involved: S1. Data collection and integration collects parameters from field equipment in the power distribution system in real time, including double-circuit incoming lines, 110kv step-down stations, transformers in each workshop, transformers in the headquarters and outsourcing. The parameters of field equipment include electrical parameters, equipment status and environmental data of multiple monitoring points in each distribution room and distribution cabinet, and integrates them into a unified data set; S2. Data preprocessing: first clean up outliers and missing values, standardize and normalize data from different equipment and locations to make them comparable and consistent, and reduce data errors caused by differences in workshops and equipment; S3, feature extraction engineering, first analyze the domain knowledge and data to extract the characteristics of the entire distribution system status and fault trend. The characteristics of the distribution system status and fault trend include the long-term voltage stability characteristics of field equipment, the coordinated operation characteristics of multiple transformers between workshops, and the periodic characteristics of power consumption of the command center; S4. Establish a big data analysis model, use learning technology to train the feature data covering the entire system, and adjust and improve the model parameters based on the evaluation results; S5. Model evaluation and optimization: Through cross-validation, the accuracy of the model in predicting different locations of workshop, headquarters and outsourcing faults is evaluated, and the parameters of the model are adjusted according to the evaluation results to improve the model performance; S6. Real-time monitoring and prediction uses the data collected in real time from the entire distribution system to input the trained model to perform real-time fault prediction and continuously update the prediction results including the positions of each workshop and the headquarters; S7. Fault warning and decision support: When it is predicted that the fault risk at a specific location exceeds the set threshold, a warning signal will be issued in a timely manner, and detailed information on the location, time and type of the fault will be provided to provide decision support for operation and maintenance personnel to formulate maintenance strategies for that location.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, the distribution system fault prediction method based on big data analysis as described in claim 8 is implemented.
10. A computer device comprising a processor, characterized in that: Also included is the computer-readable storage medium of claim 8.
Citation Information
Cited By
Urban rail power supply system fault early warning method based on multi-source data fusion
CN120742012A
A fault early warning method for urban rail power supply system based on multi-source data fusion
CN120742012B
Electrical automation power distribution cabinet maintenance system capable of rapidly replacing modular components
CN121258486A
A modular component quick-change electrical automation distribution cabinet maintenance system
CN121258486B
AI-based power distribution cabinet remote monitoring and fault prediction method and system
CN122512634A