Intelligent debugging data management method and system for high-voltage prefabricated substation
The data acquisition frequency is determined through the frequency multi-layer analysis method, and multi-source heterogeneous data acquisition and analysis are carried out, which solves the problem of insufficient data management in the existing technology, and realizes the precise maintenance and operation stability of high-voltage pre-installed substations.
Patent Information
- Application Number
- CN202510139554.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-02-08
AI Technical Summary
The existing intelligent debugging data management technology of high-voltage pre-installed substations has insufficient data management, making it difficult to achieve real-time collection, transmission and processing of debugging data, resulting in difficult to ensure the accuracy and completeness of the data, and it is impossible to provide accurate decision-making support to operation and maintenance personnel in a timely manner.
A reasonable data acquisition frequency is determined through frequency multi-layer analysis method, and multi-source heterogeneous data acquisition is carried out, including equipment operation parameters, environmental parameters, equipment assets and document data. The collected data is preprocessed and transmitted to the remote management center and stored. Multi-source data analysis is performed based on the stored data, the working conditions of the high-voltage pre-installed substation are determined, and the maintenance plan of the substation is determined based on the working conditions.
It realizes accurate and efficient equipment maintenance, reduces the probability of failure, reduces maintenance costs, improves the reliability and stability of the operation of the substation, extends the service life of the equipment, and ensures the safety and sustainability of the power supply.
Smart Images

Figure CN120013525A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data management, and in particular to a method and system for intelligent debugging data management of a high-voltage prefabricated substation. Background Art
[0002] With the rapid development of economy and society, the demand for electric energy is increasing, the construction of power grid is constantly advancing, and the application of high-voltage prefabricated substations is becoming more and more extensive. The integration of power grid control has become an important development direction of the power industry, and higher requirements have been put forward for the intelligent level of substations, including the intelligent and efficient management of debugging data. The prefabricated substation adopts a modular design method, which reasonably modularizes a fully functional substation. Each module is prefabricated in a modern factory, and the wiring and module debugging are completed inside the factory. The substation site only needs to carry out external connection and overall linkage debugging, which greatly reduces the workload of the substation site and improves the construction quality.
[0003] However, the existing high-voltage prefabricated substation intelligent commissioning data management technology has deficiencies in data management. Traditional substation commissioning data management often relies on manual recording and organization. When faced with a large amount of data, it is not only inefficient, but also prone to errors and omissions. It is difficult to ensure the accuracy and completeness of the data, and it is difficult to achieve real-time collection, transmission and processing of commissioning data, and it is impossible to provide accurate decision support for operation and maintenance personnel in a timely manner. Summary of the invention
[0004] In view of the deficiencies in the prior art, the present invention provides a method and system for intelligent debugging data management of a high-voltage prefabricated substation, which can achieve accurate and efficient equipment maintenance, reduce the probability of failure, reduce maintenance costs, and improve the reliability and stability of substation operation.
[0005] To achieve the above objectives, the present invention is implemented through the following technical solutions: A method for managing intelligent debugging data of a high-voltage prefabricated substation comprises the following steps: Determine the data collection frequency based on the frequency multi-layer analysis method, and collect multi-source heterogeneous data based on the determined data collection frequency. The multi-source heterogeneous data includes equipment operation parameters, environmental parameters, equipment assets and document data; After data preprocessing, the collected multi-source heterogeneous data is transmitted to the remote management center and stored; Perform multi-source data analysis based on pre-processed multi-source heterogeneous data stored in the remote management center to determine the working conditions of the high-voltage prefabricated substation, including equipment failure types, equipment failure development trends, and equipment failure causes; Determine the substation maintenance plan based on the operating conditions of the high-voltage prefabricated substation.
[0006] Preferably, determining the data acquisition frequency based on the frequency multi-layer analysis method comprises the following steps: Perform a frequency analysis based on equipment operating parameters and environmental parameters to obtain a data collection frequency based on the equipment operating parameters and environmental parameters. The equipment operating parameters include transformer operating data, high-voltage switchgear operating data, low-voltage switchgear operating data, and capacitor compensation device operating data. Perform secondary frequency analysis based on data change trends to obtain secondary data collection frequencies based on data change trends. Data change trends include equipment operating parameter change trends and environmental parameter change trends. Performing three-frequency analysis based on the power grid operation conditions to obtain three-frequency data collection frequencies based on the power grid operation conditions; Determine the maximum value among the primary data collection frequency, the secondary data collection frequency and the tertiary data collection frequency, and determine the maximum value as the to-be-determined collection frequency; Determine whether the pending acquisition frequency is greater than the original acquisition frequency: If the pending acquisition frequency is greater than the original acquisition frequency, the pending acquisition frequency is determined as the data acquisition frequency; If the pending acquisition frequency is not greater than the original acquisition frequency, the original acquisition frequency is determined as the data acquisition frequency.
[0007] Preferably, performing a frequency analysis based on the equipment operating parameters and the environmental parameters to obtain a data collection frequency based on the equipment operating parameters and the environmental parameters comprises the following steps: Standardize the equipment operating parameters and environmental parameters to obtain the standardized primary frequency determination data; Based on the set hierarchical structure, the normalized frequency determination data is divided to obtain the judgment matrix A. , n is the number of indicators in the frequency-determined data after standardization, Indicates the importance of the i-th indicator relative to the j-th indicator, ; Solve the characteristic equation of judgment matrix A , where I represents the identity matrix, Denotes the eigenvalue of A, and obtains the maximum eigenvalue and The corresponding eigenvector W is , T is the transposition symbol, The value representing the relative importance of the i-th indicator; After normalizing the feature vector W, we get the weight vector P of the indicator. , , represents the weight of the i-th indicator; Calculate the consistency index CI and random consistency ratio CR, perform consistency test, and judge the rationality of the matrix: , where RI is the average random consistency index corresponding to n stored in the database. If CR is less than the exceeding threshold stored in the database, the judgment matrix is considered reasonable. If CR is not less than the exceeding threshold stored in the database, the judgment matrix is considered unreasonable and the judgment matrix A is readjusted. Assume that the status level of substation equipment is divided into H levels, for the i-th indicator , its membership function Indicates the degree to which the indicator belongs to the hth state level, , let the value range of the hth state level be , then the membership function for: ; in, is the lower limit threshold of the value range of the hth state level, is the parameter threshold of the value range of the hth state level, is the upper threshold of the value range of the hth state level, , and are all constants; Based on the membership function of each indicator, the membership of each indicator to each state level is calculated, and the fuzzy evaluation matrix R is constructed. ,in, ; The weight vector P is combined with the fuzzy evaluation matrix R to obtain the comprehensive evaluation vector D. , ; According to the maximum membership principle, the state level corresponding to the element with the largest membership in the comprehensive evaluation vector D is selected as the final state level of the substation equipment; The frequency of data collection for obtaining the final status level stored in the database is defined as the frequency of data collection once.
[0008] Preferably, performing secondary frequency analysis based on the data change trend to obtain the secondary data collection frequency based on the data change trend includes the following steps: Assume that the equipment operation parameters have Q equipment indicators, which are recorded as , represents the value of the qth equipment index at time t, assuming that the environmental parameters have K environmental indicators, which are respectively , represents the value of the kth environmental index at time t, , ; For time interval , in the time interval Calculate the equipment operating parameter change rate of the qth equipment indicator , the environmental data change rate of the kth environmental indicator , the acceleration of device data change of the qth device indicator and the environmental data change acceleration of the kth environmental indicator ; Determine the comprehensive change trend assessment value : ; Among them, F1 and F2 are both transfer functions. is the weight of the qth device indicator, is the weight of the kth environmental indicator, for The weight of for The weight of for The weight of for The weight of Get The frequency of data collection corresponding to the data stored in the database is defined as the secondary data collection frequency.
[0009] Preferably, performing three-frequency analysis based on the power grid operating conditions to obtain three-frequency data acquisition frequencies based on the power grid operating conditions includes the following steps: Get the ratio of current grid load to historical maximum load value ; For each node in the power grid, the difference between its injected power and outflow power is calculated, and then the difference of all nodes is normalized to obtain the node power imbalance of the whole network. ; Get the ratio of the absolute value of the difference between the current node voltage and the rated voltage to the rated voltage ; Get the absolute value of the difference between the current grid frequency and the rated frequency ; Determine the comprehensive working condition evaluation index S: ,in, for The weight factor of for The weight factor of for The weight factor of for The weight factor of Get S corresponding to the data collection frequency stored in the database, which is defined as three data collection frequencies.
[0010] Preferably, after data preprocessing, the collected multi-source heterogeneous data is transmitted to a remote management center and stored, including the following steps: The edge computing nodes deployed near the equipment of the high-voltage prefabricated substation preprocess the collected original multi-source heterogeneous data to obtain preprocessed multi-source heterogeneous data. The preprocessing includes data cleaning and format conversion. The edge computing node transmits the pre-processed multi-source heterogeneous data to the local aggregation layer device through the industrial field bus or industrial Ethernet; The local convergence layer device aggregates and verifies the pre-processed multi-source heterogeneous data, and classifies the pre-processed original multi-source heterogeneous data according to the data type and the pre-set priority to obtain the classified multi-source heterogeneous data; The local aggregation layer equipment transmits the classified multi-source heterogeneous data to the remote management center through the encrypted optical fiber network or 5G communication link; In the remote management center, the real-time data and short-term historical data in the classified multi-source heterogeneous data are stored in an in-memory database. For long-term historical data in classified multi-source heterogeneous data, a distributed file system combined with a relational database is used for storage: The classified multi-source heterogeneous data are stored in blocks according to time series, and each block of data contains multi-source heterogeneous data within a set time range; The relational database is used to store the meta-information of classified multi-source heterogeneous data, which includes the storage location of the data block, the data collection time range, and the device identification.
[0011] Preferably, performing multi-source data analysis based on pre-processed multi-source heterogeneous data stored in a remote management center to determine the working condition of the high-voltage prefabricated substation includes the following steps: Clean the preprocessed equipment operation parameters, preprocessed environmental parameters, and preprocessed equipment assets and document data in the preprocessed multi-source heterogeneous data, remove duplicate data, correct erroneous data, process missing values, and standardize them; Based on the AE autoencoder, feature extraction is performed on the numerical data to obtain feature G1. Based on the convolutional neural network, feature extraction is performed on the image data to obtain feature G2. G1 and G2 are concatenated to obtain the fused feature vector G3. Input the fused feature vector G3 into the trained support vector machine fault classification model to determine the equipment fault type corresponding to the fused feature vector G3; Based on the equipment failure type, the time series data related to the equipment failure type is extracted from the pre-processed multi-source heterogeneous data, and the time series data related to the equipment failure type is processed according to the determined time step and input into the trained LSTM time series prediction model to obtain the future development trend of equipment failure; Input the equipment failure type, equipment failure development trend and fusion feature vector G3 into the trained random forest failure cause analysis model to obtain the equipment failure cause; The equipment failure type, equipment failure development trend and equipment failure cause are defined as the high-voltage prefabricated substation operating conditions.
[0012] Preferably, determining a substation maintenance plan based on the working condition of the high-voltage prefabricated substation includes the following steps: Compare the working condition of the high-voltage prefabricated substation with the set working condition of the high-voltage prefabricated substation stored in the database, and determine the initial maintenance plan set corresponding to the working condition of the high-voltage prefabricated substation; Obtain historical data of high-voltage prefabricated substation equipment, including historical operation data, basic equipment information and maintenance records; The initial maintenance plan set is screened based on historical data to obtain the final substation maintenance plan.
[0013] Preferably, screening the initial maintenance plan set based on historical data to obtain a final substation maintenance plan includes the following steps: Obtain the historical parameter data corresponding to each initial maintenance plan in the initial maintenance plan set, including parameter historical operation data, parameter equipment basic information and parameter maintenance records; Based on the Jaccard similarity function, the Jaccard similarity value between the basic information of the equipment and the basic information of each parameter equipment is determined, and the initial maintenance plan set is screened to obtain the maintenance plan set after preliminary screening; Determine the Euclidean distance value between the historical operation data and each of the preliminary screening maintenance plans in the preliminary screening maintenance plan set based on the Euclidean distance function, perform secondary screening on the preliminary screening maintenance plan set, and obtain a secondary screening maintenance plan set; Determine the cosine similarity value of each post-secondary screening maintenance plan in the post-secondary screening maintenance plan set based on the cosine similarity function, perform three screenings on the post-secondary screening maintenance plan set, and obtain a post-three-times screening maintenance plan set; The weighted sum of the remaining life of the substation equipment after maintenance and the inverse of the maintenance cost stored in the database corresponding to each three-times screening maintenance plan in the three-times screening maintenance plan set is performed to obtain the maintenance plan determination coefficient; The maximum value among the determination coefficients of the maintenance schemes is determined, and the maintenance scheme after three screenings corresponding to the maintenance scheme determination coefficient corresponding to the maximum value is determined as the substation maintenance scheme.
[0014] A high-voltage prefabricated substation intelligent debugging data management system, used to implement the above method, includes a data acquisition frequency determination module, a data storage module, a substation operating condition determination module and a substation maintenance plan determination module, wherein: The data collection frequency determination module is used to determine the data collection frequency based on the frequency multi-layer analysis method, and to collect multi-source heterogeneous data based on the determined data collection frequency. The multi-source heterogeneous data includes equipment operation parameters, environmental parameters, equipment assets and document data; The data storage module is used to pre-process the collected multi-source heterogeneous data, transmit it to the remote management center and store it; The substation operating condition determination module is used to perform multi-source data analysis based on the pre-processed multi-source heterogeneous data stored in the remote management center to determine the high-voltage prefabricated substation operating conditions, including equipment failure types, equipment failure development trends, and equipment failure causes; The substation maintenance plan determination module is used to determine the substation maintenance plan based on the working conditions of the high-voltage prefabricated substation.
[0015] The present invention has the following beneficial effects: The present invention determines a reasonable data collection frequency through a frequency multi-layer analysis method, and can accurately obtain multi-source heterogeneous data such as equipment operating parameters, environmental parameters, equipment assets and documents, avoiding the problem of insufficient or excessive data collection. The collected data is preprocessed and remotely stored to ensure data quality and traceability. Multi-source data analysis can accurately determine the type, development trend and cause of equipment failure, and provide comprehensive and timely equipment operating condition information for operation and maintenance personnel. Finally, the maintenance plan is determined based on the operating conditions, which can achieve accurate and efficient equipment maintenance, reduce the probability of failure, reduce maintenance costs, improve the reliability and stability of substation operation, extend the service life of equipment, and ensure the safety and continuity of power supply. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 This is a flow chart of the intelligent debugging data management method for high-voltage prefabricated substation of the present invention; Figure 2 This is a module connection diagram of the intelligent debugging data management system for a high-voltage prefabricated substation of the present invention. DETAILED DESCRIPTION
[0017] This embodiment uses a high-voltage prefabricated substation intelligent debugging data management method and system to achieve the effect of using a frequency multi-layer analysis method to reasonably determine the data collection frequency and determine the maintenance plan.
[0018] Embodiment 1: like Figure 1 As shown, a method for intelligent commissioning data management of a high-voltage prefabricated substation includes the following steps: determining a data acquisition frequency based on a frequency multi-layer analysis method, and performing multi-source heterogeneous data acquisition based on the determined data acquisition frequency, wherein the multi-source heterogeneous data includes equipment operating parameters, environmental parameters, and equipment assets and document data; Perform a frequency analysis based on equipment operating parameters and environmental parameters to obtain a data collection frequency based on the equipment operating parameters and environmental parameters. The equipment operating parameters include transformer operating data, high-voltage switchgear operating data, low-voltage switchgear operating data, and capacitor compensation device operating data. The equipment operating parameters and environmental parameters are standardized to obtain the standardized primary frequency determination data; the equipment operating parameters and environmental parameters are standardized to make data of different dimensions and magnitudes on the same scale, which is convenient for comparison and analysis, and improves the accuracy and efficiency of subsequent analysis.
[0019] Based on the set hierarchical structure, the normalized frequency determination data is divided to obtain the judgment matrix A. , n is the number of indicators in the frequency-determined data after standardization, Indicates the importance of the i-th indicator relative to the j-th indicator, . Use the 1-9 standard method to assign values, for example: Scale meaning 1 The i-th indicator is as important as the j-th indicator 3 The i-th indicator is slightly more important than the j-th indicator 5 The i-th indicator and the j-th indicator are obviously important 7 The i-th indicator is strongly important to the j-th indicator 9 The i-th indicator and the j-th indicator are extremely important 2,4,6,8 The middle value of the above adjacent judgments The 1-9 standard method is used to construct the judgment matrix A to determine the relative importance of each indicator. This can quantify people's subjective judgment on the importance of each indicator, making the analysis more scientific and objective, and avoiding the arbitrariness of determining weights based solely on experience or intuition.
[0020] Solve the characteristic equation of judgment matrix A , where I represents the identity matrix, Denotes the eigenvalue of A, and obtains the maximum eigenvalue and The corresponding eigenvector W is , T is the transposition symbol, The value representing the relative importance of the ith indicator; the weight vector P of the indicator is obtained by normalizing the feature vector W. , , represents the weight of the i-th indicator; Calculate the consistency index CI and random consistency ratio CR, perform consistency test, and judge the rationality of the matrix: , where RI is the average random consistency index corresponding to n stored in the database. If CR is less than the exceeding threshold stored in the database, the judgment matrix is considered reasonable. If CR is not less than the exceeding threshold stored in the database, the judgment matrix is considered unreasonable and the judgment matrix A is readjusted. Readjusting the judgment matrix A ensures the accuracy and reliability of weight determination, making subsequent weight-based analysis results more credible.
[0021] Assume that the status level of substation equipment is divided into H levels, for the i-th indicator , its membership function Indicates the degree to which the indicator belongs to the hth state level, , let the value range of the hth state level be , then the membership function for: ; in, is the lower limit threshold of the value range of the hth state level, is the parameter threshold of the value range of the hth state level, is the upper threshold of the value range of the hth state level, , and They are all constants and can be set according to actual needs; Based on the membership function of each indicator, the membership of each indicator to each state level is calculated, and the fuzzy evaluation matrix R is constructed. ,in, (That is, for To represent ); The fuzzy evaluation matrix R is used to comprehensively consider the membership of each indicator, which is in line with the characteristics of fuzziness and uncertainty of the substation equipment status. It can more accurately describe the actual operating status of the equipment and is more in line with the actual situation than the traditional deterministic evaluation method.
[0022] The weight vector P is combined with the fuzzy evaluation matrix R to obtain the comprehensive evaluation vector D. , ; According to the maximum membership principle, the state level corresponding to the element with the largest membership in the comprehensive evaluation vector D is selected as the final state level of the substation equipment; the data collection frequency of the final state level stored in the database is obtained, which is defined as the primary data collection frequency. The weight vector P is synthesized with the fuzzy evaluation matrix R to obtain the comprehensive evaluation vector D, and then the final state level of the equipment is determined according to the maximum membership principle. It can comprehensively consider the weights and memberships of each indicator, and obtain an evaluation result that fully reflects the overall state of the equipment, which provides a strong basis for determining a reasonable data collection frequency.
[0023] Considering multiple indicators including substation equipment and environment, and constructing a membership function for each indicator to indicate the degree to which it belongs to different status levels, it can comprehensively and meticulously characterize the operating status of the equipment, avoiding the problem of incomplete equipment status evaluation caused by focusing on a single indicator, and can adjust the collection frequency according to different equipment operating status and environmental conditions. According to the final status level, the corresponding data collection frequency is obtained, so that the data collection frequency matches the actual operating status of the equipment, realizing the precision and intelligence of data collection, which not only ensures the validity of the data, but also makes rational use of resources.
[0024] Perform secondary frequency analysis based on data change trends to obtain secondary data collection frequencies based on data change trends. Data change trends include equipment operating parameter change trends and environmental parameter change trends. Assume that the equipment operation parameters have Q equipment indicators, which are recorded as , represents the value of the qth equipment index at time t, assuming that the environmental parameters have K environmental indicators, which are respectively , represents the value of the kth environmental index at time t, , ; For the time interval , in the time interval Calculate the equipment operating parameter change rate of the qth equipment indicator , the environmental data change rate of the kth environmental indicator , the acceleration of device data change of the qth device indicator and the environmental data change acceleration of the kth environmental indicator ; Determine the comprehensive change trend assessment value : ; Among them, F1 and F2 are both transfer functions. is the weight of the qth device indicator, is the weight of the kth environmental indicator, for The weight of for The weight of for The weight of for The weight of The frequency of data collection corresponding to the data stored in the database is defined as the secondary data collection frequency; weights are assigned to equipment indicators and environmental indicators respectively, as well as to the change rate and change acceleration, so that the impact of different factors on the comprehensive change trend evaluation value can be reasonably weighed according to the actual situation. Different equipment indicators and environmental indicators have different importance to the operation of the substation, and their change rate and acceleration may also be different. The relative importance of each factor can be more scientifically reflected through the setting of weights, making the comprehensive evaluation results more reasonable.
[0025] Get The frequency of data collection corresponding to the data stored in the database is defined as the secondary data collection frequency.
[0026] In terms of data feature capture, the system not only considers the rate of change of equipment operating parameters and environmental parameters, but also introduces the acceleration of equipment data change and environmental data change, which can comprehensively capture the dynamic change characteristics of data from multiple dimensions. The rate of change can reflect the speed of data change within a certain period of time, while the acceleration can further reflect whether the trend of change is intensifying or slowing down, making the description of data change trends more detailed and accurate. By comprehensively considering multiple change factors, it can better adapt to the complex changes of data in actual operation and avoid misjudgment of data change trends due to focusing on a single factor.
[0027] Determining the data collection frequency based on the comprehensive change trend evaluation value can closely combine the collection frequency with the actual changes in the data. When the data changes dramatically, that is, the comprehensive change trend evaluation value is high, the data collection frequency is increased accordingly to capture data changes more timely and obtain more key information; when the data changes relatively smoothly, the collection frequency is reduced to save data collection and storage resources while ensuring data quality, thus achieving accurate dynamic matching between data collection frequency and data change trend.
[0028] Performing three-frequency analysis based on the power grid operation conditions to obtain three-frequency data collection frequencies based on the power grid operation conditions; Get the ratio of current grid load to historical maximum load value ; For each node in the power grid, calculate the difference between its injected power and outflow power, and then normalize the difference of all nodes to obtain the node power imbalance of the whole network ; Get the absolute value of the difference between the current node voltage and the rated voltage and the ratio of the rated voltage ; Get the absolute value of the difference between the current grid frequency and the rated frequency ; Determine the comprehensive working condition evaluation index S: ,in, for The weight factor of for The weight factor of for The weight factor of for The weight factor of S is obtained; the data collection frequency corresponding to S stored in the database is defined as three data collection frequencies.
[0029] By considering , , and , which can comprehensively reflect the operating conditions of the power grid from different angles. These indicators cover important aspects such as power grid load, power balance, voltage stability and frequency stability, making the monitoring of the overall operating status of the power grid more comprehensive and accurate. Calculating the comprehensive operating condition evaluation index S based on multiple indicators can make full use of the information contained in each indicator. Multiple indicators complement and verify each other, which can more comprehensively and accurately reflect the actual operating conditions of the power grid, reduce misjudgments caused by fluctuations or errors in individual indicators, thereby improving the reliability of the comprehensive evaluation and providing a more solid basis for determining the frequency of data collection.
[0030] Determining the data collection frequency based on the comprehensive operating condition evaluation index S can make the data collection frequency closely related to the actual operating condition of the power grid. When the operating condition of the power grid is more complex or there are potential risks, that is, when the comprehensive operating condition evaluation index S shows a large change or anomaly, the data collection frequency is increased accordingly to monitor the power grid operation data more intensively and promptly discover and deal with possible problems; when the power grid operation is relatively stable, the collection frequency is reduced. This can effectively save data collection, transmission and storage resources while ensuring the acquisition of key information, realize data collection on demand, and improve resource utilization efficiency.
[0031] Determine the maximum value among the primary data acquisition frequency, the secondary data acquisition frequency and the tertiary data acquisition frequency, and determine the maximum value as the pending acquisition frequency; judge whether the pending acquisition frequency is greater than the original acquisition frequency: if the pending acquisition frequency is greater than the original acquisition frequency, determine the pending acquisition frequency as the data acquisition frequency; if the pending acquisition frequency is not greater than the original acquisition frequency, determine the original acquisition frequency as the data acquisition frequency.
[0032] After data preprocessing, the collected multi-source heterogeneous data is transmitted to the remote management center and stored; The edge computing nodes deployed near the equipment of the high-voltage prefabricated substation preprocess the collected original multi-source heterogeneous data to obtain preprocessed multi-source heterogeneous data. The preprocessing includes data cleaning and format conversion. The original data collected by the equipment of the high-voltage prefabricated substation often have various quality problems and format differences. Direct transmission and storage of these original data will bring many difficulties to subsequent work, such as low data processing efficiency and inaccurate analysis results. By performing data cleaning and format conversion at the edge computing node, the data can be preliminarily processed at the source of the data to solve these problems.
[0033] The edge computing node transmits the pre-processed multi-source heterogeneous data to the local convergence layer device through the industrial field bus or industrial Ethernet. The industrial field bus includes Profibus and Modbus. The local convergence layer device summarizes and verifies the pre-processed multi-source heterogeneous data, and classifies the pre-processed original multi-source heterogeneous data according to the data type and pre-set priority to obtain the classified multi-source heterogeneous data. The edge computing node is connected to the local convergence layer device using the industrial field bus (such as Profibus and Modbus) or industrial Ethernet. These communication methods have high reliability, real-time and anti-interference capabilities, and can ensure that the pre-processed data is accurately and quickly transmitted to the local convergence layer device. The local convergence layer device summarizes, verifies and classifies the data, further ensuring the integrity and accuracy of the data. At the same time, classification according to data type and priority is conducive to optimizing the allocation of resources for subsequent transmission and processing.
[0034] The local aggregation layer equipment transmits the classified multi-source heterogeneous data to the remote management center through the encrypted fiber optic network or 5G communication link; the classified data is transmitted to the remote management center using the encrypted fiber optic network or 5G communication link. The encrypted fiber optic network has the characteristics of high bandwidth, low latency and high security, and the 5G communication link has the advantages of high speed and flexibility, which can meet the needs of remote data transmission in different scenarios, ensuring that a large amount of multi-source heterogeneous data can be transmitted to the remote management center for processing and storage in a timely and secure manner.
[0035] Considering the wide distribution of substation equipment, data needs to be transmitted from the site to the remote management center for centralized management and analysis. Different communication methods are suitable for different scenarios and distances. Through this design combining layered transmission and multiple communication methods, the most appropriate transmission method can be selected according to actual conditions to ensure the stability and efficiency of data transmission. At the same time, data aggregation, verification and classification can be performed to further manage and optimize data during transmission, thereby improving the quality and efficiency of data transmission.
[0036] In the remote management center, the real-time data and short-term historical data in the classified multi-source heterogeneous data are stored in an in-memory database, including Redis. The long-term historical data in the classified multi-source heterogeneous data is stored in a distributed file system combined with a relational database, including Ceph, and the relational database includes PostgreSQL. The classified multi-source heterogeneous data is divided into blocks and stored according to time series, and each block of data contains multi-source heterogeneous data within a set time range. The relational database is used to store the metadata of the classified multi-source heterogeneous data, and the metadata includes the storage location of the data block, the data collection time range, and the device identification.
[0037] Multi-source heterogeneous data has different characteristics and application requirements. Real-time data and short-term historical data require fast access and processing, while long-term historical data requires a large amount of storage space and effective organization and management. Using different storage technologies and methods can optimize storage according to the characteristics of data, meet the needs of different businesses for data storage and access, and improve the performance and scalability of the entire data storage system.
[0038] Perform multi-source data analysis based on pre-processed multi-source heterogeneous data stored in the remote management center to determine the working conditions of the high-voltage prefabricated substation, including equipment failure types, equipment failure development trends, and equipment failure causes; The preprocessed equipment operation parameters, preprocessed environmental parameters, and preprocessed equipment assets and document data in the preprocessed multi-source heterogeneous data are cleaned, duplicate data is removed, erroneous data is corrected, missing values are processed, and standardized; the numerical data is feature extracted based on the AE autoencoder to obtain feature G1, and the image data is feature extracted based on the convolutional neural network to obtain feature G2, and G1 and G2 are spliced to obtain the fusion feature vector G3; multi-source heterogeneous data includes equipment operation parameters, environmental parameters, equipment assets and document data, etc. The data quality is uneven, and different types of data have different characteristics and structures. Through targeted data preprocessing and feature extraction methods, the value of different types of data can be fully explored, the availability of data and the effectiveness of features can be improved, thereby improving the accuracy and reliability of the entire analysis process.
[0039] The fused feature vector G3 is input into the trained support vector machine fault classification model to determine the equipment fault type corresponding to the fused feature vector G3; based on the equipment fault type, the time series data related to the equipment fault type is extracted from the preprocessed multi-source heterogeneous data, and the time series data related to the equipment fault type is processed according to the determined time step and input into the trained LSTM time series prediction model to obtain the future equipment failure development trend; the equipment failure type, equipment failure development trend and fused feature vector G3 are input into the trained random forest fault cause analysis model to obtain the equipment failure cause; the equipment failure type, equipment failure development trend and equipment failure cause are defined as the working condition of the high-voltage prefabricated substation.
[0040] In high-voltage prefabricated substations, equipment failure types are diverse, and accurate identification of failure types is crucial for timely taking corresponding maintenance measures and ensuring safe operation of the power grid. As a mature classification algorithm, support vector machine can effectively handle high-dimensional data and complex classification problems. Combined with fusion feature vectors, it can better adapt to the characteristics of multi-source heterogeneous data and achieve accurate fault type classification.
[0041] The support vector machine (SVM) fault classification model is a supervised learning model that can be used to classify and identify equipment fault types. The SVM model is trained with the fused feature vector as input and the equipment fault type label as output. During the training process, the classification performance of the model is optimized by adjusting hyperparameters such as the kernel function (such as linear kernel, radial basis kernel, etc.) and the penalty parameter C. For example, using the radial basis kernel function and selecting the appropriate C value through cross-validation, the SVM model can accurately distinguish different types of faults such as overheating faults and short circuit faults of transformers.
[0042] LSTM is a recurrent neural network that is suitable for processing time series data. Time series data from multi-source data (such as equipment temperature, current, and other time-varying data) is input into the LSTM model to predict the future trend of equipment operating parameters. During the training process, appropriate hyperparameters such as the time step and the number of hidden layer neurons are determined so that the model can accurately capture the time dependency of the data. For example, the LSTM model is used to predict the changing trend of transformer oil temperature in the next few hours.
[0043] Random forest is an integrated learning model composed of multiple decision trees. The fused feature vector and the known fault cause labels are used as input to train the random forest model. During the training process, the analysis accuracy of the model is improved by adjusting hyperparameters such as the number of decision trees and the maximum depth of each decision tree. For example, the random forest model is used to analyze the possible causes of transformer overheating failure, such as cooling system failure and winding short circuit.
[0044] In a specific embodiment, the steps of obtaining the working condition of the high-voltage prefabricated substation are as follows: From the fused feature vector dataset, divide the training set and test set according to a certain ratio (such as 70%-30% or 80%-20%). Ensure that the sample distribution of each fault type in the training set and the test set is relatively balanced to avoid bias in the model. For example, if there are three types of overheating faults, short circuit faults, and insulation faults, the number of samples of each type of fault in the training set is roughly the same. Normalize the fused feature vectors of the training set and the test set so that their feature values are in the range of [0,1] or [-1,1].
[0045] Kernel function selection: Common kernel functions include linear kernel, radial basis kernel (RBF), polynomial kernel, etc. Generally, we start with the radial basis kernel because it is more adaptable to data distribution. Hyperparameter adjustment: Use cross-validation methods, such as 5-fold cross-validation, to divide the training set into 5 subsets, and use 4 subsets to train the model each time, and 1 subset for validation. For different hyperparameter combinations, calculate the accuracy of the model on the validation set. Select the hyperparameter combination with the highest accuracy as the optimal hyperparameter.
[0046] Train the SVM model using the optimal hyperparameters and training set. Call the corresponding machine learning library (such as the SVC class in Scikit-learn), pass in the training data and labels for model training. Use the test set to evaluate the trained SVM model and calculate indicators such as accuracy, precision, recall, and F1 value.
[0047] Extract time series data related to equipment operating parameters from multi-source data, such as oil temperature and current of transformers. Normalize these time series data, and you can also use minimum-maximum normalization or Z-score normalization. Convert the time series data into a format suitable for LSTM model input. For example, determine the time step X, divide the time series data into multiple sequence segments of length X, and use the last value of each sequence segment as the prediction target, and the previous X-1 values as input features.
[0048] Similarly, divide the training set and test set according to a certain ratio. Determine the hyperparameters (determine the time step, the number of neurons in the hidden layer, and the learning rate). Use deep learning frameworks such as Keras or PyTorch to build the LSTM model, and set the structure of the input layer, LSTM layer, and output layer. Use the training set to train the model, set the appropriate number of training rounds (such as 100-200 rounds) and batch size (such as 32, 64, etc.). Use the test set to evaluate the trained model, calculate indicators such as mean square error (MSE) and root mean square error (RMSE), and evaluate the prediction accuracy of the model.
[0049] Combine the fused feature vector and the known fault cause labels into a data set, and divide it into training and test sets. Feature selection can be performed on the feature vector to remove some features that contribute less to fault cause analysis and improve the training efficiency and accuracy of the model. You can use a method based on feature importance sorting, such as the feature importance evaluation function of random forest, to select the first few features with higher importance.
[0050] Perform hyperparameter adjustment (the number of decision trees and the maximum depth of each decision tree), use the RandomForestClassifier class in Scikit-learn, and pass in the training set data and labels for model training. Use the test set to evaluate the trained model, calculate indicators such as accuracy, precision, and recall, and evaluate the accuracy of the model in fault cause analysis.
[0051] After the training is completed, the normalized fusion feature vector is input into the trained SVM fault classification model. The model calculates the score or probability that the input feature vector belongs to each fault type based on the classification boundary obtained through training. The fault type with the highest score or the highest probability is selected as the preliminary diagnosis result. For example, if the calculated probability that the input feature vector belongs to an overheating fault is 0.8, the probability that it belongs to a short circuit fault is 0.1, and the probability that it belongs to an insulation fault is 0.1, then the preliminary diagnosis is that the equipment has an overheating fault.
[0052] According to the preliminary diagnosis results of the SVM model, extract the time series data related to the fault type from the multi-source data. For example, if the diagnosis is an overheating fault, extract the oil temperature time series data of the transformer. Process the extracted time series data according to the previously determined time step and input it into the trained LSTM time series prediction model. The model outputs the predicted values of the equipment operating parameters in the future, such as predicting the change trend of the transformer oil temperature in the next 3 hours. The prediction results can be plotted as a curve to intuitively show the development trend of the fault.
[0053] Combine the classification results of the SVM model (such as the fault type label), the prediction results of the LSTM model (such as the parameter prediction value in the future period), and the original fused feature vector into new input data. Input the new input data into the trained random forest fault cause analysis model. The model analyzes the specific cause of the fault according to the decision rules obtained through training. The model outputs the probability or score of each fault cause, and selects the cause with the highest probability or the highest score as the main fault cause. For example, the model outputs that the probability of the transformer overheating fault due to cooling fan damage is 0.7, the probability of winding short circuit is 0.2, and the probability of overload is 0.1, then it is determined that the main fault cause is cooling fan damage.
[0054] Determine the substation maintenance plan based on the operating conditions of the high-voltage prefabricated substation.
[0055] The working condition of the high-voltage prefabricated substation is compared with the set working condition of the high-voltage prefabricated substation stored in the database to determine the initial maintenance plan set corresponding to the working condition of the high-voltage prefabricated substation; the historical data of the high-voltage prefabricated substation equipment is obtained, and the historical data includes historical operation data (including but not limited to temperature, humidity, current, voltage, operation time, number of failures), basic equipment information (such as model, production date, installation date, etc.) and maintenance records (such as maintenance time, replacement of parts, etc.); the initial maintenance plan set is screened based on the historical data to obtain the final substation maintenance plan.
[0056] By comparing the current working conditions of the high-voltage prefabricated substation with the set working conditions stored in the database, it is possible to quickly locate the set working conditions similar to the current situation and their corresponding maintenance plans, making the initial maintenance plan set more targeted. The set working conditions and corresponding maintenance plans stored in the database are obtained through long-term accumulation and summary, and contain a lot of practical experience and professional knowledge. By comparing these historical data, we can make full use of existing resources and improve the efficiency and quality of maintenance plan formulation.
[0057] Obtain the historical parameter data corresponding to each initial maintenance plan in the initial maintenance plan set, including parameter historical operation data, parameter equipment basic information and parameter maintenance records; determine the Jaccard similarity value between the basic information of the equipment and the basic information of each parameter equipment based on the Jaccard similarity function, screen the initial maintenance plan set, and obtain a preliminary screened maintenance plan set (when the Jaccard similarity value corresponding to one of the initial maintenance plans is greater than the first threshold stored in the database, the initial maintenance plan is used as one of the preliminary screened maintenance plan set); equipment of different models, production dates and installation dates may have significant differences in performance, structure and usage, and the corresponding maintenance requirements and plans will also be different.
[0058] The Jaccard similarity function is suitable for processing classified data. The basic information of equipment (such as model, production date, installation date, etc.) is mostly classified or discrete data. By calculating the Jaccard similarity value of the basic information of the equipment and the basic information of the reference equipment, maintenance plans with similar basic conditions of the equipment can be screened out. This helps to exclude those plans that are not applicable due to differences in the characteristics of the equipment itself, so that the maintenance plan set after preliminary screening is more in line with the actual situation of the current substation equipment.
[0059] Based on the Euclidean distance function, the Euclidean distance values between the historical operation data and each preliminary screening maintenance plan in the preliminary screening maintenance plan set are determined, and the preliminary screening maintenance plan set is secondary screened to obtain a secondary screening maintenance plan set (when the Euclidean distance value corresponding to one of the preliminary screening maintenance plans is not greater than the second threshold stored in the database, the preliminary screening maintenance plan is used as one of the secondary screening maintenance plan set); even if the basic information of the equipment is similar, its operating status may vary due to the actual use environment and operating conditions. The operating data reflects the current working status and performance of the equipment.
[0060] The Euclidean distance function is often used to measure the distance between continuous numerical data. Historical operating data (such as temperature, humidity, current, voltage, operating time, number of faults, etc.) are continuous numerical data. By calculating the Euclidean distance between historical operating data and the maintenance plans after preliminary screening, secondary screening can be performed to further select plans with similar operating conditions. This ensures that the maintenance plan set after secondary screening is more compatible with the current substation in terms of equipment operating status.
[0061] Based on the cosine similarity function, the cosine similarity values between the maintenance records and each secondary screening maintenance plan in the secondary screening maintenance plan set are determined, and the secondary screening maintenance plan set is screened three times to obtain a three-time screening maintenance plan set (when the cosine similarity value corresponding to one of the secondary screening maintenance plans is greater than the third threshold stored in the database, the secondary screening maintenance plan is taken as one of the three-time screening maintenance plan set); the maintenance history of the equipment is an important reflection of its operating reliability and failure mode, and similar maintenance records may mean similar failure causes and solutions.
[0062] The cosine similarity function is used to measure the cosine value of the angle between vectors and is suitable for processing high-dimensional vector data. Maintenance records (such as maintenance time, replacement parts, etc.) can be regarded as a kind of high-dimensional vector data. By calculating the cosine similarity value of the maintenance record and the maintenance plan after the second screening, the third screening can find the plan with similar maintenance history. This helps to consider the past maintenance experience and fault handling methods of the equipment, so that the maintenance plan set after the third screening is more in line with the actual maintenance needs of the equipment.
[0063] The weighted sum of the remaining life of the substation equipment after maintenance and the inverse of the maintenance cost stored in the database corresponding to each three-times screening maintenance plan in the three-times screening maintenance plan set is performed to obtain the maintenance plan determination coefficient; the maximum value of each maintenance plan determination coefficient is determined, and the three-times screening maintenance plan corresponding to the maintenance plan determination coefficient corresponding to the maximum value is determined as the substation maintenance plan.
[0064] In the above three screening processes: If there is only one preliminary screening maintenance plan in the set of preliminary screening maintenance plans, then the preliminary screening maintenance plan is directly used as the substation maintenance plan. In addition, if there is no preliminary screening maintenance plan in the set of preliminary screening maintenance plans, then the initial maintenance plan corresponding to the largest Jaccard similarity value is used as the substation maintenance plan.
[0065] If there is only one secondary screening maintenance plan in the set of secondary screening maintenance plans, then the secondary screening maintenance plan is directly used as the substation maintenance plan. In addition, if there is no secondary screening maintenance plan in the set of secondary screening maintenance plans, then the initial screening maintenance plan corresponding to the smallest Euclidean distance value is used as the substation maintenance plan.
[0066] If there is only one three-times screening maintenance plan in the three-times screening maintenance plan set, then the three-times screening maintenance plan is directly used as the substation maintenance plan. In addition, if there is no three-times screening maintenance plan in the three-times screening maintenance plan set, then the secondary screening maintenance plan corresponding to the largest cosine similarity value is used as the substation maintenance plan.
[0067] Embodiment 2: like Figure 2 As shown, a high-voltage prefabricated substation intelligent debugging data management system is used to implement the method in Example 1, including a data acquisition frequency determination module, a data storage module, a substation operating condition determination module and a substation maintenance plan determination module, wherein: the data acquisition frequency determination module is used to determine the data acquisition frequency based on the frequency multi-layer analysis method, and perform multi-source heterogeneous data acquisition based on the determined data acquisition frequency, and the multi-source heterogeneous data includes equipment operating parameters, environmental parameters, and equipment assets and document data; the data storage module is used to perform data preprocessing on the collected multi-source heterogeneous data, and then transmit it to a remote management center and store it; the substation operating condition determination module is used to perform multi-source data analysis based on the preprocessed multi-source heterogeneous data stored in the remote management center to determine the operating condition of the high-voltage prefabricated substation, including the equipment failure type, the equipment failure development trend and the equipment failure cause; the substation maintenance plan determination module is used to determine the substation maintenance plan based on the high-voltage prefabricated substation operating condition.
[0068] Embodiment 3: An electronic device includes: a processor; and a memory, wherein computer program instructions are stored in the memory, and when the computer program instructions are executed by the processor, the processor executes the method in embodiment 1.
[0069] Embodiment 4: A computer-readable storage medium is used to store a program, and when the program is executed by a processor, the method in embodiment 1 is implemented.
Claims
1. A method for managing intelligent commissioning data of a high-voltage prefabricated substation, characterized in that: The following steps are involved: Determine the data collection frequency based on the frequency multi-layer analysis method, and collect multi-source heterogeneous data based on the determined data collection frequency. The multi-source heterogeneous data includes equipment operation parameters, environmental parameters, equipment assets and document data; After data preprocessing, the collected multi-source heterogeneous data is transmitted to the remote management center and stored; Perform multi-source data analysis based on pre-processed multi-source heterogeneous data stored in the remote management center to determine the working conditions of the high-voltage prefabricated substation, including equipment failure types, equipment failure development trends, and equipment failure causes; Determine the substation maintenance plan based on the operating conditions of the high-voltage prefabricated substation.
2. A method for managing intelligent debugging data of a high-voltage prefabricated substation according to claim 1, characterized in that: The frequency of data collection is determined based on the frequency multi-level analysis method, including the following steps: Perform a frequency analysis based on equipment operating parameters and environmental parameters to obtain a data collection frequency based on the equipment operating parameters and environmental parameters. The equipment operating parameters include transformer operating data, high-voltage switchgear operating data, low-voltage switchgear operating data, and capacitor compensation device operating data. Perform secondary frequency analysis based on data change trends to obtain secondary data collection frequencies based on data change trends. Data change trends include equipment operating parameter change trends and environmental parameter change trends. Performing three-frequency analysis based on the power grid operation conditions to obtain three-frequency data collection frequencies based on the power grid operation conditions; Determine the maximum value among the primary data collection frequency, the secondary data collection frequency and the tertiary data collection frequency, and determine the maximum value as the to-be-determined collection frequency; Determine whether the pending acquisition frequency is greater than the original acquisition frequency: If the pending acquisition frequency is greater than the original acquisition frequency, the pending acquisition frequency is determined as the data acquisition frequency; If the pending acquisition frequency is not greater than the original acquisition frequency, the original acquisition frequency is determined as the data acquisition frequency.
3. A method for managing intelligent debugging data of a high-voltage prefabricated substation according to claim 2, characterized in that: Performing a frequency analysis based on the equipment operating parameters and the environmental parameters to obtain a data collection frequency based on the equipment operating parameters and the environmental parameters includes the following steps: Standardize the equipment operating parameters and environmental parameters to obtain the standardized primary frequency determination data; Based on the set hierarchical structure, the normalized frequency determination data is divided to obtain the judgment matrix A. , n is the number of indicators in the frequency-determined data after standardization, Indicates the importance of the i-th indicator relative to the j-th indicator, ; Solve the characteristic equation of judgment matrix A , where I represents the identity matrix, Denotes the eigenvalue of A, and obtains the maximum eigenvalue and The corresponding eigenvector W is , T is the transposition symbol, The value representing the relative importance of the i-th indicator; After normalizing the feature vector W, we get the weight vector P of the indicator. , , represents the weight of the i-th indicator; Calculate the consistency index CI and random consistency ratio CR, perform consistency test, and judge the rationality of the matrix: , where RI is the average random consistency index corresponding to n stored in the database. If CR is less than the exceeding threshold stored in the database, the judgment matrix is considered reasonable. If CR is not less than the exceeding threshold stored in the database, the judgment matrix is considered unreasonable and the judgment matrix A is readjusted. Assume that the status level of substation equipment is divided into H levels, for the i-th indicator , its membership function Indicates the degree to which the indicator belongs to the hth state level, , let the value range of the hth state level be , then the membership function for: ; in, is the lower limit threshold of the value range of the hth state level, is the parameter threshold of the value range of the hth state level, is the upper threshold of the value range of the hth state level, , and are all constants; Based on the membership function of each indicator, the membership of each indicator to each state level is calculated, and the fuzzy evaluation matrix R is constructed. ,in, ; The weight vector P is combined with the fuzzy evaluation matrix R to obtain the comprehensive evaluation vector D. , ; According to the maximum membership principle, the state level corresponding to the element with the largest membership in the comprehensive evaluation vector D is selected as the final state level of the substation equipment; The frequency of data collection for obtaining the final status level stored in the database is defined as the frequency of data collection once.
4. A method for managing intelligent debugging data of a high-voltage prefabricated substation according to claim 2, characterized in that: Performing secondary frequency analysis based on the data change trend to obtain the secondary data collection frequency based on the data change trend includes the following steps: Assume that the equipment operation parameters have Q equipment indicators, which are recorded as , represents the value of the qth equipment index at time t, assuming that the environmental parameters have K environmental indicators, which are respectively , represents the value of the kth environmental index at time t, , ; For time interval , in the time interval Calculate the equipment operating parameter change rate of the qth equipment indicator , the environmental data change rate of the kth environmental indicator , the acceleration of device data change of the qth device indicator and the environmental data change acceleration of the kth environmental indicator ; Determine the comprehensive change trend assessment value : ; Among them, F1 and F2 are both transfer functions. is the weight of the qth device indicator, is the weight of the kth environmental indicator, for The weight of for The weight of for The weight of for The weight of Get The frequency of data collection corresponding to the data stored in the database is defined as the secondary data collection frequency.
5. A method for managing intelligent debugging data of a high-voltage prefabricated substation according to claim 2, characterized in that: Performing a three-frequency frequency analysis based on the power grid operation condition to obtain three-frequency data acquisition frequencies based on the power grid operation condition includes the following steps: Get the ratio of current grid load to historical maximum load value ; For each node in the power grid, the difference between its injected power and outflow power is calculated, and then the difference of all nodes is normalized to obtain the node power imbalance of the whole network. ; Get the ratio of the absolute value of the difference between the current node voltage and the rated voltage to the rated voltage ; Get the absolute value of the difference between the current grid frequency and the rated frequency ; Determine the comprehensive working condition evaluation index S: ,in, for The weight factor of for The weight factor of for The weight factor of for The weight factor of Get S corresponding to the data collection frequency stored in the database, which is defined as three data collection frequencies.
6. A method for managing intelligent debugging data of a high-voltage prefabricated substation according to claim 1, characterized in that: After data preprocessing, the collected multi-source heterogeneous data is transmitted to the remote management center and stored, including the following steps: The edge computing nodes deployed near the equipment of the high-voltage prefabricated substation preprocess the collected original multi-source heterogeneous data to obtain preprocessed multi-source heterogeneous data. The preprocessing includes data cleaning and format conversion. The edge computing node transmits the pre-processed multi-source heterogeneous data to the local aggregation layer device through the industrial field bus or industrial Ethernet; The local convergence layer device aggregates and verifies the pre-processed multi-source heterogeneous data, and classifies the pre-processed original multi-source heterogeneous data according to the data type and the pre-set priority to obtain the classified multi-source heterogeneous data; The local aggregation layer equipment transmits the classified multi-source heterogeneous data to the remote management center through the encrypted optical fiber network or 5G communication link; In the remote management center, the real-time data and short-term historical data in the classified multi-source heterogeneous data are stored in an in-memory database. For long-term historical data in classified multi-source heterogeneous data, a distributed file system combined with a relational database is used for storage: The classified multi-source heterogeneous data are stored in blocks according to time series, and each block of data contains multi-source heterogeneous data within a set time range; The relational database is used to store the meta-information of classified multi-source heterogeneous data, which includes the storage location of the data block, the data collection time range, and the device identification.
7. A method for managing intelligent debugging data of a high-voltage prefabricated substation according to claim 1, characterized in that: The multi-source data analysis is performed based on the pre-processed multi-source heterogeneous data stored in the remote management center to determine the working condition of the high-voltage prefabricated substation, including the following steps: Clean the preprocessed equipment operation parameters, preprocessed environmental parameters, and preprocessed equipment assets and document data in the preprocessed multi-source heterogeneous data, remove duplicate data, correct erroneous data, process missing values, and standardize them; Based on the AE autoencoder, feature extraction is performed on the numerical data to obtain feature G1. Based on the convolutional neural network, feature extraction is performed on the image data to obtain feature G2. G1 and G2 are concatenated to obtain the fused feature vector G3. Input the fused feature vector G3 into the trained support vector machine fault classification model to determine the equipment fault type corresponding to the fused feature vector G3; Based on the equipment failure type, the time series data related to the equipment failure type is extracted from the pre-processed multi-source heterogeneous data, and the time series data related to the equipment failure type is processed according to the determined time step and input into the trained LSTM time series prediction model to obtain the future development trend of equipment failure; Input the equipment failure type, equipment failure development trend and fusion feature vector G3 into the trained random forest failure cause analysis model to obtain the equipment failure cause; The equipment failure type, equipment failure development trend and equipment failure cause are defined as the high-voltage prefabricated substation operating conditions.
8. A method for managing intelligent debugging data of a high-voltage prefabricated substation according to claim 1, characterized in that: Determine the substation maintenance plan based on the working condition of the high-voltage prefabricated substation, including the following steps: Compare the working condition of the high-voltage prefabricated substation with the set working condition of the high-voltage prefabricated substation stored in the database, and determine the initial maintenance plan set corresponding to the working condition of the high-voltage prefabricated substation; Obtain historical data of high-voltage prefabricated substation equipment, including historical operation data, basic equipment information and maintenance records; The initial maintenance plan set is screened based on historical data to obtain the final substation maintenance plan.
9. A method for managing intelligent debugging data of a high-voltage prefabricated substation according to claim 8, characterized in that: The initial maintenance plan set is screened based on historical data to obtain the final substation maintenance plan, including the following steps: Obtain the historical parameter data corresponding to each initial maintenance plan in the initial maintenance plan set, including parameter historical operation data, parameter equipment basic information and parameter maintenance records; Based on the Jaccard similarity function, the Jaccard similarity value between the basic information of the equipment and the basic information of each parameter equipment is determined, and the initial maintenance plan set is screened to obtain the maintenance plan set after preliminary screening; Determine the Euclidean distance value between the historical operation data and each of the preliminary screening maintenance plans in the preliminary screening maintenance plan set based on the Euclidean distance function, perform secondary screening on the preliminary screening maintenance plan set, and obtain a secondary screening maintenance plan set; Determine the cosine similarity value of each post-secondary screening maintenance plan in the post-secondary screening maintenance plan set based on the cosine similarity function, perform three screenings on the post-secondary screening maintenance plan set, and obtain a post-three-times screening maintenance plan set; The weighted sum of the remaining life of the substation equipment after maintenance and the inverse of the maintenance cost stored in the database corresponding to each three-times screening maintenance plan in the three-times screening maintenance plan set is performed to obtain the maintenance plan determination coefficient; The maximum value among the determination coefficients of the maintenance schemes is determined, and the maintenance scheme after three screenings corresponding to the maintenance scheme determination coefficient corresponding to the maximum value is determined as the substation maintenance scheme.
10. A high-voltage prefabricated substation intelligent commissioning data management system, used to implement the method according to any one of claims 1 to 9, characterized in that: It includes a data acquisition frequency determination module, a data storage module, a substation operating condition determination module and a substation maintenance plan determination module, wherein: The data collection frequency determination module is used to determine the data collection frequency based on the frequency multi-layer analysis method, and to collect multi-source heterogeneous data based on the determined data collection frequency. The multi-source heterogeneous data includes equipment operation parameters, environmental parameters, equipment assets and document data; The data storage module is used to pre-process the collected multi-source heterogeneous data, transmit it to the remote management center and store it; The substation operating condition determination module is used to perform multi-source data analysis based on the pre-processed multi-source heterogeneous data stored in the remote management center to determine the high-voltage prefabricated substation operating conditions, including equipment failure types, equipment failure development trends, and equipment failure causes; The substation maintenance plan determination module is used to determine the substation maintenance plan based on the working conditions of the high-voltage prefabricated substation.
Citation Information
Patent Citations
Demand-oriented intelligent acquisition optimization method for multi-source heterogeneous data of power transmission and transformation equipment
CN116451136A
Intelligent operation and maintenance system and method for digital twin substation
CN118172040A
Substation main equipment self-checking system based on multi-source heterogeneous data fusion technology
CN119312250A