A power grid time synchronization management method based on big data analysis
Through big data analysis methods, power grid information is collected and processed, feature extraction and model training is performed, and real-time time synchronization adjustment scheme is obtained, which solves the complexity of power grid time synchronization management and realizes efficient and automated time synchronization management.
Patent Information
- Application Number
- CN202510157064.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2045-02-13
AI Technical Summary
The existing technology is difficult to effectively manage power grid time synchronization, especially in complex environments that process multi-source and multi-type data, which leads to increased data analysis difficulty, difficulty in formulating adjustment strategies, and inability to adapt to complex power grid operation environments.
The power grid time synchronization management method is adopted for big data analysis. By collecting comprehensive power grid information, standardized processing and feature extraction, data dimensionality reduction and correlation analysis are carried out, and the time synchronization scheme design model is trained using the GBT model based on CatBoost, real-time adjustment scheme is obtained, and visual interaction is carried out.
It realizes the intelligence of data processing and feature extraction, improves the automation of time synchronization prediction and solution design, improves the visualization effect of time synchronization management, and enhances the accuracy, reliability and anti-interference ability of power grid time synchronization.
Smart Images

Figure CN119624405B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power grid time synchronization management, and more specifically, to a power grid time synchronization management method for big data analysis. Background Art
[0002] Grid time synchronization management is the basis for ensuring the safe, stable, reliable and economical operation of the power system. It involves all aspects of the power grid, from data collection, fault analysis, stability control, operation scheduling, power trading to smart grid applications, all of which require accurate time synchronization. Lack of time synchronization will lead to data confusion, wrong decisions, and equipment malfunctions, thus affecting the normal operation of the power grid and even causing safety accidents. Therefore, the synchronization management of the power grid time is of irreplaceable importance.
[0003] In the process of power grid time synchronization management, it is necessary to process various types of data from multiple sources (Beidou module, ground link, power grid system, etc.), including numerical data, status data, time series data, etc. The data types are inconsistent and the structure is complex. In addition, the power grid operation generates massive data, including real-time operation status data, historical data, fault data, etc. All of these will increase the difficulty of data analysis and mining in the process of power grid time synchronization management, making it difficult to specifically analyze the main factors affecting the power grid time synchronization performance, and there may be complex nonlinear relationships between the factors, which are difficult to analyze and predict. This makes it difficult to formulate adjustment strategies, unable to adapt to the complex power grid operation environment, and difficult to ensure real-time time synchronization.
[0004] In view of this, the present invention proposes a power grid time synchronization management method based on big data analysis to solve the above problems. Summary of the invention
[0005] In order to overcome the above-mentioned defects of the prior art and to achieve the above-mentioned purpose, the present invention provides the following technical solutions: A power grid time synchronization management method for big data analysis, comprising: step S1: collecting comprehensive power grid information at the current moment, standardizing the comprehensive power grid information, and extracting feature data according to the processed comprehensive power grid information to obtain a comprehensive feature set;
[0006] Step S2: Performing data dimensionality reduction processing by analyzing the correlation of each data category in the comprehensive feature set to obtain a low-dimensional data set;
[0007] Step S3: Use the training set to train the time synchronization solution design model, import the real-time low-dimensional data set into the trained time synchronization solution design model, and obtain a real-time adjustment solution;
[0008] Step S4: Visualize and interact with the adjustment plan.
[0009] Preferably, the comprehensive power grid information includes clock status data, Beidou module data, ground link signal data, hot standby signal data, environment and external interference data and power grid operation data;
[0010] The clock status data includes power status, frequency taming status, alarm status, main clock signal status and backup clock signal status;
[0011] Beidou module is a commonly used time reference signal source, including synchronization status module status, antenna status, number of satellites, channel difference, longitude, latitude and altitude;
[0012] Ground link signals include synchronization status, time deviation, signal quality, channel difference and electrical type;
[0013] Hot standby signals include synchronization status, time deviation, signal quality, channel difference and electrical type;
[0014] Environmental and external interference data include electromagnetic interference data, physical vibrations, environmental data, deception and jamming signals;
[0015] Grid operation data includes load data, fault events, system frequency data and power area synchronization data.
[0016] Preferably, the method for standardizing the integrated power grid information includes:
[0017] Extract discrete state data and data types from comprehensive power grid information and convert them into numerical data using enumeration encoding or one-hot encoding;
[0018] Convert units of the same data type to be consistent, use statistical methods or machine learning algorithms to detect outliers and missing values in the integrated power grid information, delete outliers to form missing values, and fill in the missing values with the mean or median;
[0019] Extract all timestamps in the comprehensive power grid information and convert them into a unified format;
[0020] Construct a blank structured table, with timestamps as rows and data categories as columns, fill the values of the data categories at the timestamps at each moment into the corresponding spaces, and convert the comprehensive power grid information into a structured table representation.
[0021] Preferably, the method of extracting feature data according to the processed comprehensive power grid information to obtain a comprehensive feature set includes:
[0022] Move forward based on the current moment Time length, get a cycle time , collect comprehensive power grid information at all times during the cycle time to form a comprehensive information set;
[0023] Set the window length to ,and , represents the adjustment coefficient, Indicates the time deviation of the current moment;
[0024] Use the window to slide on the comprehensive information set, rank transform the value of each data type in the window, and for each data type, calculate the Spearman rank correlation coefficient between the rank-transformed value and the time synchronization performance index. ;
[0025] in, The value range is between -1 and 1, with a value of 1 indicating a perfect positive correlation, a value of -1 indicating a perfect negative correlation, and a value of 0 indicating no correlation. Indicates The difference between the ranks of the two variables at a certain moment in time, Indicates the number of moments;
[0026] Set a correlation coefficient threshold ,reserve data type, excluding other data types;
[0027] For each retained data type, randomly shuffle it in the comprehensive power grid information, keep the data of other data types unchanged, form a disturbed data set, and use the disturbed data set of each data category to calculate the loss function value of the GBT model, which is recorded as the replacement performance value;
[0028] Calculate the difference between the replacement performance value and the initial performance value, repeatedly obtain and average the values, arrange the average values of each data category in descending order, select the first G_Y data categories as the main feature data, and collect the selected G_Y data categories to form a comprehensive feature set.
[0029] Preferably, the training method of the GBT model includes:
[0030] The GBT model is designed based on CatBoost, with the comprehensive information set as the input of the model and the time synchronization performance index as the output of the model;
[0031] Use cross validation or grid search to obtain the optimal hyperparameter combination, set MAE as the loss function, and stop the loss function when the value does not change.
[0032] Import the sample set into the GBT model and perform training to obtain the trained GBT model.
[0033] Preferably, the method for acquiring the sample set includes:
[0034] The sample set includes E_N samples, each of which includes a comprehensive information set and a corresponding time synchronization performance indicator;
[0035] Step D1: Collect A_v samples at A_v historical moments and randomly select K_s comprehensive information sets as cluster centers;
[0036] Step D2: Calculate the distance between each comprehensive information set and the cluster center, and assign the comprehensive information set to the cluster center with the nearest distance to form K_o groups;
[0037] Step D3: Calculate the absolute difference between the cluster centers. If the absolute difference is less than the limit threshold, randomly select a cluster center of a group to remove it, and release the comprehensive information set in the group corresponding to the cluster center, and redistribute it to the groups of other cluster centers to form K_f groups;
[0038] Step D4: Update the cluster center using the average value of the comprehensive information set in each group;
[0039] Step D5: Repeat steps D2 to D4 until all cluster centers no longer change, obtain K_h groups, and then divide the comprehensive information set into K_h groups;
[0040] Step D6: Connect the time synchronization performance index with the corresponding comprehensive information set, and then divide the A_v samples into K_h sample groups, and import the sample groups into the GBT model for training.
[0041] Preferably, the method of performing data dimensionality reduction processing on the correlation analysis of each data category in the comprehensive feature set to obtain a low-dimensional data set includes:
[0042] Extract each data category in the comprehensive feature set, obtain the standard data value corresponding to the data category under time synchronization, traverse and select the standard data value to replace the corresponding data category in the comprehensive feature set, and form a reorganized data set;
[0043] Import the reorganized data set into the trained GBT model, obtain the output time synchronization performance index, and traverse and calculate the absolute difference between the two time synchronization performance indicators;
[0044] If the absolute difference is greater than the correlation threshold, it is determined that the data categories replaced by the standard data values in the reorganized data set are highly correlated, otherwise, they are determined to be uncorrelated;
[0045] Treat the related data categories as a cluster and set the representative index of each cluster ;
[0046] in, Indicates the cluster The value of the data category, represents the number of data categories in the cluster, Indicates the cluster The adjustment coefficient corresponding to each data category is: Indicates applicable parameters, is the adjustment item;
[0047] Collect individual data categories that are not related to other data categories, and aggregate the representative index of each cluster to form a pending data set. Re-import the pending data set into the trained GBT model to obtain the time synchronization performance index output at this time, and use the least squares method to adjust , and , until the time synchronization performance index of the pending data set output is the same as that of the comprehensive feature set, and then the optimal , and value;
[0048] Extract the values from the comprehensive feature set and use the optimal , and The representative index of each cluster is calculated, and all representative indices and numerical values corresponding to single unrelated data categories are collected to form a low-dimensional data set.
[0049] Preferably, the construction and training method of the time synchronization scheme design model includes:
[0050] The framework of the time synchronization scheme design model is based on neural network learning technology, including input layer, hidden layer and output layer. The input of the input layer is set as a low-dimensional data set, and the output layer sets three task heads in parallel, namely the time synchronization prediction head, the cause prediction head and the scheme design prediction head.
[0051] For each task head, a linear activation function is used and the loss function is set to cross entropy loss. The time synchronization prediction head is used to predict whether the time will be synchronized at the next moment. The cause prediction head is used to predict the cause of time asynchrony. The solution design prediction head is used to output the adjustment solution for time synchronization.
[0052] The total loss function of the time synchronization scheme design model is set as the accumulation of the loss functions of each task head, and the constraints are minimization of adjustment time and minimization of cost utilization;
[0053] The training set includes R_G training groups, each of which includes a low-dimensional data set and corresponding time synchronization labels, cause labels, and solution design labels;
[0054] The training set is input into the time synchronization scheme design model in sequence. An adaptive learning rate strategy is used. The Adam or SGD optimizer is adopted to output the loss function of each task head in sequence. The training is stopped when the total loss function value no longer changes or the maximum number of iterations is reached to obtain the trained time synchronization scheme design model.
[0055] Preferably, the visual interaction includes displaying the adjustment scheme on the end surface in the form of a chart, video or text, and the user operates through the end surface.
[0056] Preferably, the end surface refers to a terminal display interface of a computer or a display device.
[0057] The technical effects and advantages of the power grid time synchronization management method based on big data analysis of the present invention are as follows:
[0058] 1. Intelligent data processing and feature extraction:
[0059] Through standardization, data from different sources and different types are integrated into a unified format data set, eliminating the differences between heterogeneous data and facilitating subsequent analysis. Based on big data technology, it can efficiently process massive data in power grid operation. Through sliding window analysis, Spearman rank correlation coefficient, and permutation importance, it automatically selects the feature type that has the greatest impact on time synchronization performance, avoiding human subjectivity and improving the accuracy of feature selection, thereby reducing the dimension of the data and the complexity of the model, and retaining key information. Through learnable parameters and clustering ideas, the feature dimension is further compressed, and the relationship between features is considered at the same time.
[0060] 2. Automation of time synchronization prediction and solution design:
[0061] By training the GBT model based on CatBoost, the time synchronization status of the next moment can be predicted, and early warning of time synchronization can be achieved. By establishing a multi-task time synchronization solution design model, the time synchronization status, fault causes and adjustment solutions can be predicted at the same time, avoiding the shortcomings of a single task model that cannot take into account the overall situation, and enabling the model to analyze the power grid operation status more intelligently and give corresponding adjustment solutions. The model can automatically give corresponding adjustment solutions based on the current power grid status, reduce manual intervention, improve processing efficiency, and make the adjustment strategy more accurate. By adding the constraints of minimizing adjustment time and cost through the loss function, the generated adjustment solution is more in line with actual needs and ensures the effectiveness of the solution.
[0062] 3. Visualization of time synchronization management:
[0063] Through visual interaction, users can globally monitor the time synchronization status of the entire power grid. Complex power grid data and time synchronization information are displayed in an intuitive graphical way, which is convenient for users to understand and operate. An interactive interface is provided, allowing users to view detailed information, manually adjust parameters, and perform related operations, which improves the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 It is a structural schematic diagram of a power grid time synchronization management system for big data analysis of the present invention;
[0065] Figure 2 A schematic diagram of the steps of a power grid time synchronization management method for big data analysis of the present invention. DETAILED DESCRIPTION
[0066] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0067] Example 1
[0068] See also Figure 1 and Figure 2 As shown, the present embodiment provides a power grid time synchronization management method for big data analysis, including:
[0069] Grid time synchronization is the core technology for the stable operation of power systems. Especially in modern smart grids, time accuracy directly affects the coordinated operation of power generation, transmission, distribution and power consumption equipment. Based on the complex requirements of grid time synchronization and the above content, we designed a grid time synchronization management method based on big data analysis, aiming to improve the accuracy, reliability and anti-interference ability of grid time synchronization through multi-source data fusion, real-time monitoring, predictive analysis and intelligent optimization.
[0070] Step S1: collecting comprehensive power grid information at the current moment, standardizing the comprehensive power grid information, and extracting feature data according to the processed comprehensive power grid information to obtain a comprehensive feature set;
[0071] Data collection and feature processing module: used to collect comprehensive power grid information, standardize and extract features from the comprehensive power grid information, and obtain a comprehensive feature set;
[0072] Data collection is the basis of power grid time synchronization management methods. In order to ensure the accuracy, reliability and stability of the time synchronization system, it is necessary to clearly define what data to collect and how to collect it, including data categories, data items and collection methods.
[0073] Comprehensive power grid information includes clock status data, Beidou module data, ground link signal data, hot standby signal data, environment and external interference data, and power grid operation data;
[0074] Clock status data includes power status (whether it is working properly, for example, normal / abnormal), frequency taming status (whether the master / slave clock is tamed by the external reference signal, including normal / abnormal), alarm status (including specific information such as hardware alarms and synchronization alarms of equipment operation), master clock signal status and backup clock signal status;
[0075] The master clock signal status specifically includes: signal type (such as Beidou, ground link, etc.), synchronization status (synchronized / unsynchronized), time deviation (offset from reference time, μs or ms), channel difference (time difference between signal source and local time, μs), quality bit (time quality level of signal);
[0076] The status of the backup clock signal specifically includes the synchronization status (synchronous / unsynchronized), time deviation (μs) and electrical type (such as TTL, IRIG-B, RS422, etc.).
[0077] The Beidou module is a commonly used time reference signal source, including synchronization status (whether it is synchronized with the Beidou satellite), antenna status (whether the antenna is normal, normal / disconnected / short circuit), module status (Beidou module hardware operation status, normal / faulty), number of satellites (the number of Beidou satellites currently received), channel difference (the time difference between the Beidou signal and the local clock, μs), longitude, latitude and altitude; longitude, latitude and altitude are the location information returned by the Beidou module.
[0078] The ground link signal is another commonly used time reference signal source, including synchronization status (whether it is synchronized with the ground signal), time deviation (the time deviation between the ground signal and the local clock, μs), signal quality (i.e., quality bit, indicating the quality level of the ground signal), channel difference (the time difference between the ground signal and the local clock) and electrical type (signal interface type, such as TTL, RS422, etc.);
[0079] The hot standby signal is an important guarantee for system redundancy, including synchronization status (whether the hot standby signal is synchronized), time deviation (time deviation between the hot standby signal and the main clock, μs), signal quality (also called quality bit, the quality level of the hot standby signal), channel difference (time difference between the hot standby signal and the main clock) and electrical type (signal interface type, such as TTL, RS422, etc.);
[0080] The environment and interference have a great impact on the time synchronization system. The environment and external interference data include electromagnetic interference data, physical vibration, environmental data, deception and interference signals;
[0081] Electromagnetic interference data mainly include electrostatic discharge (intensity / kV), radio frequency electromagnetic field interference (frequency range, field strength / V / m), power frequency magnetic field interference (field strength / A / m), surge voltage (common mode / differential mode, kV), and fast transient pulse group interference (frequency / kHz, amplitude / kV).
[0082] Physical vibration mainly includes vibration frequency (Hz) and vibration intensity (acceleration / g).
[0083] Environmental data mainly include temperature (℃) and humidity (%RH).
[0084] Load changes and fault events during grid operation may affect time synchronization performance. Grid operation data includes load data, fault events, system frequency data and power area synchronization data.
[0085] Load data includes operating parameters such as voltage, current, power, and load fluctuations (fluctuation rate / %) of each substation.
[0086] Fault events include fault occurrence time (accurate to ms), fault type (such as short circuit, overload, ground fault, etc.), fault location (substation name, line number, etc.), and fault waveform data (such as fault current, voltage waveform).
[0087] System frequency data includes real-time frequency deviation (Hz), system frequency fluctuation range (short-term and long-term trends), and relationship data with time synchronization performance.
[0088] The power area synchronization data includes the time deviation between different power areas (synchronization difference between areas) and the quality information of the regional synchronization signal.
[0089] In order to collect data efficiently and comprehensively, it is necessary to adopt appropriate collection methods and technical methods according to different data categories:
[0090] Clock status data collection method: Use the communication interface (such as serial port, network port) provided by the device to collect equipment operation status data; use standard protocols (such as SNMP, Modbus) to periodically poll to obtain power status, frequency taming status, alarm status and other information; the master / slave clock device actively reports alarms, status changes and other information to the monitoring system through the network interface.
[0091] Beidou module data collection method: read the Beidou module's synchronization status, antenna status, number of satellites, time deviation and other information through the serial port or network interface;
[0092] Ground link signal data collection method: obtain the synchronization status, signal quality bit and electrical type of the ground link through the communication interface provided by the equipment.
[0093] Hot standby signal data collection method: real-time monitoring of the time deviation between the main clock and the hot standby signal to evaluate the reliability of the hot standby state.
[0094] Environmental and external interference data collection method: Deploy sensors in key equipment areas (such as near the main clock and Beidou module antenna) to collect temperature, humidity, vibration, electromagnetic interference and other data in real time. When electromagnetic interference (such as electrostatic discharge, radio frequency interference, surge, etc.) occurs, use interference detection equipment to record key information such as the intensity, frequency, and duration of the interference event. When collecting fast transient pulse group interference, use high-frequency sampling equipment to record the interference waveform and extract pulse amplitude, frequency and other characteristics. For vibration and impact, deploy vibration sensors to record vibration frequency and intensity (such as acceleration, amplitude).
[0095] Sources of power grid operation data: power grid SCADA (supervisory control and data acquisition system), EMS (energy management system), PMU (synchronous phasor measurement unit), and protection devices. Obtain real-time operation data of substations and distribution stations from the SCADA system, including voltage, current, power load, etc. Use the time synchronization data (such as phasors, timestamps) output by PMU devices to analyze the time synchronization accuracy of different locations. Data collection methods include PDC (phasor data concentrator) interface or standard data protocols of PMU devices (such as IEEE C37.118). Obtain fault occurrence time, type, waveform data, etc. from protection devices or fault recording devices.
[0096] In order to efficiently collect, transmit and store the above data, the following architecture is designed:
[0097] Interface type: RS485, RS232, TCP / IP, CAN bus, etc.
[0098] Communication protocols: SNMP (network management protocol), Modbus (industrial communication protocol), NMEA (Beidou module output protocol), IEEE C37.118 (PMU protocol).
[0099] Data transmission mode: analog signal (4-20mA) or digital signal (such as RS485).
[0100] Data is stored in the cloud or on a centralized big data platform using a distributed storage system such as Hadoop HDFS.
[0101] Methods for standardizing comprehensive power grid information include:
[0102] Extract discrete state data and data types from comprehensive power grid information and convert them into numerical data using enumeration encoding or one-hot encoding;
[0103] Discrete state data refers to data with finite and discontinuous values, for example:
[0104] Power supply status: normal, abnormal;
[0105] Frequency taming status: locked, unlocked, taming;
[0106] Alarm status: no alarm, alarm;
[0107] Beidou / ground wired / hot standby signal synchronization status: synchronized, out of sync, searching;
[0108] Module status: normal, abnormal;
[0109] Quality level: Q1, Q2, Q3... (different quality levels);
[0110] Electrical type: RS422, RS485, optical fiber, BNC;
[0111] Furthermore, power status, frequency taming status, alarm status, etc. are regarded as data types.
[0112] Enumeration encoding assigns a unique integer value to each discrete state. For example:
[0113] Power status: Normal (1), Abnormal (0)
[0114] Frequency taming status: locked (2), unlocked (0), taming (1)
[0115] One-hot encoding creates a new binary feature (0 or 1) for each discrete state. For example:
[0116] Frequency taming status: locked (1,0,0), unlocked (0,1,0), taming (0,0,1).
[0117] Convert units of the same data type to be consistent, such as converting all temperature data to Celsius or Fahrenheit, using statistical methods (such as box plots) or machine learning algorithms (such as isolation forests) to detect outliers and missing values in the integrated power grid information, remove outliers to form missing values, and use the mean or median to fill in the missing values;
[0118] Extract all timestamps in the comprehensive power grid information and convert them into a unified format; convert all timestamp data into a unified time zone, such as UTC or Beijing time.
[0119] Construct a blank structured table, with timestamps as rows and data categories as columns, fill the values of the data categories at the timestamps at each moment into the corresponding spaces, and convert the comprehensive power grid information into a structured table representation.
[0120] For example, the structure table of comprehensive power grid information can be:
[0121] Timestamp node Power Status Frequency Taming Status Beidou 1 synchronization status …… 1698391 001 1 2 1 …… 1698392 002 1 0 1 …… …… …… …… …… …… ……
[0122] The method of extracting feature data according to the processed comprehensive power grid information and obtaining a comprehensive feature set includes:
[0123] Move forward based on the current moment Time length, get a cycle time , collect comprehensive power grid information at all times during the cycle time to form a comprehensive information set;
[0124] Set the window length to ,and , represents the adjustment coefficient, Indicates the time deviation of the current moment. The time deviation can be the absolute difference between the actual time displayed by the power grid and the standard accurate time;
[0125] Use a window to slide on the comprehensive information set and perform rank transformation on the values of each data type in the window. Rank transformation replaces the original values with their rankings in the window (for example, the smallest value is ranked first, the second smallest is ranked second, and so on). A similar rank transformation is also performed on the time synchronization performance index, which is obtained through the GBT model. For each data type, the Spearman rank correlation coefficient between the rank-transformed value and the time synchronization performance index is calculated. ;
[0126] in, The value range is between -1 and 1, with a value of 1 indicating a perfect positive correlation, a value of -1 indicating a perfect negative correlation, and a value of 0 indicating no correlation. The larger the absolute value of is, the stronger the correlation is. Indicates The difference between the ranks of the two variables at a certain moment in time, Indicates the number of moments;
[0127] Set a correlation coefficient threshold ,reserve The data type of the data is excluded, and other data types are excluded; this step is mainly to exclude the feature types that have almost no monotonic relationship with the time synchronization performance indicators. The threshold can be determined by using an absolute value threshold or a significance level based on a statistical test, such as by using a p-value analysis to determine the threshold.
[0128] For each retained data type, randomly shuffle it in the comprehensive power grid information, keep the data of other data types unchanged, form a disturbed data set, and use the disturbed data set of each data category to calculate the loss function value of the GBT model, which is recorded as the replacement performance value;
[0129] Calculate the difference between the replacement performance value and the initial performance value. The larger the difference, the greater the performance degradation, which means that the feature type is more important. Repeat the acquisition and take the average value to reduce the error caused by randomness. Arrange the average value of each data category in descending order, select the first G_Y data categories as the main feature data, and sort the feature types according to the replacement importance value. The feature data at the front of the sequence has a greater impact on the model performance. Collect the selected G_Y data categories to form a comprehensive feature set.
[0130] The Spearman rank correlation coefficient is used to preliminarily screen out important feature types that have a monotonic relationship with the time synchronization performance. Then the GBT model is used to learn the complex nonlinear relationship between these feature types and the time synchronization performance. Finally, the permutation importance is used to evaluate the importance of each feature type to the prediction performance of the GBT model, so as to finally select the optimal feature type combination.
[0131] The training methods of GBT model include:
[0132] The GBT model is designed based on CatBoost, with the comprehensive information set as the input of the model and the time synchronization performance index as the output of the model;
[0133] Use cross validation or grid search to obtain the optimal hyperparameter combination, set MAE as the loss function, and stop the loss function when the value does not change.
[0134] Import the sample set into the GBT model and perform training to obtain the trained GBT model.
[0135] The methods for obtaining the sample set include:
[0136] The sample set includes E_N samples, each of which includes a comprehensive information set and a corresponding time synchronization performance indicator;
[0137] Step D1: Collect A_v samples at A_v historical moments and randomly select K_s comprehensive information sets as cluster centers;
[0138] Step D2: Calculate the distance between each comprehensive information set and the cluster center, and assign the comprehensive information set to the cluster center with the nearest distance to form K_o groups;
[0139] Step D3: Calculate the absolute difference between the cluster centers. If the absolute difference is less than the limit threshold, randomly select a cluster center of a group to remove it, and release the comprehensive information set in the group corresponding to the cluster center, and redistribute it to the groups of other cluster centers to form K_f groups;
[0140] The limit threshold can be set based on experimental data analysis or historical data analysis and experience.
[0141] Step D4: Update the cluster center using the average value of the comprehensive information set in each group;
[0142] Step D5: Repeat steps D2 to D4 until all cluster centers no longer change, obtain K_h groups, and then divide the comprehensive information set into K_h groups.
[0143] Step D6: Connect the time synchronization performance index with the corresponding comprehensive information set, and then divide the A_v samples into K_h sample groups, and import the sample groups into the GBT model for training.
[0144] Step S2: Performing data dimensionality reduction processing by analyzing the correlation of each data category in the comprehensive feature set to obtain a low-dimensional data set;
[0145] Correlation dimensionality reduction: used to perform correlation analysis on each data category in the comprehensive feature set to obtain a low-dimensional data set;
[0146] The correlation analysis of each data category in the comprehensive feature set is used to perform data dimensionality reduction. The method of obtaining a low-dimensional data set includes:
[0147] Extract each data category in the comprehensive feature set, obtain the standard data value corresponding to the data category under time synchronization, traverse and select the standard data value to replace the corresponding data category in the comprehensive feature set, and form a reorganized data set;
[0148] Import the reorganized data set into the trained GBT model, obtain the output time synchronization performance index, and traverse and calculate the absolute difference between the two time synchronization performance indicators;
[0149] If the absolute difference is greater than the correlation threshold, it is determined that the data categories replaced by the standard data values in the reorganized data set are highly correlated, otherwise, they are determined to be uncorrelated;
[0150] The correlation threshold can be pre-set as a very small positive number to determine whether the synchronization performance indicator has a significant change. The more significant the change, the higher the degree of correlation between the data categories.
[0151] Treat the interrelated data categories as a cluster. The number of data categories in each cluster should be greater than 1. Set the representative index of each cluster. ;
[0152] in, Indicates the cluster The value of the data category, represents the number of data categories in the cluster, Indicates the cluster The adjustment coefficient corresponding to each data category is: Indicates applicable parameters, is the adjustment item;
[0153] Collect individual data categories that are not related to other data categories, and aggregate the representative index of each cluster to form a pending data set. Re-import the pending data set into the trained GBT model to obtain the time synchronization performance index output at this time, and use the least squares method to adjust , and , until the time synchronization performance index of the pending data set output is the same as that of the comprehensive feature set, and then the optimal , and value;
[0154] Extract the values from the comprehensive feature set and use the optimal , and The representative index of each cluster is calculated, and all representative indices and numerical values corresponding to single unrelated data categories are collected to form a low-dimensional data set.
[0155] For example, the comprehensive feature set includes node=1, power state=1 and frequency taming state=2, and the corresponding standard data values are 1, 0, 0 respectively; then the reorganized data set after reorganization is 1,1,2, 1,0,2, 1,0,0, 1,1,0.
[0156] Step S3: Use the training set to train the time synchronization solution design model, import the real-time low-dimensional data set into the trained time synchronization solution design model, and obtain a real-time adjustment solution;
[0157] Solution design module: used to design models based on real-time low-dimensional data sets and time synchronization solutions to obtain adjustment solutions;
[0158] The construction and training methods of the time synchronization solution design model include:
[0159] The framework of the time synchronization scheme design model is based on neural network learning technology, including input layer, hidden layer and output layer. The input of the input layer is set as a low-dimensional data set, and the output layer sets three task heads in parallel, namely the time synchronization prediction head, the cause prediction head and the scheme design prediction head.
[0160] For each task head, a linear activation function is used and the loss function is set to cross entropy loss. The time synchronization prediction head is used to predict whether the time will be synchronized at the next moment. The cause prediction head is used to predict the cause of time asynchrony. The solution design prediction head is used to output the adjustment solution for time synchronization.
[0161] The total loss function of the time synchronization scheme design model is set as the accumulation of the loss functions of each task head, and the constraints are minimization of adjustment time and minimization of cost utilization;
[0162] The training set includes R_G training groups, each of which includes a low-dimensional data set and corresponding time synchronization labels, cause labels, and solution design labels;
[0163] The training set is input into the time synchronization scheme design model in sequence. An adaptive learning rate strategy is used. The Adam or SGD optimizer is adopted to output the loss function of each task head in sequence. The training is stopped when the total loss function value no longer changes or the maximum number of iterations is reached to obtain the trained time synchronization scheme design model.
[0164] Step S4: Visualize and interact with the adjustment plan.
[0165] Visual interaction includes displaying the adjustment plan on the end face in the form of charts, videos or texts, and the user operates through the end face.
[0166] The end face refers to the terminal display interface of a computer or display device.
[0167] Visual interaction is achieved through front-end technologies (e.g., JavaScript libraries, D3.js, Canvas, WebGL, SVG, etc., for creating interactive charts and graphics) and back-end technologies (for data processing, storage, and transmission, such as Python, Java, Node.js, etc.).
[0168] Example 2
[0169] See also Figure 1 As shown, the part not described in detail in this embodiment is described in Example 1, which provides a power grid time synchronization management system for big data analysis, including:
[0170] Data collection and feature processing module: used to collect comprehensive power grid information, standardize and extract features from the comprehensive power grid information, and obtain a comprehensive feature set;
[0171] Correlation dimensionality reduction module: used to perform correlation analysis on each data category in the comprehensive feature set to obtain a low-dimensional data set;
[0172] Solution design module: used to design models based on real-time low-dimensional data sets and time synchronization solutions to obtain adjustment solutions;
[0173] Visualization module: used for visual interaction of adjustment schemes.
[0174] Example 3
[0175] This embodiment discloses an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the operation mode of the power grid time synchronization management method for big data analysis provided above is implemented.
[0176] Since the electronic device introduced in this embodiment is an electronic device used to implement a power grid time synchronization management method for big data analysis in the embodiment of this application, based on the power grid time synchronization management method for big data analysis introduced in the embodiment of this application, the technical personnel of this field can understand the specific implementation of the electronic device of this embodiment and its various variations, so how the electronic device implements the method in the embodiment of this application is not introduced in detail here. As long as the technical personnel of this field implement the electronic device used by the power grid time synchronization management method for big data analysis in the embodiment of this application, it belongs to the scope of protection of this application.
[0177] The above formulas are all dimensionless and numerical calculations. The formula is a formula for the most recent real situation obtained by collecting a large amount of data and performing software simulation. The preset parameters and thresholds in the formula are set by technicians in this field according to actual conditions.
[0178] The above is only a preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions under the concept of the present invention belong to the protection scope of the present invention. It should be pointed out that for ordinary technical users in this technical field, some improvements and modifications without departing from the principle of the present invention should also be regarded as the protection scope of the present invention.
Claims
1. A power grid time synchronization management method based on big data analysis, characterized in that: include: Step S1: collecting comprehensive power grid information at the current moment, standardizing the comprehensive power grid information, and extracting feature data according to the processed comprehensive power grid information to obtain a comprehensive feature set; The method of extracting feature data according to the processed comprehensive power grid information to obtain a comprehensive feature set includes: Taking the current time as the base point, move forward W-1 time lengths to obtain a cycle time W, collect the comprehensive power grid information at all times within the cycle time, and form a comprehensive information set; Set the window length to L, and α represents the adjustment coefficient, ΔSj represents the time deviation at the current moment; Use the window to slide on the comprehensive information set, rank-transform the values of each data type in the window, and for each data type, calculate the Spearman rank correlation coefficient between the rank-transformed value and the time synchronization performance index. The value of ρ ranges from -1 to 1, with a value of 1 indicating a perfect positive correlation, a value of -1 indicating a perfect negative correlation, and a value of 0 indicating no correlation. i represents the difference between the ranks of two variables at the i-th moment, and n represents the number of moments; Set a correlation coefficient threshold ρ_yu, retain the data type with |ρ|≥ρ_yu, and exclude other data types; For each retained data type, randomly shuffle it in the comprehensive power grid information, keep the data of other data types unchanged, form a disturbed data set, and use the disturbed data set of each data category to calculate the loss function value of the GBT model, which is recorded as the replacement performance value; Calculate the difference between the replacement performance value and the initial performance value, repeatedly obtain and average, sort the average values of each data category in descending order, select the first G_Y data categories as feature data, and collect the selected G_Y data categories to form a comprehensive feature set; Step S2: Performing data dimensionality reduction processing by analyzing the correlation of each data category in the comprehensive feature set to obtain a low-dimensional data set; Step S3: Use the training set to train the time synchronization solution design model, import the real-time low-dimensional data set into the trained time synchronization solution design model, and obtain a real-time adjustment solution; The construction and training method of the time synchronization scheme design model includes: The framework of the time synchronization scheme design model is based on neural network learning technology, including input layer, hidden layer and output layer. The input of the input layer is set as a low-dimensional data set, and the output layer sets three task heads in parallel, namely the time synchronization prediction head, the cause prediction head and the scheme design prediction head. For each task head, a linear activation function is used and the loss function is set to cross entropy loss. The time synchronization prediction head is used to predict whether the time will be synchronized at the next moment. The cause prediction head is used to predict the cause of time asynchrony. The solution design prediction head is used to output the adjustment solution for time synchronization. The total loss function of the time synchronization scheme design model is set as the accumulation of the loss functions of each task head, and the constraints are minimization of adjustment time and minimization of cost utilization; The training set includes R_G training groups, each of which includes a low-dimensional data set and corresponding time synchronization labels, cause labels, and solution design labels; Input the training set into the time synchronization scheme design model in sequence, use the adaptive learning rate strategy, adopt the Adam or SGD optimizer, output the loss function of each task head in sequence, stop training when the total loss function value no longer changes or reaches the maximum number of iterations, and obtain the trained time synchronization scheme design model; Step S4: Visualize and interact with the adjustment plan.
2. The power grid time synchronization management method for big data analysis according to claim 1 is characterized in that: The comprehensive power grid information includes clock status data, Beidou module data, ground link signal data, hot standby signal data, environment and external interference data and power grid operation data; The clock status data includes power status, frequency taming status, alarm status, main clock signal status and backup clock signal status; Beidou module is a commonly used time reference signal source, including synchronization status, module status, antenna status, number of satellites, channel difference, longitude, latitude and altitude; Ground link signals include synchronization status, time deviation, signal quality, channel difference and electrical type; Hot standby signals include synchronization status, time deviation, signal quality, channel difference and electrical type; Environmental and external interference data include electromagnetic interference data, physical vibrations, environmental data, deception and jamming signals; Grid operation data includes load data, fault events, system frequency data and power area synchronization data.
3. The power grid time synchronization management method for big data analysis according to claim 2 is characterized in that: The method for standardizing the integrated power grid information comprises: Extract discrete state data and data types from comprehensive power grid information and convert them into numerical data using enumeration encoding or one-hot encoding; Convert units of the same data type to be consistent, use statistical methods or machine learning algorithms to detect outliers and missing values in the integrated power grid information, delete outliers to form missing values, and fill in the missing values with the mean or median; Extract all timestamps in the comprehensive power grid information and convert them into a unified format; Construct a blank structured table, with timestamps as rows and data categories as columns, fill the values of the data categories at the timestamps at each moment into the corresponding spaces, and convert the comprehensive power grid information into a structured table representation.
4. The power grid time synchronization management method for big data analysis according to claim 3 is characterized in that: The training method of the GBT model includes: The GBT model is designed based on CatBoost, with the comprehensive information set as the input of the model and the time synchronization performance index as the output of the model; Use cross validation or grid search to obtain the optimal hyperparameter combination, set MAE as the loss function, and stop the mechanism when the loss function value does not change; Import the sample set into the GBT model and perform training to obtain the trained GBT model.
5. The power grid time synchronization management method for big data analysis according to claim 4 is characterized in that: The method for obtaining the sample set includes: The sample set includes E_N samples, each of which includes a comprehensive information set and a corresponding time synchronization performance indicator; Step D1: Collect A_v samples at A_v historical moments and randomly select K_s comprehensive information sets as cluster centers; Step D2: Calculate the distance between each comprehensive information set and the cluster center, and assign the comprehensive information set to the cluster center with the nearest distance to form K_o groups; Step D3: Calculate the absolute difference between the cluster centers. If the absolute difference is less than the limit threshold, randomly select a cluster center of a group to remove it, and release the comprehensive information set in the group corresponding to the cluster center, and redistribute it to the groups of other cluster centers to form K_f groups; Step D4: Update the cluster center using the average value of the comprehensive information set in each group; Step D5: Repeat steps D2 to D4 until all cluster centers no longer change, obtain K_h groups, and then divide the comprehensive information set into K_h groups; Step D6: Connect the time synchronization performance index with the corresponding comprehensive information set, and then divide the A_v samples into K_h sample groups, and import the sample groups into the GBT model for training.
6. The power grid time synchronization management method for big data analysis according to claim 5 is characterized in that: The method of performing data dimensionality reduction processing on the correlation analysis of each data category in the comprehensive feature set to obtain a low-dimensional data set includes: Extract each data category in the comprehensive feature set, obtain the standard data value corresponding to the data category under time synchronization, traverse and select the standard data value to replace the corresponding data category in the comprehensive feature set, and form a reorganized data set; Import the reorganized data set into the trained GBT model, obtain the output time synchronization performance index, and traverse and calculate the absolute difference between the two time synchronization performance indicators; If the absolute difference is greater than the correlation threshold, it is determined that the data categories replaced by the standard data values in the reorganized data set are highly correlated, otherwise, they are determined to be uncorrelated; Treat the related data categories as a cluster and set the representative index of each cluster Among them, Exc kj represents the value of the kjth data category in the cluster, KJ represents the number of data categories in the cluster, θ kj represents the adjustment coefficient corresponding to the kj-th data category in the cluster, δ represents the applicable parameter, and ε is the adjustment term; Collect individual data categories that are not related to other data categories, and aggregate the representative index of each cluster to form a pending data set. Re-import the pending data set into the trained GBT model to obtain the time synchronization performance index output at this time, and use the least squares method to adjust θ kj , ε and δ until the time synchronization performance index of the undetermined data set output is the same as that of the comprehensive feature set, and the optimal θ is obtained. kj , ε and δ values; Extract the values from the comprehensive feature set and use the optimal θ kj , ε, and δ values are used to calculate the representative index of each cluster, and all representative indices and numerical values corresponding to single unrelated data categories are collected to form a low-dimensional data set.
7. The power grid time synchronization management method for big data analysis according to claim 6 is characterized in that: The visual interaction includes displaying the adjustment plan on the end surface in the form of charts, videos or texts, and the user operates through the end surface.
8. The power grid time synchronization management method for big data analysis according to claim 7 is characterized in that: The end face refers to the terminal display interface of a computer or a display device.
Citation Information
Patent Citations
Steady state oscillograph for power system
CN102169158A
Signature identification for power system events
US20210088563A1