Data processing method and system for power grid enterprise, medium and product
Through the bidirectional verification and dynamic correction mechanism of high- and low-frequency data, the data distortion problem caused by independent processing of data acquisition equipment of power grid enterprises is solved, the accuracy and reliability of power grid operation data are achieved, and support for data quality management and system optimization is provided.
Patent Information
- Application Number
- CN202510710950.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-10-10
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In existing technologies, the independent processing of data acquisition equipment by power grid companies causes data interruption or distortion in the event of failure or interference, affecting the accurate assessment of the power grid's operating status. In particular, extreme weather or equipment failures are difficult to detect and correct in a timely manner.
A two-way verification and dynamic correction mechanism for high- and low-frequency data is adopted. Simulated data is generated through a feature mapping model. Combined with data deviation and confidence assessment, correction strategies are dynamically selected to ensure data accuracy and reliability.
It achieves intelligent coordination of high-frequency and low-frequency data, effectively identifies and handles data quality issues, improves the accuracy and reliability of power grid operation data, and provides a decision-making basis for data quality management and system optimization.
Smart Images

Figure CN120763151A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of electrical digital data processing, and in particular to a data processing method, system, medium and product for a power grid enterprise. Background Art
[0002] With the continuous advancement of smart grid construction, power grid companies' data collection systems are becoming increasingly sophisticated. Various automated devices continuously collect technical operational data, such as power quality indicators and equipment status indicators. This data is an important foundation for assessing grid operation and ensuring power supply quality. Especially in the context of large-scale distributed energy integration and diverse load characteristics, the reliability of grid operation data is crucial for ensuring safe and stable grid operation.
[0003] In related technologies, power grid companies primarily employ a hierarchical data collection and independent processing approach. In this approach, data collection equipment is divided into different tiers based on function, with each tier independently completing its data collection tasks. For example, PMUs (synchronized phasor measurement units) installed in substations collect phasor data hundreds of times per second, dispatch automation systems collect telemetry data every five minutes, and online power quality monitoring devices record harmonic data every ten minutes. The collected data is analyzed and stored by their respective processing units before being made available to the corresponding application systems.
[0004] However, in actual applications, it has been found that when a single data acquisition device fails or is subject to external interference, data at that level may be interrupted or distorted for extended periods of time. For example, when a PMU is subject to electromagnetic interference, it may continuously output abnormal data, without the system being able to detect and correct the problem in a timely manner. This situation is particularly prominent under abnormal operating conditions such as extreme weather or equipment failures, hindering the accurate assessment of the grid's operating status. Summary of the Invention
[0005] The present application provides a data processing method, system, medium and product for a power grid enterprise, which are used to improve the data accuracy of the power grid enterprise.
[0006] In a first aspect, the application provides a data processing method of a power grid enterprise, applied to a data processing system, the method comprising: obtaining power grid operation data, extracting associated operation data in the power grid operation data based on a preset association rule; dividing the associated operation data into high-frequency data with a collection frequency higher than a preset collection frequency, and low-frequency data with a collection frequency not higher than the preset collection frequency; collecting high-frequency data in a preset time period, generating simulation data of the low-frequency data based on the high-frequency data; calculating a data deviation value of the simulation data and the low-frequency data; when the data deviation value is higher than a preset deviation threshold, calculating a data confidence value of the low-frequency data; when the data confidence value is lower than a preset confidence threshold, correcting the low-frequency data according to a historical database and the simulation data; and when the data confidence value is not lower than the preset confidence threshold, correcting the high-frequency data according to the historical database and the low-frequency data.
[0007] In the above embodiment, the data processing system divides the power grid operation data into high-frequency and low-frequency data according to the collection frequency, generates a reference value of the low-frequency data by simulation using the high-frequency data, dynamically selects a correction strategy through comparative analysis and data confidence evaluation, corrects the low-frequency data using the simulation data when the confidence of the low-frequency data is low, and in turn corrects the high-frequency data when the confidence of the low-frequency data is high. Through the bidirectional verification and correction mechanism, the accuracy and reliability of the data are improved.
[0008] In combination with some embodiments of the first aspect, in some embodiments, the step of collecting high-frequency data in a preset time period and generating simulation data of the low-frequency data based on the high-frequency data specifically comprises: extracting paired samples of the high-frequency data and the low-frequency data from a historical database; constructing a feature mapping model based on the paired samples; collecting the high-frequency data in the preset time period and inputting the high-frequency data into the feature mapping model to obtain the simulation data.
[0009] In the above embodiment, the data processing system constructs a feature mapping model using paired samples in the historical database, establishes a mapping relationship from the high-frequency data to the low-frequency data, can fully learn the internal correlation characteristics between data of different frequencies, improves the accuracy and representativeness of the simulation data, and effectively reduces errors in the data processing process.
[0010] In combination with some embodiments of the first aspect, in some embodiments, the step of calculating a data confidence value of the low-frequency data when the data deviation value is higher than a preset deviation threshold specifically comprises: when the data deviation value is higher than the preset deviation threshold, extracting a historical operation trend and a device operation state corresponding to the low-frequency data; calculating a deviation degree of the low-frequency data based on the historical operation trend, and determining a data reliability coefficient based on the device operation state; and determining the data confidence value according to the deviation degree and the data reliability coefficient.
[0011] In the above embodiment, the data processing system performs confidence assessment on low-frequency data based on historical operating trends and equipment operating status, and obtains a data confidence value by calculating the degree of data deviation and determining the reliability coefficient. This can accurately identify abnormal data, avoid unnecessary corrections to normal data, and improve the accuracy and efficiency of data processing.
[0012] In combination with some embodiments of the first aspect, in some embodiments, when the data confidence value is lower than a preset confidence threshold, after the step of correcting the low-frequency data according to the historical database and simulation data, the method also includes: when the correction amplitude of the low-frequency data exceeds the preset correction range, obtaining business events associated with the low-frequency data within a preset time period; determining the cause of the data anomaly based on the business event, and generating a data quality analysis report.
[0013] In the above embodiment, when the data correction amplitude is abnormal, the data processing system analyzes the cause of the data anomaly in combination with relevant business events and generates a quality analysis report. This can not only detect data quality problems in a timely manner, but also trace the root cause of the problem, providing a decision-making basis for data quality management and system optimization.
[0014] In combination with some embodiments of the first aspect, in some embodiments, after the steps of determining the cause of data anomaly based on business events and generating a data quality analysis report, the method also includes: extracting historical anomaly data corresponding to the cause of data anomaly from a historical database; performing pattern matching on low-frequency data and historical anomaly data to determine the anomaly development trend and impact degree; and generating anomaly handling suggestions based on the anomaly development trend and impact degree.
[0015] In the above embodiment, the data processing system uses historical abnormal data for pattern matching, analyzes abnormal development trends and impact levels, and generates targeted processing suggestions. It can predict potential risks, provide preventive solutions, and improve the system's proactive prevention and control capabilities.
[0016] In combination with some embodiments of the first aspect, in some embodiments, the step of correcting the high-frequency data according to the historical database and the low-frequency data when the data confidence value is not lower than the preset confidence threshold specifically includes: when the data confidence value is not lower than the preset confidence threshold, comparing the data distribution characteristics of the high-frequency data with similar data in the historical database, and calculating the data fluctuation deviation; generating a trend correction coefficient for the high-frequency data based on the key indicators of power grid operation and the data fluctuation deviation in the low-frequency data; and correcting the high-frequency data based on the trend correction coefficient.
[0017] In the above embodiment, the data processing system calculates the trend correction coefficient of high-frequency data by comparing data distribution characteristics and analyzing fluctuation deviations, thereby maintaining the overall distribution characteristics of the data and avoiding the introduction of new deviations during the correction process.
[0018] In combination with some embodiments of the first aspect, in some embodiments, the step of correcting the high-frequency data based on the trend correction coefficient specifically includes: calculating the product of the trend correction coefficient and the high-frequency data to obtain a preliminary corrected data sequence; based on the collection time point of the low-frequency data, dividing the preliminary corrected data sequence into time segments, and calculating the data mean and standard deviation in each time period; determining the data correction interval based on the data mean and standard deviation, and eliminating abnormal fluctuation points in the preliminary corrected data sequence that exceed the data correction interval; performing linear interpolation processing on the data sequence after eliminating the abnormal fluctuation points to obtain a correction result.
[0019] In the above embodiment, the data processing system performs segmented statistical analysis on the data, determines the correction interval based on the mean and standard deviation, and performs linear interpolation processing after removing abnormal fluctuation points. This not only ensures the continuity of the data, but also avoids the interference of abnormal values, thereby improving the accuracy and reliability of data correction.
[0020] In a second aspect, an embodiment of the present application provides a data processing system, which includes: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to enable the data processing system to execute the method described in the first aspect and any possible implementation of the first aspect.
[0021] In a third aspect, an embodiment of the present application provides a computer program product comprising instructions, which, when the computer program product is run on a data processing system, enables the data processing system to execute the method described in the first aspect and any possible implementation of the first aspect.
[0022] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium comprising instructions, which, when executed on a data processing system, enables the data processing system to execute the method described in the first aspect and any possible implementation of the first aspect.
[0023] It is understandable that the data processing system provided in the second aspect, the computer program product provided in the third aspect, and the computer storage medium provided in the fourth aspect are all used to execute the methods provided in the embodiments of the present application. Therefore, the beneficial effects that can be achieved can be referenced to the beneficial effects of the corresponding methods and will not be repeated here.
[0024] One or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages: 1. Due to the adoption of a two-way verification and dynamic correction mechanism for high- and low-frequency data, by extracting related data from power grid operation data and classifying and processing it based on preset frequencies, combined with simulation data generation and confidence assessment, intelligent verification and correction of data are achieved. Therefore, the consistency and accuracy between data of different frequencies can be effectively guaranteed, and the problems of data distortion and mutual verification difficulties caused by independent processing of each collection level in the existing technology are effectively solved, thereby achieving high-quality collection and processing of power grid operation data, and providing reliable data support for the safe and stable operation of the power grid.
[0025] 2. Due to the adoption of a deep analysis and tracing mechanism for abnormal data, by monitoring the data correction amplitude, analyzing the cause of the abnormality in combination with relevant business events, and performing pattern matching and trend analysis based on historical abnormal data, data quality problems can be discovered in a timely manner and traced back to their source. This effectively solves the problems of data anomalies being difficult to locate and lacking early warning mechanisms in existing technologies, thereby achieving rapid identification and proactive prevention of data quality issues, and improving the reliability and availability of power grid operation data.
[0026] 3. Due to the adoption of a high-frequency data correction method based on data distribution characteristics, by comparing and analyzing the distribution characteristics of historical data, combining the key indicators in the low-frequency data to calculate the correction coefficient, and performing refined data processing, it is possible to achieve accurate correction while maintaining the statistical characteristics of the data, effectively solving the problem of data distortion caused by high-frequency data correction in the existing technology, and thus achieving precise correction of high-frequency data, ensuring the accuracy and reliability of the power grid operation status assessment. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 This is a flow chart of a data processing method for a power grid enterprise according to an embodiment of the present application; Figure 2 This is another flowchart of the data processing method for a power grid enterprise according to an embodiment of the present application; Figure 3 It is a schematic diagram of the structure of a physical device of the data processing system in an embodiment of the present application. DETAILED DESCRIPTION
[0028] The terms used in the following examples of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification of the present application, the singular expressions "a", "an", "above", "the", and "this" are intended to include plural expressions as well, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in the present application refers to any or all possible combinations of one or more of the listed items.
[0029] In the following, the terms "first" and "second" are used for descriptive purposes only and should not be understood to imply or suggest relative importance or implicitly indicate the number of the technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of this application, unless otherwise specified, "plurality" means two or more.
[0030] For ease of understanding, the application scenarios of the embodiments of the present application are introduced below.
[0031] During the construction of a smart grid, a provincial power grid company deployed a large number of intelligent measurement devices, including PMUs and smart meters. The data collection frequencies of these devices varied widely, ranging from milliseconds to hours. During grid situational awareness, inconsistencies were discovered between data of different frequencies, impacting decision-making accuracy. For example, the PMU at a 330kV substation collected voltage phase angle data 50 times per second, while conventional measurement and control equipment collected voltage RMS values every five minutes. When analyzing equipment operating status, the trends displayed by the two data types deviated, making it difficult to accurately determine whether equipment was experiencing abnormal conditions. In particular, during grid disturbances, high-frequency data would reveal short-term voltage fluctuations, while low-frequency data would not reflect these fluctuations. This made it difficult for operators to promptly detect and address potential faults.
[0032] In related technologies, unified processing of high-frequency and low-frequency data can be achieved by using data averaging and linear interpolation. The following describes a scenario in which the data processing method of a power grid enterprise in related technologies is used.
[0033] To address data inconsistencies, the power grid company previously employed a simple data averaging method. For high-frequency data, the average within a time window was calculated to match the sampling period of low-frequency data. For low-frequency data, linear interpolation was used to supplement sampling points. However, in practice, this method was found to have significant flaws. For example, during a line fault analysis, the substation bus voltage experienced a -7% fluctuation within 0.5 seconds. However, due to the use of a 5-minute average, this important transient characteristic completely disappeared after data processing. Furthermore, simple interpolation of low-frequency data failed to reflect the actual physical changes, resulting in significant deviations from the actual operating conditions and affecting the accuracy of fault analysis.
[0034] The data processing method for power grid enterprises in the embodiments of this application, by establishing a feature mapping model and a multi-dimensional data reliability assessment mechanism, achieves intelligent collaboration and anomaly correction for high- and low-frequency data. This not only preserves important feature information in the data, but also effectively identifies and addresses data quality issues. The following describes scenarios in which the data processing method for power grid enterprises in this application is used.
[0035] After adopting this solution, the power grid company achieved intelligent collaboration between high-frequency and low-frequency data. In a distribution network fault analysis, the system first extracted all relevant measurement point data on the fault line based on preset association rules, including millisecond-level fault recording data and minute-level telemetry data. By establishing a feature mapping model, the system accurately captured the voltage sag characteristics when the fault occurred. When a large deviation was found between the low-frequency telemetry data and the simulation data, the system automatically evaluated the data confidence and found that the communication quality of some measurement and control devices had degraded. Based on historical data and simulation data, the system made targeted corrections to the low-frequency data, which not only retained important transient characteristics but also ensured data continuity, providing a reliable basis for fault location.
[0036] It can be seen that the data processing method of the power grid enterprise in the embodiment of the present application can not only realize the unified processing of high-frequency and low-frequency data, but also effectively solve the problems of data feature loss and inaccurate quality assessment, thereby realizing the intelligent management and application of power grid operation data.
[0037] For ease of understanding, the following describes the process of the method provided by this implementation in combination with the above scenario. Figure 1 , which is a flow chart of the data processing method of the power grid enterprise in the embodiment of the present application.
[0038] S101: Acquire power grid operation data, and extract associated operation data from the power grid operation data based on preset association rules.
[0039] Among them, power grid operation data refers to various types of data collected by power grid companies during daily operations, including but not limited to electrical parameters such as voltage, current, active power, reactive power, as well as operating data such as equipment operating status and environmental parameters; preset association rules represent a set of pre-defined rules for judging the correlation between data, including physical connection relationship rules, operation logic relationship rules, equipment subordination relationship rules, etc.; associated operation data refers to a set of data with mutual correlation selected according to preset association rules.
[0040] At the beginning of the data processing process at the power grid enterprise, the data processing system first needs to obtain complete grid operation data. Specifically, the data processing system establishes a communication connection with the power grid enterprise's various data acquisition devices to receive grid operation data in real time or in batches at regular intervals. The data processing system then traverses the preset association rule library, applying each rule in turn to analyze and filter the grid operation data, extracting data subsets that meet the association conditions to form a collection of associated operation data. For example, for data from a specific substation, the data processing system will extract related data such as circuit breaker status data and line power data on the same busbar based on device topology relationship rules.
[0041] In some embodiments, the process of extracting associated data can be implemented in a variety of ways: Optionally, the data processing system can first construct a data association graph, modeling the physical connection relationships and logical control relationships between devices as a graph structure, where nodes represent devices or measurement points and edges represent association relationships, and then use a graph traversal algorithm to identify mutually associated data sets; Optionally, the data processing system can also use a rule engine-based approach to encode various association rules into a rule set and automatically derive the association relationships between data through a rule reasoning mechanism. It is understandable that other data mining or machine learning methods can also be used to identify the associations between data, which are not limited here.
[0042] In practical applications, data processing systems can encounter problems where association rules are too strict, resulting in too little extracted data, or too loose, leading to unrelated data being incorrectly associated. Data processing systems can employ an adaptive rule adjustment mechanism: first, a target range for associated data coverage is set. If the amount of extracted data falls below the lower limit, the rule conditions are appropriately relaxed; if it exceeds the upper limit, the rule conditions are tightened. Furthermore, by calculating the correlation coefficient between the associated data and setting a correlation threshold, the correlation is filtered out, eliminating weakly correlated data pairs and ensuring the accuracy of the extracted results.
[0043] S102: Divide the associated operation data into high-frequency data whose collection frequency is higher than a preset collection frequency, and low-frequency data whose collection frequency is not higher than the preset collection frequency.
[0044] Among them, the acquisition frequency refers to the time interval for the data acquisition device to acquire data; the preset acquisition frequency represents a pre-set frequency threshold used to distinguish high-frequency data from low-frequency data; high-frequency data refers to a data subset with an acquisition frequency higher than the preset acquisition frequency; low-frequency data refers to a data subset with an acquisition frequency no higher than the preset acquisition frequency.
[0045] After extracting the associated operating data, the data processing system needs to classify and process the data. Specifically, the data processing system first obtains the acquisition frequency information for each data point. This can be done by analyzing the data timestamp sequence to calculate the actual acquisition frequency, or by directly obtaining the nominal acquisition frequency from the configuration information of the data acquisition device. The acquired acquisition frequency is then compared with the preset frequency threshold, and the data is divided into a high-frequency data set and a low-frequency data set based on the comparison result. For example, when the preset acquisition frequency is once per minute, the phasor data collected by the PMU device every second will be classified as high-frequency data, while the power quality data collected every 10 minutes will be classified as low-frequency data.
[0046] In some embodiments, the frequency division of data can be achieved in various ways: optionally, the data processing system can establish a hierarchical acquisition frequency model, layering different types of data according to their natural acquisition period, and then determining the relative high and low frequency attributes according to the hierarchical relationship; optionally, the data processing system can also use a dynamic threshold method to adaptively adjust the threshold of the preset acquisition frequency according to the time distribution characteristics of the data. It can be understood that other time series analysis methods can also be used to achieve the frequency division of data, which is not limited here.
[0047] In practical applications, the data processing system may encounter unstable data acquisition frequency, resulting in a large difference in acquisition frequency of the same data source at different time periods. The data processing system can use a sliding window frequency evaluation method: set an appropriate time window, count the average acquisition frequency and frequency fluctuation range of the data within the window, and when the fluctuation range exceeds the set threshold, perform frequency normalization processing on the data through data interpolation or resampling, etc. to ensure the stability of the frequency division.
[0048] S103, collect high-frequency data in a preset time period, and generate simulated data of low-frequency data based on the high-frequency data simulation.
[0049] Among them, the preset time period represents a predefined time interval for data analysis, which can be fixed length or dynamically adjusted according to actual needs; the simulated data refers to a data set generated by calculating high-frequency data through a mathematical model, used to replace or verify low-frequency data; the feature mapping model refers to a mathematical model used to describe the corresponding relationship between high-frequency data and low-frequency data, including but not limited to statistical regression model, neural network model, etc.
[0050] After completing the data frequency division, the data processing system needs to generate simulated data for verification based on the high-frequency data. Specifically, the data processing system first continuously collects high-frequency data within a preset time period to ensure the integrity and continuity of the data; then the feature mapping model is used to downsample and feature extract the high-frequency data to generate simulated data with the same sampling frequency and data characteristics as the low-frequency data. For example, for PMU data collected every second, statistical features such as mean and standard deviation every 5 minutes can be calculated to generate simulated data corresponding to 5-minute sampled telemetry data.
[0051] The feature mapping model is trained using paired samples of high-frequency and low-frequency data from a historical database. The input is a high-frequency data sequence within a fixed time window (e.g., voltage data sampled at 50 Hz), and the output is the low-frequency data value at the corresponding time point (e.g., voltage data sampled every 5 minutes). The training goal is to minimize the mean squared error between the model output and the actual low-frequency data, while taking into account the temporal correlation and physical constraints of the data. The model itself can adopt a deep neural network structure, consisting of a time series feature extraction layer (such as a CNN or LSTM) and a feature mapping layer, or a multi-scale decomposition and reconstruction structure based on wavelet transforms. When used, the high-frequency data for the current time period is input into the model to obtain the corresponding low-frequency data prediction value.
[0052] In some embodiments, the analog conversion of high-frequency data to low-frequency data can be achieved in a variety of ways: Optionally, the data processing system can construct a time series prediction model based on deep learning, which includes a multi-layer convolutional neural network for extracting the time series features of high-frequency data, and a recurrent neural network for capturing long-term dependencies, and finally outputs the corresponding low-frequency data prediction value through a fully connected layer; Optionally, the data processing system can also use a physical model-based method to establish a state equation based on the dynamic characteristics of the power system, and calculate the system state at different time scales through a numerical integration method. It is understandable that other data dimensionality reduction or feature extraction methods can also be used to achieve frequency conversion, which is not limited here.
[0053] In practical applications, data processing systems often encounter noise or outliers in high-frequency data, which can affect the accuracy of simulation data generation. Data processing systems can employ multi-stage filtering and anomaly detection mechanisms: first, median filtering and wavelet transform are used to denoise the high-frequency data. Then, local anomaly detection algorithms are used to identify and correct outliers before generating simulation data. Furthermore, by setting feature extraction windows at different time scales, the dynamic characteristics of the data at different frequencies can be captured, improving the accuracy of simulation results.
[0054] S104: Calculate the data deviation value between the simulation data and the low-frequency data.
[0055] The data deviation value indicates the degree of difference between the simulated data and the actual low-frequency data, which can be quantified by a variety of statistical indicators; the deviation calculation methods include but are not limited to mean square error, mean absolute error, relative error, etc.
[0056] After generating the simulated data, the data processing system needs to evaluate the accuracy of the simulation results. Specifically, the data processing system first time-aligns the simulated data and low-frequency data to ensure that the data being compared are from the same moment. It then calculates the deviation between the data at corresponding moments. Multiple statistical indicators can be used for comprehensive evaluation to obtain a deviation value that reflects the degree of data discrepancy. For example, the mean and standard deviation of the relative error can be calculated to characterize the overall deviation level and fluctuation.
[0057] In some embodiments, data deviation can be calculated using a variety of methods: Alternatively, the data processing system can employ a distance-based similarity measurement method, treating the data sequence as a vector in a high-dimensional space and measuring the differences between the data by calculating metrics such as Euclidean distance and cosine similarity. Alternatively, the data processing system can employ a distribution-based comparison method, calculating the probability distribution characteristics of the data and using information theory metrics such as KL divergence to measure the differences between distributions. It is understood that other statistical analysis methods can also be employed to assess data deviation, and these are not limited herein.
[0058] In practical applications, data processing systems often encounter significant differences in the dimensions and ranges of different data types, leading to incomparable deviation calculation results. Data processing systems can employ an adaptive normalization approach: first, different types of data are grouped. Within each group, normalization parameters are determined based on the statistical characteristics of historical data, transforming the data into a unified scale space. The normalized deviation values are then calculated, and weighted according to the importance of the data to produce a weighted average comprehensive deviation indicator.
[0059] S105 : When the data deviation value is higher than a preset deviation threshold, calculating a data confidence value of the low-frequency data.
[0060] Among them, the preset deviation threshold represents the critical value used to judge whether the data deviation is abnormal, which can be determined based on historical experience or statistical analysis; the data confidence value refers to a quantitative indicator reflecting the reliability of low-frequency data, and its value range is usually between 0 and 1.
[0061] When abnormal data deviations are detected, the data processing system needs to further evaluate the reliability of the low-frequency data. Specifically, the data processing system first determines whether the calculated data deviation value exceeds a preset threshold. If so, a multi-dimensional reliability assessment of the low-frequency data is required, including analyzing the data's conformity with historical trends, verifying the operating status of the acquisition equipment, and checking the quality of data transmission. Ultimately, a comprehensive calculation is performed to obtain a data confidence value. For example, when the data at a certain measurement point deviates significantly from the historical level for the same period, and the corresponding acquisition equipment is in an abnormal state, its data confidence value will be assessed as low.
[0062] In some embodiments, data confidence can be assessed in a variety of ways: Alternatively, the data processing system can establish an assessment model based on fuzzy reasoning, using various assessment indicators as input variables and deriving a final confidence score through fuzzy rule reasoning. Alternatively, the data processing system can employ a method based on evidence theory, treating assessment results from different dimensions as independent sources of evidence and fusing evidence using the DS evidence theory framework to obtain a comprehensive confidence judgment. It is understood that other data quality assessment methods can also be used to calculate confidence values, and these are not limited here.
[0063] In practical applications, data processing systems often encounter the problem of mutual influence between different evaluation dimensions, leading to biased confidence assessment results. Data processing systems can employ a dynamic weight adjustment mechanism: by analyzing the correlations between evaluation dimensions in historical data, a dimension influence relationship network is established. When calculating confidence values, the weight coefficients are dynamically adjusted based on the actual performance of each dimension under current working conditions, reducing assessment bias caused by mutual influence between dimensions.
[0064] S106: When the data confidence value is lower than a preset confidence threshold, correct the low-frequency data according to the historical database and simulation data.
[0065] Among them, the preset confidence threshold represents the confidence critical value for judging whether the data needs to be corrected; the historical database refers to the data warehouse that stores historical operation data, including long-term accumulated operation records; data correction refers to the process of correcting and optimizing the original data through certain algorithms.
[0066] When the confidence level of low-frequency data is low, the data processing system needs to correct it. Specifically, the data processing system first compares the calculated data confidence value with a preset threshold. If it is lower than the threshold, it indicates that the data may have quality issues and needs to be corrected. During the correction process, the system simultaneously references similar data in the historical database and previously generated simulation data. Through weighted fusion or optimized selection, it generates a corrected value to replace the original low-frequency data. For example, a data distribution model can be established based on historical data from the same period, and outliers can be corrected in combination with simulation data.
[0067] In some embodiments, correction of low-frequency data can be achieved through a variety of methods: Optionally, the data processing system can employ a correction method based on Bayesian estimation, using the distribution of historical data as prior knowledge and simulated data as observations, to obtain the optimal correction value through posterior probability estimation. Optionally, the data processing system can also employ a method based on optimization theory to establish an optimization model that considers multiple objectives such as data continuity and physical constraints, and solve for the optimal correction result. It is understood that other data fusion or optimization methods can also be employed to achieve data correction, which are not limited here.
[0068] In practical applications, the data processing system may encounter the problem of over-correction in the correction process, resulting in the loss of original characteristics of the corrected data. The data processing system can adopt a progressive correction strategy: first, set the correction step parameter to divide the correction process into multiple small steps; in each step, evaluate the deviation of the correction result from the original data, and appropriately reduce the correction amplitude when the deviation exceeds the preset range; at the same time, keep the key feature points of the data to ensure that the correction process does not destroy the essential characteristics of the data.
[0069] S107、In the case that the data confidence value is not lower than the preset confidence threshold, the high-frequency data is corrected according to the historical database and the low-frequency data.
[0070] When the low-frequency data has a high confidence, the data processing system will use these reliable low-frequency data to correct the high-frequency data. Specifically, the data processing system first confirms that the data confidence value is not lower than the preset threshold; then analyzes the distribution characteristic difference between the high-frequency data and the historical data, and combines the key indicators in the low-frequency data to calculate the trend correction coefficient; finally, based on the coefficient, the high-frequency data is corrected to make the overall trend consistent with the reliable low-frequency data. For example, the change trend of the low-frequency data can be used to adjust the mean level and fluctuation range of the high-frequency data.
[0071] In some embodiments, the correction of high-frequency data can be achieved in various ways: optionally, the data processing system can establish a state estimation model based on Kalman filtering, take the low-frequency data as the observation value, and realize dynamic correction of the high-frequency data through state prediction and update; optionally, the data processing system can also use a wavelet analysis-based method to perform multi-scale decomposition on the high-frequency data, adjust the coefficients on different scales according to the low-frequency data, and then reconstruct to obtain the corrected data. It can be understood that other signal processing or data correction methods can also be used to realize the correction of high-frequency data, which is not limited here.
[0072] In practical applications, the data processing system may encounter the problem of loss of useful transient information in high-frequency data, which may be lost by simple correction. The data processing system can adopt a feature-preserving correction strategy: first, perform feature decomposition on the high-frequency data to identify important transient features and periodic features; in the correction process, different correction methods are used for different feature components to ensure that valuable transient information is preserved; finally, the corrected feature components are recombined to obtain a correction result that conforms to the trend of the low-frequency data and preserves important features.
[0073] The method provided by the embodiment will be further described in more detail. Please refer to Figure 2 , another flowchart of the data processing method of the power grid enterprise in the embodiment of the present application.
[0074] S201: Acquire power grid operation data, and extract associated operation data from the power grid operation data based on preset association rules.
[0075] Referring to step S101 , the data processing system extracts interrelated operation data from the power grid operation data based on preset association rules (such as physical connection relationships, operation logic relationships, etc. between devices).
[0076] S202: Divide the associated operation data into high-frequency data whose collection frequency is higher than a preset collection frequency, and low-frequency data whose collection frequency is not higher than the preset collection frequency.
[0077] Referring to step S102 , the data processing system divides the associated operating data into two categories: high-frequency data and low-frequency data according to the comparison result between the collection frequency and the preset frequency.
[0078] S203: Collect high-frequency data within a preset time period, and generate simulated data of low-frequency data based on the high-frequency data.
[0079] Referring to step S103 , the data processing system collects high-frequency data within a specific time period, and uses the high-frequency data to generate corresponding low-frequency data simulation values through simulation calculation.
[0080] In some embodiments, the data processing system performs data simulation based on feature mapping; that is, the data processing system extracts paired samples of high-frequency data and low-frequency data from a historical database; constructs a feature mapping model based on the paired samples; collects high-frequency data within a preset time period and inputs it into the feature mapping model to obtain simulated data.
[0081] Among them, paired samples represent pairs of high-frequency data and low-frequency data samples corresponding to time; the feature mapping model refers to a mathematical model that describes the conversion relationship from high-frequency data to low-frequency data.
[0082] When verifying the reliability of low-frequency data, the data processing system needs to establish a mapping relationship between high- and low-frequency data based on historical data. Specifically, the data processing system first retrieves and extracts pairs of high-frequency and low-frequency data samples with matching timestamps from the historical database to ensure the representativeness and integrity of the samples. Then, based on these paired samples, a feature mapping model is constructed through machine learning or statistical modeling methods. This model can accurately describe the characteristic changes of high-frequency data during the downsampling process. Finally, the high-frequency data collected during the current period is input into the trained feature mapping model to obtain the corresponding low-frequency data simulation value. For example, for the voltage data of a transformer, a mapping model from second-level data to minute-level data can be established based on historical data.
[0083] In some embodiments, the construction and application of the feature mapping model can be achieved in a variety of ways: Optionally, the data processing system can adopt a deep neural network method, first divide the high-frequency data into time windows and extract features, build a feature extraction network containing multiple convolutional layers and pooling layers, and then realize the mapping of features to low-frequency data through a fully connected layer, and finally use the back propagation algorithm to optimize the network parameters; Optionally, the data processing system can also adopt a method based on wavelet transform, first perform wavelet decomposition on the high-frequency data to obtain characteristic coefficients of different frequency bands, and then establish a regression model between the characteristic coefficients and the low-frequency data, and finally obtain simulated data through model prediction. It is understandable that other feature extraction and mapping methods can also be used to achieve data conversion, which is not limited here.
[0084] In practical applications, data processing systems may encounter the problem of abnormal data in historical samples, which can affect the training of feature mapping models. Data processing systems can adopt a robust modeling strategy: first, perform anomaly detection and data cleaning on historical samples to remove or correct abnormal samples; then, adopt modeling methods with strong interference resistance, such as regression models with regularization terms or neural networks with attention mechanisms; and simultaneously, use cross-validation to evaluate the generalization performance of the model and select the optimal model parameters. Furthermore, a sample quality assessment mechanism can be established to assign different training weights to samples of different qualities to improve model accuracy.
[0085] S204: Calculate the data deviation value between the simulation data and the low-frequency data.
[0086] Referring to step S104 , the data processing system calculates the deviation between the simulated low-frequency data and the actually collected low-frequency data.
[0087] S205: When the data deviation value is higher than the preset deviation threshold, calculate the data confidence value of the low-frequency data.
[0088] Referring to step S105 , when the calculated data deviation value exceeds a preset deviation threshold, the data processing system performs a confidence assessment on the actually collected low-frequency data.
[0089] In some embodiments, the data processing system evaluates data confidence based on historical operating trends and equipment operating status; that is, when the data deviation value is higher than a preset deviation threshold, the data processing system extracts the historical operating trends and equipment operating status corresponding to the low-frequency data; calculates the degree of deviation of the low-frequency data based on the historical operating trends, and determines the data reliability coefficient based on the equipment operating status; and determines the data confidence value based on the degree of deviation and the data reliability coefficient.
[0090] Among them, the historical operating trend represents the changing pattern of low-frequency data in the same historical period; the equipment operating status refers to the working condition of the equipment related to data collection; the degree of deviation represents the degree to which the data deviates from the historical trend; the data reliability coefficient refers to the data credibility coefficient based on the equipment status assessment; the data confidence value represents an evaluation indicator that comprehensively reflects the reliability of the data.
[0091] When the data processing system discovers abnormal data deviations, it needs to conduct a comprehensive assessment of the reliability of the low-frequency data. Specifically, the data processing system first extracts the operating data of the measuring point in the same historical period from the historical database, constructs a standard operating trend model, and obtains the operating status information of the data acquisition equipment, including equipment health, communication quality, etc.; then calculates the degree of deviation of the current low-frequency data from the historical operating trend, taking into account the influence of factors such as seasonality and periodicity, and evaluates the reliability coefficient of data acquisition based on the equipment operating status information; finally, the degree of deviation and the reliability coefficient are weighted and fused to calculate a confidence value reflecting the overall reliability of the data. For example, when the power data of a certain line deviates significantly from the historical load curve and the corresponding acquisition equipment is in an alarm state, its data confidence value will be assessed as a low level.
[0092] In some embodiments, data reliability assessment can be achieved in a variety of ways: Optionally, the data processing system can use a method based on time series analysis to first perform seasonal decomposition and trend extraction on historical data, establish a time series model containing trend terms, seasonal terms, and random terms, and then calculate the standardized deviation between the current data and the model prediction value, and obtain the final confidence value through a fuzzy comprehensive evaluation method in combination with the equipment status index; Optionally, the data processing system can also use a method based on statistical process control to establish a multi-dimensional control chart model, analyze the out-of-control situation of the data in various dimensions, and combine the equipment reliability assessment results to comprehensively determine the data confidence level. It is understandable that other reliability assessment methods can also be used to calculate data confidence values, which are not limited here.
[0093] In practical applications, data processing systems often encounter the problem of mutual influence between different evaluation dimensions, leading to biased confidence value calculations. Data processing systems can employ a dynamic weight adjustment mechanism: first, a correlation model for evaluation indicators is established to analyze the influence relationships between them; then, indicator weights are dynamically adjusted based on the performance characteristics of each indicator under current operating conditions; finally, a feedback correction mechanism is introduced to continuously optimize the weight parameters based on the accuracy of historical evaluation results. Furthermore, a multi-level evaluation model can be established, performing reliability assessments at different levels and determining the final confidence value using the Analytic Hierarchy Process (AHP).
[0094] S206: When the data confidence value is lower than a preset confidence threshold, correct the low-frequency data according to the historical database and simulation data.
[0095] Referring to step S106 , when the data confidence value is low, the data processing system will correct the low-frequency data by combining the data in the historical database and the simulation calculation results.
[0096] S207: When the correction range of the low-frequency data exceeds a preset correction range, obtain business events associated with the low-frequency data within a preset time period.
[0097] Among them, the correction amplitude indicates the degree of change before and after the low-frequency data is corrected; the preset correction range refers to the allowable range of data correction changes; and the business event refers to the power grid operation event related to data changes, including equipment maintenance, load adjustment, fault handling, etc.
[0098] When low-frequency data undergoes a significant correction, the data processing system needs to further analyze the cause of the anomaly. Specifically, the data processing system first calculates the difference between the corrected data and the original data to determine whether it exceeds the preset correction range. If so, the system then queries the operation logs, operation records, and alarm information for that time period to extract business event information related to the low-frequency data, providing a basis for subsequent cause analysis. For example, if the power data of a certain line undergoes a significant correction, the system automatically checks to see if there are any related events such as line maintenance or equipment switching.
[0099] In some embodiments, business event correlation analysis can be implemented in a variety of ways: Optionally, the data processing system can establish an event correlation model based on a knowledge graph, constructing entities such as devices, data, and events and their relationships into a knowledge network, and performing multi-hop correlation analysis using graph algorithms to discover potential correlations between data anomalies and events. Optionally, the data processing system can also employ a method based on temporal association rules to identify key events that lead to data anomalies by mining the temporal patterns between event sequences and data changes. It is understood that other event analysis methods can also be used to achieve business correlation discovery, and these are not limited here.
[0100] In practical applications, data processing systems often encounter multiple related events within the same time period, making it difficult to accurately identify the primary influencing factors. Data processing systems can employ an event screening mechanism based on quantified impact: First, an event impact assessment model is established, taking into account factors such as event type, level, and impact range, to calculate the potential impact of each event on the data. Weight coefficients are then set based on temporal and spatial correlations, and a comprehensive assessment is performed to derive an impact score for each event. Finally, key events with high impact are screened for focused analysis.
[0101] The event impact assessment model is constructed based on evidence theory and fuzzy reasoning. During the training phase, the model classifies historical events by type, level, and impact range, constructing a multidimensional feature vector. It also extracts data change characteristics within the corresponding time period, including magnitude, duration, and degree of fluctuation. Event characteristics are mapped to the impact space using fuzzy membership functions, establishing a basic probability distribution function. Evidence from different dimensions is integrated using the Dempster-Shafer evidence theory framework to calculate confidence intervals for various possible impact levels. In spatiotemporal correlation analysis, an exponential decay function is used to describe the time delay effect of the event's impact, and a topological distance function is used to characterize the spatial propagation characteristics. The information entropy criterion is used to optimize evidence weights, enabling dynamic assessment of the event's impact. During use, the model receives feature information about the current event and, through fuzzy reasoning and evidence fusion, outputs a normalized impact score, along with an uncertainty interval for the assessment result. This model effectively addresses uncertainty and ambiguity in event impact assessment, providing a reliable quantitative basis for data anomaly analysis.
[0102] Furthermore, data processing systems must account for event propagation delays and set appropriate time windows to capture event impacts. Furthermore, a library of event types and impact patterns should be established to accumulate historical experience and continuously optimize the accuracy of event correlation analysis. For example, for equipment maintenance events, typical data change patterns can be summarized from historical data to quickly identify similar situations.
[0103] S208. Determine the cause of data anomalies based on business events and generate a data quality analysis report.
[0104] Among them, the cause of data anomaly indicates the root cause of data quality problems; the data quality analysis report refers to an analysis document containing anomaly description, cause analysis, impact assessment, etc.; the types of anomaly causes include equipment failure, communication interruption, parameter configuration error, etc.
[0105] After acquiring relevant business events, the data processing system needs to conduct an in-depth analysis of the causes of data anomalies. Specifically, the data processing system first categorizes and organizes the collected business events, establishing a correspondence between the events and data anomalies. It then analyzes the impact of the events on the data from multiple dimensions, including the timing characteristics of the events, the scope of impact, and the propagation path. Finally, it generates a data quality analysis report that includes the complete analysis process and conclusions. For example, by analyzing the maintenance operation sequence of a certain line, it is possible to explain the cause of abnormal changes in current data during that period.
[0106] In some embodiments, diagnostic analysis of the cause of anomalies can be achieved through a variety of methods: Optionally, the data processing system can construct a diagnostic model based on causal reasoning, constructing a causal network based on information such as data anomalies, device status, and operational events, and using probabilistic reasoning methods to locate the most likely cause of the anomaly while simultaneously evaluating the confidence of different causes. Optionally, the data processing system can also employ an expert system-based approach, utilizing a predefined diagnostic rule base and reasoning mechanism to automatically derive the logical chain of anomaly causes. It is understood that other intelligent diagnostic methods can also be used to implement anomaly cause analysis, which is not limited here.
[0107] In practical applications, data processing systems often encounter complex propagation paths for anomalies, making it difficult to accurately trace the root cause. Data processing systems can employ a multi-layered cause analysis strategy: first, construct an anomaly propagation tree model, using observed anomalies as leaf nodes. Through layer-by-layer analysis, the anomaly propagation chain is established. The necessity and sufficiency of each propagation layer are then verified, assessing the contribution of each factor. Finally, a pruning algorithm is used to retain the main influencing paths and identify the key root causes.
[0108] Furthermore, the data processing system needs to regularly archive analysis reports and accumulate knowledge, establish a library of exception cases, and support experience reuse. At the same time, it needs to design templates for automatic report generation to ensure the standardization and readability of analysis results.
[0109] In some embodiments, the data processing system performs pattern recognition and trend analysis based on historical abnormal data; that is, the data processing system extracts historical abnormal data corresponding to the causes of data abnormalities from the historical database; performs pattern matching on low-frequency data and historical abnormal data to determine the abnormal development trend and impact degree; and generates abnormal handling suggestions based on the abnormal development trend and impact degree.
[0110] Among them, historical anomaly data refers to data records of similar anomalies that occurred in the past; pattern matching refers to the process of comparing the similarity of the current anomaly with historical cases; the anomaly development trend indicates the possible evolution direction and law of data anomalies; the degree of impact refers to the potential impact of the anomaly on the system operation; and the anomaly handling suggestion refers to the handling plan and warning information provided for the current anomaly.
[0111] When the data processing system discovers data anomalies, in-depth analysis and early warning need to be performed based on historical experience. Specifically, the data processing system first retrieves historical case data similar to the current anomaly type in the historical database, including data characteristics, evolution process, and processing results when the anomaly occurs; then performs multi-dimensional pattern matching of the anomaly characteristics of the current low-frequency data with the historical cases, evaluates the similarity of the anomaly, and analyzes the possible development trend of the anomaly and the impact range on the system based on the matching results; finally, according to the analysis results, the historical processing experience is combined to automatically generate an anomaly processing scheme including risk warning, processing measures, and prevention suggestions. For example, when it is found that the current data of a certain line fluctuates abnormally, the system will match similar historical fault cases to predict the possible development trend of the fault.
[0112] In some embodiments, anomaly analysis and early warning can be achieved in various ways: optionally, the data processing system can establish a case-based reasoning analysis model, first build a case library containing multi-dimensional information such as anomaly characteristics, environmental conditions, processing measures, etc., then use multi-layer similarity calculation methods for case matching, adjust the historical processing scheme through case adaptability analysis, and finally form a processing suggestion for the current situation; optionally, the data processing system can also use a rule chain reasoning method to establish an anomaly evolution rule library and an impact propagation model, analyze the development path of the anomaly through forward reasoning, determine the key influencing factors through backward reasoning, and generate a targeted processing scheme. It can be understood that other intelligent analysis methods can also be used to realize anomaly early warning, which is not limited here.
[0113] In practical applications, the data processing system may encounter the problem that the historical cases differ from the current situation, and directly applying historical experience may be misleading. The data processing system can use a context-aware analysis strategy: first, establish a multi-dimensional context feature model, including device status, operating conditions, environmental conditions, and other factors; then evaluate the context similarity of the historical cases and adjust the weight of the matching results; finally, integrate the processing experience of multiple similar cases and optimize the scheme in combination with the current specific context to generate more targeted processing suggestions. At the same time, a processing effect evaluation mechanism is established to continuously accumulate and optimize processing experience.
[0114] S209、In the case where the data confidence value is not lower than the preset confidence threshold, the high-frequency data is compared with the same type of data in the historical database in terms of data distribution characteristics, and a data fluctuation deviation is calculated.
[0115] Among them, the data distribution characteristics refer to characteristic indicators describing the statistical law of the data; the same type of data refers to historical data with the same physical meaning and collection conditions; the data fluctuation deviation refers to the difference between the current data fluctuation and the historical normal fluctuation; the characteristic comparison includes the comparison of statistical quantities such as mean, variance, quantile, etc.
[0116] When low-frequency data has a high confidence level, the data processing system needs to evaluate the fluctuation characteristics of the high-frequency data. Specifically, the data processing system first extracts data samples with the same characteristics as the data to be processed from the historical database. It then calculates the statistical distribution characteristics of these samples to establish a baseline model for data fluctuation. Finally, the distribution characteristics of the current high-frequency data are compared with the baseline model to quantitatively assess fluctuation deviations. For example, by comparing the fluctuation range of voltage data at a certain measuring point with the fluctuation characteristics of historical data from the same period, abnormal fluctuation patterns can be discovered.
[0117] In some embodiments, comparative analysis of data distribution characteristics can be achieved through a variety of methods: Optionally, the data processing system can establish a distribution fitting model based on kernel density estimation, perform probability density estimation on historical data and current data, and quantify the distribution difference by calculating the JS divergence or KL divergence between the distributions; Optionally, the data processing system can also use a method based on empirical distribution functions, calculate the empirical distribution functions of two sets of data, and use a two-sample statistical test method to evaluate the significance of the distribution difference. It is understood that other statistical analysis methods can also be used to achieve distribution feature comparison, which is not limited here.
[0118] In practical applications, data processing systems often encounter the problem of data distribution characteristics changing dynamically over time, making static comparison methods incapable of accurately reflecting fluctuation anomalies. Data processing systems can employ adaptive feature extraction mechanisms: First, historical data is stratified by time scale to establish feature models at multiple time granularities. Then, appropriate time windows are selected based on the current time point, and feature benchmarks are dynamically updated. Finally, comparison results at different time scales are weighted and combined to achieve a more accurate assessment of fluctuation deviations.
[0119] S210 : Generate a trend correction coefficient for the high-frequency data based on the key power grid operation indicators and data fluctuation deviations in the low-frequency data.
[0120] Among them, the key indicators of power grid operation refer to important parameters that reflect the operating status of the power grid, including but not limited to voltage level, power flow, equipment load rate, etc.; the trend correction coefficient represents the correction parameter used to adjust the trend of high-frequency data changes; the data fluctuation deviation refers to the difference between the actual fluctuation and the expected fluctuation.
[0121] After completing the data fluctuation and deviation analysis, the data processing system needs to determine a correction scheme for the high-frequency data. Specifically, the data processing system first extracts key indicators reflecting the grid's operating status from the low-frequency data, including characteristic quantities such as the mean and extreme values of various electrical quantities. It then combines the previously calculated data fluctuation and deviation to establish a mathematical model reflecting the data's changing trends. Finally, through optimization calculations, it derives a trend correction coefficient that adjusts the high-frequency data to conform to the current grid's operating characteristics. For example, based on the hourly mean bus voltage and voltage fluctuation deviation, the corresponding correction coefficient for the high-frequency voltage data can be calculated.
[0122] In some embodiments, the trend correction coefficient can be calculated in a variety of ways: Alternatively, the data processing system can establish a correction model based on state estimation, using low-frequency data as a constraint, and solving an optimization problem that considers physical constraints to obtain a correction coefficient that ensures that the high-frequency data meets the operating rules of the power grid. Alternatively, the data processing system can also adopt a data-driven approach, using a machine learning algorithm to establish a mapping relationship between key indicators and correction coefficients, and obtain appropriate correction parameters through model prediction. It is understood that other numerical calculation methods can also be used to determine the correction coefficient, which is not limited here.
[0123] In practical applications, data processing systems often encounter coupling relationships between key indicators, leading to unstable correction coefficient calculations. This system can employ a layered decoupling calculation strategy: first, correlation analysis is performed on key indicators to identify independent groups of indicators. Then, a local correction model is established within each indicator group to calculate the grouped correction coefficients. Finally, through coordinated optimization, the correction coefficients of each group are integrated to achieve a globally consistent correction solution.
[0124] S211. Correct the high-frequency data based on the trend correction coefficient.
[0125] After calculating the trend correction coefficient, the data processing system needs to correct the high-frequency data. Specifically, the data processing system first applies the calculated trend correction coefficient to the original high-frequency data to generate a preliminary correction result. The correction result is then verified for validity in various aspects, including whether it meets physical constraints and is consistent with related data. Finally, based on the verification results, the correction parameters are fine-tuned until the final correction result meets the requirements. For example, when correcting the high-frequency power data of a certain line, it is necessary to ensure that the corrected data meets the power balance relationship.
[0126] In some embodiments, high-frequency data correction can be achieved through a variety of methods: Alternatively, the data processing system can employ an interpolation-based correction method, using key points of the low-frequency data as control points and generating a correction curve that meets trend requirements through methods such as spline interpolation or polynomial interpolation. Alternatively, the data processing system can employ a filtering-based method, designing an adaptive filter that dynamically adjusts the filtering parameters based on the trend correction coefficient to achieve smooth data correction. It is understood that other data processing methods can also be employed to achieve high-frequency data correction, and these are not limited herein.
[0127] In practical applications, data processing systems may encounter local oscillations or discontinuities during the correction process. Data processing systems can employ a multi-scale correction strategy: first, high-frequency data is decomposed into multiple scales, and trend correction is applied at each scale. The correction results at each scale are then smoothed to eliminate local anomalies. Finally, the final correction result is obtained through rescaling. Furthermore, a transition interval is set to ensure smooth data transitions across the time dimension.
[0128] Furthermore, the data processing system needs to store information about the correction process, including correction coefficients, constraints, and verification results, to support subsequent analysis and backtracking. At the same time, a correction effectiveness evaluation mechanism should be established to continuously improve the correction strategy through statistical analysis.
[0129] In some embodiments, the data processing system will accurately correct the high-frequency data through data segmentation statistics and outlier processing; that is, the data processing system will calculate the product of the trend correction coefficient and the high-frequency data to obtain a preliminary corrected data sequence; based on the collection time point of the low-frequency data, the preliminary corrected data sequence will be segmented by time, and the data mean and standard deviation in each time period will be calculated; the data correction interval will be determined based on the data mean and standard deviation, and abnormal fluctuation points in the preliminary corrected data sequence that exceed the data correction interval will be eliminated; the data sequence after eliminating the abnormal fluctuation points will be linearly interpolated to obtain a correction result.
[0130] Among them, the corrected data series refers to the high-frequency data time series that has undergone preliminary correction; the data mean and standard deviation refer to statistics that describe the data distribution characteristics; the data correction interval represents the allowable data fluctuation range; the abnormal fluctuation point refers to the data point that exceeds the reasonable fluctuation range; linear interpolation refers to the method of filling missing data through a linear function.
[0131] When the data processing system completes the calculation of the trend correction coefficient, it needs to segment the high-frequency data for correction and abnormality processing. Specifically, the data processing system first applies the trend correction coefficient to the original high-frequency data to obtain a preliminary corrected sequence; then divides the corrected sequence into multiple time periods based on the collection time points of the low-frequency data, calculates the statistical characteristics of each segment of data, including mean, standard deviation, etc.; then determines a reasonable data fluctuation interval based on these statistical characteristics, identifies and eliminates abnormal fluctuation points that exceed the interval; and finally performs linear interpolation processing on the data sequence after eliminating abnormal points to ensure data continuity and smoothness. For example, for the correction of voltage data, a reasonable fluctuation interval can be determined based on the voltage qualification rate requirement.
[0132] In some embodiments, the segmented correction of data can be implemented in various ways: optionally, the data processing system can use a method based on quantile statistics to first calculate the quantiles of each order of data in each time period, construct an adaptive abnormality detection threshold, then determine a reasonable correction interval in combination with the fluctuation characteristics of the data before correction, and finally use a segmented cubic spline interpolation method to achieve smooth transition; optionally, the data processing system can also use a dynamic correction method based on a sliding window to perform local statistical feature analysis by setting overlapping time windows, dynamically adjust the correction parameters, and achieve smooth correction of data. It can be understood that other data processing methods can also be used to implement segmented correction, which is not limited here.
[0133] In actual application, the data processing system may encounter problems such as breakpoints or excessive smoothing during data correction. The data processing system can use an adaptive correction control strategy: first, analyze the time-varying characteristics of the data to identify important change points that need to be retained; then, when processing abnormal points, set difference constraints according to the physical meaning of the data to avoid unreasonable jumps; finally, by adjusting the interpolation parameters, the necessary dynamic characteristics are retained while ensuring data smoothness. At the same time, a correction effect evaluation mechanism is established to compare the data characteristics before and after correction to ensure the reasonableness of the correction results.
[0134] In the embodiments of the present application, since the data collaborative processing method based on feature mapping is adopted, including data extraction of pre-set association rules, construction of feature mapping model, multi-dimensional data reliability evaluation, and adaptive data correction mechanism, intelligent collaboration and quality mutual verification between high-frequency and low-frequency data can be realized, effectively solving the problems of feature information loss, inaccurate data quality evaluation, and unsatisfactory correction effect in traditional data processing methods, and thus realizing intelligent management of power grid operation data; providing more reliable data support for power grid operation analysis and decision-making, which is of great significance to improving the intelligent level of the power grid.
[0135] The data processing system in the embodiment of the present invention is described below from the perspective of hardware processing. Figure 3 , is a schematic diagram of a physical device structure of a data processing system in an embodiment of the present application.
[0136] It should be noted that Figure 3 The structure of the data processing system shown is only an example and should not limit the functions and scope of use of the embodiments of the present invention.
[0137] like Figure 3 As shown, the data processing system includes a CPU 301, which can perform various appropriate actions and processes according to the programs stored in the ROM 302 or the programs loaded from the storage unit 308 into the RAM 303, such as executing the methods described in the above embodiments. The RAM 303 also stores various programs and data required for system operation. The CPU 301, ROM 302, and RAM 303 are connected to each other via a bus 304. An I / O interface 305 is also connected to the bus 304.
[0138] The following components are connected to the I / O interface 305: an input section 306 including an audio input device, push button switches, and the like; an output section 307 including a liquid crystal display (LCD), an audio output device, indicator lights, and the like; a storage section 308 including a hard disk and the like; and a communication section 309 including a network interface card such as a LAN (Local Area Network) card or a modem. The communication section 309 performs communication processing via a network such as the Internet. A drive 310 is also connected to the I / O interface 305 as needed. Removable media 311, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 310 as needed, so that computer programs read from the removable media can be installed in the storage section 308 as needed.
[0139] In particular, according to an embodiment of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present invention includes a computer program product comprising a computer program carried on a computer-readable medium, the computer program including a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 309 and / or installed from the removable medium 311. When the computer program is executed by the CPU 301, the various functions defined in the present invention are performed.
[0140] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present invention. Each box in the flowchart or block diagram can represent a module, program segment, or part of the code, and the above-mentioned module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings.
[0141] Specifically, the data processing system of this embodiment includes a processor and a memory. The memory stores a computer program. When the computer program is executed by the processor, the data processing method for the power grid enterprise provided in the above embodiment is implemented.
[0142] As another aspect, the present invention further provides a computer-readable storage medium, which may be included in the data processing system described in the above embodiments, or may exist independently and not be incorporated into the data processing system. The storage medium carries one or more computer programs, and when the one or more computer programs are executed by a processor of the data processing system, the data processing system implements the data processing method for the power grid enterprise provided in the above embodiments.
[0143] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.
[0144] As used in the above embodiments, the term “when” may be interpreted to mean “if” or “after” or “in response to determining that” or “in response to detecting that”, depending on the context. Similarly, the phrases “upon determining that” or “if (stated condition or event) is detected” may be interpreted to mean “if determining that” or “in response to determining that” or “upon detecting (stated condition or event)” or “in response to detecting (stated condition or event)”, depending on the context.
Claims
1. A data processing method for a power grid enterprise, characterized in that: Applied to a data processing system, the method comprises: Acquiring power grid operation data, and extracting associated operation data from the power grid operation data based on preset association rules; Dividing the associated operation data into high-frequency data having a collection frequency higher than a preset collection frequency and low-frequency data having a collection frequency not higher than the preset collection frequency; Collecting the high-frequency data within a preset time period, and generating simulated data of the low-frequency data based on the high-frequency data; Calculating a data deviation value between the simulation data and the low-frequency data; When the data deviation value is higher than a preset deviation threshold, calculating a data confidence value of the low-frequency data; When the data confidence value is lower than a preset confidence threshold, correcting the low-frequency data according to the historical database and the simulation data; When the data confidence value is not lower than a preset confidence threshold, the high-frequency data is corrected according to the historical database and the low-frequency data.
2. The method according to claim 1, characterized in that The step of collecting the high-frequency data within a preset time period and generating simulated data of the low-frequency data based on the high-frequency data specifically includes: extracting paired samples of the high-frequency data and the low-frequency data from the historical database; constructing a feature mapping model based on the paired samples; The high-frequency data within a preset time period is collected and input into the feature mapping model to obtain simulation data.
3. The method according to claim 1, characterized in that The step of calculating the data confidence value of the low-frequency data when the data deviation value is higher than a preset deviation threshold specifically includes: When the data deviation value is higher than a preset deviation threshold, extracting the historical operation trend and equipment operation status corresponding to the low-frequency data; Calculating the degree of deviation of the low-frequency data based on the historical operating trend, and determining the data reliability coefficient based on the equipment operating status; The data confidence value is determined according to the degree of deviation and the data reliability coefficient.
4. The method according to claim 1, wherein When the data confidence value is lower than a preset confidence threshold, after the step of correcting the low-frequency data according to the historical database and the simulation data, the method further includes: When the correction amplitude of the low-frequency data exceeds a preset correction range, obtaining a business event associated with the low-frequency data within the preset time period; Determine the cause of data anomalies based on the business events and generate a data quality analysis report.
5. The method according to claim 4, characterized in that After the step of determining the cause of the data anomaly according to the business event and generating a data quality analysis report, the method further includes: Extracting historical abnormal data corresponding to the cause of the data abnormality from the historical database; Performing pattern matching on the low-frequency data and the historical abnormal data to determine the abnormal development trend and impact degree; Generate an exception handling suggestion based on the exception development trend and the impact level.
6. The method according to claim 1, characterized in that The step of correcting the high-frequency data according to the historical database and the low-frequency data when the data confidence value is not lower than a preset confidence threshold specifically includes: When the data confidence value is not lower than a preset confidence threshold, the data distribution characteristics of the high-frequency data are compared with similar data in the historical database to calculate the data fluctuation deviation; generating a trend correction coefficient for high-frequency data based on the key grid operation indicators in the low-frequency data and the data fluctuation deviation; The high-frequency data is corrected based on the trend correction coefficient.
7. The method according to claim 6, characterized in that The step of correcting the high-frequency data based on the trend correction coefficient specifically includes: Calculating the product of the trend correction coefficient and the high-frequency data to obtain a preliminary corrected data sequence; Based on the collection time points of the low-frequency data, the preliminary corrected data sequence is divided into time segments, and the mean and standard deviation of the data in each time segment are calculated; Determine a data correction interval based on the data mean and the standard deviation, and remove abnormal fluctuation points in the preliminary corrected data sequence that exceed the data correction interval; The data sequence after removing the abnormal fluctuation points is subjected to linear interpolation processing to obtain a correction result.
8. A data processing system, characterized in that: The data processing system includes: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to cause the data processing system to execute the method according to any one of claims 1 to 7.
9. A computer-readable storage medium comprising instructions, characterized in that: When the instructions are executed on a data processing system, the data processing system is caused to execute the method according to any one of claims 1 to 7.
10. A computer program product, characterized in that When the computer program product is run on a data processing system, the computer program product causes the data processing system to execute the method according to any one of claims 1 to 7.