An optimization method and system for the efficiency of extracting the value of power data

By constructing multi-source heterogeneous data sets and cleaning and standardizing, analyzing similarity coefficients and data distribution coefficients, dynamically generating priority estimation coefficients, the problem of uneven data distribution in power data processing is solved, and efficient and accurate data extraction and system optimization are achieved.

CN119443382BActive Publication Date: 2025-08-05STATE GRID HENAN INFORMATION & TELECOMM CO
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411506539.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-25
Publication Date
2025-08-05
Estimated Expiration
2044-10-25

AI Technical Summary

Technical Problem

When the existing power data processing system faces uneven data distribution, it lacks effective response measures, resulting in inefficient extraction efficiency and unreasonable resource allocation. The existing methods fail to dynamically consider the similarity and distribution between data, affecting system efficiency and accuracy.

Method used

By constructing a multi-source heterogeneous data set, data cleaning, standardization and binning are carried out, similarity coefficients and data distribution coefficients are analyzed, priority estimation coefficients are dynamically generated, parameter extraction order is determined, and data extraction process is monitored and optimized in real time, combining efficiency prediction models and security thresholds.

Benefits of technology

It effectively avoids bottleneck problems caused by uneven data distribution, ensures priority processing of key data, improves overall extraction efficiency and accuracy, enhances the sensitivity and responsiveness of the system, and ensures the stable and efficient operation of the power system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119443382B_ABST
    Figure CN119443382B_ABST
Patent Text Reader

Abstract

The present invention discloses an optimization method and system for the efficiency of extracting the value of power data, which relates to the technical field of power data processing. This method constructs a multi-source heterogeneous data set from multiple links within the power system, and performs cleaning, standardization, and binning processing on it, making the data set more standardized and consistent before feature extraction. This preprocessing operation reduces noise, outliers, and data redundancy in the data, improves the accuracy of subsequent data analysis, and ensures that the system can process and analyze large amounts of power data more efficiently. By introducing a similarity coefficient Xsxs and a data distribution coefficient Xbxs, the system can perform similarity and distribution analysis according to the actual situation of each parameter, thereby dynamically generating a priority estimation coefficient Ygxs. This enables the system to reasonably determine the priority order of data extraction based on the dependence relationship between real-time data and historical data, as well as the dispersion state of each parameter in time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power data processing, and specifically provides an optimization method and system for the efficiency of extracting the value of power data. Background Art

[0002] With the digital development of modern power systems, the multi-source heterogeneous data from power generation, transmission, distribution, and the user side has increased sharply. The value mining and extraction efficiency of power data has become one of the core links in the optimization and management of power systems. In the specific field of the efficiency of extracting the value of power data, by analyzing the data collected from the power system at different time periods, the key features in the system can be identified, and the operation efficiency and reliability of the system can be improved. Especially in the analysis of data distribution, the distribution of power data at different time periods is often uneven. The data volume in some periods is too large, resulting in an excessive calculation load, while the data in other periods is sparse, causing local bottlenecks in the extraction process. Therefore, how to efficiently extract the value of data under this background and ensure the balanced processing of data in each time period is the key to optimizing the efficiency of extracting the value of power data.

[0003] During the process of analyzing the efficiency of extracting the value of power data, the uneven data distribution is a common phenomenon. If the data volume is too large in some time periods and too small in other time periods, it may lead to local bottlenecks in the data extraction process, resulting in unreasonable resource allocation and low efficiency when the system processes data. Existing systems often lack effective countermeasures for this uneven distribution, resulting in fluctuations in extraction efficiency. In addition, when estimating the priorities of various parameters by existing methods, they usually rely on static or fixed rules and fail to dynamically consider the similarity and distribution of data. Therefore, the unreasonable arrangement of the extraction order may lead to the key parameters not being processed in time, affecting the efficiency and accuracy of the overall system. Summary of the Invention

[0004] Aiming at the deficiencies of the prior art, the present invention provides an optimization method and system for the efficiency of extracting the value of power data, which solves the problems in the above background art.

[0005] To achieve the above objectives, the present invention is realized through the following technical solutions: An optimization method for the efficiency of extracting the value of power data includes the following steps.

[0006] S1. Obtain the power status data monitored by a number of monitoring devices at different time periods from each link in the power system in advance to construct a multi-source heterogeneous data set.

[0007] S2. Perform preprocessing operations before feature extraction on the multi-source heterogeneous data set obtained in S1, and perform data cleaning, normalization processing, and binning processing on the multi-source heterogeneous data set to generate a power data set.

[0008] S3. Based on the power dataset and combined with the power status data real-time monitored by several groups of monitoring devices, analyze the dependence between the real-time parameters and the corresponding parameters in the power dataset to construct a similarity coefficient Xsxs. And in the time series, analyze the dispersion state of the corresponding parameters in different monitoring periods to calculate the data distribution coefficient Xbxs of the corresponding parameters. By correlating the similarity coefficient Xsxs with the data distribution coefficient Xbxs of the corresponding parameters, obtain a priority estimation coefficient Ygxs. Based on the priority estimation coefficient Ygxs, determine the parameter extraction order;

[0009] S4. Based on the extraction order of each parameter in S3, analyze the timeliness in the data extraction process to obtain the data freshness Sqxd, and combined with the trained efficiency prediction model and linear normalization processing, fit and output the extraction efficiency index Txzs;

[0010] S5. Preset a safety threshold Q in advance, and by comparing and analyzing the extraction efficiency index Txzs with the safety threshold Q, comprehensively judge the extraction efficiency status of the parameters during the feature extraction in the corresponding monitoring period, and take corresponding optimization measures based on the judgment results.

[0011] Preferably, the specific steps of S1 include:

[0012] S11. Use several groups of monitoring devices to monitor the power status data in different periods to construct a multi-source heterogeneous dataset. Among them, the multi-source heterogeneous dataset includes the power status data monitored by different monitoring devices in different periods.

[0013] Preferably, the specific steps of S2 include:

[0014] S21. Perform preprocessing operations on the multi-source heterogeneous dataset obtained in S11. The preprocessing operations include data cleaning, normalization processing, and binning processing;

[0015] S211. Among them, by equally binning the parameters in the multi-source heterogeneous dataset, obtain several groups of power data for monitoring periods;

[0016] S22. Perform time-series aggregation on several groups of power data for monitoring periods to re-obtain a power dataset.

[0017] Preferably, the specific steps of S3 include:

[0018] S31. Analyze the dependence between the real-time parameters and the corresponding parameters in the power dataset based on the power dataset and in combination with the power status data output in real time in the power system, and map each parameter in the power status data to a three-dimensional space through non-linear mapping to obtain the similarity coefficient Xsxs of the corresponding parameter. The specific method for obtaining it is as follows:

[0019] In the formula, n represents the monitoring period in the power dataset; i = 1, 2, 3,..., n; represents the contribution of the parameter in the i-th monitoring period in the power dataset to the parameter output in real time in the current power system; represents the parameter in the i-th monitoring period in the power dataset; represents the parameter output in real time in the current power system; represents the kernel function;

[0020] S311. The kernel function is obtained through the following formula:

[0021] In the formula, represents the parameter in the Gaussian kernel function; represents the exponential growth function, specifically , where e is the natural base; represents the Euclidean distance between the parameter in the i-th monitoring period in the power dataset and the parameter output in real time in the current power system.

[0022] Preferably, the specific steps of S3 further include:

[0023] S32. Analyze the dispersion state of the corresponding parameters in different monitoring periods in the time series to calculate the data distribution coefficient Xbxs of the parameters in the corresponding monitoring periods. The specific method for obtaining it is as follows:

[0024] In the formula, represents the average parameter within the monitoring period.

[0025] Preferably, the specific steps of S3 further include:

[0026] S33. Based on the similarity coefficient Xsxs and the data distribution coefficient Xbxs of the corresponding parameters obtained in S31 and S32, and after linear normalization processing, map the corresponding data values to the interval to calculate and obtain the priority estimation coefficient Ygxs. The specific method for obtaining it is as follows:

[0027] In the formula, and are both weight values. is represented as the first correction constant, where and the specific values are set by the user according to the situation;

[0028] S34. Based on the method of obtaining the priority estimation coefficient Ygxs in S33, respectively obtain the priority estimation coefficients Ygxs of several groups of parameters in the power state data output in real time in the power system, and sort the priority estimation coefficients Ygxs of several groups of parameters according to the numerical size to construct a priority data column. According to the priority data column, perform feature extraction on several groups of parameters in the power state data output in real time in the power system according to the priority.

[0029] Preferably, the specific steps of S4 include:

[0030] S41. During the process of feature extraction according to the priority in S34, monitor the freshness during the data extraction process in real time to obtain relevant extraction status data information, and the relevant extraction status data information includes the preprocessing time , the data generation time and the data extraction time ;

[0031] S42. According to the relevant extraction status data information, calculate and obtain the data freshness Sqxd, and the data freshness Sqxd is obtained through the following formula:

[0032] In the formula, is the preprocessing time, indicating the time when the multi-source heterogeneous data is preprocessed; is the data generation time, indicating the moment when the data is generated by the sensor or other devices; is the data extraction time, indicating the moment when the system extracts data from the power dataset.

[0033] Preferably, the specific steps of S4 further include:

[0034] S42. Use the convolutional neural network technology to construct an efficiency prediction model, input the data freshness Sqxd and the priority estimation coefficient Ygxs into the efficiency prediction model, and after linear normalization processing, fit and output the extraction efficiency index Txzs from the efficiency prediction model, and the extraction efficiency index Txzs is obtained through the following formula:

[0035] In the formula, K represents the number of parameters, k = 1, 2, 3,..., K, represents the data freshness of the k-th parameter, represents the average data freshness, represents the priority estimation coefficient of the k-th parameter, is expressed as the average priority estimation coefficient, and are both weight values, P is expressed as the second correction constant, and and their specific values are set by the user according to the situation.

[0036] Preferably, the specific steps of S5 include:

[0037] S51. By comparing and analyzing the pre-set safety threshold Q with the extraction efficiency index Txzs, comprehensively judge the extraction efficiency state of the parameter during feature extraction in the corresponding monitoring period. The specific judgment content is as follows:

[0038] If the extraction efficiency index Txzs falls within the safety threshold Q, it is judged that the extraction efficiency of the parameter during feature extraction in the corresponding monitoring period is in a normal state. At this time, continue the current feature extraction operation on the real-time output power state data in the power system;

[0039] S52. If the extraction efficiency index Txzs does not fall within the safety threshold Q, it is judged that the extraction efficiency of the parameter during feature extraction in the corresponding monitoring period is not in a normal state. At this time, use caching technology to extract and process data in advance, and re-evaluate the priority of each parameter according to the actual application scenario.

[0040] An optimization system for the extraction efficiency of power data value, including a data acquisition module, a preprocessing module, an extraction order analysis module, a comprehensive evaluation module and a feedback module;

[0041] The data acquisition module is used to pre-obtain the power state data monitored by several groups of monitoring devices at different times from each link in the power system to construct a multi-source heterogeneous data set;

[0042] The preprocessing module is used to perform preprocessing operations on the multi-source heterogeneous data set before feature extraction, perform data cleaning, standardization processing and binning processing on the multi-source heterogeneous data set to generate a power data set;

[0043] The extraction order analysis module is used to analyze the dependence between real-time parameters and the corresponding parameters in the power data set based on the power data set and in combination with the power state data monitored by several groups of monitoring devices in real time, to construct a similarity coefficient Xsxs, and analyze the dispersion state of the corresponding parameters in different monitoring periods in the time series to calculate the data distribution coefficient Xbxs of the corresponding parameters. By associating the similarity coefficient Xsxs with the data distribution coefficient Xbxs of the corresponding parameters, obtain the priority estimation coefficient Ygxs, and determine the parameter extraction order based on the priority estimation coefficient Ygxs;

[0044] The comprehensive evaluation module is used to analyze the timeliness in the data extraction process based on the extraction order of each parameter in S3, so as to obtain the data freshness Sqxd, and combined with the trained efficiency prediction model and linear normalization processing, to fit and output the extraction efficiency index Txzs;

[0045] The feedback module is used to preset a safety threshold Q, and by comparing and analyzing the extraction efficiency index Txzs with the safety threshold Q, to comprehensively judge the extraction efficiency status of the parameter during the feature extraction in the corresponding monitoring period, and take corresponding optimization measures based on the judgment result.

[0046] The present invention provides an optimization method and system for the extraction efficiency of power data, and has the following beneficial effects:

[0047] (1) By constructing a multi-source heterogeneous data set from multiple links in the power system and performing cleaning, standardization and binning processing on it, the data set becomes more standardized and consistent before feature extraction. This preprocessing operation reduces noise, outliers and data redundancy in the data, improves the accuracy of subsequent data analysis, and ensures that the system can process and analyze large-scale power data more efficiently. By introducing the similarity coefficient Xsxs and the data distribution coefficient Xbxs, the system can perform similarity and distribution analysis according to the actual situation of each parameter, so as to dynamically generate the priority estimation coefficient Ygxs. This enables the system to reasonably determine the priority order of data extraction according to the dependency relationship between real-time data and historical data, and the dispersion state of each parameter in time, further avoiding bottleneck problems caused by uneven data distribution, ensuring that key data can be processed preferentially, and thus improving the overall extraction efficiency. In the feature extraction process, the method introduces an evaluation mechanism for data freshness Sqxd to ensure that the system can judge the timeliness of extraction according to the real-time nature of the data. Through the real-time calculation of freshness, the system can identify data delay problems and adjust the data extraction strategy according to the delay situation. This mechanism enhances the sensitivity of the system in processing real-time data and helps the power system for instant monitoring and rapid response. The fitting output of the extraction efficiency index Txzs, combined with the setting of the safety threshold Q, can help the system quickly judge whether the extraction efficiency of each parameter in the current monitoring period is in a normal state, so as to ensure that the overall extraction efficiency of the system remains at a high level.

[0048] (2) This method analyzes the dependencies between the parameters in the power dataset and the real-time output parameters, and uses kernel functions to map these parameters into a three-dimensional space. This non-linear mapping can capture complex feature patterns that cannot be extracted by linear methods in the original space. Especially in power data, there are often non-linear relationships between parameters such as voltage and load fluctuations. By mapping to a higher-dimensional space, these implicit feature patterns can be better extracted, which helps to discover potential associations between different parameters. The acquisition of the similarity coefficient Xsxs is based on this non-linear mapping of complex relationships, enabling the system to accurately measure the similarity between the parameters in the power dataset and the current real-time parameters, thus more precisely monitoring and analyzing the power system.

[0049] (3) By analyzing the dispersion state of the corresponding parameters in different monitoring periods, the system can identify the fluctuation degree of the data in each period. By calculating the data distribution coefficient Xbxs, the dispersion degree of different parameters within the monitoring period can be quantified. This helps to distinguish which parameters have larger fluctuations and which are more stable in different time periods. This in-depth analysis of the data dispersion state can help the system prioritize the processing of data segments with larger fluctuations, ensuring that during the feature extraction process, key parameters crucial for power system monitoring and scheduling are focused on, thereby improving the accuracy and effectiveness of overall feature extraction. Based on the similarity coefficient Xsxs and the data distribution coefficient Xbxs, after linear normalization, the corresponding data values are mapped to a unified interval, ensuring that parameters in different monitoring periods can be compared under the same standard. Through this comprehensive processing, the system can more accurately calculate the priority estimation coefficient Ygxs, effectively improving the rationality of the priority ranking of each parameter. This process enables the system to dynamically adjust the priorities of different parameters, ensuring that key parameters with high priorities (such as load fluctuations, voltage anomalies, etc.) can be processed first, thus avoiding low system monitoring efficiency or missing important information due to improper processing order.

[0050] (4) During the priority feature extraction process, the system can monitor the status information of data extraction in real time, including preprocessing time, data generation time, and data extraction time, thus ensuring that at each stage of feature extraction, the system has a comprehensive control over the data extraction and processing process. This function enables the system to effectively avoid problems such as delays and data obsolescence, thereby improving the timeliness of data processing and ensuring that the latest status data can be obtained in a timely manner in the power system, helping to make quick and effective responses and decisions. By comprehensively considering the preprocessing time, data generation time, and data extraction time, the system can accurately calculate the data freshness Sqxd to measure the overall delay degree of data from generation to preprocessing and then to extraction. The dynamic evaluation of data freshness helps to ensure that data can be quickly extracted and processed after generation, further avoiding data aging or obsolescence problems. Brief Description of the Drawings

[0051] Figure 1 It is a schematic flowchart of an optimization method for the efficiency of extracting the value of power data according to the present invention;

[0052] Figure 2 It is a block diagram of an optimization system for the efficiency of extracting the value of power data according to the present invention. Detailed Embodiments

[0053] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0054] Embodiment 1

[0055] Please refer to Figure 1 , the present invention provides an optimization method for the efficiency of extracting the value of power data, including the following steps,

[0056] S1. Obtain power status data monitored by several groups of monitoring devices at different time periods from each link in the power system in advance to construct a multi-source heterogeneous data set;

[0057] S2. Perform preprocessing operations before feature extraction on the multi-source heterogeneous data set obtained in S1, perform data cleaning, normalization processing and binning processing on the multi-source heterogeneous data set, that is, divide continuous variables into discrete intervals to generate a power data set;

[0058] S3. Based on the power data set and in combination with the power status data monitored by several groups of monitoring devices in real time, analyze the dependence between real-time parameters and the corresponding parameters in the power data set to construct a similarity coefficient Xsxs, and analyze the dispersion state of the corresponding parameters in different monitoring time periods in the time series to calculate the data distribution coefficient Xbxs of the corresponding parameters. By associating the similarity coefficient Xsxs with the data distribution coefficient Xbxs of the corresponding parameters, obtain a priority estimation coefficient Ygxs, and determine the parameter extraction order based on the priority estimation coefficient Ygxs;

[0059] S4. Based on the extraction order of each parameter in S3, analyze the timeliness in the data extraction process to obtain data freshness Sqxd, and in combination with the trained efficiency prediction model and linear normalization processing, fit and output an extraction efficiency index Txzs;

[0060] S5. Preset a safety threshold Q, and comprehensively judge the extraction efficiency status of parameters during feature extraction in the corresponding monitoring period by comparing and analyzing the extraction efficiency index Txzs with the safety threshold Q, and take corresponding optimization measures based on the judgment results.

[0061] In this embodiment, the method constructs a multi-source heterogeneous data set by obtaining power state data at different times from different links and multiple monitoring devices in the power system. Through preprocessing operations such as data cleaning, normalization processing, and binning processing, the consistency and availability of data are further ensured, the noise and redundant information in the data are reduced, and subsequent feature extraction is made more efficient and accurate. This processing method solves the problems of complex power data sources and inconsistent structures, effectively improves the data processing efficiency and quality, and provides a reliable data basis for feature extraction. The method constructs a similarity coefficient Xsxs by analyzing the dependence between real-time monitoring data and historical power data, analyzes the dispersion state of parameters in the time series, calculates the data distribution coefficient Xbxs, associates the similarity coefficient with the data distribution coefficient, and generates a priority estimation coefficient Ygxs, so as to dynamically determine the extraction order of each parameter according to the priority. This method not only considers the similarity between data, but also analyzes the data distribution characteristics at different times, ensuring that key parameters can be preferentially extracted during periods with a large amount of data, thus avoiding local bottleneck problems caused by uneven data distribution and improving the efficiency and accuracy of data extraction. By analyzing the timeliness in the data extraction process, the data freshness Sqxd is obtained and input into a trained efficiency prediction model. Through linear normalization processing, the extraction efficiency index Txzs is fitted and output. This step enables the system to evaluate the efficiency in the data extraction process in real time, ensures that key data can be extracted and processed in a timely manner, and further optimizes the response speed and processing capacity of the entire system. By setting the safety threshold Q, the present invention can monitor the extraction efficiency in real time during the data extraction process. When the extraction efficiency index Txzs does not fall within the safety threshold, the system can automatically identify the low-efficiency data extraction situation and take corresponding optimization measures (such as adjusting the extraction order or optimizing parameters) to ensure that the system always operates in an efficient state. This dynamic tuning mechanism not only improves the stability of the system, but also significantly enhances the processing efficiency of power data. In summary, through the construction and preprocessing of a multi-source heterogeneous data set, the combination of similarity coefficient and data distribution coefficient, the real-time analysis of data freshness, and the comparison and optimization of the extraction efficiency index, this method realizes the comprehensive improvement of the power data value extraction efficiency. This method can effectively avoid bottleneck problems caused by uneven data distribution, ensure that the system can still operate efficiently in a complex power data environment, improve the data processing capacity and response speed of the power system, and has important practical value.

[0062] Embodiment 2

[0063] Please refer to Figure 1 , specifically: S1 specific steps include:

[0064] S11. Utilize several groups of monitoring devices to monitor power status data in different time periods to construct a multi-source heterogeneous data set, wherein the multi-source heterogeneous data set includes power status data monitored by different monitoring devices in different time periods.

[0065] The specific steps of S2 include:

[0066] S21. Preprocess the multi-source heterogeneous data set obtained in S11, which includes data cleaning, standardization, and binning. Data cleaning is the first step in preprocessing, and its purpose is to remove or repair anomalies, errors, and redundant information in the data, fill missing values through linear interpolation, Lagrange interpolation, or spline interpolation, and identify and eliminate outliers (such as data points outside 3 times the standard deviation) through the mean and standard deviation.

[0067] S211, wherein, by dividing the parameters of the multi-source heterogeneous data set into equal-frequency bins, the amount of data in each interval is made roughly the same, so as to obtain power data for several groups of monitoring periods;

[0068] S22. Aggregate the power data of several groups of monitoring periods in time series to re-obtain the power data set.

[0069] In this embodiment, by deploying multiple monitoring devices, the system can obtain multi-source heterogeneous data in real time from different links of the power system. These data contain power status information collected at different time periods and by different devices, ensuring the diversity and comprehensiveness of the data sources. Combined with the data cleaning step, the system can effectively remove or repair abnormal, incorrect, and redundant data, and fill in missing values through methods such as linear interpolation, Lagrange interpolation, or spline interpolation to ensure data integrity. The 3-sigma method is also used in the data cleaning process to eliminate outliers, further improving data accuracy, thus laying a solid foundation for subsequent feature extraction. When preprocessing the multi-source heterogeneous data set, the system uses the equal-frequency binning technique to make the number of data in each bin approximately the same. This processing method effectively solves the problem of uneven data distribution, ensuring that data in each monitoring period can be evenly considered in subsequent analysis and reducing analysis biases caused by data volume differences. Through equal-frequency binning, the system can better capture the data change trends in different time periods, improving data representativeness and usability. The last step of preprocessing is to perform time-series aggregation on data from multiple monitoring periods to regenerate the power data set. This process reduces the time-series inconsistency of data by aggregating data within different time periods and lays a foundation for subsequent time-series analysis. Time-series aggregation ensures that when the system processes complex time-series data, it can maintain data continuity and consistency, further avoiding the impact of data jumps or missing values on analysis results, and enhancing the system's ability to capture and analyze data in the time dimension. Through these multi-faceted preprocessing operations, including data cleaning, outlier elimination, equal-frequency binning, and time-series aggregation, the integrity, accuracy, balance, and consistency of power data are further improved. Through these optimization steps, the system can process multi-source heterogeneous data more accurately, ensuring the quality of feature extraction and subsequent analysis, and ultimately enhancing the overall efficiency and accuracy of power data value extraction.

[0070] Embodiment 3

[0071] Please refer to Figure 1 , specifically: The specific steps of S3 include:

[0072] S31. According to the power data set and combined with the power status data output in real time in the power system, analyze the dependence between real-time parameters and the corresponding parameters in the power data set, and map each parameter in the power status data to a three-dimensional space through non-linear mapping, and extract the feature patterns that are inseparable in the original space to obtain the similarity coefficient Xsxs of the corresponding parameters. The specific method for obtaining is as follows:

[0073] In the formula, n represents the monitoring period in the power data set; i = 1, 2, 3,..., n; It is expressed as the contribution of the parameters in the \(i\)-th monitoring period within the power data set to the parameters of the real-time output in the current power system; It is expressed as the parameters in the \(i\)-th monitoring period within the power data set; It is expressed as the parameters of the real-time output in the current power system; It is expressed as a kernel function that maps the parameters in the \(i\)-th monitoring period within the power data set and the parameters of the real-time output in the current power system to a three-dimensional space to capture the non-linear relationship between the parameters in the \(i\)-th monitoring period within the power data set and the parameters of the real-time output in the current power system;

[0074] The similarity coefficient \(X_{sxs}\) of the corresponding parameters reflects the feature representation of the current data point \(x\) in the high-dimensional space.

[0075] S311, the said kernel function It is obtained through the following formula:

[0076] In the formula, It is expressed as the parameter in the Gaussian kernel function, which controls the "smoothness" or "width" of the Gaussian kernel function and determines the attenuation rate of the similarity between data points; It is expressed as an exponential growth function, specifically , where \(e\) is the natural base, approximately equal to 2.718; It is expressed as the Euclidean distance between the parameters in the \(i\)-th monitoring period within the power data set and the parameters of the real-time output in the current power system.

[0077] In this embodiment, by using a kernel function, each parameter in the power data is mapped to a three-dimensional high-dimensional space. This method effectively captures complex non-linear feature patterns that cannot be separated by linear methods in the original space. This non-linear mapping can reflect the deep relationships hidden between different parameters in the power system, especially the dependencies of complex parameters such as power load and voltage fluctuations. Through these feature extraction methods, the system can obtain the similarity coefficient Xsxs between the power state data and real-time parameters in each monitoring period, thereby achieving more accurate state monitoring and prediction, and improving the effectiveness and accuracy of data extraction. By analyzing the contribution of each parameter in the power data set to the real-time output parameter and combining the calculation of the similarity coefficient Xsxs, the system can accurately judge the mutual relationship between each parameter in different monitoring periods according to the dependency of each parameter on the current real-time power state. This dependency analysis can play a key role in data extraction at different time periods, helping to determine the priority extraction order and improving the overall response speed of the system. The use of the kernel function further improves the system's ability to capture non-linear relationships in the power data set, enabling complex feature relationships to be well characterized in the high-dimensional space. In particular, the precise non-linear mapping achieved through the Gaussian kernel function can effectively evaluate the similarity between data points, ensuring that the system can better handle non-linear relationships when processing multi-source data. The parameter σ in the kernel function controls the "smoothness" or "width" of the Gaussian kernel function, allowing the system to flexibly adjust the decay rate of similarity. This parameterized design provides greater flexibility for the system when dealing with different power state data, enabling similarity analysis to be adjusted according to specific application scenarios, thereby enhancing the adaptability of data analysis. Whether facing power load data with large fluctuations or stable voltage data, more refined similarity analysis can be performed by adjusting the smoothness of the kernel function.

[0078] Embodiment 4

[0079] Please refer to Figure 1 , specifically: The specific steps of S3 also include:

[0080] S32. In the time series, analyze the dispersion state of the corresponding parameters in different monitoring periods to calculate the data distribution coefficient Xbxs of the parameters in the corresponding monitoring period, which is specifically obtained in the following manner:

[0081] In the formula, represents the average parameter within the monitoring period.

[0082] The specific steps of S3 also include:

[0083] S33. Based on the similarity coefficient Xsxs and the data distribution coefficient Xbxs of the corresponding parameters obtained in S31 and S32, and after linear normalization processing, map the corresponding data values to the interval to calculate and obtain the priority estimation coefficient Ygxs, which is specifically obtained in the following manner:

[0084] In the formula, and are both weight values, is expressed as the first correction constant, where and The specific values are set by the user according to the situation;

[0085] S34. Based on the method of obtaining the priority estimation coefficient Ygxs in S33, respectively obtain the priority estimation coefficients Ygxs of several groups of parameters in the power state data output in real time in the power system, and sort the priority estimation coefficients Ygxs of several groups of parameters in terms of numerical magnitude to construct a priority data column. According to the priority data column, perform feature extraction on several groups of parameters in the power state data output in real time in the power system according to the priority.

[0086] In this embodiment, by analyzing the dispersion state of power data parameters in different monitoring periods in the time series and calculating the corresponding data distribution coefficient Xbxs, the volatility and distribution uniformity of data in each time period can be effectively measured. This step enables the system to identify which time periods have large data fluctuations and which time periods have more concentrated or sparse data, so as to better adjust the subsequent feature extraction strategy. This analysis process helps to avoid the extraction bottleneck caused by uneven data distribution and improves the stability and accuracy of the system when facing a large amount of data. This method combines the similarity coefficient Xsxs obtained in S31 and the data distribution coefficient Xbxs obtained in S32, and after linear normalization processing, maps the two to a unified interval, and then calculates the priority estimation coefficient Ygxs of each parameter. This processing method not only considers the similarity of each parameter to historical data, but also combines the data distribution in different time periods to ensure more accurate priority estimation. This method can effectively optimize the feature extraction order, ensure that key data is processed first, reduce data lag and resource waste, and improve the extraction efficiency. Based on the calculation of the priority estimation coefficient Ygxs, the system ranks the priority of each parameter output in real time in the power system and constructs a priority data column, so that key or high-priority parameters can be the first to perform feature extraction. This dynamic sorting mechanism can ensure that when the system processes data in real time, according to the importance and current state of the data, it reasonably allocates computing resources and processing time, thereby improving the response speed and real-time processing ability of the entire system. By constructing the priority data column, the system can avoid the processing bottleneck caused by excessive data volume or uneven distribution, ensure that the system can accurately and quickly obtain key features under high load or critical moments, and then optimize the operation and monitoring of the power system. Through the dynamic adjustment mechanism of the priority estimation coefficient, the system can automatically optimize the feature extraction order according to the real-time changes of the power state data. The system no longer relies on static or preset rules, but can adaptively adjust the feature extraction strategy in real time according to the actual situation. This flexibility enables the system to cope with different operating conditions and improves the accuracy of power system monitoring and analysis. In short, through this method, the distribution and similarity of power data can be accurately evaluated, and on this basis, the priority of different time periods and parameters can be sorted. Combining the similarity coefficient and the data distribution coefficient, after linear normalization processing, the system can dynamically adjust the data extraction order to ensure that important parameters are processed first. This mechanism not only improves the efficiency and accuracy of data extraction, but also enhances the processing ability and flexibility of the system during high-load periods, effectively improving the operation optimization ability of the power system and ensuring that the data value extraction efficiency always remains at a relatively high level.

[0087] Embodiment 5

[0088] Please refer to Figure 1 , specifically: The specific steps of S4 include:

[0089] S41. During the process of feature extraction according to the priority in S34, the freshness during the data extraction process is monitored in real time to obtain relevant extraction status data information, and the relevant extraction status data information includes the preprocessing time , the data generation time and the data extraction time ;

[0090] S42. According to the relevant extraction status data information, calculate to obtain the data freshness Sqxd, and the data freshness Sqxd is obtained through the following formula:

[0091] In the formula, is the preprocessing time, representing the time when multi-source heterogeneous data is preprocessed; is the data generation time, representing the moment when the data is generated by the sensor or other devices; is the data extraction time, representing the moment when the system extracts data from the power dataset; where, is used to measure the total time difference from data generation to preprocessing time, representing the delay from data generation to preprocessing time; is used to measure the time difference from data generation to extraction. The smaller it is, the more timely the data is extracted and the closer the data is to the latest state;

[0092] The data freshness measures the degree of delay from data generation to preprocessing and then to extraction. Through this calculation, we can evaluate: whether the data is quickly preprocessed after generation, and whether the data is quickly extracted after preprocessing, further avoiding excessive lag.

[0093] The above-mentioned preprocessing time , the data generation time , the data extraction time can be monitored through smart meters, data acquisition systems or edge computing devices.

[0094] In this embodiment, the method can dynamically evaluate the timeliness of data extraction by monitoring the relevant status information of data extraction in real time during the feature extraction process, including preprocessing time, data generation time, and data extraction time. In this way, the system can accurately calculate the data freshness Sqxd to ensure that the data can be quickly extracted and processed after generation. For scenarios with high real-time requirements such as power systems, monitoring the data freshness can effectively prevent the system from processing outdated data, thereby improving the accuracy of data analysis and the timeliness of decision-making, and timely detecting possible bottlenecks or delay links. If the time of a certain link is too long, the system can take optimization measures in time, such as adjusting the data extraction frequency or optimizing the preprocessing process, so as to reduce the data extraction delay and improve the real-time performance of data processing. By evaluating the data freshness, the system can identify which data has a delay in the extraction process, and then can dynamically adjust the data extraction strategy. Especially in the process of real-time monitoring of power status, the system can preferentially process the latest data according to the level of data freshness, reducing decision-making biases or system misjudgments caused by data lag. Through the accurate evaluation of data freshness, the system can ensure that the processing and response speed of power data is faster, further avoiding system scheduling delays caused by data lag. Especially in response to key scenarios such as power load fluctuations and equipment fault warnings, the system can make a rapid response based on the latest data information, significantly improving the stability and security of the power system.

[0095] Embodiment 6

[0096] Please refer to Figure 1 , specifically: The specific steps of S4 further include:

[0097] S42. Use the convolutional neural network technology to construct an efficiency prediction model, input the data freshness Sqxd and the priority estimation coefficient Ygxs into the efficiency prediction model, and after linear normalization processing, fit and output the extraction efficiency index Txzs from the efficiency prediction model. The extraction efficiency index Txzs is obtained through the following formula:

[0098] In the formula, K represents the number of parameters, k = 1, 2, 3,..., K, represents the data freshness of the k-th parameter, represents the average data freshness, represents the priority estimation coefficient of the k-th parameter, represents the average priority estimation coefficient, and are both weight values, P represents the second correction constant, and The specific values are set by the user according to the situation.

[0099] The specific steps of S5 include:

[0100] S51. By comparing and analyzing the pre-set safety threshold Q with the extraction efficiency index Txzs, comprehensively judge the extraction efficiency state of the parameters during the corresponding monitoring period when feature extraction is performed. The specific judgment content is as follows:

[0101] If the extraction efficiency index Txzs falls within the safety threshold Q, it is judged that the extraction efficiency of the parameters during the corresponding monitoring period when feature extraction is performed is in a normal state. At this time, continue the current feature extraction operation on the real-time output power state data in the power system;

[0102] S52. If the extraction efficiency index Txzs does not fall within the safety threshold Q, it is judged that the extraction efficiency of the parameters during the corresponding monitoring period when feature extraction is performed is not in a normal state. At this time, use caching technology to extract and process data in advance to ensure that the system can still maintain a high freshness during high-load periods, and re-evaluate the priorities of each parameter according to the actual application scenario. Some key parameters (such as load volatility, voltage stability, etc.) may be more important than other parameters in certain periods. Therefore, the priorities of each parameter can be dynamically adjusted to improve the data processing priority of important parameters, reduce irrelevant or redundant features, and in the convolutional neural network, apply multi-scale convolutional kernels to obtain data features at different resolutions, especially suitable for complex non-linear features in power data.

[0103] In this embodiment, the present invention constructs an efficiency prediction model using convolutional neural network technology. Taking data freshness Sqxd and priority estimation coefficient Ygxs as input variables, after linear normalization processing, it predicts the extraction efficiency of different parameters. The convolutional neural network can conduct in-depth analysis on the complex non-linear relationships in power data through multi-level feature extraction, thereby more accurately fitting the extraction efficiency index Txzs. Through this method, the system can more effectively evaluate the extraction efficiency of each parameter, ensuring the quality and speed of feature extraction. By comparing and analyzing the extraction efficiency index Txzs with a preset safety threshold, it can determine the extraction efficiency status of parameters in each monitoring period in real time. When the extraction efficiency index falls within the safety threshold, the system can continue with normal data extraction operations; while when the efficiency does not meet the standard, the system will automatically identify the anomaly and take corresponding optimization measures. This dynamic adjustment mechanism ensures that the system can promptly identify inefficient states during the processing, reducing potential bottleneck problems and guaranteeing the stability and efficiency of the power data extraction process. When the extraction efficiency does not meet the standard, the system will use caching technology to extract and process data in advance, thereby ensuring that the system can still maintain a high data freshness during high-load periods. This mechanism can effectively cope with data flow fluctuations and reduce data lag problems caused by excessive load. In addition, the system dynamically adjusts the priorities of each parameter, giving priority to processing certain key parameters (such as load fluctuations, voltage stability, etc.), improving the response speed to key data, thereby ensuring that important data can be processed quickly and enhancing the real-time performance of the overall system.

[0104] Embodiment 7

[0105] Please refer to Figure 2 , specifically: An optimization system for the value extraction efficiency of power data includes a data acquisition module, a preprocessing module, an extraction order analysis module, a comprehensive evaluation module, and a feedback module;

[0106] The data acquisition module is used to pre-obtain power status data monitored by several groups of monitoring devices at different times from each link within the power system to construct a multi-source heterogeneous data set;

[0107] The preprocessing module is used to perform preprocessing operations before feature extraction on the multi-source heterogeneous data set, conduct data cleaning, normalization processing, and binning processing on the multi-source heterogeneous data set to generate a power data set;

[0108] The extraction order analysis module is used to analyze the dependence between real-time parameters and corresponding parameters in the power dataset based on the power dataset and combined with the power status data real-time monitored by several groups of monitoring devices, so as to construct a similarity coefficient Xsxs. And in the time series, analyze the dispersion state of corresponding parameters in different monitoring periods to calculate the data distribution coefficient Xbxs of the corresponding parameters. By associating the similarity coefficient Xsxs with the data distribution coefficient Xbxs of the corresponding parameters, obtain a priority estimation coefficient Ygxs, and determine the parameter extraction order based on the priority estimation coefficient Ygxs;

[0109] The comprehensive evaluation module is used to analyze the timeliness in the data extraction process based on the extraction order of each parameter in S3 to obtain the data freshness Sqxd, and combine the trained efficiency prediction model and linear normalization processing to fit and output the extraction efficiency index Txzs;

[0110] The feedback module is used to preset a safety threshold Q, and by comparing and analyzing the extraction efficiency index Txzs with the safety threshold Q, comprehensively judge the extraction efficiency status of parameters during feature extraction in the corresponding monitoring period, and take corresponding optimization measures based on the judgment results.

[0111] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for optimizing the efficiency of power data value extraction, characterized by: The following steps are included: S1. Obtain power status data monitored by several groups of monitoring devices at different time periods from various links in the power system in advance to construct a multi-source heterogeneous data set; S2. Preprocessing the multi-source heterogeneous data set obtained in S1 before feature extraction, performing data cleaning, standardization, and binning on the multi-source heterogeneous data set to generate a power data set; S3. Based on the power data set and in combination with the power status data monitored in real time by several groups of monitoring devices, the dependency between the real-time parameters and the corresponding parameters in the power data set is analyzed to construct a similarity coefficient Xsxs. In the time series, the dispersion state of the corresponding parameters in different monitoring periods is analyzed to calculate the data distribution coefficient Xbxs of the corresponding parameters. By correlating the similarity coefficient Xsxs with the data distribution coefficient Xbxs of the corresponding parameters, a priority estimation coefficient Ygxs is obtained. Based on the priority estimation coefficient Ygxs, the parameter extraction order is determined. S4. Based on the extraction order of each parameter in S3, analyze the timeliness of the data extraction process to obtain the data freshness Sqxd, and combine the trained efficiency prediction model and linear normalization processing to fit the output extraction efficiency index Txzs; S5. Pre-set the safety threshold Q, and compare and analyze the extraction efficiency index Txzs with the safety threshold Q to comprehensively judge the extraction efficiency of the parameters in the corresponding monitoring period during feature extraction, and take corresponding optimization measures based on the judgment results.

2. The method for optimizing the efficiency of power data value extraction according to claim 1, characterized in that: The specific steps of S1 include: S11. Utilize several groups of monitoring devices to monitor power status data in different time periods to construct a multi-source heterogeneous data set, wherein the multi-source heterogeneous data set includes power status data monitored by different monitoring devices in different time periods.

3. The method for optimizing the efficiency of power data value extraction according to claim 2, characterized in that: The specific steps of S2 include: S21, preprocessing the multi-source heterogeneous data set obtained in S11, wherein the preprocessing includes data cleaning, standardization, and binning; S211, wherein, by dividing the internal parameters of the multi-source heterogeneous data set into equal-frequency bins, a plurality of groups of power data of monitoring periods are obtained; S22. Aggregate the power data of several groups of monitoring periods in time series to re-obtain the power data set.

4. The method for optimizing the efficiency of power data value extraction according to claim 3, characterized in that: The specific steps of S3 include: S31. Based on the power data set and in combination with the real-time power status data output from the power system, the dependency between the real-time parameters and the corresponding parameters in the power data set is analyzed, and each parameter in the power status data is mapped to a three-dimensional space through nonlinear mapping to obtain a similarity coefficient Xsxs of the corresponding parameter. The similarity coefficient Xsxs is obtained specifically in the following manner: Where n represents the monitoring period in the power data set; i = 1, 2, 3, ..., n; It is expressed as the contribution of the parameter of the i-th monitoring period in the power data set to the real-time output parameter of the current power system; It is represented as the parameter of the i-th monitoring period in the power data set; It is expressed as the real-time output parameter of the current power system; Expressed as kernel function; S311, the kernel function Obtained by the following formula: Where, Represented as parameters in the Gaussian kernel function; It is expressed as an exponential growth function, specifically , where e is the natural base; It is expressed as the Euclidean distance between the parameter of the i-th monitoring period in the power data set and the parameter output in real time in the current power system.

5. The method for optimizing the efficiency of power data value extraction according to claim 4, characterized in that: The specific steps of S3 also include: S32. In the time series, the dispersion state of the corresponding parameters in different monitoring periods is analyzed to calculate the data distribution coefficient Xbxs of the parameters in the corresponding monitoring period, which is obtained specifically in the following manner: Where, Expressed as the average parameter during the monitoring period.

6. The method for optimizing the efficiency of power data value extraction according to claim 5, characterized in that: The specific steps of S3 also include: S33, based on the similarity coefficient Xsxs obtained in S31 and S32 and the data distribution coefficient Xbxs of the corresponding parameter, after linear normalization processing, the corresponding data value is mapped to the interval The priority estimation coefficient Ygxs is calculated and obtained in the following way: Where, and are weight values, Expressed as the first correction constant, where and The specific value is set by the user according to the situation; S34. Based on the method of obtaining the priority estimation coefficient Ygxs in S33, the priority estimation coefficients Ygxs of several groups of parameters in the power status data output in real time in the power system are obtained respectively, and the priority estimation coefficients Ygxs of several groups of parameters are sorted by numerical size to construct a priority data column. According to the priority data column, features of several groups of parameters in the power status data output in real time in the power system are extracted according to priority.

7. The method for optimizing the efficiency of power data value extraction according to claim 6, characterized in that: The specific steps of S4 include: S41, in the process of feature extraction according to the priority in S34, real-time monitoring of the freshness of the data extraction process to obtain relevant extraction status data information, the relevant extraction status data information includes preprocessing time , data generation time and data extraction time ; S42. Calculate the data freshness Sqxd based on the relevant extracted state data information. The data freshness Sqxd is obtained by the following formula: Where, is the preprocessing time, which indicates the time required to preprocess multi-source heterogeneous data; The data generation time indicates the moment when the data is generated by the sensor or other equipment; is the data extraction time, which indicates the moment when the system extracts data from the power dataset.

8. The method for optimizing the efficiency of power data value extraction according to claim 7, characterized in that: The specific steps of S4 also include: S42. Construct an efficiency prediction model using convolutional neural network technology, input the data freshness Sqxd and the priority estimation coefficient Ygxs into the efficiency prediction model, and after linear normalization, fit the output extraction efficiency index Txzs from the efficiency prediction model. The extraction efficiency index Txzs is obtained by the following formula: Where K represents the number of parameters, k=1, 2, 3, ..., K, Represented as the data freshness of the kth parameter, Expressed as the average data freshness, Expressed as the priority estimation coefficient of the kth parameter, Expressed as the average priority estimation coefficient, and are all weight values, P represents the second correction constant, and The specific value is set by the user according to the situation.

9. The method for optimizing the efficiency of power data value extraction according to claim 1, characterized in that: The specific steps of S5 include: S51. By comparing and analyzing the preset safety threshold Q with the extraction efficiency index Txzs, a comprehensive judgment is made on the extraction efficiency of the parameters during the corresponding monitoring period during feature extraction. The specific judgment contents are as follows: If the extraction efficiency index Txzs falls within the safety threshold Q, it is determined that the extraction efficiency of the parameters during the corresponding monitoring period is normal during feature extraction. At this time, the feature extraction operation of the current power status data output in real time in the power system will continue. S52. If the extraction efficiency index Txzs does not fall within the safety threshold Q, it is determined that the extraction efficiency of the parameters during the corresponding monitoring period is not in a normal state during feature extraction. At this time, the cache technology will be used to extract and process the data in advance, and the priority of each parameter will be re-evaluated according to the actual application scenario.

10. A system for optimizing the efficiency of power data value extraction, for implementing the method for optimizing the efficiency of power data value extraction as claimed in any one of claims 1 to 9, characterized in that: It includes data acquisition module, pre-processing module, extraction sequence analysis module, comprehensive evaluation module and feedback module; The data acquisition module is used to obtain power status data monitored by several groups of monitoring devices at different time periods from various links in the power system in advance to construct a multi-source heterogeneous data set; The preprocessing module is used to perform preprocessing operations on the multi-source heterogeneous data sets before feature extraction, and perform data cleaning, standardization and binning on the multi-source heterogeneous data sets to generate power data sets; The extraction sequence analysis module is used to analyze the dependency between the real-time parameters and the corresponding parameters in the power data set based on the power data set and in combination with the power status data monitored in real time by several groups of monitoring devices to construct a similarity coefficient Xsxs, and analyze the dispersion state of the corresponding parameters in different monitoring periods in the time series to calculate the data distribution coefficient Xbxs of the corresponding parameters, and obtain the priority estimation coefficient Ygxs by correlating the similarity coefficient Xsxs with the data distribution coefficient Xbxs of the corresponding parameters. Based on the priority estimation coefficient Ygxs, the parameter extraction order is determined; The comprehensive evaluation module is used to analyze the timeliness of the data extraction process based on the extraction order of each parameter in S3 to obtain the data freshness Sqxd, and combine the trained efficiency prediction model and linear normalization processing to fit the output extraction efficiency index Txzs; The feedback module is used to pre-set the safety threshold Q, and compare and analyze the extraction efficiency index Txzs with the safety threshold Q to comprehensively judge the extraction efficiency of the parameters in the corresponding monitoring period during feature extraction, and take corresponding optimization measures based on the judgment results.

Citation Information

Patent Citations

  • Cloud data placement strategy management system introducing SaaS features

    CN116319815A

  • Abnormal electricity utilization detection system based on data enhancement and multi-dimensional feature extraction

    CN118503880A