Adaptive compression method for sparsity features based on electric power big data
Through methods such as time series decomposition and sparse representation, the compression strategy of power data is adaptively optimized, which solves the feature extraction and compression mapping problems caused by data sparseness in the power system, and realizes efficient data processing and monitoring.
Patent Information
- Application Number
- CN202510231460.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-07-22
AI Technical Summary
The dynamic sparse characteristics of data in power systems lead to inefficiency in traditional feature extraction algorithms and the inability to accurately capture key features, and it is difficult to establish adaptive optimization for the complex association between sparse features and compression mapping strategies.
Through time series decomposition, sparse representation, dictionary learning and compression perception methods, the perception matrix is adaptively optimized, compression strategies are dynamically adjusted, and adaptive compression and reconstruction of power data are realized in combination with multi-dimensional data correlation analysis.
Improve data integrity and availability, reduce storage and transmission costs, enhance grid data processing efficiency and grid operating status monitoring capabilities, and ensure data quality and prediction accuracy.
Smart Images

Figure CN120357908A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information technology, and in particular to an adaptive compression method based on the sparsity characteristics of electric power big data. Background Art
[0002] The data in the power system presents a dynamically changing sparsity feature, which brings great challenges to feature extraction and compression processing. Traditional fixed feature extraction algorithms are difficult to adapt to this dynamic change, resulting in low extraction efficiency and inability to accurately capture key features. At the same time, there is a complex relationship between the extracted sparse features and the compression mapping strategy. How to dynamically adjust the compression scheme according to the features becomes a key issue. Specifically, the power consumption, equipment status, meteorological data, etc. in the power system present dynamic sparsity in the temporal and spatial dimensions. This sparsity is not only reflected in the data itself, but also in the correlation between the data. For example, the power consumption in certain periods or regions may show highly sparse characteristics, while other periods or regions are relatively dense. Equipment status data may be relatively sparse during normal operation, but suddenly become dense when a fault occurs. The sparsity of meteorological data may vary with seasons and geographical locations. This multi-dimensional and multi-scale dynamic sparsity brings great challenges to feature extraction. How to design an adaptive feature extraction algorithm that can adjust the extraction strategy in real time according to the dynamic sparsity of the data while ensuring extraction efficiency and feature quality is a technical problem that needs to be solved urgently. In addition, there is a complex nonlinear relationship between the extracted sparse features and the subsequent compression mapping. How to establish a dynamic mapping model between the two and achieve adaptive optimization of the compression scheme is also an important research direction. Summary of the invention
[0003] The present invention provides an adaptive compression method based on the sparsity characteristics of power big data, which mainly includes:
[0004] Obtain power grid operation data at different sampling frequencies. According to the data missing patterns at different time scales, use time series decomposition method to decompose multi-scale data into high-frequency components and low-frequency components, perform sparse representation on high-frequency components, perform interpolation estimation on low-frequency components, and evaluate the data missing rate at each time scale to obtain preliminary completed multi-time granularity power data.
[0005] According to the physical quantity type and numerical distribution characteristics of power data, the completed multi-time granularity power data is divided into time windows, and the high-dimensional feature matrix reflecting the operating status of the equipment is extracted. The feature matrix is compressed by reducing the dimension using the frequency domain analysis method to obtain the key feature vectors at different time granularities and measure the sparsity of the data at each time scale.
[0006] The key feature vectors are screened by time-frequency analysis method to obtain the multi-granularity feature subset most relevant to load forecasting. According to the sparsity of data at each time scale and the spatial distribution characteristics of missing values, the fusion ratio of data with different time granularities is dynamically adjusted, the weights of each time scale are adaptively set, and the data quality and compression ratio are balanced.
[0007] For the selected multi-granularity feature subset, the sensing matrix is adaptively optimized by dictionary learning method, and a sparse regularization term is introduced to enhance the sparsity of data reconstruction. According to the missing pattern and context information of missing values, the mapping relationship between the compression ratio, reconstruction error and feature subset is established to determine the adaptive compression strategy.
[0008] Based on the adaptive compression strategy, the compressive sensing method is used to adaptively compress the multi-time granularity power data. The compression ratio is dynamically adjusted according to the reconstruction error feedback to balance data compression and reconstruction quality. The reconstructed multi-granularity power data is synchronized according to the timestamp accuracy and used as the input for subsequent correlation analysis.
[0009] Mine the association rules between power data with different time granularities, fuse multi-scale features, and construct a multi-dimensional data association analysis model for power grid operation state monitoring. According to the periodic characteristics of power data and the differences in data update cycles, construct a dynamic association rule update mechanism for multiple time scales to perform real-time association analysis under the change of power grid conditions.
[0010] The compressed and reconstructed multi-time granularity power data is evenly sliced and stored. The storage and computing resource allocation is dynamically optimized according to the data sparsity, missing value distribution and association strength. The reconstruction error is fed back to the feature extraction stage, and the feature matrix is iteratively optimized to incorporate the missing data estimation and multi-scale fusion error into the dynamic adjustment process.
[0011] The technical solution provided by the embodiment of the present invention may include the following beneficial effects:
[0012] The present invention discloses an adaptive compression method based on the sparsity characteristics of power big data. First, the power grid operation data with different sampling frequencies is processed through decomposition technology, and data processing and complementation are respectively carried out on the high-frequency and low-frequency components, improving the integrity and availability of the data, and solving the problem that traditional fixed feature extraction is difficult to adapt to dynamic changes and has low efficiency. Secondly, according to the physical quantity type and numerical distribution characteristics of power data, feature extraction and dimensionality reduction compression are effectively carried out on power data with multiple time granularities, reducing the complexity of data processing and the storage space requirement, and improving the resource utilization efficiency. In addition, by screening out the multi-granularity feature subsets most relevant to load forecasting, dynamically adjusting the data fusion ratio and weight, the data quality and prediction accuracy are improved, and it can adaptively adapt to different power grid operation states and data characteristics, with strong flexibility and adaptability. And by establishing an adaptive compression strategy, the balance between data compression and reconstruction quality is achieved, and at the same time, the sparsity of data reconstruction is enhanced. Finally, by synchronously processing the reconstructed power data with multiple time granularities and performing correlation analysis on power data with different time granularities, a multi-dimensional data correlation analysis model is constructed, which can monitor the power grid operation state in real time. In summary, the present invention significantly reduces the storage and transmission costs while ensuring data quality, improves the efficiency of power grid data processing and analysis, effectively solves the data missing problem, improves the data quality and compression efficiency, and enhances the power grid operation state monitoring ability. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 It is a flowchart of an adaptive compression method based on the sparsity characteristics of power big data according to the present invention.
[0014] Figure 2 It is a schematic diagram of an adaptive compression method based on the sparsity characteristics of power big data according to the present invention.
[0015] Figure 3 It is another schematic diagram of an adaptive compression method based on the sparsity characteristics of power big data according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0016] The technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0017] Such as Figures 1-3 , a specific adaptive compression method based on the sparsity characteristics of power big data in this embodiment may specifically include:
[0018] Step S101: Obtain power grid operation data with different sampling frequencies. For the data missing patterns at different time scales, use the time series decomposition method to decompose the multi-scale data into high-frequency components and low-frequency components. Perform sparse representation on the high-frequency components, perform interpolation estimation on the low-frequency components, and evaluate the data missing rate at each time scale to obtain the initially completed power data with multiple time granularities.
[0019] Obtain power grid operation data with multiple sampling frequencies, where the sampling frequencies include second-level, minute-level, and hour-level; calculate the data missing rate based on the power grid operation data, and analyze the missing pattern of the data missing rate; if the data missing rate exceeds a preset threshold, mark the power grid operation data. Perform seasonal decomposition on the marked power grid operation data to obtain a trend component, a seasonal component, and a residual component; among them, the trend component is fitted by polynomial regression, the seasonal component is expanded by Fourier series, and the residual component is modeled by autoregressive moving average. Select corresponding processing methods according to the characteristics of the trend component, seasonal component, and residual component; perform locally weighted regression smoothing on the trend component, perform periodic interpolation on the seasonal component, and estimate the residual component using Kalman filtering. Combine the processed trend component, seasonal component, and residual component; perform inverse transformation on the combined data to obtain the restored data at the original time scale; the restored data includes second-level, minute-level, and hour-level data. Calculate the root mean square error and mean absolute percentage error of the restored data; if the root mean square error or mean absolute percentage error exceeds a preset threshold, adjust the seasonal decomposition parameters, regression coefficients, and the number of Fourier series terms; re-execute the data processing process until the accuracy requirement is met or the maximum number of iterations is reached.
[0020] Specifically, obtain power grid operation data at multiple sampling frequencies, including second-level, minute-level, and hour-level data. For each time scale, calculate the data missing rate and analyze the missing patterns, such as random missing or consecutive missing. Judge the degree of data missing according to the preset threshold, and mark the data at the time scale exceeding the threshold. Perform seasonal decomposition on the marked multi-scale data, and divide it into trend component, seasonal component, and residual component. Use polynomial regression to fit the trend component, use Fourier series expansion for the seasonal component, and model the residual component through autoregressive moving average. Select appropriate processing methods according to the characteristics of different components, perform locally weighted regression smoothing on the trend component, perform periodic interpolation on the seasonal component, and estimate the residual component using the Kalman filter. Combine the processed trend component, seasonal component, and residual component. Perform inverse transformation on the combined data to restore it to the data under the original time scale. Evaluate the completion effect through the root mean square error and mean absolute percentage error. The calculation formulas are: root mean square error = √(Σ(actual value - predicted value) / n), mean absolute percentage error = Σ|actual value - predicted value| / actual value × 100% / n, where n is the number of data points. If the error value exceeds the preset threshold, adjust the seasonal decomposition parameters, regression coefficients, and the number of Fourier series terms, and re-perform the data processing process. Repeat until the accuracy requirement is met or the maximum number of iterations, which is 10 times, is reached. Finally, output the completion result of multi-time granularity power data that meets the accuracy requirement. In the collection of power grid operation data, obtain second-level, minute-level, and hour-level data, and the sampling frequencies are 100Hz, 1Hz, and 1 / 3600Hz respectively. Calculate the data missing rate of each time scale through the sliding window method, and the window size is 24 times the corresponding time scale. Analysis shows that the second-level data shows random missing, and the missing rate is 5%; the minute-level data has consecutive missing, and the missing rate is 8%; the hour-level data missing rate is 2%. Set the missing rate threshold to 6%, and mark the second-level and minute-level data. Apply the X-12-ARIMA seasonal decomposition algorithm to the marked data to decompose the data into trend, season, and irregular components. Use cubic polynomial regression to fit the trend component, use 6th-order Fourier series expansion for the seasonal component, and model the irregular component through ARIMA(2,1,2). Apply the LOESS locally weighted regression scatterplot smoothing method to smooth the trend component, and set the window size to 10% of the data length. Use the spline interpolation method to perform periodic interpolation on the seasonal component, and the number of interpolation points is 2 times the number of original data points. Use the Kalman filter to estimate the irregular component, where the observation noise covariance matrix R and the process noise covariance matrix Q are obtained through maximum likelihood estimation. Add the three processed components to obtain the combined data, and then restore it to the original time scale through inverse transformation. Calculate the root mean square error and mean absolute percentage error, and the results are 0.015 and 1.8% respectively. Since the error values are lower than the preset thresholds of 2% and 3%, there is no need to adjust the parameters and repeat the execution.
[0021] Step S102, according to the physical quantity type and numerical distribution characteristics of the power data, the completed multi-time granularity power data is divided into time windows, and a high-dimensional feature matrix reflecting the operating status of the equipment is extracted. The feature matrix is compressed by dimensionality reduction using a frequency domain analysis method to obtain key feature vectors at different time granularities, and the sparsity of data at each time scale is measured.
[0022] Receive multi-time granularity power data, and normalize the multi-time granularity power data according to the voltage, current and power physical quantity types; obtain the normalized multi-time granularity power data, and use the sliding window method to divide the time window, and the time window size is set to an integer multiple of the data period; calculate statistical features for the data in the time window, and the statistical features include mean, variance, kurtosis and skewness to obtain a high-dimensional feature matrix reflecting the operating status of the equipment; apply short-time Fourier transform to the high-dimensional feature matrix to obtain time-frequency domain features; set an energy threshold, retain the frequency components whose cumulative energy ratio exceeds the preset threshold, and eliminate other components to obtain dimensionality reduction compression The feature matrix after dimensionality reduction and compression is obtained; principal component analysis is applied to the feature matrix after dimensionality reduction and compression, and the principal component whose cumulative contribution rate reaches a preset threshold is selected as the key feature vector; the Gini coefficient of the key feature vector is calculated as a sparse measurement indicator; if the Gini coefficient is less than a first preset threshold, sparse processing is performed by increasing the time window or reducing the sampling rate; if the Gini coefficient is greater than a second preset threshold, encryption processing is performed by reducing the time window or increasing the sampling rate; the feature extraction and dimensionality reduction and compression process is repeated until the Gini coefficient falls between the first preset threshold and the second preset threshold, so as to obtain a feature representation that takes into account both information retention and computational efficiency.
[0023] Specifically, according to the types of physical quantities such as voltage, current, and power, the complemented multi-time granularity power data is normalized. The sliding window method is used to divide the time window, and the window size is set to an integer multiple of the data period. Statistical features including mean, variance, kurtosis, and skewness are calculated for the data within each window to form a high-dimensional feature matrix reflecting the operating state of the equipment. The short-time Fourier transform is applied to the feature matrix to obtain time-frequency domain features. An energy threshold is set to retain the frequency components with a cumulative energy ratio exceeding 90%, and other components are removed to achieve dimensionality reduction and compression. Through the energy distribution characteristics, the periodicity and stability of data at different time granularities are judged. The principal component analysis is applied to the dimensionality-reduced feature matrix, and the principal components with a cumulative contribution rate reaching 95% are selected as the key feature vectors. The number of principal components at each time granularity is calculated as the feature complexity index. By comparing the number of principal components at different time granularities, the information richness of data at each time scale is quantified. The Gini coefficient of the key feature vectors is calculated as the sparsity metric index. If the Gini coefficient is less than 0.6, it is considered that the data is too dense, and sparsification is performed by increasing the time window or reducing the sampling rate. If the Gini coefficient is greater than 0.9, it is considered that the data is too sparse, and encryption is performed by reducing the time window or increasing the sampling rate. The feature extraction and dimensionality reduction and compression processes are repeatedly executed until the Gini coefficient falls between 0.6 and 0.9 to obtain a feature representation that balances information retention and computational efficiency. In power data processing, first, the physical quantities such as voltage, current, and power are normalized using min-max normalization, mapping the values to the range of 0-1. The sliding window method is used to divide the time window. For a 50Hz AC power system, the window size is set to 0.02 seconds per cycle. The mean, variance, kurtosis, and skewness are calculated within each window to form a high-dimensional feature matrix. For example, for 1-hour data, there are 180,000 windows in total, and 4 features are extracted from each window, resulting in a feature matrix of 180,000×4. Subsequently, the short-time Fourier transform is applied to the feature matrix using a Hanning window function with a 50% window overlap rate to obtain time-frequency domain features. An energy threshold of 90% is set to retain the main frequency components and achieve dimensionality reduction. By comparing the energy distributions of data at different time granularities, it is found that the energy ratio of hourly data in the low-frequency band of 0-10Hz reaches 95%, while the energy distribution of second-level data is more uniform in the range of 0-25Hz. The principal component analysis is performed on the dimensionality-reduced feature matrix, and the principal components with a cumulative contribution rate of 95% are selected. The results show that only 3 principal components are required for hourly data to reach a 95% contribution rate, while 7 principal components are required for second-level data, reflecting the information complexity of data at different time scales. Finally, the Gini coefficient of the key feature vectors is calculated as the sparsity metric. The Gini coefficient of hourly data is 0.85, which is in the ideal range; while the Gini coefficient of second-level data is 0.55, and sparsification is required by increasing the window to 0.1 seconds.Repeat the feature extraction process until the Gini coefficients at all time scales fall between 0.6 and 0.9, and finally obtain a feature representation that not only retains key information but also has high computational efficiency.
[0024] Step S103, screen out the multi-granularity feature subset most relevant to load forecasting from the key feature vectors through time-frequency analysis methods. According to the sparsity of data at each time scale and the spatial distribution characteristics of missing values, dynamically adjust the fusion ratio of data with different time granularities, adaptively set the weights of each time scale, and balance data quality and compression rate.
[0025] Apply the short-time Fourier transform and continuous wavelet transform to the key feature vectors to obtain the time-frequency features; calculate the mutual information between each frequency band and the load forecasting result according to the time-frequency features, and select the top N frequency bands with the highest mutual information values as the multi-granularity feature subset; if there are missing values in the multi-granularity feature subset, use the local outlier factor algorithm to detect the spatial distribution characteristics of the missing values; calculate the local density and relative density of the missing values through the spatial distribution characteristics, and construct a data quality evaluation function in the form of a weighted sum; design a Mamdani-type fuzzy controller based on the output value of the data quality evaluation function, where the input of the fuzzy controller is the data quality score and the output is the adjustment amount of the fusion ratio; calculate the conditional entropy of data at each time scale using the adjustment amount of the fusion ratio to obtain an information quantity index; construct an objective programming model by combining the information quantity index and the data quality score; solve the objective programming model to obtain the adaptive weights of each time scale.
[0026] Specifically, apply the short-time Fourier transform and continuous wavelet transform to the key feature vector to obtain time-frequency features. Calculate the mutual information between each frequency band and the load prediction result, and select the top N frequency bands with the highest mutual information values as the multi-granularity feature subset. The value of N is determined by cross-validation to achieve preliminary dimensionality reduction while retaining key information. Use the local outlier factor algorithm to detect the spatial distribution characteristics of missing values, and calculate the local density and relative density of the missing values. Combine the aforementioned sparsity index to construct a data quality evaluation function in the form of a weighted sum. The weight coefficients are determined by the particle swarm optimization algorithm, so that the output value of the evaluation function is between 0 and 1, and the larger the value, the higher the data quality. Based on the output of the data quality evaluation function, design a Mamdani-type fuzzy controller. The input is the data quality score, and the output is the fusion ratio adjustment amount. Use the trapezoidal membership function, set 5 fuzzy rules, such as reducing the ratio if the quality is low. Perform defuzzification by the centroid method to obtain the specific fusion ratio adjustment value. Set the upper and lower limits of the fusion ratio to ensure that the minimum contribution of each time-scale data is not less than 10%. Calculate the conditional entropy of each time-scale data as the information quantity index. Combine the data quality score and the fusion ratio to construct an objective programming model. Set the weight coefficients of data quality and compression ratio to 0.6 and 0.4 respectively. Solve the objective programming model to obtain the adaptive weights of each time scale. The weight values are limited between 0.05 and 0.5 to ensure that each time-scale data makes a certain contribution. In power load forecasting, first perform time-frequency analysis on the key feature vector. Adopt the short-time Fourier transform, set the window size to 128 sampling points, and the overlap rate to 50% to obtain the time-frequency spectrogram. At the same time, use the continuous wavelet transform, select the Morlet wavelet as the mother wavelet, and set the scale range to 1 to 64 to obtain the time-frequency features. Calculate the mutual information between each frequency band and the historical load data, and select the top 10 frequency bands with the highest mutual information values as the multi-granularity feature subset. Determine the best N value as 10 through 5-fold cross-validation, reducing the feature dimension by 60% while retaining 90% of the information. Subsequently, use the local outlier factor algorithm to detect the distribution of missing values, set the neighborhood size k = 5, and calculate the local density and relative density of the missing values. Combine the previously calculated Gini coefficient as the sparsity index to construct the data quality evaluation function Q = 0.4*(1 - local density) + 0.3*relative density + 0.3*(1 - Gini coefficient). Use the particle swarm optimization algorithm, with a population size of 50 and 100 iterations, to determine the weight coefficients. Based on the output of the evaluation function, design a Mamdani-type fuzzy controller with 5x5 rules, with the input being the data quality score and the output being the fusion ratio adjustment amount. Use the trapezoidal membership function and perform defuzzification by the centroid method to obtain the fusion ratio adjustment value. Set the fusion ratio range to [0.1, 0.5] to ensure that each time-scale data accounts for at least 10%.Finally, calculate the conditional entropy of each time-scale data as the information quantity index, construct the objective programming model max(0.6 data quality + 0.4 compression rate), with the constraint that the sum of weights is 1 and the range of a single weight is [0.05, 0.5]. Solve it using the interior point method to obtain the adaptive weights of each time scale and achieve the balance between data quality and compression rate.
[0027] Step S104: For the selected multi-granularity feature subsets, adaptively optimize the sensing matrix through the dictionary learning method, introduce a sparse regularization term to enhance the sparsity of data reconstruction, and establish the mapping relationship among the compression rate, reconstruction error, and feature subsets according to the missing patterns and context information of the missing values to determine the adaptive compression strategy.
[0028] Apply the K-SVD algorithm to the selected multi-granularity feature subsets for dictionary learning; according to the dictionary learning results, use the exponentially weighted moving average method to calculate the time series trend of the missing values, and the time series trend calculation uses a 24-hour window size; obtain the mean, variance, and skewness of the time series trend and context data, and construct a 10-dimensional missing pattern description vector; use a random forest regressor to establish the mapping relationship among the compression rate, reconstruction error, and feature subsets, where the input of the random forest regressor is the feature subset and the missing pattern description vector, and the output is the compression rate and reconstruction error; according to the output of the random forest regressor, design a fuzzy control system, where the input variables of the fuzzy control system are the predicted compression rate and reconstruction error, and the output variable is the compression parameter adjustment amount; if the compression rate or reconstruction error exceeds the preset range, then monitor the compression effect in real time and trigger the re-evaluation and adjustment of the compression strategy.
[0029] Specifically, apply the K-SVD algorithm to the selected multi-granularity feature subsets for dictionary learning. Set the dictionary size to 1.5 times the feature dimension and the sparsity threshold to 0.1. Introduce the L1 norm as the sparse regularization term, and determine the regularization coefficient through cross-validation. Adaptively update the sensing matrix during the iterative optimization process to enhance the sparsity of data reconstruction. Use the exponentially weighted moving average method to calculate the time series trend of missing values, with a window size of 24 hours. Combine the mean, variance, and skewness of the context data to construct a 10-dimensional missing pattern description vector. Adopt the sliding window method to analyze the time series pattern of missing values, calculate the autocorrelation coefficient and partial autocorrelation coefficient of missing values, and further enrich the missing pattern description. Use a random forest regressor to establish the mapping relationship between the compression ratio, reconstruction error, and feature subsets. The input is the feature subset and the missing pattern description vector, and the output is the compression ratio and reconstruction error. Use grid search to optimize hyperparameters such as the number of trees and the maximum depth. Through feature importance analysis, identify the feature subsets that have the greatest impact on the compression ratio and reconstruction error. Based on the output of the random forest regressor, design a fuzzy control system. The input variables are the predicted compression ratio and reconstruction error, and the output variable is the adjustment amount of compression parameters. Set 5 fuzzy rules, such as increasing the compression intensity if the error is low and the compression ratio is high. According to the output of the fuzzy control system, dynamically adjust the compression parameters to achieve an adaptive compression strategy. Monitor the compression effect in real time. When the compression ratio or reconstruction error exceeds the preset range, trigger a re-evaluation and adjustment of the compression strategy. During the power data compression process, first apply the K-SVD algorithm to the multi-granularity feature subsets for dictionary learning. Assume the feature dimension is 100, set the dictionary size to 150, and the sparsity threshold to 0.1. Determine the L1 regularization coefficient to be 0.01 through 5-fold cross-validation. After 50 iterations, obtain the optimized sensing matrix, and the sparsity of data reconstruction is increased by 20%. Subsequently, use the exponentially weighted moving average method to analyze the missing value trend, set the weight factor to 0.9, and the window size to 24 hours. Combine the statistical features of the context data to construct a 10-dimensional missing pattern description vector, including 3 each of the mean, variance, and skewness, plus the autocorrelation coefficient (lag = 1) and partial autocorrelation coefficient (lag = 1). Use a random forest regressor to establish the mapping relationship, set 100 decision trees, and the maximum depth to 10. Optimize the hyperparameters through grid search, and finally select the number of trees as 120 and the maximum depth as 8. Feature importance analysis shows that the autocorrelation coefficient in the missing pattern description vector has the greatest impact on the compression ratio, with a contribution of 25%. Based on the random forest output, design a 5x5 rule fuzzy control system. The input variables, the compression ratio and reconstruction error, are respectively divided into 5 fuzzy sets, and the output variable, the adjustment amount of compression parameters, is divided into 5 fuzzy sets. Set the rule that if the error is very low and the compression ratio is very high, then greatly increase the compression intensity. Use the centroid method for defuzzification to obtain the specific adjustment amount of compression parameters.Real-time monitoring shows that the time compression ratio remains between 0.6 and 0.8 for 95% of the time, and the reconstruction error is controlled within 5%. When the compression ratio is lower than 0.5 or the reconstruction error exceeds 8%, the compression strategy is triggered for re-evaluation, and the adjustment period is on average once every 4 hours.
[0030] Step S105, based on the adaptive compression strategy, the compressive sensing method is used to adaptively compress the multi-time granularity power data, the compression ratio is dynamically adjusted according to the reconstruction error feedback, the data compression and the reconstruction quality are balanced, and the reconstructed multi-granularity power data is synchronized according to the timestamp accuracy and used as the input for subsequent correlation analysis.
[0031] The compressive sensing data is obtained according to the multi-time granularity power data, and the compressive sensing data is obtained by processing the multi-time granularity power data with a sparse representation algorithm; the iterative hard threshold algorithm is used to process the compressive sensing data to obtain the reconstructed data, and the root mean square error is calculated according to the reconstructed data; if the root mean square error is greater than the preset error threshold, the gradient descent method with an adaptive step size is used to adjust the compression ratio, and the gradient descent method with an adaptive step size dynamically adjusts the step size according to the change of the root mean square error; the reconstructed data is processed by the timestamp accuracy unification algorithm to obtain the multi-granularity power data with unified accuracy, and the multi-granularity power data with unified accuracy includes second-level data, minute-level data, and hour-level data; the dynamic time warping algorithm is used to process the multi-granularity power data with unified accuracy to obtain the aligned multi-granularity power data, and the dynamic time warping algorithm includes calculating the distance matrix between time series, constructing the optimal alignment path by using the dynamic programming method, and mapping the data with different time granularities to the unified time axis according to the optimal alignment path.
[0032] Specifically, according to the adaptive compression strategy, the sparse representation algorithm is applied to the multi-time granularity power data for compressive sensing. Using the pre-optimized sensing matrix and sparse dictionary, based on the data characteristics and historical compression effects, the initial compression ratio is determined through Bayesian optimization. The upper and lower limits of the compression ratio are set to 0.3 and 0.8 respectively, and optimization is carried out within this range to obtain the compressed data representation. The iterative hard threshold algorithm is used to reconstruct the compressed data, and the root mean square error is calculated as the reconstruction error index. The error tolerance is set to 5% of the standard deviation of the original data. The gradient descent method with an adaptive step size is used to adjust the compression ratio, and the initial value of the step size is set to 0.05, which is dynamically adjusted according to the error change in each iteration. If the current error is less than the error of the previous iteration, the step size is increased; otherwise, it is decreased. Repeat this process until a balance point is reached or 50 iterations are completed. The timestamps of the reconstructed multi-granularity power data are unified in precision, and the timestamps of second-level data, minute-level data, and hour-level data are uniformly converted to millisecond-level precision. For low-precision data, the cubic spline interpolation method is used to expand the time points to maintain the smoothness and continuity of the data. When interpolating, the slope of adjacent data points is considered to ensure that the interpolation result is consistent with the trend of the original data. The dynamic time warping algorithm is used to align the data of different time granularities. The Euclidean distance is used to calculate the distance matrix between time series, and the dynamic programming method is used to construct the optimal alignment path. The path slope constraint is set between 0.5 and 2 to ensure the rationality of the alignment. Through the alignment path, the data of different time granularities are mapped onto a unified time axis to achieve the synchronization of multi-granularity power data, which is used as the input for subsequent correlation analysis. During the power data compression process, first, the sparse representation algorithm is applied to the multi-time granularity data for compressive sensing. A 100x200-dimensional sensing matrix and a 200x500-dimensional sparse dictionary optimized previously are used, and the initial compression ratio is determined to be 0.65 through Bayesian optimization. During the optimization process, the upper limit of the compression ratio is set to 0.8 and the lower limit is set to 0.3. The Gaussian process regression model is used, and the optimal compression ratio is obtained after 30 iterations. Subsequently, the iterative hard threshold algorithm is used for data reconstruction, with the maximum number of iterations set to 100 and the convergence threshold set to 1e-6. The calculated root mean square error is 0.08, which is higher than the preset error tolerance of 5% of the standard deviation of the original data (0.05). The gradient descent method with an adaptive step size is started to adjust the compression ratio, and the initial step size is set to 0.05. At the 15th iteration, the error drops to 0.048, which is lower than the tolerance, and the compression ratio is 0.72 at this time. The timestamps of the reconstructed data are unified in precision, and the 3600000ms precision of the hour-level data and the 60000ms precision of the minute-level data are converted to millisecond-level precision. For the hour-level data, the cubic spline interpolation method is used for expansion, and 3599 data points are inserted within each hour to ensure the continuity of the interpolation result with the values and first-order derivatives of the original data at the whole points. Finally, the dynamic time warping algorithm is applied to align the data of different time granularities.Construct a 100x100 distance matrix, use the Sakoe-Chiba bandwidth with a constraint of 10, and solve for the optimal path through dynamic programming. Set the path slope constraint to be between 0.5 and 2. The obtained alignment results show that the hourly data and the minute-level data are synchronized at 99.5% of the time points, providing a basis for subsequent correlation analysis.
[0033] Step S106, mine the association rules between power data of different time granularities, fuse multi-scale features, construct a multi-dimensional data association analysis model for power grid operation state monitoring, and construct a multi-time scale dynamic association rule update mechanism according to the periodic characteristics of power data and the differences in data update cycles to perform real-time association analysis under changes in power grid conditions.
[0034] Use the FP-Growth algorithm to mine the association rules for the power data of different time granularities, and obtain the optimal support and confidence thresholds through grid search. According to the optimal support and confidence thresholds, extract multi-scale features from the power data, and adaptively select the best wavelet basis and decomposition level using wavelet packet transform. For the multi-scale features, analyze the time-frequency spectrum characteristics using short-time Fourier transform, and identify the main periodic components using the peak detection algorithm. According to the main periodic components, design an adaptive sliding window mechanism, and the window size is dynamically adjusted according to the main period. For the data within the adaptive sliding window, construct a real-time association analysis model based on gated recurrent units, and the input layer of the model includes multi-time scale features and association rule strengths. If the model output indicates a change in the current power grid condition, update the model parameters through online learning to perform real-time association analysis under changes in power grid conditions.
[0035] Specifically, use the FP-Growth algorithm to mine the association rules for power data of different time granularities, and determine the optimal support and confidence thresholds through grid search. Generate association rule sets for second-level, minute-level, and hourly data respectively, and then perform hierarchical clustering to merge similar rules to obtain cross-time scale association rules. Remove redundant rules through rule pruning to improve the compactness and effectiveness of the rule set.
[0036] Use wavelet packet transform to extract multi-scale features from power data, and adaptively select the best wavelet basis and decomposition level. Evaluate the information content of features through the information entropy criterion, and select the feature subset with the largest information content. Combine the association rules and multi-scale features to construct a multi-dimensional data association analysis model based on random forest for power grid operation state monitoring.
[0037] The short-time Fourier transform is used to analyze the time-frequency spectral characteristics of the data, and the peak detection algorithm is utilized to identify the main periodic components. An adaptive sliding window mechanism is designed, and the window size is dynamically adjusted according to the main period. Within each window, the incremental learning algorithm is used to update the association rules. Different update periods are set for data with different periods to construct a dynamic association rule update mechanism.
[0038] A real-time association analysis model based on gated recurrent units is constructed. The input layer includes features at multiple time scales and the strength of association rules. The Adam optimizer is used for online parameter update, and the forgetting gate mechanism is set to balance the weights of historical information and new data. The model output represents the current grid operating condition, and the model parameters are continuously updated through online learning to achieve real-time association analysis under grid condition changes, providing support for grid operating state monitoring.
[0039] In grid operating state monitoring, first, the FP-Growth algorithm is applied to mine association rules for power data with multiple time granularities. The optimal support threshold of 0.05 and confidence threshold of 0.7 are determined through grid search. Association rule sets are generated for second-level, minute-level, and hour-level data respectively, obtaining a total of 2,000 original rules. The Ward hierarchical clustering method is adopted, with the distance threshold set to 0.3, to merge similar rules, and finally 500 cross-time-scale association rules are obtained. Subsequently, wavelet packet transform is performed using the Daubechies-4 wavelet basis, and a 3-layer decomposition structure is adaptively selected. The Shannon entropy criterion is used to evaluate the feature information content, and the top 20 features with the largest information content are selected. Combining the association rules and multi-scale features, a random forest model containing 100 decision trees is constructed. The model achieves an accuracy of 95% on the test set. Then, the 512-point short-time Fourier transform, Hanning window, and 50% overlap rate are used to analyze the time-frequency spectrum of the data. Through the peak detection algorithm, three main periodic components of 24 hours, 7 days, and 30 days are identified. An adaptive sliding window is designed, and the window size is dynamically adjusted between 1 hour and 24 hours. Within each window, the Incremental FP-Growth algorithm is used to update the association rules, and the update periods are 5 minutes, 1 hour, and 6 hours respectively. Finally, a two-layer GRU network containing 128 hidden units is constructed. The input layer includes 20 multi-scale features and 500 association rule strengths. The Adam optimizer is used with a learning rate of 0.001 for online parameter update. The forgetting gate threshold is set to 0.3 to balance historical information and new data. The model outputs the current grid operating condition assessment result every 30 seconds, and in the simulated grid fault scenario, an average response time of 0.5 seconds is achieved, providing real-time and accurate support for grid operating state monitoring.
[0040] Step S107: Perform balanced sharding storage on the compressed and reconstructed multi-time granularity power data, dynamically optimize the storage and computing resource allocation according to the data sparsity degree, missing value distribution and correlation strength, feedback the reconstruction error to the feature extraction stage, iteratively optimize the feature matrix, and incorporate the missing data estimation and multi-scale fusion error into the dynamic adjustment process.
[0041] Use the Z-order curve to perform multi-dimensional indexing on the compressed and reconstructed multi-time granularity power data, and map the data to a one-dimensional space. Perform sharding storage on the one-dimensional space mapped data using the consistent hashing algorithm, and achieve load balancing by dynamically adjusting the number of virtual nodes. For the sharded stored data, construct a resource allocation model, and use the particle swarm optimization algorithm to dynamically adjust the storage and computing resources. Obtain the optimization result of the resource allocation model, design an error feedback mechanism, and use the adaptive moment estimation algorithm to update the feature matrix. If the update of the feature matrix is completed, construct a multi-task learning framework based on the attention mechanism, and use the missing data estimation and multi-scale fusion error as auxiliary tasks. The multi-task learning framework uses a soft parameter sharing strategy to allow information interaction between tasks. For the multi-task learning framework, use the method of uncertainty weighting to dynamically adjust the task weights, and introduce gradient normalization to alleviate the gradient imbalance problem. Achieve the global optimization of the feature matrix through the multi-task learning framework, dynamically adjust the weights of each task, and adapt to the changing characteristics of power data.
[0042] Specifically, the Z-order curve is adopted to perform multi-dimensional indexing on the compressed and reconstructed power data with multiple time granularities, mapping the data into a one-dimensional space. The consistent hashing algorithm is used for sharding storage, and load balancing is achieved by dynamically adjusting the number of virtual nodes. A Bloom filter is introduced to quickly determine the data storage location and improve the access efficiency. The number of virtual nodes is set according to the data sparsity, and locally sensitive hashing is used to store strongly correlated data in adjacent nodes. A resource allocation model is constructed, and the particle swarm optimization algorithm is used to dynamically adjust the storage and computing resources. The data sparsity is quantified by the Gini coefficient, the missing value distribution is represented by the entropy value, and the correlation strength is calculated by mutual information. Multiple objective constraints are set, including the response time and resource utilization rate. During the optimization process, the data characteristics are considered to dynamically configure the storage space and computing nodes, balancing the resource utilization rate and data processing efficiency. An error feedback mechanism is designed, and the adaptive moment estimation algorithm is used to update the feature matrix. Elastic net regularization is introduced to balance the L1 and L2 penalty terms, enhancing the sparsity and robustness of the features. The alternating direction multiplier method is used to solve the optimization problem, improving the convergence speed. The reconstruction error is transmitted to the feature extraction stage through this mechanism, continuously optimizing the feature matrix and improving the accuracy of feature extraction. A multi-task learning framework based on the attention mechanism is constructed, taking the missing data estimation and multi-scale fusion error as auxiliary tasks. The soft parameter sharing strategy is used to allow information interaction between tasks. The uncertainty weighted method is adopted to dynamically adjust the task weights, and the gradient normalization technique is introduced to alleviate the gradient imbalance problem. The generalization ability of the model is improved through shared representation learning, achieving the global optimization of the feature matrix, dynamically adjusting the weights of each task, and adapting to the changing characteristics of power data. In the power data management system, first, the 18-bit Z-order curve is used to index the multi-time granularity data, mapping the three-dimensional spatio-temporal data into a one-dimensional space. The consistent hashing algorithm is used for sharding, and 200 virtual nodes are set to achieve the uniform distribution of data. A 512-bit Bloom filter is introduced, and the false positive rate is set to 0.01 to quickly locate the data. Through locally sensitive hashing, the number of hash buckets is set to 1000 to achieve the adjacent storage of related data. Subsequently, a resource allocation model is constructed, and the particle swarm algorithm with 30 particles is used to optimize the storage and computing resources. The data sparsity is quantified by the Gini coefficient, and the threshold is set to 0.6; the entropy value of the missing value distribution is restricted between 0.3 and 0.7; the correlation strength is calculated by mutual information, and the threshold is 0.5. The upper limit of the response time is set to 100 ms, and the lower limit of the resource utilization rate is set to 85% as the constraint conditions. Then, an error feedback mechanism is designed, and the adaptive moment estimation algorithm is used to update the feature matrix, with the initial learning rate set to 0.01. Elastic net regularization is introduced, and α is set to 0.5 to balance the L1 and L2 penalties. The alternating direction multiplier method is used for solution, with the maximum number of iterations set to 500 and the convergence threshold set to 1e-6. Finally, a 5-layer attention multi-task learning framework is constructed, with the main task being feature extraction, and the auxiliary tasks including missing data estimation and multi-scale fusion error calculation.The soft parameter sharing rate is set to 0.3, and information interaction between tasks is allowed. The uncertainty weighting method is adopted, and the weight update period is 100 batches. A gradient clipping with a threshold of 0.1 is introduced to alleviate gradient imbalance.
[0043] The above embodiments are only one of the preferred embodiments of the present invention and should not be used to limit the protection scope of the present invention. Any meaningless modifications or polishings made on the main design concept and spirit of the present invention, as long as the technical problems solved are still consistent with those of the present invention, should be included in the protection scope of the present invention.
Claims
1. An adaptive compression method based on the sparsity characteristics of power big data, characterized in that, The method includes: Obtain power grid operation data at different sampling frequencies. For data missing patterns at different time scales, use time series decomposition method to decompose multi-scale data into high-frequency components and low-frequency components, perform sparse representation on the high-frequency components, perform interpolation estimation on the low-frequency components, and evaluate the data missing rate at each time scale to obtain preliminarily completed multi-time granularity power data; According to the physical quantity type and numerical distribution characteristics of the power data, divide the completed multi-time granularity power data into time windows, extract a high-dimensional feature matrix reflecting the equipment operation status, use frequency domain analysis method to reduce the dimension and compress the feature matrix, obtain key feature vectors at different time granularities, and measure the sparsity of data at each time scale; Through time-frequency analysis method, screen out the multi-granularity feature subset most relevant to load forecasting in the key feature vectors. According to the sparsity of data at each time scale and the spatial distribution characteristics of missing values, dynamically adjust the fusion ratio of data at different time granularities, adaptively set the weights of each time scale, and balance data quality and compression ratio; For the selected multi-granularity feature subset, adaptively optimize the sensing matrix through dictionary learning method, introduce a sparse regularization term to enhance the sparsity of data reconstruction, establish a mapping relationship between the compression ratio, reconstruction error and feature subset according to the missing pattern and context information of the missing values, and determine the adaptive compression strategy; Based on the adaptive compression strategy, use compressive sensing method to perform adaptive compression on multi-time granularity power data, dynamically adjust the compression ratio according to the reconstruction error feedback, balance data compression and reconstruction quality, synchronize the reconstructed multi-granularity power data according to the timestamp accuracy, and use it as the input for subsequent correlation analysis; Mine the association rules between multi-time granularity power data, fuse multi-scale features, construct a multi-dimensional data association analysis model for power grid operation status monitoring, and construct a multi-time scale dynamic association rule update mechanism according to the periodic characteristics of power data and the data update cycle difference, and perform real-time association analysis under the change of power grid working conditions; Perform balanced sharding storage on the compressed and reconstructed multi-time granularity power data, dynamically optimize the storage and computing resource allocation according to the data sparsity, missing value distribution and association strength, feedback the reconstruction error to the feature extraction stage, and iteratively optimize the feature matrix to incorporate the missing data estimation and multi-scale fusion error into the dynamic adjustment process.
2. The method according to claim 1, wherein The step of obtaining power grid operation data at different sampling frequencies, for data missing patterns at different time scales, using time series decomposition method to decompose multi-scale data into high-frequency components and low-frequency components, performing sparse representation on the high-frequency components, performing interpolation estimation on the low-frequency components, and evaluating the data missing rate at each time scale to obtain preliminarily completed multi-time granularity power data includes: Obtain power grid operation data at multiple sampling frequencies, where the sampling frequencies include second level, minute level and hour level; Calculate the data missing rate according to the power grid operation data, and analyze the missing pattern of the data missing rate; If the data missing rate exceeds a preset threshold, mark the power grid operation data; Performing seasonal decomposition on the marked power grid operation data to obtain a trend component, a seasonal component and a residual component; The trend component is fitted by polynomial regression, the seasonal component is expanded by Fourier series, and the residual component is modeled by autoregressive moving average; Selecting a corresponding processing method according to the characteristics of the trend component, seasonal component and residual component; Performing local weighted regression smoothing on the trend component, performing periodic interpolation on the seasonal component, and estimating the residual component using Kalman filtering; Merging the processed trend component, seasonal component and residual component; Perform an inverse transformation on the merged data to obtain the restored data at the original time scale; The restored data includes second-level, minute-level and hour-level data; Calculate the root mean square error and mean absolute percentage error of the restored data; If the root mean square error or mean absolute percentage error exceeds a preset threshold, adjusting the seasonal decomposition parameter, regression coefficient and number of Fourier series terms; The data processing process is re-executed until the accuracy requirement is met or the maximum number of iterations is reached.
3. The method according to claim 1, wherein, According to the physical quantity type and numerical distribution characteristics of the power data, the completed multi-time granularity power data is divided into time windows, a high-dimensional feature matrix reflecting the operating status of the equipment is extracted, and the feature matrix is compressed by dimensionality reduction using a frequency domain analysis method to obtain key feature vectors at different time granularities, and the sparsity of data at each time scale is measured, including: Receiving multi-time granularity power data, and normalizing the multi-time granularity power data according to voltage, current and power physical quantity types; Acquire normalized multi-time granularity power data, and use a sliding window method to divide the time window, where the time window size is set to an integer multiple of the data period; Calculating statistical features for the data in the time window, the statistical features including mean, variance, kurtosis and skewness, to obtain a high-dimensional feature matrix reflecting the operating status of the device; Applying short-time Fourier transform to the high-dimensional feature matrix to obtain time-frequency domain features; Set an energy threshold, retain the frequency components whose cumulative energy percentage exceeds the preset threshold, remove other components, and obtain the feature matrix after dimensionality reduction and compression; Applying principal component analysis to the feature matrix after dimension reduction and compression, selecting the principal component whose cumulative contribution rate reaches a preset threshold as the key feature vector; Calculating the Gini coefficient of the key feature vector as a sparse measurement indicator; If the Gini coefficient is less than a first preset threshold, performing sparse processing by increasing the time window or reducing the sampling rate; If the Gini coefficient is greater than a second preset threshold, encryption processing is performed by reducing the time window or increasing the sampling rate; The feature extraction and dimensionality reduction compression process is repeatedly performed until the Gini coefficient falls between the first preset threshold and the second preset threshold, thereby obtaining a feature representation that takes into account both information retention and computational efficiency.
4. The method according to claim 1, wherein Screen out the multi-granularity feature subset most relevant to load forecasting from the key feature vectors through time-frequency analysis methods, and dynamically adjust the fusion ratio of data at different time granularities according to the sparsity of data at each time scale and the spatial distribution characteristics of missing values, and adaptively set the weights of each time scale to balance data quality and compression ratio, including: Apply short-time Fourier transform and continuous wavelet transform to the key feature vectors to obtain the time-frequency features; Calculate the mutual information between each frequency band and the load forecasting result according to the time-frequency features, and select the top N frequency bands with the highest mutual information values as the multi-granularity feature subset; If there are missing values in the multi-granularity feature subset, use the local outlier factor algorithm to detect the spatial distribution characteristics of the missing values; Calculate the local density and relative density of the missing values through the spatial distribution characteristics, and construct a data quality evaluation function in the form of a weighted sum; Design a Mamdani-type fuzzy controller based on the output value of the data quality evaluation function. The input of the fuzzy controller is the data quality score, and the output is the fusion ratio adjustment amount; Calculate the conditional entropy of data at each time scale using the fusion ratio adjustment amount to obtain an information quantity index; Construct an objective programming model by combining the information quantity index and the data quality score; Solve the objective programming model to obtain the adaptive weights of each time scale.
5. The method according to claim 1, wherein, For the selected multi-granularity feature subset, adaptively optimize the sensing matrix through dictionary learning methods, introduce a sparse regularization term to enhance the sparsity of data reconstruction, and establish a mapping relationship between the compression ratio, reconstruction error, and feature subset according to the missing pattern and context information of the missing values to determine the adaptive compression strategy, including: Apply the K-SVD algorithm to the selected multi-granularity feature subset for dictionary learning; According to the dictionary learning result, use the exponentially weighted moving average method to calculate the time series trend of the missing values, and the time series trend calculation uses a 24-hour window size; Obtain the mean, variance, and skewness of the time series trend and context data, and construct a 10-dimensional missing pattern description vector; Use a random forest regressor to establish a mapping relationship between the compression ratio, reconstruction error, and feature subset. The input of the random forest regressor is the feature subset and the missing pattern description vector, and the output is the compression ratio and reconstruction error; According to the output of the random forest regressor, design a fuzzy control system. The input variables of the fuzzy control system are the predicted compression ratio and reconstruction error, and the output variable is the compression parameter adjustment amount; If the compression ratio or reconstruction error exceeds the preset range, then monitor the compression effect in real time and trigger a re-evaluation and adjustment of the compression strategy.
6. The method according to claim 1, wherein Based on the adaptive compression strategy, use the compressive sensing method to adaptively compress multi-time granularity power data, dynamically adjust the compression ratio according to the reconstruction error feedback, balance data compression and reconstruction quality, and synchronize the reconstructed multi-granularity power data according to the timestamp accuracy as the input for subsequent correlation analysis, including: Obtain compressive sensing data according to the multi-time granularity power data. The compressive sensing data is obtained by processing the multi-time granularity power data with a sparse representation algorithm; The compressed sensing data is processed by the iterative hard threshold algorithm to obtain the reconstructed data, and the root mean square error is calculated according to the reconstructed data; If the root mean square error is greater than the preset error threshold, the gradient descent method with an adaptive step size is used to adjust the compression ratio, and the gradient descent method with an adaptive step size dynamically adjusts the step size according to the change of the root mean square error; The reconstructed data is processed by the timestamp accuracy unification algorithm to obtain multi-granularity power data with unified accuracy, and the multi-granularity power data with unified accuracy includes second-level data, minute-level data, and hour-level data; The dynamic time warping algorithm is used to process the multi-granularity power data with unified accuracy to obtain the aligned multi-granularity power data. The dynamic time warping algorithm includes calculating the distance matrix between time series, constructing the optimal alignment path by the dynamic programming method, and mapping the data with different time granularities to the unified time axis according to the optimal alignment path.
7. The method according to claim 1, wherein, The association rules between the power data with different time granularities are mined, the multi-scale features are fused, and a multi-dimensional data association analysis model for power grid operation state monitoring is constructed. According to the periodic characteristics of the power data and the difference in data update cycles, a multi-time scale dynamic association rule update mechanism is constructed to perform real-time association analysis under the change of power grid conditions, including: The FP-Growth algorithm is used to mine the association rules of the power data with different time granularities, and the optimal support and confidence thresholds are obtained through grid search; According to the optimal support and confidence thresholds, multi-scale feature extraction is performed on the power data, and the wavelet packet transform is used to adaptively select the best wavelet basis and decomposition layers; For the multi-scale features, the short-time Fourier transform is used to analyze the time-frequency spectrum characteristics, and the peak detection algorithm is used to identify the main periodic components; According to the main periodic components, an adaptive sliding window mechanism is designed, and the window size is dynamically adjusted according to the main period; For the data within the adaptive sliding window, a real-time association analysis model based on the gated recurrent unit is constructed, and the input layer of the model includes multi-time scale features and association rule strengths; If the model output indicates that the current power grid condition has changed, the model parameters are updated through online learning to perform real-time association analysis under the change of power grid conditions.
8. The method according to claim 1, wherein The multi-time granularity power data after compression and reconstruction is evenly sliced and stored, and the storage and computing resource allocation is dynamically optimized according to the data sparsity, missing value distribution, and association strength. The reconstruction error is fed back to the feature extraction stage, and the feature matrix is iteratively optimized, and the missing data estimation and multi-scale fusion error are incorporated into the dynamic adjustment process, including: The Z-order curve is used to perform multi-dimensional indexing on the multi-time granularity power data after compression and reconstruction, and the data is mapped to a one-dimensional space; According to the one-dimensional space mapped data, the consistent hashing algorithm is used for sliced storage, and the load balance is achieved by dynamically adjusting the number of virtual nodes; For the sliced stored data, a resource allocation model is constructed, and the particle swarm optimization algorithm is used to dynamically adjust the storage and computing resources; Obtain the optimization result of the resource allocation model, design an error feedback mechanism, and update the feature matrix using the adaptive moment estimation algorithm; If the update of the feature matrix is completed, construct a multi-task learning framework based on the attention mechanism, and use missing data estimation and multi-scale fusion error as auxiliary tasks; The multi-task learning framework uses a soft parameter sharing strategy to allow information interaction between tasks; For the multi-task learning framework, adopt an uncertainty-weighted method to dynamically adjust the task weights, and introduce gradient normalization to alleviate the gradient imbalance problem; Through the multi-task learning framework, achieve the global optimization of the feature matrix, dynamically adjust the weights of each task, and adapt to the changing characteristics of power data.
Citation Information
Cited By
Unsupervised anomaly detection method and system for substation equipment
CN120833525A
Deep neural network reasoning acceleration method and system
CN120893499A
Control method for combustion system of direct injection methanol generator based on multi-objective optimization
CN120968919A
New energy station mutual inductor signal adaptive decomposition method and system
CN121522242A
Data compression method, device and equipment for regular expression HNFA model
CN121603013A