Screening and dimension reduction method for multi-source heterogeneous data based on electric power big data
By screening and reducing the multi-source heterogeneous data of power big data, the problems of wide data sources, diverse formats and serious noise interference in the power system are solved, the unity and accuracy of data are achieved, data processing efficiency and accuracy are improved, and the complex characteristics of the power system are adapted.
Patent Information
- Application Number
- CN202510339608.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-18
AI Technical Summary
In the power big data scenario, multi-source heterogeneous data is diverse in sources, inconsistent formats, high dimensions and noise-containing, resulting in low data integration efficiency and redundant feature interference analysis accuracy. It is difficult for traditional dimensionality reduction methods to effectively integrate timing and spatial characteristics, restricting the accuracy of data processing.
Multi-source heterogeneous data screening and dimensionality reduction methods based on power big data are adopted, including data preprocessing, dynamic feature selection, hybrid dimensionality reduction and secondary denoising steps: through cloud platform preprocessing, central processor timing characteristic analysis, hybrid dimensionality reduction module and denoising autoencoder processing, redundant features are eliminated, combined with principal component analysis method and long and short-term memory network to reduce dimensionality reduction data sets to generate.
It effectively solves the problems of wide data sources, diverse formats and serious noise interference in the power system, ensures data uniformity and accuracy, improves data processing efficiency, retains core information, improves data purity and accuracy, and adapts to the complex characteristics of the power system.
Smart Images

Figure CN120336837A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power data processing, and particularly to a method for screening and dimensionality reduction of multi-source heterogeneous data based on power big data. Background Art
[0002] In the scenario of power big data, due to diverse sources, inconsistent formats, high dimensions and noise in multi-source heterogeneous data, the data integration efficiency is low, redundant features interfere with the analysis accuracy, and traditional dimensionality reduction methods are difficult to effectively integrate temporal and spatial characteristics, restricting the accuracy of data processing. Summary of the Invention
[0003] Aiming at the deficiencies of the prior art, the present invention provides a method for screening and dimensionality reduction of multi-source heterogeneous data based on power big data.
[0004] A method for screening and dimensionality reduction of multi-source heterogeneous data based on power big data includes the following steps:
[0005] S1: Obtain multi-source heterogeneous data in the power system, and transmit the multi-source heterogeneous data to a cloud platform. The cloud platform preprocesses the multi-source heterogeneous data through a data preprocessing unit;
[0006] S2: Input the preprocessed multi-source heterogeneous data into a central processing unit. The central processing unit analyzes the temporal characteristics of the multi-source heterogeneous data based on a dynamic feature selection algorithm, screens key features, and eliminates redundant features;
[0007] S3: Use a hybrid dimensionality reduction module to perform dimensionality reduction processing on the screened multi-source heterogeneous data;
[0008] S4: Perform secondary denoising on the dimensionally reduced data to generate a final dimensionally reduced data set.
[0009] Preferably, the step S1 is specifically as follows,
[0010] Obtain multi-source heterogeneous data in the power system, where the multi-source heterogeneous data includes sensor data, user electricity consumption records, and device status logs;
[0011] Transmit the collected multi-source heterogeneous data to the cloud platform. After the cloud platform receives the data, it processes the multi-source heterogeneous data using a data preprocessing unit.
[0012] Preferably, the step S2 is specifically as follows,
[0013] The central processing unit performs time series segmentation on the integrated data, divides the data into multiple time windows, and the data within each time window has the same sampling frequency and time span;
[0014] Within each time window, local features of multi-source heterogeneous data are extracted based on the sliding window mechanism, and the statistics of each feature are calculated;
[0015] The extracted features are scored by a feature importance evaluation model;
[0016] According to the feature scoring results, key features within each time window are selected, and redundant features with scores lower than the preset threshold are removed;
[0017] The filtered key features are recombined in chronological order to form a new time series dataset.
[0018] Preferably, in step S3,
[0019] The new time series dataset is preprocessed and divided into a linear feature part and a non-linear feature part;
[0020] For the linear feature part, principal component analysis is used for dimensionality reduction, which specifically includes the following process:
[0021] For the non-linear feature part, a long short-term memory network is used for dimensionality reduction;
[0022] The dimensionality reduction results of principal component analysis and the long short-term memory network are integrated through a weighted fusion module to generate a dimensionality reduction dataset.
[0023] Preferably, in step S3, for the linear feature part, principal component analysis is used for dimensionality reduction, specifically:
[0024] The linear features are standardized to eliminate the influence of dimensions;
[0025] The covariance matrix is calculated, and its eigenvalues and eigenvectors are solved;
[0026] According to the preset variance contribution rate threshold, the number of principal components is selected, and the original linear features are mapped to a low-dimensional space.
[0027] Preferably, in step S3, for the non-linear feature part, a long short-term memory network is used for dimensionality reduction, specifically:
[0028] The non-linear time series features are standardized and the data sequence is divided according to a fixed time window, and each window includes the feature values of consecutive time steps;
[0029] The preprocessed non-linear time series data is input into the LSTM network, and the hidden state of the time step is extracted as the low-dimensional feature representation.
[0030] Preferably, step S4 is specifically:
[0031] The denoising autoencoder is used to reconstruct and train the dimension-reduced data, and the residual noise is identified through the reconstruction error;
[0032] Based on the isolation forest algorithm, the reconstruction error data points are filtered twice;
[0033] The data distribution after denoising is smoothed.
[0034] The present invention discloses a method for screening and dimension reduction of multi-source heterogeneous data based on power big data, and the beneficial effects thereof are as follows:
[0035] By standardizing the formats of multi-source heterogeneous data, the problems of wide data sources, diverse formats and serious noise interference in the power system can be effectively solved, ensuring the unity and accuracy of the data. By screening key features and eliminating redundant features, the features that have a greater impact on the power system state can be accurately identified, reducing data redundancy, improving data processing efficiency, and at the same time avoiding noise interference introduced by redundant features. The hybrid dimension reduction module is used to perform dimension reduction processing on the screened data, combining the principal component analysis method and the long short-term memory network, and performing dimension reduction for linear characteristics and non-linear features respectively, which can effectively reduce the data dimension, retain the core information of the data, further improve the data processing efficiency, and better adapt to the complex characteristics of the power system data. The dimension-reduced data is denoised twice. Using the denoising autoencoder and the isolation forest algorithm can further remove the residual noise, and at the same time smooth the data distribution, improving the purity and accuracy of the data. Brief Description of the Drawings
[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0037] Figure 1 It is a method flow chart of a method for screening and dimension reduction of multi-source heterogeneous data based on power big data provided by the present invention. Detailed Embodiments
[0038] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.
[0039] To better understand the above technical solution, the above technical solution will be described in detail below in conjunction with the accompanying drawings of the specification and specific implementation manners.
[0040] Referring Figure 1 , a method for screening and dimensionality reduction of multi-source heterogeneous data based on power big data provided by the present invention includes the following steps:
[0041] S1: Obtain multi-source heterogeneous data in the power system, and transmit the multi-source heterogeneous data to the cloud platform. The cloud platform preprocesses the multi-source heterogeneous data through a data preprocessing unit;
[0042] S2: Input the preprocessed multi-source heterogeneous data into a central processing unit. The central processing unit performs time-series characteristic analysis on the multi-source heterogeneous data based on a dynamic feature selection algorithm, screens key features, and eliminates redundant features;
[0043] S3: Use a hybrid dimensionality reduction module to perform dimensionality reduction processing on the screened multi-source heterogeneous data;
[0044] S4: Perform secondary denoising on the dimension-reduced data to generate a final dimension-reduced data set.
[0045] A method for screening and dimensionality reduction of multi-source heterogeneous data based on power big data provided by the present invention can effectively solve the problems of wide data sources, diverse formats, and serious noise interference in the power system data by standardizing the format and filtering noise of the multi-source heterogeneous data, ensuring the unity and accuracy of the data. By screening key features and eliminating redundant features, it can accurately identify the features that have a greater impact on the power system state, reduce data redundancy, improve data processing efficiency, and avoid noise interference introduced by redundant features. Using a hybrid dimensionality reduction module to perform dimensionality reduction processing on the screened data, combining the principal component analysis method and the long short-term memory network, and performing dimensionality reduction for linear characteristics and non-linear features respectively, can effectively reduce the data dimension, retain the core information of the data, further improve the data processing efficiency, and better adapt to the complex characteristics of the power system data. Perform secondary denoising on the dimension-reduced data, using a denoising autoencoder and an isolation forest algorithm, which can further remove residual noise and smooth the data distribution, improving the quality and accuracy of the data.
[0046] In a preferred embodiment, step S1 is specifically
[0047] Acquire multi-source heterogeneous data in the power system, including sensor data, user power consumption records and equipment status logs. Sensor data is collected in real time by various sensors in the power system, such as current, voltage, power, temperature and humidity; user power consumption records cover detailed data such as user basic information, power consumption, power consumption time and time period; equipment status logs include the operating status of power equipment, fault information, maintenance records and equipment parameters.
[0048] The collected multi-source heterogeneous data is transmitted to the cloud platform through the SSL / TLS protocol. After the cloud platform receives the data, it uses the data preprocessing unit to process the multi-source heterogeneous data. Preprocessing includes data cleaning, removing erroneous data and duplicate data, and correcting abnormal data points; data alignment, unifying data from different sources and different timestamps to the same time base; data format conversion, converting data in different formats into a unified data format for subsequent processing and analysis.
[0049] In a preferred embodiment, step S2 is specifically,
[0050] The central processing unit performs time series segmentation on the integrated data, dividing the data into multiple time windows, and the data in each time window has the same sampling frequency and time span;
[0051] In each time window, local features of multi-source heterogeneous data are extracted based on the sliding window mechanism, and the statistics of each feature are calculated, including mean, variance, and trend change rate.
[0052] Specifically, the sensor data collected every second is divided into 30-minute time windows, and data of different frequencies are interpolated and aligned. Local statistics are calculated using a 5-minute sliding window (with a step size of 1 minute), including the mean (trend), variance (fluctuation), trend change rate (linear regression slope), and frequency domain characteristics of the vibration signal (FFT dominant frequency).
[0053] The extracted features are scored by a feature importance evaluation model, where the feature importance evaluation model is based on a random forest algorithm to evaluate the importance score of each feature to the system state;
[0054] Based on the feature scoring results, the key features in each time window are dynamically selected, and redundant features with scores below the preset threshold are eliminated;
[0055] The filtered key features are reassembled in chronological order to form a new time series dataset.
[0056] For example, set the scoring threshold to 0.1. If the feature score is greater than the threshold, the key features are retained. The score of the temperature trend change rate is 0.85, and the score of the vibration main frequency offset is 0.78 are key features, while the score of the minimum temperature is 0.02, etc. are determined to be redundant, where the scoring threshold is dynamically adjusted according to the scenario.
[0057] In a preferred embodiment, step S3,
[0058] Preprocess the new time series dataset, and divide the new time series dataset into a linear feature part and a non-linear feature part;
[0059] For the linear feature part, use the principal component analysis method for dimensionality reduction;
[0060] For the non-linear feature part, use the long short-term memory network for dimensionality reduction, where the non-linear features include user power consumption data and equipment status data, etc.;
[0061] Integrate the dimensionality reduction results of the principal component analysis method and the long short-term memory network through a weighted fusion module to generate a dimensionality reduction dataset.
[0062] In a preferred embodiment, in step S3, for the linear feature part, use the principal component analysis method for dimensionality reduction, specifically:
[0063] Standardize the linear features to eliminate the influence of dimensions and make the features comparable on the same scale; among them, the linear features include the average temperature, average current, average insulation resistance, etc.
[0064] Calculate the covariance matrix and solve its eigenvalues and eigenvectors;
[0065] According to the preset variance contribution rate threshold, select the number of principal components and map the processed linear features to a low-dimensional space.
[0066] Among them, standardizing the linear features to eliminate the influence of dimensions specifically includes:
[0067] Perform mean centering on each linear feature to make its mean 0;
[0068] Perform standard deviation normalization on each linear feature to make its standard deviation 1;
[0069] Calculate the covariance matrix of the standardized linear features. The covariance matrix reflects the linear relationship between the features;
[0070] Perform eigenvalue decomposition on the covariance matrix to obtain the eigenvalues and corresponding eigenvectors. The eigenvalues represent the variance contribution rates of the principal components, and the eigenvectors represent the directions of the principal components;
[0071] Select the first k principal components according to a preset variance contribution rate threshold, so that the cumulative variance contribution rate of the selected principal components reaches or exceeds the threshold;
[0072] Project the original linear feature data into the low-dimensional space formed by the selected principal components to generate a reduced-dimensional linear feature data set;
[0073] In a preferred embodiment, for the non-linear feature part in step S3, a long short-term memory network is used for dimensionality reduction, specifically:
[0074] Perform standardization processing on the non-linear time series features, and divide the data sequence according to a fixed time window. Each window includes the feature values of consecutive time steps;
[0075] Input the preprocessed non-linear time series data into the LSTM network, and extract the hidden state at the time step as the low-dimensional feature representation. Among them, input the non-linear feature sequence into the LSTM network, and through iterative calculations at multiple time steps, obtain the hidden state at each time step; aggregate or select the hidden state to generate the low-dimensional feature representation.
[0076] In a preferred embodiment, step S4 is specifically:
[0077] Use a denoising autoencoder to perform reconstruction training on the reduced-dimensional data, and identify residual noise through the reconstruction error;
[0078] Based on the isolation forest algorithm, perform secondary filtering on the data points with high reconstruction error;
[0079] Smooth the distribution of the denoised data to ensure that it conforms to the business logic constraints of the power system.
[0080] Specifically, the encoder maps the reduced-dimensional data to a lower-dimensional latent space, and the decoder attempts to reconstruct the original input data from the latent space. By adding a certain proportion of Gaussian noise to the reduced-dimensional data, train the autoencoder so that it can restore the original data from the noisy data. During the training process, use the stochastic gradient descent method to optimize the reconstruction error. After multiple iterative trainings, make the autoencoder accurately identify and remove the residual noise in the data.
[0081] The isolation forest constructs multiple isolation trees and performs anomaly detection according to the distribution characteristics of the data points. For data points with a large reconstruction error, it may be caused by residual noise or actual abnormal data, so the isolation forest is used to identify them as abnormal points and filter them out to further improve the purity of the data.
[0082] Smooth the distribution of the denoised data, using methods such as the moving average method or the Gaussian smoothing kernel, to make the data smoother in the time series and conform to the business logic constraints of the power system.
Claims
1. A method for screening and dimensionality reduction of multi-source heterogeneous data based on power big data, characterized in that It includes the following steps: S1: Obtain multi-source heterogeneous data in the power system, and transmit the multi-source heterogeneous data to the cloud platform. The cloud platform preprocesses the multi-source heterogeneous data through a data preprocessing unit; S2: Input the preprocessed multi-source heterogeneous data into the central processing unit. The central processing unit performs time-series characteristic analysis on the multi-source heterogeneous data based on the dynamic feature selection algorithm, screens key features, and eliminates redundant features; S3: Use a hybrid dimensionality reduction module to perform dimensionality reduction processing on the screened multi-source heterogeneous data; S4: Perform secondary denoising on the data after dimensionality reduction to generate a final dimensionality-reduced data set.
2. The screening and dimensionality reduction method for multi-source heterogeneous data based on power big data according to claim 1, characterized in that, The specific content of step S1 is as follows: Obtain multi-source heterogeneous data in the power system, where the multi-source heterogeneous data includes sensor data, user electricity consumption records, and device status logs; Transmit the collected multi-source heterogeneous data to the cloud platform. After the cloud platform receives the data, it uses a data preprocessing unit to process the multi-source heterogeneous data.
3. A screening and dimensionality reduction method for multi-source heterogeneous data based on power big data according to claim 1, characterized in that The specific content of step S2 is as follows: The central processing unit performs time-series segmentation on the integrated data, divides the data into multiple time windows, and the data within each time window has the same sampling frequency and time span; Within each time window, perform local feature extraction on the multi-source heterogeneous data based on the sliding window mechanism, and calculate the statistic of each feature; Score the extracted features through a feature importance evaluation model; According to the feature scoring results, select the key features within each time window, and eliminate the redundant features with scores lower than the preset threshold; Recombine the screened key features in chronological order to form a new time-series data set.
4. A screening and dimensionality reduction method for multi-source heterogeneous data based on power big data according to claim 1, characterized in that, For step S3, Preprocess the new time-series data set, and divide the new time-series data set into a linear characteristic part and a non-linear feature part; For the linear feature part, use the principal component analysis method for dimensionality reduction, which specifically includes the following process: For the non-linear feature part, use a long short-term memory network for dimensionality reduction; Integrate the dimensionality reduction results of the principal component analysis method and the dimensionality reduction results of the long short-term memory network through a weighted fusion module to generate a dimensionality-reduced data set.
5. A screening and dimensionality reduction method for multi-source heterogeneous data based on power big data according to claim 4, characterized in that In step S3, for the linear feature part, using the principal component analysis method for dimensionality reduction is specifically as follows: Perform standardization processing on the linear features to eliminate the influence of dimension; Calculate the covariance matrix, and solve its eigenvalues and eigenvectors; According to the preset variance contribution rate threshold, select the number of principal components, and map the original linear features to a low-dimensional space.
6. A screening and dimensionality reduction method for multi-source heterogeneous data based on power big data according to claim 4, characterized in that In step S3, for the non-linear feature part, using a long short-term memory network for dimensionality reduction is specifically as follows: Perform standardization processing on the non-linear time-series features, and divide the data sequence according to a fixed time window. Each window includes the feature values of consecutive time steps; Input the preprocessed non-linear time-series data into the LSTM network, and extract the hidden state of the time step as the low-dimensional feature representation.
7. A screening and dimensionality reduction method for multi-source heterogeneous data based on power big data according to claim 1, characterized in that, The specific content of step S4 is as follows: Use a denoising autoencoder to perform reconstruction training on the data after dimensionality reduction, and identify the residual noise through the reconstruction error; Perform secondary filtering on the reconstruction error data points based on the isolation forest algorithm; Smooth the distribution of the data after denoising.
Citation Information
Cited By
Multi-dimensional feature optimization sample library construction method supporting user classification and power saving evaluation
CN121092996A