Multi-source power data collaborative fusion method considering data characteristic sensitivity analysis

Through data feature sensitivity analysis and multi-source power data fusion methods, abnormal data is identified and repaired, solving the problems of diverse data types and differentiated time resolution in multi-source power data processing, achieving efficient fusion and utilization of data, and improving the analysis and decision-making capabilities of the power system.

CN120654173APending Publication Date: 2025-09-16STATE GRID JIANGSU ELECTRIC POWER CO LTD NANJING POWER SUPPLY COMPANY +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510626094.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing technologies face problems such as diverse data types, large data volumes, measurement errors, missing data sections, and differentiated data acquisition time resolution when processing multi-source power data, resulting in low data processing and utilization efficiency.

Method used

The data feature sensitivity analysis method is adopted, and the isolation forest algorithm is used to identify and eliminate abnormal data. The influencing factors are determined by combining multiple linear regression and correlation coefficient method. The cubic spline interpolation and piecewise aggregation approximation methods are used to repair and fuse data to achieve timestamp unification.

Benefits of technology

It improves the repair accuracy and utilization rate of multi-source power data, and enhances the reliability and accuracy of power system analysis and decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120654173A_ABST
    Figure CN120654173A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-source electric power data collaborative fusion method considering data characteristic sensitivity analysis, which comprises the following steps of: 1, analyzing and detecting missing values and abnormal data in original multi-source electric power data, and removing the missing values and the abnormal data to obtain data samples after the missing values and the abnormal data are removed; step 2, performing sensitivity analysis on the data features of the data sample obtained in the step 1 after the missing value and the abnormal data are removed, and then repairing the data to obtain a repaired data sample; and step 3, performing up-sampling on low-dimensional data with low time resolution in the repaired data samples obtained in the step 2, and performing equal-time-resolution dimension reduction on high-dimensional data in the low-dimensional data, thereby realizing time stamp fusion and unification of the time sequence data samples. According to the method, a scientific basis is provided for optimization of data samples, and the utilization rate of power big data in a system is promoted to be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data fusion methods, and in particular to a multi-source power data collaborative fusion method taking into account data feature sensitivity analysis. Background Art

[0002] With the development of renewable energy generation technologies, power electronics applications, and intelligent devices, the multi-source, heterogeneous data generated during power system operation can provide more comprehensive operational information, facilitating more accurate analysis and intelligent decision-making. However, while this presents enormous opportunities, it also presents challenges in data processing and efficient utilization. Data processing often faces issues such as diverse power data types, large data volumes, measurement errors, missing data sections, and varying temporal resolutions of data acquisition. The key to current research is to repair and fuse multi-source power data to ultimately generate high-quality data samples with unified timestamps and high reliability, providing reliable data support for subsequent analysis and research. Summary of the Invention

[0003] The present invention provides a multi-source power data collaborative fusion method taking into account data feature sensitivity analysis to solve the fusion problem in the prior art multi-source power data processing.

[0004] In order to achieve the above object, the technical solution adopted by the present invention is:

[0005] A multi-source power data collaborative fusion method taking into account data feature sensitivity analysis includes the following steps:

[0006] Step 1: Obtain original multi-source power data, analyze and detect missing values ​​in the original multi-source power data and remove them, then identify and remove abnormal data in the remaining data, thereby obtaining a data sample after removing missing values ​​and abnormal data, providing support for the subsequent abnormal section data repair and fusion;

[0007] Step 2: Perform sensitivity analysis on the data features of the data sample obtained in step 1 after removing missing values ​​and abnormal data, screen out the factors that have a more significant impact on the data sample, and then calculate the correlation coefficient between the influencing factor data group corresponding to the data abnormal moment and the influencing factor data group at other normal moments, and obtain the influencing factor data group with the largest correlation coefficient value. The value corresponding to this data group is the repair value at the abnormal moment, thereby obtaining the repaired data sample;

[0008] Step 3: Upsample the low-dimensional data with low temporal resolution in the repaired data samples obtained in step 2, and perform dimensionality reduction with equal temporal resolution on the high-dimensional data in the repaired data samples obtained in step 2, thereby achieving the timestamp fusion and unification of the time series data samples.

[0009] Furthermore, in step 1, the Pandas data analysis package is used to analyze and detect missing values ​​in the original multi-source power data and remove them.

[0010] Furthermore, in step 1, the isolation forest algorithm is used to identify abnormal data in the remaining data.

[0011] Furthermore, the isolation forest algorithm used mainly includes two parts: training isolation trees and anomaly scoring. The steps of constructing an isolation tree in the isolation forest algorithm are as follows:

[0012] (1) A dataset X with N samples = {x i |i=1,2,…,N}, x i The data dimension is m. N sample points are randomly selected from the data set X to form a training set as the root node of the isolation tree. The data set X is constructed by removing the missing values ​​in step 1.

[0013] (2) Randomly select a j dimension from the m dimensions of the dataset X and randomly select a split value p between the maximum and minimum values ​​of the dimension j in the training set;

[0014] (3) Based on the hyperplane formed by the segmentation value p, the training set is divided into two subsets, the left subtree and the right subtree, to form leaf nodes;

[0015] (4) Recursively perform steps (2) and (3) in the left and right subsets of the leaf node, and continuously generate new leaf nodes until the leaf node has only one data sample or the number of splits reaches the set value, then stop splitting, thus forming an isolated tree;

[0016] Based on the above steps (1)-(4), multiple isolated trees are constructed to form an isolation forest for training, and the anomaly score is used to comprehensively judge the training results. For a data set containing n sample points, the anomaly score s(x,n) of the data sample x is shown in the following formula:

[0017]

[0018] c(n)=2(ln(n-1)+δ)-(2(n-1) / n) (2)

[0019] In formulas (1) and (2), h(x) is the height of data sample x, that is, the number of segmentations experienced; E(h(x)) is the average value of h(x); δ is the correction term; c(n) is the normalization factor when the number of samples is n;

[0020] The closer the s(x,n) value is to 1, the higher the probability of anomaly, and the closer it is to 0, the normal data. If most of the data scores are close to 0.5, it means there is no abnormal data.

[0021] Therefore, the isolation forest algorithm is used to identify abnormal data in the remaining data after removing missing values, and the abnormal data is removed to obtain a data sample after removing missing values ​​and abnormal data.

[0022] Furthermore, in step 2, a multiple linear regression method is used to perform sensitivity analysis on the data characteristics of the data samples obtained in step 1.

[0023] Furthermore, the process of data feature sensitivity analysis using the multiple linear regression method is as follows:

[0024] First, the mathematical expression of the MLR model is established as shown below:

[0025] y=β0+β1x1+β2x2+…+β n x n +ε (3)

[0026] In formula (3), y is the dependent variable; x1, x2,…, x n are independent variables; β0,β1,…,β n is the regression coefficient; ε is the error term;

[0027] Suppose there are m groups of data samples, and the mth group of dependent variables in the m groups of data samples is y m , the independent variable is x m1 、x m2 、……、x mn , then formula (3) can be converted into the following formula:

[0028]

[0029] Among them, ε1, ε2, …, ε m are the error terms of each group, and ε1,ε2,…,ε m Independent of each other;

[0030] set up are β0,β1,…,β n The least squares estimate of is the observed value y k The estimated value of is as follows:

[0031]

[0032] Where: k = 1, 2, ..., m; e k is the estimation error; the independent variable is x k1 、x k2 ,……,x kn ; ε k is the error term for each group;

[0033] The least squares method is used to estimate the dependent variable parameters of the MLR model to estimate the error e k The goal is to minimize the sum of squares Q, as shown in the following formula:

[0034]

[0035] The estimated value of the regression coefficient β can be obtained by solving the extreme value principle and the matrix equation; β0,β1,…,β n In the matrix solution, first construct a matrix H consisting of m groups of sample independent variable data as shown below:

[0036]

[0037] Among them, the independent variable is x 11 、x 12 、x 1n 、x 21 ,……,x mn ;

[0038] Secondly, construct a column vector Y=[y1,y2,…,y m ] T , the regression coefficient β can be calculated using the following formula:

[0039] β=(H T H) -1 H T Y=[β0,β1,β2,…,β n ] T (10)

[0040] Among them: dependent variables are y1, y2, ..., y m ;

[0041] Using the data obtained in step 1 after removing missing values ​​and abnormal data, the regression coefficients β0, β1,…, β n , which can characterize the data feature sensitivity of the data sample obtained in step 2 after removing missing values ​​and abnormal data.

[0042] Furthermore, the correlation coefficient methods in step 2 are the Pearson correlation coefficient method and the Spearman correlation coefficient method. The Pearson correlation coefficient is used to analyze the linear correlation between data and is applicable to linearly correlated continuous variable data; the Spearman correlation coefficient is used for nonlinearly correlated or non-normally distributed data, focusing on measuring the monotonic correlation between data.

[0043] Furthermore, in the Pearson correlation coefficient method, the sample Pearson correlation coefficient ρ between two sets of data X and Y is X,Y The calculation formula is as follows:

[0044]

[0045] In formula (11), cov(X,Y) is the covariance of data sets X and Y; σ X , σ Y are the standard deviations of X and Y respectively; n is the number of data points in the data set; x i and y i are the i-th data point in the data sets X and Y respectively; E(X) represents the variance of X; E(Y) represents the variance of Y; E(XY) represents the variance of XY;

[0046] ρ X,Y The closer the value is to 1, the stronger the positive correlation between the data groups is; the closer it is to -1, the stronger the negative correlation between the data groups is.

[0047] In the Spearman correlation coefficient method, the Spearman correlation coefficient ρ between two sets of data X and Y s The calculation formula is as follows:

[0048]

[0049] In formula (12), n is the number of data points in the data set; d i is the position difference between the corresponding i-th data point in X and Y after they are arranged in ascending order;

[0050] ρ s The closer the value is to 1, the stronger the positive correlation between the data groups is; the closer it is to -1, the stronger the negative correlation between the data groups is.

[0051] After finding the influencing data group with the greatest correlation with the influencing factor group of the data to be repaired, the corresponding data is used as the data to be repaired, thereby obtaining a repaired data sample.

[0052] Furthermore, in step 3, a cubic spline interpolation algorithm is used to upsample the low-dimensional data with low temporal resolution in the repaired data samples obtained in step 2.

[0053] Furthermore, in step 3, a segmented aggregation approximation method is used to perform dimensionality reduction of the high-dimensional data in the repaired data sample obtained in step 2 with equal time resolution.

[0054] The present invention introduces the isolation forest theory and data correlation analysis method into the abnormal data identification and data repair / filling of power grid data. Since time series data may be affected by some other factors, such as photovoltaic output will be affected by the meteorological conditions at the corresponding time, the present invention takes into account the data feature sensitivity analysis, and uses the random forest algorithm combined with the multivariate linear regression method and the correlation coefficient method to extract the key influencing factors affecting the power data, and screen the data features with higher contribution for data repair, so as to improve the accuracy of data repair. In addition, the present invention adopts the cubic spline interpolation and PAA methods to address the problem of differentiated time resolution of data acquisition, unify the data dimensions according to the time granularity requirements of the joint analysis of different types of data sets, and finally realize data fusion. The method of the present invention will provide a scientific basis for the optimization of data samples and promote the improvement of the utilization rate of big data of electric power in the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 Flowchart of the present invention.

[0056] Figure 2 Flowchart for data repair that takes into account the correlation of data features.

[0057] Figure 3 The figure shows the comparison results of different methods in single sampling point data restoration.

[0058] Figure 4 The comparison results of different methods in repairing continuous sampling point data are shown in the figure.

[0059] Figure 5 This is the load time granularity upsampling result diagram.

[0060] Figure 6 This is the result of dimensionality reduction of photovoltaic output time granularity. DETAILED DESCRIPTION

[0061] The present invention will be further described below with reference to the accompanying drawings and examples.

[0062] like Figure 1 As shown, this embodiment discloses a multi-source power data collaborative fusion method taking into account data feature sensitivity analysis, including the following steps:

[0063] Step 1: Obtain the original multi-source power data, use the Pandas data analysis package to analyze and detect missing values ​​in the original multi-source power data and remove them, then use the isolation forest algorithm to identify abnormal data in the remaining data and remove the abnormal data, thereby obtaining a data sample after removing missing values ​​and abnormal data, providing support for subsequent abnormal section data repair.

[0064] Specifically, in this embodiment, the original multi-source power data is imported into a two-dimensional table DataFrame through the Pandas data analysis package, which facilitates row-by-row and column-by-column operations on the two-dimensional table DataFrame. The DataFrame.isnull().sum() function in the Pandas data analysis package is used to detect whether there are null values ​​in each column of the DataFrame and give the number of null values. Depending on the importance of different types of data sets, the next step of operation is performed on the null values, thereby detecting missing values ​​and eliminating them.

[0065] In this embodiment, the isolation forest algorithm used mainly includes two parts: training an isolation tree and scoring anomalies. The steps of constructing an isolation tree in the isolation forest algorithm are as follows:

[0066] (1) A dataset X with N samples = {x i |i=1,2,…,N}, x i The data dimension is m. N sample points are randomly selected from the dataset X to form a training set as the root node of the isolation tree. The dataset X is constructed by removing the missing values ​​in step 1.

[0067] (2) Randomly select a j dimension from the m dimensions of the dataset X, and randomly select a split value p between the maximum and minimum values ​​of the dimension j in the training set.

[0068] (3) According to the hyperplane formed by the segmentation value p, the training set is divided into two subsets, the left subtree and the right subtree, to form leaf nodes.

[0069] (4) Recursively perform steps (2) and (3) in the left and right subsets of the leaf node, and continuously generate new leaf nodes until the leaf node has only one data sample or the number of splits reaches the set value, then stop splitting, thus forming an isolated tree.

[0070] Based on the above steps (1)-(4), multiple isolated trees are constructed to form an isolation forest for training, and the anomaly score is used to comprehensively judge the training results. For a data set containing n sample points, the anomaly score s(x,n) of the data sample x is shown in the following formula:

[0071]

[0072] c(n)=2(ln(n-1)+δ)-(2(n-1) / n) (2)

[0073] In formulas (1) and (2), h(x) is the height of data sample x, that is, the number of segmentations experienced; E(h(x)) is the average value of h(x); δ is the correction term; and c(n) is the normalization factor when the number of samples is n.

[0074] The closer the s(x,n) value is to 1, the higher the probability of anomaly, while the closer it is to 0, the more normal the data. If the majority of the data scores are close to 0.5, there are no anomalies. Using the Isolation Forest algorithm, we can identify and remove anomalies from the remaining data after removing missing values, thus obtaining a data sample after removing missing and anomaly values.

[0075] In this embodiment, the effect of Isolation Forest anomaly data recognition is verified by using some user load data. The data contains 100 data samples, including 5 anomaly samples. Typical anomaly detection evaluation indicators such as True Positive Rate (TPR), False Positive Rate (FPR), Precision (P) and Accuracy (Acc) are used to verify the effectiveness of the algorithm. TPR represents the proportion of samples correctly detected as normal to the actual normal samples. FPR represents the proportion of samples that are not successfully detected among the actual anomaly samples. P represents the proportion of correctly detected samples to all normal detection samples. Due to slight deviations in the program running results, the present invention provides the average accuracy (avgA) obtained from multiple simulations. Among them, the closer TPR, P, Acc and avgA are to 1, and the closer FPR is to 0, the better the detection effect of the method. The results obtained are shown in Table 1.

[0076] Table 1 Isolation Forest Anomaly Detection Results

[0077]

[0078] In Table 1, N_itree represents the number of isolated trees, and ψ represents the size of the training data. The default values ​​for N_itree and ψ are 100 and 256, respectively. As shown in the table, the anomaly detection performance decreases as N_itree and ψ decrease. When the default values ​​for the number of training trees and the training data size are used, anomaly detection accuracy is higher. Various indicators also demonstrate the effectiveness of the Isolation Forest algorithm in identifying anomaly data.

[0079] Step 2: Use the Multiple Linear Regression (MLR) method to perform sensitivity analysis on the data characteristics of the data samples obtained in step 1 after removing missing values ​​and abnormal data, and screen out the factors that have a more significant impact on the data samples. Then calculate the Pearson correlation coefficient between the influencing factor data group corresponding to the data abnormal moment and the influencing factor data group at other normal moments, and obtain the influencing factor data group with the largest correlation coefficient value. The value corresponding to this data group is the repair value at the abnormal moment, and thus the repaired data sample is obtained. The specific process is as follows: Figure 2shown.

[0080] Specifically, in this embodiment, the process of performing data feature sensitivity analysis using the Multiple Linear Regression (MLR) method is as follows:

[0081] First, the mathematical expression of the MLR model is established as shown below:

[0082] y=β0+β1x1+β2x2+…+β n x n +ε (3)

[0083] In formula (3), y is the dependent variable; x1, x2,…, x n are independent variables; β0,β1,…,β n is the regression coefficient; ε is the error term.

[0084] Suppose there are m groups of data samples, and the mth group of dependent variables in the m groups of data samples is y m , the independent variable is x m1 、x m2 ,……,x mn , then formula (3) can be converted into the following formula:

[0085]

[0086] Among them, ε1, ε2, …, ε m are the error terms of each group, and ε1,ε2,…,ε m Independent of each other.

[0087] set up are β0,β1,…,β n The least squares estimate of is the observed value y k The estimated value of is as follows:

[0088]

[0089] Where: k = 1, 2, ..., m; e k is the estimation error; the independent variable is x k1 、x k2 ,……,x kn ; ε k is the error term for each group.

[0090] The least squares method is used to estimate the dependent variable parameters of the MLR model to estimate the error e k The goal is to minimize the sum of squares Q, as shown in the following formula:

[0091]

[0092] The estimated value of the regression coefficient β can be obtained by solving the extreme value principle and the matrix equation. β0,β1,…,β n In the matrix solution, first construct a matrix H consisting of m groups of sample independent variable data as shown below:

[0093]

[0094] Among them, the independent variable is x 11 、x 12 、x 1n 、x 21 ,……,x mn .

[0095] Secondly, construct a column vector Y=[y1,y2,…,y m ] T , the regression coefficient β can be calculated using the following formula:

[0096] β=(H T H) -1 H T Y=[β0,β1,β2,…,β n ] T (10)

[0097] Among them: dependent variables are y1, y2, ..., y m .

[0098] Using the data obtained in step 1 after removing missing values ​​and abnormal data, the regression coefficients β0, β1,…, β n , which can characterize the data feature sensitivity of the data sample obtained in step 2 after removing missing values ​​and abnormal data.

[0099] After determining the sensitivity, we screen out the significant factors related to the dependent variable, and then calculate the correlation coefficient of the significant factors.

[0100] In this embodiment, the correlation coefficient method used is the Pearson correlation coefficient method or the Spearman correlation coefficient method. The Pearson correlation coefficient is used to analyze the linear correlation between data and is more suitable for linearly correlated continuous variable data. It works better for normally distributed data. The Spearman correlation coefficient is often used for nonlinearly correlated or non-normally distributed data, focusing on measuring the monotonic correlation between data. Among them, the sample Pearson correlation coefficient ρ between two sets of data X and Y is X,Y The calculation formula is as follows:

[0101]

[0102] In formula (11), cov(X,Y) is the covariance of data sets X and Y; σ X , σ Y are the standard deviations of X and Y respectively; n is the number of data points in the data set; x i and y i are the i-th data point in the data sets X and Y respectively; E(X) represents the variance of X; E(Y) represents the variance of Y; and E(XY) represents the variance of XY.

[0103] ρ X,Y The closer the value is to 1, the stronger the positive correlation between the data groups is; the closer it is to -1, the stronger the negative correlation between the data groups is.

[0104] The Spearman correlation coefficient ρ between two sets of data X and Y s The calculation formula is as follows:

[0105]

[0106] In formula (12), n is the number of data points in the data set; d i It is the position difference between the corresponding i-th data point in X and Y after they are arranged in ascending order.

[0107] Similarly, ρ s The closer the value is to 1, the stronger the positive correlation between the data groups is; the closer it is to -1, the stronger the negative correlation between the data groups is.

[0108] After finding the influencing data group with the greatest correlation with the influencing factor group of the data to be repaired, the corresponding data is used as the data to be repaired, thereby obtaining a repaired data sample.

[0109] This embodiment uses meteorological data and photovoltaic output data from the same period to perform data feature sensitivity analysis, and compares and verifies the proposed data repair method with the traditional data repair method. The data set uses relevant data collected from distributed photovoltaics for one month (data collection resolution 5min, a total of 8928 data samples). Meteorological data include pm2.5 content, wind speed, wind direction, rainfall, humidity, air pressure, temperature, illumination, radiation, cumulative radiation and noise. Different meteorological factors have different sensitivities to photovoltaic output. Therefore, MLR analysis is used to extract the key meteorological factors affecting photovoltaic output, with meteorological factors as independent variables and photovoltaic output as dependent variables. In order to calculate the multiple regression coefficient, standard error, t value and P value, the least squares method is used in Python language to solve the problem. The results are shown in Table 2.

[0110] Table 2 Primary analysis parameters of MLR1 model

[0111]

[0112]

[0113] If the P value of a variable is less than 0.05, the confidence interval for the variable being a significant influencing variable is 95%. Based on the P values ​​of the test results, it can be found that, except for the constant term, PM2.5 content, rainfall, and air pressure do not have a statistically significant impact on photovoltaic output. Excluding these three variables, the remaining variables were further subjected to MLR analysis. The calculation results are shown in Table 3.

[0114] Table 3 Secondary analysis parameters of MLR1 model

[0115]

[0116] The secondary analysis results of MLR1 show that wind speed, wind direction, humidity, temperature, radiation, cumulative radiation, and noise remain significant variables influencing PV output. Radiation has the greatest impact on PV output, and the two are positively correlated. To clarify the impact of illumination and other factors on PV output, the MLR2 model was constructed, and the analytical results are shown in Table 4.

[0117] Table 4 Secondary analysis parameters of MLR1 model

[0118]

[0119] According to the test results, except for the constant term, PM2.5 content, wind direction, rainfall and air pressure have no statistically significant impact on photovoltaic output. In addition to these four variables, the remaining variables continue to be subjected to MLR analysis, and the calculation results are shown in Table 5.

[0120] Table 5. Secondary analysis parameters of MLR2 model

[0121]

[0122] The secondary analysis results of the MLR2 model show that wind speed, humidity, temperature, illuminance, cumulative radiation, and noise remain significant variables affecting PV output. Illuminance has the greatest impact on PV output, and the two are positively correlated. Combining the analysis results of the MLR1 and MLR2 models, wind speed, humidity, temperature, illuminance, radiation, cumulative radiation, and noise are significant factors affecting PV output. The correlation coefficients between the selected significant factors and actual PV output are shown in Table 6.

[0123] Table 6 Correlation coefficients between significant factors and photovoltaic output

[0124]

[0125] The data repair comparison simulation method selects linear interpolation and cubic spline interpolation to compare with the correlation interpolation method proposed in the present invention, and selects two groups of data for simulation verification, one group is a single sampling point and the other group is continuous sampling points.

[0126] Group 1: A missing data section was set at 11:20 on July 5, 2021 (sampling point 137). Pearson coefficient correlation analysis showed that the time point closest to 11:20 on July 5, 2021, in dataset C2 was 12:35 on July 26, 2021. The correlation coefficient between the two sets of meteorological factors exceeded 0.999. Therefore, the photovoltaic output value corresponding to 12:35 on July 26, 2021 was used as the repair value for 11:20 on July 5, 2021. The method comparison results are shown in the figure below. Figure 3 As shown in the figure, this set of simulation results shows that the error of linear interpolation is 14.883kW, the error of cubic spline interpolation is 11.925kW, and the error of correlation interpolation (the method proposed in this invention) is 9.866kW. It can be seen intuitively from the results that the data repair effect of the interpolation method based on correlation analysis is better than the comparison method.

[0127] The second group: set a continuous missing data section from 9:45 to 10:05 on July 5, 2021 (sampling points 118-122). Through Pearson coefficient correlation analysis, the time points closest to 9:45-10:05 on July 5, 2021 in dataset C2 were 12:55 on July 26, 2021 and 13:15, 14:35, 9:35, and 11:55 on July 5, 2021. The correlation coefficients between each group of meteorological factors exceeded 0.999. Figure 4 As shown in the figure, this set of simulation results shows that the average error of linear interpolation is 78.378kW, the error of cubic spline interpolation is 118.411kW, and the error of correlation interpolation (the method proposed in this invention) is 28.045kW. It can be seen intuitively from the results that the data repair effect of the interpolation method based on correlation analysis is better than the comparison method.

[0128] It can also be seen from the two sets of simulations that when the missing data sections are small, spline interpolation can still play a good role, but it is not effective when the missing sections of continuous data are large. At this time, the data repair effect of the interpolation method based on correlation analysis is significantly better than the comparison method.

[0129] Step 3: Use the cubic spline interpolation algorithm to upsample the low-dimensional data with low temporal resolution in the repaired data samples obtained in step 2, and use the piecewise aggregate approximation (PAA) method to reduce the dimensionality of the high-dimensional data in the repaired data samples obtained in step 2 to equal temporal resolution, thereby achieving the timestamp fusion and unification of time series data samples.

[0130] Specifically, in this embodiment, the upsampling process of using the cubic spline interpolation algorithm for low-dimensional data with low temporal resolution is as follows:

[0131] Divide the interval [a, b] into n small intervals, that is, a=x0 <x1<…<x n =b, where x1, x2, x n are the values ​​of each interval respectively. The function value f(x i )=f i , i=0,1,…,n, f(x) is a cubic spline interpolation function composed of n cubic polynomials and is second-order continuously differentiable in [a,b]. To solve f(x), we can assume that the function y i (x) is shown in the following formula:

[0132] y i (x) = a i x 3 +b i x 2 +c i x+d i ,x∈[x i ,x i+1 ](i=0,1,…,n-1) (13)

[0133] Among them: a i ,b i ,c i and d i There are 4n unknown numbers, and y i (x) The following conditions are met:

[0134]

[0135] Formula (14) includes interpolation conditions, continuity conditions, first-order derivative continuity conditions, second-order derivative continuity conditions and boundary conditions. The above conditions can be used to solve y i (x) thus obtaining f(x).

[0136] In this embodiment, the process of using PAA to perform equal time resolution dimensionality reduction on high-dimensional data is as follows:

[0137] For a time series data with dimension L, Y={y j |j=1,2,…,L}, and we want to reduce the dimension to a time series data of dimension l Y'={y' i |i=1,2,…,l}, we need to divide the sequence Y into l equal segments, where y j is the time series data of the i-th segment in L segments, y' i is the time series data of the i-th segment in the l-th segment.

[0138] y' i The specific calculation process is as follows:

[0139]

[0140] The data after being divided into l segments Y'={y' i |i=1,2,…,l} is the dimensionality reduction data.

[0141] This embodiment uses a user load dataset with a sampling time resolution of 15 minutes and a photovoltaic output dataset with a sampling time resolution of 5 minutes to verify the data dimension unification method.

[0142] Figure 5 The figure shows the upsampling results of user load data. For data that needs to be upsampled, two sampling points are equidistantly interpolated between the original two adjacent sampling points using the cubic spline interpolation algorithm. Figure 6 The figure shows the dimensionality reduction results of photovoltaic output data. For data requiring dimensionality reduction, the PAA method was used to divide the original 288-hour time series into 96 equal segments, each containing three sampling points. The average of each segment was taken to obtain the reduced-dimensionality series. Upsampling or dimensionality reduction can generate datasets with different time granularities required for analysis, enabling data fusion.

[0143] The preferred embodiments of the present invention are described in detail above with reference to the accompanying drawings. The embodiments described in the present invention are merely descriptions of the preferred embodiments of the present invention and do not limit the concept and scope of the present invention. The various specific technical features described in the above specific embodiments can be combined in any suitable manner unless there is any contradiction. Such combinations should also be regarded as the contents disclosed in this disclosure as long as they do not violate the concept of the present invention. In order to avoid unnecessary repetition, the present invention will not further describe various possible combinations.

[0144] The present invention is not limited to the specific details of the above-mentioned embodiments. Within the scope of the technical concept of the present invention and without departing from the design concept of the present invention, various modifications and improvements made to the technical solution of the present invention by those skilled in the art should fall within the scope of protection of the present invention. The technical contents for which protection is sought in the present invention have been fully recorded in the claims.

Claims

1. A multi-source power data collaborative fusion method taking into account data feature sensitivity analysis, characterized in that: The following steps are involved: Step 1: Obtain original multi-source power data, analyze and detect missing values ​​in the original multi-source power data and remove them, then identify and remove abnormal data in the remaining data, thereby obtaining a data sample after removing missing values ​​and abnormal data, providing support for the subsequent abnormal section data repair and fusion; Step 2: Perform sensitivity analysis on the data features of the data sample obtained in step 1 after removing missing values ​​and abnormal data, screen out the factors that have a more significant impact on the data sample, and then calculate the correlation coefficient between the influencing factor data group corresponding to the data abnormal moment and the influencing factor data group at other normal moments, and obtain the influencing factor data group with the largest correlation coefficient value. The value corresponding to this data group is the repair value at the abnormal moment, thereby obtaining the repaired data sample; Step 3: Upsample the low-dimensional data with low temporal resolution in the repaired data samples obtained in step 2, and perform dimensionality reduction with equal temporal resolution on the high-dimensional data in the repaired data samples obtained in step 2, thereby achieving the timestamp fusion and unification of the time series data samples.

2. A multi-source power data collaborative fusion method taking into account data feature sensitivity analysis according to claim 1, characterized in that: In step 1, the Pandas data analysis package is used to analyze and detect missing values ​​in the original multi-source power data and remove them.

3. The multi-source power data collaborative fusion method taking into account data feature sensitivity analysis according to claim 1 is characterized in that: In step 1, the isolation forest algorithm is used to identify abnormal data in the remaining data.

4. The multi-source power data collaborative fusion method taking into account data feature sensitivity analysis according to claim 3 is characterized in that: The isolation forest algorithm used mainly includes two parts: training isolation trees and anomaly scoring. The steps to build an isolation tree in the isolation forest algorithm are as follows: (1) A dataset X with N samples = {x i |i=1,2,…,N}, x i The data dimension is m. N sample points are randomly selected from the data set X to form a training set as the root node of the isolation tree. The data set X is constructed by removing the missing values ​​in step 1. (2) Randomly select a j dimension from the m dimensions of the dataset X and randomly select a split value p between the maximum and minimum values ​​of the dimension j in the training set; (3) Based on the hyperplane formed by the segmentation value p, the training set is divided into two subsets, the left subtree and the right subtree, to form leaf nodes; (4) Recursively perform steps (2) and (3) in the left and right subsets of the leaf node, and continuously generate new leaf nodes until the leaf node has only one data sample or the number of splits reaches the set value, then stop splitting, thus forming an isolated tree; Based on the above steps (1)-(4), multiple isolated trees are constructed to form an isolation forest for training, and the anomaly score is used to comprehensively judge the training results. For a data set containing n sample points, the anomaly score s(x,n) of the data sample x is shown in the following formula: c(n)=2(ln(n-1)+δ)-(2(n-1) / n) (2) In formulas (1) and (2), h(x) is the height of data sample x, that is, the number of segmentations experienced; E(h(x)) is the average value of h(x); δ is the correction term; c(n) is the normalization factor when the number of samples is n; The closer the s(x,n) value is to 1, the higher the probability of anomaly, and the closer it is to 0, the normal data. If most of the data scores are close to 0.5, it means there is no abnormal data. Therefore, the isolation forest algorithm is used to identify abnormal data in the remaining data after removing missing values, and the abnormal data is removed to obtain a data sample after removing missing values ​​and abnormal data.

5. The multi-source power data collaborative fusion method taking into account data feature sensitivity analysis according to claim 1 is characterized in that: In step 2, a multiple linear regression method is used to perform sensitivity analysis on the data characteristics of the data samples obtained in step 1.

6. The multi-source power data collaborative fusion method taking into account data feature sensitivity analysis according to claim 5 is characterized in that: The process of data feature sensitivity analysis using the multiple linear regression method is as follows: First, the mathematical expression of the MLR model is established as shown below: y=β0+β1x1+β2x2+…+β n x n +ε (3) In formula (3): y is the dependent variable; x1, x2,…, x n are independent variables; β0,β1,…,β n is the regression coefficient; ε is the error term; Suppose there are m groups of data samples, and the mth group of dependent variables in the m groups of data samples is y m , the independent variable is x m1 、x m2 ,……,x mn , then formula (3) can be converted into the following formula: Among them, ε1, ε2, …, ε m are the error terms of each group, and ε1,ε2,…,ε m Independent of each other; set up are β0,β1,…,β n The least squares estimate of is the observed value y k The estimated value of is as follows: Where: k = 1, 2, ..., m; e k is the estimation error; the independent variable is x k1 、x k2 ,……,x kn ; ε k is the error term for each group; The least squares method is used to estimate the dependent variable parameters of the MLR model to estimate the error e k The goal is to minimize the sum of squares Q, as shown in the following formula: The estimated value of the regression coefficient β can be obtained by solving the extreme value principle and the matrix equation; β0,β1,…,β n In the matrix solution, first construct a matrix H consisting of m groups of sample independent variable data as shown below: Among them, the independent variable is x 11 、x 12 、x 1n 、x 21 ,……,x mn ; Secondly, construct a column vector Y=[y1,y2,…,y m ] T , the regression coefficient β can be calculated using the following formula: β=(H T H) -1 H T Y=[β0,β1,β2,…,β n ] T (10) Where: dependent variables are y1, y2, ..., y m ; Using the data obtained in step 1 after removing missing values ​​and abnormal data, the regression coefficients β0, β1,…, β n , which can characterize the data feature sensitivity of the data sample obtained in step 2 after removing missing values ​​and abnormal data.

7. The multi-source power data collaborative fusion method taking into account data feature sensitivity analysis according to claim 1 is characterized in that: The correlation coefficient methods in step 2 are the Pearson correlation coefficient method and the Spearman correlation coefficient method. The Pearson correlation coefficient is used to analyze the linear correlation between data and is applicable to linearly correlated continuous variable data; the Spearman correlation coefficient is used for nonlinearly correlated or non-normally distributed data, focusing on measuring the monotonic correlation between data.

8. The multi-source power data collaborative fusion method taking into account data feature sensitivity analysis according to claim 7 is characterized in that: In the Pearson correlation coefficient method, the sample Pearson correlation coefficient ρ between two sets of data X and Y X,Y The calculation formula is as follows: In formula (11), cov(X,Y) is the covariance of data sets X and Y; σ X , σ Y are the standard deviations of X and Y respectively; n is the number of data points in the data set; x i and y i are the i-th data point in the data sets X and Y respectively; E(X) represents the variance of X; E(Y) represents the variance of Y; E(XY) represents the variance of XY; ρ X,Y The closer the value is to 1, the stronger the positive correlation between the data groups is; the closer it is to -1, the stronger the negative correlation between the data groups is. In the Spearman correlation coefficient method, the Spearman correlation coefficient ρ between two sets of data X and Y s The calculation formula is as follows: In formula (12), n is the number of data points in the data set; d i is the position difference between the corresponding i-th data point in X and Y after they are arranged in ascending order; ρ s The closer the value is to 1, the stronger the positive correlation between the data groups is; the closer it is to -1, the stronger the negative correlation between the data groups is. After finding the influencing data group with the greatest correlation with the influencing factor group of the data to be repaired, the corresponding data is used as the data to be repaired, thereby obtaining a repaired data sample.

9. The multi-source power data collaborative fusion method taking into account data feature sensitivity analysis according to claim 1 is characterized in that: In step 3, the cubic spline interpolation algorithm is used to upsample the low-dimensional data with low temporal resolution in the repaired data samples obtained in step 2.

10. The multi-source power data collaborative fusion method taking into account data feature sensitivity analysis according to claim 1 is characterized in that: In step 3, the segmented aggregation approximation method is used to perform dimensionality reduction of the high-dimensional data in the repaired data sample obtained in step 2 with equal time resolution.