Power failure recovery time prediction method based on big data analysis
By analyzing the influencing factors of power outage recovery from multiple dimensions and building an XGBoost-GAFT prediction model, the problem of dimensional focus on singularity and data imbalance in the existing technology is solved, and more efficient and accurate prediction of power outage recovery time is achieved.
Patent Information
- Application Number
- CN202510088517.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-30
AI Technical Summary
In the prediction of power outage recovery time, the dimensions focus on singularity, lack of systematicity and hierarchy, resulting in feature redundancy, low computing efficiency and high complexity, difficulty in processing unbalanced data, and ignoring the sub-region data characteristics with low power outage frequency.
The influencing factors of power outage recovery were analyzed from three dimensions: time dependence, space dependence and power facility status. The characteristics were screened through the mRMR method, a simple and efficient feature set was generated, and an XGBoost-GAFT prediction model was constructed to deal with complex power outage recovery time prediction and distribution fitting problems.
It effectively reduces complexity and computing costs, improves model generalization capabilities and accuracy, and can better handle unbalanced data and ignore sub-region data characteristics with low power outage frequency.
Smart Images

Figure CN120069181A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of big data technology, and specifically refers to a method for predicting power outage restoration time based on big data analysis. Background Art
[0002] The prediction of power outage restoration time based on big data analysis mainly relies on advanced data collection, storage, processing, and analysis technologies. By processing a large amount of complex data, it provides more accurate predictions, thereby effectively improving the stability and reliability of the power grid. Generally, the dimensions of power outage factors are too single, lacking systematicness and hierarchy. There is a lot of redundant information in the selected features, resulting in low computational efficiency and high complexity. Traditional power outage restoration time prediction methods are difficult to handle imbalanced data and easily ignore the data characteristics of sub-regions with low power outage frequencies, and have limitations in processing large-scale data. Summary of the Invention
[0003] In view of the above situation, to overcome the defects of the prior art, the present invention provides a method for predicting power outage restoration time based on big data analysis. Aiming at the problems that the dimensions of general power outage factors are too single, lacking systematicness and hierarchy, and there is a lot of redundant information in the selected features, resulting in low computational efficiency and high complexity, this solution analyzes the influencing factors of power outage time restoration in the target area from three dimensions: time dependence, space dependence, and power facility status, ensuring the comprehensiveness of feature construction and the diversity of data. Through the mRMR (Maximum Relevance Minimum Redundancy) method for feature selection, sub-features that are strongly correlated with the target variable but have low redundancy are effectively selected to generate a concise and efficient feature set, effectively reducing complexity and computational cost. Aiming at the problems that traditional power outage restoration time prediction methods are difficult to handle imbalanced data, easily ignore the data characteristics of sub-regions with low power outage frequencies, and have limitations in processing large-scale data, this solution constructs an XGBoost-GAFT prediction model to effectively handle complex power outage restoration time prediction and distribution fitting problems. By learning the residuals between the feature set and the target variable, it provides high model generalization ability and accuracy for situations with data imbalance and high uncertainty.
[0004] The technical solution adopted by the present invention is as follows: A method for predicting power outage restoration time based on big data analysis provided by the present invention, the method comprising the following steps:
[0005] Step S1: Data collection and preprocessing. Select the target area for which the power outage restoration time needs to be predicted, collect historical power outage data, power facility data, and geographical location data in the target area and integrate them into a data set. Preprocess the data set to obtain a preprocessed data set, and divide the preprocessed data set into a training set and a test set;
[0006] Step S2: Feature set construction. Divide the target area into sub-areas, set the target variable as the power outage restoration time, extract the feature vectors of each sub-area from the historical power outage data, power facility data, and geographical location data in the training set, and generate a feature set;
[0007] Step S3: Model construction. Construct an XGBoost-GAFT prediction model, input the feature set, and output the predicted values and their probabilities of the power outage restoration time for each sub-area;
[0008] Step S4: Model evaluation. Evaluate the performance of the XGBoost-GAFT prediction model on the test set, verify the clustering pattern of the predicted values of the power outage restoration time, and further optimize the parameters of the prediction model;
[0009] Step S5: Decision support. Use the optimized prediction model to predict the power outage restoration time, and provide support for power grid operation and maintenance, dispatching, and decision-making based on the prediction results.
[0010] Furthermore, in Step S2, the feature set construction specifically includes the following steps:
[0011] Step S21: Time feature construction. Divide the target area into sub-areas, calculate the historical power outage restoration time of each sub-area according to the historical power outage data in the training set, and use the time series analysis method to extract the long-term change trend of the restoration time from the historical power outage restoration time of each sub-area as the time feature of each sub-area;
[0012] Step S22: Spatial feature construction. Construct a regional adjacency matrix and a regional spatial weight matrix according to the geographical location data in the training set, and extract the spatial features of each sub-area;
[0013] Step S23: Power facility feature construction. Map the power facility data in the training set into sub-areas and convert it into a numerical vector as the power facility feature of each sub-area;
[0014] Step S24: Feature combination. Combine the time features, spatial features, and power facility features of each sub-area into a feature vector;
[0015] Step S25: Feature screening. Use the mRMR method to check the relationship between the feature vectors of each sub-area and the target variable, and select the features strongly related to the target variable in the feature vectors of each sub-area to generate a feature set. The formula used is as follows: ;
[0016] In the formula, represents the maximum correlation minimum redundancy value of the sub-area, represents the th eigenvalue in the feature vector of the sub-area, The th eigenvalue in the eigenvector representing the sub-region, represents the set of eigenvectors containing all sub-regions, represents the subset of eigenvectors of sub-regions filtered from the set , represents the target variable, represents the feature and the mutual information value with the target variable , represents the mutual information value between the feature and the feature , represents selecting the eigenvector subset in the set to maximize the mutual information value and minimize it.
[0017] Furthermore, in step S222, the construction of the spatial feature specifically includes the following steps:
[0018] Step S221: Determine the adjacency relationship between sub-regions. If sub-region i and sub-region j have a common boundary line, then sub-region i and sub-region j are considered adjacent regions;
[0019] Step S222: Obtain the total number n of sub-regions and construct a matrix of size as the regional adjacency matrix, and the element value in the matrix represents the adjacency relationship between sub-regions;
[0020] Step S223: Calculate the distance between each pair of sub-regions and convert it into a regional spatial weight. The formula used is as follows: ;
[0021] In the formula, represents the regional spatial weight, represents the distance between sub-region i and sub-region j;
[0022] Step S224: Construct a matrix of size as the regional spatial weight matrix, and the element value in the matrix represents the regional spatial weight;
[0023] Step S225: Normalize the regional spatial weight matrix so that the sum of the element values in each row of the regional spatial weight matrix is 1. The formula used is as follows: ;
[0024] In the formula, represents the normalized regional geographical location weight;
[0025] Step S226: Calculate the spatial autocorrelation for each sub-region based on the regional adjacency matrix and the regional spatial weight matrix, and output the spatial autocorrelation coefficient as the spatial feature of each sub-region. The formula used is as follows: ;
[0026] In the formula, represents the spatial autocorrelation coefficient, respectively represent the power outage restoration times of sub-region i and sub-region j, represents the average value of the historical power outage restoration times of all sub-regions.
[0027] Furthermore, in step S3, the model construction specifically includes the following steps:
[0028] Step S31: Input the feature set and obtain the feature dimension T of the feature set;
[0029] Step S32: Construct an XGBoost model, initialize the number of decision trees, and let the decision trees learn the residuals between the feature set and the target variable and predict the power outage restoration time, and output the predicted values of each layer of decision trees;
[0030] Step S33: Perform hierarchical iteration, superimpose the predicted values of the th layer and the output of the th decision tree, and fit the residuals between the output of the previous layer and the target variable through the gradient boosting method;
[0031] Step S34: Optimize the parameters, update the predicted values, and when the fitting effect reaches the best, output the predicted values of the power outage restoration time of the last layer;
[0032] Step S35: Construct a GAFT model, integrate the predicted values of the power outage restoration time output by the XGBoost model with the feature set as the input features of the GAFT model, and the feature dimension of the input features is T + 1;
[0033] Step S36: Construct a survival function to fit the survival probability distribution of the power outage restoration time;
[0034] Step S37: Introduce random effects to capture the individual differences between sub-regions by introducing the random effects between sub-regions;
[0035] Step S38: Output the survival probability distribution to describe the predicted values of the power outage restoration time of sub-regions and their probabilities.
[0036] The beneficial effects achieved by the present invention using the above solution are as follows:
[0037] (1) Regarding the problem that the dimension of attention to general power outage factors is too single, lacking systematicness and hierarchy, with a lot of redundant information in the selected features, resulting in low calculation efficiency and high complexity, this solution analyzes the influencing factors of power outage time recovery in the target area from three dimensions: time dependence, spatial dependence, and the condition of power facilities, ensuring the comprehensiveness of feature construction and the diversity of data. Through the mRMR method for feature screening, sub-features that are strongly correlated with the target variable but have low redundancy are effectively selected, generating a concise and efficient feature set, effectively reducing complexity and calculation costs.
[0038] (2) Regarding the problem that traditional power outage recovery time prediction methods are difficult to handle unbalanced data, easily ignore the data characteristics of sub-regions with low power outage frequencies, and have limitations in processing large-scale data, this solution constructs an XGBoost-GAFT prediction model, effectively dealing with complex power outage recovery time prediction and distribution fitting problems. By learning the residuals between the feature set and the target variable, it provides high model generalization ability and accuracy for scenarios with data imbalance and high uncertainty. Brief Description of the Drawings
[0039] Figure 1 It is a schematic flowchart of a power outage recovery time prediction method based on big data analysis proposed by the present invention;
[0040] Figure 2 It is a schematic flowchart of step S2.
[0041] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention. Specific Embodiments
[0042] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments; based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0043] Example 1, refer to Figure 1 A power outage recovery time prediction method based on big data analysis provided by the present invention, the method includes the following steps:
[0044] Step S1: Data collection and preprocessing. Select the target area for power outage restoration time prediction, collect historical power outage data, power facility data, and geographical location data within the target area from the publicly available power information database and integrate them into a dataset. The historical power outage data includes the power outage area, power outage start time, and power outage end time. The power facility data includes the number of substations, grid density, and power load. The geographical location data includes coordinate information and population density. Preprocess the dataset to obtain a preprocessed dataset. Randomly select 70% of the preprocessed dataset as the training set and 30% as the test set.
[0045] Step S2: Feature set construction. Divide the target area into sub-areas, set the target variable as the power outage restoration time, extract the feature vectors of each sub-area from the historical power outage data, power facility data, and geographical location data in the training set, and generate a feature set.
[0046] Step S3: Model construction. Construct an XGBoost-GAFT prediction model, input the feature set, and output the predicted power outage restoration time value and its probability for each sub-area.
[0047] Step S4: Model evaluation. Evaluate the performance of the XGBoost-GAFT prediction model on the test set, verify the clustering pattern of the predicted power outage restoration time values, and further optimize the parameters of the prediction model.
[0048] Step S5: Decision support. Use the optimized prediction model to predict the power outage restoration time and provide support for power grid operation and maintenance, scheduling, and decision-making based on the prediction results.
[0049] Example 2. Refer to Figure 1 , this example is based on the above example. In step S1, preprocess the dataset, which specifically includes the following steps:
[0050] Step S11: Data cleaning. Remove redundant and invalid data from the dataset to obtain a cleaned dataset.
[0051] Step S12: Missing value handling. Fill the missing values in the dataset using the mean method to obtain a filled dataset.
[0052] Step S13: Outlier detection. Detect the outliers in the filled dataset and process the outliers using the standard deviation method to obtain a preprocessed dataset.
[0053] Example 3. Refer to Figure 1 and Figure 2 , this example is based on the above example. In step S2, feature set construction specifically includes the following steps:
[0054] Step S21: Time feature construction. Divide the target area into sub-areas, calculate the historical power outage restoration time of each sub-area according to the historical power outage data in the training set, and use the time series analysis method to extract the long-term change trend of the restoration time from the historical power outage restoration time of each sub-area as the time feature of each sub-area;
[0055] Step S22: Space feature construction. Construct the space features of each sub-area according to the geographical location data in the training set, including the following steps:
[0056] Step S221: Determine the adjacency relationship between sub-areas. If sub-area i and sub-area j have a common boundary line, it is considered that sub-area i and sub-area j are adjacent areas;
[0057] Step S222: Obtain the total number n of sub-areas, and construct a matrix of size as the area adjacency matrix. The element value in the matrix represents the adjacency relationship between sub-areas; if sub-area i and sub-area j are adjacent areas, the element value is 1, otherwise the element value is 0;
[0058] Step S223: Calculate the distance between each sub-area and convert it into the area space weight. The formula used is as follows: ;
[0059] In the formula, represents the area space weight, represents the distance between sub-area i and sub-area j;
[0060] Step S224: Construct a matrix of size as the area space weight matrix. The element value in the matrix represents the area space weight;
[0061] Step S225: Normalize the area space weight matrix so that the sum of the element values in each row of the area space weight matrix is 1. The formula used is as follows: ;
[0062] In the formula, represents the normalized area geographical location weight;
[0063] Step S226: Calculate the spatial autocorrelation for each sub-area based on the area adjacency matrix and the area space weight matrix, and output the spatial autocorrelation coefficient as the space feature of each sub-area. The formula used is as follows: ;
[0064] In the formula, represents the spatial autocorrelation coefficient, respectively represent the power outage restoration times of sub-region i and sub-region j, represents the average value of the historical power outage restoration times of all sub-regions;
[0065] Step S23: Power facility feature construction. Map the power facility data in the training set to the sub-regions and convert it into a numerical vector as the power facility feature of each sub-region;
[0066] Step S24: Feature combination. Combine the time features, space features, and power facility features of each sub-region into a feature vector;
[0067] Step S25: Feature screening. Use the mRMR method to check the relationship between the feature vector of each sub-region and the target variable, and select the features strongly related to the target variable of each sub-region to generate a feature set. The formula used is as follows: ;
[0068] In the formula, represents the maximum correlation and minimum redundancy value of the sub-region, represents the th eigenvalue in the feature vector of the sub-region, represents the th eigenvalue in the feature vector of the sub-region, represents the set containing the feature vectors of all sub-regions, represents the feature subset of the sub-region selected from the set , represents the target variable, represents the feature and the target variable 's mutual information value, represents the feature and the feature 's mutual information value, represents selecting the feature subset in the set to maximize the mutual information value and minimize it.
[0069] By performing the above operations, for the problem that the dimensions concerned by general power outage factors are too single, lacking systematicness and hierarchy, and the screened features have a lot of redundant information, resulting in low computational efficiency and high complexity, this solution analyzes the influencing factors of the power outage time restoration in the target area from three dimensions: time dependence, space dependence, and power facility status, ensuring the comprehensiveness of feature construction and the diversity of data. Through the mRMR method for feature screening, it effectively selects sub-features that are strongly related to the target variable but have low redundancy, generates a concise and efficient feature set, and effectively reduces complexity and computational cost.
[0070] Example 4, refer to Figure 1 , based on the above example, in step S21, the time series analysis method adopts the polynomial fitting method, and uses polynomial fitting to analyze the changing trend of the historical power outage recovery time of each sub-region over time, and extracts the trend value as the time feature of each sub-region.
[0071] Example 5, refer to Figure 1 , based on the above example, in step S3, model construction specifically includes the following steps:
[0072] Step S31: Input the feature set and obtain the feature dimension T of the feature set;
[0073] Step S32: Construct an XGBoost model, initialize the number of decision trees. The nodes of the decision tree represent the thresholds of the features, the branches represent the decision paths, and the leaves represent the predicted values of the power outage recovery time. The decision tree learns the residuals between the feature set and the target variable and predicts the power outage recovery time, and outputs the predicted values of each layer of the decision tree. The formula used is as follows: ;
[0074] In the formula, represents the residual between the features of sub-region i and the target variable, represents the actual value of the target variable, represents the layer's predicted value;
[0075] Step S33: Hierarchical iteration, add the predicted value of the layer and the output of the th decision tree, and fit the residual between the output of the previous layer and the target variable through the gradient boosting method. The formula used is as follows: ;
[0076] In the formula, represents the learning rate, represents the predicted value of the layer, represents the output of the th tree;
[0077] Step S34: Parameter optimization, adjust the depth and learning rate of the decision tree, update the predicted value, and when the fitting effect reaches the best, output the predicted value of the power outage recovery time of the last layer;
[0078] Step S35: Construct a GAFT model, integrate the predicted value of the power outage recovery time output by the XGBoost model with the feature set as the input feature of the GAFT model, and the feature dimension of the input feature is T + 1;
[0079] Step S36: Construct a survival function to fit the survival probability distribution of the power outage recovery time. The formula used is as follows: ; ;
[0080] In the formula, represents the survival probability of the power outage recovery time at time represents the exponential function, represents the input feature, represents the input feature of the linear combination, represents the coefficient vector, represents the shape parameter of the survival probability distribution;
[0081] Step S37: Introduce random effects. Introduce the random effects between sub-regions to capture the individual differences between sub-regions. The formula used is as follows: ;
[0082] In the formula, represents the rate function of the power outage recovery time, represents the predicted rate based on the input features, represents the random effect of sub-region i;
[0083] Step S38: Output the survival probability distribution, describing the predicted value of the power outage recovery time in the sub-region and its probability.
[0084] By performing the above operations, for the problem that traditional power outage recovery time prediction methods are difficult to handle unbalanced data, easily ignore the data characteristics of sub-regions with low power outage frequencies, and have limitations in processing large-scale data, this solution constructs an XGBoost-GAFT prediction model, effectively dealing with complex power outage recovery time prediction and distribution fitting problems. By learning the residuals between the feature set and the target variable, it provides high model generalization ability and accuracy for scenarios with data imbalance and high uncertainty.
[0085] Example 6, refer to Figure 1 , this example is based on the above example. In step S4, the Moran's I method is used to verify the clustering pattern of the predicted values of the power outage recovery time, and the model parameters are further adjusted through the calculation results of the spatial distance between the predicted values and the actual values of the power outage recovery time in the test set.
[0086] It should be noted that, in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device.
[0087] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
[0088] The above description of the present invention and its embodiments is not restrictive. What is shown in the drawings is only one of the embodiments of the present invention, and the actual structure is not limited thereto. In general, if those of ordinary skill in the art are inspired by it and, without departing from the purpose of the present invention, design similar structural forms and embodiments to this technical solution without creative efforts, they should all fall within the protection scope of the present invention.
Claims
1. A method for predicting power outage restoration time based on big data analysis, characterized in that: The method comprises the following steps: Step S1: data collection and preprocessing, selecting a target area for power outage restoration time prediction, collecting historical power outage data, power facility data, and geographic location data in the target area and integrating them into a data set, preprocessing the data set to obtain a preprocessed data set, and dividing the preprocessed data set into a training set and a test set; Step S2: constructing a feature set, dividing the target area into sub-areas, setting the target variable as the power outage recovery time, extracting the feature vector of each sub-area from the historical power outage data, power facility data, and geographic location data of the training set, and generating a feature set; Step S3: Model construction, constructing the XGBoost-GAFT prediction model, inputting the feature set, and outputting the predicted value and probability of power outage restoration time for each sub-region; Step S4: Model evaluation: evaluate the performance of the XGBoost-GAFT prediction model on the test set, verify the clustering pattern of the power outage restoration time prediction value, and further optimize the parameters of the prediction model; Step S5: Decision support, using the optimized prediction model to predict the power outage restoration time, and providing support for power grid operation and maintenance, scheduling and decision-making based on the prediction results.
2. The method for predicting power outage restoration time based on big data analysis according to claim 1 is characterized in that: In step S2, the feature set is constructed, including the following steps: Step S21: constructing time features, dividing the target area into sub-areas, calculating the historical power outage recovery time of each sub-area based on the historical power outage data in the training set, and using the time series analysis method to extract the long-term change trend of the recovery time from the historical power outage recovery time of each sub-area as the time feature of each sub-area; Step S22: constructing spatial features, constructing a regional adjacency matrix and a regional spatial weight matrix based on the geographical location data in the training set, and extracting the spatial features of each sub-region; Step S23: constructing the power facility features, mapping the power facility data in the training set to the sub-regions and converting them into numerical vectors as the power facility features of each sub-region; Step S24: feature combination, combining the time features, spatial features, and power facility features of each sub-region into a feature vector; Step S25: Feature screening: Use the mRMR method to check the relationship between the feature vector of each sub-region and the target variable, and select the features that are strongly correlated with the target variable in the feature vector of each sub-region to generate a feature set. The formula used is as follows: ; In the formula, represents the maximum relevant minimum redundancy value of the sub-region, The first feature vector of the subregion eigenvalues, The first feature vector of the subregion eigenvalues, represents the set of feature vectors containing all sub-regions, Represents from the set Filter out the feature subset of the sub-region, represents the target variable, Representation characteristics With the target variable The mutual information value of Representation characteristics With features The mutual information value of Indicates in the collection Select feature subsets Make the mutual information value Maximize and minimize.
3. The method for predicting power outage restoration time based on big data analysis according to claim 2 is characterized in that: In step S22, the spatial feature construction includes the following steps: Step S221: determining the adjacency relationship between sub-regions. If sub-region i and sub-region j have a common boundary line, then sub-region i and sub-region j are considered to be adjacent regions. Step S222: Obtain the total number of sub-regions n and construct a The matrix is used as the regional adjacency matrix, and the element values in the matrix represent the adjacency relationship between sub-regions; Step S223: Calculate the distance between each sub-region and convert it into regional spatial weight. The formula used is as follows: ; In the formula, represents the regional spatial weight, represents the distance between sub-region i and sub-region j; Step S224: Construct a The matrix is used as the regional spatial weight matrix, and the element values in the matrix represent the regional spatial weights; Step S225: normalize the regional spatial weight matrix so that the sum of the element values in each row of the regional spatial weight matrix is 1. The formula used is as follows: ; In the formula, Indicates the normalized regional geographic location weight; Step S226: Calculate the spatial autocorrelation of each sub-region based on the regional adjacency matrix and the regional spatial weight matrix, and output the spatial autocorrelation coefficient as the spatial feature of each sub-region. The formula used is as follows: ; In the formula, represents the spatial autocorrelation coefficient, denote the power outage recovery time of sub-area i and sub-area j respectively, Represents the average historical power outage restoration time of all sub-areas.
4. The method for predicting power outage restoration time based on big data analysis according to claim 1 is characterized in that: In step S3, the model is constructed, including the following steps: Step S31: input a feature set and obtain a feature dimension T of the feature set; Step S32: construct an XGBoost model, initialize the number of decision trees, learn the residual between the feature set and the target variable, and predict the power outage restoration time, and output the predicted value of each layer of decision trees; Step S33: Layered iteration, The predicted value of the layer and the The outputs of the decision trees are superimposed, and the residual between the output of the previous layer and the target variable is fitted by the gradient boosting method; Step S34: Optimize parameters and update the predicted value. When the fitting effect reaches the best, output the predicted value of the power outage restoration time of the last layer. Step S35: construct a GAFT model, integrate the predicted value of the power outage restoration time output by the XGBoost model with the feature set as the input feature of the GAFT model, and the feature dimension of the input feature is T+1; Step S36: constructing a survival function to fit the survival probability distribution of power outage recovery time; Step S37: introducing random effects, introducing random effects between sub-regions to capture personalized differences between sub-regions; Step S38: Output the survival probability distribution, describing the predicted value of the power outage restoration time in the sub-area and its probability.