Power failure classification prediction method and system based on LSTM and improved XGBOOST

Through the power failure classification prediction method based on LSTM and improved XGBOOST, the space-time alignment and fault tracing of multi-source equipment data in the new power system are realized, and the fault prediction problems caused by equipment heterogeneity and new energy fluctuations are solved, and the accuracy of fault identification and the real-time control capabilities of the power grid are improved.

CN120429693APending Publication Date: 2025-08-05NANJING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510596482.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

In the new power system, the degree of equipment heterogeneity has deepened and the system shape is complex. The existing fault prediction methods are difficult to achieve fault data traceability and precise correction of equipment layer, and they are weak when the new energy power fluctuates randomly, and data repair accuracy and timeliness are difficult to meet real-time control needs.

Method used

The power fault classification prediction method based on LSTM and improved XGBOOST is adopted, and the LSTM prediction model and XGBOOST fault analysis model are trained to perform spatiotemporal alignment of multi-source equipment data, and fault traceability is traced using the power fault knowledge graph, and the data importance is identified through feature weighting and hysteresis feature changes, which improves the accuracy of fault identification.

Benefits of technology

It improves the accuracy and efficiency of power failure prediction, reduces the error rate of traditional models, enhances the real-time monitoring of the status of power grid equipment, and meets the safe and stable operation needs of the new power system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120429693A_ABST
    Figure CN120429693A_ABST
Patent Text Reader

Abstract

The invention discloses a power fault classification prediction method and system based on LSTM and improved XGBOOST. The method comprises the following steps: training an LSTM prediction model and an XGBOOST fault analysis model according to historical data of a power system; the problems that the number of heterogeneous devices is large, data space and time are not aligned, sampling data are missing, electrical fault features are prone to being confused and time sequence features are not obvious in power faults under a novel power system are solved, on the basis that an existing LSTM-XGBOOST static model prediction result is accurate, time sequence data input into an XGBOOST fault analysis model are subjected to targeted processing, and the fault analysis accuracy is improved. And the identification accuracy is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power system fault analysis, and in particular to a power fault classification and prediction method and system based on LSTM and improved XGBOOST. Background Art

[0002] As a vital component of the energy system, regional power grids have become a key area of focus in the development of new power systems. Compared to traditional power systems, these new grids exhibit three disruptive characteristics: critical penetration of renewable energy, large-scale deployment of power electronics, and ubiquitous access to distributed, adjustable resources. These changes have led to increased heterogeneity in grid equipment and a more complex system architecture, placing stricter demands on advanced grid applications.

[0003] Faced with the exponential growth of device terminals and massive data streams, spatiotemporal synchronization deviations and data loss are becoming increasingly prominent. Existing expert-rule-based pre-verification mechanisms and robust estimation methods have significant limitations. Mainstream solutions often employ weight adjustment or erroneous data exclusion strategies. While these strategies can improve the robustness of state estimation, they cannot accurately trace fault data and accurately correct errors at the device level. Decision-support systems based primarily on power flow optimization struggle to correct inherent errors in device models, posing the risk of misjudgment. Traditional linear interpolation methods struggle to cope with random fluctuations in renewable energy power, and existing algorithms struggle to exhaust complex scenarios. In key areas such as power forecasting and device state inversion, data repair accuracy and timeliness struggle to meet real-time control requirements. The nonlinear mapping relationship between massive amounts of heterogeneous data and device states has not yet been effectively decoupled, leaving technical gaps in device health diagnosis and event identification based on measurement data. This directly impacts the accuracy of fault warnings and the efficiency of response, hindering lean grid operations.

[0004] It is noteworthy that the reliable operation of new power systems is highly dependent on the integrity and accuracy of data links. There is an urgent need to develop an intelligent verification system for the spatiotemporal alignment of multi-source data, develop dynamic data repair algorithms adapted to the characteristics of new energy sources, and construct a new power system data feature extraction and event data traceability system. These technological breakthroughs will directly determine the grid's situational awareness capabilities and risk prevention and control capabilities, providing a core guarantee for the safe and stable operation of new power systems. Summary of the Invention

[0005] Technical purpose: In response to the shortcomings of existing power system fault prediction, the present invention discloses a power fault classification prediction method and system based on LSTM and improved XGBOOST.

[0006] Technical solution: To achieve the above technical objectives, the present invention adopts the following technical solution:

[0007] A power fault classification prediction method based on LSTM and improved XGBOOST includes the following steps:

[0008] S01. Train the LSTM prediction model and XGBOOST fault analysis model based on historical power system data.

[0009] S02. Acquire measurement data of multiple source devices in the power system;

[0010] S03. Sampling the measured data to obtain a sampling number and performing spatiotemporal alignment, inputting the sampling data into the LSTM prediction model, generating LSTM prediction data, and completing fault feature extraction of the sampling data;

[0011] S04. Generate a residual sequence based on the difference between the LSTM predicted data and the corresponding sampled data, use the obtained residual sequence to construct a power fault knowledge graph, and use the power fault knowledge graph to trace the fault;

[0012] S05. Input the fault characteristics into the XGBOOST fault analysis model to obtain a fault classification result, use the fault classification result to verify the power fault knowledge graph, and output the fault type according to the verification result.

[0013] Preferably, in step S01 of the present invention, the process of training the LSTM prediction model includes: constructing a power grid data matrix based on the historical data of the power system, and when constructing the power grid data matrix, first normalizing the power data matrix, and mapping the time data corresponding to the power data to a sine function, and constructing a training array with the time data represented by the sine function and the corresponding normalized power data matrix, and inputting them into the LSTM model for training.

[0014] Preferably, in step S03 of the present invention, performing spatiotemporal alignment on the sampled data includes: establishing a cubic spline interpolation function of the sampled data, and performing interpolation processing on the sampled data so that the synchronization error of the sampled data of the multi-source devices is zero after the processing.

[0015] Preferably, in step S05 of the present invention, when the fault feature is input into the XGBOOST fault analysis model, the XGBOOST fault analysis model is first used to extract the importance of data related to the fault feature in the power system historical data, and the input fault feature is feature weighted according to the data importance, and weights are assigned to the data corresponding to the fault feature.

[0016] Preferably, when the LSTM prediction model generates prediction data based on the sampled data, the present invention performs missing check on the sampled data through the LSTM prediction model, uses the prediction data to fill the missing data in the sampled data, and after filling the sampled data, extracts the fault characteristics of the sampled data.

[0017] Preferably, when training the XGBOOST fault analysis model based on the power system historical data, the present invention sorts the power system historical data by time, periodically encodes the time parameters of the power system historical data, cuts the power system historical data according to the periodic coding, and uses the cut power system historical data arranged in time sequence to train the XGBOOST fault analysis model.

[0018] Preferably, when the present invention periodically encodes the time data of the power system historical data, the hour value of each day is mapped to the corresponding unit circle with a period of 24 and the date corresponding to the data is mapped to a period of 7 days, and the time data of the power system historical data is encoded using a sine function, and the time data is corresponded to the power data of the power system historical data.

[0019] Preferably, when the present invention trains the XGBOOST fault analysis model, the importance of the data to the power fault is identified through the changes in hysteresis characteristics and the data changes in rolling window statistics.

[0020] Preferably, when the present invention outputs the fault type according to the verification result, if the fault tracing verification of the power fault knowledge graph passes, the output of the power fault knowledge graph tracing is used as the result. If the verification fails, the fault type obtained by the knowledge graph tracing is judged. If the fault type is a characteristic easily confused fault, the output of the XGBO0ST fault analysis model is used as the result. If the fault type is an instantaneous fault with little dependence on time, the output of the power fault knowledge graph tracing is used as the result.

[0021] The present invention also discloses a power fault classification prediction system based on the above-mentioned power fault classification prediction method based on LSTM and improved XGBO0ST.

[0022] Beneficial effects: The present invention, through the disclosed LSTM-based and improved XGBOOST-based power fault classification and prediction method and system, performs spatiotemporal alignment on data from multiple source devices, and can perform synchronous analysis and processing. According to the time series and PQI features obtained by LSTM processing the data, the data input into the XGBOOST fault analysis model is feature weighted, and the time series data is preprocessed, so that the model can directly understand the periodicity of linear time encoding and obtain fault classification results; introducing lag features and rolling window statistics when training the XGBOOST fault analysis model can enhance the time series expression capability of power data and better identify some faults with similar power data changes but obvious time series features. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for describing the embodiments or the prior art.

[0024] Figure 1 This is a flow chart of the present invention for fault classification and prediction of power system data streams;

[0025] Figure 2 This is a flow chart of the spatiotemporal alignment process of measurement data according to the present invention;

[0026] Figure 3 The PQI, true PQI, and residual curves predicted by the LSTM prediction model when there is no spatiotemporal alignment;

[0027] Figure 4 The PQI, true PQI, and residual curves predicted by the LSTM prediction model after spatiotemporal alignment;

[0028] Figure 5 This is a schematic diagram of the power failure knowledge graph of the present invention. DETAILED DESCRIPTION

[0029] Reference will now be made in detail to the embodiments of the present disclosure, one or more examples of which are set forth herein below. Each embodiment and example is provided by way of explanation of the apparatus, composition, and materials of the present disclosure, and is not intended to be limiting. On the contrary, the following description provides a convenient illustration of exemplary embodiments for implementing the present disclosure. In fact, it will be apparent to those skilled in the art that various modifications and variations can be made within the teachings of the present disclosure without departing from the scope or spirit of the present disclosure.

[0030] like Figure 1 and Figure 2 As shown, the present invention discloses a power fault classification prediction method based on LSTM and improved XGBOOST, comprising the steps of:

[0031] S01. Train the LSTM prediction model and XGBOOST fault analysis model based on historical power system data.

[0032] The process of training the LSTM prediction model includes: constructing a power grid data matrix based on the historical data of the power system. When constructing the power grid data matrix, the power data matrix is first normalized, and the time data corresponding to the power data is mapped to a sine function. The time data represented by the sine function and the corresponding normalized power data matrix are used to construct a training array and input into the LSTM model for training.

[0033] When mapping the time data corresponding to the power data to a sine function, [month, day] and [hour, minute] are mapped to the sine functions day and clock, respectively. Month represents the month, day represents the date, and hour and minute represent the hour and minute corresponding to the power data, respectively. Finally, the array [day, clock, [P, Q, I]] is constructed and converted into a two-dimensional array as training data for the LSTM long short-term memory neural network. P represents the active power of the power system, Q represents the reactive power of the power system, and I represents the current of the power system.

[0034] When training the XGBOOST fault analysis model based on the power system historical data, the present invention sorts the power system historical data by time, periodically encodes the time parameters of the power system historical data, segments the power system historical data according to the periodic encoding, and uses the segmented time-series power system historical data to train the XGBOOST fault analysis model.

[0035] When the present invention periodically encodes the time data of the power system historical data, the hour value of each day is mapped to the corresponding unit circle with a period of 24 and the date corresponding to the data is mapped to a period of 7 days. The sine function is used to encode the time data of the power system historical data, and the time data is corresponded to the power data of the power system historical data.

[0036] Specifically, the hour value is mapped to the unit circle with a period of 24. For example, the difference between Hour = 12 (noon) and Hour = 0 (midnight) on the unit circle is π, and the signs of their sine and cosine values are opposite. Map the days of the week onto the unit circle with a period of 7. For example, the difference between day_week = 3 (Thursday) and day_week = 6 (Sunday) on the unit circle is

[0037] Through sinusoidal function encoding, the code values of adjacent times change continuously on the unit circle, preserving the continuity of the data, avoiding the timing adjustment from being regarded as discrete linear features, and enhancing the learning of periodicity. Therefore, the periodicity of linear time encoding can be directly understood through the XGBOOST model, and the fault classification results can be obtained based on the input power data.

[0038] When the present invention trains the XGBOOST fault analysis model, the importance of data to power faults is identified through changes in hysteresis features and data statistics of rolling window statistics. The hysteresis is in hours, and the current time is taken as t, then t-1 represents the previous hour, and t-24 represents the same time of the previous day. The autoregressive model is used to predict the current time using the previous hour and the same time of the previous day, which can capture the periodic changes of the data. The rolling window statistics time window takes the average value and mean square error of the 12 hours before and after. By combining rolling window statistics with hysteresis features, the model's sensitivity to trends and volatility of event sequence data can be enhanced, and the identification of power data faults with time characteristics can be strengthened. Some faults with similar power data changes but obvious time series characteristics can be better identified, such as abnormal mutations and short-term strong fluctuations in power data. The former is sporadic, and the latter is a drastic and repeated change over a period of time. The data statistical identification method of the present invention can ensure the analysis accuracy of the training model, and improve the parameter weight of the power data corresponding to the fault characteristics in the subsequent power system fault analysis.

[0039] After completing the training of the LSTM prediction model and the XGBOOST fault analysis model, the fault analysis and prediction of the power system can be performed according to the process of steps S02-S05.

[0040] S02. Acquire measurement data of multiple source devices in the power system;

[0041] S03. Sampling the measured data to obtain sampled data and performing spatiotemporal alignment, inputting the sampled data into the LSTM prediction model to generate LSTM prediction data and extracting fault features from the sampled data, wherein the fault features include time series features and power data features;

[0042] The process of performing spatiotemporal alignment of the sampled data includes: establishing a cubic spline interpolation function of the sampled data, performing interpolation processing on the sampled data, and making the synchronization error of the sampled data of the multi-source devices after the processing be zero.

[0043] The interpolation process is: cubic spline interpolation function:

[0044]

[0045] Among them, i and T are sampling number and sampling time respectively;

[0046] S i-1 、S i are the i-1th and i-th sampling data;

[0047] S′ i-1 , S′ i is the first-order derivative of the i-1th and i-th sampling data;

[0048] T i-1 ≤T≤T i , ΔT i is the sampling interval between the i-1th and i-th sampling data points, ΔT i =T i -T i-1 .

[0049] After normalization, the cubic spline interpolation function changes to:

[0050] S i (T) = k1S i-1 +k2S i +k3S′ i-1 +k4S′ i ;

[0051] A WGAN network is constructed and trained based on the data characteristics of multi-source heterogeneous equipment in a new power system. The trained WGAN network is then used to generate interpolated data to supplement the sampled data. The original sampling sequence f(s) is first input into the system. Based on this original sampling sequence f(s), a new data sequence F(s) with a sampling interval of 10 seconds is constructed using the WGAN-generated data. The optimal correction value S' is calculated based on this new data sequence, resulting in a new sampling sequence f(s, s'). The time difference between the data time and the sampling and transmission processes is obtained. A cubic spline interpolation function is then used to calculate the optimal correction value for this new sampling sequence, correcting the previous and next data. After the correction is completed, a corrected sampling sequence f'(s, s') is generated, and the corrected data is output.

[0052] S04. Generate a residual sequence based on the difference between the LSTM predicted data and the corresponding sampled data, use the obtained residual sequence to construct a power fault knowledge graph, and use the power fault knowledge graph to trace the fault;

[0053] After the spatiotemporally aligned sampling data is input into the LSTM prediction model, the LSTM prediction model outputs a normalized prediction matrix, which is then inversely mapped to obtain the predicted value, thereby obtaining the LSTM predicted data and the actual sampled data. Fault features are then extracted from the data, and the difference between the LSTM predicted data and the actual sampled data is used to generate a residual sequence:

[0054] r t =y t -p t ;

[0055] y t : the true value at time t;

[0056] p t : The predicted value of the LSTM model;

[0057] r t: Residual, reflecting the error of model prediction;

[0058] When the model prediction is accurate and the data is normal, the residual should satisfy: r t It follows a normal distribution with mean 0.

[0059] like Figure 3 Figure 4 As shown, Figure 4 for Figure 3 The actual value, predicted value, and residual curves of the LSTM prediction model output after the data is processed by time and space alignment. Both the actual value and the predicted value include active power P, reactive power Q, and current characteristics I. It can be seen that when the system is working normally, the actual active power P, reactive power Q, and current I of the real sampling data are basically consistent with the predicted values. The residual curve shows that when the system model changes, the residual will change significantly. After time and space alignment, the characteristics of the data are still retained and have little impact on the prediction results of the LSTM prediction model.

[0060] like Figure 5 As shown, the non-Gaussian characteristics of the residual distribution of data under fault conditions can be used to extract the dynamic evolution patterns of multiple operating parameters. Relying on the power fault knowledge graph built with the Neo4j graph database, a Match statement is used to implement subgraph matching of feature vectors, generate preliminary traceability hypotheses, and identify fault events based on the sampled data and the power fault knowledge graph input into the model. The sampled data is then input into the XGBOOST fault analysis model for analysis and verification. When the LSTM prediction model generates predicted data based on the sampled data, it performs a missing check on the sampled data, uses the predicted data to fill in the missing data in the sampled data, extracts the fault features of the sampled data after filling in the sampled data, and marks the time points of the missing data. Finally, the data obtained after the missing check is output to the XGBOOST fault analysis model along with the previously extracted time series features.

[0061] S05. Input the fault characteristics into the XGBOOST fault analysis model to obtain a fault classification result, use the fault classification result to verify the power fault knowledge graph, and output the fault type according to the verification result.

[0062] In step S05, when the fault features are input into the XGBOOST fault analysis model, the XGBOOST fault analysis model first extracts the importance of data related to the fault features from the power system historical data. The input fault features are then weighted according to the data importance, and the corresponding data of the fault features are weighted. The weights can also be set based on the weights of relevant features set based on experience and combined with the weights corresponding to the data importance. After the data is input into the XGBOOST fault analysis model, the corresponding analysis results are output, and the output of the LSTM prediction model is verified based on the analysis results.

[0063] When outputting the fault type according to the verification result, if the fault tracing verification of the power fault knowledge graph passes, the output of the power fault knowledge graph tracing is used as the result. If the verification fails, the fault type obtained by the knowledge graph tracing is judged. If the fault type is a characteristic easily confused fault, the output of the XGBOOST fault analysis model is used as the result. If the fault type is an instantaneous fault with little dependence on time, the output of the power fault knowledge graph tracing is used as the result.

[0064] The present invention also discloses a power fault classification prediction system based on the above-mentioned power fault classification prediction method based on LSTM and improved XGBOOST.

[0065] The following is a simulated fault classification prediction based on the direction of the present invention, using the power load data for 2022 and 2023 provided by a new energy power grid in a certain region, with a data sampling period of 5 minutes. The data from January to March 2022 is used as training data to predict the type of power grid data fault from January to March 2023. At the same time, it is compared with the prediction model of the traditional LSTM-XGBOOST. Table 1 shows the prediction accuracy of different models for different fault types. Short-term strong fluctuations indicate that the data fluctuates violently in a short period of time, and abnormal mutations indicate abnormal changes in the PI ratio in the data. The total data volume is 12147, of which the normal data volume is 6108, the model change data volume is 3977, the short-term strong fluctuation data volume is 1648, and the abnormal mutation data volume is 1414.

[0066]

[0067] Table 1 Prediction statistics of different models for different fault types

[0068] As shown in Table 1, the present invention has the highest accuracy for fault identification, reducing the error rate by 31.5% compared to the traditional LSTM-XGBOOST. It can be clearly concluded that the present invention can improve the accuracy of power grid fault diagnosis.

[0069] Model Traditional XGBOOST model XGBOOST fault analysis model of the present invention 1 I:4.62 I:9.02 2 P:3.03 I_lag1: 5.92 3 Hour: 1.01 P:5.89 4 Q:0.59 Hour_cos: 2.52 5 Minute: 0.27 P_lag1: 1.85 6 Day: 0.25 I_rolling_mean: 1.31 7 Year: 0 hour_sin: 1.01 8 Month: 0 Q_lag1: 0.88 9 Second: 0 Q_rolling_mean: 0.76 10 P_rolling_mean: 0.75

[0070] Table 2 Comparison of the importance of fault features of different models

[0071] P_lag1 represents the data of parameter P lagged by one hour, and P_rolling_mean represents the 24-hour rolling window statistics of parameter P. Similarly, I_lag1 and Q_lag1 represent the data of the corresponding parameters I and Q lagged by one hour, respectively. Q_rolling_mean and I_rolling_mean represent the 24-hour rolling window statistics of the corresponding parameters Q and I, respectively.

[0072] As can be seen from Table 2, the traditional XGBOOST model has no time-dependent features and only uses the original timestamps (Year, Month, Day, Hour, etc.) and parameters (I, P, Q), without considering temporal dependencies. The importance of features such as Hour and Minute is low (Hour = 1.01, Minute = 0.27) because it is difficult for the model to directly understand the periodicity of linear time coding. The XGBOOST fault analysis and prediction model of the present invention is an improvement on time series feature engineering. I_lag1 (1 hour lag) and P_lag1 are ranked 2nd and 5th with an importance of 5.92 and 1.85 respectively, indicating that historical values are crucial to current predictions. Hour_cos (hourly periodic coding) is 2.52 important, much higher than the original Hour feature (1.01). Periodic coding maps time to the unit circle, making it easier for the model to identify periodic patterns (such as day and night differences). The core parameter I is enhanced, and the importance of the original feature I is increased from 4.62 to 9.02. This is because the newly added hysteresis and rolling features enhance its timing expression capability, thereby improving the accuracy of the model for fault analysis.

Claims

1. The power fault classification prediction method based on LSTM and improved XGBOOST is characterized by: Including steps: S01. Train the LSTM prediction model and XGBOOST fault analysis model based on historical power system data. S02. Acquire measurement data of multiple source devices in the power system; S03. Sampling the measured data to obtain sampled data and performing spatiotemporal alignment, inputting the sampled data into the LSTM prediction model, generating LSTM prediction data and completing fault feature extraction of the sampled data; S04. Generate a residual sequence based on the difference between the LSTM predicted data and the corresponding sampled data, use the obtained residual sequence to construct a power fault knowledge graph, and use the power fault knowledge graph to trace the fault; S05. Input the fault characteristics into the XGBOOST fault analysis model to obtain a fault classification result, use the fault classification result to verify the power fault knowledge graph, and output the fault type according to the verification result.

2. The power fault classification prediction method based on LSTM and improved XGBOOST according to claim 1 is characterized in that: In step S01, the process of training the LSTM prediction model includes: constructing a power grid data matrix based on the historical data of the power system. When constructing the power grid data matrix, the power data matrix is first normalized, and the time data corresponding to the power data is mapped to a sine function. The time data represented by the sine function and the corresponding normalized power data matrix are used to construct a training array and input into the LSTM model for training.

3. The power fault classification prediction method based on LSTM and improved XGBOOST according to claim 1 is characterized in that: In step S03 , the process of performing spatiotemporal alignment on the sampled data includes: establishing a cubic spline interpolation function of the sampled data, and performing interpolation processing on the sampled data so that the synchronization error of the measured data of the multi-source devices is zero after the processing.

4. The power fault classification prediction method based on LSTM and improved XGBOOST according to claim 1 is characterized in that: In step S05, when the fault characteristics are input into the XGBOOST fault analysis model, the XGBOOST fault analysis model is first used to extract the importance of data related to the fault characteristics from the power system historical data, and then the input fault characteristics are feature-weighted according to the data importance, and weights are assigned to the data related to the fault characteristics.

5. The power fault classification prediction method based on LSTM and improved XGBOOST according to claim 4 is characterized in that: When the LSTM prediction model generates prediction data based on the sampled data, the sampled data is checked for missing data through the LSTM prediction model, and the prediction data is used to fill in the missing data in the sampled data. After the sampled data is filled, the fault characteristics of the sampled data are extracted.

6. The power fault classification prediction method based on LSTM and improved XGBOOST according to claim 4 is characterized in that: When training the XGBOOST fault analysis model based on the historical data of the power system, the historical data of the power system is sorted by time, the time parameters of the historical data of the power system are periodically encoded, and the historical data of the power system is cut according to the periodic coding. The XGBOOST fault analysis model is trained using the cut historical data of the power system arranged in time sequence.

7. The power fault classification prediction method based on LSTM and improved XGBOOST according to claim 6 is characterized in that: When periodically encoding the time data of the power system historical data, the hourly value of the time data per day is mapped to the corresponding unit circle with a period of 24, and the date corresponding to the time data is mapped to a period of 7 days. The time data of the power system historical data is encoded using a sine function, and the time data is corresponded to the power data of the power system historical data.

8. The power fault classification prediction method based on LSTM and improved XGBOOST according to claim 7 is characterized in that: When training the XGBOOST fault analysis model, the importance of data to power faults is identified through changes in hysteresis features and rolling window statistics.

9. The power fault classification prediction method based on LSTM and improved XGBOOST according to claim 1 is characterized in that: When outputting the fault type according to the verification result, if the fault tracing verification of the power fault knowledge graph passes, the output of the power fault knowledge graph tracing is used as the result. If the verification fails, the fault type obtained by the knowledge graph tracing is judged. If the fault type is a characteristic easily confused fault, the output of the XGBOOST fault analysis model is used as the result. If the fault type is an instantaneous fault with little dependence on time, the output of the power fault knowledge graph tracing is used as the result.

10. The power fault classification prediction system based on LSTM and improved XGBOOST is characterized by: Use the power fault classification prediction method based on LSTM and improved XGBOOST as described in any one of claims 1-9.