Characteristic importance-based power time sequence prediction model analysis method, device, equipment, medium and product

By analyzing global and local importance and combining the SHAP and LIME algorithms, the internal decision-making logic of the power time series prediction model is revealed, which solves the problem of low reliability of traditional models and improves the reliability of power grid dispatching and operation.

CN121786435APending Publication Date: 2026-04-03CHINA SOUTHERN POWER GRID ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-17
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Traditional deep learning models struggle to explain their internal decision-making logic in power time series forecasting, resulting in low reliability.

Method used

By acquiring the global and local importance of power grid operation data characteristics, and combining the SHAP and LIME algorithms, the behavior of power time series prediction models is analyzed to reveal their internal decision-making logic.

Benefits of technology

This improves the decision-making transparency and interpretability of power time-series forecasting models, and enhances the credibility of forecast results in power grid dispatching and operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121786435A_ABST
    Figure CN121786435A_ABST
Patent Text Reader

Abstract

The invention relates to an electric power time sequence prediction model analysis method and device based on feature importance, equipment, a medium and a product, and relates to the technical field of power grids. Comprising the following steps: in a process of performing electric power prediction on a target power grid by an electric power time sequence prediction model, acquiring multiple types of operation data characteristics and original characteristic values corresponding to the operation data characteristics at prediction moments; performing global importance analysis on each type of operation data features based on the original feature values corresponding to the operation data features at the prediction moments; superposing interference information on the basis of each original feature value to obtain an interference feature value corresponding to the operation data feature at each prediction moment; based on each interference characteristic value, carrying out local importance analysis on the operation data characteristics; and according to the global importance and each local importance, carrying out prediction behavior analysis on the power time sequence prediction model. By adopting the method, the credibility of the power time sequence prediction model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of power grid technology, and in particular to a method, apparatus, computer equipment, computer-readable storage medium, and computer program product for analyzing power time series prediction models based on feature importance. Background Technology

[0002] In power grid operation and dispatch, various operational indicators exhibit distinct time-series characteristics. To achieve safe and efficient power grid operation, power time-series forecasting has become an important technical means.

[0003] Traditional technologies often rely on deep learning models for power time-series forecasting, such as LSTM (Long Short-Term Memory), GRU (Gated Recurrent Unit), or Transformer models. These models can capture nonlinear relationships in complex time-series data, improving the accuracy of power forecasting. However, these models are black-box characteristics, and their internal decision-making logic is difficult to interpret, resulting in low reliability. Summary of the Invention

[0004] Therefore, it is necessary to provide a feature-based power time series prediction model analysis method, apparatus, computer equipment, computer-readable storage medium, and computer program product that can improve the reliability of power time series prediction models by addressing the above-mentioned technical problems.

[0005] Firstly, this application provides a power time-series forecasting model analysis method based on feature importance, including:

[0006] In the process of power prediction of the target power grid by the power time series prediction model, multiple types of operating data features in the target power grid as prediction samples are obtained, as well as the original feature values ​​of each operating data feature at each prediction time.

[0007] For each type of operational data feature, based on the original feature values ​​corresponding to the operational data feature at each prediction time, a global importance analysis is performed on the operational data feature to obtain the global importance of the operational data feature;

[0008] Interference information is superimposed on each original feature value to obtain the interference feature values ​​corresponding to the running data features at each prediction time.

[0009] Based on the various interference feature values, a local importance analysis is performed on the operational data features to obtain the local importance of each operational data feature at each prediction time.

[0010] Based on the global importance and the importance of each locality, the predictive behavior of the power time series prediction model is analyzed, and the behavioral analysis results of the power time series prediction model are obtained.

[0011] In one embodiment, for each type of operational data feature, based on the original feature values ​​corresponding to the operational data feature at each prediction time, a global importance analysis is performed on the operational data feature to obtain its global importance, including:

[0012] For each type of operational data feature, obtain at least one initial feature combination that does not contain operational data features at each prediction time.

[0013] For each initial feature combination, the running data features are added to the initial feature combination to obtain the target feature combination;

[0014] Based on the original feature values ​​corresponding to the initial feature combination and the target feature combination, the feature contribution of the running data features in the target feature combination is analyzed.

[0015] The global importance of operational data features is obtained by integrating the feature contribution of operational data features in multiple combinations of target features.

[0016] In one embodiment, based on the original feature values ​​corresponding to the initial feature combination and the target feature combination, the feature contribution of the running data features in the target feature combination is analyzed, including:

[0017] The original feature values ​​corresponding to the initial feature combination and the target feature combination are respectively input into the power time series prediction model to obtain the first prediction result corresponding to the initial feature combination and the second prediction result corresponding to the target feature combination.

[0018] Obtain the difference between the first prediction result and the second prediction result. Based on the difference, determine the feature contribution of the running data features in the target feature combination. The feature contribution is positively correlated with the absolute value of the difference.

[0019] In one embodiment, based on each interference feature value, a local importance analysis is performed on the running data features to obtain the local importance of each running data feature at each prediction time, including:

[0020] For each prediction time, the disturbance characteristic value corresponding to the prediction time is input into the power time series prediction model to obtain the power prediction result at the prediction time.

[0021] Based on power forecast results and disturbance characteristic values, linear regression analysis is performed on the operational data characteristics to obtain the local importance of the operational data characteristics at the forecast time.

[0022] In one embodiment, the power time-series prediction model analysis method based on feature importance further includes:

[0023] Obtain the prediction error and prediction resources of the power time series prediction model;

[0024] Based on prediction error and prediction resources, the performance of the power time series prediction model is evaluated, and the performance evaluation results of the power time series prediction model are obtained.

[0025] Based on global and local importance, the predictive behavior of the power time-series forecasting model is analyzed, yielding the following results:

[0026] Based on the global importance, local importance, and performance evaluation results, the predictive behavior of the power time series prediction model is analyzed, and the behavioral analysis results of the power time series prediction model are obtained.

[0027] In one embodiment, based on global importance and local importance, a predictive behavior analysis is performed on the power time series prediction model to obtain the behavior analysis results of the power time series prediction model, including:

[0028] Based on the global importance and the importance of each locality, a decision analysis is performed on the power time series forecasting model to obtain the decision path diagram of the power time series forecasting model when forecasting power for the target power grid;

[0029] Based on the decision path diagram, the predictive behavior of the power time series prediction model is analyzed, and the behavioral analysis results of the power time series prediction model are obtained.

[0030] Secondly, this application also provides a power time series prediction model analysis device based on feature importance, comprising:

[0031] The data acquisition module is used to acquire multiple types of operational data features in the target power grid as prediction samples, as well as the original feature values ​​of each operational data feature at each prediction time, during the process of power prediction by the power time series prediction model for the target power grid.

[0032] The global importance analysis module is used to perform global importance analysis on each type of running data feature based on the original feature values ​​corresponding to the running data feature at each prediction time, and to obtain the global importance of the running data feature.

[0033] The feature interference module is used to superimpose interference information on each original feature value to obtain the interference feature values ​​corresponding to the running data features at each prediction time.

[0034] The local importance analysis module is used to perform local importance analysis on the features of the running data based on each interference feature value, and to obtain the local importance of each feature of the running data at each prediction time.

[0035] The model behavior analysis module is used to perform predictive behavior analysis on the power time series prediction model based on global importance and local importance, and obtain the behavior analysis results of the power time series prediction model.

[0036] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described power time series prediction model analysis method based on feature importance.

[0037] Fourthly, this application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps in the above-described power time-series prediction model analysis method based on feature importance.

[0038] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps in the above-described power time-series prediction model analysis method based on feature importance.

[0039] The aforementioned power time-series forecasting model analysis method, apparatus, computer equipment, computer-readable storage medium, and computer program product based on feature importance, in the process of power time-series forecasting model predicting power in a target power grid, first acquires multiple types of operational data features in the target power grid as prediction samples, and the original feature values ​​corresponding to each operational data feature at each prediction time. For each type of operational data feature, based on the original feature values ​​corresponding to the operational data feature at each prediction time, a global importance analysis is performed on the operational data feature to obtain its global importance. In addition, by superimposing interference information on each original feature value, interference feature values ​​corresponding to each operational data feature at each prediction time are generated. Based on these interference feature values, a local importance analysis is further performed on the operational data feature to obtain its local importance at each prediction time. Finally, by combining the global importance and the local importance, a prediction behavior analysis of the power time-series forecasting model is performed to obtain the model's behavior analysis results. Thus, by integrating the global and local importance of various features, this scheme reveals the internal decision-making logic of the power time series prediction model, effectively improving the model's decision transparency and interpretability. This significantly enhances the credibility of its prediction results in power grid dispatch and operation, providing a reliable basis for subsequent intelligent power grid dispatch, safety early warning, and optimization decisions based on the prediction results. Attached Figure Description

[0040] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0041] Figure 1 This is a diagram illustrating the application environment of a power time-series prediction model analysis method based on feature importance in one embodiment.

[0042] Figure 2 This is a flowchart illustrating a power time-series prediction model analysis method based on feature importance in one embodiment;

[0043] Figure 3 This is a flowchart illustrating the global importance analysis process in one embodiment;

[0044] Figure 4 This is a flowchart illustrating the feature contribution analysis in one embodiment;

[0045] Figure 5 This is a flowchart illustrating the local importance analysis in one embodiment;

[0046] Figure 6 This is a schematic diagram of the model performance evaluation process in one embodiment;

[0047] Figure 7 This is a flowchart illustrating the model behavior analysis process in one embodiment;

[0048] Figure 8 This is a structural block diagram of a power time series prediction model analysis device based on feature importance in one embodiment;

[0049] Figure 9 This is an internal structural diagram of a computer device in one embodiment;

[0050] Figure 10 This is a diagram of the internal structure of a computer device in another embodiment. Detailed Implementation

[0051] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0052] It should be noted that the terms "first," "second," etc., used in this application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "comprising" and "having," and any variations thereof, used in this application, are intended to cover non-exclusive inclusion. The term "multiple" used in this application refers to two or more. The term "and / or" used in this application refers to one of the embodiments, or any combination of multiple embodiments.

[0053] The power time series prediction model analysis method based on feature importance provided in this application can be applied to, for example... Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104, or it can be located in the cloud or on another network server.

[0054] The terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, drones, low-altitude aircraft, IoT devices, etc. The server 104 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud computing services.

[0055] Specifically, during the power time-series forecasting model's power prediction of the target power grid, server 104 acquires multiple types of operational data features from the target power grid as prediction samples, as well as the original feature values ​​corresponding to each operational data feature at each prediction time. For each type of operational data feature, server 104 performs a global importance analysis based on the original feature values ​​corresponding to the operational data feature at each prediction time to obtain the global importance of the operational data feature. By superimposing interference information on each original feature value, server 104 can obtain the interference feature values ​​corresponding to the operational data feature at each prediction time. Based on each interference feature value, local importance analysis is performed on the operational data feature to obtain the local importance of the operational data feature at each prediction time. Finally, server 104 performs a prediction behavior analysis on the power time-series forecasting model based on the global importance and the local importance of each feature to obtain the behavior analysis results of the power time-series forecasting model.

[0056] It should be noted that the method provided in this application embodiment can be executed by the server 104 or the terminal 102 alone, or it can be executed interactively by the server 104 and the terminal 102.

[0057] In one exemplary embodiment, such as Figure 2As shown, a power time series prediction model analysis method based on feature importance is provided, and this method is applied to... Figure 1 Taking server 104 as an example, the following steps are included:

[0058] Step S202: During the process of power prediction of the target power grid by the power time series prediction model, multiple types of operating data features in the target power grid as prediction samples, and the original feature values ​​of each operating data feature at each prediction time are obtained.

[0059] In this context, the power time-series forecasting model refers to a pre-trained machine learning model used to predict the power parameters of a target power grid at a certain point in the future. Essentially, it is a black-box function that receives structured input and outputs prediction results. Power parameters can be, but are not limited to, at least one of the following: load, power, and renewable energy output. The target power grid refers to a specific power network object, which can be a regional power grid or a specific distribution network. The power forecasting process refers to the process of using the power time-series forecasting model to predict the power of the target power grid based on the input prediction samples and outputting the corresponding prediction results. This embodiment intervenes after this process to analyze the model's decision-making behavior during the forecasting process. A prediction sample is the smallest data unit processed by the model in a single prediction. In time-series forecasting, one prediction sample typically corresponds to one prediction time and consists of all operational data features within that time and the preceding time window. Operational data features are the data features corresponding to the operational parameters of the target power grid during the forecasting phase, such as at least one of the following: load features, voltage features, current features, power features, temperature features, humidity features, wind speed features, and renewable energy output features. Operational parameters can be, but are not limited to, at least one of the following: load, voltage, current, power, temperature, humidity, wind speed, and renewable energy output. The prediction time refers to the specific future time point at which the model needs to output a predicted value. Multiple consecutive prediction times constitute the prediction time series. The original feature value refers to the specific numerical value corresponding to each type of running data feature at a specific prediction time. It is the true input value used by the model when making the original prediction.

[0060] For example, the server can obtain multi-dimensional time-series operating parameters from the target power grid's multi-source data platform. For instance, it can obtain at least one type of power data, such as load, voltage, and current, from the real-time acquisition system of the power grid dispatch center, with a sampling frequency of once every 5 minutes. Information such as temperature, humidity, and wind speed can be obtained from meteorological monitoring stations, with a sampling frequency of once per hour. As for new energy output data, which includes the output power curves of renewable energy power plants such as wind power and photovoltaics, it is characterized by strong fluctuations and nonlinearity, and its data usually comes from the real-time monitoring platform on the power grid side. It should be noted that the above sampling frequency can be adjusted according to the actual application scenario, and this embodiment does not impose any limitations on it. Furthermore, the server can also determine the length of the time window covered by each prediction sample according to the model training requirements, such as 7 to 30 days, which can be determined based on the actual situation.

[0061] After completing the acquisition of multi-source operating parameters, the corresponding operating data features for each parameter can be obtained. Since the data sources and sampling frequencies of different features may differ, the server can further align them using timestamps. Specifically, high-frequency data, such as power data, can be sampled directly at a specified frequency, while low-frequency data, such as meteorological data, can be padded at specific time points using linear interpolation or forward padding, thus ensuring that all features are aligned in the time dimension. Subsequently, to eliminate dimensional differences and improve model convergence efficiency and analytical stability, various features can be normalized. For example, a minimum-maximum scaling method can be used to map feature values ​​to the [0,1] interval. Features such as load, power, and renewable energy output can use a unified scaling standard based on historical extreme values, while meteorological data features are normalized independently according to their physical properties. Finally, the server outputs a structured data feature set. In this dataset, row indices correspond to prediction times, column indices correspond to operating data features, and each cell stores the original feature value of that feature at that time. This step provides a standardized and consistent data foundation for subsequent importance analysis.

[0062] Step S204: For each type of operational data feature, based on the original feature values ​​corresponding to the operational data feature at each prediction time, perform a global importance analysis on the operational data feature to obtain the global importance of the operational data feature.

[0063] Global importance analysis is a process that quantifies the contribution of each type of operational data feature to the overall prediction result of the model based on the original feature values ​​at all prediction times. Its goal is to reveal the influence of features on long-term prediction. In this embodiment, global importance may include the feature importance distribution that characterizes the dynamic change of the contribution of each type of operational data feature over time, and the global average importance extracted from this distribution for feature ranking.

[0064] For example, the server uses a pre-trained power time-series forecasting model as a basis and employs a game theory-based feature importance analysis algorithm to perform global importance analysis. In a preferred example, this algorithm is the SHAP (SHapley Additive Explanation) algorithm. Based on the Shapley value theory in cooperative game theory, the SHAP algorithm treats the forecasting model as a cooperative game system and the various input operational data features as participants in the game. It performs analysis by calculating the marginal contribution value, i.e., the SHAP value, feature-by-feature and time-by-time. The marginal contribution value refers to the change in the model's prediction result caused by adding the current target feature to a certain feature combination from all possible feature combinations. This change is obtained by weighted averaging of all possible combinations, reflecting the net impact of the feature at that time. For each type of operational data feature, the server repeats the above marginal contribution calculation at all prediction times, ultimately generating an intermediate analysis result, namely a feature importance distribution matrix. The rows of this matrix correspond to each type of operational data feature, and the columns correspond to each prediction time. Each element in the matrix represents the marginal contribution value of a certain feature at a specific time, ranging from [0,1]. This matrix comprehensively records the dynamic changes in the contribution of each feature over time. Furthermore, by taking the absolute value of the contribution value sequence of each row in the matrix (i.e., each feature at all prediction times) and calculating the average, the global average importance score of that feature over the entire prediction period can be obtained. This scalar indicator can be used for comprehensive comparison and ranking among features.

[0065] In some embodiments, to enable feature importance analysis to adapt to the temporal evolution of the power grid state, thereby providing more timely feature contribution tracking, the server can also set a fixed-length time window. The length depends on the data sampling frequency and the model prediction granularity, typically set to 24 time points, corresponding to a 24-hour cycle. This window slides forward with the current prediction time. For each operational data feature within the current window, the server extracts the contribution value sequence of that feature at each time point within the window from its corresponding feature importance distribution matrix. Subsequently, the feature importance change rate within the current window is calculated, i.e., the absolute value of the difference between the contribution values ​​at adjacent time points in the sequence is calculated sequentially, and then the average of these absolute values ​​is calculated. The larger this change rate value, the more drastic the recent fluctuations in the importance of the feature, which may mean that its impact on the power grid is unstable or in a changing state. Next, the server assigns a weight adjustment coefficient to each feature based on the change rate value according to a preset mapping rule. The mapping rule typically divides the change rate into several intervals and assigns a coefficient to each interval, with the principle that the higher the change rate, the larger the coefficient corresponding to the interval. For example, features with the highest rate of change (top 20%) have a weight adjustment coefficient of 1.5; features with the highest rate of change (middle 60%) have a coefficient of 1; and features with the lowest rate of change (bottom 20%) have a coefficient of 0.8. This amplifies the contribution of highly volatile features within the current time window, making them more readily noticeable. Finally, the server multiplies the original contribution value of each feature at each time point within the window by its corresponding weight adjustment coefficient to obtain a weighted feature importance value. This value is then updated in the feature importance distribution matrix, or a new dynamic weighted matrix is ​​generated. This process continues as the time window slides, allowing the evaluation of feature importance to dynamically respond to the latest state of the system. This enables multi-scale, adaptive tracking of feature contributions, improving the understanding of model decision-making behavior.

[0066] In other embodiments, the weight adjustment coefficient can also be set in a reverse time order, that is, the closer a feature is to the current time, the larger its weight coefficient is, and vice versa, the farther a feature is from the current time, the smaller its weight coefficient is.

[0067] Step S206: Add interference information to each original feature value to obtain the interference feature values ​​corresponding to the running data features at each prediction time.

[0068] In this context, perturbation information refers to small, random variations injected into the original feature values ​​to explore the local behavior of the model. Its purpose is to systematically construct a new set of similar input points within the feature space corresponding to the original feature values, rather than introducing errors or noise. Perturbation feature values ​​are the new values ​​obtained by superimposing perturbation information onto the original feature values. For a single original feature value, multiple different perturbation feature values ​​can be generated by superimposing different random perturbation information; the set of these values ​​constitutes a new set of sample points within the local neighborhood.

[0069] For example, the server can employ the LIME (Local Interpretable Model-agnostic Explanations) algorithm as a specific implementation of local importance analysis. Specifically, for a specific prediction time to be interpreted, the server first obtains the original feature values ​​of each class of running data features constituting the prediction samples at that time, forming an original feature vector. Then, to explore the model's behavior in this local region, the server invokes the LIME algorithm. The LIME algorithm generates multiple (e.g., 500) neighborhood perturbation samples in the feature space surrounding the original feature vector. Specifically, the LIME algorithm operates independently on each running data feature in the original feature vector, superimposing an independently generated perturbation information onto its original feature value. This perturbation information is typically obtained by uniformly or normally distributed random sampling within a preset relative proportion range (e.g., ±5% or ±10% of the original value). This process is performed on all features simultaneously, thereby generating new data points, i.e., neighborhood perturbation samples. Repeating this independent random perturbation process multiple times yields a set of perturbation samples surrounding the original feature vector; the values ​​of these samples are the perturbation feature values.

[0070] Step S208: Based on each interference feature value, perform local importance analysis on the running data features to obtain the local importance of each running data feature at each prediction time.

[0071] Local importance analysis is a feature importance analysis process for a single, specific prediction time. Its goal is to quantify the specific contribution of each type of operational data feature to the prediction result at that local prediction time. Local importance refers to a numerical index calculated for each type of operational data feature at a specific prediction time through local importance analysis. This index is usually expressed in the form of a regression coefficient, whose sign (positive / negative) represents the direction of the feature's influence on the prediction result, i.e., promoting or inhibiting it, and whose absolute value represents the degree or intensity of the influence.

[0072] For example, the LIME algorithm combines all the disturbance feature values ​​generated at the current prediction time to form a disturbance sample set, with each disturbance feature value being a disturbance sample. This disturbance sample set is then input into the original power time series prediction model to obtain the corresponding disturbance prediction result. Next, the Euclidean distance between the feature values ​​of each disturbance sample and the original sample is calculated, and a weight is assigned to each disturbance sample based on this distance; the closer the distance, the greater the weight, with the weight ranging from [0,1]. Based on the feature values ​​of the disturbance samples (i.e., the disturbance feature values), the disturbance prediction result, and the aforementioned sample weights, the LIME algorithm uses a weighted linear regression method to fit a linear regression model. After fitting, the regression coefficient corresponding to each running data feature is extracted from the linear regression model. This regression coefficient is defined as the local importance of the feature at the current prediction time. A positive coefficient indicates that the feature has a positive pulling effect on the prediction result at the current time; a negative coefficient indicates an inhibitory effect. The larger the absolute value of the coefficient, the more critical the feature is in the current decision. Finally, the set of local importance values ​​for all running data features at the prediction time is output. This process is performed independently for each prediction time, thus providing a personalized, traceable interpretation of the decision for each prediction point in the entire time series.

[0073] Step S210: Based on the global importance and the importance of each locality, perform prediction behavior analysis on the power time series prediction model to obtain the behavior analysis results of the power time series prediction model.

[0074] Predictive behavior analysis refers to the process of systematically interpreting and summarizing the decision-making patterns, characteristic dependencies, and time-series evolution patterns of power time-series forecasting models by integrating global and local importance. The results of the behavioral analysis are summary information obtained through predictive behavior analysis that comprehensively reflects the model's predictive logic, and are typically presented in the form of an interpretable analysis report.

[0075] For example, after completing the feature importance analysis, the server can perform decision analysis on the power time series prediction model based on the global importance and the importance of each local part, so as to obtain information such as the decision mode, feature dependency relationship and time series evolution law of the power time series prediction model, that is, the behavior analysis results.

[0076] In this embodiment, during the power time-series forecasting model's power prediction of the target power grid, firstly, multiple types of operational data features serving as prediction samples in the target power grid, along with the original feature values ​​corresponding to each operational data feature at each prediction time, are acquired. For each type of operational data feature, based on the original feature values ​​corresponding to the operational data feature at each prediction time, a global importance analysis is performed to obtain the global importance of the operational data feature. Furthermore, by superimposing interference information on each original feature value, interference feature values ​​corresponding to each operational data feature at each prediction time are generated. Based on these interference feature values, a local importance analysis is further performed on the operational data feature to obtain its local importance at each prediction time. Finally, by combining the global importance and the local importance, a prediction behavior analysis of the power time-series forecasting model is performed to obtain the model's behavior analysis results. Thus, by integrating the global and local importance of various features, this embodiment reveals the internal decision-making logic of the power time series prediction model, effectively improving the model's decision transparency and interpretability, thereby significantly enhancing the credibility of its prediction results in power grid dispatch and operation, and providing a reliable basis for subsequent intelligent power grid dispatch, safety early warning and optimization decision-making based on the prediction results.

[0077] In one exemplary embodiment, such as Figure 3 As shown, for each type of operational data feature, based on the original feature values ​​corresponding to the operational data feature at each prediction time, a global importance analysis is performed on the operational data feature to obtain its global importance, including:

[0078] Step S302: For each type of running data feature, obtain at least one initial feature combination that does not contain running data features at each prediction time.

[0079] The initial feature combination refers to any subset of features other than the currently running data features. For example, if the set of running data features at a certain prediction time is {A, B, C, D}, and the feature currently being analyzed for importance is A, then the initial feature combination can be an empty set {}, {B}, {C}, {D}, {B, C}, {B, D}, {C, D}, {B, C, D}, etc., with each combination representing a specific feature participation scenario.

[0080] For example, during global importance analysis, for the feature to be analyzed (denoted as feature i), the server first removes feature i itself from the complete feature set corresponding to the current prediction time, thus obtaining a remaining feature set. Then, it systematically enumerates all possible subsets of the remaining feature set, including the empty set, and these enumerated subsets are the initial feature combinations.

[0081] Step S304: For each initial feature combination, add the running data features to the initial feature combination to obtain the target feature combination.

[0082] The target feature combination refers to a new feature combination formed by adding the features of the currently running data to the initial feature combination. For example, if the initial feature combination is {B, C} and the current feature is A, then the target feature combination is {A, B, C}.

[0083] For example, for each generated initial feature combination S, the server adds the currently analyzed feature i to form a new feature subset S∪{i}, i.e., the target feature combination. By systematically constructing all possible target feature combinations, the foundation is laid for the next step of quantifying the incremental contribution of feature i in each specific feature participation scenario.

[0084] Step S306: Based on the original feature values ​​corresponding to the initial feature combination and the target feature combination, analyze the feature contribution of the running data features in the target feature combination.

[0085] Among them, the feature contribution is the change in the model prediction result caused by the addition of running data features before and after the initial feature combination.

[0086] For example, for each pair of initial feature combinations S and target feature combinations S∪{i}, the server needs to calculate two different model predictions: the prediction value f(S) of the initial feature combination and the prediction value f(S∪{i}) of the target feature combination S∪{i}. The difference between the two is the contribution of feature i to the initial feature combination S, which reflects the pure incremental impact of adding feature i given the initial feature combination S.

[0087] Step S308: Integrate the feature contribution of the running data features in the combination of multiple target features to obtain the global importance of the running data features.

[0088] For example, after obtaining the feature contribution, the server can perform a weighted summation based on the Shapley value theory. The weight is determined by the number of features contained in the initial feature combination S. Specifically, the weight is inversely proportional to the number of features; the fewer the features in the initial feature combination S, the greater its corresponding weight, and vice versa. Multiplying each feature contribution by its corresponding weight and summing them yields the SHAP value of feature i at the current prediction time. This value quantifies the degree and direction of feature i's influence on the prediction result at the current time. Finally, the server takes the absolute value of the SHAP values ​​of feature i at all prediction times and calculates their arithmetic mean to obtain the global importance of feature i. This score reflects the average influence strength of feature i throughout the entire prediction period. Because the SHAP algorithm evaluates contributions by traversing all possible feature combinations during the calculation process, it not only captures the independent role of each feature but also naturally integrates the interactive influence between different features. Therefore, this embodiment can simultaneously quantify the individual contribution and interactive contribution of features, ensuring a comprehensive reflection of feature effects.

[0089] In this embodiment, by simulating the addition of all possible initial feature combinations and calculating the resulting changes, not only is the independent effect of the feature evaluated, but the interaction between the feature and other features is naturally integrated by traversing all combination scenarios. This ensures that the final global importance index reliably reveals the true impact of the feature on the overall model decision-making, providing a solid and reliable basis for subsequent model behavior analysis.

[0090] In one exemplary embodiment, such as Figure 4 As shown, based on the original feature values ​​corresponding to the initial feature combination and the target feature combination, the feature contribution of the running data features in the target feature combination is analyzed, including:

[0091] Step S402: Input the original feature values ​​corresponding to the initial feature combination and the target feature combination into the power time series prediction model respectively to obtain the first prediction result corresponding to the initial feature combination and the second prediction result corresponding to the target feature combination.

[0092] In this embodiment, each original feature value refers to the complete model input vector formed after a specific filling rule. The specific rule is as follows: for features within the feature combination, their true original feature values ​​at the current prediction time are used; for features not within the feature combination, preset background values ​​are used for filling, typically the statistical average or random sample value of the feature in the training set. The first prediction result refers to the output value obtained by inputting the original feature values ​​corresponding to the initial feature combination into the power time series prediction model. The second prediction result refers to the output value obtained by inputting the original feature values ​​corresponding to the target feature combination into the power time series prediction model.

[0093] For example, the server constructs two complete feature vectors conforming to the model input dimensions for each determined initial feature combination S and target feature combination S∪{i}. Specifically, for features present in the initial feature combination S, they are assigned the true feature index at the current prediction time; for all features outside the combination, they are uniformly filled with a baseline value sampled from the background dataset, such as the feature's average value across all samples. These two filled complete vectors are sequentially fed into the trained power time series prediction model to obtain the corresponding first prediction result f(S) and second prediction result f(S∪{i}).

[0094] Step S404: Obtain the difference between the first prediction result and the second prediction result. Based on the difference, determine the feature contribution of the running data features in the target feature combination. The feature contribution is positively correlated with the absolute value of the difference.

[0095] Here, the difference refers to the numerical difference between the second prediction result and the first prediction result, i.e., Δ = f(S∪{i}) - f(S). The feature contribution is directly proportional to the absolute value of the difference; that is, the larger |Δ| is, the stronger the effect of feature i under the initial feature combination S, and the smaller |Δ| is, the weaker the effect. The sign of the difference indicates the direction of contribution: a positive value indicates a positive pull, and a negative value indicates a negative inhibition.

[0096] For example, after obtaining f(S) and f(S∪{i}), the server directly calculates the difference Δ between them. This difference Δ is the feature contribution determined in this embodiment.

[0097] In this embodiment, by simulating the addition of all possible combinations of initial features and calculating the resulting changes, the contribution of the feature can be quantified fairly and comprehensively.

[0098] In one exemplary embodiment, such as Figure 5 As shown, based on each interference feature value, a local importance analysis is performed on the operational data features to obtain the local importance of each operational data feature at each prediction time, including:

[0099] Step S502: For each prediction time, input the disturbance characteristic value corresponding to the prediction time into the power time series prediction model to obtain the power prediction result at the prediction time.

[0100] Among them, the power prediction result refers to the output value obtained after inputting the disturbance sample composed of disturbance feature values ​​into the original prediction model, that is, the disturbance prediction result. These results will be used as the dependent variable in the subsequent linear regression analysis.

[0101] For example, the server inputs the disturbance feature values ​​of each type of operational data at the current prediction time into the trained power time series prediction model. The model performs forward computation on each disturbance sample and outputs the corresponding predicted value, such as load or power prediction value.

[0102] Step S504: Based on the power forecast results and disturbance characteristic values, perform linear regression analysis on the operating data characteristics to obtain the local importance of the operating data characteristics at the forecast time.

[0103] Linear regression analysis, also known as weighted linear regression, is the core step of the LIME algorithm. It aims to approximate the behavior of complex models in local neighborhoods using a simple linear model.

[0104] For example, the LIME algorithm uses all disturbance features as independent variables, their corresponding power prediction results as dependent variables, and sample weights as a measure of the importance of observations. It then uses weighted least squares to fit a linear regression model. After fitting, the regression coefficients corresponding to each type of operational data feature are extracted from the linear model. These regression coefficients are defined as the local importance of the feature at the current prediction time.

[0105] In this embodiment, by analyzing the relationship between the input features of the interference data and the model output, and fitting a local linear model, the black-box decision-making of the complex model at a specific moment can be transformed into an intuitive and quantifiable feature contribution, providing a traceable explanation for the prediction results at each moment, thereby significantly enhancing the transparency of the model and the credibility of the decision.

[0106] In one exemplary embodiment, such as Figure 6 As shown, the power time series prediction model analysis method based on feature importance also includes:

[0107] Step S602: Obtain the prediction error and prediction resources of the power time series prediction model.

[0108] In this context, prediction error refers to the deviation between the output value of the power time-series prediction model and the actual observed value, used to quantify the model's prediction accuracy. This embodiment typically uses at least one statistical indicator such as mean squared error, root mean square error, or mean absolute error. Mean squared error is the average of the squares of all prediction errors, and is more sensitive to larger errors. Root mean square error is the square root of the mean squared error. Mean absolute error is the average of the absolute values ​​of prediction errors, reflecting the average degree of deviation. Prediction resources refer to the resources consumed by the model during the prediction task, including but not limited to at least one of computational resources and time costs. Time costs can be divided into model training time costs and model prediction time costs. Model training time costs are the total time taken from loading prediction samples to the complete completion of model training, measured in seconds; model prediction time costs are the average or total time required for the model to complete the prediction, measured in milliseconds. Computational resources refer to the overall resources consumed during model prediction, specifically obtained by calculating the average of multiple computational resource indicators, including but not limited to at least two such indicators such as CPU utilization, GPU utilization, and memory usage.

[0109] In some embodiments, to quantify the impact of the model behavior analysis mechanism itself on prediction, the power time series prediction model can be independently evaluated before and after the introduction of the mechanism. Specifically, the performance of the original model is evaluated first, and then evaluated again after integrating the model behavior analysis mechanism. The performance indicators of the two evaluations (such as prediction error, computation time, resource utilization, etc.) will be organized into a structured comparison data table. Through comparative analysis, the actual impact of introducing the model behavior analysis mechanism on the overall prediction efficiency, response speed, and resource consumption of the model will be clearly revealed.

[0110] For example, during the power time-series forecasting model's power prediction of the target power grid, the server obtains the model's output prediction results and compares them with the corresponding actual observations. By calling a statistical computing library, the server programmatically calculates three core prediction error indicators: mean square error, root mean square error, and mean absolute error. Simultaneously, during model training and operation, the server uses system monitoring tools to accurately record the time consumed from model training to the time consumed from starting prediction to outputting results, and samples and records the computational resources consumed during model operation.

[0111] Step S604: Based on the prediction error and prediction resources, perform a performance evaluation on the power time series prediction model to obtain the performance evaluation results of the power time series prediction model.

[0112] Performance evaluation refers to the process of making a quantitative judgment on the model's performance based on its prediction error and prediction resources. The performance evaluation result is the specific output of the above performance evaluation process, typically a structured comparative report or data summary that clearly demonstrates the model's performance across different evaluation dimensions, with a particular focus on comparing the changes in accuracy and efficiency before and after the introduction of model behavior analysis mechanisms.

[0113] In some embodiments, the predictive behavior analysis of the power time series prediction model is performed based on global importance, local importance, and performance evaluation results to obtain the behavioral analysis results of the power time series prediction model.

[0114] For example, this embodiment aims to integrate global, local, and performance dimensions to complete a systematic and comprehensive diagnosis of the model's predictive behavior. The server first performs a correlation analysis between the global importance ranking and the local importance at each time point, identifying globally critical and locally stable core features, globally minor features that play a dominant role at specific times, and the correspondence between local contribution patterns and system operating conditions (such as daily peaks and valleys, holidays). Simultaneously, combined with performance evaluation results, it analyzes the distribution of local importance of the model during periods of high error, locates the key feature combinations leading to prediction bias, and assesses whether the model's computational resource consumption is within acceptable limits.

[0115] In this embodiment, the decision-making basis of the model is revealed not only from the global and local levels, but also the model performance is quantified, thus providing comprehensive decision support that takes into account reliability, accuracy and efficiency for the engineering deployment, optimization iteration and reliable application of the model in power dispatch.

[0116] In one exemplary embodiment, such as Figure 7 As shown, based on global importance and local importance, the predictive behavior analysis of the power time series prediction model is performed, and the behavioral analysis results of the power time series prediction model are obtained, including:

[0117] Step S702: Based on the global importance and the importance of each local area, perform decision analysis on the power time series forecasting model to obtain the decision path diagram of the power time series forecasting model when forecasting power for the target power grid.

[0118] Step S704: Based on the decision path diagram, perform prediction behavior analysis on the power time series prediction model to obtain the behavior analysis results of the power time series prediction model.

[0119] Among them, the decision path diagram is a comprehensive visualization or structured graph used to systematically show the decision logic of the power time series prediction model in the time series dimension. It is usually composed of a combination of various related charts, such as feature importance ranking table, local prediction interpretation diagram of key prediction time, feature contribution time series curve, feature time series heat map, prediction error distribution map, and performance comparison statistics chart, etc.

[0120] In some embodiments, the feature contribution time-series curve represents the changing trend of the importance of each type of operational data feature throughout the entire prediction period. The horizontal axis represents continuous prediction time, and the vertical axis represents the global average importance score of the feature. A curve is plotted for each input feature, and its trend indicates the changing trend of the feature's contribution within the prediction period. The curve height reflects its relative influence on the prediction result. Typically, the top ten features with the highest average contribution are selected for display to ensure image clarity and highlight key features. Furthermore, a feature contribution heatmap can be constructed, supporting segmented and layered display by day, week, month, etc. Additionally, a reference line function can be optionally added to the graph, such as marking load peaks and temperature critical points, to assist users in associating feature contribution changes with the actual operating status of the power system.

[0121] The feature interaction heatmap reflects the increase or decrease in the impact of any two feature combinations on the prediction results during the model prediction process. Specifically, several features with the highest contribution ranking are selected, and the contribution trends between each pair of features at the same prediction time are calculated. The Pearson correlation coefficient or cosine similarity between the two trends is then calculated as an interaction strength index, generating a symmetric matrix as the input data for the heatmap. The horizontal and vertical axes of the matrix represent the feature names, and the color intensity of each cell corresponds to the correlation strength value; the larger the value, the darker the color, visually demonstrating the degree of synergy and structural relationship between features.

[0122] The horizontal axis of the prediction error distribution plot represents the prediction time, and the vertical axis represents the prediction error value. The plot can also highlight the locations of samples with high errors. For a more in-depth assessment of error concentration, an error histogram or probability density function curve can be overlaid on the plot to analyze whether the errors exhibit a skewed distribution or clustering of abnormal peaks.

[0123] The feature importance ranking table displays the global average importance score of all features in the prediction model in tabular form, sorted from highest to lowest. Each row shows the feature name, average importance score (average contribution value), range of contribution value variation, and weight fluctuations over different time periods. The local prediction interpretation plot is used to show the specific values ​​of each feature at typical prediction times and their directional impact on the prediction results. Each feature is represented by a horizontal bar, with the horizontal length representing the degree of influence, and colors distinguishing between positive and negative directions.

[0124] Performance comparison charts, in the form of bar charts or line charts, compare prediction error, training time, prediction time, and computational resource consumption before and after the introduction of model behavior analysis mechanisms. Each indicator is accompanied by absolute value, percentage change, and key conclusions. For example, the introduction of model behavior analysis mechanisms increases the average absolute error by 1.2%, but increases the prediction latency by 15 milliseconds, thus providing a direct basis for decision-making on model deployment and engineering trade-offs.

[0125] For example, based on global importance, local importance, and performance evaluation results, a decision analysis is performed on the power time-series forecasting model. This yields a feature importance ranking table, local prediction interpretation diagrams for key prediction moments, feature contribution time-series curves, feature time-series heatmaps, prediction error distribution maps, and performance comparison statistics. These curves and graphs collectively constitute the decision path diagram of the power time-series forecasting model. Based on this decision path diagram, information such as the decision-making pattern, feature dependencies, and time-series evolution patterns of the power time-series forecasting model can be interpreted—that is, the behavioral analysis results.

[0126] In this embodiment, by integrating the global laws abstracted by the model with specific local decisions into an intuitive and coherent decision path diagram, the black-box prediction process of the model is transformed into a traceable behavioral analysis report, which directly supports the trust judgment of the prediction results and the decision-making process.

[0127] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.

[0128] Based on the same inventive concept, this application also provides a feature-importance-based power time-series prediction model analysis device for implementing the feature-importance-based power time-series prediction model analysis method described above. The solution provided by this device is similar to the implementation scheme described in the above method. Therefore, the specific limitations of one or more feature-importance-based power time-series prediction model analysis device embodiments provided below can be found in the limitations of the feature-importance-based power time-series prediction model analysis method described above, and will not be repeated here.

[0129] In one exemplary embodiment, such as Figure 8 As shown, a power time series prediction model analysis device based on feature importance is provided, comprising:

[0130] The data acquisition module 802 is used to acquire multiple types of operational data features in the target power grid as prediction samples, as well as the original feature values ​​of each operational data feature at each prediction time, during the process of power prediction of the target power grid by the power time series prediction model.

[0131] The global importance analysis module 804 is used to perform global importance analysis on each type of running data feature based on the original feature values ​​corresponding to the running data feature at each prediction time, and to obtain the global importance of the running data feature.

[0132] The feature interference module 806 is used to superimpose interference information on each original feature value to obtain the interference feature values ​​corresponding to the running data features at each prediction time.

[0133] The local importance analysis module 808 is used to perform local importance analysis on the running data features based on each interference feature value, and to obtain the local importance of each running data feature at each prediction time.

[0134] The model behavior analysis module 810 is used to perform predictive behavior analysis on the power time series prediction model based on global importance and local importance, and obtain the behavior analysis results of the power time series prediction model.

[0135] In one embodiment, the global importance analysis module 804 is further configured to:

[0136] For each type of operational data feature, obtain at least one initial feature combination that does not contain operational data features at each prediction time.

[0137] For each initial feature combination, the running data features are added to the initial feature combination to obtain the target feature combination;

[0138] Based on the original feature values ​​corresponding to the initial feature combination and the target feature combination, the feature contribution of the running data features in the target feature combination is analyzed.

[0139] The global importance of operational data features is obtained by integrating the feature contribution of operational data features in multiple combinations of target features.

[0140] In one embodiment, the global importance analysis module 804 is further configured to:

[0141] The original feature values ​​corresponding to the initial feature combination and the target feature combination are respectively input into the power time series prediction model to obtain the first prediction result corresponding to the initial feature combination and the second prediction result corresponding to the target feature combination.

[0142] Obtain the difference between the first prediction result and the second prediction result. Based on the difference, determine the feature contribution of the running data features in the target feature combination. The feature contribution is positively correlated with the absolute value of the difference.

[0143] In one embodiment, the local importance analysis module 808 is further configured to:

[0144] For each prediction time, the disturbance characteristic value corresponding to the prediction time is input into the power time series prediction model to obtain the power prediction result at the prediction time.

[0145] Based on power forecast results and disturbance characteristic values, linear regression analysis is performed on the operational data characteristics to obtain the local importance of the operational data characteristics at the forecast time.

[0146] In one embodiment, the device is further configured to:

[0147] Obtain the prediction error and prediction resources of the power time series prediction model;

[0148] Based on prediction error and prediction resources, the performance of the power time series prediction model is evaluated, and the performance evaluation results of the power time series prediction model are obtained.

[0149] In one embodiment, the model behavior analysis module 810 is further configured to:

[0150] Based on the global importance, local importance, and performance evaluation results, the predictive behavior of the power time series prediction model is analyzed, and the behavioral analysis results of the power time series prediction model are obtained.

[0151] In one embodiment, the model behavior analysis module 810 is further configured to:

[0152] Based on the global importance and the importance of each locality, a decision analysis is performed on the power time series forecasting model to obtain the decision path diagram of the power time series forecasting model when forecasting power for the target power grid;

[0153] Based on the decision path diagram, the predictive behavior of the power time series prediction model is analyzed, and the behavioral analysis results of the power time series prediction model are obtained.

[0154] The modules in the aforementioned power time-series prediction model analysis device based on feature importance can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.

[0155] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 9 As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media. The database stores power time-series forecasting model analysis data based on feature importance. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When the computer program is executed by the processor, it implements a power time-series forecasting model analysis method based on feature importance.

[0156] In one exemplary embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 10As shown, the computer device includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interface. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides the environment for the operation of the operating system and computer programs in the non-volatile storage media. The input / output interface is used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a power time-series prediction model analysis method based on feature importance.

[0157] Those skilled in the art will understand that Figure 9 or Figure 10 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0158] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.

[0159] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.

[0160] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0161] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0162] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0163] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0164] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A power time series prediction model analysis method based on feature importance, characterized in that, The method includes: In the process of power prediction of the target power grid by the power time series prediction model, multiple types of operational data features in the target power grid as prediction samples are obtained, as well as the original feature values ​​of each of the operational data features at each prediction time. For each type of operational data feature, based on the original feature values ​​corresponding to the operational data feature at each prediction time, a global importance analysis is performed on the operational data feature to obtain the global importance of the operational data feature; Interference information is superimposed on each of the original feature values ​​to obtain the interference feature values ​​corresponding to the running data features at each of the prediction times. Based on each of the aforementioned interference feature values, a local importance analysis is performed on the operational data features to obtain the local importance of each operational data feature at each of the aforementioned prediction times; Based on the global importance and the local importance of each power time series prediction model, a prediction behavior analysis is performed on the power time series prediction model to obtain the behavior analysis results of the power time series prediction model.

2. The method according to claim 1, characterized in that, For each type of operational data feature, based on the original feature values ​​corresponding to the operational data feature at each prediction time, a global importance analysis is performed on the operational data feature to obtain its global importance, including: For each type of operational data feature, obtain at least one initial feature combination that does not contain the operational data feature at each prediction time. For each of the initial feature combinations, the running data features are added to the initial feature combination to obtain the target feature combination; Based on the original feature values ​​corresponding to the initial feature combination and the target feature combination, the feature contribution of the running data features in the target feature combination is analyzed. The global importance of the operational data features is obtained by integrating the feature contribution values ​​of the operational data features in multiple combinations of target features.

3. The method according to claim 2, characterized in that, The step of analyzing the feature contribution of the running data features in the target feature combination based on the original feature values ​​corresponding to the initial feature combination and the target feature combination includes: The original feature values ​​corresponding to the initial feature combination and the target feature combination are respectively input into the power time series prediction model to obtain the first prediction result corresponding to the initial feature combination and the second prediction result corresponding to the target feature combination. Obtain the difference between the first prediction result and the second prediction result, and determine the feature contribution of the running data feature in the target feature combination based on the difference; the feature contribution is positively correlated with the absolute value of the difference.

4. The method according to claim 1, characterized in that, The step of performing local importance analysis on the running data features based on each of the aforementioned interference feature values ​​to obtain the local importance of each running data feature at each of the aforementioned prediction times includes: For each predicted time, the disturbance feature value corresponding to the predicted time is input into the power time series prediction model to obtain the power prediction result at the predicted time. Based on the power forecast results and the disturbance characteristic values, a linear regression analysis is performed on the operating data characteristics to obtain the local importance of the operating data characteristics at the forecast time.

5. The method according to claim 1, characterized in that, The method further includes: Obtain the prediction error and prediction resources of the power time series prediction model; Based on the prediction error and the prediction resources, the performance of the power time series prediction model is evaluated to obtain the performance evaluation result of the power time series prediction model. The step of performing predictive behavior analysis on the power time series prediction model based on the global importance and the local importance, and obtaining the behavior analysis results of the power time series prediction model, includes: Based on the global importance, the local importance, and the performance evaluation results, the predictive behavior analysis of the power time series prediction model is performed to obtain the behavior analysis results of the power time series prediction model.

6. The method according to claim 1, characterized in that, The step of performing predictive behavior analysis on the power time series prediction model based on the global importance and each of the local importance, and obtaining the behavior analysis results of the power time series prediction model, includes: Based on the global importance and the local importance of each of the above, a decision analysis is performed on the power time series prediction model to obtain the decision path diagram of the power time series prediction model when making power predictions for the target power grid. Based on the decision path diagram, the predictive behavior analysis of the power time series prediction model is performed to obtain the behavior analysis results of the power time series prediction model.

7. A power time series prediction model analysis device based on feature importance, characterized in that, The device includes: The data acquisition module is used to acquire multiple types of operational data features in the target power grid as prediction samples, as well as the original feature values ​​of each operational data feature at each prediction time, during the process of power prediction by the power time series prediction model for the target power grid. The global importance analysis module is used to perform global importance analysis on each type of operational data feature based on the original feature values ​​corresponding to the operational data feature at each prediction time, and to obtain the global importance of the operational data feature. The feature interference module is used to superimpose interference information on each of the original feature values ​​to obtain the interference feature values ​​corresponding to the running data features at each of the prediction times. The local importance analysis module is used to perform local importance analysis on the running data features based on each of the interference feature values, and to obtain the local importance of each of the running data features at each of the prediction times. The model behavior analysis module is used to perform predictive behavior analysis on the power time series prediction model based on the global importance and the local importance of each power time series prediction model, and obtain the behavior analysis results of the power time series prediction model.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.