Electric power material demand prediction analysis method and system based on big data analysis
Through big data analysis and dynamic optimization of the power material demand forecasting method, the problems of pseudo-correlation interference and static modeling are solved, and accurate prediction and dynamic response of power material demand are achieved to meet the power grid operation and maintenance needs.
Patent Information
- Application Number
- CN202510774152.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-09-19
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing power material demand forecasting methods are unable to effectively capture true causal relationships due to pseudo-correlation interference and static modeling methods, resulting in prediction results deviating from actual demand and making it difficult to meet the real-time and accuracy requirements of material allocation for dynamic power grid operation and maintenance.
A method based on big data analysis is adopted, through feature screening, coupling analysis, causal verification and dynamic optimization, combined with time domain-frequency domain joint analysis and Granger causality test, to dynamically adjust the lag order, establish a power material demand forecasting model, and use the Actor-Critic structure for real-time update.
Accurately screening real causal variables improves the ability to capture sudden demands and medium- and long-term trends, ensures the stability and accuracy of the forecasting model under complex working conditions, and supports precise allocation of materials.
Smart Images

Figure CN120672058A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power material demand forecasting, and in particular to a power material demand forecasting and analysis method based on big data analysis and a system thereof. Background Art
[0002] Existing power material demand forecasting methods mainly rely on traditional statistical models or single machine learning algorithms. Their key limitation is that the causal variable identification mechanism has significant defects. Traditional methods usually screen variables based on correlation analysis, such as measuring the correlation between features and target variables through the Pearson correlation coefficient or mutual information. However, the complex correlation of multi-source heterogeneous data in the power system (such as equipment failure records, load curves, and environmental parameters) often leads to pseudo-correlation interference. For example, the accidental synchronization of load fluctuations and equipment failures may be mistakenly judged as a causal relationship. The introduction of such pseudo-correlated variables will weaken the model's ability to capture true causal relationships, and ultimately cause the material demand forecast results to deviate from actual demand.
[0003] In addition, electricity data has significant non-stationary and cyclical characteristics, such as differences in peak and valley periods, seasonal load fluctuations, and sudden demands caused by extreme weather (such as a surge in emergency repair materials). However, existing forecasting models often use fixed lag orders or single difference strategies to process time series data, such as fixed-order ARIMA models or sliding window methods based on empirical settings. This type of static modeling method cannot adapt to the time-varying characteristics of adaptive data, resulting in a weak ability to capture medium- and long-term trends in the model, and a slow response in sudden demand scenarios, causing the forecast results to lag behind actual changes in material demand, making it difficult to meet the real-time and accuracy requirements of dynamic power grid operation and maintenance for material allocation. Summary of the Invention
[0004] The purpose of this section is to summarize some aspects of the embodiments of the present invention and briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this section and the abstract and title of this application to avoid obscuring the purpose of this section, the abstract and the title of the invention, and such simplifications or omissions should not be used to limit the scope of the present invention.
[0005] To solve the above technical problems, the present invention provides the following technical solution: a method for predicting and analyzing power material demand based on big data analysis, comprising the following steps:
[0006] Acquiring target power-related data and preprocessing the target power-related data, wherein the preprocessing includes feature screening;
[0007] performing a coupling analysis on the failure rate of power grid equipment and load fluctuation based on the pre-processed target power-related data to obtain a set of candidate intervention variables;
[0008] Performing a screening operation on the candidate intervention variable set to obtain a final intervention variable set, wherein the screening operation includes screening through a causal validation process;
[0009] Anomaly detection is performed on the final set of intervention variables, and the anomalies are classified according to the degree of the detected anomalies. The filling strategy is dynamically adjusted through Bayesian optimization to minimize the data correction error, and finally the cleaned training data set is output. The anomaly detection is performed using a variational autoencoder combined with Mahalanobis distance.
[0010] As a preferred solution of the power material demand forecasting and analysis method based on big data analysis described in the present invention, it also includes:
[0011] Based on the training data set, an initial power material demand forecasting model is established and dynamically optimized. During the dynamic optimization process, an Actor-Critic structure is introduced, and the experience replay mechanism is used to update the strategy.
[0012] Input real-time power grid data into the optimized forecasting model to generate forecasts of power material demand in future time periods.
[0013] As a preferred solution of the power material demand forecasting and analysis method based on big data analysis described in the present invention, wherein: based on the preprocessed target power-related data, the failure rate of power grid equipment and load fluctuation are coupled analyzed to obtain a set of candidate intervention variables, including quantifying the coupling relationship between equipment failure rate and load fluctuation through a time domain-frequency domain joint analysis method.
[0014] As a preferred solution of the power material demand forecasting and analysis method based on big data analysis described in the present invention, the time domain analysis includes the following steps:
[0015] The load data is divided into peak, flat and valley periods according to the power dispatch cycle, and different weighting factors are assigned to each period;
[0016] Calculate the absolute difference between the failure rate series and the load series at each time point, introduce the time period weight factor to weight the difference, calculate the grey correlation degree based on the weighted difference, and retain the variables with grey correlation degree greater than the first set value as the time domain candidate set;
[0017] The frequency domain analysis further performs frequency domain coupling verification based on the time domain candidate set output by the time domain analysis, including the following steps:
[0018] Perform Morlet wavelet transform on the load sequence and failure rate sequence in the time domain candidate set;
[0019] Extract wavelet coefficients of characteristic frequency bands;
[0020] Calculating the wavelet coherence within the characteristic frequency band, and retaining the variables whose coherence is greater than the second set value and whose phase difference is less than the third set value;
[0021] The time domain grey correlation value and the frequency domain coherence coefficient are weighted and integrated to generate the final coupling strength score. The candidate intervention variables are sorted in descending order according to the coupling strength score, and the first b variables with the highest scores are selected as the final candidate intervention variable set.
[0022] As a preferred solution of the power material demand forecasting and analysis method based on big data analysis described in the present invention, the causal verification process includes the following steps:
[0023] Do-Calculus was used to calculate the intervention effect value θ of power material demand on the intervention variable to quantify the causal impact;
[0024] Propensity score matching was implemented to control the matching bias between the experimental group and the control group;
[0025] Adjust the lag order L for the Granger causality test.
[0026] As a preferred solution of the power material demand forecasting and analysis method based on big data analysis described in the present invention, wherein: the implementation of propensity score matching to control the matching deviation between the experimental group and the control group includes the following steps:
[0027] The three power characteristics of equipment operation age, ambient temperature, and load rate in the power-related data are mandatory and included as matching variables;
[0028] The radius matching method was used, and the matching radius was set to N times the standard deviation of the propensity score;
[0029] After the matching is completed, check whether the experimental group and the control group have reached a balance. If so, the matched samples are used as input data to enter the subsequent causal effect analysis stage.
[0030] As a preferred embodiment of the power material demand forecasting and analysis method based on big data analysis described in the present invention, after the matching is completed, a two-sample t-test is used to evaluate the significance of the difference between the groups, and the standardized mean difference method is combined to jointly determine the balance between the experimental group and the control group. The joint determination rules and results are described as follows:
[0031] When the standardized mean difference is less than or equal to the first threshold, and the significance value is greater than or equal to the second threshold, the match is considered successful.
[0032] When the standardized mean difference is greater than the first threshold and the significance value is less than the second threshold, the match is determined to be unsuccessful and remedial measures are triggered. If the conditions are still not met, the intervention variable is eliminated.
[0033] When the standardized mean difference is less than or equal to the first threshold and the significance value is less than the second threshold, further judgment process is required:
[0034] If the standardized mean difference is less than or equal to the third threshold, the difference is considered significant but negligible, and the match is considered successful. If the third threshold is less than or equal to the standardized mean difference and less than or equal to the first threshold, the match is considered unsuccessful, triggering remedial measures. If the condition is still not met, the intervention variable is eliminated.
[0035] When the standardized mean difference is greater than the first threshold and the significance value is greater than the second threshold, it means that the difference is large but not statistically significant. The match is judged as unsuccessful and remedial measures are triggered. If the condition still does not meet the requirements, the intervention variable is eliminated.
[0036] When the initial matching fails to meet the balance requirement, the following remedial measures are triggered in sequence:
[0037] Add the number of equipment maintenance times as an additional matching variable and execute the joint decision rule to determine whether the matching success criteria are met;
[0038] Expand the matching radius;
[0039] If the conditions are still not met, the intervention variable will be automatically eliminated.
[0040] As a preferred solution of the power material demand forecasting and analysis method based on big data analysis of the present invention, wherein: the adjusting the lag order L of the Granger causality test includes the following steps:
[0041] Determine the initial lag order according to the periodic characteristics of power load data;
[0042] Perform ADF test on the original sequence and obtain the significance probability value during the test;
[0043] The significance probability value is used to determine whether a time series is a stationary series. The judgment rules are as follows:
[0044] If the significance probability value is less than the fourth threshold, the sequence is determined to be stationary and the difference order d=0 is recorded;
[0045] If the significance probability value is greater than or equal to the fourth threshold, the sequence is judged to be non-stationary and the step-by-step difference process is entered. The ADF test is repeated after each difference until the sequence becomes stationary.
[0046] Record the stabilized sequence and its optimal difference order d as the benchmark parameter for subsequent order adjustment;
[0047] Use the Bayesian Information Criterion to preliminarily select the upper limit L of the lag order max ;
[0048] Dynamically adjust the lag order L according to the stationarity results;
[0049] Based on the dynamically adjusted lag order L, a Granger causality test is performed to calculate the F statistic and significance probability value. If the significance probability value is less than the fourth threshold, it is considered that a Granger causality relationship exists; otherwise, it is considered that no Granger causality relationship exists. The variables judged to have a Granger causal relationship are used to generate an optimized feature set, that is, the final intervention variable set.
[0050] A rolling window is used to test Granger causality to observe whether it changes over time. If the Granger causality is unstable, the data preprocessing process is re-evaluated and the lag order selection is optimized.
[0051] As a preferred solution of the power material demand forecasting and analysis method based on big data analysis of the present invention, the process of the step-by-step difference process and the strategy of dynamically adjusting the lag order L are as follows:
[0052] The step-by-step difference process is as follows:
[0053] Perform step-by-step difference processing on the non-stationary series, and re-perform the ADF test after each difference until the series is stationary;
[0054] Record the stabilized sequence and its optimal difference order d as the benchmark parameter for subsequent order adjustment;
[0055] The strategy for dynamically adjusting the lag order L is as follows:
[0056] If the sequence is a stationary sequence, that is, the ADF test rejects the null hypothesis, then the optimal lag order L selected by the Bayesian Information Criterion is directly used. opt Conduct Granger causality test;
[0057] If the sequence is non-stationary but the order after stationarization is d=1, a larger lag order is used to compensate for the effect of the difference operation on information loss;
[0058] If the series is non-stationary and d ≥ 2 after stationarity, increase or decrease the lag order to avoid information loss due to over-differentiation.
[0059] The present invention also provides a power material demand forecasting and analysis system based on big data analysis, which is applied to the above-mentioned power material demand forecasting and analysis method based on big data analysis, including:
[0060] Data acquisition and preprocessing module, used to collect power-related data and perform preprocessing and feature screening on the data;
[0061] Candidate intervention variable identification module, used to identify candidate intervention variables with causal effects and provide key input features for subsequent predictions;
[0062] Causal verification and feature optimization module, used to dynamically adjust the lag order and optimize the feature set;
[0063] The anomaly detection and data filling module uses a variational autoencoder combined with Mahalanobis distance to detect anomalies, classify abnormal data, and dynamically adjust the filling strategy through Bayesian optimization to minimize data correction errors.
[0064] The power material demand forecasting model module establishes a power material demand forecasting model based on the optimized feature set, adopts the Actor-Critic structure for dynamic optimization, and uses the experience replay mechanism to continuously update the forecasting strategy;
[0065] The real-time data input and forecast output module receives real-time power grid data, inputs it into the optimized forecast model, generates forecast results of power material demand in future time periods, and provides visual display and decision support;
[0066] The matching and balance test module uses a two-sample t-test to evaluate the matching balance between the experimental group and the control group after propensity score matching, and determines whether the matching is successful or unsuccessful based on the set thresholds of the standardized mean difference and significance value;
[0067] The dynamic adjustment and optimization module triggers remedial measures when the matching fails to reach equilibrium, optimizes the lag order selection strategy, and dynamically evaluates the stability of Granger causality using the rolling window method.
[0068] Beneficial effects of the present invention:
[0069] 1. This invention utilizes dynamic coupling analysis and a triple causal verification mechanism. Do-Calculus quantifies intervention effects, propensity score matching eliminates confounding bias, and Granger tests confirm temporal causality. This allows for precise screening of true causal variables and avoids spurious correlations. Furthermore, by combining the ADF stationarity test with the Bayesian Information Criterion to dynamically optimize the lag order, the model adapts to the time-varying characteristics of power data, significantly improving its ability to capture sudden demand and medium- and long-term trends, providing reliable support for the precise allocation of supplies.
[0070] 2. The present invention adopts a dynamic weight mechanism of joint analysis of time domain and frequency domain to distinguish the impact of peak and valley periods, and ensures the stability of the model in data imbalance scenarios through remedial measures (variable addition / matching radius adjustment). The adaptive lag order compensation technology of non-stationary sequences significantly reduces the information loss caused by differential operations, so that the prediction model still maintains superior performance under complex working conditions and improves prediction accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. Those skilled in the art can also derive other drawings based on these drawings without inventive effort. Among them:
[0072] Figure 1 This is a flowchart of the power material demand forecasting and analysis method based on big data analysis of the present invention.
[0073] Figure 2 This is a flowchart of the causal verification process of the power material demand forecasting and analysis method based on big data analysis of the present invention. DETAILED DESCRIPTION
[0074] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0075] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0076] Secondly, the term "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in various places throughout this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive of other embodiments.
[0077] Furthermore, the present invention is described in detail with reference to schematic diagrams. For ease of illustration, when describing the embodiments of the present invention, cross-sectional views illustrating device structures may be partially enlarged and not to scale. Furthermore, the schematic diagrams are merely illustrative and should not limit the scope of protection of the present invention. Furthermore, in actual production, the three-dimensional dimensions of length, width, and depth should be included.
[0078] Example 1
[0079] Reference Figure 1-2 , as an embodiment of the present invention, provides a method for predicting and analyzing power material demand based on big data analysis, comprising the following steps:
[0080] S1: Acquire target power-related data and preprocess the target power-related data, where the preprocessing includes feature screening.
[0081] Specifically, the target power-related data includes but is not limited to grid operation data, load data, historical power material consumption data, weather information, and other factors affecting power system operation. After acquiring the data, it is cleaned and normalized to a uniform scale, laying a solid foundation for subsequent analysis. The Maximum Information Coefficient (MIC) method is also used to assess feature correlation and eliminate low-correlation features to reduce data redundancy.
[0082] It should be noted that the Maximum Information Coefficient (MIC) is a powerful statistical tool that works by exploring all possible two-dimensional spatial partitioning methods between variables to find the grid partitioning that maximizes mutual information, thereby quantifying the complex dependencies between variables. Compared to the traditional correlation coefficient, the MIC is not limited to linear relationships and can capture various correlation patterns such as nonlinearity and periodicity. In power data analysis, the MIC algorithm is used to calculate the MIC value between each feature and the target variable. This is used as a quantitative indicator of feature correlation. Feature selection is performed based on the quantitative indicator, and features with low correlation with the prediction target are eliminated.
[0083] In an optional embodiment, the feature screening module is mainly carried out for load data, that is, from the acquired power-related target data set, the feature variables closely related to the power load are extracted. The screened load features include but are not limited to: load fluctuation rate, load rate, temperature change in the power grid operation area, load proportion in typical time periods (such as peak / valley), historical extreme load records and other key indicators. By calculating the maximum information coefficient (MIC) between the above features and the target prediction variable (i.e., the demand for power materials), the core factors that have a greater impact on the fluctuation of material demand can be effectively identified, thereby providing high-quality input variables for subsequent prediction model training.
[0084] S2: Based on the pre-processed target power-related data, a coupling analysis is performed on the failure rate of power grid equipment and the load fluctuation to obtain a set of candidate intervention variables.
[0085] Specifically, based on the preprocessed target power-related data, a coupling analysis of the power grid equipment failure rate and load fluctuation is performed to obtain a set of candidate intervention variables, including quantifying the coupling relationship between equipment failure rate and load fluctuation through a joint analysis method in the time domain and frequency domain.
[0086] Further, through time domain analysis, the following steps are included:
[0087] S21-1: Divide the pre-processed load data into peak, flat and valley periods according to the power dispatch cycle, and assign different weighting factors to each period. For example, the weighting factor for peak period is 1.2, for flat period is 1, and for valley period is 0.8.
[0088] S21-2: Calculate the absolute difference between the failure rate sequence and the load sequence at each time point, introduce a time period weight factor to weight the difference, calculate the grey correlation degree based on the weighted absolute difference, and use the variables with a grey correlation degree greater than the first set value as the time domain candidate set.
[0089] It should be noted that by solving the absolute difference between the failure rate series and the load series at each time point, the degree of deviation between the failure rate series and the load series at the same time point can be measured. When the absolute difference is small, it means that the values of the failure rate series and the load series at each time point are very close. In other words, when the load changes, the failure rate changes almost synchronously, and there is a strong coupling between the two. When the absolute difference is large, it means that the values of the failure rate series and the load series at each time point differ greatly. In other words, the load increases sharply but the failure rate remains basically unchanged. The coupling between the two is weak, indicating a pseudo-correlation, which needs to be eliminated. By eliminating pseudo-correlations, the accuracy and credibility of the final intervention variable screening can be improved, thereby enhancing the prediction model's sensitivity to changes in power material demand and its ability to resist interference.
[0090] Furthermore, weighting the absolute differences is used to "amplify" the sensitivity of key periods to the coupled relationship between equipment failures and load fluctuations, thereby improving the accuracy of variable identification and, in turn, increasing the contribution of key period variables to the forecasting model. For example, during peak periods, when loads are high and grid operating pressures are high, material failures are more frequent and have a more severe impact. Meanwhile, during off-peak periods, when loads are relatively low, even if failures occur, they have little impact on the overall forecast.
[0091] It should also be noted that the gray correlation degree typically ranges from 0 to 1. This paper sets a first value of 0.6, which can screen for characteristic variables with moderate or higher correlations with grid failure rates or load fluctuations. This value is determined after extensive training and verification based on a large amount of historical data, and can avoid the incorporation of noise variables due to setting it too low, or the omission of true variables due to setting it too high.
[0092] The calculation formula of grey relational degree is as follows:
[0093]
[0094] in:
[0095] i is a variable;
[0096] γ i The weighted grey correlation between the i-th variable (such as a load node) and the equipment failure rate is used to measure the strength of the coupling relationship between them. The value range is usually between [0,1], and the closer to 1, the stronger the coupling;
[0097] ρ resolution coefficient, with a value of 0.5;
[0098] ω(k) is the weighting factor of the time period corresponding to time k, which is used to express the relative importance of the time period. For example, the weighting factor of the peak time period is 1.2, the weighting factor of the normal time period is 1, and the weighting factor of the valley time period is 0.8.
[0099] Δ i (k) represents the absolute difference of the i-th candidate variable at time k;
[0100] minΔ i (k) represents Δ at all times k i (k) minimum value;
[0101] maxΔ i (k) represents Δ at all times k i The maximum value of (k);
[0102] n represents the total number of time periods, i.e. the sample length. For example, if sampling is done hourly in one day, then n = 24.
[0103] For example, to evaluate the correlation between load variables and the failure rate series of a 220kV transformer, load data from three typical periods within a day were selected for analysis. These periods are peak, normal, and off-peak.
[0104] The weighting factors for each period are set as follows:
[0105] Peak period weighting factor ω=1.2
[0106] The weighting factor ω in normal period is 1
[0107] Valley period weighting factor ω=0.8
[0108] The load variables and the corresponding equipment failure rate target sequences in the three periods are extracted from the transformer SCADA system, and the original difference sequences are calculated as follows:
[0109] Original difference sequence Δ=[0.2,0.5,0.3]
[0110] minΔ i (k)=0.2,maxΔ i (k) = 0.5
[0111] Set the resolution coefficient ρ to 0.5
[0112] Calculate the numerator term:
[0113] For example, during the peak period, k = 1 (Δ1 = 0.2):
[0114]
[0115] For normal periods, k = 2 (Δ2 = 0.5):
[0116]
[0117] For the valley period, k = 3 (Δ2 = 0.3):
[0118]
[0119] Compute the weighted sum:
[0120] Assume the number of data points n = 3
[0121]
[0122] γ i =0.818>0.6, included in the time domain candidate set.
[0123] In summary, load data is divided into peak, flat, and valley periods according to the power dispatch cycle, and each period is assigned a different weighting factor (e.g., 1.2, 1, and 0.8). In the calculation of the grey correlation degree, the aforementioned time period weighting factors are introduced to perform a weighted average of the grey correlation values at each time point, thereby obtaining a comprehensive weighted grey correlation index. This can more accurately measure the dynamic coupling strength between power grid equipment failure rates and load fluctuations.
[0124] When performing frequency domain analysis, frequency domain coupling verification is further performed based on the time domain candidate set output by the time domain analysis, including the following steps:
[0125] S22-1: Perform Morlet wavelet transform on the load sequence and fault rate sequence in the time-domain candidate set, with a basis function parameter of 6. It should be noted that the engineering rationale for using a parameter of 6 is that, according to testing, when the parameter = 6, the time-frequency resolution product reaches an optimal value (Δt·Δf = 0.08), which is better than the Gabor wavelet (0.12).
[0126] S22-2: Extract the wavelet coefficients of the characteristic frequency band and use them to calculate the wavelet coherence between the power load signal and the equipment failure rate signal. Wavelet coherence measures the similarity and coupling strength of two signals in the frequency domain. It should be noted that according to the "DL / T 1630-2016 Guidelines for Vibration Testing of Power Equipment," the main vibration frequency band of the transformer core is 0.3-0.8 Hz, and the vibration frequency band of the circuit breaker operating mechanism is 0.1-0.5 Hz. Therefore, 0.1-1 Hz was selected as the characteristic frequency band.
[0127] S22-3: Calculate wavelet coherence within the characteristic frequency band and retain variables with coherence greater than the second set value and phase difference less than the third set value. By analyzing wavelet coherence, the frequency-domain coupling strength between the power load and the equipment failure rate within the physical equipment vibration frequency band can be assessed. A high coherence indicates that changes in load have a significant impact on the equipment failure rate; a low coherence indicates that changes in load have a small impact on the equipment failure rate.
[0128] It should be noted that the second setting value is set to 0.7, based on the typical empirical threshold of wavelet coherence indicators in actual signal analysis, which can effectively distinguish sporadic weak correlations from systematic strong coupling behavior. The third setting value is set to π / 4, based on the commonly used engineering criteria for judging signal coherence synchronization. It can effectively exclude variables that are coupled but not responding in real time, ensuring that the selected variables have high response consistency under actual operating conditions.
[0129] The calculation formula of cross-wavelet coherence is as follows:
[0130]
[0131] Among them, WTC X,Y (f) represents the coherence of variables X and Y at frequency f (range 0-1); W X (f) and W Y (f) Wavelet transform coefficients of variables X and Y respectively; W Y The complex conjugate of (f); S{·} is a smoothing operator;
[0132] The phase difference calculation formula is as follows:
[0133]
[0134] Among them, arg(·) represents the argument of the complex number, that is, the phase; It is the phase difference between signals X and Y at frequency f, which is very important for understanding the synchronization, lag or lead between the two signals.
[0135] In summary, for the variables that satisfy γ>0.6, the variables with coherence>0.7 and phase difference<π / 4 in the 0.1-1 Hz frequency band are retained for further analysis of the coupling relationship between equipment failure rate and load fluctuation.
[0136] For example, it is assumed that this application selects a specific vibration frequency band of the transformer (such as 100-200Hz) for analysis. Through wavelet transform, this application extracts the wavelet coefficients within this frequency band and calculates the wavelet coherence between the power load signal and the transformer failure rate signal. The results show that within this frequency band, the wavelet coherence is high, indicating that changes in power load have a significant impact on the failure rate of the transformer. This may be because changes in load lead to changes in mechanical stress or thermal stress inside the transformer, which in turn affects the operating state and failure rate of the transformer.
[0137] Through this example, this application can see that wavelet coefficient extraction and wavelet coherence analysis within the characteristic frequency band can provide a deeper understanding of the relationship between power load and equipment failure rate, providing strong support for equipment status monitoring and fault prediction.
[0138] S22-4: A weighted synthesis of the time-domain grey correlation value and the frequency-domain coherence coefficient (weight ratio 6:4, time-domain: frequency-domain) is performed. The benefit of this synthesis is that the grey correlation value reflects the degree of synchronization between the variable and the target variable in the time series, making it suitable for exploring "long-term trends." Cross-wavelet coherence can identify resonance or co-fluctuation relationships between variables at specific frequencies, making it more suitable for reflecting "periodic causal patterns." The weighted fusion of the two achieves complementary time-frequency information, improving the comprehensiveness and robustness of intervention variable identification. After this weighted synthesis, a final coupling strength score is generated. The candidate intervention variables are ranked according to the coupling strength score, and the top b variables with the highest scores are selected as the final candidate intervention variable set. This means that all coupling strength scores are calculated and sorted in descending order. The top b variables are selected as the final candidate intervention variable set, providing a high-quality variable foundation for subsequent causal verification.
[0139] S3: Screening the candidate intervention variable set to obtain the final intervention variable set. The screening operation includes screening through a causal validation process.
[0140] Specifically, the causal verification process includes:
[0141] S31: Do-Calculus is used to calculate the intervention effect value θ of power material demand on the intervention variable to quantify the causal impact. The formula is as follows:
[0142]
[0143] Among them, Y is the demand for power materials; X is the intervention variable, such as policy changes, grid compliance fluctuations, etc. represents the expected value of Y after the intervention on X; It represents the expected value of Y in the absence of intervention. This formula can be used to quantify the impact of intervention variables on the demand for electricity materials.
[0144] When used specifically, take the intervention variable X to be verified (such as a sudden load change event), extract the material procurement records for 180 days before and after the event as observation data; calculate the daily average value of material demand after the event occurs. As the intervention group, the average demand for the same equipment in the same historical period As a control group; When the threshold δ is exceeded, X is deemed to have a significant causal influence, where δ = 15%. For example:
[0145] Take the load restriction requirements during the peak summer period in 2023 in a certain province as an example:
[0146] Intervention group (policy implementation period): average daily transformer purchases tower;
[0147] Control group (same period in previous years): tower;
[0148] The effect value θ = 1.1 units, and the increase of 52.4% > δ, confirming that the policy is a key intervention variable.
[0149] S32: When implementing propensity score matching, controlling the matching bias between the experimental group and the control group includes the following steps:
[0150] S32-1: The three power characteristics of equipment operation age, ambient temperature, and load rate in the power-related data are mandatory and included as matching variables.
[0151] It's important to note that when performing propensity score matching, it's crucial to ensure that the experimental group (samples receiving the intervention) and the control group (samples not receiving the intervention) maintain similarity on certain key variables to reduce the impact of confounding factors. This step requires that, regardless of the selection of other variables, the equipment's operating age, ambient temperature, and load factor must be included in the matching process, as they significantly impact power material demand forecasts.
[0152] S32-2: Use the radius matching method, setting the matching radius to N times the standard deviation of the propensity score. N is preferably set to 0.1. Assuming the standard deviation of the propensity score for all samples is 0.15, then the matching radius = 0.1 × 0.15 = 0.015. The value of N can be adjusted between 0.05 and 0.2 times the data size. For example, if the sample size is greater than 100,000, a smaller value such as 0.05 can be used; if the sample size is less than 10,000, a larger value such as 0.15 can be used. This balances matching accuracy and sample size, avoiding overmatching that results in insufficient samples.
[0153] Among them, the radius matching method is to find a control group individual whose propensity score is within a certain radius for each individual in the experimental group during the matching process to ensure the closeness of the propensity scores of the two groups.
[0154] S32-3: After matching is complete, check whether the experimental and control groups are truly balanced (i.e., whether their differences in key characteristics are sufficiently small) to ensure the quality of the matching. If balance is achieved, the matched samples are used as input data for the subsequent causal effect analysis phase.
[0155] In an optional embodiment of the present invention, after matching is completed, a two-sample t-test is used to evaluate the significance of the difference between the groups, and the standardized mean difference method is combined to jointly determine the balance between the experimental group and the control group.
[0156] The significance value is obtained by performing a two-sample t-test on the sample means of the experimental group and the control group on each key feature (such as equipment operation years, ambient temperature, load rate, etc.). The calculation formula is as follows:
[0157]
[0158] Where t is the significance value, are the sample means of the experimental group and the control group respectively; S1 2 , S2 2 are the sample variances of the experimental group and the control group respectively; n1 and n2 are the sample sizes of the experimental group and the control group respectively.
[0159] The significance value calculated based on this statistic is used to determine whether there is a statistically significant difference between the experimental group and the control group in this feature.
[0160] The standardized mean difference is calculated as follows:
[0161]
[0162] in, are the sample means of the experimental group and the control group respectively; S1 2 , S2 2 are the sample variances of the experimental group and the control group, respectively.
[0163] The calculation of the standardized mean difference (SMD) shows that it is unaffected by sample size. It is a unitless measure of standardized differences in variables and is suitable for comparing different characteristics. Based on the SMD, balance judgments are made according to pre-defined judgment rules. Based on the judgment results, decisions are made regarding whether to trigger remedial measures or remove intervening variables.
[0164] The joint determination rules are as follows:
[0165] When the standardized mean difference is less than or equal to the first threshold, and the significance value is greater than or equal to the second threshold, the match is considered successful.
[0166] When the standardized mean difference is greater than the first threshold and the significance value is less than the second threshold, the match is determined to be unsuccessful and remedial measures are triggered. If the conditions are still not met, the intervention variable is eliminated.
[0167] When the standardized mean difference is less than or equal to the first threshold and the significance value is less than the second threshold, further judgment process is required:
[0168] If the standardized mean difference is less than or equal to the third threshold, the difference is considered significant but negligible, and the match is considered successful. If the third threshold is less than or equal to the standardized mean difference and less than or equal to the first threshold, the match is considered unsuccessful, triggering remedial measures. If the condition is still not met, the intervention variable is eliminated.
[0169] When the standardized mean difference is greater than the first threshold and the significance value is greater than the second threshold, it means that the difference is large but not statistically significant. The match is judged as unsuccessful and remedial measures are triggered. If the condition still does not meet the requirements, the intervention variable is eliminated.
[0170] The first threshold is set to 0.1, which is based on the Cohen standard in statistics. When the standardized mean difference is less than 0.1, it is widely considered to be a very small effect size, which means that the difference between the two groups is extremely small and acceptable.
[0171] The second threshold is set to 0.01. The setting basis is: the traditional standard for statistical significance is generally 0.05 or 0.01. The present invention adopts a more stringent 0.01, which can improve the credibility of causal inference.
[0172] The third threshold is set to 0.05. The reason for this setting is that in high-precision fields such as medicine and biostatistics, a standardized mean difference of <0.05 is considered to be "extremely small and negligible", especially when dealing with unbalanced sample matching.
[0173] It can be understood that when the standardized mean difference is less than 0.1, it means that the difference in the variable between the two groups is small, and the matching can be considered successful; when the significance value is greater than 0.01, it means that the difference between the two groups on the variable is not significant and is acceptable.
[0174] Preferably, when the initial matching fails to meet the balance requirement, the following remedial measures are triggered in sequence:
[0175] Add the number of equipment maintenance times as an additional matching variable and execute the joint decision rule to determine whether the matching success criteria are met;
[0176] Expand the matching radius, for example, expand the matching radius to 0.15 times the standard deviation;
[0177] If the conditions are still not met, the intervention variable will be automatically eliminated.
[0178] It should be noted that the present invention can effectively reduce the interference of potential confounding factors when processing observational data using the propensity score matching method, thereby improving the reliability of the estimation of the causal effect of the intervention variable. In addition, through propensity score matching, the similarity of the key characteristics of the experimental group and the control group can be ensured, so that the real impact of the intervention variable on the demand for power materials can be evaluated. This method is particularly suitable for scenarios where the intervention factors in the power system are complex and changeable.
[0179] It should be noted that the present invention preferably adopts the Propensity Score Matching method (PSM) to effectively reduce the interference of potential confounding factors when processing observational data, thereby improving the reliability of the estimation of the causal effect of the intervention variable. By constructing the similarity between the experimental group and the control group in key characteristics (such as equipment operation years, ambient temperature, load rate, etc.), this method can achieve more comparable inter-group comparisons, and then scientifically evaluate the real impact of the intervention variables on the demand for power materials, and can assist in accurately identifying the key variables that affect the changes in the demand for power materials, providing favorable support for subsequent power material demand forecasts, helping to optimize inventory management, reduce operating costs and improve the reliability of power material supply.
[0180] S33: Adjusting the lag order L of the Granger causality test to adapt the identification of causal relationships to the time-varying characteristics of power load data, specifically including the following steps:
[0181] S33-1: Determine the initial lag order based on the periodic characteristics of the power load data.
[0182] It should be noted that the method for determining the initial lag order has been widely used in the fields of power system load forecasting, causal analysis, etc. It is a mature and effective technical means and will not be described in detail here.
[0183] S33-2: Perform an ADF test on the original load sequence. If the original sequence is non-stationary, perform step-by-step differentiation until the sequence becomes stationary (i.e., the unit root test rejects the null hypothesis), and record the stabilized sequence and its optimal difference order d.
[0184] It should be noted that the ADF test is a unit root test method based on the time series regression model. A significance probability value can be obtained during the test process. The significance probability value is calculated by statistical software according to a specific ADF distribution. Based on the significance probability value, it is judged whether a time series is a stationary series. If the test result rejects the null hypothesis (that is, the sequence has a unit root), the series can be considered stationary; otherwise, it is a non-stationary series.
[0185] The rules for judging whether a time series is a stationary series by the significance probability value are as follows:
[0186] If the significance probability value is less than the fourth threshold, the sequence is determined to be stationary and the difference order d=0 is recorded;
[0187] If the significance probability value is greater than or equal to the fourth threshold, the sequence is judged to be non-stationary and the step-by-step difference process is entered (for example, if the first-order difference process is performed first and it is still not stationary, the second-order difference process is performed). The ADF test is repeated after each difference until the sequence is stationary.
[0188] It should be noted that the fourth threshold is set to 0.05, and the setting basis is the traditional standard of significance test in statistics. Specifically, when the significance probability value of the ADF test is less than 0.05, it means that there is sufficient statistical evidence to reject the null hypothesis of the existence of a unit root, so that the time series can be considered to be in a stationary state. By reasonably setting this threshold, the present application can effectively identify the changes in structural characteristics in power load data, improve the accuracy of causal relationship identification, thereby providing more reliable data support for power material forecasting, and further enhancing the system's adaptability to load fluctuations and the optimization level of resource allocation.
[0189] The stabilized sequence and its optimal difference order d (i.e., the order that first satisfies the significance probability value of < 0.05) are recorded as the benchmark parameters for subsequent order adjustments.
[0190] S33-3: Using the Bayesian Information Criterion to Preliminarily Select an Upper Limit L for the Lag Order max .
[0191] S33-4: Dynamically adjust the lag order L according to the stationary sequence determination result. The adjustment strategy is as follows:
[0192] If the sequence is a stationary sequence, that is, the ADF test rejects the null hypothesis, then the optimal lag order L selected by the Bayesian Information Criterion is directly used. opt Conduct Granger causality test;
[0193] If the sequence is non-stationary but the order after stabilization is d=1, a larger lag order (such as L=L opt +1) to compensate for the information loss caused by the differential operation;
[0194] If the sequence is non-stationary and d≥2 after stationary, increase or decrease the lag order (such as L=L opt -1) to avoid information loss caused by over-differentiation.
[0195] S33-5: Based on the dynamically adjusted lag order L, a Granger causality test is performed to calculate the F statistic and significance probability value. If the significance probability value is <0.05, it is considered that a Granger causal relationship exists; otherwise, it is considered that there is no Granger causal relationship. The variables judged to have a Granger causal relationship are used to generate an optimized feature set, that is, the final intervention variable set.
[0196] It should be noted that the F statistic used in the Granger causality test assesses whether the model fit significantly improves after introducing an external variable. Its value is calculated by comparing the sum of squares of the residuals after constructing the full regression model with the controlled model, reflecting the statistical strength of the causal influence between the variables. The F value and its corresponding significance probability can be automatically calculated by statistical analysis software, ensuring the accuracy and efficiency of causal identification. The significance probability here refers to the significance probability of the F statistic under the F distribution. Its meaning is the same as the significance probability used in the ADF test to determine the presence of a unit root; both are used to determine whether to reject the null hypothesis. However, because the assumptions set by the Granger test differ from those of the ADF test, the significance probability must be recalculated to ensure the rigor and scientific nature of the causal judgment. This significance judgment rule can further enhance the reliability of causal identification and the targeted feature selection in power load analysis.
[0197] The larger the F value, the better the unconstrained model fits the data than the constrained model, that is, the greater the possibility of Granger causality.
[0198] S33-6: Use a rolling window to test Granger causality and observe whether it changes over time. If Granger causality is unstable, re-evaluate the data preprocessing process and optimize the lag order selection.
[0199] S4: Perform anomaly detection on the final set of intervention variables, classify the anomalies according to the degree of detected anomalies, and dynamically adjust the filling strategy through Bayesian optimization to minimize the data correction error, and finally output the cleaned training dataset.
[0200] Specifically, anomaly detection is performed on the final set of intervention variables, anomalies are classified according to the degree of detected anomalies, and the filling strategy is dynamically adjusted through Bayesian optimization to minimize the data correction error. Finally, the cleaned training dataset is output, which includes the following steps:
[0201] S41: Use variational autoencoder (VAE) to extract a low-dimensional representation of the final set of intervention variables and calculate the reconstruction error.
[0202] It should be noted that VAE, as an unsupervised learning method, has been widely used in the field of anomaly detection. By maximizing the likelihood of estimating the latent variable distribution of learning data, it can identify abnormal data points that do not conform to normal patterns (reference: Kingma & Welling, 2014, Auto-Encoding Variational Bayes).
[0203] S42: In order to further improve the detection accuracy, this method combines the Mahalanobis distance to measure the degree of deviation of data points from the normal distribution, sets a threshold to identify abnormal data points, and by further introducing the Mahalanobis distance, it can supplement the shortcomings of VAE in identifying boundary samples.
[0204] It should be noted that a threshold is set to distinguish normal data from abnormal data. The threshold can be determined based on the statistical distribution of historical data or a dynamic adjustment strategy. When both the reconstruction error and the Mahalanobis distance of a data point exceed the set threshold, the data point is considered an outlier. The Mahalanobis distance is a statistical method for measuring the similarity of multidimensional data and has been widely used for anomaly detection. Its advantage lies in its ability to consider the correlation between variables, effectively identifying outliers that significantly deviate from the normal data distribution (see: De Maesschalck et al., 2000, The Mahalanobis distance).
[0205] S43: For detected anomaly errors, the anomaly data is classified based on the joint distribution of the reconstruction error and the Mahalanobis distance, distinguishing between minor and major anomalies for more targeted subsequent processing. This classification method combines existing unsupervised learning anomaly detection techniques (such as VAE) with statistical anomaly detection methods (such as Mahalanobis distance). It is a common data anomaly processing strategy and has been widely used in fields such as time series data analysis and power load forecasting.
[0206] Data points identified as severely outliers are directly removed; data points identified as slightly outliers are corrected using Bayesian optimization to select the optimal interpolation strategy. Existing research has shown that Bayesian optimization has excellent global search capabilities for high-dimensional non-convex optimization problems. It can find the optimal solution among various interpolation methods (such as local weighted regression and time series interpolation), ensuring smoothness and consistency of the corrected data (see: Snoek et al., 2012, Practical Bayesian Optimization of Machine Learning Algorithms).
[0207] S44: Integrate the corrected data and remove serious outliers to generate a cleaned training data set to ensure the stability of data quality and provide high-quality input data for subsequent causal analysis and demand forecasting.
[0208] It should be noted that this application introduces an anomaly detection strategy based on a joint recognition mechanism of a variational autoencoder and Mahalanobis distance, combined with a Bayesian optimization interpolation correction method. This effectively improves the stability and representativeness of the data training set and significantly enhances the accuracy and adaptability of subsequent Granger causality identification in processing power load time series data. In particular, in scenarios where data quality is uncontrollable and power system loads fluctuate drastically, this solution can effectively suppress the misleading influence of abnormal disturbances on causal inference results, improve the system's reliability in identifying key variables, and provide a more robust model foundation for the optimal allocation of power resources.
[0209] S5: Based on the training dataset, a time series forecasting model is used to construct an initial power material demand forecasting model and the model is dynamically optimized. During the dynamic optimization process, the Actor-Critic reinforcement learning framework is introduced, and the experience replay mechanism and TD-error are used to update the Actor strategy parameters.
[0210] It should be noted that time series forecasting is an important method for predicting future data points. It is widely used in scenarios such as power load forecasting and material demand forecasting. When using it, the time series forecasting model is trained using historical power material demand data to obtain preliminary forecast results.
[0211] In order to improve prediction accuracy and adaptability, the Actor-Critic structure in reinforcement learning is used for dynamic optimization. Its core idea is:
[0212] Actors are responsible for making decisions (forecasting material demand) and making adjustments based on environmental feedback;
[0213] Critic is responsible for evaluating the decision quality of Actor and providing improvement directions.
[0214] In addition, during the model optimization process, an experience replay mechanism is introduced to improve sample utilization and reduce model overfitting.
[0215] S6: Input the real-time power grid data into the optimized forecasting model to generate the power material demand forecast for the future period.
[0216] In summary, the present invention uses dynamic coupling analysis and triple causal verification mechanism, quantifies intervention effects through Do-Calculus, eliminates confounding bias through propensity score matching, and confirms time series causality through Granger test, which can accurately screen real causal variables and avoid pseudo-correlation interference. In addition, combining ADF stationarity test with Bayesian information criterion to dynamically optimize the lag order, the model can adapt to the time-varying characteristics of power data, significantly improve the ability to capture sudden demands and medium- and long-term trends, and provide reliable support for the precise allocation of materials. The present invention adopts a time domain-frequency domain joint analysis dynamic weight mechanism to distinguish the impact of peak and valley periods, and ensures the stability of the model in data imbalance scenarios through remedial measures (variable addition / matching radius adjustment), while the adaptive lag order compensation technology of non-stationary series significantly reduces the information loss caused by differential operations, so that the prediction model still maintains superior performance under complex working conditions and improves prediction accuracy.
[0217] Example 2, an embodiment of the present invention, provides a power material demand forecasting and analysis system based on big data analysis, including:
[0218] Data acquisition and preprocessing module, used to collect power-related data and perform preprocessing and feature screening on the data;
[0219] Candidate intervention variable identification module, used to identify candidate intervention variables with causal effects and provide key input features for subsequent predictions;
[0220] Causal verification and feature optimization module, used to dynamically adjust the lag order and optimize the feature set;
[0221] The anomaly detection and data filling module uses a variational autoencoder combined with Mahalanobis distance to detect anomalies, classify abnormal data, and dynamically adjust the filling strategy through Bayesian optimization to minimize data correction errors.
[0222] The power material demand forecasting model module establishes a power material demand forecasting model based on the optimized feature set, adopts the Actor-Critic structure for dynamic optimization, and uses the experience replay mechanism to continuously update the forecasting strategy;
[0223] The real-time data input and forecast output module receives real-time power grid data, inputs it into the optimized forecast model, generates forecast results of power material demand in future time periods, and provides visual display and decision support;
[0224] The matching and balance test module uses a two-sample t-test to evaluate the matching balance between the experimental group and the control group after propensity score matching, and determines whether the matching is successful or unsuccessful based on the set thresholds of the standardized mean difference and significance value;
[0225] The dynamic adjustment and optimization module triggers remedial measures when the matching fails to reach equilibrium, optimizes the lag order selection strategy, and dynamically evaluates the stability of Granger causality using the rolling window method.
[0226] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A method for predicting and analyzing power material demand based on big data analysis, characterized in that: The following steps are involved: Obtain target power related data and pre-process the target power related data. Includes feature screening; performing a coupling analysis on the failure rate of power grid equipment and load fluctuation based on the pre-processed target power-related data to obtain a set of candidate intervention variables; Performing a screening operation on the candidate intervention variable set to obtain a final intervention variable set, wherein the screening operation includes screening through a causal validation process; Anomaly detection is performed on the final set of intervention variables, and the anomalies are classified according to the degree of the detected anomalies. The filling strategy is dynamically adjusted through Bayesian optimization to minimize the data correction error, and finally the cleaned training data set is output. The anomaly detection is performed using a variational autoencoder combined with Mahalanobis distance.
2. The method for predicting and analyzing power material demand based on big data analysis according to claim 1, characterized in that: Also includes, Based on the training data set, an initial power material demand forecasting model is established and dynamically optimized. During the dynamic optimization process, an Actor-Critic structure is introduced, and the experience replay mechanism is used to update the strategy. Input real-time power grid data into the optimized forecasting model to generate forecasts of power material demand in future time periods.
3. The method for predicting and analyzing power material demand based on big data analysis according to claim 1, characterized in that: Based on the preprocessed target power-related data, a coupling analysis is performed on the power grid equipment failure rate and load fluctuation to obtain a set of candidate intervention variables, including quantifying the coupling relationship between the equipment failure rate and load fluctuation through a time domain-frequency domain joint analysis method.
4. The method for predicting and analyzing power material demand based on big data analysis according to claim 3, characterized in that: The time domain analysis includes the following steps: The pre-processed load data is divided into peak, flat and valley periods according to the power dispatch cycle, and different weighting factors are assigned to each period; Calculate the absolute difference between the failure rate series and the load series at each time point, introduce the time period weight factor to weight the difference, calculate the grey correlation degree based on the weighted difference, and retain the variables with grey correlation degree greater than the first set value as the time domain candidate set; The frequency domain analysis further performs frequency domain coupling verification based on the time domain candidate set output by the time domain analysis, including the following steps: Perform Morlet wavelet transform on the load sequence and failure rate sequence in the time domain candidate set; Extract wavelet coefficients of characteristic frequency bands; Calculating the wavelet coherence within the characteristic frequency band, and retaining the variables whose coherence is greater than the second set value and whose phase difference is less than the third set value; The time domain grey correlation value and the frequency domain coherence coefficient are weighted and integrated to generate the final coupling strength score. The candidate intervention variables are sorted in descending order according to the coupling strength score, and the first b variables with the highest scores are selected as the final candidate intervention variable set.
5. The method for predicting and analyzing power material demand based on big data analysis according to claim 1, characterized in that: The causal verification process includes the following steps: Do-Calculus was used to calculate the intervention effect value θ of power material demand on the intervention variable to quantify the causal impact; Propensity score matching was implemented to control the matching bias between the experimental group and the control group; Adjust the lag order L for the Granger causality test.
6. The method for predicting and analyzing power material demand based on big data analysis according to claim 5, characterized in that: The implementation of propensity score matching to control the matching bias between the experimental group and the control group includes the following steps: The three power characteristics of equipment operation age, ambient temperature, and load rate in the power-related data are mandatory and included as matching variables; The radius matching method was used, and the matching radius was set to N times the standard deviation of the propensity score; After the matching is completed, check whether the experimental group and the control group have reached a balance. If so, the matched samples are used as input data to enter the subsequent causal effect analysis stage.
7. The method for predicting and analyzing power material demand based on big data analysis according to claim 6, characterized in that: After the matching is completed, the two-sample t-test is used to evaluate the significance of the difference between the groups, and the standardized mean difference method is combined to jointly determine the balance between the experimental group and the control group. The joint determination rules and results are as follows: When the standardized mean difference is less than or equal to the first threshold, and the significance value is greater than or equal to the second threshold, the match is considered successful. When the standardized mean difference is greater than the first threshold and the significance value is less than the second threshold, the match is determined to be unsuccessful and remedial measures are triggered. If the conditions are still not met, the intervention variable is eliminated. When the standardized mean difference is less than or equal to the first threshold and the significance value is less than the second threshold, further judgment process is required: If the standardized mean difference is less than or equal to the third threshold, the difference is considered significant but negligible, and the match is considered successful. If the third threshold is less than or equal to the standardized mean difference and less than or equal to the first threshold, the match is considered unsuccessful, triggering remedial measures. If the condition is still not met, the intervention variable is eliminated. When the standardized mean difference is greater than the first threshold and the significance value is greater than the second threshold, it means that the difference is large but not statistically significant. The match is judged as unsuccessful and remedial measures are triggered. If the condition still does not meet the requirements, the intervention variable is eliminated. When the initial matching fails to meet the balance requirement, the following remedial measures are triggered in sequence: Add the number of equipment maintenance times as an additional matching variable and execute the joint decision rule to determine whether the matching success criteria are met; Expand the matching radius; If the conditions are still not met, the intervention variable will be automatically eliminated.
8. The method for predicting and analyzing power material demand based on big data analysis according to claim 5, characterized in that: The method of adjusting the lag order L of the Granger causality test comprises the following steps: Determine the initial lag order according to the periodic characteristics of power load data; Perform ADF test on the original sequence and obtain the significance probability value during the test; The significance probability value is used to determine whether a time series is a stationary series. The judgment rules are as follows: If the significance probability value is less than the fourth threshold, the sequence is determined to be stationary and the difference order d=0 is recorded; If the significance probability value is greater than or equal to the fourth threshold, the sequence is judged to be non-stationary and the step-by-step difference process is entered. The ADF test is repeated after each difference until the sequence becomes stationary. Record the stabilized sequence and its optimal difference order d as the benchmark parameter for subsequent order adjustment; Use the Bayesian Information Criterion to preliminarily select the upper limit L of the lag order max ; Dynamically adjust the lag order L according to the stationarity results; Based on the dynamically adjusted lag order L, a Granger causality test is performed to calculate the F statistic and significance probability value. If the significance probability value is less than the fourth threshold, it is considered that a Granger causality relationship exists; otherwise, it is considered that no Granger causality relationship exists. The variables judged to have a Granger causal relationship are used to generate an optimized feature set, that is, the final intervention variable set. A rolling window is used to test Granger causality to observe whether it changes over time. If the Granger causality is unstable, the data preprocessing process is re-evaluated and the lag order selection is optimized.
9. The method for predicting and analyzing power material demand based on big data analysis according to claim 8, characterized in that: The process of the step-by-step difference process and the strategy for dynamically adjusting the lag order L are as follows: The step-by-step difference process is as follows: Perform step-by-step difference processing on the non-stationary series, and re-perform the ADF test after each difference until the series is stationary; Record the stabilized sequence and its optimal difference order d as the benchmark parameter for subsequent order adjustment; The strategy for dynamically adjusting the lag order L is as follows: If the sequence is a stationary sequence, that is, the ADF test rejects the null hypothesis, then the optimal lag order L selected by the Bayesian Information Criterion is directly used. opt Conduct Granger causality test; If the sequence is non-stationary but the order after stationarization is d=1, a larger lag order is used to compensate for the effect of the difference operation on information loss; If the series is non-stationary and d ≥ 2 after stationarity, increase or decrease the lag order to avoid information loss due to over-differentiation.
10. A system for forecasting and analyzing power material demand based on big data analysis, based on the method for forecasting and analyzing power material demand based on big data analysis according to any one of claims 1 to 9, characterized in that: include, Data acquisition and preprocessing module, used to collect power-related data and perform preprocessing and feature screening on the data; Candidate intervention variable identification module, used to identify candidate intervention variables with causal effects and provide key input features for subsequent predictions; Causal verification and feature optimization module, used to dynamically adjust the lag order and optimize the feature set; The anomaly detection and data filling module uses a variational autoencoder combined with Mahalanobis distance to detect anomalies, classify abnormal data, and dynamically adjust the filling strategy through Bayesian optimization to minimize data correction errors. The power material demand forecasting model module establishes a power material demand forecasting model based on the optimized feature set, adopts the Actor-Critic structure for dynamic optimization, and uses the experience replay mechanism to continuously update the forecasting strategy; The real-time data input and forecast output module receives real-time power grid data, inputs it into the optimized forecast model, generates forecast results of power material demand in future time periods, and provides visual display and decision support; The matching and balance test module uses a two-sample t-test to evaluate the matching balance between the experimental group and the control group after propensity score matching, and determines whether the matching is successful or unsuccessful based on the set thresholds of the standardized mean difference and significance value; The dynamic adjustment and optimization module triggers remedial measures when the matching fails to reach equilibrium, optimizes the lag order selection strategy, and dynamically evaluates the stability of Granger causality in combination with the rolling window method.
Citation Information
Cited By
Flow filling rate gap complement prediction and correction method
CN120952881A
Flow filling rate gap filling prediction and correction method
CN120952881B
Engineering material purchasing optimization method and system considering weather and price fluctuation
CN121660175A
Supply and demand management and control method and system for electric power material storage
CN122222535A