Power load prediction method, device and system based on large model, and storage medium
By constructing a time-varying causal correlation topology and a large-model causal gradient transfer mechanism, the influencing factors of the power system are separated and the prediction parameters are optimized, which solves the problem of incomplete consideration of factors in existing power load prediction methods and achieves higher accuracy and reliability in prediction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ELECTRIC POWER RESEARCH INSTITUTE OF STATE GRID SHANDONG ELECTRIC POWER COMPANY
- Filing Date
- 2025-12-16
- Publication Date
- 2026-04-21
AI Technical Summary
Existing power load forecasting methods fail to fully consider various influencing factors, resulting in low forecast accuracy and an inability to adapt to the complexity and dynamism of power systems.
By constructing an initial causal correlation topology with time-varying causal strength, separating direct and indirect influencing factors, adjusting prediction parameters using the causal gradient transfer mechanism within the large model, and performing dimensional deviation causal tracing, the power load prediction results are optimized.
It improves the accuracy and reliability of power load forecasting, can adapt to the complexity and dynamics of power systems, and generates forecast results that are more in line with the actual situation.
Smart Images

Figure CN121906398A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of power load forecasting technology, specifically to a power load forecasting method, device, system, and storage medium based on a large model. Background Technology
[0002] In the field of electricity load forecasting, most existing forecasting methods have certain limitations. Traditional electricity load forecasting methods, such as time series analysis, mainly rely on historical electricity load data for modeling and forecasting. These methods typically assume that changes in electricity load are only related to their own historical data, ignoring the effects of many other influencing factors. However, electricity load is affected by a combination of factors, including the operating status of power system equipment, external environmental conditions, and user electricity consumption behavior. Relying solely on historical data makes it difficult to comprehensively and accurately capture the changing patterns of electricity load, resulting in low forecast accuracy.
[0003] In recent years, some methods have begun to consider incorporating external influencing factors to improve forecast accuracy. However, these methods often employ simple and fixed correlation methods when dealing with the relationship between influencing factors and power load. They fail to delve into the impact mechanisms of different influencing factors on power load at different time scales, cannot distinguish between direct and indirect influencing factors, and cannot accurately characterize the hierarchical relationships and time-varying characteristics among influencing factors. This makes it impossible to dynamically adjust the weights of each influencing factor according to actual conditions during the forecasting process, making it difficult to adapt to the complexity and dynamism of the power system, thus limiting the accuracy and reliability of power load forecasting. Summary of the Invention
[0004] This application provides a method, apparatus, system, and storage medium for predicting power load based on a large model.
[0005] The first aspect of this application provides a power load forecasting method based on a large model, comprising:
[0006] Based on the power system operation data set, the power load influencing factors at various time scales are obtained. By constructing an initial causal association topology containing time-varying causal strength, a causal association structure of power load influencing factors containing hierarchical relationships and time-varying characteristics of influencing factors is obtained. Based on the time-varying characteristics of the causal association structure, the data in the operation data set is classified according to time scale to obtain a classified operation data set. The classified operation data set and the time-varying causal strength parameters of the causal association structure are input into a pre-trained power load prediction model. The prediction parameters of each layer are adjusted through the causal gradient transfer mechanism within the large model to obtain the initial power load prediction result. Based on the hierarchical relationship of the causal association structure, the initial power load prediction result is subjected to dimensional deviation causal tracing processing to obtain the hierarchical deviation contribution value. Based on the hierarchical deviation contribution value, the power load prediction processing with time scale matching is re-performed to obtain the optimized power load prediction result.
[0007] In an optional embodiment of this application, based on a power system operation data set, power load influencing factors at various time scales of the power system are obtained. By constructing an initial causal association topology containing time-varying causal strength, a causal association structure of power load influencing factors containing hierarchical relationships and time-varying characteristics is obtained, including:
[0008] The power system operation dataset includes historical power load data, power system equipment operation data, external environmental impact data, and user electricity consumption behavior data.
[0009] In an optional embodiment of this application, based on a power system operation data set, power load influencing factors at various time scales of the power system are obtained. By constructing an initial causal association topology containing time-varying causal strength, a causal association structure of power load influencing factors containing hierarchical relationships and time-varying characteristics is obtained, including:
[0010] A dynamic causal discovery algorithm is used to mine the power load influencing factors of the operational data set in order to separate and obtain the direct and indirect influencing factors of the power system at various time scales.
[0011] In an optional embodiment of this application, based on a power system operation data set, power load influencing factors at various time scales of the power system are obtained. By constructing an initial causal association topology containing time-varying causal strength, a causal association structure of power load influencing factors containing hierarchical relationships and time-varying characteristics is obtained, including:
[0012] Based on the multi-timescale time series characteristics and attribute characteristics of various types of data in the dataset, the correlation degree between various types of data and historical power load data is calculated at different time scales. Data with a correlation degree higher than a preset correlation degree threshold at each time scale are selected to form a dataset of potential influencing factors at different time scales. Pairwise interaction analysis is performed on the data in the dataset of potential influencing factors at each time scale to obtain pairs of influencing factors with unidirectional influence relationships at different time scales. An initial causal association edge set is constructed based on the pairs of influencing factors at different time scales, and an initial causal association topology is constructed based on the initial causal association edge set. The nodes in the initial causal association topology are divided into direct and indirect influencing factors. Based on the hierarchical division results and the time scale attributes of the nodes, the causal association structure is obtained.
[0013] In an optional embodiment of this application, pairwise interaction analysis is performed on the data in the dataset of potential influencing factors at each time scale to obtain pairs of influencing factors with unidirectional influence relationships at different time scales. Based on these pairs of influencing factors, an initial causal relationship edge set for each time scale is constructed. An initial causal relationship topology is then constructed based on this initial causal relationship edge set, including:
[0014] For the data in the potential influencing factor dataset, a time delay variable is set to determine whether there is a unidirectional influence relationship between the data that changes with the time scale, so as to obtain the influencing factor pairs; in the initial causal association edge set, each causal association edge corresponds to a pair of influencing factors with a unidirectional influence relationship, and the direction of influence and the initial influence intensity that changes with the time scale are marked to construct the initial causal association topology. The initial causal association topology uses influencing factors as topology nodes and causal association edges of different time scales as connections between nodes.
[0015] In an optional embodiment of this application, the nodes in the initial causal relationship topology are hierarchically divided into direct and indirect influencing factors. Based on the hierarchical division results and the node time scale attributes, a causal relationship structure is obtained, including:
[0016] Redundancy removal is performed on the causal links in the causal relationship structure, removing duplicate causal links and causal links whose influence intensity is lower than the preset intensity threshold at each time scale; the change rate of influence intensity of the remaining causal links at different time scales is calculated, and the time-varying causal intensity parameters of each link are labeled to obtain the final causal relationship structure of power load influencing factors that includes the hierarchical relationship and time-varying characteristics of influencing factors.
[0017] In an optional embodiment of this application, based on the time-varying characteristics of the causal relationship structure, the data in the runtime dataset is classified according to time scale to obtain a classified runtime dataset, including:
[0018] Based on the time-varying characteristics of the causal relationship structure, the operational data set is divided into causal dimensions according to time scale matching. The data is classified according to the time scale corresponding to the direct and indirect influencing factors, resulting in the classified operational data set.
[0019] In an optional embodiment of this application, based on the time-varying characteristics of the causal relationship structure, the data in the runtime dataset is classified according to time scale to obtain a classified runtime dataset, including:
[0020] From the causal relationship structure, extract all names of influencing factors at both the direct and indirect influencing factor levels at each time scale. Arrange them according to time scale to form separate lists of direct and indirect influencing factors. Based on the influencing factor category and time scale attribute corresponding to each data point in the dataset, data belonging to the direct influencing factor category at the same time scale are grouped into the corresponding direct influencing factor subset, and data belonging to the indirect influencing factor category at the same time scale are grouped into the corresponding indirect influencing factor subset. Standardize the data format of the direct and indirect influencing factor subsets at each time scale. Combine the standardized direct and indirect influencing factor subsets at each time scale and sort them according to time scale order to form a categorized dataset that includes time scale relationships.
[0021] In an optional embodiment of this application, the classified set of operational data and the time-varying causal strength parameters of the causal relationship structure are input into a pre-trained large-scale power load prediction model. The prediction parameters of each layer are adjusted through the causal gradient transfer mechanism within the large-scale model to obtain the initial power load prediction result, including:
[0022] By employing a causal gradient transfer mechanism, the correlation gradient between the parameters of each layer of the large model and the causal feature vector is calculated to determine the adjustment direction of the parameters of each layer of the large model. Based on the correlation gradient and adjustment direction, the weight ratio of the direct influencing factors at each time scale is enhanced. Feature enhancement processing is performed on the data of direct influencing factors, and feature filtering processing is performed on the data of indirect influencing factors. The feature-enhanced data of direct influencing factors and the feature-filtered data of indirect influencing factors are fused across scales to generate fused feature data containing time scale correlations. Load forecasting is performed on the fused feature data according to the large model, and preliminary power load forecasting data at different time scales is output. The preliminary power load forecasting data is then integrated into a time series, and the preliminary power load forecasting data at each scale is arranged according to a unified time axis to obtain the initial power load forecasting results.
[0023] In an optional embodiment of this application, the correlation gradient between the parameters of each layer of the large model and the causal feature vector is calculated through a causal gradient transfer mechanism to determine the adjustment direction of the parameters of each layer of the large model, including:
[0024] The time-varying causal strength parameters are converted into time-varying causal codes for large-scale model identification. The time-varying causal codes contain information on the hierarchical information of influencing factors at each time scale, the influence direction of causal association edges, and time-varying causal strength data. Key causal information in the time-varying causal codes is extracted through a multi-scale causal feature parsing algorithm of the large model to generate causal feature vectors at each time scale. Based on the causal feature vectors at each time scale, the correlation gradient between the parameters of each layer of the large model and the causal feature vectors is calculated through a causal gradient transfer mechanism.
[0025] In an optional embodiment of this application, the correlation gradient between the parameters of each layer of the large model and the causal feature vector is calculated through a causal gradient transfer mechanism to determine the adjustment direction of the parameters of each layer of the large model, including:
[0026] The causal feature vectors at each time scale are input into the large model. The parameter gradient matrices of each layer of the large model are initialized, and the current weight parameters of each layer of the large model are extracted. The weight parameters of each layer of the large model are associated and mapped with the causal feature vectors of the corresponding time scale according to the time scale correspondence. The dot product of the weight parameters of each layer of the large model and the causal feature vector of the corresponding time scale is calculated to obtain the causal correlation value of each layer of the large model. Based on the causal correlation value, the partial derivatives of the parameters of each layer of the large model with respect to the causal feature vector are calculated to generate the local gradient vector of each layer of the power load prediction large model. The local gradient vector is passed in reverse according to the model hierarchy, and the gradient information of adjacent layers is fused to obtain the global correlation gradient of each layer of the large model. The direction of parameter adjustment of the layer is determined according to the positive or negative sign of the global correlation gradient, and the adjustment magnitude of the parameter of the layer is determined according to the absolute value of the global correlation gradient.
[0027] In an optional embodiment of this application, based on the hierarchical relationship of the causal association structure, the initial power load forecast result is subjected to dimensional deviation causal tracing processing to obtain the hierarchical deviation contribution value. Based on the hierarchical deviation contribution value, the power load forecast is reprocessed with time-scale matching to obtain the optimized power load forecast result, including:
[0028] Based on the contribution value of hierarchical deviation, the inter-level parameter transfer weights and the time-varying causal strength of the causal relationship structure of the large model are adjusted to re-process the power load forecasting with time-scale matching.
[0029] In an optional embodiment of this application, based on the hierarchical relationship of the causal association structure, the initial power load forecast result is subjected to dimensional deviation causal tracing processing to obtain the hierarchical deviation contribution value. Based on the hierarchical deviation contribution value, the power load forecast is reprocessed with time-scale matching to obtain the optimized power load forecast result, including:
[0030] Obtain the actual power load data corresponding to the initial power load forecast results for the same period, calculate the deviation between the initial power load forecast results and the actual power load data at each time scale, and obtain abnormal deviation data; based on the hierarchical relationship of the causal relationship structure, extract all causal relationship paths within the time interval corresponding to the abnormal deviation data at the time scale, extract the influencing factor data in each causal relationship path, and obtain the actual value data of each level of influencing factors within the abnormal deviation time interval according to the hierarchical classification; calculate the relative difference between the actual value data and the historical influencing factor value data of the same time scale, and obtain the abnormal influencing factors; based on the time-varying causal intensity change and the influence chain of the causal relationship path containing abnormal influencing factors, calculate the contribution of each causal relationship path containing abnormal influencing factors to the prediction deviation, and determine the hierarchical deviation contribution value of causal relationship paths containing abnormal influencing factors at different levels.
[0031] A second aspect of the embodiments of this application provides a power load forecasting device based on a large model, including a processor and a memory storing program instructions. The processor is configured to execute the power load forecasting method based on a large model as described in the first aspect of the embodiments of this application when running the program instructions.
[0032] A third aspect of this application provides a system comprising:
[0033] The system itself; and,
[0034] The large-model-based power load forecasting device, as described in the second aspect of this application, is installed on the system body.
[0035] A fourth aspect of the embodiments of this application provides a computer-readable storage medium storing program instructions that, when executed, cause a computer to perform a large-model-based power load forecasting method as described in the first aspect of the embodiments of this application.
[0036] The power load forecasting method, apparatus, system, and storage medium based on large models provided in the embodiments of this application have the following beneficial effects:
[0037] This application embodiment utilizes a dynamic causal discovery algorithm to mine and process power system operation-related data sets. This process accurately separates direct and indirect influencing factors at different time scales, constructs an initial causal relationship topology with time-varying causal strength, and thus presents the causal relationship structure of power load influencing factors. Then, based on the time-varying characteristics of the causal relationship structure, causal dimension partitioning is performed, making data classification more scientific and reasonable, and better reflecting the effects of different influencing factors at different time scales. A pre-trained power load prediction model is invoked, and its internal causal gradient transfer mechanism is used to adjust prediction parameters. Combined with the time-varying causal strength parameters of the causal relationship structure, an initial power load prediction result that better reflects the actual situation can be generated. Dimensional deviation causal tracing processing is performed on the initial prediction result to locate the causal relationship paths of different levels of influencing factors leading to prediction deviations, generating a deviation causal tracing report. Based on the deviation tracing report, the inter-layer parameter transfer weights and the time-varying causal strength of the causal relationship structure are adjusted, and prediction processing is performed again. Finally, an optimized power load prediction result is obtained, greatly improving the accuracy and reliability of the prediction and effectively adapting to the complexity and dynamism of the power system. Attached Figure Description
[0038] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments of this application and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0039] Figure 1 This is a schematic diagram of the power load forecasting method based on a large model provided in the embodiments of this application;
[0040] Figure 2 This is a schematic diagram of a power load forecasting device based on a large model provided in an embodiment of this application.
[0041] Figure label:
[0042] 800: Power load forecasting device based on large model; 801: Processor; 802: Memory; 803: Communication interface; 804: Bus. Detailed Implementation
[0043] To make the technical solutions and advantages of the embodiments of this application clearer, the exemplary embodiments of this application will be described in further detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not an exhaustive list of all embodiments. It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of this application can be combined with each other.
[0044] Figure 1This is a schematic diagram of a power load forecasting method based on a large model provided in an embodiment of this application. The method can be executed in the system or in a server or terminal device that is connected to the system.
[0045] Combination Figure 1 As shown in the figure, this application provides a power load forecasting method based on a large model, including:
[0046] S100: Based on the power system operation data set, obtain the power load influencing factors at various time scales of the power system. By constructing an initial causal relationship topology containing time-varying causal strength, obtain the causal relationship structure of power load influencing factors that includes the hierarchical relationship and time-varying characteristics of the influencing factors.
[0047] In one optional embodiment of this application, the power system operation dataset includes historical power load data, power system equipment operation data, external environmental impact data, and user electricity consumption behavior data.
[0048] In an optional embodiment of this application, a dynamic causal discovery algorithm is used to mine the power load influencing factors of the operating data set in order to separate and obtain the direct and indirect influencing factors of the power system at various time scales.
[0049] In this embodiment, the application scenario is short-term power load forecasting for a power grid in the core area of a city. This area includes a high-density commercial building complex, residential communities, and a small number of industrial users. The power system operation-related data set is stored in a hybrid architecture of relational and time-series databases in a local data center. Historical power load data consists of active power records collected every 15 minutes, including three consecutive years of data. Power system equipment operation data covers real-time monitoring data such as transformer load rate and circuit breaker opening / closing status of 110kV substations within the area. External environmental impact data integrates hourly temperature, humidity, precipitation data, and holiday schedule data from the city's meteorological bureau. User electricity consumption behavior data is categorized by industry, including commercial users' electricity consumption records during business hours and residential users' time-of-use pricing response data.
[0050] In this embodiment, a dynamic causal discovery algorithm is used to mine and process the power load influencing factors in the power system operation-related data set, separating the direct and indirect influencing factors at different time scales, constructing an initial causal association topology containing time-varying causal strength, and obtaining a causal association structure of power load influencing factors that includes the hierarchical relationship and time-varying characteristics of influencing factors. The power system operation-related data set includes historical power load data, power system equipment operation data, external environmental impact data, and user electricity consumption behavior data.
[0051] In an optional embodiment of this application, based on the multi-timescale time series characteristics and attribute characteristics of various types of data in the dataset, the correlation degree between various types of data and historical power load data is calculated at different time scales. Data with a correlation degree higher than a preset correlation degree threshold at each time scale are selected to form a dataset of potential influencing factors at different time scales. Pairwise interaction analysis is performed on the data in the dataset of potential influencing factors at each time scale to obtain pairs of influencing factors with unidirectional influence relationships at different time scales. An initial causal association edge set for each time scale is constructed based on the pairs of influencing factors. An initial causal association topology is constructed based on the initial causal association edge set. The nodes in the initial causal association topology are divided into direct and indirect influencing factors. Based on the results of the hierarchical division and the time scale attributes of the nodes, the causal association structure is obtained.
[0052] In an optional embodiment of this application, for the data in the potential influencing factor dataset, a time delay variable is set to determine whether there is a one-way influence relationship between the data that changes with the time scale, so as to obtain the influencing factor pairs; in the initial causal association edge set, each causal association edge corresponds to a pair of influencing factors with a one-way influence relationship, and the direction of influence and the initial influence intensity that changes with the time scale are marked, so as to construct an initial causal association topology. The initial causal association topology uses influencing factors as topology nodes and causal association edges of different time scales as connections between nodes.
[0053] In an optional embodiment of this application, redundancy removal is performed on the causal association edges in the causal association structure, deleting duplicate causal association edges and causal association edges whose influence intensity is lower than a preset intensity threshold at each time scale; the change rate of influence intensity of the remaining causal association edges at different time scales is calculated, and the time-varying causal intensity parameters of each edge are labeled to obtain the final causal association structure of power load influencing factors that includes the hierarchical relationship and time-varying characteristics of influencing factors.
[0054] Furthermore, in the embodiments of this application:
[0055] S110: Extract multi-timescale time series features and attribute features of various types of data in the power system operation-related data set. The multi-timescale time series features include the trend features of data changes over time at multiple time scale levels, and the attribute features include the data type features and value range features. Establish a multi-dimensional data feature index table.
[0056] Specifically, for historical power load data, the extraction of time series features across multiple time scales is achieved through a sliding time window technique. At the 15-minute time scale, the maximum, minimum, and average load values within each sliding window are calculated to reflect short-term load fluctuations. At the daily time scale, the shape of the daily load curve is calculated, such as the timing and duration of peaks and troughs. At the weekly time scale, the load differences between different days within a week are analyzed. Trend features are represented by the slope of the fitted line obtained through linear fitting of the series data at different time scales, indicating the overall trend of the data. Regarding attribute features, type features are determined through metadata information of data fields, such as marking the "active power" field as numerical continuous data. Value range features are determined by the frequency of statistical data occurrences within historical periods, with the range of values within a specific proportion of cumulative frequency taken as the valid value range. Attribute features of power system equipment operation data also include the rated parameters of the equipment, such as the rated capacity of transformers, which are obtained by linking to an equipment archive database. The multi-dimensional data feature index table is stored in the form of a database table. Each record in the table corresponds to the feature information of a type of data at a time scale, including fields such as data identifier, time scale type, feature name, feature description, and feature value range.
[0057] S120, based on a multi-dimensional data feature index table, calculates the correlation between various types of data and historical power load data at different time scales, filters out data with a correlation higher than a preset correlation threshold at each time scale, and forms a data set of potential influencing factors at different time scales.
[0058] Specifically, based on the multi-dimensional data feature index table, for each time scale, time series from various data types that coincide with historical power load data are selected. The correlation calculation employs a distance-based similarity metric, treating the time series of the two data types as two vectors, calculating the standardized Euclidean distance between the vectors, and then converting the distance value into a correlation score; the smaller the distance, the greater the correlation. The preset correlation threshold is determined based on the actual distribution of the data. Through statistical analysis of the correlation scores of historical data, a value that represents a reasonable distribution of correlation scores is selected as the threshold. For a 15-minute time scale, the correlation score between each type of 15-minute sampled data and the 15-minute power load data is calculated; for a daily time scale, the correlation score between daily statistical data and daily power load statistical data is calculated, and so on. Data with correlation scores higher than the corresponding threshold at each time scale are selected and categorized by time scale to form a potential influencing factor data set. Each set is stored in a folder, with the folder name containing the time scale identifier and data category information.
[0059] S130 performs pairwise interaction analysis on the data in the dataset of potential influencing factors at each time scale, sets a time delay variable to determine whether there is a unidirectional influence relationship between the data that changes with the time scale, and marks the pairs of influencing factors that have a unidirectional influence relationship at different time scales.
[0060] Specifically, the dataset of potential influencing factors at a 15-minute timescale includes 15-minute temperature data, humidity data, and commercial user electricity consumption data. Any two types of data are selected from this dataset, such as 15-minute temperature data and 15-minute commercial user electricity consumption data, to form a data pair to be analyzed. The interaction between these two types of data is analyzed to determine whether a unidirectional influence relationship exists. The time delay variable is set considering the temporal resolution of the data and potential lag effects. For 15-minute data, multiple different delay steps are set, such as one sampling interval, two sampling intervals, and four sampling intervals. The direction of influence between the data is determined through lag processing and correlation analysis.
[0061] Furthermore, regarding step S130:
[0062] S131, Select the first data and the second data from the dataset of potential influencing factors at any time scale to form the data pair to be analyzed at that time scale.
[0063] Specifically, in the dataset of potential influencing factors at the daily time scale, the data is stored on a daily basis and includes daily average temperature data, daily total precipitation data, and daily industrial electricity consumption data. Daily average temperature data is selected as the first data set, and daily industrial electricity consumption data is selected as the second data set, forming a data pair to be analyzed at this time scale. During the selection process, it is ensured that the time ranges of the two types of data are completely consistent, both containing the same number of days.
[0064] S132, extract the time series data of the first data and the time series data of the second data in the data pair to be analyzed, so that the time granularity of the time series data of the first data and the time series data of the second data is consistent with the time scale to which they belong.
[0065] Specifically, a time series of daily average temperature data is extracted, where each data point corresponds to the average temperature value for one day, with a time granularity of one day. Similarly, a time series of daily industrial electricity consumption data is extracted, where each data point represents the total industrial electricity consumption for one day, also with a time granularity of one day. The two time series are then time-aligned to ensure that each time point has corresponding temperature and electricity consumption data. For missing time points, interpolation methods are used to supplement the data, ensuring the completeness of the time series and the consistency of the time granularity.
[0066] S133, set a time delay variable that matches the time scale, set multiple different delay steps, perform lag processing on the time series data of the first data, and generate multiple first data lag sequences corresponding to different delay steps.
[0067] Specifically, for the daily time scale, the delay step for the time delay variable is set to 1 day, 2 days, 3 days, and 7 days. Based on the time series of daily average temperature data, when the delay step is 1 day, each data point in the original series is shifted forward by 1 day to generate a lag series of daily average temperature data with a 1-day lag. Similarly, lag series with lags of 2 days, 3 days, and 7 days are generated. The length of the lag series is the same as the original time series. For any missing data at the beginning of the lag series, historical average data from that series is used to fill in the gaps.
[0068] S134, calculate the cross-correlation between the time series data of each first data lag sequence and the second data to obtain the cross-correlation coefficient sequence at that time scale.
[0069] Specifically, the generated first data lag sequences are cross-correlation calculated with the time series of the second data. During the calculation, for each lag sequence, starting from the first common time point, the correlation coefficient between the two sequences is calculated for all corresponding data points at and after that time point. The correlation coefficients calculated for each lag step are arranged in order to form a cross-correlation coefficient sequence for that time scale.
[0070] S135, analyze the cross-correlation coefficient sequence to determine if there is a delay step corresponding to the maximum value. If the maximum value appears at the delay step where the first data lags behind the second data, it is preliminarily judged that the second data may affect the first data at this time scale.
[0071] Specifically, the cross-correlation coefficient sequence is scanned to find the maximum value in the sequence and the corresponding lag step. If the maximum value occurs when the first data lags behind the second data, that is, the time series of the second data changes before the lagged series of the first data, then it is preliminarily considered that at this time scale, the change in the second data may cause the change in the first data, that is, the second data may affect the first data.
[0072] S136, if the maximum value appears when the second data lags behind the delay step of the first data, it is preliminarily determined that the first data may affect the second data at this time scale.
[0073] Specifically, when the maximum value of the cross-correlation coefficient sequence appears when the second data lags behind the delay step of the first data, it indicates that the time series of the first data changes before the time series of the second data. Therefore, it can be preliminarily judged that the first data may have an impact on the second data at this time scale.
[0074] S137, Replace the time series data of the first and second data under this time scale with another part, repeat the lag processing and cross-correlation calculation steps, and verify the consistency of the preliminary judgment results under this time scale.
[0075] Specifically, to verify the reliability of the preliminary judgment, time series of the first and second data points from another time period are selected from the dataset of potential influencing factors at this time scale. For example, if the initial analysis used data from the first half of the year, then data from the second half of the year is selected this time. Following steps S133 to S136, lag processing and cross-correlation calculations are performed again to obtain new judgment results, which are then compared with the preliminary judgment results.
[0076] For example, regarding step S137:
[0077] S137(1), determine the first data type, the second data type and the time scale corresponding to the preliminary judgment result, and find another part of time series data that is the same as the first data type and the second data type and belongs to the same time scale.
[0078] Based on the preliminary assessment, the primary data type (e.g., daily average temperature data), the secondary data type (e.g., daily industrial electricity consumption data), and the corresponding time scale (e.g., daily) were identified. The database was then used to search for similar data within another time period based on the data type and time scale identifier, ensuring that the time range of the new data did not overlap with the initial analysis data, and that the data collection conditions and environment were essentially consistent.
[0079] S137(2), extract the first data new sequence and the second data new sequence from another part of the time series data, so that the time length and time granularity of the first data new sequence and the second data new sequence are consistent with the time series data of the first data and the second data in the initial analysis.
[0080] From the retrieved time series data, extract the first and second new data sequences. Check the time length of the new sequences to ensure they contain the same number of time points as the initial analysis. Maintain the same daily time granularity as the initial analysis. For missing data in the new sequences, use the same interpolation method as the initial analysis to ensure data quality.
[0081] S137(3), set the same time delay variable and delay step as the initial analysis, perform lag processing on the first data new sequence, and generate multiple first data new lag sequences with different delay steps.
[0082] The time delay variables were set exactly the same as in the initial analysis, including the number and specific values of the delay steps. The first new data series was lagged to generate a new lagged series with the same number and delay steps as in the initial analysis. The lag processing method and the method of filling in missing data were also consistent with the initial analysis.
[0083] S137(4), calculate the cross-correlation between each new lag sequence of the first data and the new sequence of the second data to obtain the new cross-correlation coefficient sequence at this time scale.
[0084] The new lagged sequence of the first data was cross-correlation calculated with the new sequence of the second data. The calculation process and parameter settings were the same as in the initial analysis, resulting in a new cross-correlation coefficient sequence.
[0085] S137(5), extract the delay step corresponding to the maximum value from the new cross-correlation coefficient sequence, and determine the influence direction corresponding to the delay step, which is the new judgment result.
[0086] The new cross-correlation coefficient sequence is analyzed to find the maximum value and its corresponding delay step, and the new direction of influence is determined based on the delay step to form a new judgment result.
[0087] S137(6), compare the influence direction and corresponding delay step length of the initial judgment result and the new judgment result. If the influence direction is the same and the difference in delay step length is within the preset range, then the initial judgment result is consistent at this time scale.
[0088] Compare the initial and new judgment results to see if their influence directions are the same. If the influence directions are the same, further calculate the difference in the corresponding delay step lengths between the two results. A preset range for the delay step length difference is established, determined by the size of the time scale; for a daily time scale, the preset range is 1 day. If the delay step length difference falls within this range, the initial judgment result is considered consistent at that time scale.
[0089] S137(7), if the direction of influence is different or the difference in delay step exceeds the preset range, then select the time series data of the first and second data of the third part under the time scale, repeat the lag processing and cross-correlation calculation steps, and obtain the third judgment result.
[0090] When the initial judgment result and the new judgment result have different directions of influence, or when the difference in delay step size exceeds the preset range, time series data of the first and second data of the third time period are selected from the dataset of potential influencing factors at that time scale. Lag processing and cross-correlation calculations are then performed according to the same procedure described above to obtain the third judgment result.
[0091] S137(8), compare the initial judgment result, the new judgment result and the third judgment result. If at least two of the three have the same influence direction and the difference in delay step size is within the preset range, then the initial judgment result is determined to be valid.
[0092] The initial judgment result, the new judgment result, and the third judgment result are compared pairwise. If two or three of the judgment results have the same direction of influence and the difference in their delay step size is within a preset range, the preliminary judgment result is considered valid and acceptable.
[0093] S137(9), if the three factors have different directions of influence or the difference in delay step size exceeds the preset range, return the dataset of potential influencing factors at this time scale and reselect the data pair to be analyzed.
[0094] If the three judgment results have different directions of influence, or if the difference in delay step size between any two results exceeds the preset range, it indicates that the currently selected data pair to be analyzed may not have a stable influence relationship. It is necessary to return to the data set of potential influencing factors at this time scale and reselect other data pairs for analysis.
[0095] S137(10), record the judgment result of successful verification, the corresponding time scale and delay step.
[0096] For verified results, the direction of influence, the time scale, and the corresponding delay step size will be recorded in the analysis report.
[0097] In this embodiment of the application, after performing the above steps, further:
[0098] S138. If the two judgment results are consistent, it is determined that there is a unidirectional influence relationship between the data pairs to be analyzed at this time scale. The direction of influence and time scale attribute are marked to form the influencing factor pair at this time scale.
[0099] Specifically, when the influence directions of the two judgment results are the same and the difference in delay step size is within a preset range, it is confirmed that there is a unidirectional influence relationship between the data pairs to be analyzed at that time scale. In the representation of the data pairs, the direction of influence is marked with an arrow, such as "A→B" indicating that A influences B, and the time scale attribute, such as day level, is marked next to the arrow, thus forming a pair of influencing factors at that time scale.
[0100] S139. If the two judgment results are inconsistent, select the time series data of the first and second data of the third part under this time scale, repeat the lag processing and cross-correlation calculation steps again, and determine whether there is a one-way influence relationship based on the majority conclusion of the three calculation results.
[0101] Specifically, when the two judgment results are inconsistent, the third part of the data is selected for analysis according to the method in step S1137 to obtain the third judgment result. The influence direction that appears more frequently in the three judgment results is used to determine whether there is a one-way influence relationship between the data pairs to be analyzed. If a certain influence direction appears two or three times, the influence relationship is determined to exist.
[0102] Furthermore, the data in the dataset of potential influencing factors across all time scales are combined in pairs, and the above analysis process is repeated to identify pairs of influencing factors with unidirectional influence relationships at each time scale.
[0103] The dataset of potential influencing factors at each time scale is traversed, and each pair of data types is paired to form multiple data pairs to be analyzed. For each data pair, the analysis process from steps S1131 to S1139 is followed to determine whether a unidirectional influence relationship exists, and the influencing factor pairs with influence relationships, their direction of influence, and time scale attributes are marked.
[0104] In this embodiment of the application, after performing the above steps, further:
[0105] S140. Based on the pairs of influencing factors at different time scales, construct an initial causal relationship topology for each time scale. Each causal relationship edge corresponds to a pair of influencing factors with a one-way influence relationship, and the direction of influence and the initial influence intensity that changes with the time scale are marked.
[0106] Specifically, for each time scale, the influencing factors at that scale are treated as nodes in the topology graph, and the unidirectional influence relationships between pairs of influencing factors are treated as directed edges, i.e., causal relationship edges. In the topology graph, the direction of influence is indicated by the arrows. The initial influence strength is determined based on the maximum value of the cross-correlation coefficient sequence of the influencing factor pair, and this maximum value is used as the initial influence strength value and marked on the corresponding causal relationship edge. For example, at the daily time scale, the initial influence strength marked on the causal relationship edge for the influencing factor pair of daily average temperature data and daily industrial electricity consumption data is the maximum value of their cross-correlation coefficients.
[0107] S150, based on the initial causal relationship edge set of time scale, constructs an initial causal relationship topology with time scale dimension, takes influencing factors as topology nodes, and uses the causal relationship edges of time scale as connections between nodes, forming a topology structure containing nodes, edges and time scale identifiers.
[0108] Specifically, the initial causal relationship edges at different time scales are integrated into a unified topology. Each influencing factor is treated as an independent node, with its name containing data type and time scale information. Causal relationship edges at different time scales serve as directed connections between nodes, with each edge labeled with its corresponding time scale identifier. For example, causal relationship edges at the hourly time scale are labeled "hourly," and those at the daily time scale are labeled "daily," etc. This forms an initial causal relationship topology containing nodes, edges, and edge time scale identifiers.
[0109] S160, the nodes in the initial causal relationship topology are hierarchically divided. Nodes that have direct causal relationship edges with historical power load data at each time scale are divided into the direct influencing factor level, and nodes that have indirect relationship with historical power load data through other nodes are divided into the indirect influencing factor level. The time scale attributes of each node are labeled.
[0110] Specifically, in the initial causal relationship topology, nodes directly connected to historical power load data nodes at each time scale are identified. These nodes represent influencing factors that directly impact power load, and are categorized into the direct influencing factor level. Nodes not directly connected to historical power load data nodes but connected to nodes in the direct influencing factor level via one or more causal relationships are categorized into the indirect influencing factor level. A level identifier and time scale attribute, such as "Direct Influencing Factor - Hourly" or "Indirect Influencing Factor - Daily," are added to the attribute information of each node.
[0111] Step S170: Based on the hierarchical division results and time scale attributes, adjust the labeling information of the causal association edges in the initial causal association topology, supplement the hierarchical identifier, time scale identifier, and association path between levels, and form a preliminary causal association structure of power load influencing factors that includes the hierarchical relationship and time-varying characteristics of influencing factors.
[0112] Specifically, based on the hierarchical division of nodes, hierarchical identifiers are added to the causal relationship edges. For example, "direct → direct" indicates the influence relationship within the direct influencing factor hierarchy, and "indirect → direct" indicates the influence relationship between the indirect influencing factor hierarchy and the direct influencing factor hierarchy. Simultaneously, the accuracy of the time scale identifier on each edge is ensured. Furthermore, the association path information between different levels is added to the topology, i.e., the complete path from the indirect influencing factor node through the direct influencing factor node to the power load node, thus forming a preliminary causal relationship structure for power load influencing factors. This causal relationship structure reflects the hierarchical relationship and time scale characteristics of the influencing factors.
[0113] S180 performs redundancy removal on the causal relationship edges in the preliminary causal relationship structure of the factors affecting power load, deleting duplicate causal relationship edges and causal relationship edges whose influence intensity is lower than the preset intensity threshold at each time scale.
[0114] Specifically, all causal relationships in the preliminary causal relationship structure of factors affecting power load are examined. If two or more edges have the same start point, end point, direction of influence, and time scale, they are considered duplicate edges. One of these duplicate edges is retained, and the rest are deleted. A threshold for influence intensity is preset. This threshold is determined based on historical data analysis and expert experience. Causal relationships with influence intensity below this threshold at all time scales are considered to have a small impact on power load and are deleted to simplify the causal relationship structure.
[0115] S190, calculate the rate of change of influence intensity of the remaining causal association edges at different time scales, label the time-varying causal intensity parameters of each edge, and obtain the final causal association structure of power load influencing factors that includes the hierarchical relationship and time-varying characteristics of influencing factors.
[0116] Specifically, for each causal relationship edge remaining after redundancy removal, preliminary influence intensity data are collected at different time scales. The ratio of the change in influence intensity at adjacent time scales to the influence intensity at the previous time scale is calculated as the rate of change of influence intensity of that edge between the two time scales. The influence intensity value and the rate of change of influence intensity at each time scale are labeled as time-varying causal intensity parameters on the corresponding causal relationship edges. After the above processing, the final causal relationship structure of power load influencing factors is obtained. This causal relationship structure of power load influencing factors fully includes the hierarchical relationship of influencing factors, time scale attributes, and time-varying causal intensity parameters.
[0117] See also Figure 1 :
[0118] S200: Based on the time-varying characteristics of the causal relationship structure, the data in the running dataset is classified according to the time scale to obtain the classified running dataset.
[0119] In an optional embodiment of this application, based on the time-varying characteristics of the causal relationship structure, the running data set is divided into causal dimensions according to time scale matching, and the data is classified according to the time scales corresponding to the direct influencing factors and the indirect influencing factors to obtain the classified running data set.
[0120] In an optional embodiment of this application, the names of all influencing factors at the direct and indirect influencing factor levels at each time scale are extracted from the causal relationship structure, and arranged by time scale to form a scale-specific direct influencing factor list and a scale-specific indirect influencing factor list. Based on the influencing factor category and time scale attribute corresponding to each data point in the dataset, data belonging to the direct influencing factor category at the same time scale are grouped into the corresponding time scale direct influencing factor data subset, and data belonging to the indirect influencing factor category at the same time scale are grouped into the corresponding time scale indirect influencing factor data subset. The direct and indirect influencing factor data subsets at each time scale are processed to unify their data formats, and the format-unified direct and indirect influencing factor data subsets at each time scale are combined and sorted by time scale order to form a classified dataset containing time scale associations.
[0121] In this embodiment, based on the time-varying characteristics of the causal relationship structure of power load influencing factors, the power system operation-related data set is divided into causal dimensions by time scale matching. The data is classified according to the time scale corresponding to the direct influencing factors and the indirect influencing factors to obtain the classified power system operation-related data set.
[0122] Specifically, the time-varying characteristics of the causal relationship structure of power load influencing factors reflect the changes in the impact of these factors on power load at different time scales. Based on these characteristics, the time scale and influencing factor hierarchy of each data category are determined. For each data record in the power system operation-related data set, based on the position of its corresponding influencing factor in the causal relationship structure, it is determined whether it is a direct or indirect influencing factor, and its time scale is determined. The data are then classified according to time scale and influencing factor hierarchy, with data belonging to the same time scale and influencing factor hierarchy grouped together to form a classified power system operation-related data set.
[0123] Furthermore, in the embodiments of this application:
[0124] S210: Extract all names of direct influencing factors at each time scale from the causal relationship structure of factors affecting power load, and arrange them according to time scale to form a subscale direct influencing factor list.
[0125] Specifically, the causal relationship structure of factors affecting electricity load is traversed. At each time scale, nodes belonging to the direct influencing factor level are selected, and the names of the influencing factors represented by these nodes are extracted. The names of direct influencing factors at the same time scale are arranged in a certain order, such as from largest to smallest influencing intensity, to form a list of direct influencing factors for that time scale. The lists of direct influencing factors at different time scales are arranged in ascending order of time scale to obtain a multi-scale list of direct influencing factors.
[0126] S220 extracts the names of all influencing factors at each time scale from the causal relationship structure of factors affecting power load, and arranges them according to time scale to form a list of indirect influencing factors at different scales.
[0127] Specifically, following a method similar to step S210, the node names of the indirect influencing factor levels are extracted at each time scale, i.e., the names of the indirect influencing factors. These names are arranged in a certain order to form a list of indirect influencing factors for each time scale. Then, the lists are arranged in order of time scale to obtain a multi-scale list of indirect influencing factors.
[0128] S230: Traverse each data point in the power system operation-related data set, identify the category and time scale attribute of each data point's influencing factors, and determine whether it belongs to a direct or indirect influencing factor at any time scale.
[0129] Specifically, for each data record in the power system operation-related data set, its corresponding influencing factor category, such as temperature and humidity, is determined through data identification information and metadata description. Based on the data sampling interval or statistical period, its time scale attribute is determined, such as 15-minute or hourly. Then, by consulting the subscale direct influencing factor list and the subscale indirect influencing factor list, it is determined whether the influencing factor is a direct or indirect influencing factor at the corresponding time scale.
[0130] S240, data belonging to the same category of direct influencing factors at the same time scale are grouped into the corresponding subset of direct influencing factor data at the time scale. Each subset of direct influencing factor data at the time scale contains all data corresponding to each influencing factor in the list of direct influencing factors at that scale.
[0131] Specifically, based on the judgment result of step S230, data records belonging to the same category of direct influencing factors at the same time scale are stored together to form a subset of direct influencing factor data for that time scale. Within this subset, data is organized according to the name of the influencing factor, with each influencing factor name containing all data records for that factor at that time scale.
[0132] S250 categorizes data belonging to the same indirect influencing factor category at the same time scale into a subset of indirect influencing factor data for the corresponding time scale. Each subset of indirect influencing factor data for each time scale contains all data corresponding to each influencing factor in the indirect influencing factor list for that scale.
[0133] Similarly, data records belonging to the same category of indirect influencing factors at the same time scale are grouped into the corresponding subset of indirect influencing factor data at the same time scale, and the data in the subset are stored according to the name of the influencing factor.
[0134] S260 performs data format unification processing on the direct and indirect influencing factor data subsets for each time scale, combines the unified direct and indirect influencing factor data subsets for each time scale, and labels the influencing factor hierarchy information and time scale identifier corresponding to each direct and indirect influencing factor data subset.
[0135] Specifically, data format standardization includes standardizing the timestamp format, numerical units, and data storage format. Timestamps are standardized to a standard date and time format; numerical units are standardized according to the data type, such as temperature to degrees Celsius and electricity consumption to kilowatt-hours; and data storage is standardized to CSV format. After format standardization, the subsets of direct and indirect influencing factors for each time scale are stored in the same folder, with the folder name and file metadata indicating the hierarchy of influencing factors (direct or indirect) and the time scale identifier, such as "Daily Level_Direct Influencing Factor Data" and "Hourly Level_Indirect Influencing Factor Data".
[0136] S270: Sort all the data subsets of direct influencing factors and indirect influencing factors after combination according to time scale order to form a classified set of power system operation-related data containing time scale correlation.
[0137] Specifically, all the combined subsets of direct and indirect influencing factors are arranged in ascending order of time scale, such as 15-minute, hourly, daily, and weekly scales. Time scale correlation information is added to the arranged datasets, recording the correspondence between data subsets at different time scales, such as a daily data subset corresponding to multiple hourly data subsets. Through this process, a categorized power system operation-related dataset is formed. The data in this dataset is arranged in order of time scale, facilitating the input and processing of subsequent large-scale models.
[0138] See also Figure 1 :
[0139] S300 inputs the classified set of operational data and the time-varying causal strength parameters of the causal relationship structure into the pre-trained power load prediction model. The prediction parameters of each layer are adjusted through the causal gradient transfer mechanism inside the large model to obtain the initial power load prediction results.
[0140] In an optional embodiment of this application, the correlation gradient between the parameters of each layer of the large model and the causal feature vector is calculated through a causal gradient transfer mechanism to determine the adjustment direction of the parameters of each layer of the large model. Based on the correlation gradient and adjustment direction, the weight ratio of the direct influencing factors at each time scale is enhanced, and the direct influencing factor data is subjected to feature enhancement processing, while the indirect influencing factor data is subjected to feature filtering processing. The feature-enhanced direct influencing factor data and the feature-filtered indirect influencing factor data are then fused across scales to generate fused feature data containing time scale correlations. The fused feature data is used to perform load forecast calculations based on the large model to output preliminary power load forecast data for each time scale. The preliminary power load forecast data is then integrated into a time series and arranged according to a unified time axis to obtain the initial power load forecast results.
[0141] In an optional embodiment of this application, the time-varying causal strength parameters are converted into time-varying causal codes for large-scale model identification. These time-varying causal codes include hierarchical information of influencing factors at each time scale, the direction of influence of causal relationships, and time-varying causal strength data. Key causal information in the time-varying causal codes is extracted using a multi-scale causal feature parsing algorithm for the large-scale model, generating causal feature vectors for each time scale. Based on the causal feature vectors at each time scale, the correlation gradient between the parameters of each layer of the large-scale model and the causal feature vectors is calculated using a causal gradient transfer mechanism.
[0142] In an optional embodiment of this application, causal feature vectors at each time scale are input into a large model, the parameter gradient matrix of each layer of the large model is initialized, the current weight parameters of each layer of the large model are extracted, and the weight parameters of each layer of the large model are associated and mapped with the causal feature vectors of the corresponding time scale according to the time scale correspondence. The dot product of the weight parameters of each layer of the large model and the causal feature vector of the corresponding time scale is calculated to obtain the causal correlation value of each layer of the large model. Based on the causal correlation value, the partial derivatives of the parameters of each layer of the large model with respect to the causal feature vector are calculated to generate the local gradient vector of each layer of the power load prediction large model. The local gradient vector is passed in reverse according to the model hierarchy, and the gradient information of adjacent layers is fused to obtain the global correlation gradient of each layer of the large model. The direction of adjustment of the parameter of the layer is determined according to the positive or negative sign of the global correlation gradient, and the adjustment magnitude of the parameter of the layer is determined according to the absolute value of the global correlation gradient.
[0143] In this embodiment, a pre-trained large-scale power load prediction model is invoked, and the time-varying causal strength parameters of the causal relationship structure of the power load influencing factors are input into the classified power system operation-related data set and the causal relationship structure of the power load influencing factors. The prediction parameters of each layer are adjusted through the causal gradient transfer mechanism inside the large-scale power load prediction model to obtain the initial power load prediction result.
[0144] Specifically, the pre-trained large-scale power load forecasting model is deployed on a high-performance computing server. This model is built on a deep learning framework and features multiple hidden layers and a complex network structure. Through the API provided by the pre-trained model, the time-varying causal strength parameters of the classified power system operation-related dataset and the causal relationship structure of power load influencing factors are input into the model. Upon receiving the input data, the model activates its internal causal gradient propagation mechanism. This mechanism calculates the gradient contribution of each influencing factor to the prediction result based on the time-varying causal strength parameters, and then adjusts the weight and bias parameters of each layer of the model to better capture the causal relationship between influencing factors and power load. After parameter adjustment, the model processes the input data and outputs an initial power load forecast result, which includes power load forecast values at different time scales.
[0145] Furthermore, in the embodiments of this application:
[0146] S310: Input the subset of direct influencing factors at each time scale in the classified power system operation-related data set to the first data receiving port of the power load forecasting big model at the corresponding time scale, and input the subset of indirect influencing factors at each time scale to the second data receiving port of the power load forecasting big model at the corresponding time scale.
[0147] Specifically, the large-scale power load forecasting model has multiple data receiving ports, with two ports corresponding to each time scale: the first data receiving port and the second data receiving port. For the categorized power system operation-related data set, the subset of data on direct influencing factors at each time scale is transmitted over the network to the first data receiving port corresponding to the model's time scale, while the subset of data on indirect influencing factors is transmitted to the second data receiving port corresponding to the model's time scale. Encryption protocols are used during data transmission to ensure data security. Each data receiving port of the model has a data verification function, capable of checking the format and integrity of the input data. If data anomalies are detected, an error message is returned to the data sending end.
[0148] S320 converts the time-varying causal intensity parameters of the causal relationship structure of power load influencing factors into a time-varying causal code that can be recognized by the large model. The time-varying causal code contains the hierarchical information of influencing factors at each time scale, the influence direction of causal relationship edges, and time-varying causal intensity data.
[0149] Specifically, the time-varying causal strength parameters of the causal relationship structure of power load influencing factors are stored in an XML file. A dedicated encoding conversion program reads this XML file and parses the hierarchical information of influencing factors at each time scale, the influence direction of causal relationships, and the time-varying causal strength data. This information is then converted into a binary data stream that the model can recognize according to preset encoding rules—that is, time-varying causal encoding. During encoding, the hierarchical information of influencing factors is represented by specific binary bits, the influence direction of causal relationships is represented by a single binary digit (0 for positive influence, 1 for negative influence), and the time-varying causal strength data is converted into a two's complement form of a certain number of bits. After encoding, a file header containing an encoding version number and a checksum is generated, which, together with the encoded data, forms a complete time-varying causal encoded data stream.
[0150] S330 inputs the time-varying causal code into the causal information processing unit of the large-scale power load forecasting model. The key causal information in the code is extracted by the multi-scale causal feature parsing algorithm inside the causal information processing unit, and causal feature vectors at each time scale are generated.
[0151] Specifically, time-varying causal coding is input to the causal information processing unit of the large-scale power load forecasting model via a data bus. The multi-scale causal feature parsing algorithm within this unit first decodes the encoded data stream, separating the file header and the encoded data. Then, based on the version number in the file header, the corresponding parsing method is selected to extract the hierarchical information of influencing factors, the direction of influence of causal relationships, and the time-varying causal intensity data at each time scale from the encoded data. For the extracted information, feature processing is performed, converting the hierarchical information into the dimensions of a feature vector, with the direction of influence and the time-varying causal intensity data serving as the values for the corresponding dimensions. For each time scale, the processed feature information is combined into a vector, namely the causal feature vector, which characterizes the key features of the causal relationship structure at that time scale.
[0152] S340, based on the causal feature vectors at each time scale, activates the causal gradient transfer mechanism within the large-scale power load forecasting model, calculates the correlation gradient between the parameters of each layer of the large-scale power load forecasting model and the causal feature vector, and determines the adjustment direction of the parameters of each layer of the large-scale power load forecasting model.
[0153] Specifically, causal feature vectors at each time scale are input into the model's parameter tuning module, activating the causal gradient propagation mechanism. This mechanism first maps the causal feature vectors to the parameters of each layer of the model, determining which parameters are related to which causal features. Then, it calculates the partial derivatives of the model's prediction results with respect to the parameters of each layer, i.e., the gradients. It then combines these gradients with the causal feature vectors to calculate the degree of correlation between these gradients and the causal features, obtaining the correlation gradients. The sign of the correlation gradient determines the direction of parameter adjustment: a positive correlation gradient indicates that increasing the parameter value improves the model's ability to capture causal features; a negative correlation gradient requires decreasing the parameter value.
[0154] Furthermore, regarding step S340:
[0155] S341, input the causal feature vectors of each time scale into the gradient calculation unit of the power load forecasting model, and initialize the parameter gradient matrix of each layer of the power load forecasting model.
[0156] The causal feature vectors at each time scale are transmitted to the gradient calculation unit of the large-scale power load forecasting model through an internal data interface. The gradient calculation unit assigns a parameter gradient matrix to each layer of the model, with the same dimension as the parameter matrix of the corresponding layer. During initialization, all elements in the parameter gradient matrix are set to zero to prepare for subsequent gradient calculations.
[0157] S342, extract the current weight parameters of each layer of the large-scale power load forecasting model, and associate and map the weight parameters of each layer of the large-scale power load forecasting model with the causal feature vectors of the corresponding time scale according to the time scale correspondence.
[0158] The current weight parameters of each layer are read from the model's parameter storage area; these parameters are stored in matrix form. Based on the time scale correspondence, the weight parameters of the layers processing data at a certain time scale are associated with the causal feature vector of that time scale, establishing a correspondence table between parameters and causal features, clarifying which dimension of the causal feature vector each weight parameter is associated with.
[0159] S343, calculate the dot product of the weight parameters of each layer of the power load forecasting model with the causal feature vector of the corresponding time scale to obtain the causal correlation value of each layer of the power load forecasting model. The causal correlation value reflects the degree of correlation between the parameters of that layer and the causal features.
[0160] For each layer's weight parameter matrix, each element in the matrix is multiplied by the corresponding element in the causal feature vector. Then, all products are summed to obtain the dot product of the weight parameter matrix and the causal feature vector, which is the causal correlation value for that layer. The magnitude of the causal correlation value reflects the degree of linear correlation between the layer's weight parameters and the causal feature vector; a larger value indicates a stronger correlation.
[0161] S344, based on the causal correlation value, calculates the partial derivatives of the parameters of each layer of the large-scale power load forecasting model with respect to the causal feature vector, and generates the local gradient vector of each layer of the large-scale power load forecasting model.
[0162] Using the causal correlation value as the dependent variable and the weight parameters of each layer as independent variables, the partial derivatives of the causal correlation value with respect to each weight parameter are calculated using numerical differentiation. These partial derivatives are then arranged in order of their position within the matrix according to the weight parameters, forming the local gradient vector for that layer. Each element in the local gradient vector represents the change in the causal correlation value caused by a small change in the corresponding weight parameter.
[0163] S345, the local gradient vectors of each layer of the large power load forecasting model are passed in reverse order according to the model hierarchy. During the transmission process, the gradient information of adjacent layers is fused to obtain the global correlation gradient of each layer of the large power load forecasting model.
[0164] Starting from the output layer of the model, the local gradient vector of that layer is passed to the layer above. During this process, the local gradient vector of the current layer is weighted and summed with the gradient vector passed down from the previous layer. The weights are determined by the connection strength between the two layers; the stronger the connection, the greater the weight of the corresponding local gradient vector. This process is repeated until the gradient information is passed to the input layer of the model. After backpropagation and fusion, each layer obtains a globally correlated gradient that integrates the gradient information from all subsequent layers.
[0165] S346. Analyze the positive and negative signs of the global correlation gradients of each layer in the large-scale power load forecasting model. If the global correlation gradient is positive, it indicates that the parameters of that layer need to be increased to enhance the correlation with causal features; if the global correlation gradient is negative, it indicates that the parameters of that layer need to be decreased to optimize the correlation with causal features.
[0166] For each element in the global correlation gradient vector of each layer, determine the sign. If the element value is positive, it means that increasing the weight parameter value corresponding to that element can improve the correlation between that parameter and causal features; if the element value is negative, decreasing the weight parameter value can optimize the correlation and avoid the model's ability to capture causal features decreasing due to the parameter being too large or too small.
[0167] S347. Based on the absolute value of the global correlation gradient, determine the adjustment range of parameters for each layer of the large-scale power load forecasting model. The larger the absolute value, the larger the adjustment range.
[0168] The absolute value of the global correlation gradient reflects the necessity and urgency of parameter adjustment. A larger absolute value indicates that the parameter's correlation with causal features deviates further from the ideal state, requiring a more significant adjustment; a smaller absolute value results in a smaller adjustment. The specific adjustment range is determined by multiplying the absolute value of the global correlation gradient by a preset learning rate, the size of which is set based on model training experience.
[0169] S348 integrates the adjustment direction and adjustment range of each layer of the large-scale power load forecasting model to generate an adjustment scheme for the parameters of each layer of the large-scale power load forecasting model. The adjustment scheme includes the specific adjustment direction, adjustment range and time scale adaptation requirements of each layer parameter.
[0170] Based on the adjustment direction (increase or decrease) and magnitude of each layer's parameters, as well as the time scale of the data processed by that layer, an adjustment plan for each layer's parameters is formulated. The adjustment plan is presented in tabular form, with columns including layer number, parameter name, parameter location, adjustment direction, adjustment magnitude, and time scale adaptation requirements. The time scale adaptation requirements mean that the parameter adjustment magnitude should consider the changes in time-varying causality strength at the corresponding time scale. For example, at time scales with higher time-varying causality strength, the parameter adjustment magnitude can be appropriately increased.
[0171] S349 inputs the adjustment scheme into the parameter adjustment unit of the power load forecasting model.
[0172] The generated parameter adjustment plan is sent to the parameter adjustment unit of the power load forecasting model via an internal communication protocol. The parameter adjustment unit parses and verifies the received adjustment plan to ensure its integrity and correctness. If an error is found in the plan, the parameter adjustment unit will return an error message and request a resend; if the plan is correct, it will prepare to adjust the model parameters according to the plan.
[0173] In this embodiment, after performing the above steps:
[0174] S350 adjusts the weight parameters of the direct and indirect influencing factors in the large power load forecasting model according to the correlation gradient and adjustment direction, thereby enhancing the weight ratio of the direct influencing factors at each time scale. The weight ratio of the direct influencing factors also adapts to the time-varying causal strength as the time scale changes.
[0175] The parameter adjustment unit adjusts the weight parameters of each layer of the model according to the associated gradient and adjustment direction in the adjustment scheme. For weight parameters corresponding to direct influencing factors, their values are increased according to the adjustment magnitude to increase their weight proportion in the model. For weight parameters corresponding to indirect influencing factors, appropriate adjustments are made according to the adjustment scheme, which can be increased or decreased, but overall, the weight proportion of direct influencing factors is higher than that of indirect influencing factors. At the same time, based on the time-varying causal strength parameter, when the time-varying causal strength of the direct influencing factor increases at a certain time scale, the proportion of the weight parameter corresponding to the direct influencing factor at that time scale is further increased; when the strength decreases, the weight proportion is appropriately reduced, so that the weight proportion can dynamically adapt to the time-varying causal strength as the time scale changes.
[0176] S360 uses adjusted weight parameters to perform feature enhancement processing on the subset of direct influencing factor data input from the first data receiving port at each time scale, highlighting the key load influence features in the subset of direct influencing factor data that match the time-varying causal strength of that time scale.
[0177] The adjusted weight parameters are loaded into the model's feature processing layer. For each timescale's first data receiving port's direct influencing factor data subset, the feature processing layer performs weighted calculations on each feature in the data according to the weight parameters. Features with higher weight parameters are considered more important in the data, i.e., the features are strengthened. Specifically, each feature value in the data subset is multiplied by its corresponding weight parameter to obtain a weighted feature value. The weighted feature value highlights key load influence features that match the time-varying causal strength at that timescale, allowing the model to focus more on these features in subsequent processing.
[0178] S370 performs feature filtering on the subset of indirect influencing factor data input from the second data receiving port at each time scale, retaining indirect features in the subset of indirect influencing factor data that are related to direct influencing factors at the same time scale, and eliminating redundant features in the subset of indirect influencing factor data that are not related.
[0179] A subset of indirect influencing factor data is input into the model's feature filtering layer. The feature filtering layer filters features within the indirect influencing factor data subset based on the relationships between indirect and direct influencing factors recorded in the causal relationship structure of electricity load influencing factors. Features that are not associated with any direct influencing factor at the same time scale are deemed redundant and discarded; indirect features that are associated with direct influencing factors are retained. For example, if "humidity" in the indirect influencing factors is associated with "temperature" in the direct influencing factors, the "humidity" feature is retained; otherwise, it is discarded.
[0180] S380 performs cross-scale fusion processing on the subset of direct influencing factors data after feature enhancement and the subset of indirect influencing factors data after feature filtering at each time scale to generate fused feature data containing time scale correlation.
[0181] Cross-scale fusion processing is performed in the model's fusion layer. The fusion layer first aligns the subsets of direct and indirect influencing factors across each time scale along the time dimension, ensuring temporal correspondence between data from different time scales. Then, an attention-based fusion method is employed, calculating attention weights for data at different time scales based on the time-varying causal strength parameters. Time scales with higher time-varying causal strength receive higher attention weights. The data from each time scale are multiplied by their corresponding attention weights and concatenated into a multi-dimensional feature vector, i.e., the fused feature data. This fused feature data contains correlation information between data from different time scales, providing a more comprehensive reflection of the combined impact of influencing factors on power load.
[0182] S390 inputs the fused feature data into the prediction calculation unit of the large power load prediction model. The multi-layer calculation components inside the prediction calculation unit perform load prediction calculations in sequence according to the time scale, and output preliminary power load prediction data by time scale.
[0183] The fused feature data is input into a prediction calculation unit, which contains multiple calculation components arranged sequentially by time scale. Each component is responsible for load prediction calculation at one time scale. First, the fused feature data enters the calculation component corresponding to the smallest time scale (e.g., 15 minutes). This component performs convolution, pooling, and other operations on the input data to extract local features and outputs a preliminary load prediction value for that time scale through a fully connected layer. Then, this preliminary load prediction value, along with data from the fused feature data corresponding to larger time scales (e.g., hours), is input into the hourly calculation component for more macroscopic feature extraction and prediction calculation. This process continues until prediction calculations for all time scales are completed, outputting preliminary power load prediction data for each time scale.
[0184] Finally, the preliminary power load forecast data at different time scales are integrated into a time series, and the preliminary power load forecast data at each scale are arranged according to a unified time axis to obtain the initial power load forecast results.
[0185] Preliminary power load forecast data, distributed across different time scales, have varying time resolutions. For example, 15-minute data may have a forecast value every 15 minutes, while hourly data may have a forecast value every hour. Time series integration processing transforms this data onto a unified time axis, with the smallest time unit being 15 minutes. For forecast data with lower time resolution, such as hourly data, the forecast values are copied to each 15-minute time point within the corresponding hour; for data with higher time resolution, they are directly arranged in chronological order. After integration, a continuous time series is obtained, which is the initial power load forecast result. This initial power load forecast result contains power load forecast values at multiple 15-minute intervals starting from the forecast start time.
[0186] See also Figure 1 :
[0187] S400, based on the hierarchical relationship of the causal association structure, performs dimensional deviation causal tracing processing on the initial power load forecast results to obtain the hierarchical deviation contribution value. Based on the hierarchical deviation contribution value, the power load forecast is reprocessed with time scale matching to obtain the optimized power load forecast results.
[0188] In an optional embodiment of this application, the inter-layer parameter transfer weights and time-varying causal strength of the causal relationship structure of the large model are adjusted according to the hierarchical deviation contribution value, so as to re-perform the time-scale matching power load forecasting process.
[0189] In an optional embodiment of this application, the actual power load data corresponding to the initial power load forecast result is obtained, the deviation value between the initial power load forecast result and the actual power load data at each time scale is calculated, and abnormal deviation data is obtained; based on the hierarchical relationship of the causal relationship structure, according to the time scale to which the abnormal deviation data belongs, all causal relationship paths within the time interval corresponding to the abnormal deviation data at that time scale are extracted, the influencing factor data in each causal relationship path is extracted, and the actual value data of each level of influencing factors within the abnormal deviation time interval is obtained according to the hierarchical classification; the relative difference between the actual value data and the historical influencing factor value data of the same time scale is calculated, and abnormal influencing factors are obtained; based on the time-varying causal intensity change and the influence chain of the causal relationship path containing abnormal influencing factors, the contribution of each causal relationship path containing abnormal influencing factors to the prediction deviation is calculated, and the hierarchical deviation contribution value of causal relationship paths containing abnormal influencing factors at different levels is determined.
[0190] In the embodiments of this application:
[0191] S400A, based on the hierarchical relationship of the causal relationship structure of the factors affecting power load, performs multi-dimensional deviation causal tracing processing on the initial power load forecast results, locates the causal relationship path of different levels of influencing factors that lead to forecast deviations, and generates a deviation causal tracing report containing the contribution of hierarchical deviations.
[0192] Specifically, there is a certain deviation between the initial power load forecast and the actual power load data. This study analyzes the deviation from two dimensions: the hierarchy of direct and indirect influencing factors, utilizing the hierarchical relationship of the causal relationship structure of power load influencing factors. By comparing the forecast results with the actual data, the deviation values at different time scales are calculated. Then, based on the correlation paths in the causal relationship structure, the influencing factors leading to the deviation and their respective hierarchical levels are traced. The contribution of each influencing factor to the deviation is determined, i.e., the hierarchical deviation contribution. This information is compiled into a report, forming a deviation causal traceability report. This report clearly shows the paths and contributions of influencing factors at different levels to the forecast deviation.
[0193] In the embodiments of this application:
[0194] S400B adjusts the inter-level parameter transfer weights and time-varying causal strength of the causal relationship structure of power load influencing factors in the large-scale power load forecasting model based on the hierarchical deviation contribution of the deviation causal tracing report, and re-performs the power load forecasting process with time scale matching to obtain the optimized power load forecasting results.
[0195] Specifically, based on the hierarchical deviation contribution values of each level in the deviation causal tracing report, the model layers and causal association edges that require key adjustments are identified. For levels with high hierarchical deviation contribution values, the inter-layer parameter transfer weights of the corresponding model layers are adjusted to increase the influence of these layer parameters on the prediction results. Simultaneously, based on the time-varying causal intensity change data of key causal association paths, the time-varying causal intensity parameters of the corresponding causal association edges in the causal association structure of power load influencing factors are adjusted to more accurately reflect the actual impact. After adjustment, the data classification, model input, parameter adjustment, and prediction calculation are reprocessed according to steps S200 to S400 to obtain the optimized power load prediction results.
[0196] Specifically, in this embodiment of the application, for step S400A:
[0197] S410A acquires the actual power load data corresponding to the initial power load forecast results, splits the actual power load data by time scale, calculates the deviation value between the initial power load forecast results and the actual power load data at each time scale, and forms a deviation value sequence by time scale.
[0198] Actual power load data contemporaneous with the initial power load forecast is obtained from the real-time database of the power system. The time resolution of the actual data is consistent with the forecast, at the 15-minute level. The actual power load data is then broken down into different time scales, such as 15-minute, hourly, and daily data sequences. For each time scale, the difference between the forecast data and the broken-down actual data at the corresponding time scale in the initial power load forecast is calculated point-by-point to obtain the deviation value at each time point. These deviation values are then arranged in chronological order to form a sequence of deviation values by time scale.
[0199] S420A filters out abnormal deviation data from the deviation value sequence of each time scale that exceeds the preset deviation threshold of that time scale, and records the time interval, deviation value and time scale to which the abnormal deviation data belongs.
[0200] A preset deviation threshold is established for each time scale. This threshold is determined based on the accuracy requirements of power load forecasting and historical deviation data. For each time scale's deviation value sequence, each deviation value is checked to see if it exceeds the preset deviation threshold. If it does, it is determined to be abnormal deviation data, and the corresponding time interval (e.g., a 15-minute interval, an hour, etc.), the specific deviation value, and the time scale to which it belongs are recorded. For cases where multiple consecutive time points show abnormal deviations, they are merged into a single abnormal deviation time interval for recording.
[0201] S430A, based on the hierarchical relationship of the causal relationship structure of the factors affecting power load, extracts all causal relationship paths within the time interval of the abnormal deviation data according to the time scale to which the abnormal deviation data belongs. Each causal relationship path contains the complete influence chain from the influencing factor to the power load and its hierarchical affiliation.
[0202] Based on the time scale of the abnormal deviation data, all causal paths ending at the power load node are identified within the causal relationship structure of power load influencing factors at that time scale. Each path starts from one or more influencing factor nodes, passes through a series of intermediate nodes (which may be direct or indirect influencing factor nodes), and finally reaches the power load node, forming a complete influence chain. While extracting the paths, the hierarchical affiliation (direct or indirect influencing factor level) of each node in the path is recorded to facilitate subsequent analysis of the impact of different levels on the deviation.
[0203] S440A extracts the influencing factor data in each causal path and obtains the actual value data of each level of influencing factor within the abnormal deviation time interval by classifying them according to hierarchical affiliation.
[0204] For each extracted causal path, based on the names of the influencing factor nodes and the time interval of the abnormal deviation, the actual values of the corresponding influencing factors within that time interval are extracted from the power system operation-related data set. The extracted data is then categorized according to the hierarchical affiliation of the influencing factor nodes into two categories: actual values of direct influencing factors and actual values of indirect influencing factors. Each category of data is arranged in chronological order.
[0205] S450A calculates the relative difference between the actual values of influencing factors at each level and the historical values of influencing factors at the same time scale. The relative difference is the ratio of the absolute value of the difference between the actual value and the historical value to the historical value. Abnormal influencing factors with a relative difference higher than the preset relative difference threshold in each level are selected.
[0206] The system retrieves influencing factor values from the historical database that coincide with the time interval of the abnormal deviation; these are the historical values for the same period. For each influencing factor at each level, the absolute value of the difference between the actual value and the historical value at each time point is calculated. This absolute value is then divided by the historical value for the same period to obtain the relative difference. A preset relative difference threshold is established. Influencing factors with a relative difference higher than this threshold are identified as abnormal influencing factors, and their abnormal values are considered to be one of the causes of prediction deviation.
[0207] S460A tracks causal paths containing anomalous influencing factors, analyzes the time-varying causal intensity changes of influencing factors in causal paths containing anomalous influencing factors hierarchically, and determines whether causal paths containing anomalous influencing factors directly lead to power load forecasting deviations at the corresponding time scale.
[0208] For the identified anomalous influencing factors, all causal paths containing these factors are identified within the causal relationship structure. The time-varying causal intensity of each influencing factor within these paths is analyzed over the time interval of the anomalous deviation, noting whether the intensity significantly increases or decreases. If the time-varying causal intensity of the anomalous influencing factor in a path changes significantly within the anomalous interval, and the path directly reaches the power load node, then it is determined that the path may directly cause the power load prediction deviation at the corresponding time scale.
[0209] S470A calculates the contribution of each causal path containing abnormal influencing factors to the prediction bias based on the time-varying causal intensity changes and the influence chain of causal association paths containing abnormal influencing factors, and determines the hierarchical bias contribution value of causal association paths containing abnormal influencing factors at different levels.
[0210] For causal paths containing anomalous influencing factors that are identified as potentially causing prediction bias, the time-varying causal strength of each causal edge in the path is calculated within the anomalous interval. Based on the length of the influence chain, the strength changes of each edge are weighted and summed, with edges at the beginning of the chain having higher weights and edges at the end having lower weights, resulting in the overall strength change value of the path. The relative difference of the anomalous influencing factors is multiplied by the overall strength change value to obtain the initial bias contribution value. The initial contribution value is then weighted according to the weight coefficient of the path's level (the level coefficient for direct influencing factors is higher than that for indirect influencing factors) to obtain the level-weighted contribution value. The intermediate bias contribution value is then calculated by combining the proportion of the anomalous bias value corresponding to this path to the total bias value. If the path has cross-influences with other paths, the redundant contribution from the cross-influences is deducted to obtain the net bias contribution value, i.e., the bias contribution degree of this path. The bias contribution degrees of all causal paths containing anomalous influencing factors at the same level are summed to obtain the level-based bias contribution value.
[0211] Furthermore:
[0212] S471A: Extract the time-varying causal intensity change data of each causal link in each causal path containing abnormal influencing factors, calculate the intensity change amplitude of each causal link in the abnormal deviation time interval, and the intensity change amplitude is the absolute value of the difference between the intensity in the abnormal interval and the intensity in the normal interval.
[0213] For each causal link in a causal path containing anomalous influencing factors, extract the time-varying causal intensity data of that link within the anomalous deviation time interval from the causal correlation structure of power load influencing factors. Simultaneously, extract the time-varying causal intensity data of that link within the normal time interval (the same length as the anomalous interval) preceding the anomalous interval. Calculate the absolute value of the difference between the average intensity of the anomalous interval and the average intensity of the normal interval, using this as the intensity variation amplitude of the causal link.
[0214] S472A: Based on the length of the influence chain of the causal relationship path containing abnormal influencing factors, the intensity change amplitude of each causal relationship edge is weighted and summed according to the propagation order of the causal relationship path containing abnormal influencing factors. The weight of the causal relationship edge at the beginning of the causal relationship path containing abnormal influencing factors is higher than the weight of the causal relationship edge at the end of the causal relationship path containing abnormal influencing factors, thus obtaining the overall intensity change value of the causal relationship path containing abnormal influencing factors.
[0215] The length of the influence chain refers to the number of causal links in the path. Each causal link is assigned a weight, with the edge at the beginning of the path having a weight of 1.0. The weight of each subsequent edge is multiplied by a decay factor (less than 1, such as 0.8) to ensure that the weight of the initial edge is higher than that of the final edge. The magnitude of the change in strength of each edge is multiplied by its corresponding weight, and all products are summed to obtain the overall strength change value of the causal path containing the anomalous influencing factor.
[0216] S473A: Obtain the relative difference data of abnormal influencing factors in the causal relationship path containing abnormal influencing factors, calculate the relative difference magnitude of abnormal influencing factors, the relative difference magnitude is the ratio of the absolute value of the difference between the actual value and the historical value in the same period to the historical value in the same period, multiply the relative difference magnitude by the overall strength change value of the causal relationship path containing abnormal influencing factors, and obtain the initial deviation contribution value of the causal relationship path containing abnormal influencing factors.
[0217] From the relative difference calculated in step S450A, the relative difference data of the abnormal influencing factors in the causal relationship path containing abnormal influencing factors are extracted, and the average value of the relative difference within the abnormal deviation time interval is taken as the relative difference magnitude. This relative difference magnitude is multiplied by the overall intensity change value of the path, and the product is the initial deviation contribution value of the causal relationship path containing abnormal influencing factors.
[0218] S474A: Based on the level to which the causal path containing abnormal influencing factors belongs, set the level weight coefficient. The weight coefficient of the level of direct influencing factors is set higher than the weight coefficient of the level of indirect influencing factors. Multiply the initial deviation contribution value of the causal path containing abnormal influencing factors by the weight coefficient of the corresponding level.
[0219] Based on the hierarchical affiliation of the initial influencing factor node in the causal path containing abnormal influencing factors, hierarchical weight coefficients are assigned to the path. The path weight coefficient for the level of direct influencing factors is set to a higher value (e.g., 0.9), and the path weight coefficient for the level of indirect influencing factors is set to a lower value (e.g., 0.6). The initial deviation contribution value of the causal path containing abnormal influencing factors is multiplied by the weight coefficient of its respective level to obtain the hierarchically weighted deviation contribution value.
[0220] S475A: Calculate the ratio of the deviation value of the abnormal deviation data corresponding to the causal relationship path containing abnormal influencing factors to the total deviation value of all abnormal deviation data at the same time scale. Multiply the ratio value by the contribution value of the causal relationship path containing abnormal influencing factors after hierarchical weighting to obtain the intermediate deviation contribution value of the causal relationship path containing abnormal influencing factors.
[0221] For anomalous deviation data corresponding to causal paths containing anomalous influencing factors, calculate the sum of the absolute values of their deviations. Divide this sum by the sum of the absolute values of the deviations of all anomalous deviation data at the same time scale to obtain a proportion. Multiply this proportion by the hierarchically weighted deviation contribution value to obtain the intermediate deviation contribution value of the causal path containing anomalous influencing factors.
[0222] S476A analyzes whether there are cross-influences of other related paths in the causal relationship path containing abnormal influencing factors. If there are cross-influences, the redundant contribution caused by the cross-influences is deducted from the intermediate deviation contribution value of the causal relationship path containing abnormal influencing factors to obtain the net deviation contribution value of the causal relationship path containing abnormal influencing factors.
[0223] Check whether the causal path containing the anomalous influencing factor shares some nodes or edges with other causal paths, i.e., whether there is cross-influence. If cross-influence exists, calculate the redundant contribution value by analyzing the contribution of the cross portion to the deviation. Subtract the redundant contribution value from the intermediate deviation contribution value to obtain the net deviation contribution value of the causal path containing the anomalous influencing factor.
[0224] S477A compares the net deviation contribution value of the causal path containing abnormal influencing factors with the preset contribution benchmark value of the time scale. If the net deviation contribution value is higher than the benchmark value, the net deviation contribution value is retained as the final deviation contribution value of the causal path containing abnormal influencing factors; if it is lower than the benchmark value, the net deviation contribution value is adjusted to the set proportion of the benchmark value.
[0225] The preset contribution benchmark value is determined based on historical deviation contribution data at this time scale. If the net deviation contribution value of a causal path containing abnormal influencing factors is higher than the preset contribution benchmark value, it is used as the final deviation contribution value of that path; if it is lower than the benchmark value, the net deviation contribution value is multiplied by a set ratio (such as 0.5) to obtain the final deviation contribution value, so as to avoid the interference of an excessively small contribution value on the overall analysis.
[0226] S478A: The final deviation contribution value is classified and statistically analyzed according to the level to which the causal association path containing abnormal influencing factors belongs. The sum of the final deviation contribution values of all key causal association paths at each level is calculated to obtain the hierarchical deviation contribution value of causal association paths containing abnormal influencing factors at different levels.
[0227] The final deviation contribution values of all causal paths containing anomalous influencing factors are classified according to their respective levels. Path contribution values from the level of direct influencing factors are assigned to the direct level, and those from the level of indirect influencing factors are assigned to the indirect level. The contribution values of each level are summed to obtain the level deviation contribution value for that level.
[0228] S479A records the hierarchical deviation contribution value of each level and the number of key causal paths contained in that level.
[0229] The calculated hierarchical deviation contribution values of the direct and indirect influencing factor levels, as well as the number of key causal paths (paths whose final deviation contribution value is higher than the preset contribution benchmark value) contained in each level, are recorded as important basis for deviation analysis.
[0230] After performing the above steps:
[0231] S480A marks causal paths whose determined hierarchical deviation contribution values exceed the preset contribution threshold for that hierarchical level, identifies them as key causal paths that cause prediction deviations at the corresponding time scale, and marks the hierarchical level and time scale to which the key causal path belongs.
[0232] A contribution threshold is preset for each level, determined based on the level's importance and historical level deviation contribution data. Causal paths with level deviation contributions exceeding the preset threshold are marked as critical causal paths, with the level and time scale indicated in the markings. This allows for subsequent focused monitoring and adjustments to these critical paths.
[0233] S490A records detailed information on key causal pathways at various time scales, including the names of influencing factors, their direction of influence, changes in time-varying causal intensity, differential data of anomalous influencing factors, and hierarchical deviation contribution values in key causal pathways.
[0234] For each critical causal path at each time scale, detailed records are made of the names of all influencing factors in the path, the direction of influence of each causal link, the changes in time-varying causal strength within the abnormal interval, the relative difference data of abnormal influencing factors (such as relative difference degree, actual value and historical value of the same period), and the hierarchical deviation contribution value of the path, forming a detailed record table of critical paths.
[0235] Finally, the key causal relationship path information, abnormal deviation data information, influencing factor difference data information, and hierarchical deviation contribution values at each time scale are compiled into a structured report, forming a deviation causal tracing report containing hierarchical deviation contributions.
[0236] The key causal relationship path information, abnormal deviation data (time interval, deviation value), influencing factor difference data (relative difference, actual and historical values, etc.), and hierarchical deviation contribution values recorded at each time scale are organized according to a preset report template. The report includes a summary, introduction, deviation analysis at each time scale, details of key causal relationship paths, hierarchical deviation contribution analysis, and conclusions, forming a structured deviation causal tracing report.
[0237] Specifically, in this embodiment of the application, for step S400B:
[0238] S410B extracts the hierarchical deviation contribution values and corresponding key causal relationship path information at different levels under each time scale from the deviation causal tracing report, and determines the levels and time scales that need to be adjusted.
[0239] Read the deviation causal tracing report and extract the hierarchical deviation contribution values for the levels of direct and indirect influencing factors at each time scale, as well as the key causal relationship path information corresponding to each level. Compare the hierarchical deviation contribution values at different levels and time scales, and identify the levels and corresponding time scales with hierarchical deviation contribution values higher than the preset adjustment threshold as the objects requiring key adjustment. The preset adjustment threshold is determined based on the target for improving prediction accuracy.
[0240] S420B calculates the inter-layer parameter transfer weight adjustment coefficient of the corresponding power load forecasting large model layer based on the layer deviation contribution value of the layer that needs to be adjusted. The higher the layer deviation contribution value, the larger the adjustment coefficient.
[0241] For levels requiring focused adjustment, their level deviation contribution value is compared with the historical average level deviation contribution value for that level, and the growth rate of the deviation contribution value is calculated. Based on this growth rate, a base adjustment coefficient (e.g., 1.0) is multiplied to obtain the inter-level parameter transfer weight adjustment coefficient. The higher the growth rate of the level deviation contribution value, the larger the adjustment coefficient. The upper limit of the adjustment coefficient is set according to the model stability requirements.
[0242] S430B calls the inter-layer weight management unit of the power load forecasting model, extracts the inter-layer parameter transfer weights of each layer of the current power load forecasting model, multiplies the inter-layer parameter transfer weights of the power load forecasting model layers corresponding to the layers that need to be adjusted by the adjustment coefficient, and obtains the adjusted inter-layer parameter transfer weights.
[0243] The model calls the inter-layer weight management unit via its provided interface. This unit extracts the current inter-layer parameter transfer weight matrix for each layer from the model's parameter storage area. It then identifies the model layer corresponding to the layer requiring significant adjustment, multiplies each element of that layer's inter-layer parameter transfer weight matrix by the calculated adjustment coefficient, and obtains the adjusted inter-layer parameter transfer weight matrix.
[0244] S440B extracts the time-varying causal intensity change data of the causal edges of each key causal path from the key causal path information, and adjusts the time-varying causal intensity parameters of the corresponding causal edges in the causal relationship structure of power load influencing factors according to the intensity change data, so that the time-varying causal intensity parameters match the actual influence intensity.
[0245] The time-varying causal intensity variation data of each causal link on each critical path is obtained from the critical causal path information, i.e., the difference between the intensity in the abnormal interval and the intensity in the normal interval. For the causal link corresponding to the causal link structure of power load influencing factors, the value of its time-varying causal intensity parameter in the abnormal deviation time interval is adjusted to the average value of the intensity in the abnormal interval, so that the adjusted time-varying causal intensity parameter can more accurately reflect the influence intensity of the causal link under actual conditions.
[0246] S450B extracts the data characteristics of direct influencing factors at each time scale within the abnormal deviation time interval in the deviation causal tracing report, and determines the key data types that need to be strengthened for input at each time scale.
[0247] Analyze the data characteristics of direct influencing factors at each time scale within the abnormal deviation time interval in the deviation causal tracing report, such as the data fluctuation amplitude and trend. Identify data types that are highly correlated with abnormal deviations and have a significant impact on prediction results as key data types requiring enhanced input. For example, if temperature data fluctuates significantly within the abnormal interval and contributes significantly to the deviation, then temperature data is identified as a key data type.
[0248] S460B adjusts the input parameters of the large-scale power load forecasting model, increases the proportion of key data types in the input data at each time scale, and reduces the proportion of data corresponding to abnormal influencing factors.
[0249] Modify the model's input configuration file. For key data types at each time scale, increase their sampling frequency or data volume proportion in the input data. For example, increase temperature data that was originally sampled once per hour to be sampled once every half hour. For data types corresponding to abnormal influencing factors identified in the deviation causal tracing report, appropriately reduce their proportion in the input data, such as reducing the weight of this data in the input feature vector or reducing its sampling frequency.
[0250] S470B converts the adjusted causal relationship structure of power load influencing factors into a new time-varying causal code, which is then input into the causal information processing unit of the power load forecasting model.
[0251] Following the method in step S320, the time-varying causal intensity parameters of the adjusted causal relationship structure of the power load influencing factors are converted into new time-varying causal codes, which are then input into the causal information processing unit of the power load prediction large model through the data interface for the readjustment of model parameters.
[0252] The S480B inputs the adjusted set of classified power system operation data to the corresponding data receiving port of the power load forecasting model according to the time scale.
[0253] The classified power system operation-related data set after adjusting the input parameters in step S460B is then input into the data receiving port of the model corresponding to the time scale, following the method in step S310, with the direct influencing factor data subsets and indirect influencing factor data subsets for each time scale respectively.
[0254] S490B, through the causal gradient transfer mechanism within the large power load forecasting model, readjusts the weight parameters of each layer of the large power load forecasting model based on the new time-varying causal coding and the adjusted inter-layer parameter transfer weights.
[0255] After receiving the new time-varying causal code and input data, the model initiates the causal gradient transfer mechanism. Following the process in step S340, based on the new time-varying causal code and the adjusted inter-layer parameter transfer weights, it recalculates the correlation gradient and adjustment direction of the parameters of each layer, and then adjusts the weight parameters of each layer of the model.
[0256] Furthermore, the input data undergoes feature enhancement, filtering, and cross-scale fusion processing at different time scales to generate new fused feature data.
[0257] Following the methods in steps S360 to S380, the input data undergoes feature enhancement (for direct influencing factors), feature filtering (for indirect influencing factors), and cross-scale fusion processing to generate new fused feature data, which contains adjusted influencing factor feature information.
[0258] Furthermore, the new fused feature data is input into the prediction calculation unit of the large-scale power load prediction model. The prediction calculation unit performs re-prediction calculations in sequence according to the time scale to obtain secondary power load prediction data for different time scales.
[0259] The new fused feature data is input into the prediction calculation unit, and the prediction calculation is performed again according to the method in step S139 to obtain the secondary power load prediction data by time scale.
[0260] Furthermore, the new deviation values between the secondary power load forecast data and the actual power load data at each time scale are calculated. If the new deviation values at all time scales are lower than the corresponding preset deviation thresholds, the secondary power load forecast data at each time scale are integrated to obtain the optimized power load forecast results.
[0261] For each time scale, the secondary power load forecast data is compared with the actual power load data of the same period, and a new deviation value is calculated. If the new deviation values of all time scales are lower than the preset deviation threshold for that scale, the forecast result is considered to have been optimized. The secondary power load forecast data of each time scale are then integrated into a time series according to the method in step S300 to obtain the optimized power load forecast result.
[0262] Finally, if new deviation values at any time scale are still higher than the preset deviation threshold, repeat the above adjustment steps until new deviation values at all time scales are lower than the threshold, and obtain the final optimized power load forecast result.
[0263] If the new deviation values at some time scales are still higher than the preset deviation threshold, return to step S151, re-extract the deviation contribution value and critical path information, perform a new round of model parameter and causal relationship structure adjustment, and prediction calculation until the new deviation values at all time scales are lower than the preset threshold, and obtain the final optimized power load prediction result.
[0264] Combination Figure 2 As shown in the figure, this application provides a power load forecasting device 800 based on a large model, including a processor 801 and a memory 802. Optionally, the device may further include a communication interface 803 and a bus 804. The processor 801, communication interface 803, and memory 802 can communicate with each other via the bus 804. The communication interface 803 can be used for information transmission. The processor 801 can call logical instructions in the memory 802 to execute the power load forecasting method based on a large model described in the above embodiment.
[0265] Furthermore, the logic instructions in the aforementioned memory 802 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium.
[0266] The memory 802, as a computer-readable storage medium, can be used to store software programs and computer-executable programs, such as the program instructions / modules corresponding to the methods in the embodiments of this application. The processor 801 executes functional applications and data processing by running the program instructions / modules stored in the memory 802, thereby realizing the power load forecasting method based on a large model in the above embodiments.
[0267] The memory 802 may include a program storage area and a data storage area. The program storage area may store the operating system and application programs required for at least one function; the data storage area may store data created based on the use of the terminal device. Furthermore, the memory 802 may include high-speed random access memory and may also include non-volatile memory.
[0268] This application provides a system comprising: a system body and the aforementioned large-model-based power load forecasting device 800. The large-model-based power load forecasting device 800 is installed within the system body. The installation relationship described herein is not limited to placement within the system, but also includes installation connections with other components of the system, including but not limited to physical connections, electrical connections, or signal transmission connections. Those skilled in the art will understand that the large-model-based power load forecasting device 800 can be adapted to feasible system bodies to achieve other feasible embodiments.
[0269] This application provides a computer-readable storage medium storing computer-executable instructions configured to execute the above-described power load forecasting method based on a large model.
[0270] The technical solutions of this application embodiment can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes one or more instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in this application embodiment. The aforementioned storage medium can be a non-transitory storage medium, including: USB flash drive, portable hard drive, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, and other media capable of storing program code.
[0271] The technical solutions of this application embodiment can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes one or more instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in this application embodiment. The aforementioned storage medium can be a non-transitory storage medium, including: USB flash drive, portable hard drive, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, and other media capable of storing program code.
[0272] The foregoing description and accompanying drawings fully illustrate embodiments of this application to enable those skilled in the art to practice them. Other embodiments may include structural, logical, electrical, procedural, and other changes. The embodiments represent only possible variations. Individual components and functions are optional unless explicitly required, and the order of operation may vary. Parts and features of some embodiments may be included in or replace parts and features of other embodiments. Moreover, the terminology used in this application is for describing embodiments only and is not intended to limit the claims. As used in the description of embodiments and claims, the singular forms “a,” “an,” and “the” are intended to equally include the plural forms unless the context clearly indicates otherwise. Similarly, the term “and / or,” as used herein, means including one or more of the associated listed items and all possible combinations thereof. Additionally, when used in this application, the term "comprise" and its variations "comprises" and / or "comprising" refer to the presence of stated features, integrals, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof. Without further limitations, an element defined by the phrase "comprising a..." does not exclude the presence of other identical elements in the process, method, or apparatus that includes said element. In this document, each embodiment may focus on the differences from other embodiments, and similar or identical parts between embodiments can be referred to mutually. For methods, products, etc., of the embodiments claimed, if they correspond to the method section of the embodiments claimed, then the relevant parts can be referred to the description of the method section.
[0273] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments claimed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the embodiments of this application. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0274] The methods and products (including but not limited to devices and equipment) disclosed in the embodiments herein can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For instance, the division of units may be merely a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces, and the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to implement this embodiment according to actual needs. In addition, the functional units in the embodiments of this application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0275] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than that shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. In the descriptions corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks may also occur in a different order than disclosed in the description; sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. Each block in a block diagram and / or flowchart, and combinations of blocks in a block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
Claims
1. A power load forecasting method based on a large model, characterized in that, include: Based on the power system operation data set, the power load influencing factors at each time scale of the power system are obtained. By constructing an initial causal association topology containing time-varying causal strength, a causal association structure of power load influencing factors containing hierarchical relationships and time-varying characteristics of influencing factors is obtained. Based on the time-varying characteristics of the causal relationship structure, the data in the running data set are classified according to the time scale to obtain the classified running data set; The classified set of operational data and the time-varying causal strength parameters of the causal relationship structure are input into the pre-trained power load prediction model. The prediction parameters of each layer are adjusted through the causal gradient transfer mechanism inside the model to obtain the initial power load prediction result. Based on the hierarchical relationship of the causal association structure, the initial power load forecast result is subjected to dimensional deviation causal tracing processing to obtain the hierarchical deviation contribution value. Based on the hierarchical deviation contribution value, the power load forecast processing is re-processed with time scale matching to obtain the optimized power load forecast result.
2. The method according to claim 1, characterized in that, Based on the power system operation data set, the power load influencing factors at various time scales of the power system are obtained. By constructing an initial causal association topology containing time-varying causal strength, a causal association structure of power load influencing factors containing hierarchical relationships and time-varying characteristics is obtained, including: The power system operation data set includes historical power load data, power system equipment operation data, external environmental impact data, and user electricity consumption behavior data.
3. The method according to claim 1 or 2, characterized in that, Based on the power system operation data set, the power load influencing factors at various time scales of the power system are obtained. By constructing an initial causal association topology containing time-varying causal strength, a causal association structure of power load influencing factors containing hierarchical relationships and time-varying characteristics is obtained, including: A dynamic causal discovery algorithm is used to mine the power load influencing factors of the operational data set in order to separate and obtain the direct and indirect influencing factors of the power system at each time scale.
4. The method according to claim 3, characterized in that, Based on the power system operation data set, the power load influencing factors at various time scales of the power system are obtained. By constructing an initial causal association topology containing time-varying causal strength, a causal association structure of power load influencing factors containing hierarchical relationships and time-varying characteristics is obtained, including: Based on the multi-timescale time series characteristics and attribute characteristics of various types of data in the dataset, the correlation between various types of data and historical power load data is calculated according to different time scales. Data with a correlation higher than a preset correlation threshold at each time scale are selected to form a dataset of potential influencing factors by time scale. A pairwise interaction analysis is performed on the data in the dataset of potential influencing factors at each time scale to obtain pairs of influencing factors with unidirectional influence relationships at different time scales. An initial causal association edge set for each time scale is constructed based on the pairs of influencing factors, and an initial causal association topology is constructed based on the initial causal association edge set. The nodes in the initial causal relationship topology are divided into direct and indirect influencing factors. Based on the division results and the time scale attributes of the nodes, the causal relationship structure is obtained.
5. The method according to claim 4, characterized in that, Pairwise interaction analysis is performed on the data in the dataset of potential influencing factors at each time scale to obtain pairs of influencing factors with unidirectional influence relationships at different time scales. Based on these pairs of influencing factors, an initial causal relationship edge set for each time scale is constructed. Based on this initial causal relationship edge set, an initial causal relationship topology is constructed, including: For the data in the dataset of potential influencing factors, a time delay variable is set to determine whether there is a unidirectional influence relationship between the data that changes with the time scale, so as to obtain the pairs of influencing factors; In the initial causal relationship edge set, each causal relationship edge corresponds to a pair of influencing factors with a one-way influence relationship, and the direction of influence and the initial influence intensity changing with the time scale are marked to construct the initial causal relationship topology. The initial causal relationship topology uses influencing factors as topological nodes and causal relationship edges of different time scales as connections between nodes.
6. The method according to claim 5, characterized in that, The nodes in the initial causal relationship topology are hierarchically divided into direct and indirect influencing factors. Based on the hierarchical division results and the node time scale attributes, the causal relationship structure is obtained, including: Redundancy removal is performed on the causal relationship edges in the causal relationship structure, deleting duplicate causal relationship edges and causal relationship edges whose influence intensity is lower than a preset intensity threshold at each time scale. Calculate the rate of change of influence intensity of the remaining causal association edges at different time scales, label the time-varying causal intensity parameters of each edge, and obtain the final causal association structure of power load influencing factors that includes the hierarchical relationship and time-varying characteristics of influencing factors.
7. The method according to claim 6, characterized in that, Based on the time-varying characteristics of the causal relationship structure, the data in the operational data set are classified according to time scale to obtain the classified operational data set, including: Based on the time-varying characteristics of the causal relationship structure, the operational data set is divided into causal dimensions according to time scale matching. The data is classified according to the time scales corresponding to the direct and indirect influencing factors to obtain the classified operational data set.
8. The method according to claim 7, characterized in that, Based on the time-varying characteristics of the causal relationship structure, the data in the operational data set are classified according to time scale to obtain the classified operational data set, including: Extract all the names of the influencing factors at each time scale from the causal relationship structure, including the direct influencing factor level and the indirect influencing factor level, and arrange them according to the time scale to form a subscale direct influencing factor list and a subscale indirect influencing factor list. Based on the category of influencing factors and time scale attribute corresponding to each data point in the dataset, data belonging to the category of direct influencing factors under the same time scale are grouped into the direct influencing factor data subset of the corresponding time scale, and data belonging to the category of indirect influencing factors under the same time scale are grouped into the indirect influencing factor data subset of the corresponding time scale. The data formats of the direct and indirect influencing factor subsets at each time scale are standardized. The standardized direct and indirect influencing factor subsets at each time scale are then combined and sorted according to the time scale order to form the classified data set containing time scale associations.
9. The method according to claim 8, characterized in that, The classified operational data set and the time-varying causal strength parameters of the causal relationship structure are input into a pre-trained large-scale power load prediction model. The prediction parameters of each layer are adjusted through the causal gradient transfer mechanism within the large-scale model to obtain initial power load prediction results, including: Through the causal gradient transfer mechanism, the correlation gradient between the parameters of each layer of the large model and the causal feature vector is calculated, the adjustment direction of the parameters of each layer of the large model is determined, and the weight ratio of the direct influencing factors at each time scale is enhanced according to the correlation gradient and adjustment direction. Feature enhancement processing is performed on the data of the direct influencing factors, and feature filtering processing is performed on the data of the indirect influencing factors. The data of direct influencing factors after feature enhancement and the data of indirect influencing factors after feature filtering are fused across scales to generate fused feature data containing time scale correlation. The fused feature data is then used to perform load forecasting calculations based on the large model to output preliminary power load forecasting data by time scale. The preliminary power load forecast data is integrated into a time series, and the preliminary power load forecast data at each scale are arranged according to a unified time axis to obtain the initial power load forecast results.
10. The method according to claim 9, characterized in that, Through the aforementioned causal gradient transfer mechanism, the correlation gradient between the parameters of each layer of the large model and the causal feature vector is calculated to determine the adjustment direction of the parameters of each layer of the large model, including: The time-varying causal strength parameters are converted into time-varying causal codes for the identification of the large model. The time-varying causal codes include the hierarchical information of influencing factors at each time scale, the influence direction of causal association edges, and time-varying causal strength data. The key causal information in the time-varying causal coding is extracted by the multi-scale causal feature parsing algorithm of the large model, and causal feature vectors at each time scale are generated. Based on the causal feature vectors at each time scale, the correlation gradient between the parameters of each layer of the large model and the causal feature vectors is calculated through the causal gradient transfer mechanism.
11. The method according to claim 10, characterized in that, Through the aforementioned causal gradient transfer mechanism, the correlation gradient between the parameters of each layer of the large model and the causal feature vector is calculated to determine the adjustment direction of the parameters of each layer of the large model, including: The causal feature vectors at each time scale are input into the large model, the parameter gradient matrix of each layer of the large model is initialized, the current weight parameters of each layer of the large model are extracted, and the weight parameters of each layer of the large model are associated and mapped with the causal feature vectors of the corresponding time scale according to the time scale correspondence. Calculate the dot product of the weight parameters of each layer of the large model with the causal feature vector of the corresponding time scale to obtain the causal correlation value of each layer of the large model. Based on the causal correlation value, calculate the partial derivative of the parameters of each layer of the large model with respect to the causal feature vector to generate the local gradient vector of each layer of the power load prediction large model. The local gradient vectors are passed in reverse order according to the model hierarchy, and the gradient information of adjacent layers is fused to obtain the global associated gradients of each layer of the large model. The direction of adjustment of the parameters of the layer is determined according to the positive or negative sign of the global associated gradients, and the adjustment magnitude of the parameters of the layer is determined according to the absolute value of the global associated gradients.
12. The method according to claim 11, characterized in that, Based on the hierarchical relationship of the causal association structure, the initial power load forecast results are subjected to dimensional deviation causal tracing processing to obtain hierarchical deviation contribution values. Based on these hierarchical deviation contribution values, a new time-scale matched power load forecast is performed to obtain optimized power load forecast results, including: Based on the hierarchical deviation contribution value, the inter-level parameter transfer weights of the large model and the time-varying causal strength of the causal relationship structure are adjusted to re-perform time-scale matching power load forecasting.
13. The method according to claim 12, characterized in that, Based on the hierarchical relationship of the causal association structure, the initial power load forecast results are subjected to dimensional deviation causal tracing processing to obtain hierarchical deviation contribution values. Based on these hierarchical deviation contribution values, a new time-scale matched power load forecast is performed to obtain optimized power load forecast results, including: Obtain the actual power load data corresponding to the initial power load forecast result, calculate the deviation value between the initial power load forecast result and the actual power load data at each time scale, and obtain the abnormal deviation data; Based on the hierarchical relationship of the causal association structure, according to the time scale to which the abnormal deviation data belongs, all causal association paths within the time interval corresponding to the abnormal deviation data at that time scale are extracted, and the influencing factor data in each causal association path are extracted. The actual value data of each level of influencing factors within the time interval of the abnormal deviation are obtained by classifying them according to their hierarchical affiliation. Calculate the relative difference between the actual value data and the historical influencing factor values at the same time scale, and obtain the abnormal influencing factors; Based on the time-varying causal intensity changes and the influence chains of causal association paths containing the aforementioned anomalous influencing factors, the contribution of each causal association path containing anomalous influencing factors to the prediction bias is calculated, and the hierarchical bias contribution value of causal association paths containing anomalous influencing factors at different levels is determined.
14. A power load forecasting device based on a large model, comprising a processor and a memory storing program instructions, characterized in that, The processor is configured to execute, when running the program instructions, the power load forecasting method based on a large model as described in any one of claims 1 to 13.
15. A system, characterized in that, include: System body; as well as, The power load forecasting device based on a large model as described in claim 14 is installed on the system body.
16. A computer-readable storage medium storing program instructions, characterized in that, When the program instructions are executed, they cause the computer to perform the power load forecasting method based on a large model as described in any one of claims 1 to 13.