Power system source load prediction-oriented pre-training model construction method

By designing a pretrained model framework that includes expert sub-models of space-time mode and forecast information, the problem of difficulty in building a pretrained model suitable for multi-space-time scale prediction tasks in the prior art is solved, and the source load prediction effect with high precision and high generality is achieved.

CN120012957AActive Publication Date: 2025-05-16ZHEJIANG UNIV

Patent Information

Application Number
CN202510488795.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-05-16
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

It is difficult for the existing technology to build a pre-trained model with strong learning ability, high prediction accuracy, strong versatility, and suitable for multi-time and spatial scale prediction tasks to improve the predictive perception of the operating risk situation of the power system distribution network.

Method used

A pre-trained model framework is designed, including a space-time and space-time model mixed expert submodel and a forecast information mixed expert submodel. It adopts a self-attention mechanism and a heterogeneous hybrid expert submodel, which can learn general knowledge from massive time series data and adapt to different prediction scenarios.

Benefits of technology

It achieves stronger generalization performance and broader downstream application scenarios, which can effectively improve the accuracy and versatility of source load prediction and adapt to the needs of different application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012957A_ABST
    Figure CN120012957A_ABST
Patent Text Reader

Abstract

The invention discloses a pre-training model construction method for power system source load prediction, and belongs to the field of power system time sequence prediction. According to the pre-training model, a heterogeneous combination architecture of a space-time mode mixed expert sub-model and a forecast information mixed expert sub-model is constructed for the problem of source-load double-side uncertainty caused by new energy access and novel load development, and the pre-training model can adapt to rich downstream prediction scenes and meet the prediction requirements of multi-time scale and multivariable input. According to the method, the general spatio-temporal features are extracted from mass data through the pre-training model, the precision and generalization performance of source load prediction are effectively improved, the multi-scene requirements of planning, scheduling and the like are adapted, and reliable technical support is provided for operation risk situation awareness of the power distribution network containing high-proportion distributed resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for constructing a pre-training model for power system source-load prediction. The method is suitable for constructing a pre-training model for a multi-temporal and spatial scale prediction task of power system source-load time series data, and belongs to the field of power system time series prediction. Background Art

[0002] The source-load forecasting of power systems is the core basic technology supporting the safe, economical and low-carbon operation of modern power systems, and can provide a strong guarantee for the accuracy of dispatching decisions and the stability of operation control. After a large number of new energy sources are connected to the power grid, the uncertainty faced by the power supply side of the system has increased significantly. With the increase in the ownership rate of electric vehicles and the development of new loads such as household energy storage, the uncertainty on the load side of the system is also increasing.

[0003] With the massive data and high-dimensional features, the difficulty of achieving accurate source-load prediction increases, and higher requirements are placed on the parameter scale and performance of the prediction model. In addition, the requirements for source-load prediction in different application scenarios such as planning, scheduling, optimization and control are different, and the prediction tasks in different time and space are quite different. Traditional time series prediction methods usually choose to build dedicated small-parameter scale prediction models for different prediction tasks and scenarios, which is time-consuming and labor-intensive and has poor performance in terms of prediction versatility and generalization. In the context of a large number of distributed resources connected to the power system, it is urgent to build a pre-trained model with strong learning ability, high prediction accuracy, strong versatility, and suitable for multi-time and space scale prediction tasks to improve the prediction and perception capabilities of the distribution network operation risk situation. Summary of the invention

[0004] In order to solve the problems summarized in the background technology, the present invention proposes a method for constructing a pre-training model for power system source and load prediction, and designs a pre-training model framework suitable for power system source and load time series data. The pre-training model structurally supports the input and prediction output of single variables and multivariate variables; it can perform prediction tasks of multiple time scales such as ultra-short-term, short-term, and medium- and long-term. During the pre-training process, based on the self-attention mechanism, the prediction model composed of heterogeneous hybrid expert sub-models can learn general knowledge from massive time series data, and has stronger generalization performance and broader downstream application scenarios.

[0005] In the field of power systems, source-load data specifically includes photovoltaic power, wind power, and load data of different voltage levels. The classification of historical data and forecast data and the data input of the pre-trained model in different prediction scenarios are summarized in the following table:

[0006]

[0007] A method for constructing a pre-trained model for power system source and load prediction is as follows:

[0008] From the framework, the pre-trained model contains two heterogeneous hybrid expert sub-models: the spatiotemporal pattern hybrid expert sub-model and the forecast information hybrid expert sub-model, which are used to process historical data and forecast data in the forecast task respectively, and can adapt to complex downstream data scenarios. When designing the sub-model, the different forecasting scenarios of single variable and multi-variable input should be fully considered to adapt to different forecasting tasks. The construction of the pre-trained model specifically includes the following steps:

[0009] Step 1: Construct the spatiotemporal pattern hybrid expert sub-model and the forecast information hybrid expert sub-model respectively.

[0010] The spatiotemporal mode hybrid expert sub-model uses the Transformer decoder as the model backbone structure and adopts the causal mask attention mechanism to learn the correlation between the input time series data fragments; the characteristics of the spatiotemporal mode hybrid expert sub-model are reflected in the output structure of the model. This part adopts a multi-degree-of-freedom hybrid expert Student T distribution output head, and each output head uses an independent parameter mapping network. The multi-degree-of-freedom specifically includes high, medium, and low degrees of freedom. The high-degree-of-freedom Student T distribution output head is used to fit the distribution characteristics of load-type stable data; the medium-degree-of-freedom Student T distribution output head is used to fit the scenario with periodic fluctuations and a lighter distribution tail; the low-degree-of-freedom Student T distribution output head is used to fit the thick-tailed characteristics with frequent extreme values. The Student T distribution is selected as the distribution type. This type of distribution is more stable in the scenario where there are abnormal points or offset points in the time series data, and is more suitable for pre-training scenarios with massive data input. The adaptability of the model in pre-training is improved by introducing distribution output heads with different degrees of freedom.

[0011] The forecast information hybrid expert sub-model uses a single feedforward neural network as the basic expert unit and is composed of multiple (e.g., 54) expert models. It can be further divided into three independent expert model groups: photovoltaic, wind power, and load, each of which processes the forecast information in its respective field. The structural design of the load, photovoltaic, and wind power expert model groups introduces many basic forecast units. In pre-training, the complex meteorological-power mapping relationship in different scenarios can be learned, so that it can also show strong generalization performance in cross-task and cross-scenario migration. The forecast information hybrid expert sub-model adopts a two-stage routing network design to achieve model guidance from task type to forecast fine-grained features, breaking through the limitations of the traditional single forecast model and better meeting the forecast requirements of multi-level, multi-factor, and complex working conditions of the power system. In the process of executing prediction and training, the forecast information input to the model is diverted to the expert model groups in different fields through the main routing network in the first stage. In the second stage, the sub-routing network in the current expert model group further selects the optimal expert model to perform prediction, and the final prediction result is output in parameterized Gaussian form.

[0012] Step 2: Create a pre-training data set. The pre-training data set used for pre-training model training consists of two parts: one is an open source time series data set available on the Internet (specific sources include energy, Internet of Things, health, network, transportation and environment); the other is real photovoltaic, wind power and load data of different voltage levels in the power system. Among them, most open source data sets do not contain forecast information, and the information that can be provided in the forecast task is mainly historical data; the real operating data in the power system includes real data of photovoltaic, wind power and load, as well as supporting meteorological forecast information (including wind speed, irradiance, temperature and humidity data). In the construction of the pre-training data set, historical data is extracted from the open source data and the power system data to construct a pre-training data set for the spatiotemporal pattern hybrid expert sub-model, and forecast data is extracted to construct a pre-training data set for the forecast information hybrid expert sub-model.

[0013] Step 3: Based on the pre-training dataset constructed in step 2, pre-train the two sub-models separately, using the negative log-likelihood loss function. During the pre-training process, verify the prediction effect of the pre-training model until it reaches the optimal value in the validation set. After the training is completed, save the parameters of the two sub-models to obtain the pre-training model. Using datasets from different fields during the pre-training of the sub-models can learn different spatiotemporal features.

[0014] Furthermore, the final pre-trained model structure and fine-tuning scheme for deployment can be determined based on the hardware resources, forecasting requirements, and available data of the downstream deployment application scenario. Among them, the data that the downstream application scenario can provide for model prediction is the key to determining the pre-trained model structure and fine-tuning scheme.

[0015] The specific scheme is summarized in the following table. In the table: STP represents the spatiotemporal mode hybrid expert sub-model; MIM represents the forecast information hybrid expert sub-model; when the application scenario can only provide historical data for model prediction (for example, small-capacity distributed photovoltaic operators, who cannot afford the high price of weather forecast services), the spatiotemporal mode hybrid expert sub-model is selected, the pre-trained weights of the sub-model are loaded, and the model is fine-tuned using the target scenario historical data; when the application scenario has no available historical data or less historical data (for example, a newly commissioned wind farm station, with little available historical wind power data but meteorological forecast services purchased), the forecast information hybrid expert sub-model is selected, the sub-model parameters are loaded, and the model is fine-tuned using the target scenario forecast data; when both historical data and forecast data are available and sufficient, the combined structure of the spatiotemporal mode hybrid expert sub-model and the forecast information hybrid expert sub-model is selected to achieve the highest prediction accuracy, and the integrated fine-tuning technology of the pre-trained model is used to complete the scale alignment and complementary advantages of the two sub-models.

[0016]

[0017] In terms of hardware resources and prediction timeliness, the network depth of the model can be adjusted in combination with the computing power and memory limitations of the server deployed in the downstream application scenarios. The solution of partially loading pre-trained model parameters can be used to reduce model complexity and computing power requirements and improve prediction speed.

[0018] Based on the deployment scheme and model fine-tuning selected above, the model is tested using real data from application scenarios to ensure that it meets the acceptance requirements in terms of running speed and prediction accuracy. The model is deployed offline to the server and the time series prediction service is launched.

[0019] The present invention also provides an electronic device, comprising: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the pre-training model construction method for power system source and load prediction.

[0020] The present invention also provides a computer-readable storage medium having computer instructions stored thereon, wherein the computer instructions are used to enable a computer to execute the method for constructing a pre-training model for power system source and load prediction.

[0021] The beneficial effects of the present invention are:

[0022] The present invention proposes a method for constructing a pre-trained model for power system source and load prediction. The pre-trained model proposed in the present invention is composed of heterogeneous hybrid expert sub-models. While maintaining the simplicity of the sub-model structure and the uniqueness of the function, it can also realize parallel training, and effectively speed up the training iteration cycle of the pre-trained model while ensuring the model prediction accuracy. In addition, the present invention also proposes a complete pre-trained model application scheme, which fully considers the complex data scenarios, prediction requirements and computing power limitations in the actual downstream application scenarios, and is compatible with a variety of downstream application scenarios by selecting different model structures and fine-tuning schemes, and has strong versatility. The pre-trained model of the present application supports the input and prediction output of single variables and multivariate variables, and can perform prediction tasks of multiple time scales such as ultra-short-term, short-term and medium- and long-term. The present invention extracts universal spatiotemporal features from massive data through a pre-trained model, effectively improves the accuracy and generalization performance of source and load prediction, adapts to the needs of multiple scenarios such as planning and scheduling, and provides reliable technical support for the awareness of the operation risk situation of distribution networks containing a high proportion of distributed resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 The process of building a pre-trained model suitable for multi-temporal and spatial scale prediction of power system source and load data;

[0024] Figure 2 shows the structure of the spatiotemporal mode hybrid expert submodel;

[0025] Figure 3 is a schematic diagram of the attention score causal mask

[0026] Figure 4 shows the forecast information hybrid expert sub-model structure; DETAILED DESCRIPTION

[0027] The present invention is further described below with reference to the accompanying drawings and implementation examples.

[0028] Example 1 provides a method for constructing a pre-trained model for power system source and load prediction.

[0029] The construction and deployment process of the pre-trained model is as follows Figure 1 As shown, the following steps are included:

[0030] Step 1: Construct the forecast information hybrid expert sub-model and the spatiotemporal pattern hybrid expert sub-model. This includes the following two parallel steps:

[0031] (1) Construct a spatiotemporal hybrid expert sub-model. The specific structure of the model is as follows: Figure 2 As shown. Assume that the input historical data is , B represents the number of input batches; S represents the length of the input historical sequence; C represents the number of input channels. The input data is first sliced ​​and divided into N time series segments of the same length P along the S dimension. At this time, the input data becomes the sliced ​​historical data :

[0032]

[0033] The dimension N should satisfy: Each data segment after segmentation is labeled to form Tag information ,in Used to mark which input channel the current sequence segment belongs to; Used to mark the input order of the current clip in the corresponding channel. Perform dimension changes and merge N and C dimensions. The dimension changes to The input embedding layer is mapped to the Transformer decoder (the dimension of the backbone network is D):

[0034]

[0035] In the formula is the mapping weight of the linear transformation of the embedding layer, is the bias coefficient. The feature mapping results It then inputs the backbone network of the multi-layer decoder Transformer architecture.

[0036] Input Transformer architecture After three sets of independent linear mappings, we can get , and Matrix, the size is . Split the matrix into NH heads along dimension D, where NH is the number of heads in the preset multi-head attention mechanism. The splitting process satisfies , where is the dimension of the attention head. , and The matrices are , and Indicates that its dimensions become . Use the split matrix to perform attention calculation:

[0037]

[0038] is the calculated attention score. To ensure the causality of the calculation process and adapt to multi-channel data input, the attention score also needs to perform the operation of adding bias and mask according to the label information. Attention scores, each row and column has a corresponding number Label information of . Select a data point in the attention score , let the label information of the row be , the label information of its column is The logic of adding bias is:

[0039]

[0040] In the formula is the data point after adding the bias, is a conditional item, and the value is 1 when the condition in the brackets is met, and 0 when it is not met. The logic of this formula is: when the channel dimensions of the row and column label information of a data point are equal, it means that the point comes from the attention score calculation results of two segments of the same channel, and add When the row and column label information of a data point is different, it means that the point comes from the attention score calculation results of two different channel fragments, and add Through such processing, the attention mechanism can distinguish and learn different pattern features in the time dimension (i.e., the same channel input) and the spatial dimension (i.e., different channel input).

[0041] The logic for adding the mask is:

[0042]

[0043] In the formula is the data point after the mask operation. The logic of this formula is: when the time series dimensions of the row and column label information of the data point are not equal and the input sequence number of the column label information is greater than the input sequence number of the row channel, it means that the point is the attention score calculated by the future time series segment and the historical time series segment, which violates the causality (that is, the time series segment can only calculate the attention with the historical sequence with an input sequence number less than itself, but cannot calculate the attention with the future segment with an input sequence number greater than itself), so it needs to be assigned a value of 0. Otherwise, it is a reasonable result and the attention score is retained. Figure 3 As shown in the figure, the process of calculating the attention scores of the fragment (1,3) and all historical fragments is marked with solid arrows, which conforms to causality, so no mask is required; while the attention scores calculated by the fragment (1,3) and future fragments are marked with dotted lines, which do not conform to causality and need to be masked to zero.

[0044] Traversal For each element in , add the bias and mask operations according to the above process. Then, Perform normalization and introduce The weighted summation gives the attention output:

[0045]

[0046] In the formula is the attention output. The multi-head output results are merged and the size becomes Attention Output After residual connection and layer normalization, it is input into the feedforward neural network (FFN):

[0047]

[0048]

[0049] in, Output for attention The result after residual connection and layer normalization; is the output of the feedforward neural network; is the feature input to the decoder by the embedding layer, used for and attention output Residual connection, that is ;GELU is the activation function; and Represents the weight coefficient of the linear mapping inside and outside the FFN network, and Represents the bias coefficient of the inner and outer linear mapping of the FFN network.

[0050] The output of the feedforward neural network is again processed by residual connection and layer normalization before output.

[0051] like Figure 2 As shown in the figure, the decoder structure adopts the same model structure with multiple layers stacked, and returns the output features of the final output backbone network after being processed by N layers of networks. :

[0052]

[0053] The output features of the backbone network Input the multi-degree-of-freedom hybrid expert student T distribution output head, and use the routing network to calculate the weight g of each student T distribution output head. The process is expressed in the form of a formula as follows:

[0054]

[0055] In the formula, represents the output head weight of the Student-T distribution calculated by the routing network, is the routing network weight matrix, is the routing bias coefficient, is the smoothing coefficient, which is set to 1 by default. The larger the smoothing coefficient, the smaller the calculated weight difference.

[0056] is the output feature of the backbone network Select the Student T distribution output head with the largest weight:

[0057]

[0058] Filter edge weights for the output head of the Student's T distribution The last dimension is performed. is the output feature of the backbone network The corresponding optimal Student's T distribution output head number, The integer values ​​correspond to the Student-T distribution output heads with low degrees of freedom, medium degrees of freedom, and high degrees of freedom, respectively. Included The length is Data fragment, output head number according to the optimal Student's T distribution Each data segment is assigned an output head. Inside the output head, a separate linear mapping layer is used to map the input data segment to the parameters of the Student's T distribution in the prediction space. Select a fragment , assuming that the fragment is assigned to the output header , then the formula of this distribution parameter mapping process is expressed as follows:

[0059]

[0060]

[0061]

[0062] In the formula, , and Indicates output header The corresponding degrees of freedom, standard deviation and mean of the Student's T distribution; , and Respectively represent the output head The corresponding weight matrix of the Student's T distribution's degrees of freedom, standard deviation, and mean parameter mapping, , and For output header The corresponding bias coefficient of the Student's T distribution's degrees of freedom, standard deviation, and mean parameter mapping; For output header The lower limit of the degrees of freedom of the Student's T distribution is used to set the range of degrees of freedom of the Student's T distribution output head.

[0063]

[0064] The three groups of Student's T distribution output heads with different degrees of freedom use the same structural design, but are independent of each other in terms of model parameters. The degree of freedom, standard deviation and mean parameters are used to instantiate the Student's T distribution, i.e., predict the distribution.

[0065]

[0066] Student's T distribution based on prediction implement Subsampling, you can get an array ,in, To predict the distribution Execute The result obtained by sampling. The data distribution in the statistical array can be used to obtain the probability prediction result in the form of quantiles; the median of each sampling moment of the array is selected as the deterministic prediction result output .

[0067] (2) Construct a hybrid expert sub-model of forecast information.

[0068] Assume that the input forecast data is , B represents the number of input batches; S represents the length of the input historical sequence; C represents the number of input channels. The input data also needs to be sliced, and it is cut into N time series segments of the same length P along the S dimension. At this time, the input data becomes the sliced ​​forecast data :

[0069]

[0070] The dimension N should satisfy: The data identification tag is used to distinguish the input data of different channels, including load (tag=01), photovoltaic (tag=02) and wind power (tag=03). Map the time series segments to a unified dimension and add the corresponding data identification bias coefficient to the time series segments according to the data type. , thereby incorporating data identification features into time series segments:

[0071]

[0072]

[0073] In the formula, , and are the bias coefficients under three data identifications, and They are the linear mapping weight coefficient and bias coefficient of the feature mapping layer respectively; is the data identification bias coefficient. The forecast data that incorporates the data identification features It is then input into the main routing network of the forecast information hybrid expert sub-model, and the main routing network selects the corresponding expert model group according to the data identifier. This process is shown as follows:

[0074]

[0075]

[0076] In the formula, Score the expert model, and The mapping matrix and bias weights for the main routing network; , indicating the number of the expert model, there are a total of expert models, and the optimal expert model number selected after the sub-routing network operation is .

[0077] Subsequently, the sub-routing network within the expert model group will further select the optimal expert model for the input Perform calculations and output the Gaussian distribution mean of the prediction space respectively With standard deviation .

[0078]

[0079]

[0080] Among them, an expert model includes two FFN networks with independent parameters, which are used to calculate the mean and standard deviation of Gaussian distribution respectively. and Represents the optimal expert model The two weight matrices of the FFN network used to calculate the standard deviation, and Represents the optimal expert model Two bias coefficients of the FFN network used to calculate the standard deviation; and Represents the optimal expert model The two weight matrices of the FFN network used to calculate the mean of the Gaussian distribution are: and Represents the optimal expert model Two bias coefficients of the FFN network used to calculate the mean of the Gaussian distribution.

[0081] This instantiation of a Gaussian distribution is done using a mean and standard deviation:

[0082]

[0083] Gaussian distribution based on prediction implement Subsampling, you can get an array ,in, To predict the distribution Execute The result obtained by sampling. The data distribution in the statistical array can be used to obtain the probability prediction result in the form of quantiles; the median of each sampling moment of the array is selected as the deterministic prediction result output .

[0084] Step 2: Based on the model constructed in step 1, further construct a pre-training dataset. Standardize the time series open source dataset and normalize it to a sequence with a mean of 0 and a variance of 1. In the power system dataset, align the weather forecast data (wind speed, wind direction, irradiance, temperature, humidity) with the timestamps of the corresponding photovoltaic, wind power and load data.

[0085] For the spatiotemporal pattern hybrid expert sub-model, 70% of the open source data set and 30% of the power history data were selected when constructing the pre-training data set, which were divided in chronological order. The first 80% were the training set, the middle 10% were the validation set, and the last 10% were the test set.

[0086] For the forecast information hybrid expert sub-model, only paired samples of forecast and actual values ​​in the power system data are used when constructing the pre-training data set, and stratified sampling is performed according to the site to ensure that the ratio of each site data in the training set, validation set and test set is 7:2:1. In the photovoltaic forecast scenario, the input features include forecast information such as irradiance and temperature in the forecast interval; in the wind power forecast scenario, the input features include forecast information such as hub high wind speed and wind direction in the forecast interval; in the load forecast scenario, the input features also include forecast information such as temperature, holiday signs, weather, humidity, etc.

[0087] Step 3: Based on the pre-training dataset constructed in step 2, complete the parameter initialization and pre-training of the model. The spatiotemporal hybrid expert sub-model sets a 24-layer Transformer decoder, each layer contains 8 attention heads, the model backbone dimension is 1024, the attention head dimension is 128, and the hybrid expert distribution output head part uses three groups of student T distribution output heads with different degrees of freedom. Each student T distribution output head uses 3 independent linear mapping networks to obtain distribution parameters. The AdamW optimizer is used in the training process, the initial learning rate is set to 5e-5, the learning rate attenuation coefficient is set to 0.01, and the cosine annealing learning rate scheduling strategy is adopted. The negative log-likelihood is used as the loss function for model training.

[0088] The forecast information hybrid expert sub-model has a total of 54 FFN expert models, which are equally divided into three expert model groups, responsible for photovoltaic, wind power and load forecasting respectively. The training is divided into two stages. In the first stage, the routing network parameters are fixed, and only the expert models in each field are trained to have the prediction performance in different professional fields. The learning rate is set to 1e-4; in the second stage, the routing network is introduced to participate in the training, and the entire hybrid expert model is jointly optimized. At this time, the learning rate is set to 3e-5; the AdamW optimizer is also used in the training process, and the input data batch can be selected from the range of [32-4096] according to the size of the data set and the computing power limit.

[0089] Based on the above pre-trained model, the power system source and load prediction can be realized. The available data obtained from the specific application scenario is input into the fine-tuned pre-trained model for prediction to obtain the prediction result.

[0090] Based on the available data of downstream application scenarios, the final model structure and fine-tuning scheme for deployment can be confirmed, which includes three types in total:

[0091] (1) Fine-tune and deploy the forecast information hybrid expert sub-model. Use the forecast-actual value pairing data of the target scenario to build a fine-tuning dataset. Load the pre-trained model parameters after training, freeze the forecast hybrid expert model parameters, and only fine-tune the main routing network and sub-routing network. The trainable parameters are The base learning rate is set to 1e-5, and it decreases by 5% in each epoch during fine-tuning. If the negative log-likelihood loss of the validation set does not decrease for 5 consecutive epochs, fine-tuning is stopped.

[0092] (2) Fine-tune and deploy the spatiotemporal pattern hybrid expert sub-model. During the fine-tuning process, the trainable parameters are , including the multi-degree-of-freedom hybrid expert student T distribution output head and the last three layers of the Transformer network, and freeze the other parameters of the model. Load the pre-trained model parameters after training, use the actual historical data of the downstream to build a fine-tuning dataset, and choose to turn off the redundant student T distribution output head according to the validation set test effect during the fine-tuning process. Set the basic learning rate to 1e-5, and decrease it by 5% for each epoch during fine-tuning. If the negative log-likelihood loss of the validation set does not decrease for 5 consecutive epochs, stop fine-tuning.

[0093] (3) Integrated fine-tuning and deployment of dual models. Construct a hybrid architecture based on dual model prediction and perform joint fine-tuning. Suppose the trainable parameters of the spatiotemporal pattern hybrid expert sub-model are , the trainable parameters of the forecast information hybrid expert sub-model are The prediction output of the ensemble model is expressed as:

[0094]

[0095]

[0096]

[0097] In the formula, predict dynamic weights for a trainable ensemble, and They are the prediction results of the spatiotemporal pattern hybrid expert sub-model and the forecast information hybrid expert sub-model respectively. The prediction output of the ensemble model is as follows. The ensemble fine-tuning is divided into two stages: the first stage uses 30% of the training samples and freezes the weights. and , only adjust the dynamic weights of the ensemble predictions In the second stage, 70% of the training samples are used to jointly optimize all model parameters, and the loss function is set as:

[0098]

[0099] In the formula and are the negative log-likelihood loss functions of the spatiotemporal pattern mixed expert sub-model and the forecast information mixed expert sub-model, respectively. represents the KL divergence operation, which is used to align the prediction distributions of the two sub-models. The attention coefficient for integrated fine-tuning is gradually increased from 0.3 to 1.0 during the fine-tuning process, guiding the attention of fine-tuning parameter updates to gradually shift from the prediction effect of the sub-model to the prediction effect of the integrated model.

[0100] Furthermore, the number of model layers and experts can be reduced in the model parameter loading phase to adapt to limited hardware resources. Perform model validation tests and calculate probability evaluation indicators such as the prediction error band coverage and average interval width of the fine-tuned model in real time. When the fine-tuned model meets the accuracy requirements of downstream applications for three consecutive days, it can be put online for application. On the contrary, if the model does not meet the accuracy requirements in continuous testing, the model fallback mechanism is activated to adjust the fine-tuning parameters and reorganize the training.

[0101] Example 2

[0102] The present invention also provides an electronic device, comprising: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the pre-training model construction method for power system source and load prediction.

[0103] Example 3

[0104] The present invention also provides a computer-readable storage medium having computer instructions stored thereon, wherein the computer instructions are used to enable a computer to execute the method for constructing a pre-training model for power system source and load prediction.

[0105] The above describes the specific implementation methods of the present invention in conjunction with the accompanying drawings, which is not intended to limit the scope of protection of the present invention. All equivalent models or equivalent algorithm processes made using the contents of the present invention specification and drawings, directly or indirectly applied to other related technical fields, are within the scope of patent protection of the present invention.

Claims

1. A method for constructing a pre-trained model for power system source and load prediction, characterized in that: The following steps are involved: A spatiotemporal mode hybrid expert sub-model and a forecast information hybrid expert sub-model are constructed respectively; the spatiotemporal mode hybrid expert sub-model adopts a multi-degree-of-freedom hybrid expert Student T distribution output head, and the multi-degree-of-freedom hybrid expert Student T distribution output head includes a high-degree-of-freedom Student T distribution output head, a medium-degree-of-freedom Student T distribution output head and a low-degree-of-freedom Student T distribution output head. The high-degree-of-freedom Student T distribution output head is used to fit the distribution characteristics of load-type stable data; the medium-degree-of-freedom Student T distribution output head is used to fit the scene with periodic fluctuations and a lighter distribution tail; the low-degree-of-freedom Student T distribution output head is used to fit the thick-tail characteristics of frequent extreme values; the forecast information hybrid expert sub-model adopts a two-stage routing network design. In the first stage, the main routing network distributes the input data to the corresponding expert model group according to the data type. In the second stage, the sub-routing network in the expert model group further selects the optimal expert model to process the input data; Integrate open source time series data sets with real power system source and load data to create pre-training data sets; Pre-training the spatiotemporal pattern hybrid expert sub-model and the forecast information hybrid expert sub-model based on the pre-training data set; After the training is completed, the model parameter file formed during the pre-training process is saved to obtain the pre-trained model.

2. The method for constructing a pre-trained model for power system source and load prediction according to claim 1, characterized in that: The pre-trained data set includes a pre-trained data set of a spatiotemporal pattern hybrid expert sub-model and a pre-trained data set of a forecast information hybrid expert sub-model; the pre-trained data set of the spatiotemporal pattern hybrid expert sub-model specifically includes historical data extracted from an open source time series data set and real source and load data of the power system; the pre-trained data set of the forecast information hybrid expert sub-model includes forecast data and actual data extracted from real source and load data of the power system.

3. The method for constructing a pre-trained model for power system source and load prediction according to claim 1, characterized in that: The working method of the forecast information hybrid expert sub-model is: Assume that the input forecast data is , B represents the number of batches of input; S represents the length of the historical sequence of input; C represents the number of channels of input; along the S dimension Divide into N time series segments of the same length P to obtain the sliced ​​forecast data : ; Among them, the dimension N satisfies: ; Introduce data identification tags to distinguish different types of input data, including load, photovoltaic and wind power; map time series segments to a unified dimension, and add corresponding data identification bias coefficients to the time series segments according to the data type , thereby integrating data identification features into the time series segments and obtaining forecast data integrated with data identification features : ; in, and They are the linear mapping weight coefficient and bias coefficient of the feature mapping layer respectively; The main routing network of the forecast information hybrid expert sub-model selects and fuses forecast data with data identification features according to the data identification. The corresponding expert model group; the process is expressed as follows: ; ; In the formula, Score the expert model, and is the mapping matrix and bias weight of the sub-routing network; , indicating the number of the expert model, there are a total of expert models, and the optimal expert model number selected after the sub-routing network operation is ; Then, the sub-routing network in the expert model group further selects the optimal expert model for the input Each expert model includes two FFN networks with independent parameters, which are used to calculate the mean of the Gaussian distribution of the prediction space. With standard deviation : ; ; In the formula, and Represents the optimal expert model The two weight matrices of the FFN network used to calculate the standard deviation, and Represents the optimal expert model Two bias coefficients of the FFN network used to calculate the standard deviation; and Represents the optimal expert model The two weight matrices of the FFN network used to calculate the mean of the Gaussian distribution are: and Represents the optimal expert model Two bias coefficients of the FFN network used to calculate the mean of the Gaussian distribution; Gaussian distribution prediction based on mean and standard deviation: ; In the formula, is the predicted Gaussian distribution; Gaussian distribution based on prediction implement Subsample, get array ,in, To predict the distribution Execute The result is obtained by sampling times; the data distribution in the statistical array can be used to obtain the probability prediction result in the form of quantiles; the median of each sampling moment of the array is selected as the deterministic prediction result.

4. A method for predicting source and load of a power system, characterized in that: The prediction method is implemented using a pre-trained model trained based on the method according to any one of claims 1 to 3; based on the available data of the specific application scenario, the pre-trained model structure and fine-tuning scheme ultimately used for deployment are confirmed; The available data obtained from the application scenario is input into the fine-tuned pre-trained model for prediction to obtain the prediction results.

5. The power system source load prediction method according to claim 4, characterized in that: The pre-trained model structure and fine-tuning scheme for deployment are determined based on the available data of the specific application scenario; there are three specific cases: When the application scenario can only provide historical data for model prediction, the spatiotemporal pattern hybrid expert sub-model is selected, its pre-trained parameters are loaded, and the model is fine-tuned using the historical data of the target scenario; When there is no available historical data or less historical data for the application scenario, the forecast information hybrid expert sub-model is selected, its pre-trained parameters are loaded, and the model is fine-tuned using the target scenario forecast data; When both historical data and forecast data are available and sufficient, the combined structure of the spatiotemporal pattern hybrid expert sub-model and the forecast information hybrid expert sub-model is selected to achieve the highest prediction accuracy, and the integrated fine-tuning technology of the pre-trained model is used to complete the scale alignment and complementary advantages of the two sub-models.

6. The power system source load prediction method according to claim 5, characterized in that: When the application scenario can only provide historical data for model prediction, the spatiotemporal mode hybrid expert sub-model is selected, its pre-trained parameters are loaded, and the model is fine-tuned using the historical data of the target scenario; the specific method is: The spatiotemporal mode hybrid expert sub-model uses the Transformer decoder as the model backbone structure; during the fine-tuning process, the training parameters include the multi-degree-of-freedom hybrid expert student T distribution output head and the last three layers of the Transformer network, and the other parameters of the spatiotemporal mode hybrid expert sub-model are frozen; Loading the parameter file of the spatiotemporal mode hybrid expert sub-model, and constructing a fine-tuning dataset using the target scene historical data; According to the test results of the validation set during fine-tuning, we choose to close the redundant Student-T distribution output head; The base learning rate is set to 1e-5, and it decreases by 5% in each epoch during fine-tuning. If the negative log-likelihood loss of the validation set does not decrease for 5 consecutive epochs, fine-tuning is stopped.

7. The power system source load prediction method according to claim 5, characterized in that: When there is no available historical data or less historical data for the application scenario, the forecast information hybrid expert sub-model is selected, its pre-trained parameters are loaded, and the model is fine-tuned using the target scenario forecast data; the specific method is: Construct paired samples of forecast data and actual data for the target scenario, and use the paired samples to construct a fine-tuning dataset; Loading the parameter file of the forecast information hybrid expert sub-model, freezing the parameters of the forecast information hybrid expert sub-model, and only fine-tuning the main routing network and the sub-routing network; The base learning rate is set to 1e-5, and it decreases by 5% in each epoch during fine-tuning. If the negative log-likelihood loss of the validation set does not decrease for 5 consecutive epochs, fine-tuning is stopped.

8. The power system source load prediction method according to claim 5, characterized in that: When both historical data and forecast data are available and sufficient, a combination structure of a spatiotemporal pattern hybrid expert sub-model and a forecast information hybrid expert sub-model is selected to achieve the highest prediction accuracy, and the integrated fine-tuning technology of the pre-trained model is used to complete the scale alignment and complementary advantages of the two sub-models; the specific method is: Assume that the trainable parameters of the spatiotemporal pattern mixture expert sub-model are , the trainable parameters of the forecast information hybrid expert sub-model are ; then the prediction output of the integrated model composed of two sub-models is It is expressed as: ; ; ; In the formula, Predict dynamic weights for trainable ensembles; and They are the prediction results of the spatiotemporal pattern hybrid expert sub-model and the forecast information hybrid expert sub-model respectively; Ensemble fine-tuning is divided into two stages: the first stage uses 30% of the training samples and freezes the weights and , only adjust the dynamic weights of the ensemble predictions ; In the second stage, 70% of the training samples are used to jointly optimize all parameters of the model, and the loss function is set to: ; In the formula, and are the negative log-likelihood loss functions of the spatiotemporal pattern mixed expert sub-model and the forecast information mixed expert sub-model respectively; represents the KL divergence operation, which is used to align the prediction distributions of the two sub-models. The attention coefficient for ensemble fine-tuning; is the predicted Student's T distribution.

9. An electronic device, characterized in that: include: one or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 3.

10. A computer-readable storage medium having computer instructions stored thereon, characterized in that: The computer instructions are used to enable a computer to execute the steps of the method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Operation state anomaly detection method of electric power system and related equipment

    CN119202967A

  • Coal-fired power plant unit total coal feed quantity and air quantity prediction method based on multi-task learning and related device

    CN119227887A

  • Multivariable source load data prediction method and system based on pre-trained big language model

    CN119377397A

  • Electric power field operation safety detection method and system based on hybrid expert model

    CN119478626A

  • Multi-level and multi-label content classification using unsupervised and ensemble machine learning techniques

    US12026626B1

Cited By

  • Dynamic multi-scale coding source load prediction method and system based on prompt

    CN121124038A