Generative pre-training-based power data simulation method
Through the generative pre-training method based on Transformer and graph neural network, the problems of spatiotemporal characteristics and large-scale data processing in power data simulation are solved, high-quality power data is generated, and the simulation and optimization capabilities of the power system are improved.
Patent Information
- Application Number
- CN202510488693.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-18
AI Technical Summary
The existing graph data generation model is difficult to effectively capture the spatio-temporal characteristics and global characteristics of data in power data simulation, and is inefficient when processing large-scale data, which cannot meet the deep application and value mining needs of power systems.
Generative pre-training method based on Transformer and graph neural network is adopted to build a timing graph structure, multi-attention head mechanism and adversarial loss function to generate power data that meets the actual scenario, retain global features and reduce the computational complexity.
The generated power data accurately captures complex correlation characteristics and spatiotemporal characteristics, improves the quality and practicality of simulated data, can effectively process large-scale data, and provide high-quality tools for power system analysis, prediction and optimization.
Smart Images

Figure CN120336787A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of power data generation, and particularly to a power data simulation method based on generative pre-training. Background Art
[0002] In the power market, various participants (such as power suppliers and electricity consumers) need to obtain power load simulation data within a specific time period that conforms to the periodic characteristics, so as to more reasonably and accurately formulate relevant plans such as the task optimization scheduling of each power generation unit in the power system and the optimal allocation of power energy. Among them, time series is an important feature that needs to be considered in power data simulation. By making a more detailed time division of the load demand, such as data simulation at the hour or minute level, system managers can more accurately understand the volatility and peak periods of power demand, and thus better arrange power generation plans, scheduling, and energy allocation. In current technologies, deep learning and neural network technologies can be used to learn the complex features and laws of the power system from a large amount of data and generate power data that conforms to the actual situation. However, existing graph data generation models such as GraphRNN, Graphite, NetGAN, and VGAE, although they can generate high-quality static graph data, are mainly designed for static graphs and are difficult to fully reflect the complex correlation features and distributions of data, especially unable to effectively capture the spatio-temporal characteristics of data. This limitation restricts the in-depth application and value mining of data. In addition, existing time series graph data simulation methods generally have the problem of long training time and are difficult to process large-scale data. And most generative models use GAT or other variants, focusing on a few neighbors with relatively close distances and unable to effectively retain global features.
[0003] With the continuous development of graph neural network (GNN) technology, various architectures and variants have emerged and are widely used in fields such as social network analysis, recommendation systems, and bioinformatics. Current research hotspots include the scalability of the model and dynamic graph processing, etc. At the same time, the Transformer architecture, with its self-attention mechanism-based design, can process sequence data in parallel, significantly improving efficiency and better focusing on global features. However, based on the Transformer and graph neural network architecture technologies, they have not been applied in the power time series data pre-training model and cannot solve the limitations of existing methods in aspects such as spatio-temporal feature capture and large-scale data processing. Summary of the Invention
[0004] To overcome the drawbacks that existing Transformer and graph neural network architecture technologies have not been applied in the pre-training model of power time-series data and cannot solve the limitations in aspects such as spatio-temporal feature capture and large-scale data processing, the present invention provides a power data simulation method based on generative pre-training. Under the action of relevant processes, with an innovative model architecture and training strategy, it can efficiently process large-scale data, providing important support for the in-depth application and value mining of power data. It not only significantly improves the quality and usability of simulated data, but also provides a new tool for the analysis, prediction, and optimization of power systems, and provides strong data support for better generating real-world power data.
[0005] The technical solution adopted by the present invention to solve its technical problems is as follows:
[0006] A power data simulation method based on generative pre-training includes four steps: obtaining a historical power data set and performing data preprocessing and preliminary analysis, training a model, generating simulated data, and evaluating the quality of the generated simulated data; the step of obtaining a historical power data set and performing data preprocessing and preliminary analysis includes the following processes: S1: Loading time-series graph data, deeply analyzing the time-series characteristics of historical power data, and designing and constructing a graph structure suitable for maintaining power data characteristics; S2: Extracting the feature vectors of nodes, performing normalization or standardization processing, and converting the graph data into a sparse adjacency matrix for subsequent model training; the step of training the model includes the following processes: A1: Initializing the model and setting parameters such as input and output dimensions, number of hidden layers, and number of attention layers; A2: Obtaining the loss value of the model through the difference between power simulation data and actual power consumption data, and evaluating the performance of the model based on this; the step of generating simulated data is to assemble the generated subgraphs into a complete time-series simulated power graph and convert the generated edges into the adjacency matrix of the graph or other required data formats; the step of evaluating the quality of the generated simulated data is to evaluate the quality of the generated time-series power data by calculating the statistical difference between the real data and the generated data. Specifically, first compare the mean and median of indicators such as the average maximum degree, local clustering coefficient, and number of edges in the real graph and the generated graph, and then adjust the model parameters according to the evaluation results, repeating the training and evaluation process until the model performance meets the requirements.
[0007] Further, in step S1 of the step of obtaining a historical power data set and performing data preprocessing and preliminary analysis, the specific graph structure includes nodes and edges, and their time-series information.
[0008] Further, in step A1 of the step of training the model, the central graph sampling method of the temporal graph is used to first extract the local temporal structure, which specifically includes the following processes: (1) Given the spatio-temporal graph adjacency matrix A t=1:T , first load the node features of each snapshot corresponding to its timestamp t; (2) Select representative time nodes as the central nodes of each central graph, and recursively sample neighboring nodes. The new nodes are sampled from the neighboring nodes of the early sampled nodes: (3) Build a neural network structure for time series data based on Transformer, fully learn the distribution pattern of power data, extract numerical features, use the multi-attention head mechanism to encode each sampled graph, and through a linear layer, reconstruct the output of the encoder into the structure of the graph.
[0009] Further, in step A1 of the training model, the calculation formula of the multi-head self-attention mechanism is as follows.
[0010] The calculation formula of each attention head is as follows.
[0011] The calculation formula of the multi-head attention output is as follows.
[0012] MultiHead(Q,K,V)=Concat(head1,head2,...,head h )W O , and each encoder layer consists of a multi-head self-attention mechanism and a feed-forward neural network. The feed-forward neural network consists of two linear transformations and an activation function, and the calculation formula is as follows.
[0013] FFN(x)=max(0,xW1+b1)W2+b2.
[0014] Further, in step A2 of the training model, an adversarial loss function is used. The role of the loss function is to minimize the difference between the generated graph and the real graph, thereby improving the quality of the generated graph. Through backpropagation, the model parameters will be updated according to the loss function to reduce the difference between the generated graph and the real graph.
[0015] Further, in step A2 of the training model, the adversarial loss formula is:
[0016] The adversarial loss formula uses the Adam optimizer for momentum adaptation and weight decay.
[0017] The beneficial effects of the present invention compared with the prior art are as follows: The present invention can not only generate power data that conforms to the actual scenario, but also accurately capture the complex correlation features, distribution laws, and spatio-temporal characteristics of the data, while retaining the global features; it can effectively capture the non-linear dynamic behavior and long-distance dependence relationships in the power system, thereby generating more realistic and reliable simulation data; it significantly improves the understanding and simulation ability of the complex structure of the power system; while extracting local temporal information, it can effectively reduce the computational complexity and improve the operation efficiency; it enhances the generalization ability of the model; the generated power data not only has practical application value, but also provides high-quality basic data for deep learning and data mining. In summary, with the innovative model architecture and training strategy, the present invention can efficiently process large-scale data, providing important support for the in-depth application and value mining of power data, not only significantly improving the quality and practicality of simulation data, but also providing new tools and methods for the analysis, prediction, and optimization of power systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 It is a flowchart showing the process of a power data simulation method based on generative pre-training of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] Figure 1 As shown, a power data simulation method based on generative pre-training includes four steps: obtaining a historical power data set and performing data preprocessing and preliminary analysis, training a model, generating simulation data, and evaluating the quality of the generated simulation data.
[0020] Figure 1 As shown, obtaining the historical power data set and performing data preprocessing and preliminary analysis includes the following processes: (1) Loading time-series graph data, deeply analyzing the time-series characteristics of historical power data, and designing and constructing a graph structure suitable for maintaining the characteristics of power data. The specific graph structure includes nodes and edges, as well as their time-series information (the function of this step is to provide a topological framework with physical significance for the model by constructing a graph structure that reflects the spatio-temporal correlation characteristics of the power system, ensuring the accuracy of subsequent feature learning and data generation); (2) Extracting the feature vectors of the nodes, performing normalization or standardization processing, and converting the graph data into a sparse adjacency matrix (the function of this step is to unify the data scale, reduce the impact of numerical differences on model training, optimize the storage efficiency, and provide a standardized and efficient computational representation for model input) for subsequent model training.
[0021] Figure 1 As shown, training the model includes the following processes: (1) First, initialize the model, set parameters such as input and output dimensions, number of hidden layers, number of attention layers, etc. Specifically, in order to extract local temporal structures, the present invention uses the central graph sampling method of temporal graphs, given the spatio-temporal graph adjacency matrix A t=1:T, first load the node features of each snapshot corresponding to its timestamp t. Then, select representative time nodes as the central nodes of each central graph, and recursively sample neighboring nodes. The new nodes are sampled from the neighboring nodes of the early sampled nodes (the function of this step is to focus on key spatio-temporal regions through the central node sampling strategy, retain the local dynamic characteristics of power data while reducing computational complexity, and provide structured temporal neighborhood information for the subsequent attention mechanism). (2) Build a neural network structure for time series data based on Transformer to fully learn the distribution pattern of power data and extract numerical features; among them, use the multi-attention head mechanism to encode each sampled graph, and through a linear layer, reconstruct the output of the encoder into the structure of the graph; specifically, the calculation formula of the multi-head self-attention mechanism is as follows:
[0022] where Q, K, and V represent the query, key, and value matrices respectively, and d k is the dimension of the key vector; the multi-head attention mechanism can capture information in different subspaces by calculating multiple attention heads in parallel; the calculation of each attention head can be expressed as: where, is the learnable parameter matrix of the i-th attention head; the outputs of multiple attention heads are obtained through concatenation and linear transformation to get the final multi-head attention output, and the formula is as follows: MultiHead(Q,K,V)=Concat(head1,head2,...,head h )W O , where h is the number of attention heads, and W O is the weight matrix of the output linear transformation; each encoder layer consists of a multi-head self-attention mechanism and a feed-forward neural network. The feed-forward neural network usually consists of two linear transformations and an activation function, and the formula is as follows: FFN(x)=max(0,xW1+b1)W2+b2 (the function of this step is to parallelly mine the spatio-temporal correlation patterns of power data in different feature subspaces through the multi-head self-attention mechanism, combine the feed-forward neural network to enhance the non-linear representation ability, and finally generate dynamic graph data that conforms to the characteristics of the real power grid topology through structure reconstruction, realizing the accurate modeling of the complex spatio-temporal dependence relationship of the power system); (3) Finally, obtain the loss value of the model through the difference between the power simulation data and the actual power consumption data to evaluate the performance of the model. Among them, the present invention uses adversarial loss, and the function of this loss function is to minimize the difference between the generated graph and the real graph, thereby improving the quality of the generated graph; through backpropagation, the model parameters will be updated according to the loss function to reduce the difference between the generated graph and the real graph; the adversarial loss formula is:
[0023]
[0024] , at the same time, the Adam optimizer is used for momentum adaptation and weight decay (the function of this step is to quantify the topological and statistical differences between the generated power data and the real data through the adversarial loss function, and combine the adaptive learning rate and weight decay strategy of the Adam optimizer to drive the update of model parameters to minimize the loss value, thereby improving the spatio-temporal feature fidelity of the generated data).
[0025] Figure 1 As shown, in generating the simulated data, the generated subgraphs are assembled into a complete time-series simulated power graph, and the generated edges are converted into the adjacency matrix of the graph or other required data formats (the function of this step is to reconstruct the locally sampled time-series subgraphs into a complete power network graph through topological splicing and format conversion techniques, ensuring the spatio-temporal coherence and structural integrity of the generated data, and at the same time adapting to the data input requirements of power system analysis tools to achieve seamless docking from model output to engineering applications).
[0026] Figure 1 As shown, in evaluating the quality of the generated simulated data, the quality of the generated time-series power data is evaluated by calculating the statistical differences between the real data and the generated data; for example: comparing the means and medians of indicators such as the average maximum degree, local clustering coefficient, and number of edges in the real graph and the generated graph. Adjust the model parameters according to the evaluation results, and repeat the training and evaluation process until the model performance meets the requirements (the function of this step is to quantify and evaluate the compliance of the statistical characteristics and physical laws of the generated data based on multi-dimensional indicators such as the average maximum degree and local clustering coefficient, and through iterative optimization of model parameters until the generated data meets the power system simulation accuracy standard, providing reliable data support that conforms to the real situation for power grid planning and scheduling decisions).
[0027] The following content conducts an empirical evaluation of the effectiveness of a power data simulation method based on generative pre-training. Specifically, the experimental settings are introduced first, and then the experimental results are presented.
[0028] The experimental settings are as follows. (1) Dataset: Taking the DBLP dataset as an example, this dataset is a co-author network composed of 317,080 nodes and 1,049,866 edges. (2) Baseline: In order to evaluate the model of the present invention, the present invention compares it with the following existing advanced models;
[0029] TagGen (Tag-based Graph Representation Learning): A tag-based graph representation learning method that enhances the learning of graph structures through tag information.
[0030] NetGAN(GraphNeural Network GenerativeAdversarial Networks): Graph neural network generative adversarial networks that generate graph structures through generative adversarial training.
[0031] E-R(Edge Prediction via GraphNeural Networks): Edge prediction using graph neural networks to evaluate changes in graph structures.
[0032] B-A(Bias-Aware GraphAttention Network): Bias-aware graph attention network that learns the relationships between nodes through an attention mechanism.
[0033] VGAE(Variational GraphAutoencoder): Variational graph autoencoder that captures the graph structure of a graph through variational learning.
[0034] Graphite(Graphite:AGraphNeural Network): A graph neural network for representation learning and generation of graph structures.
[0035] SBMGNN(Spectral-Based GraphNeural Network): A spectral-based graph neural network that learns using the spectral properties of a graph. (3) Parameter settings. In the experiment, first set the input dimension to the number of nodes multiplied by the number of time steps, the hidden layer dimension to 128, the number of attention heads to 4, and the output dimension to the number of nodes; the model uses the Adam optimizer with a learning rate of 4e-3, a weight decay of 1e-4, and a maximum number of training epochs of 500; the model is trained using a multi-layer full neighbor sampler with a batch size of 128, and the learning rate is adjusted using a cosine annealing learning rate scheduler. (4) Evaluation metrics,
[0036] , The experimental evaluation metrics are Mean Degree, Local Clustering Coefficient (LCC), Wedge Count, Claw Count, Power Law Exponent (PLE), and Number of Components, which are used to measure multiple aspects of the similarity between the generated graph and the real graph structure.
[0037] Taking the DBLP dataset as an example, the specific experimental results are as follows:
[0038] Table 1 Comparison of evaluation metric values of different models
[0039]
[0040]
[0041] The present invention has the following advantages. (1) By combining an advanced graph neural network architecture and an attention mechanism, it can not only generate power data that conforms to the actual scenario, but also accurately capture the complex correlation features, distribution laws, and spatio-temporal characteristics of the data, while retaining global features (in the present invention, first, a dynamic graph adjacency matrix At=1:T containing temporal information and node features X(t) are constructed in the data preprocessing stage to explicitly encode the spatio-temporal coupling relationship of the power system; second, the global calculation characteristics of self-attention are used for modeling). (2) Compared with the prior art, the present invention innovatively combines a generative adversarial network (GAN) with a Transformer model, which can effectively capture the non-linear dynamic behavior and long-distance dependence relationship in the power system, thereby generating more realistic and reliable simulation data (in the present invention, the statistical characteristics matching between the generated data and the real data is achieved through the adversarial loss function in the generative adversarial network. This loss design combined with the Transformer architecture effectively solves the problem of modeling long-distance dependence relationships, making the generated simulation data more realistic and reliable). (3) In addition, by introducing a multi-attention head mechanism, the model can extract information from multiple representation subspaces, fully mine the diverse patterns and complex relationships in the data, and significantly improve the understanding and simulation capabilities of the complex structure of the power system (in the present invention, the multi-attention head mechanism is used to process the information in different subspaces in parallel). (4) By adopting the central graph sampling method, the model can focus on the local structure in the graph, effectively reduce the computational complexity while extracting local temporal information, and improve the operation efficiency (in the present invention, representative time nodes are selected as the central nodes, and the sampling range is gradually expanded. This method ensures the extraction of local temporal information and significantly reduces the computational complexity). (5) During the training process, the Adam optimizer is adopted, and through the momentum adaptation and weight decay strategies, the overfitting problem is effectively prevented, and the generalization ability of the model is further enhanced. (6) The overall loss function balances the optimization objectives between the authenticity and structural similarity of the generated data through weighted summation, making the generated power data not only have practical application value, but also provide high-quality basic data for deep learning and data mining. In summary, the generative pre-training power simulation data method proposed by the present invention, relying on the innovative model architecture and training strategy, can efficiently process large-scale data, provides important support for the in-depth application and value mining of power data. This method not only significantly improves the quality and practicality of the simulation data, but also provides new tools and methods for the analysis, prediction, and optimization of the power system.
[0042] The foregoing has shown and described the basic principles, main features and advantages of the present invention. For a person skilled in the art, it is obvious that the present invention is limited to the details of the above-mentioned exemplary embodiments, and without departing from the spirit or basic features of the present invention, the present invention can be implemented in other specific forms. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, it is intended to embrace all changes that fall within the meaning and scope of the equivalent elements of the claims in the present invention.
[0043] In addition, it should be understood that although this specification is described in terms of embodiments, not every embodiment only contains an independent technical solution. This narrative manner of the specification is only for clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in the embodiments can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A power data simulation method based on generative pre-training, characterized in that, It includes four steps: obtaining a historical power dataset and performing data preprocessing and preliminary analysis, training a model, generating simulated data, and evaluating the quality of the generated simulated data; the step of obtaining a historical power dataset and performing data preprocessing and preliminary analysis includes the following processes. S1: Load time series graph data, deeply analyze the time series characteristics of historical power data, and design and construct a graph structure suitable for maintaining power data characteristics. S2: Extract the feature vectors of nodes, perform normalization or standardization processing, and convert the graph data into a sparse adjacency matrix for subsequent model training; the step of training the model includes the following processes. A1: Initialize the model and set parameters such as input and output dimensions, number of hidden layers, and number of attention layers. A2: Obtain the loss value of the model through the difference between the power simulation data and the actual power consumption data, and evaluate the performance of the model based on this; in the step of generating simulated data, assemble the generated subgraphs into a complete time series simulated power graph, and convert the generated edges into the adjacency matrix of the graph or other required data formats; the step of evaluating the quality of the generated simulated data is to evaluate the quality of the generated time series power data by calculating the statistical difference between the real data and the generated data. Specifically, first compare the means and medians of indicators such as the average maximum degree, local clustering coefficient, and number of edges in the real graph and the generated graph, and then adjust the model parameters according to the evaluation results, and repeat the training and evaluation process until the model performance meets the requirements.
2. The power data simulation method based on generative pre-training according to claim 1, wherein In step S1 of obtaining a historical power dataset and performing data preprocessing and preliminary analysis, the specific graph structure includes nodes and edges, as well as their time series information.
3. A power data simulation method based on generative pre-training according to claim 1, characterized in that In step A1 of the training model, the central graph sampling method of the temporal graph is used to extract the local temporal structure first. The specific process is as follows: (1) Given the spatio-temporal graph adjacency matrix A t=1:T , first load the node features of each snapshot corresponding to its timestamp t; (2) Select representative time nodes as the central nodes of each central graph, and recursively sample neighboring nodes. The new nodes are sampled from the neighboring nodes of the early sampled nodes: (3) Based on Transformer, build a neural network structure applied to time series data, fully learn the distribution pattern of power data, extract numerical features, use the multi-attention head mechanism to encode each sampled graph, and through a linear layer, reconstruct the output of the encoder into the structure of the graph.
4. A power data simulation method based on generative pre-training according to claim 3, characterized in that In step A1 of the training model, the calculation formula of the multi-head self-attention mechanism is as follows: The calculation formula for each attention head is as follows: The calculation formula for the multi-head attention output is as follows, MultiHead(Q, K, V) = Concat(head1, head2,..., head h )W O , each encoder layer consists of a multi-head self-attention mechanism and a feed-forward neural network. The feed-forward neural network consists of two linear transformations and an activation function, and the calculation formula is as follows, FFN(x) = max(0, xW1 + b1)W2 + b2.
5. A power data simulation method based on generative pre-training according to claim 1, characterized in that In step A2 of training the model, an adversarial loss function is used. The role of the loss function is to minimize the difference between the generated graph and the real graph, thereby improving the quality of the generated graph. Through backpropagation, the model parameters will be updated according to the loss function to reduce the difference between the generated graph and the real graph.
6. A method for simulating power data based on generative pre-training according to claim 1, characterized in that, In step A2 of the training model, the adversarial loss formula is as follows: The adversarial loss formula uses the Adam optimizer for momentum adaptation and weight decay.