A method and system for generatively constructing a digital twin model of a power grid
The pre-trained language model of the power system is constructed through the LlAMA2 model and reinforcement learning algorithm, which solves the problems of duplicate modeling and insufficient accuracy in power grid digital modeling, and realizes efficient and accurate generation of power grid digital twin models.
Patent Information
- Application Number
- CN202311465927.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-06
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2043-11-06
AI Technical Summary
The existing power grid digital modeling technology has serious repetitive modeling work and limited modeling accuracy, and cannot effectively explore the implicit relationships between equipment, resulting in waste of resources and insufficient modeling accuracy.
The LlAMA2 model is used to extract multimodal data features and build a pre-trained language model of the power system. By comparing learning losses and dot product calculations, combined with reinforcement learning and knowledge fine-tuning algorithms, a high-precision grid digital twin model is generated.
It realizes high-precision and real-time modeling of power system equipment, reduces duplicate modeling work, improves modeling efficiency and accuracy, and can effectively understand the implicit relationships between equipment.
Smart Images

Figure CN117454766B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of power model construction, and in particular to a method and system for generatively constructing a power grid digital twin model. Background Art
[0002] At present, for the digital modeling of power grids, the hybrid construction technology of digital twin models of power equipment is mainly used. This technology combines the existing mechanism models represented by physical and mathematical formulas, and the data-driven model built based on serialized monitoring data and AI algorithms to realize the construction of equipment digital twin hybrid models. Based on the existing digital twin models of power system equipment, the implicit relationship between devices is mined through AI algorithms, and the existing digital twin models are adjusted and combined to realize the construction of digital twin systems. However, there are the following disadvantages:
[0003] 1. Repeated modeling: For each different modeling task, data preprocessing, modeling, and debugging work need to be repeated on the corresponding set of equipment, which seriously consumes computing power, time, and personnel;
[0004] 2. Limited modeling accuracy: It is impossible to perform corresponding knowledge mining based on the entire set of devices, and it is impossible to accurately model small disturbances between devices. Summary of the invention
[0005] In view of the above technical problems, the present invention provides a method for generating a digital twin model of a power grid, comprising:
[0006] Acquire a full set of multimodal data in the power system; extract data features of the full set of multimodal data based on the LIAMA2 model;
[0007] Based on the full set of multimodal data and corresponding data features, construct a pre-trained language model of the power system for digital twins;
[0008] The pre-trained language model of the power system is trained using contrastive learning loss InfoNCE, part of the data in the multimodal data set is selected as training data, and the text is embedded into the corresponding image according to the data characteristics of the training data and the cosine similarity between the image and the text, and each node of the equipment operation of the power system is obtained, and the major nodes are connected to output the equipment operation topology structure of the power system;
[0009] Selecting part of the data in the multimodal data set as test data, using dot product calculation to query the similarity between the image or text in the test data and the image or text in the multimodal data set, and if the similarity meets a preset value, completing the test of the pre-trained language model of the power system;
[0010] The power system pre-trained language model outputs a corresponding power grid digital twin model according to pre-constructed prompt words.
[0011] Furthermore, it also includes:
[0012] Based on the knowledge fine-tuning algorithm, the fine-tuning cycle and the batch size are set, the power system pre-trained language model is fine-tuned using an optimizer, and the parameters of the power system pre-trained language model are updated according to the fine-tuning results.
[0013] Furthermore, the data features of the entire set of multimodal data are extracted based on the LIAMA2 model, including:
[0014] By introducing a graph attention layer into the LIAMA2 model, image features of the graph data in the multimodal data set are extracted;
[0015] By extracting text features of text data in the multimodal data set in the LIAMA2 model;
[0016] By introducing a graph attention layer into the LIAMA2 model, the topological features of the topological data in the entire set of multimodal data are extracted.
[0017] Furthermore, the contrastive learning loss InfoNCE is used to train the pre-trained language model of the power system, and part of the data in the multimodal data set is selected as training data. The text is embedded into the corresponding image according to the data characteristics of the training data and the cosine similarity between the image and the text, and each node of the equipment operation of the power system is obtained, including:
[0018] According to the contrastive learning loss, the cosine similarity of the image and text embedding of the correct matching pair is maximized, while the cosine similarity of the wrong matching pair is minimized, the text is embedded into the corresponding image, and each node of the equipment operation of the power system is obtained. The contrastive learning loss InfoNCE is as follows:
[0019]
[0020] Among them, Ltotal represents the contrastive learning loss, z Ij represents the embedding vector of image j, z Tj represents the embedding vector of text j, p Tj and p Ij Denote the positive sample embedding vectors of text j and image j respectively, Q T and Q I Represent the negative sample embedding vector queues of text and image respectively, and τ is a temperature parameter that controls the refinement of the embedding vector.
[0021] Further, by using dot product calculation, query the similarity between the images or texts in the test data and the images or texts in the multimodal data set, including:
[0022] Use dot product calculation to query the similarity between the images or texts in the test data and the images or texts in the multimodal data set, so as to retrieve relevant data. The specific formula is:
[0023]
[0024] Among them, similarity represents the cosine similarity between two embedding vectors. A and B are the embedding vectors of the query image or text and the data in the multimodal data set respectively. ∣A∣∣2∣∣B∣∣2 represent the L2 norms of A and B respectively.
[0025] Further, after the test of the pre-trained language model for the power system, it also includes:
[0026] By constructing a reward model, improve the prediction ability of the pre-trained language model for the power system, specifically including:
[0027] Define the feedback rules of human trainers. The standard interval of the device operation state is [S lower , S upper , then the rules are as follows:
[0028] {If the device state s satisfies S lower ≤ s ≤ S upper , then the human feedback value H = 1}
[0029] {If the device state s does not satisfy S lower ≤ s ≤ S upper , then the human feedback value H = -1}
[0030] Among them, Slower and Supper represent the lower and upper limits of the interval respectively, and s represents the current state of the device;
[0031] Select the deep Q network as the reinforcement learning model and initialize a Q network: Q(s, a; θ), where s is the state, a is the action, and θ is the network parameter;
[0032] At each time step t, select the action a according to the current state s and the Q network, and then obtain the new state s' and the human feedback value H through interaction with the environment;
[0033] Build a reward function. The first part uses the square of the difference between the human feedback value H and the predicted value Q(s, a; θ) of the Q-network as the reward function R. The second part uses an additional reward based on the device state. If the device state is within the normal range, a positive reward is given; otherwise, if the device state is abnormal, a negative reward is given. The completed reward function is:
[0034] R = (H - Q(s, a; θ)) 2 + β * |s - s′| 2
[0035] where β is a hyperparameter used to balance the importance of the two parts, and s and s' are the device states before and after the action respectively;
[0036] Update the Q-network using experience replay and the target network strategy. In experience replay, the state, action, reward, and new state of each step are stored in a replay buffer D, i.e., D = {(s, a, R, s')}. At each time step, a batch of data is randomly retrieved from D, and the loss function is calculated using this data and the target network, and the parameters θ of the Q-network are updated. The specific loss function is:
[0037] L(θ) = E {s,a,R,s′} [(R + γmax {a′} Q(s′, a′; θ′) - Q(s, a; θ)) 2
[0038] where γ is the discount factor and θ' are the parameters of the target network. At each time step, θ is used to update θ' for a selected part to keep the target network stable.
[0039] Furthermore, it also includes:
[0040] Construct the prompt words of the pre-trained language model for the power system according to the application scenarios and / or construction tasks of the pre-trained language model for the power system; the prompt words set the context of the conversation and the expected output form of the pre-trained language model for the power system.
[0041] Furthermore, the prompt words include: prompt words for the device digital twin modeling mode, prompt words for the unit-level digital twin model construction mode, prompt words for the system-level digital twin model construction mode, prompt words for the digital twin system operation mode, prompt words for the optimization mode based on digital twins, and prompt words for the operation and inspection generation mode.
[0042] The present invention also provides a generative construction system for a power grid digital twin model, including:
[0043] A data feature extraction module, used to obtain a full set of multimodal data in the power system; extract data features of the full set of multimodal data based on the LIAMA2 model;
[0044] A model building module, used to build a pre-trained language model of the power system for digital twins based on the full set of multimodal data and corresponding data features;
[0045] A training module is used to train the pre-trained language model of the power system using contrastive learning loss InfoNCE, select part of the data in the multimodal data set as training data, embed the text into the corresponding image according to the data characteristics of the training data and the cosine similarity between the image and the text, obtain each node of the equipment operation of the power system, connect the major nodes, and output the equipment operation topology structure of the power system;
[0046] A testing module is used to select part of the data in the multimodal data set as test data, and use dot product calculation to query the similarity between the image or text in the test data and the image or text in the multimodal data set. If the similarity meets a preset value, the test of the pre-trained language model of the power system is completed.
[0047] The model output module is used for the power system pre-trained language model to output the corresponding power grid digital twin model according to the pre-built prompt words.
[0048] Further, including:
[0049] An updating module is used to set a fine-tuning cycle and a batch size based on a knowledge fine-tuning algorithm, use an optimizer to fine-tune the pre-trained language model of the power system, and update the parameters of the pre-trained language model of the power system according to the fine-tuning result.
[0050] The present invention discloses a generative construction method and system for a digital twin model of a power grid, including: obtaining a full set of multimodal data in a power system; extracting data features of the full set of multimodal data based on an LIAMA2 model; constructing a pre-trained language model for a power system oriented to digital twins based on the full set of multimodal data and corresponding data features; the pre-trained language model for the power system outputs a corresponding digital twin model of a power grid according to pre-constructed prompt words. High-precision and real-time modeling of new physical equipment of a power system is achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 It is a flow chart of a method for generating a digital twin model of a power grid provided by an embodiment of the present invention;
[0052] Figure 2It is a structural schematic diagram of a power grid digital twin model generative construction system provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0053] Many specific details are set forth in the following description to facilitate a full understanding of the present invention. However, the present invention can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the connotation of the present invention, so the present invention is not limited to the specific implementation disclosed below.
[0054] Embodiment 15
[0055] Figure 1 This is a flow chart of a method for constructing a power grid digital twin model according to an embodiment of the present invention. Figure 1 The method provided by the present invention is described in detail.
[0056] Step S101, obtaining a full set of multimodal data in the power system; extracting data features 5 of the full set of multimodal data based on the LIAMA2 model.
[0057] The full set of multimodal data in the power system is the basis for model construction. The present invention obtains data based on the power grid data center and completes corresponding preprocessing, so it is not specifically explained.
[0058] Extracting data features of the multimodal data set includes: extracting image features of graph data in the multimodal data set by introducing a graph attention layer in the LlAMA2 model; extracting text features of text data in the multimodal data set by introducing a graph attention layer in the LlAMA2 model; extracting topological features of topological data in the multimodal data set by introducing a graph attention layer in the LlAMA2 model. The specific process is as follows:
[0059] The goal of this step is to learn two encoders that can embed image and text samples into the same space for effective image-text retrieval. The specific steps are as follows:
[0060] 1. Feature extraction: Extract relevant modal data features of power grid data based on the LIAMA2 model:
[0061] 1) Image feature extraction: Because LlAMA2 is not directly applicable to graph data, this method proposes to use a graph-based attention mechanism to modify the LlAMA2 architecture so that it can process graph data. Specifically, by introducing a graph attention layer in LlAMA2, the model can focus on and capture the relationship between nodes in the graph. In training and actual reasoning, the self-attention mechanism in LlAMA2 can be replaced with a graph attention mechanism to achieve modal compatibility with the power grid topology map;
[0062] 2) Text feature extraction: The present invention optimizes the LIAMA2 architecture in two aspects to make it more compatible with the representation and reasoning of power grid data:
[0063] a) Category embedding solves the embedding problem of category variables of large-scale power grid equipment by introducing category embedding technology. Specifically, by adding an embedding layer in LlAMA2 to train an embedding vector for each category variable and map it to a continuous vector space, the ability to capture similarities between categories is improved;
[0064] b) Standardization of numerical variables: Because standardization can eliminate the scale differences between different numerical variables, the problem of continuous numerical variables can be solved by standardization. Specifically, each numerical variable is subtracted from its mean before entering the LIAMA2 model, and then divided by its standard deviation.
[0065] 3) Topological feature extraction: refer to the figure feature extraction part.
[0066] Step S102: construct a pre-trained language model of the power system for digital twins based on the full set of multimodal data and corresponding data features.
[0067] Step S103, using contrastive learning loss InfoNCE to train the pre-trained language model of the power system, selecting part of the data in the multimodal data set as training data, embedding the text into the corresponding image according to the data characteristics of the training data and the cosine similarity between the image and the text, obtaining each node of the equipment operation of the power system, connecting the major nodes, and outputting the equipment operation topology structure of the power system.
[0068] Model training is based on contrastive learning. For a given image embedding, the goal is to get the best text embedding from a batch of image embeddings, and vice versa. According to the contrastive learning loss, the cosine similarity of the image and text embeddings of the correct matching pairs is maximized, while the cosine similarity of the wrong matching pairs is minimized. The text is embedded into the corresponding image to obtain each node of the equipment operation of the power system. The contrastive learning loss InfoNCE is as follows:
[0069]
[0070] Among them, Ltotal represents the contrastive learning loss, z Ij represents the embedding vector of image j, z Tj represents the embedding vector of text j, p Tj and p Ij Denote the positive sample embedding vectors of text j and image j respectively, Q T and Q ISeparate queues represent negative sample embedding vectors of text and images respectively. τ is a temperature parameter that controls the fineness of the embedding vectors.
[0071] Step S104: Select part of the data in the multimodal data set as test data. Use dot product calculation to query the similarity between the images or texts in the test data and the images or texts in the multimodal data set. If the similarity meets the preset value, the testing of the pre-trained language model for the power system is completed.
[0072] Use dot product calculation to query the similarity between the images or texts in the test data and the images or texts in the multimodal data set, so as to retrieve relevant data. The specific formula is:
[0073]
[0074] Among them, similarity represents the cosine similarity between two embedding vectors. A and B are the embedding vectors of the queried image or text and the embedding vectors of the data in the multimodal data set respectively. ∣A∣∣2∣∣B∣∣2 represent the L2 norms of A and B respectively.
[0075] Furthermore, for the pre-trained language model of the power system, after the testing is completed, it also includes:
[0076] By constructing a reward model, improve the prediction ability of the pre-trained language model of the power system, specifically including:
[0077] 1. Rule setting: First, define the feedback rules of human trainers. The standard interval of the device operation state is [S lower , S upper , then the rules are as follows:
[0078] {If the device state s satisfies S lower ≤ s ≤ S upper , then the human feedback value H = 1}
[0079] {If the device state s does not satisfy S lower ≤ s ≤ S upper , then the human feedback value H = -1}
[0080] Among them, Slower and Supper represent the lower and upper limits of the interval respectively, and s represents the current state of the device;
[0081] 2. Initialization of the reinforcement learning model: Select the deep Q network as the reinforcement learning model and initialize a Q network: Q(s, a; θ), where s is the state, a is the action, and θ is the network parameter;
[0082] 3. Selection of state-action pairs: At each time step t, an action a is selected based on the current state s and the Q network, and then a new state s' and a human feedback value H are obtained through interaction with the environment;
[0083] 4. Reward function: In the first part, the square of the difference between the human feedback value H and the predicted value Q(s, a; θ) of the Q network is used as the reward function R, i.e., R = (H - Q(s, a; θ)) 2 ; In the second part, an additional reward based on the device state is used. For example, if the device state is within the normal range, a positive reward can be given; conversely, if the device state is abnormal, a negative reward is given. This can guide the model to pay more attention to changes in the device state, thereby improving the model's prediction ability. The complete reward function is:
[0084] R = (H - Q(s, a; θ)) 2 + β * |s - s′| 2
[0085] where β is a hyperparameter used to balance the importance of the two parts, and s and s' are the device states before and after the action, respectively;
[0086] 5. Network update: The Q network is updated using experience replay and the target network strategy. In experience replay, the state, action, reward, and new state of each step are stored in a replay buffer D, i.e., D = {(s, a, R, s')}. At each time step, a batch of data is randomly sampled from D, and the loss function is calculated using this data and the target network, and the parameters θ of the Q network are updated.
[0087] The specific loss function is:
[0088] L(θ) = E {s,a,R,s′} [(R + γ max {a′} Q(s′, a′; θ′) - Q(s, a; θ)) 2
[0089] where γ is the discount factor and θ' are the parameters of the target network. At each time step, θ is used to update θ' for a selected part to keep the target network stable.
[0090] Furthermore, in the power system pre-trained language model, "environment" refers to the grid device status and topology structure generated by the digital twin pre-trained language model; "action" refers to the pre-output optimization of the text generated by the digital twin pre-trained language model; "reward" is the effect of these modifications in realizing grid device maintenance decisions. The specific steps for fine-tuning the model PPO are as follows:
[0091] 1. Initialize the policy network π and the value network V: This method uses a multi-layer perceptron (MLP) neural network to implement the policy network π. The specific steps are as follows:
[0092] a) Define the input data: First, define the structure of the input data. This method converts the power grid device status and topological structure generated by the digital twin pre-trained language model into a format that can be processed by the MLP. Specifically, it includes converting text data into numerical data, such as through word embedding or other text encoding methods;
[0093] b) Initialize the MLP neural network: The MLP usually consists of an input layer, several hidden layers, and an output layer. The number of nodes in the input layer should be equal to the number of features, that is, the dimension of the converted power grid device status and topological structure data. The number of nodes in the hidden layer is determined through experiments. The number of nodes in the output layer is equal to the dimension of the action space, that is, the number of all possible modification actions. This method uses a softmax layer to interpret the output of the network as the probability of selecting each action in a given state;
[0094] c) Define the loss function: The negative log-likelihood loss is used to encourage the network to select actions with higher rewards.
[0095] The corresponding value network V is also defined as an MLP accordingly. Its input is the power grid device status and topological structure generated by the Digital-Power LanguageModel, and the output is the value estimate of the corresponding state.
[0096] V(s; θ v ) = MLP(s; θ v )
[0097] where s is the state and θv is the parameter of the value network.
[0098] 2. Data collection: Use the current policy π to operate on the output of the digital twin pre-trained language model. For each training cycle, interact with the power grid environment according to the current policy network π to collect trajectory data (s, a, r, s'), where s, a, r, s' represent the current state, the action taken, the immediate reward obtained, and the next state respectively.
[0099] 3. Calculate the advantage estimate: Use the GAE method to calculate the advantage of each state-action pair based on the collected data and the prediction of the current value network, which helps to judge how a modification action compares to the average effect. The formula for the advantage function is A(s, a), where γ is the discount factor and V(s') is the value prediction of the next state:
[0100] A(s, a) = r + γV(s′; θ v ) - V(s; θv )
[0101] 4. Incorporate human feedback into advantage calculation: In the TAMER framework, the engineers within the system provide the corresponding human feedback R. Specifically, intuitive feedback is given on the goodness or badness of certain modification actions. Then, these feedbacks R can be incorporated into the advantage calculation and can be regarded as an additional reward signal:
[0102] A′(s,a) = A(s,a) + R
[0103] 5. Update the policy and value networks: Use the collected data and advantage estimates to update the policy and value networks. Specifically, update the parameters of the policy network and value network according to the new advantage estimate A'(s,a), as shown in the formula, where θ π and θ v represent the parameters of the policy network and value network respectively, and η is the learning rate:
[0104]
[0105] Among them, θπ: the parameters of the policy network. The policy network outputs the probabilities of taking various actions in a given state.
[0106] θv: the parameters of the value network. The value network attempts to predict the expected return or value in a given state.
[0107] η: the learning rate. This is a hyperparameter that determines the step size of each parameter update. A higher learning rate may lead to faster convergence but may cause unstable training; a lower learning rate may lead to more stable training but the convergence speed may be slower.
[0108] π(a∣s;θπ): the probability of taking action a given state s and the current policy parameters θπ. πold(a∣s;θπ): the action probability under the old policy. This is the policy before the policy update. A′(s,a): the new advantage estimate. It tells us how much better it is to take action a in state s relative to the average.
[0109] ∈: the hyperparameter that determines the clipping range of the policy update. It prevents the new policy from deviating too far from the old policy, thus increasing the stability of training.
[0110] V(s;θv): the predicted value of the value network in state s and parameters θv.
[0111] V target is the target value function and can be calculated using the TD algorithm.
[0112] 6. Loop through steps 2 to 5: After convergence or reaching a predetermined number of iterations, end the algorithm. The resulting policy network can output optimal modification actions given the power grid device status and topology generated by the digital twin pre-trained language model.
[0113] In step S105, the power system pre-trained language model outputs a corresponding digital twin model of the power grid according to the pre-constructed prompt words.
[0114] A prompt is a set of instructions provided to the digital twin pre-trained language model, which enhances its capabilities in a customized way to adapt to the requirements of digital twin system construction. The role of the prompt can be achieved by setting initial rules for the digital twin pre-trained language model dialogue, thereby affecting subsequent interactions and the output generated by the digital twin pre-trained language model. Specifically, the prompt sets the context of the dialogue, indicates which information is important to the digital twin pre-trained language model, and the expected output form. For example, the prompt can guide the digital twin pre-trained language model to only generate code that follows a certain specific application scenario or digital twin system. At the same time, the prompt can also require the digital twin pre-trained language model to mark certain keywords or phrases in the generated document and provide additional information. Based on these guiding rules, the prompt can better assist in handling a large number of digital twin system construction tasks for power grid systems.
[0115] In addition, appropriate prompts can create new interaction modes, such as allowing the digital twin pre-trained language model to generate and provide reports related to power grid system engineering, or even simulate the operating environment of the power grid system. At the same time, the prompt also has the ability to self-expand, and can propose other prompts to collect more information or generate related tasks.
[0116] Construct the prompt words of the power system pre-trained language model according to the application scenario and / or construction task of the power system pre-trained language model; the prompt words set the context of the dialogue and the expected output form of the power system pre-trained language model.
[0117] The prompt words include: prompt words for device digital twin modeling mode, prompt words for unit-level digital twin model construction mode, prompt words for system-level digital twin model construction mode, prompt words for digital twin system operation mode, prompt words for optimization mode based on digital twin, and prompt words for operation and maintenance generation mode. The construction methods of various prompt words are introduced in detail below.
[0118] 1. Construction of prompt words for device digital twin modeling mode
[0119] In this mode, the digital twin pre-trained language model acts as a power grid equipment expert. The expert needs to understand various characteristics and behaviors of power grid equipment and be able to generate a digital twin model of the corresponding equipment based on real-time data. For example, when given real-time data of a transformer, the power grid equipment expert can generate a detailed digital twin model of the transformer, including information such as its internal structure, working status, and historical faults. The following are some prompt examples:
[0120] 1) Forward prompt: Digital twin pre-trained language model, please create a digital twin model of a 10kV distribution transformer based on the following real-time data (including information such as voltage, current, and frequency). If necessary, also create the code for the operating environment of this model and the MySQL script for me.
[0121] 2) Reverse prompt: From now on, I hope you will ask me questions to create a digital twin model of a 10kV distribution transformer. When you have enough information to build the model, create a Python script and output it to me.
[0122] 2. Construction of prompt words for the unit-level digital twin model building mode
[0123] In this mode, the digital twin pre-trained language model acts as a power grid design engineer. The power grid design engineer needs to understand how various parts of the power grid are connected together and be able to build a unit-level digital twin model based on the device-level digital twin models. For example, a digital twin model of a substation level can be built, including digital twin models of all devices such as transformers, switches, and busbars, and showing their connection relationships.
[0124] 1) Forward prompt: Digital twin pre-trained language model, I have a set of device-level digital twin models, including devices such as transformers, switches, and busbars. Can you help me build a digital twin model of a substation level? If necessary, also create the code for the operating environment of this model and the MySQL script for me.
[0125] 2) Reverse prompt: From now on, I hope you will ask me questions to create a digital twin model of a 10kV substation level. When you have enough information to build the model, create a Python script and output it to me.
[0126] 3. Construction of prompt words for the system-level digital twin model building mode
[0127] In this mode, the digital twin pre-trained language model acts as a power grid system planning engineer. The system planning engineer needs to have a global understanding of the entire power grid and be able to construct a digital twin model of the entire power grid system based on the unit-level digital twin models. For example, a digital twin model of the power grid system covering all links such as power generation, power transmission, power transformation, and power distribution can be constructed. The following are some prompting examples:
[0128] 1) Forward prompting: Digital twin pre-trained language model, the temperature of a certain transformer in my digital twin model exceeds the preset threshold. What measures should I take?
[0129] 2) Reverse prompting: From now on, I hope you will ask me questions to create a digital twin model of the global power grid system. When you have enough information to implement the model construction, create a Python script and output it to me.
[0130] 4. Construction of prompting words for the operation mode of the digital twin system
[0131] In this mode, the digital twin pre-trained language model acts as a power grid operation and maintenance personnel. The power grid operation and maintenance personnel need to monitor the operation status of the power grid system and make decisions based on the information provided by the digital twin model. For example, when the digital twin model shows that the temperature of a certain transformer exceeds the preset threshold, the power grid operation and maintenance personnel can decide to take measures such as load shedding and maintenance. The following are some prompting examples:
[0132] 1) Forward prompting: Digital twin pre-trained language model, the temperature of a certain transformer in my digital twin model exceeds the preset threshold. What measures should I take?
[0133] 2) Reverse prompting: From now on, I hope you will ask me questions to create countermeasures for the situation where the temperature of a certain transformer exceeds the preset threshold. When you have enough information to output the measures, create the corresponding code and output it to me.
[0134] 5. Construction of prompting words for the optimization mode based on digital twins
[0135] In this mode, the digital twin pre-trained language model acts as a power grid optimization engineer. The power grid optimization engineer needs to understand various performance indicators of the power grid system and be able to optimize the power grid system based on the information of the digital twin model. For example, by adjusting the operation parameters of the power grid system, the performance indicators such as the operation efficiency, stability, and reliability of the system can be optimized. The following are some prompting examples:
[0136] 1) Forward prompting: Digital twin pre-trained language model, based on our digital twin model, what suggestions do you have to optimize the operation efficiency and stability of the power grid system?
[0137] 2) Reverse prompt: From now on, I hope you will ask me questions to create suggestions for optimizing the operation efficiency and stability of the power grid system. When you have enough information to output suggestions, create a word file and output it to me.
[0138] 6. Construction of prompt words for the power grid dispatching generation mode
[0139] In this mode, the role of the digital twin pre-trained language model is a power grid dispatcher. The power grid dispatcher needs to understand the operation of the power grid system and be able to answer various dispatching questions. For example, when asked about the power supply situation in a certain area, the power grid dispatcher can give an answer based on the information of the digital twin model. The prompt examples are as follows:
[0140] 1) Forward prompt: Digital twin pre-trained language model, according to our digital twin model, what is the current power supply situation in a certain area?
[0141] 2) Reverse prompt: From now on, I hope you will ask me questions to create a report on the current power supply situation in a certain area. When you have enough information to write the report, create a report file and output it to me.
[0142] 7. Construction of prompt words for the operation and maintenance generation mode
[0143] In this mode, the role played by the digital twin pre-trained language model is a power grid operation and maintenance personnel. The power grid operation and maintenance personnel need to understand the operation and maintenance of power grid equipment and be able to answer various operation and maintenance questions. For example, when asked about the maintenance situation of a certain transformer, the power grid operation and maintenance personnel can give an answer based on the information of the digital twin model. The prompt examples are as follows:
[0144] 1) Forward prompt: Digital twin pre-trained language model, according to our digital twin model, what is the maintenance situation of a certain transformer?
[0145] 2) Reverse prompt: From now on, I hope you will ask me questions to create a report on the maintenance situation of a certain transformer. When you have enough information to generate the report, create a report file and output it to me.
[0146] To avoid the corresponding time and computing power resources being difficult to bear due to the update of the digital twin pre-trained language model through retraining, the present invention is based on a knowledge fine-tuning algorithm, sets a fine-tuning period and batch size, uses an optimizer to fine-tune the pre-trained language model of the power system, and updates the parameters of the pre-trained language model of the power system according to the fine-tuning results. The specific method is as follows:
[0147] 1. Load the pre-trained digital twin pre-trained language model and the Adam optimizer;
[0148] 2. Set appropriate fine-tuning epochs and batch size. If the dataset is very large, set the epochs to 5 - 10; otherwise, set them to 50 - 100. At the same time, if the GPU memory is large, the batch can be set to 128; otherwise, set it to 32 or 64.
[0149] 3. For each epoch, perform the following operations:
[0150] 1) Randomly select a batch of fine-tuning datasets;
[0151] 2) Input the data into the model to obtain the model output;
[0152] 3) Calculate the cross-entropy loss between the model output and the true value;
[0153] 4) Use the backpropagation algorithm to calculate the gradient of each model parameter;
[0154] 5) Use the optimizer to update the model parameters according to the gradient
[0155] The formula for the cross-entropy loss function is as follows:
[0156]
[0157] where N is the number of data points in the batch, y i is the true label of the i-th data point, and p(y i ) is the probability that the model predicts the label of the i-th data point.
[0158] 4. Evaluation and optimization: After fine-tuning, obtain the accuracy, recall rate, and F1 score on the test set. If the evaluation metrics do not meet the expectations, the training parameters such as the learning rate and batch size can be adjusted. The specific adjustment methods are as follows:
[0159] 1) Learning rate: If the loss of the model decreases too slowly, increase the learning rate; if the loss of the model decreases and then increases, decrease the learning rate;
[0160] 2) Batch size: If the memory occupation of the model is too large, decrease the batch size; if the model training is too slow, increase the batch size.
[0161] Example 2
[0162] Taking a 10 kV substation as an example, a typical construction step of a question-and-answer grid digital twin based on a digital twin pre-trained language model is as follows:
[0163] 1. Inform the digital twin pre-trained language model of the specific modeling purpose, the type of data, the amount of data, the target deployment environment and other related information;
[0164] 2. Ask the digital twin to pre-train the language model and generate the deployment environment code, including the target programming language and the corresponding SQL statements;
[0165] 3. Complete the deployment environment configuration according to the code;
[0166] 4. Ask the digital twin pre-trained language model to read the corresponding training data and output the modeling results. The general result is in the form of a model file of the weight matrix;
[0167] 5. Please output the specific deployment method and corresponding code of the digital twin pre-trained language model;
[0168] 6. Deploy the modeling results in the deployment environment;
[0169] 7. Please use the digital twin pre-trained language model to generate corresponding simulation operation data and test the corresponding digital twin system;
[0170] 8. Ask the digital twin pre-trained language model to give optimization suggestions for the corresponding system and output them in a specified format;
[0171] 9. Test the optimization suggestions given by the digital twin pre-trained language model on the digital twin system;
[0172] 10. Please use the digital twin pre-trained language model to evaluate the corresponding optimization results;
[0173] 11. Apply the corresponding optimization suggestions to the actual physical system;
[0174] 12. Please use the digital twin pre-trained language model to determine the corresponding optimization effect.
[0175] Example 3
[0176] Based on the same inventive concept, the present invention also provides a system for generating and building a digital twin model of a power grid. Figure 2 As shown, including:
[0177] The data feature extraction module 210 is used to obtain the full set of multimodal data in the power system; extract the data features of the full set of multimodal data based on the LIAMA2 model;
[0178] A model building module 220, for building a pre-trained language model of a power system for digital twins based on the full set of multimodal data and corresponding data features;
[0179] A training module 230, configured to train the pre-trained language model for the power system by using the contrastive learning loss InfoNCE, select a part of the data in the multi-modal data set as training data, embed the text into the corresponding image according to the data characteristics of the training data and the cosine similarity between the image and the text, obtain each node of the equipment operation of the power system, connect the major nodes, and output the equipment operation topology structure of the power system;
[0180] A testing module 240, configured to select a part of the data in the multi-modal data set as test data, perform dot product calculation to query the similarity between the image or text in the test data and the image or text in the multi-modal data set, and if the similarity meets a preset value, complete the test of the pre-trained language model for the power system.
[0181] A model output module 250, configured to enable the pre-trained language model for the power system to output a corresponding power grid digital twin model according to a pre-constructed prompt word.
[0182] Furthermore, it includes:
[0183] An update module, configured to set a fine-tuning period and a batch size based on a knowledge fine-tuning algorithm, use an optimizer to fine-tune the pre-trained language model for the power system, and update the parameters of the pre-trained language model for the power system according to the fine-tuning result.
[0184] A method and system for generating and constructing a power grid digital twin model provided by the present invention realizes high-precision and real-time modeling of physical devices of a new power system through the construction of a new power system digital twin hybrid model, which is described in detail as follows:
[0185] 1. Improvement in modeling accuracy and scope: First, the pre-trained language model for the power system digital twin is based on the multi-modal data set of the power system data, and completes the extraction and understanding of the overall knowledge of the power system, including the operation knowledge of individual devices, device sets, and systems at all levels. The traditional digital twin modeling mainly focuses on individual devices and device sets, and most of its modeling processes are based on partial relevant data of the devices and device sets. Therefore, its ability to learn and understand the implicit knowledge of a wide-area power grid is limited. Second, compared with the traditional method of first modeling devices and then assembling them into a digital twin system, learning the power system as a whole implicitly includes the extraction of knowledge about the relationships between a large number of devices, and realizes a more essential and comprehensive understanding of the relevant knowledge.
[0186] 2. Reduction in modeling costs: Traditional methods require re-training and debugging for the modeling of different individual devices, device sets, and system levels at all levels. The time and computing power costs are high and difficult to reduce. After the training of the digital twin pre-trained language model for the power system proposed in this patent is completed, only appropriate prompt words need to be used to achieve the output of the generative model, saving a large amount of costs for re-training, debugging, deployment, and operation. At the same time, this patent provides a knowledge fine-tuning method with low time and computing power costs. This method can efficiently update the knowledge of the existing model without the need to re-train the complete pre-trained language model.
[0187] 3. Endowing business departments with modeling capabilities: The purpose of building a digital twin system is to solve practical business problems. However, business departments are often restricted by factors such as knowledge and computing power and cannot complete the construction of the digital twin system independently. The method proposed in this patent provides multi-scale and multi-type digital twin on-demand modeling capabilities based on natural language Q&A for various business departments at all levels of the power system, which helps to greatly improve the actual application effect of grid digitization.
[0188] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that the specific implementation manners of the present invention can still be modified or equivalently replaced. Any modification or equivalent replacement without departing from the spirit and scope of the present invention shall be covered by the scope of the claims of the present invention.
Claims
1. A method for generatively constructing a digital twin model of a power grid, characterized in that Including: Obtain the full set of multimodal data in the power system; Extract the data features of the full set of multimodal data based on the LlAMA2 model; Based on the full set of multimodal data and the corresponding data features, construct a pre-trained language model for the power system oriented to digital twin; Train the pre-trained language model of the power system using the contrastive learning loss InfoNCE. Select part of the data in the full set of multimodal data as training data. Embed the text into the corresponding image according to the data features of the training data and the cosine similarity between the image and the text, obtain each node of the equipment operation in the power system, connect the nodes, and output the topological structure of the equipment operation in the power system; Select part of the data in the full set of multimodal data as test data, use dot product calculation to query the similarity between the image or text in the test data and the image or text in the full set of multimodal data. If the similarity meets the preset value, complete the test of the pre-trained language model of the power system; The pre-trained language model of the power system outputs the corresponding power grid digital twin model according to the pre-constructed prompt words; After the test of the pre-trained language model of the power system, it also includes: Improve the prediction ability of the pre-trained language model of the power system by constructing a reward model, specifically including: Define the feedback rules for human trainers. The standard interval for the device operating state is [S lower , S upper . Then the rules are as follows: If the device state s satisfies S lower ≤ s ≤ S upper , then the human feedback value H = 1; If the device state s does not satisfy S lower ≤ s ≤ S upper , then the human feedback value H = -1; Among them, Slower and Supper respectively represent the lower and upper limits of the interval, and s represents the current state of the device; Select the deep Q network as the reinforcement learning model and initialize a Q network: Q(s,a;θ), where s is the state, a is the action, and θ is the network parameter; At each time step t, select the action a according to the current state s and the Q network, and then obtain the new state s' and the human feedback value H through interaction with the environment; Construct a reward function. The first part uses the square of the difference between the human feedback value H and the predicted value Q(s,a;θ) of the Q network as the reward function R. The second part uses an additional reward based on the device state. If the device state is within the normal range, give a positive reward; otherwise, if the device state is abnormal, give a negative reward. The completed reward function is: R = (H - Q(s, a; θ)) 2 + β * |s - s'| 2 Among them, β is a hyperparameter used to balance the importance of the two parts, and s and s' are the device states before and after the action respectively; Use experience replay and target network strategy to update the Q network. In experience replay, store the state, action, reward, and new state of each step in a replay buffer D, that is, D ={(s,a,R,s')}. At each time step, randomly take out a batch of data from D and use this data and the target network to calculate the loss function and update the parameters θ of the Q network. The specific loss function is: L(θ) = E {s,a,R,s′} [(R + γmax {a′} Q(s′, a′; θ′) - Q(s, a; θ)) 2 Among them, γ is the discount factor, and θ' is the parameter of the target network. At each time step, use part of the θ to update θ' to keep the target network stable.
2. The method according to claim 1, wherein It also includes: Based on the knowledge fine-tuning algorithm, set the fine-tuning period and batch size, use the optimizer to fine-tune the pre-trained language model of the power system, and update the parameters of the pre-trained language model of the power system according to the fine-tuning results.
3. The method according to claim 1, characterized in that, Extract data features from the multi-modal data set based on the LlAMA2 model, including: Extract the image features of the graph data in the multi-modal data set by introducing a graph attention layer into the LlAMA2 model; Extract the text features of the text data in the multi-modal data set by the LlAMA2 model; Extract the topological features of the topological data in the multi-modal data set by introducing a graph attention layer into the LlAMA2 model.
4. The method according to claim 1, characterized in that Use the contrastive learning loss InfoNCE to train the pre-trained language model for the power system. Select part of the data in the multi-modal data set as training data, and embed the text into the corresponding image according to the data features of the training data and the cosine similarity between the image and the text, to obtain each node of the equipment operation in the power system, including: According to the contrastive learning loss, maximize the cosine similarity of the image and text embeddings of the correct matching pairs, and at the same time minimize the cosine similarity of the wrong matching pairs, embed the text into the corresponding image, to obtain each node of the equipment operation in the power system. The contrastive learning loss InfoNCE has the following formula: Among them, Ltotal represents the contrastive learning loss, and z Ij represents the embedding vector of image j, and z Tj represents the embedding vector of text j, p Tj and p Ij respectively represent the positive sample embedding vectors of text j and image j, Q T and Q I respectively represent the queues of negative sample embedding vectors of text and image. τ is a temperature parameter that controls the fineness of the embedding vectors.
5. The method according to claim 1, wherein Use dot product calculation to query the similarity between the image or text in the test data and the image or text in the multi-modal data set, including: Use dot product calculation to query the similarity between the image or text in the test data and the image or text in the multi-modal data set, so as to retrieve relevant data. The specific formula is: Where, similarity represents the cosine similarity between two embedding vectors, A and B are the embedding vectors of the query image or text and the data in the multi-modal data set respectively, and ||A||2 and ||B||2 represent the L2 norms of A and B respectively.
6. The method according to claim 1, characterized in that It also includes: Construct prompt words for the pre-trained language model of the power system according to the application scenario and / or construction task of the pre-trained language model of the power system; The prompt words set the context of the conversation and the expected output form of the pre-trained language model of the power system.
7. The method according to claim 6, characterized in that, The prompt words include: prompt words for the device digital twin modeling mode, prompt words for the unit-level digital twin model construction mode, prompt words for the system-level digital twin model construction mode, prompt words for the digital twin system operation mode, prompt words for the optimization mode based on digital twins, and prompt words for the operation and inspection generation mode.
8. A generative construction system for a power grid digital twin model, characterized in that, It includes: A data feature extraction module for obtaining the multi-modal data set in the power system; Extract data features from the multi-modal data set based on the LlAMA2 model; A model construction module for constructing a pre-trained language model for the power system facing digital twins based on the multi-modal data set and the corresponding data features; A training module for training the pre-trained language model of the power system using the contrastive learning loss InfoNCE. Select part of the data in the multi-modal data set as training data, embed the text into the corresponding image according to the data features of the training data and the cosine similarity between the image and the text, obtain each node of the equipment operation in the power system, connect the nodes, and output the equipment operation topological structure of the power system; A test module for selecting part of the data in the multimodal data set as test data, calculating using the dot product to query the similarity between the image or text in the test data and the image or text in the multimodal data set, and if the similarity meets the preset value, completing the test of the pre-trained language model for the power system; A model output module for the pre-trained language model of the power system to output the corresponding power grid digital twin model according to the pre-constructed prompt words; After the test of the pre-trained language model of the power system, it further includes: Improving the prediction ability of the pre-trained language model of the power system by constructing a reward model, specifically including: Define the feedback rules for human trainers. The standard interval for the device operating state is [S lower , S upper , and the rules are as follows: If the device state s satisfies S lower ≤ s ≤ S upper , then the human feedback value H = 1; If the device state s does not satisfy S lower ≤ s ≤ S upper , then the human feedback value H = -1; Where Slower and Supper represent the lower and upper limits of the interval respectively, and s represents the current state of the device; Select the deep Q network as the reinforcement learning model and initialize a Q network: Q(s,a;θ), where s is the state, a is the action, and θ is the network parameter; At each time step t, select the action a according to the current state s and the Q network, and then obtain the new state s' and the human feedback value H by interacting with the environment; Construct a reward function. The first part uses the square of the difference between the human feedback value H and the predicted value Q(s,a;θ) of the Q network as the reward function R. The second part uses an additional reward based on the device state. If the device state is within the normal range, a positive reward is given; otherwise, if the device state is abnormal, a negative reward is given. The completed reward function is: R = (H - Q(s,a;θ)) 2 + β * |s - s'| 2 Where β is a hyperparameter used to balance the importance of the two parts, and s and s' are the device states before and after the action respectively; Use experience replay and target network strategy to update the Q network. In experience replay, store the state, action, reward, and new state of each step in a replay buffer D, that is, D ={(s,a,R,s')}. At each time step, randomly extract a batch of data from D and use this data and the target network to calculate the loss function and update the parameters θ of the Q network. The specific loss function is: L(θ) = E {s,a,R,s′} [(R + γmax {a′} Q(s′, a′; θ′) - Q(s, a; θ)) 2 Where γ is the discount factor and θ' is the parameter of the target network. At each time step, use the selected part of θ to update θ' to keep the target network stable.
9. The system according to claim 8, wherein It includes: An update module for fine-tuning the pre-trained language model of the power system based on the knowledge fine-tuning algorithm, setting the fine-tuning period and batch size, using an optimizer to fine-tune the pre-trained language model of the power system, and updating the parameters of the pre-trained language model of the power system according to the fine-tuning results.
Citation Information
Patent Citations
Sweeping robot obstacle avoidance method based on deep reinforcement learning
CN116439626A
Method and Apparatus for Training Information Adjustment Model of Charging Station, and Storage Medium
US20230229913A1