Power distribution network state estimation method and system based on fine tuning large language model

By fine-tuning the large language model and combining measurement and topology information, the problem of insufficient universality and robustness of traditional methods in distribution network state estimation is solved, achieving high accuracy and stability in state estimation, which is applicable to distribution networks of different sizes and structures.

CN120952088APending Publication Date: 2025-11-14SOUTHEAST UNIV +3
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511063746.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Traditional deep learning methods lack versatility and scalability in power distribution network state estimation. They require a large amount of labeled data and have insufficient generalization ability under abnormal operating conditions. They also lack a deep understanding and reasoning ability of power system operation knowledge, resulting in insufficient accuracy and robustness when data is incomplete or noisy.

Method used

A fine-tuned large language model is used to construct a distribution network state estimation dataset by combining measurement information, topology information, and constraint information. The GPT-2 model is used for training and validation. The text data is processed through a multi-head self-attention mechanism and a location feedforward network to generate state estimation results.

Benefits of technology

It improves the accuracy and stability of distribution network condition estimation, exhibiting state-of-the-art performance indicators, including reduced mean absolute error, root mean square error, and voltage over-limit rate, and is adaptable to distribution networks of different sizes and structures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120952088A_ABST
    Figure CN120952088A_ABST
Patent Text Reader

Abstract

The invention discloses a power distribution network state estimation method and system based on a fine tuning large language model, and belongs to the field of power system state estimation. The method comprises the steps of collecting node data of a power distribution network, constructing a power distribution network state estimation data set which is suitable for a large language model and comprises measurement information, topological information and constraint information, and dividing the data set into a training set, a verification set and a test set; training the large language model by using the training set data, finely adjusting parameters of the large language model to obtain a power distribution network state estimation model, and verifying the power distribution network state estimation model by using the verification set; inputting the test set data into the power distribution network state estimation model to obtain a power distribution network state estimation result; the method comprises the following steps: integrating unstructured power distribution system information with measurement data; the effectiveness and the robustness of the method are verified on an IEEE (Institute of Electrical and Electronic Engineers) 33-node power distribution network; compared with the prior art, the method provided by the invention shows the most advanced performance on all verification indexes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of power system state estimation, specifically relating to a distribution network state estimation method and system based on fine-tuning a large language model. Background Technology

[0002] Accurate distribution system state estimation plays a fundamental role in ensuring the safe operation and coordinated control of modern complex power systems. Traditional weighted least squares (WLS)-based distribution system state estimation methods suffer from accuracy degradation when dealing with distributed renewable energy grid integration. Deep learning methods have emerged as a promising solution to the challenges of distribution system state estimation. Recent research has restructured state estimation as a sequence-to-sequence problem, solving it using deep learning algorithms. This innovative framework enables neural networks to learn the mapping between historical measurements of monitored nodes and unmeasured system states. Furthermore, researchers have implemented various deep learning model frameworks, particularly graph neural networks and physical information neural networks. These advanced models incorporate grid topology information and physical constraints, thereby optimizing neural network performance, resulting in improved accuracy and computational efficiency.

[0003] The recent emergence of ChatGPT has garnered widespread attention due to the exceptional accuracy of large language models (LLMs) across various sequence generation tasks. Several research institutions have open-sourced their pre-trained large models, trained on billions of data points, and validated their fundamental capabilities in causal inference. Current research has begun fine-tuning these pre-trained LLMs on domain-specific datasets for sequence-to-sequence tasks, achieving state-of-the-art performance in multiple areas, including protein structure prediction, wind speed prediction, and load prediction.

[0004] However, traditional deep learning methods often require specialized network structures designed for specific grid topologies and operating scenarios when dealing with distribution network state estimation. This limits the model's generality and scalability, making it difficult to adapt to distribution networks of different sizes and structures. Furthermore, traditional deep learning models typically require large amounts of labeled data for training, but acquiring distribution network operating data is costly, and samples under abnormal operating conditions are scarce, resulting in insufficient generalization ability when facing new operating scenarios. In addition, existing methods lack a deep understanding and reasoning ability regarding power system operation knowledge, making it difficult to maintain the accuracy and robustness of state estimation under conditions of incomplete or noisy data. To address these technical problems, this invention proposes a distribution network state estimation method based on fine-tuning a large language model. Summary of the Invention

[0005] To address the shortcomings of existing technologies, the present invention aims to provide a distribution network state estimation method and system based on a fine-tuned large language model, thereby solving the problems in the prior art.

[0006] The objective of this invention can be achieved through the following technical solutions:

[0007] A distribution network state estimation method based on fine-tuning a large language model includes the following steps:

[0008] Data from distribution network nodes is collected to construct a distribution network state estimation dataset suitable for large language models, which includes measurement information, topology information, and constraint information. The dataset is then divided into training set, validation set, and test set.

[0009] The large language model is trained using the training set data, the parameters of the large language model are fine-tuned to obtain the distribution network state estimation model, and the distribution network state estimation model is validated using the validation set.

[0010] The test set data is input into the distribution network state estimation model to obtain the distribution network state estimation results.

[0011] Furthermore, the measurement information includes: the node's voltage, phase angle, active power, and reactive power;

[0012] The topology information includes: the topological relationship between the nodes that need to be estimated and the nodes with measured values;

[0013] The constraint information includes: the upper and lower limits of voltage regulation for each node itself.

[0014] Furthermore, the large language model is GPT-2, which is composed of multiple stacked Transformer decoder blocks. Each Transformer decoder block includes a multi-head self-attention mechanism, a position feedforward network, and a layer normalization layer.

[0015] Furthermore, after the text data is input into GPT-2, GPT-2 uses a byte-pair encoding tokenizer to segment the text data into a series of tokens; these tokens are mapped to word embedding vectors and combined with positional encoding, and then processed sequentially through multiple Transformer decoder blocks;

[0016] Each Transformer decoder block uses a multi-head self-attention mechanism to capture long-range dependencies in the text data and applies a non-linear transformation through a position feedforward network. Layer normalization modules are set after the multi-head self-attention mechanism and the position feedforward network to standardize the feature vectors. Finally, a softmax function is used in the output layer to generate the probability distribution of the next label.

[0017] Furthermore, the process of fine-tuning the parameters of the large oracle model is as follows:

[0018]

[0019]

[0020] Where, θ new D represents the adjusted parameter. DSSE Let represent the distribution network state estimation dataset, θ represent the original pre-trained parameters, L represent the loss function, η is the learning rate, and P represents the probability calculation.

[0021] Furthermore, the indicators for verifying the distribution network state estimation model include Mean Absolute Error (MAE), Root Mean Square Error (RMSE), Mean Absolute Percentage Error (MAPE), and Voltage Limit Exceedance Rate (VVR). 7. A distribution network state estimation system based on a fine-tuned large language model, characterized in that it includes:

[0022] Dataset construction module: Collects data from distribution network nodes, constructs a distribution network state estimation dataset suitable for large language models and including measurement information, topology information and constraint information, and divides the dataset into training set, validation set and test set;

[0023] Model building module: Train the large language model using training set data, fine-tune the parameters of the large language model to obtain the distribution network state estimation model, and use the validation set to validate the distribution network state estimation model;

[0024] State estimation module: Input the test set data into the distribution network state estimation model to obtain the distribution network state estimation results.

[0025] A computer storage medium storing a readable program that, when executed, instructs a computing device to perform a power distribution network state estimation method based on a fine-tuned large language model, as described above.

[0026] An electronic device includes: a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other through the communication bus;

[0027] The memory is used to store at least one executable instruction, which causes the processor to perform an operation corresponding to the above-described distribution network state estimation method based on a fine-tuned large language model.

[0028] A computer program product includes computer instructions that instruct a computing device to perform operations corresponding to a power distribution network state estimation method based on a fine-tuned large language model as described above.

[0029] The beneficial effects of this invention are:

[0030] This invention proposes a distribution network state estimation method based on a fine-tuned large language model, integrating unstructured distribution system information with measurement data. The effectiveness and robustness of the proposed method are verified through a comprehensive case study on an IEEE 33-node distribution network. Compared with existing methods, the method of this invention exhibits state-of-the-art performance across all verification metrics. Attached Figure Description

[0031] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0032] Figure 1 This is a schematic diagram of the power distribution network state estimation principle of the present invention;

[0033] Figure 2 This is a topology diagram of the IEEE 33-node distribution network system.

[0034] Figure 3 This is a comparison chart of the voltage state estimation results of the distribution network in this invention;

[0035] Figure 4 This is a comparison chart of the phase angle state estimation results of the distribution network in this invention. Detailed Implementation

[0036] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0037] Example 1

[0038] like Figure 1 As shown, a distribution network state estimation method based on fine-tuning a large language model includes the following steps:

[0039] S1. Collect data from distribution network nodes, construct a distribution network state estimation dataset suitable for large language models that includes measurement information, topology information, and constraint information, and divide the dataset into training set, validation set, and test set;

[0040] Large language models are designed for processing text data; therefore, the distribution network state estimation dataset constructed in this embodiment is in text data format. This invention generates the distribution network state estimation dataset and corresponding labels by combining measured values ​​with unstructured network topology information. Specifically, for each node requiring estimation, the distribution network state estimation dataset contains three key aspects:

[0041] 1) Measurement information: including voltage, phase angle, active power and reactive power of all nodes with measurement devices installed.

[0042] 2) Topology information: This includes the topological relationships between the nodes that need to be estimated and the nodes with measured values.

[0043] 3) Constraint information: including the upper and lower limits of voltage regulation for each node itself.

[0044] The labels only contain the voltage magnitude or phase angle of the node that needs to be estimated.

[0045] In this embodiment, the power distribution network state estimation dataset is divided into a training set, a validation set, and a test set in a ratio of 8:1:1.

[0046] S2, use the training set data to train the large language model, fine-tune the model parameters to obtain the distribution network state estimation model, and use the validation set to validate the distribution network state estimation model;

[0047] First, the traditional state estimation process based on the Weighted Least Squares (WLS) method can be described as an optimization problem. Specifically, based on the power flow equations of the power system, the relationship between the state variable v and the measured value z can be expressed as:

[0048] z = h(v) + e (1)

[0049] Among them, h i (·) represents the power flow equation, and e represents the measurement error.

[0050] In the WLS-based state estimation method, the state variable v can be obtained by solving the following optimization problem:

[0051]

[0052] Where: w i Let h be the weight of the i-th measurement, m be the total number of state variables to be solved, and J(v) be the objective function; i (v) represents the power flow equations related to the i-th variable, z i Let be the i-th measurement value.

[0053] This embodiment proposes that when using deep learning models to handle state estimation problems, it can be constructed as a supervised sequence-to-sequence training paradigm. The deep learning method used in this embodiment has the following technical features: 1) it establishes a direct mapping relationship between measurement values ​​and state variables; 2) it eliminates the need for iteratively solving complex optimization problems; and 3) it can automatically learn the intrinsic features of the system. The mathematical form of the direct mapping relationship between measurement values ​​and state variables is as follows:

[0054] v = f θ (z) (3)

[0055] Among them, f θ (·) is a deep learning model with parameter θ.

[0056] Based on the loss function, the network parameters are updated through backpropagation, and its expression is:

[0057]

[0058] Where ||·||2 represents the L2 norm, θ t+1 Here are the updated parameters, t is the number of training samples, η is the learning rate, and L(θ) is the training rate. t Let θ be the loss function. t These are the parameters during the t-th training iteration.

[0059] In this embodiment, the GPT-2 model is used as the base model for the distribution network state estimation task. GPT-2 is a pre-trained large-scale language model developed and open-sourced by OpenAI. This model consists entirely of stacked Transformer decoder blocks, omitting the encoder component found in traditional Transformers. The model consists of multiple identical decoder blocks (12, 24, or 36 layers), each of which includes three sub-layers: a multi-head self-attention mechanism, a position feedforward network, and a layer normalization layer.

[0060] For a given input text, the GPT-2 model uses a byte-pair encoding (BPE) tokenizer to segment the text into a series of tokens. These tokens are mapped to word embedding vectors and combined with positional encodings, then processed sequentially by multiple decoders. Each decoder block uses a multi-head self-attention mechanism to capture long-range dependencies in the sequence and applies a nonlinear transformation through a positional feedforward network. Layer normalization modules are set after both the multi-head self-attention mechanism and the positional feedforward network to normalize the feature vectors, stabilize gradient propagation during training, accelerate model convergence, and improve the accuracy and stability of power distribution network state parameter prediction. Finally, the model uses a softmax function at the output layer to generate the probability distribution of the next token.

[0061] In the distribution network state estimation task involved in this embodiment, it is also constructed as a supervised learning process. Unlike the method of training with randomly initialized weights, this embodiment fine-tunes the parameters of the pre-trained model on the distribution network state estimation dataset to specifically adjust the model to solve the state estimation problem. The process is as follows:

[0062]

[0063]

[0064] Where, θ new D represents the adjusted parameter. DSSE Let represent the distribution network state estimation dataset, θ represent the original pre-trained parameters, L represent the loss function, η is the learning rate, and P represents the probability calculation.

[0065] After training and tuning the Big Prophet Model (GPT-2), a distribution network state estimation model can be obtained. This model is then validated using validation set data. The metrics involved in the validation process include Mean Absolute Error (MAE), Root Mean Square Error (RMSE), Mean Absolute Percentage Error (MAPE), and Voltage Violation Rate (VVR), defined as follows:

[0066]

[0067] Where N is the sample size, Y i Y represents the i-th prediction result. i,gt Represents the corresponding true value, VVR∈(0,1) represents the voltage over-limit rate, n v This indicates the number of samples that exceeded the voltage limit.

[0068] S3. Input the test set data into the distribution network state estimation model to obtain the distribution network state estimation results.

[0069] Based on a similar inventive concept, embodiments of the present invention also provide a computer storage medium storing a readable program that, when run by a processor, can execute the aforementioned method for estimating the state of a distribution network based on a fine-tuned large language model.

[0070] Based on a similar inventive concept, this invention provides an electronic device, including: a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other through the communication bus;

[0071] The memory is used to store at least one executable instruction, which causes the processor to perform the operation corresponding to the above-described distribution network state estimation method based on a fine-tuned large language model.

[0072] Based on a similar inventive concept, embodiments of the present invention also provide a computer program product, including computer instructions, which instruct a computing device to perform the operations corresponding to the above-described distribution network state estimation method based on a fine-tuned large language model.

[0073] Example 2

[0074] This embodiment selects the IEEE 33-bus distribution system as a case study to verify the effectiveness of the state estimation method of the present invention;

[0075] Following the measurement point selection method described in the references, phasor measurement units (PMUs) were installed at nodes 3, 6, and 28 to measure voltage amplitude and phase angle. Additionally, equipment was installed on lines 10-11, 12-13, and 24-25 to measure active and reactive power, such as... Figure 2 As shown in Table 1, taking node 5 as an example, detailed information about a training data sample and its corresponding true value is presented. All parameters enclosed in {·} serve as placeholders for the measured values.

[0076] Table 1 Data Input and Label Examples

[0077]

[0078] Data generation was performed using pandapower (2.13.1) in a Python 3.9 environment. To simulate the inherent measurement uncertainties of PMUs, Gaussian noise with a standard deviation of 0.5 was introduced into the measurements. Different physical quantities were assigned different error proportions, as detailed in Table 2. By systematically varying the load range, a dataset containing 15,000 samples was generated. The dataset was divided into three subsets: 80% for training, 10% for validation, and 10% for testing. All computations were performed on Nvidia RTX 4090 GPUs.

[0079] Table 2 Installation locations and standard deviations of different measuring devices

[0080]

[0081] Since the output of the large language model is in text format natural language, the numerical values ​​of the state estimation results are extracted from the natural language and compared with the actual measurement values ​​to calculate the accuracy of the state estimation in order to analyze the model's performance. The robustness of the estimation results is evaluated by the voltage limit violation rate (VVR). In addition, the mean absolute error (MAE), root mean square error (RMSE), and mean absolute percentage error (MAPE) are also used to measure the accuracy of the state estimation results, defined as follows:

[0082]

[0083] Where N is the sample size, Y i Y represents the i-th prediction result. i,gt Represents the corresponding true value, VVR∈(0,1) represents the voltage over-limit rate, n v This indicates the number of samples that exceeded the voltage limit.

[0084] Example 3

[0085] The difference between this embodiment and Embodiment 2 is that it uses the smallest version of the GPT-2 model family, which contains 117 million trainable parameters and 12 decoder units. This model uses a vocabulary of 50,257 tokens, each token mapped to a unique index, and then converted into a 768-dimensional word embedding vector. The model implementation is accessible through the Hugging Face Transformers library using the keyword: openai-community / gpt2.

[0086] To verify the effectiveness of the proposed distribution network state estimation method, this embodiment conducted a comparative experiment using three methods: Weighted Least Squares (WLS), Multilayer Perceptron (MLP), and Graph Neural Network (GNN). The MLP model contains five linear layers with dimensions [32, 64, 128, 64, 32]. Similarly, the GNN model contains five Graph Convolutional Network (GCN) layers with the same dimensions [32, 64, 128, 64, 32]. Both models used the Adam optimizer with a mean squared error (MSE) loss function and a learning rate of 0.01. For voltage and phase angle estimation, this invention independently trained two models based on a pre-trained GPT-2 model on the voltage estimation dataset and phase angle estimation dataset, respectively. Each model was trained for 20 epochs with a batch size of 16 and a learning rate of 1×10⁻⁶. -5 .

[0087] Table 3 shows that the GPT-2-based method of this invention significantly outperforms the MLP and GNN methods. Specifically, compared with the GNN method (the second-best performer), the method of this invention improves MAPE by 64.7% and 65.6% in voltage and phase angle estimation, respectively. To clearly demonstrate the performance of different methods, the following tables are provided: Figure 3 and Figure 4 The estimated voltage and phase angle results are visualized. Furthermore, this embodiment evaluates the over-limit rate of voltage violations across all state estimates to verify the compliance of the safety information. A 0% violation rate demonstrates the effectiveness of the text-based constraint implementation.

[0088] Table 3 Comparison of the accuracy of state estimation by various methods

[0089]

[0090]

[0091] Example 4

[0092] In this embodiment, a distribution network state estimation system based on a fine-tuned large language model is proposed, specifically including:

[0093] Dataset construction module: Collects data from distribution network nodes, constructs a distribution network state estimation dataset suitable for large language models and including measurement information, topology information and constraint information, and divides the dataset into training set, validation set and test set;

[0094] Model building module: Train the large language model using training set data, fine-tune the model parameters to obtain the distribution network state estimation model, and validate the distribution network state estimation model using the validation set;

[0095] State estimation module: Input the test set data into the distribution network state estimation model to obtain the distribution network state estimation results.

[0096] The methods of the present invention can be implemented in hardware, firmware, or as software or computer code that can be stored in a recording medium (such as a CD-ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or as computer code originally stored on a remote recording medium or a non-transitory machine-readable medium and subsequently stored on a local recording medium, downloaded via a network. Thus, the methods described herein can be processed by software stored on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA). It is understood that the computer, processor, microprocessor controller, or programmable hardware includes storage components (e.g., RAM, ROM, flash memory, etc.) capable of storing or receiving software or computer code that, when accessed and executed by the computer, processor, or hardware, implements the methods described herein. Furthermore, when a general-purpose computer accesses the code used to implement the methods shown herein, the execution of the code transforms the general-purpose computer into a dedicated computer for performing the methods shown herein.

[0097] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention.

Claims

1. A method for estimating the state of a distribution network based on a fine-tuned large language model, characterized in that, Includes the following steps: Data from distribution network nodes is collected to construct a distribution network state estimation dataset suitable for large language models, which includes measurement information, topology information, and constraint information. The dataset is then divided into training set, validation set, and test set. The large language model is trained using the training set data, the parameters of the large language model are fine-tuned to obtain the distribution network state estimation model, and the distribution network state estimation model is validated using the validation set. The test set data is input into the distribution network state estimation model to obtain the distribution network state estimation results.

2. The distribution network state estimation method based on a fine-tuned large language model according to claim 1, characterized in that, The measurement information includes: node voltage, phase angle, active power, and reactive power; The topology information includes: the topological relationship between the nodes that need to be estimated and the nodes with measured values; The constraint information includes: the upper and lower limits of voltage regulation for each node itself.

3. The distribution network state estimation method based on a fine-tuned large language model according to claim 1, characterized in that, The large language model is GPT-2, which is composed of multiple stacked Transformer decoder blocks. Each Transformer decoder block includes: a multi-head self-attention mechanism, a position feedforward network, and a layer normalization layer.

4. The distribution network state estimation method based on a fine-tuned large language model according to claim 3, characterized in that, After text data is input into GPT-2, GPT-2 uses a byte-pair encoding tokenizer to segment the text data into a series of tokens; these tokens are mapped to word embedding vectors and combined with positional encoding, and then processed sequentially through multiple Transformer decoder blocks; Each Transformer decoder block uses a multi-head self-attention mechanism to capture long-range dependencies in the text data and applies a non-linear transformation through a position feedforward network. Layer normalization modules are set after the multi-head self-attention mechanism and the position feedforward network to standardize the feature vectors. Finally, a softmax function is used in the output layer to generate the probability distribution of the next label.

5. The distribution network state estimation method based on a fine-tuned large language model according to claim 1, characterized in that, The process of fine-tuning the parameters of the large oracle model is as follows: Where, θ new D represents the adjusted parameter. DSSE Let represent the distribution network state estimation dataset, θ represent the original pre-trained parameters, L represent the loss function, η is the learning rate, and P represents the probability calculation.

6. The distribution network state estimation method based on a fine-tuned large language model according to claim 1, characterized in that, The metrics used to validate the power distribution network state estimation model include mean absolute error (MAE), root mean square error (RMSE), mean absolute percentage error (MAPE), and voltage overrun rate (VVR).

7. A distribution network state estimation system based on a fine-tuned large language model, characterized in that, include: Dataset construction module: Collects data from distribution network nodes, constructs a distribution network state estimation dataset suitable for large language models and including measurement information, topology information and constraint information, and divides the dataset into training set, validation set and test set; Model building module: Train the large language model using training set data, fine-tune the parameters of the large language model to obtain the distribution network state estimation model, and use the validation set to validate the distribution network state estimation model; State estimation module: Input the test set data into the distribution network state estimation model to obtain the distribution network state estimation results.

8. A computer storage medium storing a readable program, characterized in that, When the program runs, it can instruct the computing device to execute a distribution network state estimation method based on a fine-tuned large language model as described in any one of claims 1-6.

9. An electronic device, characterized in that, include: The processor, memory, communication interface, and communication bus are provided, wherein the processor, memory, and communication interface communicate with each other via the communication bus. The memory is used to store at least one executable instruction, which causes the processor to perform the operation corresponding to the power distribution network state estimation method based on fine-tuning large language model as described in any one of claims 1-6.

10. A computer program product comprising computer instructions, characterized in that, The computer instructions instruct the computing device to perform the operation corresponding to the power distribution network state estimation method based on fine-tuning large language model as described in any one of claims 1-6.

Citation Information

Cited By

  • Power distribution network man-machine interaction regulation and control method and device based on edge side large model

    CN121216483A