Power load model parameter identification method and device based on reinforcement learning and storage medium

By building a multi-type load model library and a directed collaborative reinforcement learning environment, combined with initial value quality assessment and cross-model parameter migration mechanism, the problems of multi-model adaptation and initial value dependence in power load model parameter identification are solved, and efficient and robust parameter identification is achieved to adapt to complex power grid conditions.

CN120675054APending Publication Date: 2025-09-19NARI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510790772.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing power load model parameter identification methods have limitations, strong initial value dependence, serious waste of training resources, and insufficient consideration of physical constraints. They are difficult to adapt to different types of load models and achieve efficient and robust parameter identification.

Method used

Build a multi-type load model library, adopt a directed and collaborative reinforcement learning environment, achieve multi-model adaptation and dynamic stability through initial value quality assessment and heterogeneous environment matrix architecture, use reinforcement learning agents for parameter identification, and combine two-dimensional reward functions and cross-model parameter migration mechanisms to optimize the parameter identification process.

Benefits of technology

Automatic model identification and parameter optimization are achieved under unknown load types, which improves the consistency and reproducibility of identification results, enhances the robustness and engineering applicability of the system, and reduces computational complexity and training time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120675054A_ABST
    Figure CN120675054A_ABST
Patent Text Reader

Abstract

The invention discloses a power load model parameter identification method and device based on reinforcement learning and a storage medium, and belongs to the technical field of power system modeling and simulation, and the method comprises the steps: building a multi-type load model library according to a power system modeling standard and an actual application scene demand, determining a to-be-identified parameter vector of each model, a physical constraint range of the to-be-identified parameter vector and an input parameter of the load model; constructing a directional and collaborative reinforcement learning environment, and performing directional agent training when a target model is specified and collaborative agent training when the model is not specified; and performing parameter identification on the specified or unspecified model by using the trained reinforcement learning agent, and outputting a parameter identification result. According to the method, the problems that a traditional parameter identification method is prone to falling into local optimum, sensitive to initial values, low in calculation efficiency, poor in robustness, difficult to effectively process strong nonlinearity and dynamic characteristics and the like can be solved, and automatic and high-precision identification of load model parameters is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method and device for identifying parameters of an electric load model, belonging to the technical field of power system modeling and simulation, and in particular to a method, device and storage medium for identifying parameters of an electric load model based on reinforcement learning. Background Art

[0002] Power load models are the foundation of power system transient simulation and analysis. Model accuracy directly impacts the reliability of simulation results, which in turn influences the quality of decision-making for grid planning, operation, and control. Load models typically contain multiple parameters to be identified. These parameters are difficult to measure directly and require dynamic response data such as voltage, current, active power, and reactive power at the load point to identify them.

[0003] Traditional load parameter online identification methods mainly include the following categories:

[0004] 1) Gradient-based methods (such as the least squares method and its variants): They are sensitive to the initial values ​​of parameters and easily fall into local optimal solutions; they require the convexity of the objective function and have difficulty handling strongly nonlinear models; and calculating gradients may be complex or infeasible.

[0005] 2) Heuristic optimization algorithms (e.g., genetic algorithms, particle swarm optimization): They have strong global search capabilities but are typically computationally expensive (requiring a large number of individual evaluations and iterations); parameter settings (e.g., population size, crossover / mutation probability) have a significant impact on performance and are difficult to tune; convergence can be slow, especially in high-dimensional parameter spaces; and online identification capabilities are limited.

[0006] In addition, existing technologies also face the following technical bottlenecks:

[0007] 1) A single model has certain limitations: it cannot dynamically adapt to different types of load models. When faced with different scenarios such as static load, motor load, or mixed load, retraining is required, and there is a lack of automatic recognition and switching capabilities for model types.

[0008] 2) Strong dependence on initial values: The system is sensitive to initial parameter values ​​and relies on the experience of engineers. It lacks a systematic mechanism for evaluating the quality of initial values. Improper initial values ​​can lead to training failure or convergence to a local optimum.

[0009] 3) Serious waste of training resources: Existing methods require independent training of intelligent agents for each load model, resulting in a large number of repeated feature learning processes, especially the repeated learning of common electrical characteristics (such as voltage-power relationship) existing in different load models in the power system;

[0010] 4) Insufficient consideration of physical constraints: Existing methods mainly focus on minimizing data fitting errors, and insufficiently consider the physical rationality constraints of load model parameters, which may result in physically unfeasible parameter combinations. Summary of the Invention

[0011] Objectives of the invention: The objective of the present invention is to provide a method for identifying power load model parameters that supports multi-model adaptation, exhibits high dynamic stability and robustness, and exhibits high computational efficiency, suitable for real-world load scenarios. The second and third objectives of the present invention are to provide an electronic device and storage medium based on this method.

[0012] Technical solution: The power load model parameter identification method based on reinforcement learning described in the present invention includes the following steps:

[0013] (1) According to the power system modeling standards and actual application scenario requirements, a multi-type load model library is established, the parameter vectors to be identified for each model and their physical constraint ranges are clarified, and the input parameters of the load model are clarified;

[0014] (2) Constructing a directed and collaborative reinforcement learning environment;

[0015] (3) training a reinforcement learning agent, wherein the training includes: directed training when a target model is specified and collaborative training when no model is specified;

[0016] (4) Using the trained reinforcement learning agent, perform parameter identification on the specified or unspecified model and output the parameter identification results.

[0017] Optionally, the multi-type load model library includes but is not limited to: static models, induction motor models, comprehensive models, and various comprehensive model revision models with special requirements.

[0018] Optionally, clarifying the input parameters of the load model includes: determining whether the input of the load model is voltage, or voltage and frequency (ie, whether frequency response is considered).

[0019] Optionally, step (2) includes:

[0020] When a target model is specified, a directional environment is constructed: directional parameter matching is performed, and the parameter vector to be identified of the model is initialized within the constraints; the observation domain is expanded; and the reward function is designed by integrating the physical properties of the specified target model.

[0021] When the target model is not specified, a collaborative environment is constructed: synchronously generate initial values ​​for multiple models and evaluate the quality of the initial values; build a heterogeneous environment matrix architecture and establish a cross-model parameter migration channel.

[0022] Optionally, the extended observation domain includes: in addition to the basic observation of the historical sequence window of voltage, frequency, active power, and reactive power, the observation domain is increased according to the special state variables of the specified model. For example, the static model only describes the steady-state characteristics and has no dynamic state variables. The dynamic state variables of the induction motor model include: rotor slip s, which is used to reflect the difference between the motor speed and the synchronous speed; internal electromotive forces Ed and Eq, which are used to describe the internal electromagnetic state of the motor; rotor speed deviation, which is used to characterize the dynamic response process; and electromagnetic torque, which is used to reflect the mechanical characteristics of the motor. The dynamic state variables of the composite load model include: static load ratio λ, that is, the proportional relationship between the static load and the dynamic load; rotor slip s, which is used to reflect the difference between the motor speed and the synchronous speed; internal electromotive forces Ed and Eq, which are used to describe the internal electromagnetic state of the motor; and composite impedance modulus Z, which is used for the equivalent impedance characteristics of the composite load.

[0023] Optionally, the reward function adopts a two-dimensional reward function based on power fitting degree and dynamic stability, which can force the parameters to be optimized within the engineering feasible range, rather than just focusing on the task objective.

[0024] Optionally, generating the initial values of multiple models and evaluating the quality of the initial values includes:

[0025] Generating initial values for the model library synchronously;

[0026] Constructing an initial value quality evaluation matrix to score the feasibility of the initial value parameters of different models.

[0027] Optionally, the initial value quality evaluation matrix is defined as:

[0028]

[0029] where m represents the number of models in the model library; n represents the number of evaluation index dimensions; q

[0024] , , , ,

[0029] , 2 ,

[0032] , , k ,

[0027] ,

[0026] ,

[0025] , , ,

[0031] ,

[0030] , , , k ,

[0028] , ij , , , , represents the quality score of the i-th model under the j-th evaluation index; w k represents the weight coefficient of the k-th evaluation index, satisfying M k represents the score matrix of the k-th evaluation index.

[0030] The calculation formula for the comprehensive quality score of model i is:

[0031]

[0032] By evaluating the quality of the initial values, the present invention can pre-eliminate the parameter combinations that violate the basic physical laws before entering the computationally intensive reinforcement learning training, avoiding ineffective calculations; reducing the computational complexity: through the evaluation of the initial value quality, reducing the m candidate models to k high-quality models (k < m), making the computational complexity of the subsequent collaborative training from O(m 2) is reduced to O(k2); Training stability is guaranteed: high-quality initial values ​​provide a better starting point for training, reduce exploration blind spots in the reinforcement learning process, and improve convergence stability.

[0033] In distribution network load modeling projects, initial value debugging relies heavily on engineer experience and can take days. However, in multi-model identification research, the initial values ​​are often assumed to be known. The initial value quality assessment matrix, along with the two-stage co-evolution and heterogeneous environment matrix, forms a solution for unspecified identification, improving initial value quality and significantly increasing the efficiency of the co-training phase.

[0034] Optionally, constructing a heterogeneous environment matrix architecture includes: defining a scalable state vector, including basic measurement data and a dynamic expansion area; wherein the basic measurement data includes voltage, frequency and power history sequences, and the dynamic expansion area is adaptively filled according to the model type.

[0035] Optionally, after establishing the cross-model parameter migration channel, convergence is accelerated by sharing the strategy network feature extraction layer.

[0036] Since the present invention constructs a multi-type load model library, including load models with large differences in physical characteristics such as static models, induction motor models, and comprehensive models (for example, static models have no dynamic parameters, and motor models contain differential equations), if the target model is not specified, the calculation cost of using the traditional exhaustive method is high. Therefore, the present invention proposes a two-stage co-evolution mechanism for heterogeneous models, which dynamically screens the optimal model type from the model library and optimizes the parameters. Its state space contains model-specific states and is dynamically expanded through a heterogeneous environment matrix to solve the multi-model adaptation problem when the load characteristics are unknown. Then, dynamic screening and collaboration of multiple types of heterogeneous models are realized in power load modeling, thereby achieving efficient identification and dynamic switching of multiple model types.

[0037] Optionally, the directed training when specifying the target model includes: constructing a dedicated reinforcement learning network structure, inputting a directed state vector constructed on an extended observation domain, and outputting parameter adjustment actions, wherein the parameter adjustment actions include reduction, maintenance, and increase; and completing the directed training of the intelligent agent through exploration-experience collection, network update, and convergence judgment.

[0038] The exploration-experience collection includes: randomly selecting an action for each training, otherwise selecting a maximum action, executing a parameter adjustment action to obtain a new state and reward, and storing an experience tuple.

[0039] The network update includes: sampling batch data at fixed intervals, calculating the target reward value and minimizing the loss function, and updating the target network weight according to the loss function.

[0040] Optionally, the collaborative training when no model is specified adopts a two-stage collaborative evolution to complete model redundancy elimination and collaborative optimization acceleration. The two-stage collaborative evolution includes:

[0041] First, perform a rough screening: execute multiple environment explorations in parallel, with each environment running independently for the same number of steps; eliminate models whose fitting error decrease rate is less than a threshold;

[0042] The eliminated models are used to construct an elite environment group, and the intelligent agents in each environment synchronously output action vectors.

[0043] Because the initial quality assessment matrix essentially encompasses physical constraints, the coarse screening phase, based on both physical plausibility and data fit, can preemptively eliminate models that violate physical laws. This coarse screening eliminates models with low error reduction rates. After collaborative training, an elite environment group is constructed. Sample priorities are calculated for experience replay, and similar parameters are periodically checked across models for weight synchronization. Finally, collaborative convergence is determined, and agent training is completed when all elite environments meet the criteria.

[0044] Optionally, the synchronous output of action vectors by the intelligent agents between the environments includes: calculating sample priorities for experience playback, periodically detecting similar parameters for weight synchronization across models, and the periodic detection is based on the overlap of the physical parameter constraint range of the power load model, the similarity of electrical characteristics, and the dynamic response time constant. When there are physically related parameters between different models (such as the inertia constant H of the motor model and the H of the motor part in the comprehensive model are less than 10% different within the engineering allowable range), cross-model weight synchronization is triggered; using the physical hierarchical structure of the power load model, a parameter importance hierarchical protection mechanism is established to ensure that the physical basic characteristics are not destroyed; finally, a collaborative convergence judgment is made, and the intelligent agent training is completed when all elite environment groups meet the indicators.

[0045] Optionally, when performing parameter identification on a specified model, the following processing flow is performed:

[0046] Construct an initial state vector based on the input measurement data;

[0047] Call the trained dedicated reinforcement learning agent, load the network weights and hyperparameters; perform actions and update the state, calculate the reward value in real time and determine whether it exceeds the physical constraints, make convergence judgments and output parameter identification results and visual comparisons.

[0048] Optionally, when performing parameter identification for an unspecified model, the following processing flow is performed:

[0049] The candidate models are loaded from the multi-type load model library and the parameters of each model are initialized; the trained collaborative intelligent agent is loaded and the shared strategy network feature extraction layer and the dedicated output head are activated; the models are quickly screened to obtain an elite environment, the optimal model is selected through comprehensive evaluation indicators, and the parameter identification result is output according to the optimal model selection result.

[0050] The elite environment group achieves cross-model parameter migration (such as steady-state impedance of a static model to dynamic optimization of a comprehensive model) by sharing the policy network feature extraction layer, avoiding repeated learning of common features of heterogeneous models without relying on intelligent agent communication or policy imitation.

[0051] The electronic device described in the present invention includes a memory, a processor, and a computer program stored in the memory and capable of running on the processor. When the processor executes the program, it implements some or all of the steps in the above-mentioned reinforcement learning-based power load model parameter identification method.

[0052] The computer-readable storage medium of the present invention stores a computer program thereon, and when the computer program is executed by a processor, it implements part or all of the steps in the above-mentioned method for identifying power load model parameters based on reinforcement learning.

[0053] Beneficial effects: Compared with the prior art, the present invention has the following significant advantages:

[0054] 1. Traditional methods require pre-determining the load model type, while the load components in actual power systems are complex and changeable. This invention, by building a multi-type load model library and a collaborative identification mechanism for unspecified target models, can automatically identify the optimal model without knowing the load type in advance. This completely solves the model selection problem in engineering applications and transforms load modeling from parameter determination based on known models to automatic identification of unknown models and simultaneous parameter optimization.

[0055] 2. The innovative two-stage co-evolution mechanism of this invention avoids blind training of all candidate models, quickly eliminates unsuitable models through coarse screening, centrally optimizes the elite environment group, and uses a shared network structure to avoid repeated feature learning, truly achieving intelligent and efficient identification;

[0056] 3. Traditional methods are greatly affected by initial parameter values ​​and are prone to falling into local optimal solutions. This invention provides a scientific parameter starting point through an initial value quality assessment mechanism. A dual-dimensional reward function ensures a balance between physical constraints and fitting accuracy. The cross-model parameter migration mechanism uses existing knowledge to accelerate convergence, significantly improving the consistency and reproducibility of parameter identification results. This effectively solves the reliability problem of "different results for the same data" in traditional methods.

[0057] 4. The heterogeneous environment matrix architecture of the present invention can dynamically adjust the state space according to the characteristics of different load models. The collaborative training mechanism enables the system to have the ability to handle load mutations and model switching. When facing complex power grid conditions such as motor starting and load switching, it can automatically adjust the model type and parameters without manual intervention, which greatly improves the robustness and engineering applicability of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 It is a schematic diagram of the process of the present invention;

[0059] Figure 2 A schematic diagram of the process of directional training when specifying a target model in one embodiment of the present invention;

[0060] Figure 3 This is a schematic diagram of the overall process of Example 2 of the present invention;

[0061] Figure 4 A heterogeneous environment matrix architecture constructed in one embodiment of the present invention;

[0062] Figure 5 Schematic diagram of the process of collaborative agent training when no target model is specified in one embodiment of the present invention;

[0063] Figure 6 Schematic diagram of the two-stage co-evolution process when no target model is specified in one embodiment of the present invention. DETAILED DESCRIPTION

[0064] The technical solution of the present invention will be further described below with reference to the accompanying drawings.

[0065] like Figure 1 As shown, a method for identifying power load model parameters based on reinforcement learning includes: first, according to the power system modeling standard and the actual application scenario requirements, a multi-type load model library is established, the parameter vector to be identified for each model and its physical constraint range are clarified, and the input parameters of the load model are clarified; then, a directed and collaborative reinforcement learning environment is constructed; the agent parameters are initialized and the reinforcement learning agent is trained; the trained reinforcement learning agent is used to perform parameter identification on a specified or unspecified model, and when the reward function converges, the parameter identification result is output.

[0066] Example 1: Parameter identification of power system load model with specified target model

[0067] Requirement: An accurate equivalent load model of a motor at a load station must be established for transient stability analysis.

[0068] Step 1: Construction of load model library.

[0069] According to the Technical Requirements for Load Modeling for Power System Simulation, a third-order induction motor model is established and the parameter vector to be identified is defined as follows:

[0070] θ=[Rs,Xs,Xm,Rr,Xr,H]

[0071] Where Rs represents the stator resistance, Xs represents the stator reactance, Xm represents the excitation impedance, Rr represents the rotor resistance, Xr represents the rotor reactance, and H represents the equivalent inertia time constant of the third-order induction motor.

[0072] Add physical constraints for the above parameters:

[0073] Rs∈[0.01,0.05]pu

[0074] Xs∈[0.15,0.22]pu

[0075] Xm∈[2.0,4.0]pu

[0076] Rr∈[0.01,0.08]pu

[0077] Xr∈[0.07,0.18]pu

[0078] H∈[0.5,5.0]s

[0079] Wherein, pu represents per-unit value and s represents second.

[0080] Define input and output, including:

[0081] Input quantities: voltage amplitude U(k) and system frequency f(k);

[0082] Output: active power P_sim(k), reactive power Q_sim(k).

[0083] Where k represents the discrete time sampling step index, that is, the kth sampling moment.

[0084] Step 2: Construction of reinforcement learning environment.

[0085] Step 21: Construct the state space.

[0086] First, the basic observation domain includes:

[0087] Voltage history sequence: U(k-119) to U(k), frequency history sequence: f(k-119) to f(k), active power sequence: P(k-119) to P(k), reactive power sequence: Q(k-119) to Q(k);

[0088] The model extension domain includes:

[0089] Rotor slip sequence: s(k-119) to s(k), direct axis internal potential sequence: Ed(k-119) to Ed(k), quadrature axis internal potential sequence: Eq(k-119) to Eq(k).

[0090] Step 22: Design the reward function R.

[0091] Considering that active power fluctuates more than reactive power under voltage and frequency disturbances, the score for correct active power judgment is improved:

[0092] R=-[0.6×|P_sim-P_meas|+0.3×|Q_sim-Q_meas|]

[0093] Among them, P_sim represents the model simulated active power, P_meas represents the measured active power, Q_sim represents the model simulated reactive power, and Q_meas represents the measured reactive power.

[0094] Step 3: Perform agent training, that is, directional training when specifying the target model, such as Figure 2 shown.

[0095] Step 31: Initialize the network.

[0096] First, a dedicated reinforcement learning network structure is constructed. The input layer receives the constructed directional state vector (basic observation domain + model extension domain), and the output layer dimension is 3×n (n=number of parameters). Each parameter corresponds to three actions (decrease, maintain and increase).

[0097] In one embodiment, the DQN algorithm (Deep Q-Network) is adopted, and the network architecture is configured as follows: input layer: 120*7 (corresponding to the state vector dimension), hidden layer: first fully connected layer: 128 neurons (ReLU activation function), second fully connected layer: 64 neurons (ReLU activation function), output layer: 18 neurons (6 parameters × 3 discrete actions: decrease / maintain / increase).

[0098] Step 32: Conduct exploration experience collection.

[0099] Initialize motor parameters: Rs (0) =0.03, Xs (0) =0.2, Rr (0) =0.06, Xr (0) =0.15, H (0) =2.5,Xm (0)=3.0; select a random action with probability ε, otherwise select the action with maximum Q value; perform parameter adjustment: action encoding 0: parameter reduction 0.5% × parameter range, action encoding 1: parameter remains unchanged, action encoding 2: parameter increase 0.5% × parameter range; store experience tuple (current state, action, reward, new state).

[0100] Step 33: Update network parameters.

[0101] Every 100 training steps, 32 sets of data are randomly sampled from the experience pool and the target Q value is calculated:

[0102] Q_target=r+γ×max_a'Q(s_{t+1},a';θ - )

[0103] Q_target represents the target Q value, which is the expected cumulative reward of the reinforcement learning agent under the current state-action pair; r represents the immediate reward, which is the immediate feedback obtained by the agent after performing the parameter adjustment action; γ represents the discount factor, which balances the weight of immediate reward and future reward. The larger γ is, the more emphasis is placed on long-term benefits; max_a' represents the maximum value operation of all possible actions; Q(s_{t+1},a';θ - ) represents the Q value of the target network in the next state; a' represents all possible actions in the next state; θ - Represents the target network parameters (updated periodically from the main network).

[0104] Optimize the loss function to minimize the mean square error:

[0105] L=E[(Q_target-Q(s_t,a_t;θ)) 2 ]

[0106] Q(s_t,a_t;θ) represents the Q value predicted by the main network; s_t represents the current state vector; a_t represents the currently executed action; θ represents the weight parameter of the main network.

[0107] Step 34: Perform convergence judgment and stop training when the following conditions are met: the average reward change rate for 20 consecutive steps is <1%, the maximum parameter change is <0.0001, and the total number of training steps is ≥300.

[0108] Step 4: Parameter identification implementation.

[0109] Input data preparation: The voltage disturbance data of the load site is collected, and the active power, reactive power and frequency data of the load site corresponding to the voltage disturbance are recorded at the same time, and filtering and data standardization are performed.

[0110] Agent identification: Initialize the state matrix, iterative execution: DQN agent outputs action a t, update the parameter θ t+1 =θ t +Δθ t , drive the motor model to calculate P_sim, Q_sim, and calculate the reward r t And update the state s t+1 , if the parameter change is less than 0.0001 for 10 consecutive steps, the identification is terminated and the identification result is output.

[0111] Example 2: Parameter identification of power system load model without specified target model

[0112] Requirement: A load model needs to be established for the distribution network in a commercial area of ​​a city for voltage stability judgment, but a fixed model cannot be given.

[0113] Step 1: If Figure 3 As shown, the model library is initialized first.

[0114] Step 11: Synchronously generate initial values ​​for the model library. For example, for static models, use the least squares method to fit the steady-state point; for generator models, use the step response characteristic to match the transfer function. Then, construct a set of candidate models: static ZIP model: θ_ZIP = [Zp, Zq, Ip, Iq, Pp, Pq]; induction motor model: θ_Motor = [Rs, Xs, Xm, Rr, Xr, H]; composite model: θ_Comp = [λ_ZIP, Rs, Xs, Xm, Rr, Xr, H]. Initialize the parameters of each candidate model. Zp and Zq represent the active and reactive coefficients of a constant-impedance load; Ip and Iq represent the active and reactive coefficients of a constant-current load; Pp and Pq represent the active and reactive coefficients of a constant-power load; and λ_ZIP represents the static load ratio.

[0115] Step 12: Construct an initial value quality assessment matrix to score the feasibility of the initial value parameters of different models.

[0116] Take the model library containing three candidate models as an example:

[0117] Model 1 is a static ZIP model with initial parameters of P0 = 0.6, Q0 = 0.4, α = 1.2, β = 2.1, and γ = 0.8;

[0118] Model 2 is an induction motor model, and the initial parameter is R s =0.03,X s =0.2, Rr=0.06, Xr=0.15, H=2.5, Xm=3.0;

[0119] Model 3 is a composite model with initial parameters of λ=0.7, Rs=0.04, Xs=0.18, Rr=0.05, Xr=0.16, H=2.2, and Xm=2.8.

[0120] The initial quality assessment matrix is ​​defined as:

[0121]

[0122] Among them, m represents the number of models in the model library; n represents the number of evaluation indicator dimensions; q ij represents the quality score of the i-th model under the j-th evaluation indicator; w k Represents the weight coefficient of the kth evaluation index, satisfying M k Represents the score matrix of the k-th evaluation metric.

[0123] The physical rationality evaluation index is taken as follows:

[0124] q i1 =C bound +C relation +C physics

[0125] Among them, C bound The maximum score is 40 points: 40 points when all parameters are within the physical constraint range, 30 points when 90% of the parameters are within the constraint, 20 points when 80% of the parameters are within the constraint, and 0 points in other cases; C relation The score for correlation constraint is 35 points: 35 points are awarded when the correlation coefficient between parameters is within a reasonable range [0.1, 0.8], 20 points are awarded when the correlation coefficient between some parameters is abnormal, and 0 points are awarded when the correlation coefficient is seriously abnormal; C physics It is the typical value experience score, with a full score of 25 points: 25 points when it fully complies with the typical value range, 15 points when it deviates from the typical value by within 10%, 5 points when it deviates by 10-30%, and 0 points when it deviates by more than 30%.

[0126] The convergence prediction evaluation index is taken as follows:

[0127] q i2 =R conv +R stable

[0128] Among them, R conv The convergence speed prediction score is 60 points: 60 points are given when the gradient analysis predicts fast convergence (less than 50 steps), 40 points are given when the gradient analysis predicts moderate convergence (50-100 steps), and 20 points are given when the gradient analysis predicts slow convergence (more than 100 steps). stable The convergence stability score is 40 points: 40 points when the Hessian matrix condition number is less than 100, 25 points when the condition number is between 100 and 1000, and 10 points when the condition number is greater than 1000.

[0129] The numerical stability evaluation index is taken as follows:

[0130] q i3 =N sensitivity +N numerical

[0131] Among them, N sensitivity The initial value sensitivity score is 50 points: 50 points when the parameter change is less than 1% under 5% disturbance, 30 points when the parameter change is between 1-5%, and 10 points when the parameter change is greater than 5%. numerical The numerical calculation stability score is 50 points: 50 points when there is no overflow or underflow in the calculation process, 30 points when there is a minor numerical problem, and 0 points when there is a serious numerical problem.

[0132] According to the above scoring rules, a quality assessment matrix is ​​constructed:

[0133] Model 1 scored 85, 78, and 82 points under the three evaluation indicators respectively;

[0134] Model 2 scored 92, 85, and 88 points under the three evaluation indicators respectively;

[0135] Model 3 scored 78, 72, and 75 points under the three evaluation indicators respectively.

[0136] The comprehensive quality score calculation formula of model i is:

[0137]

[0138] The number of evaluation indicator dimensions is set to 3, and the weights of each indicator are assigned as follows: physical rationality evaluation weight w1 = 0.30, convergence prediction evaluation weight w2 = 0.25, and numerical stability evaluation weight w3 = 0.20. That is, the weighted calculation is performed using the weight vector W = [0.30, 0.25, 0.20] to obtain the comprehensive quality score of each model:

[0139] Model 1 comprehensive score:

[0140] S1=0.30×85+0.25×78+0.20×82=61.4 points

[0141] Model 2 comprehensive score:

[0142] S2 = 0.30 × 92 + 0.25 × 85 + 0.20 × 88 = 66.4 points

[0143] Model 3 comprehensive score:

[0144] S3=0.30×78+0.25×72+0.20×75=56.4 points

[0145] The ranking results by comprehensive score are: Model 2 (61.4 points) > Model 1 (66.4 points) > Model 3 (56.4 points). Setting the screening threshold to 60 points means that Models 1 and 2 enter the subsequent collaborative training phase, and Model 3 is pre-eliminat- ed , thereby improving overall training efficiency.

[0146] Step 2: Build a collaborative environment.

[0147] Step 21: Design the dynamic state container.

[0148] Build a heterogeneous environment matrix architecture: define a scalable state vector, including basic measurement data (voltage, frequency and power history series) and dynamic expansion area (adaptively filled according to model type), such as Figure 4 shown.

[0149] In one embodiment, the sampling frequency is set to 50 Hz, and an observation sequence of 120 points is constructed. The state observation domain of the static ZIP model includes:

[0150] Voltage history series: U(k-119) to U(k), frequency history series: f(k-119) to f(k), active power series: P(k-119) to P(k), reactive power series: Q(k-119) to Q(k), excluding the extended domain.

[0151] The basic observation domain of the induction motor model includes:

[0152] The voltage history sequence is from U(k-119) to U(k), the frequency history sequence is from f(k-119) to f(k), the active power sequence is from P(k-119) to P(k), and the reactive power sequence is from Q(k-119) to Q(k). The extended domain of this model includes the rotor slip sequence from s(k-119) to s(k), the direct-axis potential sequence from Ed(k-119) to Ed(k), and the quadrature-axis potential sequence from Eq(k-119) to Eq(k).

[0153] The composite model basic observation domain includes:

[0154] The voltage history sequence is from U(k-119) to U(k), the frequency history sequence is from f(k-119) to f(k), the active power sequence is from P(k-119) to P(k), and the reactive power sequence is from Q(k-119) to Q(k). The extended domain of this model includes the rotor slip sequence from s(k-119) to s(k), the direct-axis potential sequence from Ed(k-119) to Ed(k), the quadrature-axis potential sequence from Eq(k-119) to Eq(k), and the static load ratio λ(k).

[0155] Step 22: Build a cross-environment communication mechanism. When the Q value difference between the motor model and the composite model is less than 0.1, the inertia constant of the pure motor model is transferred to the inertia constant of the motor part in the composite model.

[0156] Step 3: Conduct collaborative agent training.

[0157] Step 31: Figure 5 As shown, the DQN algorithm is selected for training. First, the network architecture is constructed: the shared feature extraction layer includes the input layer: 120*7+1 neurons (zero padding to handle insufficient dimensions), fully connected layer 1: 128 neurons (ReLU activation), fully connected layer 2: 64 neurons (ReLU activation), dedicated output heads: ZIP model head: 9 neuron outputs (3 actions × 3 parameters), motor model head: 18 neuron outputs (3 actions × 6 parameters), composite model head: 21 neuron outputs (3 actions × 7 parameters).

[0158] Step 32: Perform training process.

[0159] A two-stage co-evolution is adopted, such as Figure 6 shown.

[0160] First, a rough screening is performed, and each environment runs independently for the same number of steps. Models with a fitting error reduction rate less than a threshold (such as an error reduction rate higher than -0.3) are eliminated; the eliminated models are used to construct an elite environment group, and the intelligent agents in each environment synchronously output action vectors.

[0161] Coarse screening stage (first 200 steps): All environments are explored in parallel and the fitting error reduction rate is calculated every 50 steps: η = (ε_t-ε_{t-50}) / ε_{t-50}, ε_t represents the fitting error of the current step, ε_{t-50} represents the fitting error 50 steps ago, and models with η>-0.3 are eliminated (the error reduction is less than 30%).

[0162] Elite stage (steps 201-800): Perform sample priority experience replay: sampling probability P(i)∝|Q_target-Q_pred|, Q_target represents the target Q value, Q_pred represents the predicted Q value, detect and perform parameter migration every 100 steps, and the dedicated output head is updated independently.

[0163] Among them, periodic detection is based on the following three dimensions:

[0164] 1) Overlap of physical parameter constraint ranges: Calculate the degree of overlap of the physical constraint ranges of parameters between different models. For example, if the overlap of the constraint ranges of the motor model and the motor part of the composite model exceeds 80%, the parameters are considered to be physically correlated.

[0165] 2) Electrical characteristic similarity: The judgment is based on the electrical characteristic similarity matrix S, S_ij = exp(-||x_i - x_j||^2 / σ^2), where x_i and x_j are the electrical characteristic vectors of the two models, and σ is the similarity threshold. Synchronization is triggered when S_ij>0.7;

[0166] 3) Inertia time constant H: Calculate the time constant difference of different models. When the time constant difference is less than 20%, the dynamic characteristics are considered similar.

[0167] Establish a parameter importance classification protection mechanism, including:

[0168] 1) Core physical parameters (such as the inertia time constant H of the motor model): only small adjustments within the physical constraints are allowed, and the adjustment range does not exceed 5%;

[0169] 2) Secondary physical parameters (such as stator resistance Rs, stator reactance Xs, etc.): Moderate adjustments are allowed within the physical constraints, with the adjustment range not exceeding 15%;

[0170] 3) Auxiliary parameters (such as the static load ratio λ in the composite model): large adjustments are allowed within the physical constraints, and the adjustment range does not exceed 30%.

[0171] When the above conditions are met, cross-model weight synchronization is triggered, and the synchronization process is as follows:

[0172] 1) Calculate the weight difference between the corresponding parameters of the source model and the target model: Δw = w_source - w_target, where w_source represents the parameter weight of the synchronization source and w_target represents the parameter weight of the synchronization target;

[0173] 2) Determine the synchronization coefficient α according to the parameter importance level:

[0174] -Core physical parameter synchronization coefficient: α = 0.1

[0175] -Secondary physical parameter synchronization coefficient: α = 0.3

[0176] -Auxiliary parameter synchronization coefficient: α = 0.5

[0177] 3) Update the target model weight: w_target_new = w_target + α·Δw

[0178] 4) Verify that the updated parameters meet the physical constraints. If not, roll back the update.

[0179] Step 33: Finally, perform convergence judgment. When the elite group meets the following conditions simultaneously: average reward change rate < 0.5%, maximum parameter change < 0.0005, and Q value difference between models < 0.05, stop agent training.

[0180] Step 4: Parameter identification implementation

[0181] Input data preparation: The voltage disturbance data of the commercial area is collected, and the active power, reactive power and frequency data of the load sites corresponding to the voltage disturbance are recorded at the same time, and filtering and data standardization are performed.

[0182] Agent identification: First, a quick screening is performed, 50 steps of exploration are performed in parallel, and models with an error reduction rate higher than -0.3 are eliminated. Then, elite collaborative identification is performed, and the screened models are collaboratively trained, periodically migrating parameters such as the inertia constant, and a comprehensive assessment is conducted to determine whether the requirements are met and output.

[0183] The present invention also provides an electronic device for executing a method for identifying power load model parameters based on reinforcement learning, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, some or all of the steps in the above-mentioned method for identifying power load model parameters based on reinforcement learning are implemented.

[0184] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements some or all of the steps in the above-mentioned method for identifying power load model parameters based on reinforcement learning.

Claims

1. A method for identifying power load model parameters based on reinforcement learning, characterized in that: The method comprises the following steps: (1) According to the power system modeling standards and actual application scenario requirements, a multi-type load model library is established, the parameter vectors to be identified for each model and their physical constraint ranges are clarified, and the input parameters of the load model are clarified; (2) Constructing a directed and collaborative reinforcement learning environment; (3) training a reinforcement learning agent, wherein the training includes: directed training when a target model is specified and collaborative training when no model is specified; (4) Using the trained reinforcement learning agent, perform parameter identification on the specified or unspecified model and output the parameter identification results.

2. The power load model parameter identification method based on reinforcement learning according to claim 1 is characterized in that: In step (1), the multi-type load model library includes at least one of the following: a static model, an induction motor model, a comprehensive model, and various comprehensive model revision models with special requirements.

3. The method for identifying power load model parameters based on reinforcement learning according to claim 1, characterized in that: The step (2) comprises: When a target model is specified, a directional environment is constructed: directional parameter matching is performed, and the parameter vector to be identified of the model is initialized within the constraints; the observation domain is expanded; and the reward function is designed by integrating the physical properties of the specified target model. When the target model is not specified, a collaborative environment is constructed: synchronously generate initial values ​​for multiple models and evaluate the quality of the initial values; build a heterogeneous environment matrix architecture and establish a cross-model parameter migration channel.

4. The method for identifying power load model parameters based on reinforcement learning according to claim 1, characterized in that: The directed training when specifying the target model includes: constructing a dedicated reinforcement learning network structure, inputting a directed state vector constructed on an extended observation domain, and outputting parameter adjustment actions, wherein the parameter adjustment actions include reduction, maintenance, and increase; and completing the directed training of the intelligent agent through exploration-experience collection, network update, and convergence judgment.

5. The method for identifying power load model parameters based on reinforcement learning according to claim 1, characterized in that: The collaborative training when no model is specified adopts two-stage collaborative evolution: First, perform a rough screening: perform multi-environment exploration in parallel, run each environment independently for the same number of steps, and eliminate models whose fitting error decrease rate is less than a threshold; The eliminated models are used to construct an elite environment group, and the intelligent agents in each environment synchronously output action vectors.

6. The method for identifying power load model parameters based on reinforcement learning according to claim 5, characterized in that: The synchronous output of action vectors by the intelligent agents between the environments includes: calculating sample priorities for experience playback, periodically detecting similar parameters across models for weight synchronization, the periodic detection being based on the overlap of the physical parameter constraint ranges of the power load model, the similarity of electrical characteristics, and the dynamic response time constant. When physically related parameters exist between different models, cross-model weight synchronization is triggered; utilizing the physical hierarchical structure of the power load model, a parameter importance hierarchical protection mechanism is established to ensure that the basic physical characteristics are not destroyed; finally, a collaborative convergence judgment is made, and the intelligent agent training is completed when all elite environment groups meet the indicators.

7. The method for identifying power load model parameters based on reinforcement learning according to claim 1, characterized in that: When performing parameter identification for a specified model, the following processing flow is performed: Construct an initial state vector based on the input measurement data; Call the trained dedicated reinforcement learning agent, load the network weights and hyperparameters; perform actions and update the state, calculate the reward value in real time and determine whether it exceeds the physical constraints, make convergence judgments and output parameter identification results and visual comparisons.

8. The method for identifying power load model parameters based on reinforcement learning according to claim 1, characterized in that: When performing parameter identification for an unspecified model, the following processing flow is performed: The candidate models are loaded from the multi-type load model library and the parameters of each model are initialized; the trained collaborative intelligent agent is loaded and the shared strategy network feature extraction layer and the dedicated output head are activated; the models are quickly screened to obtain an elite environment, the optimal model is selected through comprehensive evaluation indicators, and the parameter identification result is output according to the optimal model selection result.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the power load model parameter identification method based on reinforcement learning as described in any one of claims 1 to 8 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the power load model parameter identification method based on reinforcement learning as described in any one of claims 1 to 8 are implemented.