PID parameter optimization method and device for AC / DC system voltage drop

By using enhanced learning algorithms to optimize PID parameters in AC/DC systems, the problem of unstable PID parameter optimization in the prior art is solved, and faster and more accurate voltage drop management effect is achieved.

CN120197577APending Publication Date: 2025-06-24TSINGHUA UNIVERSITY +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510257765.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The prior art is difficult to effectively optimize PID parameters in AC/DC systems, resulting in unstable voltage drop management and prone to over-regulation and under-regulation problems. In severe cases, it can lead to interruption of the power supply system.

Method used

By obtaining the target input sequence and inputting it into a pre-constructed enhanced learning algorithm model, the target output sequence is generated and compared with the set of PID control parameters in the simulation circuit model to determine the optimal PID parameters.

Benefits of technology

It significantly improves the reaction speed and accuracy of the circuit, avoids the defects of slow speed and obvious oscillation of traditional algorithms, and ensures the stability and safety of AC/DC systems in the case of voltage drops.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197577A_ABST
    Figure CN120197577A_ABST
Patent Text Reader

Abstract

The invention provides a PID parameter optimization method and device for voltage drop of an AC / DC system, and relates to the technical field of power systems, and the method comprises the steps: obtaining a target input sequence, inputting the target input sequence into a pre-constructed reinforcement learning algorithm model, and obtaining a corresponding target output sequence; wherein the reinforcement learning algorithm model is constructed based on a target linear sample set of the AC / DC system under the condition of voltage drop; comparing a performance result of the target output sequence in the simulation circuit model with a performance result set of the PID control parameter set in the simulation circuit model, and determining an optimal PID parameter based on a comparison result; wherein the simulation circuit model is a simulation model comprising a PID (Proportion Integration Differentiation) controller. Through the method provided by the invention, the optimal PID control parameter is obtained, the defects of low speed and obvious oscillation of a traditional algorithm are overcome, and the reaction speed and precision of the circuit are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power systems, and in particular to a PID parameter optimization method and device for AC / DC system voltage drop. Background Art

[0002] In order to cope with the challenges of environmental pollution, energy crisis and global warming currently faced by mankind, new energy systems based on energy internet have received widespread attention. Such systems are represented by energy internet, and their operation is inseparable from a large number of power electronic equipment and new renewable energy sources, and AC / DC system is one of the typical application scenarios. In AC / DC system, the volatility of current and voltage and its operation mechanism are more complex than traditional AC system, which makes it more difficult to control it in a timely and stable manner. In the AC / DC operation scenario of energy internet, with the popularization and widespread use of renewable energy and equipment, more and more controllers, power monitoring equipment and various power-consuming equipment are connected to energy internet, and the instability of AC / DC network is significantly enhanced. When there is a grid failure or insufficient supply, cross-grid voltage drop is prone to occur. AC / DC voltage drop has become a potential serious problem threatening the overall operation of energy internet.

[0003] At present, the AC / DC voltage drop control equipment is represented by UPQC and SVG. Its software control method generally adopts PID control method, which can be arranged at the AC end and before the PWM control device. PID control methods include P, PI, PD and PID, among which P represents proportional control, I represents integral control, D represents differential control, and PI, PD and PID represent corresponding combinations. These methods are sensitive to the setting of control parameters, and are prone to over-adjustment and under-adjustment. Unreasonable parameter settings further aggravate the insecurity and instability of the AC / DC network, and in severe cases may cause power supply system interruption. Compared with the first three methods, PID combines integral and differential functions at the same time, which speeds up the voltage compensation speed, but at the same time, the voltage oscillation problem becomes more serious, which may cause voltage limit exceeding and reduce equipment life. Therefore, the selection and optimization of PID related parameters becomes more important, and its intelligent parameter optimization has become a research hotspot.

[0004] How to select the optimal PID parameters for AC / DC system voltage drop is a technical problem that needs to be solved at present. Summary of the invention

[0005] The present invention provides a PID parameter optimization method and device for AC / DC system voltage drop, so as to solve the defects in the prior art.

[0006] The present invention provides a PID parameter optimization method for AC / DC system voltage drop, comprising the following steps: Obtain a target input sequence, and input the target input sequence into a pre-constructed reinforcement learning algorithm model to obtain a corresponding target output sequence; wherein, the reinforcement learning algorithm model is constructed based on a target linear sample set in the case of voltage sag in the AC / DC system. Compare the performance result of the target output sequence in the simulation circuit model with the performance result set of the PID control parameter set in the simulation circuit model, and determine the optimal PID parameter based on the comparison result; wherein, the simulation circuit model is a simulation model including a PID controller.

[0007] According to a PID parameter optimization method for voltage sag in an AC / DC system provided by the present invention, the construction process of the reinforcement learning algorithm model includes: In the case of voltage sag in the AC / DC system, collect basic training sample data through PID circuit simulation, and screen the basic training sample data based on preset constraint conditions to obtain initial linear samples. Based on the initial linear samples, obtain the target linear sample set by means of linear interpolation. Construct the reinforcement learning algorithm model based on the target linear sample set; wherein, the reinforcement learning algorithm model includes: system state, possible actions and policies, and corresponding reward functions.

[0008] According to a PID parameter optimization method for voltage sag in an AC / DC system provided by the present invention, the step of collecting basic training sample data through PID circuit simulation in the case of voltage sag in the AC / DC system includes: In the case of voltage sag in the AC / DC system, obtain an input data sequence and an initial sag voltage; wherein, the input data sequence is an external current sequence in the first preset time period before the voltage sag. Based on the input data sequence and the initial sag voltage, iteratively calculate an output data sequence through a preset formula.

[0009] According to a PID parameter optimization method for voltage sag in an AC / DC system provided by the present invention, the step of screening the basic training sample data based on preset constraint conditions to obtain initial linear samples includes: Obtain the simulation circuit model including a PID controller established in advance, and determine the preset constraint conditions of the basic training sample data according to the voltage recovery situation of the circuit in the second preset time period; wherein, the simulation circuit model includes the situation of voltage sag. Pairwise sampling is performed on the input data sequence and the output data sequence based on the preset constraint conditions to obtain a sampling sequence pair, which is used as the initial linear sample.

[0010] According to a PID parameter optimization method for voltage sag in an AC / DC system provided by the present invention, based on the initial linear sample, the target linear sample set is obtained by linear interpolation, including: Performing linear interpolation on the sampling sequence pair to obtain a plurality of to-be-verified sequence pairs, and respectively substituting each to-be-verified sequence pair in the plurality of to-be-verified sequence pairs into the simulation circuit model. When the operation result meets the preset constraint conditions, determining the corresponding to-be-verified sequence pair as the target linear sample; The target linear sample set is constituted by the target linear sample and the initial linear sample.

[0011] According to a PID parameter optimization method for voltage sag in an AC / DC system provided by the present invention, after comparing the performance result of the target output sequence in the simulation circuit model with the performance result set of the PID control parameter set in the simulation circuit model, the method further includes: When the comparison result indicates that any performance result in the performance result set is better than the performance result of the target output sequence in the simulation circuit model, changing the parameters of the reinforcement learning algorithm model or increasing the amount of basic training sample data.

[0012] The present invention also provides a PID parameter optimization device for voltage sag in an AC / DC system, including the following modules: An acquisition module, configured to acquire a target input sequence and input the target input sequence into a pre-constructed reinforcement learning algorithm model to obtain a corresponding target output sequence; wherein, the reinforcement learning algorithm model is constructed based on a target linear sample set in the case of voltage sag in the AC / DC system; An optimization module, configured to compare the performance result of the target output sequence in the simulation circuit model with the performance result set of the PID control parameter set in the simulation circuit model, and determine the optimal PID parameter based on the comparison result; wherein, the simulation circuit model is a simulation model including a PID controller.

[0013] According to a PID parameter optimization device for voltage sag in an AC / DC system provided by the present invention, the device further includes a training module, specifically configured to: In the case of voltage sag in the AC / DC system, collecting basic training sample data through PID circuit simulation, and screening the basic training sample data based on preset constraint conditions to obtain an initial linear sample; Based on the initial linear samples, the target linear sample set is obtained by means of linear interpolation; Based on the target linear sample set, the reinforcement learning algorithm model is constructed; wherein, the reinforcement learning algorithm model includes: system state, possible actions and policies, and corresponding reward functions.

[0014] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the PID parameter optimization method for voltage sag of the AC / DC system as described in any one of the above is implemented.

[0015] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the PID parameter optimization method for voltage sag of the AC / DC system as described in any one of the above is implemented.

[0016] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the PID parameter optimization method for voltage sag of the AC / DC system as described in any one of the above is implemented.

[0017] A PID parameter optimization method and device for voltage sag of the AC / DC system provided by the present invention obtain a target input sequence and input the target input sequence into a pre-constructed reinforcement learning algorithm model to obtain a corresponding target output sequence; wherein, the reinforcement learning algorithm model is constructed based on a target linear sample set in the case of voltage sag of the AC / DC system; compare the performance result of the target output sequence in the simulation circuit model with the performance result set of the PID control parameter set in the simulation circuit model, and determine the optimal PID parameter based on the comparison result; wherein, the simulation circuit model is: a simulation model including a PID controller. It can be seen that, aiming at the problem of insufficient network training samples, the present invention greatly improves the number of samples through linear sampling technology. The reinforcement learning algorithm model built based on the target linear samples ensures the performance of the neural network while avoiding overfitting and underfitting caused by insufficient samples, obtains the optimal PID control parameters, overcomes the defects of slow speed and obvious oscillation of the traditional algorithm, and significantly improves the response speed and accuracy of the circuit. Description of the Drawings

[0018] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0019] Figure 1 It is a schematic flow chart of the PID parameter optimization method for voltage sag in the AC / DC system provided by the present invention.

[0020] Figure 2 It is a complete flow chart of the PID parameter optimization method for voltage sag in the AC / DC system provided by the present invention.

[0021] Figure 3 It is a schematic diagram of linear sample generation provided by the present invention.

[0022] Figure 4 It is a schematic diagram of the reinforcement learning algorithm model provided by the present invention.

[0023] Figure 5 It is a flow chart of the A3C algorithm provided by the present invention.

[0024] Figure 6 It is a schematic diagram of the PID control principle provided by the present invention.

[0025] Figure 7 It is a schematic structural diagram of the PID parameter optimization device for voltage sag in the AC / DC system provided by the present invention.

[0026] Figure 8 It is a schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners

[0027] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.

[0028] The following combines Figures 1 - 8 to describe a PID parameter optimization method and device for voltage sag in the AC / DC system according to the present invention.

[0029] Figure 1 It is a schematic flow chart of the PID parameter optimization method for voltage sag in the AC / DC system provided by the present invention. As Figure 1 shown, the method includes the following: Step 100, obtain a target input sequence, and input the target input sequence into a pre-constructed reinforcement learning algorithm model to obtain a corresponding target output sequence; wherein, the reinforcement learning algorithm model is constructed based on a target linear sample set when the AC / DC system has a voltage sag.

[0030] Figure 2 This is the complete flowchart of the PID parameter optimization method for voltage dips in the AC / DC system provided by the present invention. The following will describe the PID parameter optimization method for voltage dips in the AC / DC system provided by the present invention in conjunction with Figure 2 explain the PID parameter optimization method for voltage dips in the AC / DC system provided by the present invention.

[0031] It should be noted that when facing complex learning tasks, reinforcement learning algorithms can be adopted. The reinforcement learning model consists of two parts: an agent and an environment. During the interaction process, the agent first selects a series of actions from the set of possible actions, and the environment calculates the state changes of the system and the expected system rewards (rewards) generated after executing the action sequence. Based on the degree of change in the expected rewards, it learns the corresponding system states (mainly including the probability distribution of actions and their corresponding action policy gradients), and guides the agent to learn better actions and strategies from the set of actions. This process is iterated until the iteration ends, ultimately maximizing the system rewards.

[0032] Currently existing machine learning algorithms have great potential for optimizing PID control circuits. The PID system, that is, the Proportional-Integral-Derivative control system, especially the reinforcement learning algorithm. By setting a reasonable reward function, the system can achieve smaller overshoots and quickly achieve circuit stability. To address the problem of insufficient network training samples, linear sampling technology can greatly increase the number of samples to ensure the performance of relevant neural networks and avoid overfitting and underfitting caused by insufficient samples. Based on this, this embodiment provides a PID parameter optimization method for voltage dips in the AC / DC system and will be specifically described.

[0033] Specifically, when a voltage dip occurs ( ), it is necessary to use the PID circuit to quickly restore the voltage and avoid large voltage fluctuations. This embodiment collects the original training model input data by combining linear sampling and uses a deep reinforcement learning network to obtain the optimal PID parameters in this scenario. Before step 100, first, the construction process of the reinforcement learning algorithm model provided in this embodiment will be described.

[0034] The construction process of the reinforcement learning algorithm model includes: Step S110: In the case of a voltage dip in the AC / DC system, collect basic training sample data through PID circuit simulation, and screen the basic training sample data based on preset constraint conditions to obtain initial linear samples.

[0035] Step S110 specifically includes: Step S111: In the case of a voltage dip in the AC / DC system, obtain the input data sequence and the initial dip voltage; wherein, the input data sequence is the external current sequence in the first preset time period before the voltage dip.

[0036] Step S112: Based on the input data sequence and the initial dip voltage, iteratively calculate the output data sequence through a preset formula.

[0037] Step S113: Obtain the pre-established simulation circuit model including a PID controller, and determine the preset constraint conditions of the basic training sample data according to the voltage recovery situation of the circuit in the second preset time period; wherein, the simulation circuit model includes the situation of voltage dip.

[0038] Step S114: Based on the preset constraint conditions, perform pairwise sampling on the input data sequence and the output data sequence to obtain a sampling sequence pair as the initial linear sample.

[0039] In one embodiment, the above steps are specifically described.

[0040] 1. Input data sequence format.

[0041] Since the input voltage dip occurs, sample the external current sequence for a period of time before the voltage dip , (the sequence values can be measured and stored in advance) and the initial dip voltage .

[0042] 2. Output data sequence format.

[0043] PID controller parameters, that is, when the controller runs: , where (1) The corresponding parameters and (note that the traditional expression expresses as , which is for convenient parameter calculation here).

[0044] When the sampling period is , and the window length is , the above formula (1) can be discretized as: ( (2) And calculate the new current controller input through the feedback circuit: ( (3) Assume that within a period of time Remain unchanged, and the interval between voltage dips is long enough. Based on the above formula (3), by sliding a window (with a window length of ) over the input sequence, and the number of sliding times is ), a series of output voltages are obtained through iterative calculation. In the formula, is the window length for variance calculation.

[0045] Step S120: Based on the initial linear samples, obtain the set of target linear samples through linear interpolation.

[0046] Step S120 specifically includes: Step S121: Perform linear interpolation on the sampling sequence pairs to obtain multiple pairs of sequences to be verified. Substitute each pair of sequences to be verified in the multiple pairs of sequences to be verified into the simulation circuit model respectively. When the operation result meets the preset constraint conditions, determine the corresponding pair of sequences to be verified as the target linear samples.

[0047] Step S122: Use the target linear samples and the initial linear samples to form the set of target linear samples.

[0048] Based on the above embodiments, the above steps are specifically described.

[0049] 3. Generation of linear samples.

[0050] 3.1. Through simulation software (such as Matlab, etc.), establish a simulation circuit model including a PID controller. Assume that the model includes the situation of voltage dip (voltage dip time t0). By inputting different PID controller parameters, test whether the circuit can achieve voltage recovery within the specified time td and maintain stability, and use it as the filtering condition for samples (i.e., the above preset constraint conditions), as shown in formulas (4) and (5): (4) And (5) 3.2. Collect the input-output sequences that meet formulas (4) and (5), and perform pairwise sampling on this sequence. Perform linear interpolation on the sampling sequence pairs to obtain new input-output sequences. For example: ( And ( (6) 3.3. Calculate the output voltage sequence corresponding to the new sequence based on Equation (6), substitute the new sequence into the circuit model, and determine whether its operation result satisfies Equation (4) and Equation (5). If it satisfies, add it to the input-output sequence set (i.e., the above-mentioned target linear sample set).

[0051] Step S130. Construct the reinforcement learning algorithm model based on the target linear sample set; wherein, the reinforcement learning algorithm model includes: system state, possible actions and policies, and the corresponding reward function.

[0052] Based on the above embodiments, step S130 is specifically described.

[0053] 4. Establish a reinforcement learning algorithm model.

[0054] Use the input-output sequence set to train the corresponding input-output sequence model to obtain the reinforcement learning algorithm model between the input sequence and the output sequence.

[0055] 5. Parameter estimation.

[0056] Generate a new input sequence, and based on the feedforward network model, obtain the corresponding output sequence, that is, the corresponding PID control parameters.

[0057] Step 200. Compare the performance result of the target output sequence in the simulation circuit model with the performance result set of the PID control parameter set in the simulation circuit model, and determine the optimal PID parameters based on the comparison result; wherein, the simulation circuit model is: a simulation model including a PID controller.

[0058] Specifically, after step 200 compares the performance result of the target output sequence in the simulation circuit model with the performance result set of the PID control parameter set in the simulation circuit model, the method further includes: In the case where the comparison result indicates that any performance result in the performance result set is better than the performance result of the target output sequence in the simulation circuit model, change the parameters of the reinforcement learning algorithm model or increase the amount of basic training sample data.

[0059] Based on the above embodiments, the above steps are specifically described.

[0060] 6. Performance comparison.

[0061] Send the relevant input sequence and output sequence into the simulation circuit model, and compare with the performance generated by other PID control parameter sets. If the algorithm performance is not as good as other algorithms, change the parameters of the neural network (number of layers, number of neural nodes in each layer), or increase the amount of input sample data.

[0062] In addition, it should be noted that: For step 4, the reinforcement learning includes system state data , and the system action parameters include and .

[0063] Input the three control parameters into the circuit model, as shown in Eqs. (7) and (8).

[0064] (7) (8) Obtain the output sequence for a future period through iteration and sliding window, and perform training optimization based on the relevant metrics obtained from Eqs. (9), (10), and (11). Use Eq. (12) as the reward model: (9) (10) (11) Among them, Eq. (9) ensures that the system reaches stability in the shortest time, Eq. (10) ensures that the system does not exhibit excessive oscillations, and Eq. (11) ensures that the steady-state system voltage returns to a reasonable range. represents the entire voltage training space.

[0065] The reward of the reinforcement learning can be set as: (12) The smaller this reward is, the better, that is, in the case of the system restoring stability, a shorter stabilization time and oscillation amplitude are obtained.

[0066] For step 4, the reinforcement learning can adopt value-based or policy-based reinforcement learning algorithms, such as A3C (Asynchronous Advantage Actor-Critic), which currently has the optimal performance. The Actor network selects actions, that is, changes the PID control parameter values, and the Critic evaluates the actions. Through the reward function, the transient voltage oscillation amplitude of the PID circuit is minimized and it enters the steady state in a timely manner (the steady-state voltage fluctuation is lower than the threshold, calculated based on the variance of the voltage sequence), and the relevant reward corresponds to this function.

[0067] The specific learning process of the A3C enhanced network is as follows: The enhanced network includes two roles, namely the agent and the environment. The agent consists of a combination of multiple distributed Actor networks and Critic networks, and the environment includes a central control unit and a cache unit. First, the agent inputs the system parameters into the Actor network. Each agent Actor selects an appropriate action, and the action is evaluated by the Critic network. The state change caused by the action is transmitted to the environment side. The environment side calculates the relevant rewards and caches the relevant information. After obtaining a sufficient number of samples, the environment side samples the cached information, and the action and reward information of the relevant samples are transmitted to the agent side for further processing. The Critic network of the agent side changes the action distribution probability according to the reward situation, returns the relevant evaluation to the Actor network, thereby generating a new action, and updates the relevant reward evaluation and cached samples through the environment side. This iteration continues until the condition for ending the action is met, and finally, the optimal action and relevant policies are generated.

[0068] To further enhance the performance of reinforcement learning, a target Actor and target Critic learning module can be added below the Actor network and Critic network. By regularly updating the network information of the Actor network and Critic network to the target Actor and target Critic modules, the reset and update of relevant policies can be achieved, reducing the impact of old policies on the network learning ability, and thus generating optimal policies.

[0069] The Actor network and Critic network can be composed of multiple layers of deep feedforward networks. This network includes an input layer, multiple hidden layers, and an output layer. A fully connected connection is used between each layer. A residual network can be used from the input layer to the hidden layer to avoid the problem of gradient disappearance. The activation function from the input layer to the hidden layer and between hidden layers can use the Sigmoid or ReLU function, and the softmax function can be used from the hidden layer to the output layer.

[0070] Regarding step 6, the final reward of the reinforcement learning model is , corresponding to the optimal PI system parameters and .

[0071] The PID parameter optimization method for voltage sag in the AC / DC system provided in this embodiment collects basic training sample data through PID circuit simulation, retains the data samples that pass the constraint conditions, forms new samples for randomly selected sample pairs based on linear interpolation, and filters out the data samples that do not meet the constraint conditions. Using these samples, an enhanced learning algorithm model is built. The algorithm model includes system states, possible actions and strategies, and corresponding reward functions. Combining linear interpolation and enhanced learning, the relevant model should achieve the following ultimate goals: 1. The overshoot oscillation of the model should be minimized as much as possible.

[0072] 2. The model should restore the stable state as early as possible.

[0073] 3. The voltage after stabilization is within a reasonable range.

[0074] Then, the trained network model is used to optimize the PID system parameters, obtain the estimated values of the PID parameters in the corresponding scenario, and compare the performance with other similar algorithms. The comparison criteria are the same as the above goals.

[0075] The above is the step description of the PID parameter optimization method for voltage sag in the AC / DC system provided by the present invention. It can be seen from the description of the above steps that according to the PID parameter optimization method for voltage sag in the AC / DC system provided by the present invention, by obtaining a target input sequence and inputting the target input sequence into a pre-constructed enhanced learning algorithm model, a corresponding target output sequence is obtained; wherein, the enhanced learning algorithm model is constructed based on a target linear sample set when the AC / DC system has a voltage sag; comparing the performance result of the target output sequence in the simulation circuit model with the performance result set of the PID control parameter set in the simulation circuit model, and determining the optimal PID parameters based on the comparison result; wherein, the simulation circuit model is a simulation model including a PID controller. It can be seen that the present invention greatly improves the number of samples through linear sampling technology for the problem of insufficient network training samples. The enhanced learning algorithm model built based on the target linear samples ensures the performance of the neural network while avoiding overfitting and underfitting caused by insufficient samples, obtains the optimal PID control parameters, overcomes the defects of slow speed and obvious oscillation of traditional algorithms, and significantly improves the response speed and accuracy of the circuit.

[0076] Based on the above embodiment, in this embodiment, another embodiment of the PID parameter optimization method for voltage sag in the AC / DC system provided by the present invention is specifically described.

[0077] When there is an AC / DC voltage dip in the energy Internet, the system needs to perform reactive voltage compensation based on UPQC or SVC. The compensation follows the PID control principle and can be deployed at the AC end, that is, the input source end of the PWM device. When performing PID control, it is first necessary to determine the relevant control coefficients of proportional, integral, and differential. While ensuring that the voltage is restored to stability in a timely manner, reduce the overshoot oscillation of the voltage curve to ensure the safety and stability of the system and equipment and extend the equipment life.

[0078] According to the algorithm flow designed in this embodiment, as Figure 2 shown, the system first determines the input-output sequence format, then generates new qualified samples by means of linear interpolation, trains the reinforcement learning network model with sample learning, designs the model actions and rewards reasonably, obtains the optimal parameter estimation of the PID control coefficients according to the model, and compares the results with other reference algorithms. If the performance is inferior to the reference algorithm, the relevant model parameters can be changed to further improve the system performance.

[0079] The samples for determining the PID-related control coefficients can be generated by simulation software, but the complex constraint conditions greatly reduce the generation efficiency of system samples. Therefore, a small number of reference samples that meet the constraints can be generated first, and then a large number of qualified reference samples can be generated based on the method of linear interpolation. Figure 3 is the schematic diagram of linear sample generation provided by the present invention. The reference samples are as Figure 3 shown.

[0080] Then, a high-performance reinforcement learning algorithm can be used to analyze the complex corresponding relationship between the sample input part (current sequence and initial voltage) and the output part (PID control parameters), and based on the relevant reward function, obtain the optimal optimization result of the PID parameters. Figure 4 is the schematic diagram of the reinforcement learning algorithm model provided by the present invention. Figure 5 is the flowchart of the A3C algorithm provided by the present invention. The reinforcement learning method can adopt the most advanced A3C reinforcement learning network model, and its conceptual model is as Figure 4 shown, and its execution steps can refer to Figure 5 .

[0081] Figure 6 is the schematic diagram of the PID control principle provided by the present invention. The overall system model is as Figure 6 shown. Here, the system output control unit is omitted, which is a simplification of the control process of the voltage dip system. If the influence of the control unit on the PID parameter values needs to be considered, the relevant control functions can be represented by a black box or a white box for iterative estimation of voltage and current in the future for a period of time. If there is a complex system output control unit that is difficult to describe, the digital twin technology can be further used to generate relevant voltage and current sequence estimations for calculation and optimization.

[0082] To verify the model performance, the obtained PID parameters are used in the actual operating system model, and the performance of this algorithm is evaluated based on the objectives such as real-time performance, stability, and robustness set in this embodiment, pointing the way for the next step of research.

[0083] The PID parameter optimization method for voltage sags in AC / DC systems provided by the embodiments of the present invention overcomes the defects of slow speed and obvious oscillation of traditional algorithms, significantly improving the response speed and accuracy of the circuit; generating a sufficient number of network training samples based on linear interpolation and making the samples meet the necessary constraint conditions, avoiding overfitting while improving performance, and significantly increasing the generation speed of effective samples; ensuring that the trained samples reach the ideal optimal performance by setting reasonable optimization objectives and constraint conditions; designing a high-performance A3C learning network based on reinforcement learning to obtain the optimal PID control parameters.

[0084] The PID parameter optimization device for voltage sags in AC / DC systems provided by the present invention will be described below. The PID parameter optimization device for voltage sags in AC / DC systems described below can be correspondingly referred to the PID parameter optimization method for voltage sags in AC / DC systems described above.

[0085] Figure 7 It is a schematic structural diagram of the PID parameter optimization device for voltage sags in AC / DC systems provided by the present invention. As Figure 7 shown, the PID parameter optimization device for voltage sags in AC / DC systems provided by the present invention includes: An acquisition module 701, configured to acquire a target input sequence and input the target input sequence into a pre-constructed reinforcement learning algorithm model to obtain a corresponding target output sequence; wherein, the reinforcement learning algorithm model is constructed based on a target linear sample set when a voltage sag occurs in the AC / DC system; An optimization module 702, configured to compare the performance result of the target output sequence in the simulation circuit model with the performance result set of the PID control parameter set in the simulation circuit model, and determine the optimal PID parameter based on the comparison result; wherein, the simulation circuit model is a simulation model including a PID controller.

[0086] The PID parameter optimization device for voltage sag in the AC / DC system provided by the present invention obtains a target input sequence and inputs the target input sequence into a pre-constructed reinforcement learning algorithm model to obtain a corresponding target output sequence; wherein, the reinforcement learning algorithm model is constructed based on a target linear sample set of the AC / DC system under the condition of voltage sag; comparing the performance result of the target output sequence in the simulation circuit model with the performance result set of the PID control parameter set in the simulation circuit model, and determining the optimal PID parameter based on the comparison result; wherein, the simulation circuit model is a simulation model including a PID controller. It can be seen that the present invention greatly improves the number of samples through linear sampling technology for the problem of insufficient network training samples. The reinforcement learning algorithm model built based on the target linear samples ensures the performance of the neural network while avoiding overfitting and underfitting caused by insufficient samples, obtains the optimal PID control parameters, overcomes the defects of slow speed and obvious oscillation of the traditional algorithm, and significantly improves the reaction speed and accuracy of the circuit.

[0087] Based on the above embodiment, in this embodiment, the device further includes a training module, which is specifically used for: In the case of voltage sag in the AC / DC system, collect basic training sample data through PID circuit simulation, and screen the basic training sample data based on preset constraint conditions to obtain initial linear samples; Based on the initial linear samples, obtain the target linear sample set by means of linear interpolation; Construct the reinforcement learning algorithm model based on the target linear sample set; wherein, the reinforcement learning algorithm model includes: system state, possible actions and strategies, and corresponding reward functions.

[0088] Based on the above embodiment, in this embodiment, the device further includes a collection module, which is specifically used for: In the case of voltage sag in the AC / DC system, obtain an input data sequence and an initial sag voltage; wherein, the input data sequence is an external current sequence in the first preset time period before the voltage sag; Based on the input data sequence and the initial sag voltage, iteratively calculate to obtain an output data sequence through a preset formula.

[0089] Based on the above embodiment, in this embodiment, the device further includes a screening module, which is specifically used for: Obtain the simulation circuit model including the PID controller established in advance, and determine the preset constraint conditions of the basic training sample data according to the voltage recovery situation of the circuit in the second preset time period; wherein, the simulation circuit model includes the situation of voltage sag; Pairwise sampling is performed on the input data sequence and the output data sequence based on the preset constraint conditions to obtain a sampling sequence pair, which serves as the initial linear sample.

[0090] Based on the above embodiments, in this embodiment, the device further includes a linear interpolation module, which is specifically configured to: Perform linear interpolation on the sampling sequence pair to obtain a plurality of sequence pairs to be verified, and respectively input each sequence pair to be verified in the plurality of sequence pairs to be verified into the simulation circuit model. When the operation result meets the preset constraint conditions, determine the corresponding sequence pair to be verified as the target linear sample; Form the target linear sample set through the target linear sample and the initial linear sample.

[0091] Based on the above embodiments, in this embodiment, the device further includes an adjustment module, which is specifically configured to: After comparing the performance result of the target output sequence in the simulation circuit model with the performance result set of the PID control parameter set in the simulation circuit model, when the comparison result indicates that any performance result in the performance result set is better than the performance result of the target output sequence in the simulation circuit model, change the parameters of the reinforcement learning algorithm model or increase the amount of basic training sample data.

[0092] Figure 8 An example of the physical structure diagram of an electronic device is shown as Figure 8 shown. The electronic device can be a robot or other electronic device. The electronic device may include: a processor 810, a communication interface 820, a memory 830, and a communication bus 840. Among them, the processor 810, the communication interface 820, and the memory 830 complete mutual communication through the communication bus 840. The processor 810 can call the logical instructions in the memory 830 to execute the PID parameter optimization method for the voltage dip of the AC / DC system, including: Obtain a target input sequence, and input the target input sequence into a pre-constructed reinforcement learning algorithm model to obtain a corresponding target output sequence; wherein, the reinforcement learning algorithm model is constructed based on the target linear sample set when the AC / DC system has a voltage dip; Compare the performance result of the target output sequence in the simulation circuit model with the performance result set of the PID control parameter set in the simulation circuit model, and determine the optimal PID parameter based on the comparison result; wherein, the simulation circuit model is a simulation model including a PID controller.

[0093] In addition, when the logical instructions in the above-mentioned memory 830 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0094] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the PID parameter optimization method for the voltage sag of the AC / DC system provided by the above-mentioned various methods, including: Obtain a target input sequence and input the target input sequence into a pre-constructed reinforcement learning algorithm model to obtain a corresponding target output sequence; wherein, the reinforcement learning algorithm model is constructed based on a target linear sample set of the AC / DC system under the condition of voltage sag. Compare the performance result of the target output sequence in the simulation circuit model with the performance result set of the PID control parameter set in the simulation circuit model, and determine the optimal PID parameter based on the comparison result; wherein, the simulation circuit model is a simulation model including a PID controller.

[0095] On yet another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the PID parameter optimization method for the voltage sag of the AC / DC system provided by the above-mentioned various methods, including: Obtain a target input sequence and input the target input sequence into a pre-constructed reinforcement learning algorithm model to obtain a corresponding target output sequence; wherein, the reinforcement learning algorithm model is constructed based on a target linear sample set of the AC / DC system under the condition of voltage sag. Compare the performance results of the target output sequence in the simulation circuit model with the set of performance results of the PID control parameter set in the simulation circuit model, and determine the optimal PID parameters based on the comparison results; wherein, the simulation circuit model is: a simulation model including a PID controller.

[0096] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative effort.

[0097] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0098] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A PID parameter optimization method for AC / DC system voltage drop, characterized in that: include: Obtain a target input sequence, and input the target input sequence into a pre-built reinforcement learning algorithm model to obtain a corresponding target output sequence; wherein the reinforcement learning algorithm model is constructed based on a target linear sample set of an AC / DC system in the event of a voltage drop; The performance results of the target output sequence in the simulation circuit model are compared with the performance result set of the PID control parameter set in the simulation circuit model, and the optimal PID parameters are determined based on the comparison results; wherein the simulation circuit model is: a simulation model including a PID controller.

2. The PID parameter optimization method for AC / DC system voltage drop according to claim 1, characterized in that: The construction process of the reinforcement learning algorithm model includes: In the case of voltage drop in the AC / DC system, basic training sample data is collected through PID circuit simulation, and the basic training sample data is screened based on preset constraints to obtain initial linear samples; Based on the initial linear samples, obtaining the target linear sample set by linear interpolation; The reinforcement learning algorithm model is constructed based on the target linear sample set; wherein the reinforcement learning algorithm model includes: system state, possible actions and strategies, and corresponding benefit function.

3. The PID parameter optimization method for AC / DC system voltage drop according to claim 2, characterized in that: In the case of voltage drop in the AC / DC system, basic training sample data is collected through PID circuit simulation, including: When a voltage drop occurs in the AC / DC system, an input data sequence and an initial drop voltage are obtained; wherein the input data sequence is: an external current sequence in a first preset time period before the voltage drop; Based on the input data sequence and the initial drop voltage, an output data sequence is obtained by iterative calculation using a preset formula.

4. The PID parameter optimization method for AC / DC system voltage drop according to claim 3, characterized in that: The step of screening the basic training sample data based on the preset constraint conditions to obtain an initial linear sample includes: Acquire the pre-established simulation circuit model including the PID controller, and determine the preset constraint condition of the basic training sample data according to the voltage recovery of the circuit within the second preset time period; wherein the simulation circuit model includes the voltage drop condition; Based on the preset constraint condition, the input data sequence and the output data sequence are sampled in pairs to obtain sampling sequence pairs as the initial linear samples.

5. The PID parameter optimization method for AC / DC system voltage drop according to claim 4, characterized in that: The step of obtaining the target linear sample set by linear interpolation based on the initial linear samples includes: Performing linear interpolation on the sampling sequence pair to obtain a plurality of sequence pairs to be verified, bringing each of the plurality of sequence pairs to be verified into the simulation circuit model, and determining the corresponding sequence pair to be verified as a target linear sample when the operation result meets the preset constraint condition; The target linear sample set is formed by the target linear sample and the initial linear sample.

6. The PID parameter optimization method for AC / DC system voltage drop according to claim 1, characterized in that: After comparing the performance result of the target output sequence in the simulation circuit model with the performance result set of the PID control parameter set in the simulation circuit model, the method further includes: When the comparison result indicates that any performance result in the performance result set is better than the performance result of the target output sequence in the simulation circuit model, the parameters of the reinforcement learning algorithm model are changed or the amount of basic training sample data is increased.

7. A PID parameter optimization device for AC / DC system voltage drop, characterized in that: include: An acquisition module is used to acquire a target input sequence and input the target input sequence into a pre-built reinforcement learning algorithm model to obtain a corresponding target output sequence; wherein the reinforcement learning algorithm model is constructed based on a target linear sample set of an AC / DC system in the event of a voltage drop; An optimization module is used to compare the performance results of the target output sequence in the simulation circuit model with the performance result set of the PID control parameter set in the simulation circuit model, and determine the optimal PID parameters based on the comparison results; wherein the simulation circuit model is: a simulation model including a PID controller.

8. The PID parameter optimization device for AC / DC system voltage drop according to claim 7, characterized in that: The device also includes a training module, which is specifically used for: In the case of voltage drop in the AC / DC system, basic training sample data is collected through PID circuit simulation, and the basic training sample data is screened based on preset constraints to obtain initial linear samples; Based on the initial linear samples, obtaining the target linear sample set by linear interpolation; The reinforcement learning algorithm model is constructed based on the target linear sample set; wherein the reinforcement learning algorithm model includes: system state, possible actions and strategies, and corresponding benefit function.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the PID parameter optimization method for AC / DC system voltage drop according to any one of claims 1 to 6 is implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the PID parameter optimization method for AC / DC system voltage drop according to any one of claims 1 to 6 is implemented.