Static synchronous compensator control method and device based on BERT-Transform collaborative architecture DDPG, computer equipment and medium
Through the BERT-Transformer collaborative architecture DDPG method, the problem of inaccurate control of the stationary synchronous compensator during dynamic changes in the power grid is solved, the stability of the power grid and the quality of the power grid are improved, and the reactive power compensation is optimized.
Patent Information
- Application Number
- CN202510721010.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-05-30
AI Technical Summary
The existing reinforcement learning control method has poor reactive power compensation control effect when the power grid changes dynamically, and the action network and the critic network are underfitted, resulting in inaccurate control and low grid stability and power quality.
The deep deterministic strategy gradient DDPG method based on the BERT-Transformer collaborative architecture is adopted, and the voltage and current data of the static synchronization compensator is processed in a coordinated manner through the BERT network as the action network and the Transformer network as the critic network, and the voltage and current data of the static synchronization compensator are processed accurately, and reactive power compensation is optimized.
Accurate control of static synchronous compensator is achieved, reactive power compensation effect of the power grid is optimized, grid stability and power quality are improved, and energy consumption is reduced.
Smart Images

Figure CN120262452A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of new energy, large models, power system control, power electronics technology, and smart grid, and relates to a reinforcement learning control method, which is applicable to the reactive power compensation control of a static synchronous compensator in a power system. Background Art
[0002] In existing reinforcement learning control methods, there is a problem of poor reactive power compensation control effect of a static synchronous compensator when the power grid dynamically changes.
[0003] In addition, there is an underfitting problem in the action network and the critic network of existing reinforcement learning control methods when facing dynamic control problems. Summary of the Invention
[0004] Based on this, in view of the above technical problems, it is necessary to provide a static synchronous compensator control method, device, computer device, computer-readable storage medium, and computer program product based on the BERT-Transformer collaborative architecture DDPG.
[0005] In a first aspect, the present application provides a static synchronous compensator control method based on the BERT-Transformer collaborative architecture DDPG. The method includes:
[0006] Collect voltage and current data and cumulative value data from a static synchronous compensator, and normalize the voltage and current data to obtain normalized voltage and current data, where the voltage and current data includes d-axis current, q-axis current, DC voltage, positive sequence voltage on the AC side, positive sequence voltage error on the AC side, DC voltage error, d-axis current error, and q-axis current error, and the cumulative value data includes the cumulative value of the positive sequence voltage error on the AC side, the cumulative value of the DC voltage error, the cumulative value of the d-axis current error, and the cumulative value of the q-axis current error;
[0007] Use the cumulative value data and the normalized voltage and current data as state vectors to input into the current action network to obtain an action vector, and apply the action vector output by the current action network to the static synchronous compensator. The current action network is a BERT network, and the BERT network is a bidirectional encoder representation based on Transformer;
[0008] Input the state vector and the action vector together into the current critic network to evaluate the value of the current state-action pair. The current critic network is a Transformer network;
[0009] Use deep deterministic policy gradient DDPG to update the weights and bias parameters of the current critic network and the current action network.
[0010] In a second aspect, the present application also provides a static synchronous compensator control device based on the BERT-Transformer collaborative architecture DDPG. The device includes:
[0011] A data preprocessing module, configured to collect voltage and current data and cumulative value data from a static synchronous compensator, and normalize the voltage and current data to obtain normalized voltage and current data, where the voltage and current data includes d-axis current, q-axis current, DC voltage, positive-sequence voltage on the AC side, positive-sequence voltage error on the AC side, DC voltage error, d-axis current error, and q-axis current error, and the cumulative value data includes the cumulative value of the positive-sequence voltage error on the AC side, the cumulative value of the DC voltage error, the cumulative value of the d-axis current error, and the cumulative value of the q-axis current error;
[0012] An action network output action module, configured to use the cumulative value data and the normalized voltage and current data as state vectors to input into the current action network to obtain an action vector, and apply the action vector output by the current action network to the static synchronous compensator, where the current action network is a BERT network, and the BERT network is a bidirectional encoder representation based on Transformer;
[0013] A critic network evaluation module, configured to input the state vector and the action vector together into the current critic network to evaluate the value of the current state-action pair, where the current critic network is a Transformer network;
[0014] A network update and optimization module, configured to update the weights and bias parameters of the current critic network and the current action network by using deep deterministic policy gradient DDPG.
[0015] In a third aspect, the present application also provides a computer device. The computer device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0016] Collect voltage and current data and cumulative value data from a static synchronous compensator, and normalize the voltage and current data to obtain normalized voltage and current data, where the voltage and current data includes d-axis current, q-axis current, DC voltage, positive-sequence voltage on the AC side, positive-sequence voltage error on the AC side, DC voltage error, d-axis current error, and q-axis current error, and the cumulative value data includes the cumulative value of the positive-sequence voltage error on the AC side, the cumulative value of the DC voltage error, the cumulative value of the d-axis current error, and the cumulative value of the q-axis current error;
[0017] The cumulative value data and the normalized voltage and current data are used as state vectors to input the current action network, obtaining an action vector. The action vector output by the current action network is applied to the static synchronous compensator. The current action network is a BERT network, and the BERT network is a bidirectional encoder representation based on Transformer;
[0018] The state vector and the action vector are input into the current critic network together to evaluate the value of the current state-action pair. The current critic network is a Transformer network;
[0019] The weights and bias parameters of the current critic network and the current action network are updated using Deep Deterministic Policy Gradient (DDPG).
[0020] Fourthly, the present application also provides a computer-readable storage medium. On the computer-readable storage medium, there is a computer program stored, and when the computer program is executed by a processor, the following steps are implemented:
[0021] Voltage and current data and cumulative value data are collected from the static synchronous compensator, and the voltage and current data are normalized to obtain the normalized voltage and current data. The voltage and current data include d-axis current, q-axis current, DC voltage, AC side positive sequence voltage, AC side positive sequence voltage error, DC voltage error, d-axis current error, and q-axis current error. The cumulative value data includes the cumulative value of the AC side positive sequence voltage error, the cumulative value of the DC voltage error, the cumulative value of the d-axis current error, and the cumulative value of the q-axis current error;
[0022] The cumulative value data and the normalized voltage and current data are used as state vectors to input the current action network, obtaining an action vector. The action vector output by the current action network is applied to the static synchronous compensator. The current action network is a BERT network, and the BERT network is a bidirectional encoder representation based on Transformer;
[0023] The state vector and the action vector are input into the current critic network together to evaluate the value of the current state-action pair. The current critic network is a Transformer network;
[0024] The weights and bias parameters of the current critic network and the current action network are updated using Deep Deterministic Policy Gradient (DDPG).
[0025] Fifthly, the present application also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the following steps are implemented:
[0026] Collect voltage and current data and cumulative value data from a static synchronous compensator, normalize the voltage and current data to obtain the normalized voltage and current data, where the voltage and current data include d-axis current, q-axis current, DC voltage, positive-sequence voltage on the AC side, positive-sequence voltage error on the AC side, DC voltage error, d-axis current error, and q-axis current error; the cumulative value data includes the cumulative value of the positive-sequence voltage error on the AC side, the cumulative value of the DC voltage error, the cumulative value of the d-axis current error, and the cumulative value of the q-axis current error;
[0027] Take the cumulative value data and the normalized voltage and current data as the state vector and input them into the current actor network to obtain the action vector, and apply the action vector output by the current actor network to the static synchronous compensator. The current actor network is a BERT network, and the BERT network is a bidirectional encoder representation based on Transformer;
[0028] Input the state vector and the action vector together into the current critic network to evaluate the value of the current state-action pair. The current critic network is a Transformer network;
[0029] Use Deep Deterministic Policy Gradient (DDPG) to update the weights and bias parameters of the current critic network and the current actor network.
[0030] The above static synchronous compensator control method, device, computer device, storage medium and computer program product based on the BERT-Transformer collaborative architecture DDPG collect voltage and current data and cumulative value data from the static synchronous compensator, normalize the voltage and current data to obtain the normalized voltage and current data, where the voltage and current data include d-axis current, q-axis current, DC voltage, positive-sequence voltage on the AC side, positive-sequence voltage error on the AC side, DC voltage error, d-axis current error and q-axis current error; the cumulative value data includes the cumulative value of the positive-sequence voltage error on the AC side, the cumulative value of the DC voltage error, the cumulative value of the d-axis current error and the cumulative value of the q-axis current error; input the cumulative value data and the normalized voltage and current data as the state vector into the current action network to obtain the action vector, and apply the action vector output by the current action network to the static synchronous compensator, where the current action network is a BERT network, and the BERT network is a bidirectional encoder representation based on Transformer; input the state vector and the action vector into the current critic network to evaluate the value of the current state-action pair, where the current critic network is a Transformer network; use the deep deterministic policy gradient DDPG to update the weights and bias parameters of the current critic network and the current action network; the bidirectional encoder representation BERT network based on Transformer and the Transformer network work together to achieve precise control of the static synchronous compensator, have the function of optimizing the reactive power compensation effect of the power grid, can improve the stability and power quality of the power grid, and reduce the energy consumption of the power grid. Brief Description of the Drawings
[0031] Figure 1 is a flowchart of the method of the present invention.
[0032] Figure 2 is a control framework diagram of the static synchronous compensator of the method of the present invention.
[0033] Figure 3 is an internal structure diagram of a computer device in an embodiment. Detailed Description of the Embodiments
[0034] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0035] In one embodiment, as Figure 1As shown, a flowchart of a control method for a static synchronous compensator based on the BERT-Transformer collaborative architecture DDPG is provided. In this embodiment, this method is exemplified by being applied to a terminal. It can be understood that this method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:
[0036] The present invention proposes a control method for a static synchronous compensator based on the BERT-Transformer collaborative architecture DDPG. The deep deterministic policy gradient DDPG has an action network and a critic network. The action network includes a current action network and a target action network, and the critic network includes a current critic network and a target critic network; the Bidirectional Encoder Representations from Transformers (BERT) network based on Transformer is used as the action network of the deep deterministic policy gradient DDPG, and the Transformer network is used as the critic network of the deep deterministic policy gradient DDPG. The BERT network based on Transformer and the Transformer network work together to achieve precise control of the static synchronous compensator, having the function of optimizing the reactive power compensation effect of the power grid, improving the stability and power quality of the power grid, and reducing the energy consumption of the power grid; the steps in the use process are as follows:
[0037] Step (1): Data preprocessing;
[0038] Collect voltage and current data and cumulative value data from the static synchronous compensator, and normalize the voltage and current data to obtain the normalized voltage and current data. The voltage and current data include d-axis current, q-axis current, DC voltage, positive sequence voltage on the AC side, positive sequence voltage error on the AC side, DC voltage error, d-axis current error, and q-axis current error. The cumulative value data includes the cumulative value of the positive sequence voltage error on the AC side, the cumulative value of the DC voltage error, the cumulative value of the d-axis current error, and the cumulative value of the q-axis current error;
[0039] The voltage and current data collected from the static synchronous compensator include , , , , , , , , , , and ; is the d-axis current, is the q-axis current, is the DC voltage, is the positive-sequence voltage on the AC side, is the positive-sequence voltage error on the AC side, is the DC voltage error, is the d-axis current error, is the q-axis current error, is the cumulative value of the positive-sequence voltage error on the AC side, is the cumulative value of the DC voltage error, is the cumulative value of the d-axis current error, is the cumulative value of the q-axis current error;
[0040] The normalization operation normalizes the non-normalized input by rescaling it to the interval between 0 and 1; the non-normalized input is That is, normalize the voltage and current data;
[0041] while , , and do not need to be normalized. The cumulative value data includes , , and ;
[0042] The process of normalizing the non-normalized input is: .
[0043] Among them, is the set of the minimum values of each element of the non-normalized input, is the set of the maximum values of each element of the non-normalized input, is the normalized input, that is, the normalized , , , , , , and , that is, the normalized voltage and current data;
[0044] Step (2): The action network outputs an action;
[0045] Take the cumulative value data and the normalized voltage and current data as the state vector and input them into the current action network to obtain an action vector. Apply the action vector output by the current action network to the static synchronous compensator. The current action network is a BERT network, and the BERT network is a bidirectional encoder representation based on Transformer;
[0046] , , , and the normalized input After being input into the Deep Deterministic Policy Gradient (DDPG), a state vector is constructed. The state vector is: The state vector is input into the current action network, which is a Bidirectional Encoder Representations from Transformers (BERT) network; in the Bidirectional Encoder Representations from Transformers (BERT) network, the state vector undergoes neural network transformation to abstract and transform the information of the state vector layer by layer, extracting a feature representation to provide a basis for subsequent action decisions; after being processed by the Bidirectional Encoder Representations from Transformers (BERT) network, an action vector is output. The action vector is: , is the q-axis current reference value, is the d-axis current reference value, is the d-axis voltage, is the q-axis voltage; that is, the previous action network outputs the q-axis current reference value, d-axis current reference value, d-axis voltage, and q-axis voltage. That is, the Bidirectional Encoder Representations from Transformers (BERT) network outputs the q-axis current reference value, d-axis current reference value, d-axis voltage, and q-axis voltage. That is, the action vector obtained after inputting the cumulative value data and the normalized voltage and current data into the current action network is the q-axis current reference value, d-axis current reference value, d-axis voltage, and q-axis voltage.
[0047] The Bidirectional Encoder Representations from Transformers (BERT) network includes:
[0048] An input layer configured to receive the original feature input and map it into a feature embedding vector, where the original features include at least one of numerical features, categorical features, or sequential features;
[0049] A position encoding module configured to generate a position encoding vector matching the dimension of the feature embedding vector and add it to the feature embedding vector to represent the position information of the feature;
[0050] A multi-layer bidirectional Transformer encoder, where each layer of the Transformer encoder includes a multi-head self-attention mechanism and a feed-forward neural network. The multi-head self-attention mechanism allows each feature to interact bidirectionally with all other features and calculates the attention weights through query matrices, key matrices, and value matrices. The feed-forward neural network performs a non-linear transformation on the attention output;
[0051] The layer normalization and residual connection module is configured to perform layer normalization operations before and after the multi-head self-attention mechanism and the feed-forward neural network respectively, and introduce residual connections to enhance feature transmission;
[0052] The output layer is configured to aggregate the outputs of the multi-layer bidirectional Transformer encoder to generate a processed feature output, and the aggregation method includes at least one of global average pooling, global max pooling, or weighted summation.
[0053] Apply the action vector output by the current action network To the static synchronous compensator to control the operation of the static synchronous compensator. The static synchronous compensator adjusts the reactive power output according to the control action, thereby controlling the voltage level and power flow of the power grid; After the current action network outputs the action vector The operating state of the power grid changes, and the state information is updated And the corresponding reward signal; the updated state information is obtained by re-collection and after the data preprocessing in step (1) to obtain the target state vector As the input of the current action network in the next control loop; the reward signal is set according to the control target and the power grid operation index, and is used to measure the contribution degree of the current control action to achieving the control target. For example, in implementation, the reward signal Can be designed as: Or Etc.
[0054] Step (3): Critic network evaluation;
[0055] Input the state vector and the action vector into the current critic network to evaluate the value of the current state-action pair. The current critic network is a Transformer network;
[0056] Input the state vector And the action vector Together into the current critic network. The current critic network is a Transformer network. The Transformer network can process the information of the state vector And the action vector To evaluate the value of the current state-action pair; in the Transformer network, the state-action pair is jointly analyzed and evaluated. The Transformer network outputs a Value, that is, the value of the current state-action pair. This Value reflects the action vector output by the BERT network based on the Transformer's bidirectional encoder representation under the current state vector The long-term expected return that can be obtained, that is, the action vector output by the BERT network based on the Transformer bidirectional encoder representation is quantitatively evaluated for its advantages and disadvantages; the value output by the Transformer network is used as a feedback signal to guide the update and optimization of the BERT network based on the Transformer bidirectional encoder representation; The higher the value, the more conducive the current action vector
[0057] is to achieve the control goal under the evaluation of the Transformer network;
[0058] The Transformer network includes but is not limited to:
[0059] An input encoding module configured to receive raw feature data and generate feature embedding vectors. The raw feature data includes at least one of numerical, categorical, or time-series features. The input encoding module includes a feature embedding layer and a position encoding layer. The position encoding layer generates position encoding vectors through sine or cosine functions and adds them to the feature embedding vectors;
[0059] An encoder stack composed of multiple encoder layers with the same structure stacked together. Each encoder layer includes:
[0060] A multi-head self-attention sublayer configured to calculate the correlation weights between input features. The multi-head self-attention sublayer generates query, key, and value matrices by linearly transforming the input features and calculates attention scores through a scaled dot-product attention mechanism;
[0061] Multiple residual connection and layer normalization sublayers configured to apply residual connection and layer normalization operations before and after the multi-head self-attention sublayer respectively;
[0062] A feed-forward neural network sublayer including two linear transformation layers and a non-linear activation function;
[0063] A decoder stack composed of multiple decoder layers with the same structure stacked together. Each decoder layer includes:
[0064] A masked multi-head self-attention sublayer configured to prevent attention to features at future positions;
[0065] An encoder-decoder attention sublayer configured to associate the encoder output with the current decoder state;
[0066] A feed-forward neural network sublayer;
[0067] An output decoding module configured to convert the output of the decoder stack into a target feature representation. The output decoding module includes a linear transformation layer and an optional activation function layer.
[0068] Step (4): Network update and optimization;
[0069] The weights and bias parameters of the current critic network and the current actor network are updated using Deep Deterministic Policy Gradient (DDPG).
[0070] Based on the value output by the current critic network and the reward signal generated by the change in the grid operation state , the weights and bias parameters of the current actor network are optimized. The weights and bias parameters of the current actor network after optimization using DDPG are: , where is the learning rate, is the gradient operator, is the number of samples;
[0071] The goal of optimization is to make the optimized current actor network output a better action vector than the non-optimized current actor network, so as to maximize the long-term cumulative reward, thereby achieving precise control of the static synchronous compensator;
[0072] Based on the reward signal generated by the change in the grid operation state and the value output by the target critic network , the weights and bias parameters of the current critic network are updated using the loss function . This loss function can be designed as: , where is the discount factor; the update process of the value output by the Transformer network of the target critic network and the current critic network is both .
[0073] The process of updating the weights and bias parameters of the current critic network is: .
[0074] The updated current critic network can evaluate the value of the state-action pair more accurately than the non-updated current critic network; value;
[0075] The weights and bias parameters of the target actor network and the target critic network gradually approach those of the current actor network and the current critic network with a smoothing factor to achieve smooth synchronization between the target actor network and the target critic network and the current actor network and the current critic network; the synchronization process is: and , where are the weight and bias parameters of the target critic network, and
[0076]
[0077] By synchronizing the target actor network and the target critic network, the control strategy is optimized to make the control of the static synchronous compensator stable and reliable; in subsequent control loops, the target actor network and the target critic network provide accurate target values for the update of the current actor network and the current critic network, promoting the improvement of the control performance of the entire static synchronous compensator control system. The present invention has the following advantages and effects compared with the prior art:
[0078] (1) Existing traditional control methods for controlling a static synchronous compensator have problems of poor control effect and inaccurate control actions when facing grid dynamic problems. The present invention proposes a control method for a static synchronous compensator based on the BERT-Transformer collaborative architecture DDPG, which can solve the control problems of the static synchronous compensator when facing grid dynamic problems and achieve precise control of the static synchronous compensator.
[0079] (2) Existing traditional control methods for controlling a static synchronous compensator have problems of poor reactive power compensation effect, poor grid stability, and low power quality. The present invention proposes a control method for a static synchronous compensator based on the BERT-Transformer collaborative architecture DDPG, which can optimize the reactive power compensation effect of the grid, improve the stability of the grid, and improve the power quality when controlling the static synchronous compensator.
[0080] (3) Existing reinforcement learning methods have problems of underfitting in the actor network and the critic network when controlling a static synchronous compensator. The bidirectional encoder representation BERT network and the Transformer network based on Transformer can solve the problem of underfitting in the reinforcement learning actor network and the critic network.
[0081] Figure 2 is the control framework diagram of the static synchronous compensator of the method of the present invention. The input of the control method based on the BERT-Transformer collaborative architecture DDPG is the positive sequence voltage error on the AC side obtained from the static synchronous compensator , the DC voltage error , the d-axis current error and the q-axis current error , and the output is the q-axis current reference value , the d-axis current reference value , the d-axis voltage and the q-axis voltage to the static synchronous compensator to achieve the control of the static synchronous compensator.
[0082] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as shown in Figure 3 . The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a static synchronous compensator control method based on the BERT-Transformer collaborative architecture DDPG. The display unit of the computer device is used to form a visually visible picture, which may be a display screen, a projection device, or a virtual reality imaging device. The display screen may be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device may be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0083] Those skilled in the art can understand that Figure 3 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0084] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0085] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0086] In one embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0087] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.
[0088] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0089] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.
[0090] The embodiments described above merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation to the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A control method for a static synchronous compensator based on the BERT-Transformer collaborative architecture DDPG, characterized in that: Voltage and current data and cumulative value data are collected from the static synchronous compensator, and the voltage and current data are normalized to obtain normalized voltage and current data. The voltage and current data include d-axis current, q-axis current, DC voltage, positive-sequence voltage on the AC side, positive-sequence voltage error on the AC side, DC voltage error, d-axis current error, and q-axis current error. The cumulative value data includes the cumulative value of the positive-sequence voltage error on the AC side, the cumulative value of the DC voltage error, the cumulative value of the d-axis current error, and the cumulative value of the q-axis current error; The cumulative value data and the normalized voltage and current data are used as state vectors to be input into the current action network to obtain an action vector, and the action vector output by the current action network is applied to the static synchronous compensator. The current action network is a BERT network, and the BERT network is a bidirectional encoder representation based on Transformer; The state vector and the action vector are input into the current critic network together to evaluate the value of the current state-action pair. The current critic network is a Transformer network; The weights and bias parameters of the current critic network and the current action network are updated using the deep deterministic policy gradient DDPG.
2. The control method of a static synchronous compensator based on the BERT-Transformer collaborative architecture DDPG according to claim 1, characterized in that The action vector is the q-axis current reference value, the d-axis current reference value, the d-axis voltage, and the q-axis voltage.
3. A static synchronous compensator control method based on the BERT-Transformer collaborative architecture DDPG according to claim 1, characterized in that, The BERT network includes an input layer, a position encoding module, a multi-layer bidirectional Transformer encoder, a layer normalization and residual connection module, and an output layer.
4. A static synchronous compensator control method based on the BERT-Transformer collaborative architecture DDPG according to claim 1, characterized in that, The Transformer network includes an input encoding module, an encoder stack, a multi-head self-attention sublayer, multiple residual connection and layer normalization sublayers, a feed-forward neural network sublayer, a decoder stack, a masked multi-head self-attention sublayer, an encoder-decoder attention sublayer, and an output decoding module.
5. A static synchronous compensator control device based on the BERT-Transformer collaborative architecture DDPG, characterized in that, The device includes: A data preprocessing module, configured to collect voltage and current data and cumulative value data from the static synchronous compensator, and normalize the voltage and current data to obtain normalized voltage and current data. The voltage and current data include d-axis current, q-axis current, DC voltage, positive-sequence voltage on the AC side, positive-sequence voltage error on the AC side, DC voltage error, d-axis current error, and q-axis current error. The cumulative value data includes the cumulative value of the positive-sequence voltage error on the AC side, the cumulative value of the DC voltage error, the cumulative value of the d-axis current error, and the cumulative value of the q-axis current error; An action network output action module, configured to use the cumulative value data and the normalized voltage and current data as state vectors to be input into the current action network to obtain an action vector, and apply the action vector output by the current action network to the static synchronous compensator. The current action network is a BERT network, and the BERT network is a bidirectional encoder representation based on Transformer; A critic network evaluation module for inputting a state vector and an action vector into the current critic network together to evaluate the value of the current state-action pair, where the current critic network is a Transformer network; A network update and optimization module for updating the weights and bias parameters of the current critic network and the current action network by using Deep Deterministic Policy Gradient (DDPG).
6. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the static synchronous compensator control method based on the BERT-Transformer collaborative architecture DDPG according to any one of claims 1 to 4.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the static synchronous compensator control method based on the BERT-Transformer collaborative architecture DDPG according to any one of claims 1 to 4.
Citation Information
Patent Citations
Steering compensation control method and device of steer-by-wire system based on DDPG
CN112977606A
Static synchronous compensator control method based on composite deep reinforcement learning
CN115912389A
Server for controlling waybill output and method for controlling waybill output using the same
KR102767830B1