DC load MOS tube voltage regulation method based on AC reinforcement learning
By applying AC reinforcement learning strategy network and value network in MOS tube voltage regulation, the problem of inaccurate adjustment of traditional methods in complex environments is solved, and more efficient and accurate voltage regulation is achieved.
Patent Information
- Application Number
- CN202510457965.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-14
AI Technical Summary
The traditional MOS tube voltage regulation method based on fixed parameter PID control or open-loop modulation is difficult to adapt to under load a sudden change or nonlinear operating conditions, resulting in overshoot or response delays, and it is difficult to achieve refined control.
Using AC reinforcement learning method, the policy network A and value network C are designed, and the model parameters are optimized using the environmental reward and TD algorithm to achieve accurate adjustment of the gate voltage of the MOS tube.
It improves the reliability and fault diagnosis capabilities of the motor, realizes the accuracy and efficiency of MOS tube voltage regulation, and adapts to complex dynamic environments.
Smart Images

Figure CN119995353A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the intersection field of power electronic control technology and artificial intelligence, and specifically relates to a DC load MOS tube voltage regulation method based on AC reinforcement learning. Background Art
[0002] In recent years, with the rapid development of power electronics technology, DC load systems have been widely used in the fields of renewable energy generation, electric vehicles, industrial automation, etc. MOS tubes are often used as core devices for DC load voltage regulation due to their fast switching speed and low conduction loss. However, the traditional MOS tube voltage regulation method based on fixed parameter PID control or open-loop modulation has significant limitations.
[0003] The limitations come from different aspects. When there is a sudden load change or nonlinear working condition, such as temperature change causing resistance drift, frequent manual parameter adjustment is required, which is prone to overshoot or response delay. The rule-based control strategy has limited coverage scenarios and cannot adapt to complex dynamic environments.
[0004] Some studies have tried to introduce artificial intelligence technology, but there are still obvious defects: the traditional CNN neural network relies on convolution operations to extract spatial features, but current, voltage and other signals are essentially time series data, and CNN is difficult to capture dynamic time series dependencies. The classical reinforcement learning action space design is rigid, making it difficult to achieve fine control of continuous current, and lacks explicit modeling of physical constraints.
[0005] In view of the above problems, the present invention proposes a DC load MOS tube voltage regulation method based on AC reinforcement learning, which provides important support for the intelligent and automated development of industrial production and promotes the intelligent and efficient development of industrial production. Summary of the invention
[0006] The purpose of the present invention is to cope with complex situations and provide a DC load MOS tube voltage regulation method based on AC reinforcement learning, so as to effectively improve the reliability of the motor, timely and accurately diagnose the complex faults in the motor, and increase the service life of the motor.
[0007] In order to solve the above technical problems, the present invention provides a method for adjusting the gate voltage of a MOS tube in an electronic load based on reinforcement learning, comprising: AC reinforcement learning designs a strategy network A and a value network C, wherein the strategy network A and the value network C are composed of a transformer and a fully connected layer; Adjust the action Interact with the environment Environmental rewards in state , the reward value is designed based on the difference between the target current and the current value of the regulated DC load; The new environment state Input policy network A to get But it is not executed. , Input the value network C to get the evaluation at time t+1 ; Using the output of the value network C and environmental rewards , the TD error is obtained through the TD algorithm and the loss function is calculated, and the model parameters of the value network C are updated using the gradient descent algorithm to make the evaluation of the value network C more accurate; Use gradient ascent to update the model parameters of the policy network A so that the output action accumulates the discounted reward Higher directional optimization, resulting in more accurate regulation of gate voltage.
[0008] Optionally, AC reinforcement learning designs a strategy network A and a value network C, wherein the strategy network A and the value network C are composed of a transformer and a fully connected layer, including: DC load current value and MOS tube gate voltage value Composition of environmental status , the environmental state sequence { } Input strategy network A Output regulation action ;Will{ } and action sequence { } Input value network C output pair Reviews .
[0009] Optionally, the adjustment action Interact with the environment Environmental rewards in state , the reward value is designed based on the difference between the target current and the current value of the regulated DC load, including: Policy network A outputs the optimal action And interact with the environment to get a new environment state and environmental rewards ; In order to accurately track the target value of the DC load current, it is necessary to design a reward rule: ① When the DC load current value adjusted by the strategy network A is closer to the target current, the environment reward is a positive number and will be larger, otherwise it will be smaller. ② If overshoot occurs, the reward value is negative, and the more the target current value is exceeded, the greater the penalty. ③ Set the difference between the target current and the current value of the adjusted DC load to , the specified reward value is ,in, is a step function.
[0010] Optionally, the new environment state Input policy network A to get But it is not executed. , Input the value network C to get the evaluation at time t+1 ,include: definition For passing The adjusted MOS tube gate voltage value, For passing The adjusted DC load current value will be the new environmental state Input the policy network A, get the action distribution probability value and select the optimal action But it is not executed. and Input Value Network C Output ,in, , C represents the value network C, Represents the C model parameters.
[0011] Optionally, the utilization value network C outputs and environmental rewards , the TD error is obtained through the TD algorithm and the loss function is calculated. The model parameters of the value network C are updated using the gradient descent algorithm to make the evaluation of the value network C more accurate, including: based on , and ,The TD algorithm is used to calculate the TD error and loss function, and the scoring target of the TD algorithm is designed as , the loss function is ,in is the TD error. The smaller the loss function, the more accurate the evaluation of the value network C. Use gradient descent to update the model parameters of the value network C. Right now ,in, are the updated model parameters, The learning rate can be adjusted by yourself. , the value network C will reward the environment As a scoring target It is an important part of the neural network, which allows the neural network to simulate the actual environment more accurately by continuously updating the model parameters.
[0012] Optionally, the gradient ascent is used to update the model parameters of the strategy network A so that the output action moves toward the discounted cumulative reward Higher directional optimization, resulting in more accurate regulation of gate voltage, including: Using Participating TD Error Update the model parameters of the strategy network A and use the gradient ascent algorithm to update the parameter formula: ,in, is the learning rate, is the policy network A model parameter, , is the policy function; repeat the above steps to maximize the discounted cumulative reward from the initial state to the final state by updating the parameters ,Design discount cumulative reward in, is a discount factor used to balance immediate rewards with future rewards, usually , The closer it is to 1, the more emphasis is placed on long-term rewards and the less short-term problems are avoided. The value is 0.99. Representatives in The environmental rewards at each moment are used to maximize the expected value of long-term cumulative rewards through the coordinated optimization of the value network C and the strategy network A, so as to achieve precise regulation of the gate voltage of the electronic DC load MOS tube.
[0013] The present invention provides a DC load MOS tube voltage regulation method based on AC reinforcement learning, comprising: AC reinforcement learning design strategy network A and value network C; Interact with the environment Environmental rewards in state , the reward value is designed based on the difference between the target current and the current value of the adjusted DC load; the new environmental state Input policy network A to get But it is not executed. , Input value network C to get ; Using the output of value network C , and environmental rewards , the TD error is obtained through the TD algorithm and the loss function is calculated. The gradient descent algorithm is used to update the model parameters of the value network C to make the evaluation of the value network C more accurate; the gradient ascent is used to update the model parameters of the strategy network A so that the output action accumulates the reward towards the discount Higher directional optimization, resulting in more accurate regulation of gate voltage. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to more clearly illustrate the embodiments of the present invention or the technical solutions of the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0015] Figure 1 A schematic diagram of a DC load MOS tube gate voltage regulation process provided in an embodiment of the present application; Figure 2 A schematic diagram of the interaction relationship between the strategy network, value network and environment provided in the embodiment of the present application; Figure 3 A schematic diagram of the strategy network and value network model architecture provided for the embodiments of the present application; Figure 4 A schematic diagram of the cyclic flow of the AC reinforcement learning algorithm and neural network update provided in an embodiment of the present application. DETAILED DESCRIPTION
[0016] In order to enable those skilled in the art to better understand the scheme of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific implementation methods. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0017] As shown in FIG. 1 , FIG. 1 is a schematic diagram of a DC load MOS tube gate voltage regulation process provided by an embodiment of the present application, which specifically includes five contents.
[0018] S11: AC reinforcement learning designs a strategy network A and a value network C, wherein the strategy network A and the value network C are composed of a transformer and a fully connected layer.
[0019] It should be noted that the strategy network A is used to output precise MOS tube voltage regulation action. , the value network C functions as the output of the evaluation of the regulatory action , guiding the decision of the optimization strategy network A; using transformer to extract environmental features to realize time series processing and capture long-distance dependencies.
[0020] Step 11: Construct a collaborative optimization architecture of the policy network A and the value network C. The policy network A is responsible for generating the voltage regulation strategy, and the value network C performs the strategy value evaluation. The two networks extract environmental features through the Transformer encoder respectively, and finally output the decision results through the fully connected layer;
[0021] Step 12: The specific process of constructing the strategy network A is as follows: ① Input layer, the sampled DC load current sequence { } and the corresponding MOS tube gate voltage sequence { } as a feature vector input Transformer, ② Transformer encoding layer, build three sets of parallel self-attention modules , and Perform multi-head attention calculation and model the input features, as shown in formula (1), and then perform multi-head feature fusion, as shown in formula (2): In the formula, is the Query matrix of the i-th attention head, is the Key matrix of the i-th attention head, is the Value matrix of the i-th attention head, is the input feature matrix, Concat is the multi-head output concatenated along the feature dimension, is the output projection matrix. ③ The fully connected decision layer consists of three layers of deep neural network. The first layer uses ReLU activation function to realize 256→128 dimensional nonlinear transformation, as shown in formula (3). The second layer uses Layer Normalization for feature normalization, as shown in formula (4). The third output layer uses Softmax function to generate voltage regulation action probability distribution, as shown in formula (5): in, , is the fully connected weight matrix, is the action space mapping matrix, , and is the bias term, Softmax is the normalized exponential function;
[0022] Step 13: The specific process of constructing the value network C is as follows: ① The joint input layer is responsible for receiving the action vector { } and the environment status { },②Feature interaction layer: Use state features as query and action features as key-value, establish state-action association model, and output interaction features , as shown in formula (6), the LayerNorm function in formula (6) is shown in formula (7), and the attention feature calculation in formula (6) is shown in formula (8): In the formula, , To learn the scaling and offset parameters, and is the mean and standard deviation of the feature dimension, is the scaling factor, , , , s is the state feature vector, a is the action feature vector, ③ value regression layer, based on the interaction feature As input, a residual fully connected structure is used to set 128→64→32 dimension reduction channels, as shown in formulas (9) and (10), and the linear activation function is used to output the current action Reviews , as shown in formula (11): in, is the dimension reduction weight matrix, is the residual layer weight matrix, is the value regression weight, and b is the bias term.
[0023] Based on the above discussion, in an optional embodiment of the present application, AC reinforcement learning designs a strategy network A and a value network C, wherein the strategy network A and the value network C are composed of a transformer and a fully connected layer and specifically include: The transformer multi-head attention mechanism sets the attention head k=8, the hidden layer dimension d_model=512, the number of layers n_layers=6, the feedforward network ffn_hidden=2048, , initialized as an orthogonal matrix, the scaling factor is , corresponding to 8 attention heads.
[0024] S12: Adjust the action Interact with the environment Environmental rewards in state , the reward value is designed based on the difference between the target current and the adjusted DC load current value.
[0025] It should be noted that in order to accurately track the target value of the DC load current, a reward rule needs to be designed, which is as follows: when the DC load current value adjusted by the strategy network A is closer to the target current, the environment reward The larger the current value is, the smaller it will be. If overshoot occurs, the reward value will be negative, and the greater the current value exceeds the target value, the greater the penalty will be.
[0026] Step 21: During initialization, set the initial current value of the DC load And the corresponding MOS tube gate voltage value As Input strategy network A to obtain the probability value of the adjusted action and select the optimal action And execute, as shown in formula (12): in, is the policy network A parameter, It is calculated by the above Transformer encoding layer and fully connected decision layer formula;
[0027] Step 22: Interact with the environment to get a new environment state and environmental rewards , based on the above regulations, the reward value is designed, and the formula is: In the formula, is the difference between the target current and the current value of the DC load after adjustment, is a step function;
[0028] Step 23: and Input value network C output pair action Reviews The formula used to evaluate the value of the decision action of the policy network A is: In the formula, C represents the value network C, represents the initialization parameters of the value network C, It is calculated and generated through the above-mentioned feature interaction layer and value regression layer formula.
[0029] Based on the above discussion, in an optional embodiment of the present application, j will adjust the action Interact with the environment Environmental rewards in state The specific process of designing the reward value based on the difference between the target current and the current value of the regulated DC load includes: As shown in the above formula (3), the weight matrix Using He normal initialization, as shown in the above formula (4), the weight matrix Using LeCun normal initialization, as shown in the above formula (5), the weight matrix Zero-mean Gaussian initialization is used to limit the entropy of the initial action probability distribution. The optimizer type is AdamW and the initial learning rate is , the weight decay coefficient is 0.01.
[0030] S13: The new environment state Input policy network A to get But it is not executed. , Input the value network C to get the evaluation at time t+1 ;
[0031] It should be noted that in order to update the parameters of the A value network C model, it is necessary to use Calculate TD error; policy network A only makes one decision output in each cycle and execute, while Only its value is used but not executed.
[0032] Step 31: Collect new state sequence { , }and{ , Input strategy network A, output the adjustment action with the largest probability value But it is not executed. The formula is: In the formula, are policy network parameters, Generated by the above Transformer encoding layer calculation formula and the fully connected decision layer calculation formula;
[0033] Step 32: Transform the action vector { , } Input the value network C, extract the action features in the joint input layer and concatenate them with the environment features output by the transformer to form a joint feature, and output the pair Reviews , used to update the A value network C parameters, the formula is: In the formula, C represents the value network C, Represents the C model initialization parameters, Generated through the above-mentioned feature interaction layer and value regression layer.
[0034] Based on the above discussion, in an optional embodiment of the present application, for the new environment state Input policy network A to get But it is not executed. , Input the value network C to get the evaluation at time t+1 Specifically include: It is necessary to set up an action temporary storage buffer to block before the value evaluation is completed. The transmission of the value regression layer weight matrix Use Glorot uniform initialization, range , , the weight matrix is initialized as a scaling of the identity matrix, Initialized to a zero-mean Gaussian distribution with a standard deviation of 0.1. Initialization is 0.1, the optimizer type is RMSprop, and the initial learning rate is , the attenuation factor is 0.99.
[0035] S14: Using the output of value network C and environmental rewards , the TD error is obtained through the TD algorithm and the loss function is calculated, and the gradient descent algorithm is used to update the model parameters of the value network C to make the evaluation of the value network C more accurate.
[0036] It should be noted that the TD algorithm can update network parameters without waiting for the end of a full round, and can achieve online real-time updates, so that the system can adjust strategies within nanosecond time scales to avoid adjustment lags; it should be noted that the smaller the TD error, the smaller the loss function, and the stronger the evaluation ability of the value network C; the core of the value network C is to use environmental rewards As an important part of the TD goal, by continuously updating model parameters and optimizing the scoring system, the neural network can simulate the actual environment more accurately.
[0037] Step 41: First set the TD target, the formula is: In the formula, is the reward value of the environment for the current action, is the discount rate, For the value network C Reviews ;
[0038] Step 42: Specify the TD error, which is: In the formula, For the value network C The evaluation value of For TD target;
[0039] Step 43: Set the loss function, whose formula is: In the formula, For the value network C Reviews , For TD target;
[0040] Step 44: Loss function In Status Next search about The partial derivative formula is: In the formula, For the value network C Adjustment action in status evaluate, is the model parameter of the value network C at time t;
[0041] Step 45: Update for the gradient descent algorithm , the formula is: in, is the TD error, is the learning rate, is the model parameter of the value network C.
[0042] Based on the above discussion, in an optional embodiment of the present application, for using the output of the value network C and environmental rewards , the TD error is obtained through the TD algorithm and the loss function is calculated, and the model parameters of the value network C are updated using the gradient descent algorithm to make the evaluation of the value network C more accurate. The specific methods include: As shown in formula (17), is the discount rate, which is 0.99 here, which can optimize the long-term stability. As shown in formula (21), the gradient descent algorithm initializes the learning rate Range is .
[0043] S15: Use the gradient ascent algorithm to update the model parameters of the strategy network A so that the output action is cumulatively rewarded at a discount Higher directional optimization enables more accurate regulation of gate voltage.
[0044] It should be noted that in order to meet the needs of high-precision power device control, the gradient ascent algorithm is used to update the model parameters of the strategy network A, so that the voltage tracking error reaches sub-microsecond accuracy, significantly improving the training efficiency; the reinforcement learning AC algorithm uses the value network C evaluation, and the ultimate goal is to maximize the long-term cumulative rewards through the strategy network A, rather than simply optimizing the rewards of a single adjustment. This design ensures that the system takes into account both immediacy and long-term reliability in a dynamic environment.
[0045] Step 51: Using Participating TD Error , using the gradient ascent algorithm formula: In the formula, is the policy function, is the policy function with respect to the parameter The partial guide, is the TD error, is the learning rate, is the policy network A model parameter;
[0046] Step 52: Repeat the above steps, and after multiple MOS tube gate voltage adjustments, obtain the environmental reward generated by the interaction with the environment after each adjustment. , by updating the policy network A model parameters each time to maximize the discounted cumulative reward from the initial state to the final state ;
[0047] Step 53: Design the long-term cumulative reward formula as follows: In the formula, is a discount factor used to balance immediate rewards with future rewards, Representatives in Momentary environmental rewards.
[0048] Based on the above discussion, in an optional embodiment of the present application, the model parameters of the strategy network A are updated using the gradient ascent algorithm so that the output action is directed to the discounted cumulative reward The specific process of higher directional optimization, thus achieving more accurate regulation of gate voltage, includes: generally , The closer it is to 1, the more emphasis is placed on long-term rewards. To avoid short-term problems, here The value is 0.99, the learning rate of the gradient ascent algorithm for The training data of the A-value network C is dynamically generated through the real-time interaction between the agent and the environment. Its core is the interactive experience, rather than a fixed data set collected in advance, that is, the action a generated by the policy network A, the state s of the environment, and the reward r given by the environment. The data is generated in real time during the training process and is directly used to update the network.
[0049] The present application uses specific examples to illustrate the principles and implementation methods of the present invention, and the description of the above embodiments is only used to help the method and core ideas of the present invention. It should be pointed out that for ordinary people in the technical field, without departing from the principles of the present invention, the present invention can also be improved and modified, and these improvements and modifications also fall within the scope of protection of the claims of the present invention.
Claims
1. A DC load MOS tube voltage regulation method based on AC reinforcement learning, characterized in that: include: AC reinforcement learning designs a strategy network A and a value network C, wherein the strategy network A and the value network C are composed of a transformer and a fully connected layer; Adjust the action Interact with the environment Environmental rewards in state , designing the reward value based on the difference between the target current and the current value of the regulated DC load; The new environment state Input policy network A to get But it is not executed. , Input the value network C to get the evaluation at time t+1 ; Using the output of the value network C and environmental rewards , the TD error is obtained through the TD algorithm and the loss function is calculated, and the gradient descent algorithm is used to update the model parameters of the value network C to make the evaluation of the value network C more accurate; Use gradient ascent to update the model parameters of the policy network A so that the output action accumulates the discounted reward Higher directional optimization, resulting in more accurate regulation of gate voltage.
2. The DC load MOS tube voltage regulation method based on AC reinforcement learning as described in claim 1, characterized in that: AC reinforcement learning designs a strategy network A and a value network C. The strategy network A and the value network C are composed of a transformer and a fully connected layer, including: DC load current value and MOS tube gate voltage value Composition of environmental status , the environmental state sequence { } Input strategy network A Output regulation action ;Will{ } and action sequence { } Input value network C output pair Reviews .
3. The DC load MOS tube voltage regulation method based on AC reinforcement learning as described in claim 1, characterized in that: Adjust the action Interact with the environment Environmental rewards in state , the reward value is designed based on the difference between the target current and the current value of the regulated DC load, including: Policy network A outputs the optimal action And interact with the environment to get a new environment state and environmental rewards ; In order to accurately track the target value of the DC load current, it is necessary to design a reward rule: ① When the DC load current value adjusted by the strategy network A is closer to the target current, the environment reward is a positive number and will be larger, otherwise it will be smaller. ② If overshoot occurs, the reward value is negative, and the more the target current value is exceeded, the greater the penalty. ③ Set the difference between the target current and the current value of the adjusted DC load to , the specified reward value is ,in, is a step function.
4. The DC load MOS tube voltage regulation method based on AC reinforcement learning as described in claim 1, characterized in that: The new environment state Input policy network A to get But it is not executed. , Input the value network C to get the evaluation at time t+1 ,include: definition For passing The adjusted MOS tube gate voltage value, For passing The adjusted DC load current value will be the new environmental state Input the policy network A, get the action distribution probability value and select the optimal action But it is not executed. and Input Value Network C Output ,in, , C represents the value network C, Represents the C model parameters.
5. The DC load MOS tube voltage regulation method based on AC reinforcement learning as described in claim 1, characterized in that: Using the output of the value network C Environmental rewards , the TD error is obtained through the TD algorithm and the loss function is calculated. The model parameters of the value network C are updated using the gradient descent algorithm to make the evaluation of the value network C more accurate, including: based on , and ,The TD algorithm is used to calculate the TD error and loss function, and the scoring target of the TD algorithm is designed as , the loss function is ,in is the TD error, is the weight value. The smaller the loss function, the more accurate the evaluation of the value network C. Use gradient descent to update the model parameters of the value network C. Right now ,in, are the updated model parameters, The learning rate can be adjusted by yourself. , the value network C will reward the environment As a scoring target It is an important part of the neural network, which allows the neural network to simulate the actual environment more accurately by continuously updating the model parameters.
6. The DC load MOS tube voltage regulation method based on AC reinforcement learning as described in claim 1, characterized in that: Use gradient ascent to update the model parameters of the policy network A so that the output action accumulates the discounted reward Higher directional optimization, resulting in more accurate regulation of gate voltage, including: Using Participating TD Error Update the model parameters of the strategy network A and use the gradient ascent algorithm to update the parameter formula: ,in is the TD error, is the learning rate, is the policy network A model parameter, , is the policy function; repeat the above steps to maximize the discounted cumulative reward from the initial state to the final state by updating the parameters ,Design discount cumulative reward in, is a discount factor used to balance immediate rewards with future rewards, usually , The closer it is to 1, the more emphasis is placed on long-term rewards and the less short-term problems are avoided. The value is 0.
99. Representatives in The environmental rewards at each moment are used to maximize the expected value of long-term cumulative rewards through the coordinated optimization of the value network C and the strategy network A, so as to achieve precise regulation of the gate voltage of the electronic DC load MOS tube.
Citation Information
Patent Citations
DCDC buck converter discrete sliding mode control algorithm based on genetic algorithm
CN116760289A
Cloud edge computing task scheduling method based on reinforcement learning
CN118740835A
Network information age optimization method and system based on meta-deep reinforcement learning
CN119135551A
Current emulation in a power supply
EP3902130A1
Cited By
Chip layout model training method, chip layout method and related device
CN122174775A