DC Load MOSFET Voltage Regulation Method Based on AC Reinforcement Learning

Through the policy network and value network based on AC reinforcement learning, combined with the timing characteristics of Transformer and full-connection layer processing, the precise adjustment problem of MOS tube voltage regulation in complex environments is solved, and the motor is efficient fault diagnosis and long-term reliability is achieved.

CN119995353BActive Publication Date: 2025-07-08HUNAN NEXT GENERATION INSTRUMENTAL T&C TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510457965.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-08
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

Traditional MOS tube voltage regulation methods are difficult to adapt to complex dynamic environments under load sudden or nonlinear operating conditions, resulting in frequent manual parameter adjustment, overshoot or response delays, and existing artificial intelligence technologies are difficult to achieve refined control of continuous currents and explicit modeling of physical constraints.

Method used

Using AC reinforcement learning method, the strategy network A and value network C are designed, and the model parameters are updated through the TD algorithm and gradient optimization algorithm to realize the precise adjustment of the gate voltage of the MOS tube. Combined with the timing characteristics of the Transformer and the full connection layer to process the timing characteristics, the reward rules are optimized and adjusted.

Benefits of technology

It improves the accuracy of MOS tube voltage regulation and the reliability of the motor, realizes real-time fault diagnosis and efficient current control in complex environments, and improves the service life of the motor and system stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119995353B_ABST
    Figure CN119995353B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for regulating the voltage of a MOS transistor of a DC load based on AC reinforcement learning, including: designing a policy network A and a value network C for AC reinforcement learning; obtaining an environmental reward in a state by interacting an adjustment action with the environment, and designing a reward value based on the difference between the target current and the current value of the DC load after adjustment; inputting a new environmental state into the policy network A but not executing it, and then inputting it into the value network C to obtain an evaluation at the (t + 1)-th moment; using the evaluation output by the value network C and the environmental reward, obtaining a TD error through the TD algorithm and calculating a loss function, and adopting a gradient descent algorithm to update the model parameters of the value network C to make the evaluation of the value network C more accurate; adopting gradient ascent to update the model parameters of the policy network A to optimize the output action in the direction of a higher discounted cumulative reward, so as to realize more accurate regulation of the gate voltage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the cross - field of power electronics control technology and artificial intelligence, and particularly relates to a method for regulating the voltage of a MOS transistor in a DC load based on AC reinforcement learning. Background Art

[0002] In recent years, with the rapid development of power electronics technology, DC load systems have been widely used in fields such as new energy power generation, electric vehicles, and industrial automation. Due to its characteristics such as fast switching speed and low conduction loss, MOS transistors are often used as the core devices for regulating the voltage of DC loads. However, traditional MOS transistor voltage regulation methods based on fixed - parameter PID control or open - loop modulation have significant limitations.

[0003] The limitations come from different aspects. In the case of load mutations or non - linear working conditions, such as when resistance drift is caused by temperature changes, frequent manual parameter adjustment is required, which is prone to overshoot or response delay. Rule - based control strategies have limited coverage scenarios and cannot adapt to complex dynamic environments.

[0004] Some studies have tried to introduce artificial intelligence technology, but there are still obvious defects: traditional CNN neural networks rely on convolutional operations to extract spatial features, but signals such as current and voltage are essentially time - series data, and it is difficult for CNN to capture dynamic time - series dependencies. The action space design of classical reinforcement learning is rigid, making it difficult to achieve refined control of continuous current and lacking explicit modeling of physical constraints.

[0005] To address the above problems, the present invention proposes a method for regulating the voltage of a MOS transistor in a DC load based on AC reinforcement learning, which provides important support for the intelligent and automated development of industrial production and promotes the intelligent and efficient development of industrial production. Summary of the Invention

[0006] The purpose of the present invention is to provide a method for regulating the voltage of a MOS transistor in a DC load based on AC reinforcement learning to effectively improve the reliability of the motor, timely and accurately diagnose compound faults in the motor, and extend the service life of the motor in response to complex situations.

[0007] To solve the above - mentioned technical problems, the present invention provides a method for regulating the gate voltage of a MOS transistor in an electronic load based on reinforcement learning, including:

[0008] Designing a policy network A and a value network C for AC reinforcement learning, where the policy network A and the value network C are composed of a transformer and fully - connected layers;

[0009] Taking the adjustment action Interacting with the environment to obtain The environmental reward in the state , design a reward value based on the difference between the target current and the current value of the regulated DC load;

[0010] Input the new environmental state into the policy network A to obtain but do not execute it. Then, , Input it into the value network C to obtain the evaluation at time t+1 ;

[0011] Utilize the output of the value network C and the environmental reward , obtain the TD error through the TD algorithm and calculate the loss function, and use the gradient descent algorithm to update the model parameters of the value network C to make the evaluation of the value network C more accurate;

[0012] Use gradient ascent to update the model parameters of the policy network A to optimize the output action towards the direction of higher discounted cumulative reward , thereby achieving more accurate regulation of the gate voltage.

[0013] Optionally, the AC reinforcement learning designs the policy network A and the value network C, and the policy network A and the value network C are composed of a transformer and a fully connected layer, including:

[0014] The DC load current value and the MOS transistor gate voltage value constitute the environmental state , input the environmental state sequence { } into the policy network A to output the regulation action ; input { } and the action sequence { } into the value network C to output the evaluation of . .

[0015] Optionally, the regulation action interacts with the environment to obtain the environmental reward in the state , and design a reward value based on the difference between the target current and the current value of the regulated DC load, including:

[0016] The policy network A outputs the optimal action and interacts with the environment to obtain the new environmental state and the environmental reward ; to achieve accurate tracking of the target value of the DC load current, a reward rule needs to be designed: ① when the DC load current value regulated by the policy network A is closer to the target current, then the environmental reward is a positive number and will be larger, otherwise it will be smaller. ② If overshoot occurs, the reward value is negative, and the more it exceeds the target current value, the greater the penalty. ③ Set the difference between the target current and the current value of the regulated DC load as , and the reward value is specified as , where is a step function.

[0017] Optionally, input the new environmental state into the policy network A to obtain but do not execute it. Then, input and into the value network C to obtain the evaluation at time t + 1 , including:

[0018] Define as the gate voltage value of the MOS transistor after being regulated by , Define as the DC load current value after being regulated by . Input the new environmental state into the policy network A to obtain the action distribution probability value and select the optimal action but do not execute it. Input and into the value network C to output , where represents the C model parameters of the value network C.

[0019] Optionally, using the output by the value network C and the environmental reward , calculate the TD error and the loss function through the TD algorithm, and use the gradient descent algorithm to update the model parameters of the value network C to make the evaluation of the value network C more accurate, including:

[0020] Based on , and , calculate the TD error and the loss function using the TD algorithm. Design the scoring target of the TD algorithm as , and the loss function is , where is the TD error. The smaller the loss function, the more accurate the evaluation of the value network C; update the model parameters of the value network C using the gradient descent that is , where are the updated model parameters, is the learning rate that can be adjusted by itself, , and the value network C will receive the environmental reward As a scoring target As an important part, by continuously updating the model parameters, the neural network can more accurately simulate the actual environment.

[0021] Optionally, using gradient ascent to update the model parameters of the policy network A makes the output action optimize in the direction of higher discounted cumulative reward to achieve more accurate regulation of the gate voltage, including:[[]]

[0022] Using the TD error involved to update the model parameters of the policy network A. The parameter update formula using the gradient ascent algorithm is , where is the learning rate, are the model parameters of the policy network A, , is the policy function; repeat the above steps to maximize the discounted cumulative reward from the initial state to the final state by updating the parameters , design the discounted cumulative reward where is the discount factor, used to balance immediate rewards and future rewards. Usually , the closer it is to 1, the more it emphasizes long-term rewards and avoids short-sighted problems, takes a value of 0.99, represents the environmental reward at time. Through the collaborative optimization of the value network C and the policy network A, the expected value of the long-term cumulative reward is maximized to achieve precise regulation of the gate voltage of the MOS transistor of the electronic DC load.

[0023] The present invention provides a method for regulating the voltage of a MOS transistor of a DC load based on AC reinforcement learning, including: designing a policy network A and a value network C for AC reinforcement learning; obtaining the environmental reward by interacting the adjustment action with the environment in the state, designing a reward value based on the difference between the target current and the current value of the regulated DC load; inputting the new environmental state into the policy network A to obtain but not executing it, and then inputting , into the value network C to obtain ; using the output of the value network C, and the environmental reward , the TD error is obtained through the TD algorithm and the loss function is calculated. The model parameters of the value network C are updated using the gradient descent algorithm to make the evaluation of the value network C more accurate; the model parameters of the policy network A are updated using gradient ascent to optimize the output action towards the direction of higher discounted cumulative reward so as to achieve more accurate regulation of the gate voltage. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solutions of the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0025] Figure 1 It is a schematic diagram of the DC load MOS transistor gate voltage regulation process provided by the embodiment of the present application;

[0026] Figure 2 It is a schematic diagram of the interaction relationship between the policy network, the value network and the environment provided by the embodiment of the present application;

[0027] Figure 3 It is a schematic diagram of the model architectures of the policy network and the value network provided by the embodiment of the present application;

[0028] Figure 4 It is a schematic diagram of the loop process of the AC reinforcement learning algorithm and neural network update provided by the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0029] In order to enable those skilled in the art to better understand the solution of the present invention, the present invention will be further described in detail below with reference to the drawings and specific embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0030] As shown in FIG. 1, FIG. 1 is a schematic diagram of the DC load MOS transistor gate voltage regulation process provided by the embodiment of the present application, which specifically includes five contents.

[0031] S11: The AC reinforcement learning designs a policy network A and a value network C, and the policy network A and the value network C are composed of a transformer and a fully connected layer.

[0032] It should be noted that the function of the policy network A is to output accurate MOS transistor voltage regulation actions , the value network C functions to output the evaluation of the adjustment action , guiding the decision-making of the optimization policy network A; using a transformer to extract environmental features to achieve time series processing and capture long-distance dependencies.

[0033] Step 11: Construct a collaborative optimization architecture for the policy network A and the value network C. The policy network A is responsible for generating the voltage regulation strategy, and the value network C performs policy value evaluation. The two networks respectively extract environmental features through the Transformer encoder and finally output the decision result through the fully connected layer;

[0034] Step 12: The specific process of constructing the policy network A is as follows: ① Input layer, taking the sampled DC load current sequence { } and the corresponding MOS transistor gate voltage sequence { } as feature vectors and inputting them into the Transformer. ② Transformer encoding layer, constructing three groups of parallel self-attention modules 、 and to perform multi-head attention calculation and model the input features, as shown in formula (1), and then perform multi-head feature fusion, as shown in formula (2):

[0035]

[0036]

[0037] In the formula, is the Query matrix of the i-th attention head, is the Key matrix of the i-th attention head, is the Value matrix of the i-th attention head, is the input feature matrix, Concat is the multi-head output concatenated along the feature dimension, is the projection matrix of the output. ③ Fully connected decision layer, composed of three layers of deep neural networks. The first layer uses the ReLU activation function to achieve a 256→128-dimensional non-linear transformation, as shown in formula (3). The second layer uses LayerNormalization for feature standardization, as shown in formula (4). The third layer output layer uses the Softmax function to generate the voltage regulation action probability distribution, as shown in formula (5):

[0038]

[0039]

[0040]

[0041] Among them, , is the fully connected weight matrix, is the action space mapping matrix, , and are the bias terms, and Softmax is the normalized exponential function;

[0042] Step 13: The specific process of constructing the value network C is as follows: ① The joint input layer is responsible for receiving the cumulative action vectors { } from the policy network A from the initial time to time t and the environmental state { }, ② Feature interaction layer: Using the state feature as the Query and the action feature as the Key-Value, establish a state-action association model and output the interaction feature , as shown in Equation (6). The LayerNorm function in Equation (6) is shown in Equation (7), and the attention feature calculation in Equation (6) is shown in Equation (8):

[0043]

[0044]

[0045]

[0046] In the formula, , are the learnable scaling and offset parameters, and are the mean and standard deviation of the feature dimension, is the scaling factor, , , , s is the state feature vector, a is the action feature vector, ③ Value regression layer, using the interaction feature as the input, adopting a residual fully connected structure to set the 128→64→32 dimensionality reduction channels, as shown in Equations (9) and (10), and output the evaluation of the current action through a linear activation function , as shown in Equation (11):

[0047]

[0048]

[0049]

[0050] Among them, is the dimensionality reduction weight matrix, is the residual layer weight matrix, is the value regression weight, and b is the bias term.

[0051] Based on the above discussion, in an alternative embodiment of the present application, an AC reinforcement learning designs a policy network A and a value network C, and the policy network A and the value network C are composed of a transformer and a fully connected layer, specifically including:

[0052] The transformer multi-head attention mechanism sets the number of attention heads k = 8, the hidden layer dimension d_model = 512, the number of layers n_layers = 6, and the feed-forward network ffn_hidden = 2048. , initialized as an orthogonal matrix, and the scaling factor is , corresponding to 8 attention heads.

[0053] S12: The adjustment action interacts with the environment to obtain the environmental reward at the state, and designs the reward value based on the difference between the target current and the adjusted DC load current value.

[0054] It should be noted that to accurately track the target value of the DC load current, a reward rule needs to be designed, and its content is: when the DC load current value adjusted by the policy network A is closer to the target current, then the environmental reward is a positive number and will be larger, otherwise it will be smaller; if there is an overshoot phenomenon, then the reward value is negative, and the more it exceeds the target current value, the greater the penalty.

[0055] Step 21: At initialization, the initial current value of the DC load and the corresponding MOS transistor gate voltage value are used as the input to the policy network A to obtain the adjustment action probability value and select the optimal action and execute it, as shown in formula (12):

[0056]

[0057] where are the parameters of the policy network A, generated by the above Transformer encoding layer and fully connected decision layer formula calculations;

[0058] Step 22: Interacts with the environment to obtain a new environmental state and environmental reward , and designs the reward value based on the above regulations, and its formula is:

[0059]

[0060] In the formula, The difference between the target current and the current value of the regulated DC load, is a step function;

[0061] Step 23: Input and into the value network C and output the evaluation of the action for evaluating the value of the decision-making action of the policy network A. The formula is:

[0062] In the formula, C represents the value network C, represents the initial parameters of the value network C, which is calculated and generated through the above feature interaction layer and value regression layer formula.

[0063] Based on the above discussion, in an alternative embodiment of the present application, j will perform the adjustment action and interact with the environment to obtain the environmental reward in the state. The specific process of designing the reward value based on the difference between the target current and the current value of the regulated DC load includes:

[0064] As shown in the above formula (3), the weight matrix is initialized using He normal distribution. As shown in the above formula (4), the weight matrix is initialized using LeCun normal distribution. As shown in the above formula (5), the weight matrix is initialized using zero-mean Gaussian distribution to limit the entropy value of the initial action probability distribution. The optimizer type is AdamW, and the initial learning rate is , and the weight decay coefficient is 0.01.

[0065] S13: Input the new environmental state into the policy network A to obtain but do not execute it. Then, input , into the value network C to obtain the evaluation at time t+1;

[0066] It should be noted that to update the model parameters of the value network C of A, the TD error needs to be calculated using . The policy network A only makes one decision output and executes it in each loop, while only uses its value but does not execute it.

[0067] Step 31: Collect the new state sequences { , } and { , ​Input the policy network A and output the adjustment action with the maximum probability value But do not execute it. Its formula is:

[0068]

[0069] In the formula, Are the policy network parameters, Generated by the above Transformer encoding layer calculation formula and the fully connected decision layer calculation formula;

[0070] Step 32: Input the action vector { , } into the value network C. Extract the action features at the joint input layer and splice them with the environmental features output by the transformer to form joint features, and output the evaluation of , which is used to update the parameters of the A value network C. Its formula is: In the formula, C represents the value network C,

[0071]

[0072] In the formula, C represents the value network C, Represents the initialization parameters of the C model, Generated by the above feature interaction layer and value regression layer.

[0073] Based on the above discussion, in an optional embodiment of the present application, for inputting the new environmental state into the policy network A to obtain but not execute it, and then input , into the value network C to obtain the evaluation at time t + 1 Specifically include:

[0074] It is necessary to set an action temporary buffer to block the transmission of before the value evaluation is completed. The weight matrix of the value regression layer is initialized using Glorot uniform initialization, with a range of , , the weight matrix is initialized as a scaled unit matrix, is initialized as a zero-mean Gaussian distribution with a standard deviation of 0.1, is initialized as 0.1, the optimizer type is RMSprop, and the initial learning rate is , and the decay factor is 0.99.

[0075] S14: Utilize the output by the value network C and the environmental reward , the TD error is obtained through the TD algorithm and the loss function is calculated, and the model parameters of the value network C are updated using the gradient descent algorithm to make the evaluation of the value network C more accurate.

[0076] It should be noted that the TD algorithm can update the network parameters without waiting for the end of a complete round, enabling online real-time updates, allowing the system to adjust strategies on a nanosecond time scale and avoiding adjustment lags; it should be noted that the smaller the TD error, the smaller the loss function, and the stronger the evaluation ability of the value network C; the core of the value network C lies in using the environmental reward as an important part of the TD target, and by continuously updating the model parameters and optimizing the scoring system, the neural network can more accurately simulate the actual environment.

[0077] Step 41: First, set the TD target, and its formula is:

[0078]

[0079] In the formula, is the reward value of the environment for the current action, is the discount rate, is the evaluation of the value network C for ; ;

[0080] Step 42: Define the TD error, and its formula is:

[0081]

[0082] In the formula, is the evaluation value of the value network C for ; is the TD target;

[0083] Step 43: Set the loss function, and its formula is:

[0084]

[0085] In the formula, is the evaluation of the value network C for ; , is the TD target;

[0086] Step 44: Take the partial derivative of the loss function with respect to under the state , and its formula is:

[0087]

[0088] In the formula, is the evaluation of the value network C for Adjustment action in the state Evaluate, is the parameter of the value network C model at time t;

[0089] Step 45: Update for the gradient descent algorithm , and its formula is:

[0090]

[0091] Among them, is the TD error, is the learning rate, is the parameter of the value network C model.

[0092] Based on the above discussion, in an optional embodiment of the present application, for the output by the value network C and the environmental reward , the specific method of obtaining the TD error through the TD algorithm and calculating the loss function, and updating the model parameters of the value network C using the gradient descent algorithm to make the evaluation of the value network C more accurate includes:

[0093] As shown in formula (17), is the discount rate, and taking 0.99 here can make the long-term stability optimal. As shown in formula (21), the gradient descent algorithm initializes the learning rate The range is .

[0094] S15: Update the model parameters of the policy network A using the gradient ascent algorithm to optimize the output action in the direction of higher discounted cumulative reward to achieve more accurate regulation of the gate voltage.

[0095] It should be noted that to meet the high-precision power device control requirements, the model parameters of the policy network A are updated using the gradient ascent algorithm to make the voltage tracking error reach sub-microsecond-level accuracy, significantly improving the training efficiency; the reinforcement learning AC algorithm relies on the evaluation of the value network C, and the ultimate goal is to maximize the long-term cumulative reward through the policy network A, rather than simply optimizing the reward for a single adjustment. This design ensures that the system takes into account both immediacy and long-term reliability in a dynamic environment.

[0096] Step 51: Use the TD error participated by , and adopt the gradient ascent algorithm formula:

[0097]

[0098]

[0099] In the formula, is the policy function, is the partial derivative of the policy function with respect to the parameter , is the TD error, is the learning rate, are the model parameters of the policy network A;

[0100] Step 52: Repeat the above steps. After adjusting the gate voltage of the MOS transistor multiple times, obtain the environmental rewards generated by interacting with the environment after each adjustment , and maximize the discounted cumulative reward from the initial state to the final state by updating the model parameters of the policy network A each time ;

[0101] Step 53: Design the long-term cumulative reward formula as:

[0102]

[0103] In the formula, is the discount factor, used to balance the immediate reward and the future reward, represents the environmental reward at time.

[0104] Based on the above discussion, in an alternative embodiment of the present application, for the above process of updating the model parameters of the policy network A using the gradient ascent algorithm to optimize the output action towards the discounted cumulative reward in a higher direction, thereby realizing more accurate adjustment of the gate voltage, specifically includes:

[0105] Generally , the closer is to 1, the more attention is paid to the long-term reward. To avoid the short-sighted problem, here takes a value of 0.99, and the learning rate of the gradient ascent algorithm is

[0106] In the present application, specific examples are used to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.

Claims

1. A method for regulating the voltage of a MOS transistor in a DC load based on AC reinforcement learning, characterized in that, Including: The AC reinforcement learning designs a policy network A and a value network C, and the policy network A and the value network C are composed of a transformer and a fully connected layer; Take the DC load current value and the MOS transistor gate voltage value as the environmental state s t Input the policy network A to output the MOS transistor regulation voltage a t , and use a t to update the original gate voltage to obtain the new environmental state s t+1 and the regulation reward r t ; Input the new environmental state s t+1 into the policy network A to obtain the new regulated voltage a t+1 but do not execute it. Then input s t+1 and a t+1 into the value network C to output the evaluation q t+1 of the regulated voltage a t+1 ; Using q t+1 and adjusting the reward r t , the TD error is obtained through the TD algorithm and the loss function is calculated, and the gradient descent algorithm is used to update the model parameters of the value network C to make the evaluation of the value network C more accurate; The model parameters of the policy network A are updated by gradient ascent to optimize the output regulated voltage in the direction of a higher discounted cumulative reward R t so as to achieve more accurate regulation of the gate voltage.

2. The method for regulating the voltage of the MOS transistor of the DC load based on AC reinforcement learning according to claim 1, characterized in that The AC reinforcement learning designs a policy network A and a value network C, and the policy network A and the value network C are composed of a transformer and a fully connected layer, including: DC load current value i t and MOS transistor gate voltage value u t constitute the environmental state s t , and input the environmental state sequence {s1, s2, …, s t} into the policy network A to output the regulated voltage a t ; input {s1, s2, …, s t} and the regulated voltage sequence {a1, a2, …, a t} into the value network C to output the evaluation q t of a t .

3. A method for regulating the voltage of a MOS transistor of a DC load based on AC reinforcement learning as described in claim 1, characterized in that, Take the DC load current value and the MOS transistor gate voltage value as the environmental state s t The input policy network A outputs the MOS transistor regulation voltage a t , and use a t to update the original gate voltage to obtain the new environmental state s t+1 and the regulation reward r t , including: The policy network A outputs the optimal regulated voltage a t and interacts with the environment to obtain a new environmental state s t+1 and environmental reward r t ; To accurately track the target value of the DC load current, a reward rule needs to be designed: ① When the DC load current value adjusted by the policy network A is closer to the target current, the reward r t is a positive number and will be larger, otherwise it will be smaller. ② If there is an overshoot phenomenon, the reward value is negative, and the more it exceeds the target current value, the greater the penalty. ③ Let the difference between the target current and the current value of the adjusted DC load be Δ, and the reward value is specified as r t = e -Δ ·θ(Δ) + Δ·θ(-Δ), where θ(Δ) is a step function.

4. A method for regulating the voltage of a MOS transistor of a DC load based on AC reinforcement learning as described in claim 1, characterized in that Input the new environmental state s t+1 into the policy network A to obtain the new regulated voltage a t+1 but do not execute it. Then input s t+1 and a t+1 into the value network C to output the evaluation q t+1 of the regulated voltage a t+1 , including: Define u t+1 as the MOS transistor gate voltage value after being adjusted by a t , i t+1 as the DC load current value after being adjusted by a t . Input the new environmental state s t+1 into the policy network A, obtain the adjusted voltage distribution probability value and select the optimal adjustment voltage a t+1 , but do not execute it. Input a t+1 and s t+1 into the value network C and output q t+1 , where q t+1 = C(s t+1 , a t+1 ; w t+1 ). C represents the value network C, and w represents the C model parameters.

5. The method for regulating the voltage of the MOS transistor of the DC load based on AC reinforcement learning as described in claim 1, wherein Using q t+1 and adjusting the reward r t , obtaining the TD error through the TD algorithm and calculating the loss function, and updating the model parameters of the value network C using the gradient descent algorithm to make the evaluation of the value network C more accurate, including: Based on q t , q t+1 and r t , use the TD algorithm to calculate the TD error and the loss function. The scoring target of the TD algorithm is y t = r t + ρ·q t+1 , and the loss function is: where δ t = q t - y t is the TD error, ρ is the weight value, and the smaller the loss function, the more accurate the evaluation of the value network C; use gradient descent to update the model parameters w of the value network C, that is, w t+1 = w t - α·δ t ·d w,t , where w t+1 is the updated model parameter, α is the learning rate that can be adjusted by oneself, The value network C takes the environmental reward r t as an important part of the scoring target y t . By continuously updating the model parameters, the neural network can more accurately simulate the actual environment.

6. A method for regulating the voltage of a MOS transistor in a DC load based on AC reinforcement learning as described in claim 1, characterized in that The model parameters of the policy network A are updated by gradient ascent to optimize the output regulated voltage in the direction of a higher discounted cumulative reward R, so as to achieve a more accurate regulation of the gate voltage, including: t Optimizing in the direction of a higher discounted cumulative reward R, so as to achieve a more accurate regulation of the gate voltage, including: Using the TD error δ t participated by r t to update the model parameters of the policy network A. The parameter update formula using the gradient ascent algorithm is θ t+1 = θ t + β · δ t · d θ,t , where δ t = q t - y t is the TD error, β is the learning rate, θ is the model parameter of the policy network A, is the policy function; Repeat the above steps to maximize the discounted cumulative reward R t from the initial state to the final state by updating the parameters. Design the discounted cumulative reward where is the discount factor, used to balance the immediate reward and the future reward. The closer it is to 1, the more it focuses on the long-term reward and avoids the myopia problem. r t+k represents the environmental reward at time t + k. Through the collaborative optimization of the value network C and the policy network A, maximize the expected value of the long-term cumulative reward to achieve precise regulation of the gate voltage of the MOS tube of the electronic DC load.

Citation Information

Patent Citations

  • DCDC buck converter discrete sliding mode control algorithm based on genetic algorithm

    CN116760289A

  • Current emulation in a power supply

    EP3902130A1