Method for regulating pulling current of electronic load based on deep reinforcement learning
By simulating the electronic load pull-load current regulation process into the Markov decision-making process, the multi-scale jump can be used to separate the convolutional neural network model and reinforcement learning architecture, optimize the gate voltage regulation of MOS tubes, and solve the problems of long and low accuracy of electronic DC load current regulation in the existing technology, and achieve fast and high-precision current regulation.
Patent Information
- Application Number
- CN202510300585.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-03-14
AI Technical Summary
In the prior art, electronic DC load current regulation mainly relies on PID algorithms or fuzzy PID control, and requires manual parameters, which makes the design time-consuming and difficult to ensure control performance, making it difficult to achieve fast and accurate pull-load current regulation.
The electronic load pull-load current regulation process is simulated as a Markov decision-making process, and a multi-scale jump separable convolutional neural network model is used to optimize the gate voltage regulation of the MOS tube through a reinforcement learning architecture, establish a reward objective function, and optimize the agent parameters using a gradient descent algorithm to achieve fast and high-precision current regulation.
The training samples without manual labels are realized, manual participation is reduced, and the learning efficiency and control accuracy of electronic load current regulation are improved, ensuring short adjustment time, small overshoot and small steady-state errors.
Smart Images

Figure CN119834619B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of intelligent control based on deep learning, and specifically relates to a method for regulating the pulling current of an electronic load based on deep reinforcement learning. Background Art
[0002] Controlling the working current of a MOS transistor flowing from the drain to the source by controlling the gate voltage of the MOS transistor is the basic working principle of an electronic DC load. This working current is called the pulling current. The larger the pulling current, the greater the load under the same test voltage. By adjusting the magnitude of the pulling current, the size of the simulated load is controlled. Controlling the pulling current by adjusting the gate voltage of the MOS transistor is an iterative adjustment process. It is necessary to calculate the most appropriate adjustment amount sequence according to the feedback current magnitude by a control algorithm to meet the requirements of short adjustment time, small overshoot, and small steady-state error for pulling current regulation.
[0003] Currently, the regulation of the current of an electronic DC load mainly uses the PID algorithm or the fuzzy PID control algorithm. It is necessary to manually set three parameters: integral, differential, and proportional. Designing an algorithm to ensure an optimal adjustment process takes a long time and it is difficult to guarantee the control performance. With the development of deep neural networks, the machine can achieve self-learning based on historical operation data to solve the above problems, and reinforcement learning with a deep network as the agent can perfectly solve the problem of regulating the pulling current. Summary of the Invention
[0004] The purpose of this application is to design a control algorithm that can autonomously learn based on historical data to achieve fast and accurate regulation of the pulling current, reduce the design time of the control algorithm, and improve its control accuracy.
[0005] To solve the above technical problems, this application proposes a method for regulating the pulling current of an electronic load based on deep reinforcement learning, including:
[0006] Simulate the regulation process of the pulling current of the electronic load as a Markov decision process, use the high-frequency time series signal of the pulling current as the environmental state, and solve the optimal MOS transistor gate voltage adjustment amount sequence by the agent;
[0007] The agent of the reinforcement learning architecture selects a multi-scale jump separable convolutional neural network model;
[0008] Establish an adjustment amount reward value according to the difference between the actual current and the target current, and establish a training data set through multi-step adjustment by the agent;
[0009] Apply weights to the reward value sequence based on the influence degree on the pulling current to establish a reward objective function, and optimize the agent parameters with the maximum reward objective function value as the goal based on the gradient descent algorithm;
[0010] Use the trained agent as a controller, take the environmental state as input, and give the discrete sequence of the optimal adjustment amount of the MOS transistor gate voltage to achieve fast and high-precision adjustment of the load current.
[0011] Optionally, simulate the adjustment process of the electronic load current as a Markov decision process. Take the high-frequency timing signal of the load current as the environmental state, and the agent solves the optimal sequence of MOS transistor gate voltage adjustment amounts, including:
[0012] Take the current and target load currents as the initial state and the final state respectively. Take the high-frequency sampling timing signal of the load current as the observable quantity of the environmental state. Taking the observable quantity as input, the agent gives the discrete sequence of the MOS transistor gate voltage adjustment amounts for controlling the load current size Under the control of the discrete sequence, form the state transition set of the load current Thus, simulate the adjustment process of the load current as a Markov decision process.
[0013] Optionally, the agent of the reinforcement learning architecture selects a multi-scale jump separable convolutional neural network model, including:
[0014] The agent selects a neural network model. After converting the high-frequency current sampling timing signal into a grayscale image, it enters the multi-core convolutional layer. The multi-core convolution selects four different sizes of convolutional kernels: 1*1, 1*1 convolution followed by 3*3 convolution, 1*1 convolution followed by 5*5 convolution, and 1*1 convolution followed by 7*7 convolution. In the second layer of the model, first use 1*1 pointwise convolution to adjust the number of features, then select depthwise separable convolution to extract information, and then use 1*1 pointwise convolution to fuse information to obtain feature information. The skip connection method is used between the input and the feature information to achieve information intercommunication. The feature information enters a channel attention module to optimize the weights of each feature information vector in the second layer.
[0015] Optionally, establish the adjustment amount reward value according to the difference between the actual current and the target current, and establish the training data set through multi-step adjustment of the agent, including:
[0016] To achieve the state transition of the load current from to , the agent needs to complete times of MOS transistor gate voltage adjustments to form the discrete sequence of adjustment amounts . After times of cyclic adjustments, the agent forms discrete sequences. The feedback value of the load current after applying each adjustment amount to the circuit is . Calculate the target current Difference from the feedback currentValue If after adjustment is close to the target value, then give multiply by the ratio + Otherwise, multiply by the ratio - to obtain the reward value for each adjustment amount Sum up all the reward values of an adjustment amount sequence after weighting to obtain the reward value , For each adjustment amount sequence, obtain a reward value , The reward value and the adjustment amount sequence form the training data set for reinforcement learning.
[0017] Optionally, establish a reward objective function by applying weights to the reward value sequence based on the influence degree on the pulling current, and optimize the agent parameters with the maximum value of the reward objective function as the goal based on the gradient descent algorithm, including:
[0018] Propose an optimization method for the agent model parameters with a double-loop structure. The outer loop is current states, and the inner loop is adjustment amount sequences. At each state, it is necessary to sample a set of training sample sequences, When optimizing the model parameters in a state, do not consider the reward values before a certain moment , and at the same time consider t the reward values after a certain moment For the influence of the state on the model parameters gradually weakens. According to the influence degree on the model parameters, add weights to the reward value sequence to obtain the reinforcement learning reward objective function , represents the influence weight of the operation amount on the reward value, represents the state occurrence moment, with a total of states, is a constant variable, means take the current state, represents that when entering the state, only when the current state is the case, the reward value weight is the largest at 1, and the reward weights of the subsequent states gradually decrease; then use to calculate the reward gradient value in the state, where represents the sample, with a total of sampling samples, represents the th sample in Adjustment amount in a state The probability of occurrence, based on the parameter update formula Update Model parameters in a state, Represents the average value of the reward, preventing all rewards from being positive.
[0019] Optionally, use the trained agent as a controller, take the environmental state as the input, and give the discrete sequence of the optimal adjustment amount of the MOS transistor gate voltage to achieve fast and high-precision adjustment of the load current, including:
[0020] Take The high-frequency sampling timing signal of the load current in a state is input into the trained agent model, and the model outputs the optimal adjustment amount of the MOS transistor gate voltage , after applying this adjustment amount, the current enters a state, and through continuous iteration, the model outputs a voltage adjustment sequence with the shortest adjustment time, the least overshoot, and the smallest steady-state error .
[0021] This application proposes a method for adjusting the load current of an electronic load based on deep reinforcement learning. Based on the reinforcement learning framework, the adjustment process of the load current of the electronic load is simulated as a Markov decision process, taking the high-frequency timing signal of the load current as the environmental state, and the agent solves the optimal sequence of the adjustment amount of the MOS transistor gate voltage; the agent selects a multi-scale jump separable convolutional neural network model; after applying the adjustment amount predicted by the agent to the electronic load, establish an adjustment amount reward value according to the difference between the actual current and the target current, and establish a training data set through multi-step adjustment of the agent; establish a reward objective function by applying weights to the reward value sequence based on the influence degree on the load current, and optimize the agent parameters with the maximum value of the reward objective function as the goal based on the gradient descent algorithm. The deep reinforcement learning method proposed in this application does not require labeling training samples, reduces manual participation, and improves learning efficiency. Brief Description of the Drawings
[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0023] Figure 1 It is a schematic flow chart of a method for optimizing the adjustment of the load current of an electronic load based on a deep reinforcement learning method provided by an embodiment of this application;
[0024] Figure 2The electronic load current pulling circuit provided by the embodiment of the present invention;
[0025] Figure 3 The deep reinforcement learning model architecture provided by the embodiment of the present invention;
[0026] Figure 4 The multi-scale jump separable convolutional neural network model provided by the embodiment of the present invention. Detailed implementation manners
[0027] In order to enable those skilled in the art to better understand the solution of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0028] As Figure 1 shown, Figure 1 It is a schematic flowchart of a method for regulating the pulling current of an electronic load based on deep reinforcement learning provided by an embodiment of the present application, which specifically includes five contents.
[0029] S11: Simulate the regulation process of the electronic load pulling current as a Markov decision process, and use a reinforcement learning architecture to achieve current regulation. Take the high-frequency timing signal of the pulling current as the environmental state, and the agent solves the optimal sequence of MOS tube gate voltage regulation amounts.
[0030] It should be noted that as long as it can be described as a Markov decision process, a reinforcement learning model architecture can be used. Adjusting the gate voltage of the electronic load MOS tube will change the magnitude of the pulling current. For example, Figure 2 , the process of adjusting the pulling current to the target value is a current state transition process driven by the adjustment amount. The adjustment amount is calculated by the agent according to the environmental state. This process can be described as a Markov decision process. The electronic load reinforcement learning current regulation architecture proposed in the present application is as Figure 3 , and mainly includes the following steps.
[0031] Step 11: Environmental perception. Use a Hall open-close current transformer to collect Figure 2 the drain current of BG3 in
[0032] to form a time series, perform outlier removal and missing value interpolation processing on the time series signal, and form environmental observation values after filtering and denoising and sliding slicing; If the pulling current is adjusted from to , then is defined as the initial state Defined as the final state , the state is defined as the pull current after the first adjustment of the gate voltage of MOS transistor BG3. By analogy, the state set of the adjustment process is ;
[0033] Step 13: Agent. For the Markov decision process, the decision model gives the adjustment strategy for the current state according to the environmental observation. This decision model is called an agent in reinforcement learning. In this application, a deep learning model is used to give the gate adjustment voltage of MOS transistor BG3 in different states based on the high-frequency time series signal of the pull current after filtering and denoising;
[0034] Step 14: Adjustment amount. Starting from the initial state , the first adjustment amount of the gate voltage of MOS transistor BG3 is defined as , and its value is given by the agent according to the environmental observation. It can be described as "+0.1v, +0.2v, -0.2v, -0.15v,..., +0.2v". The last adjustment amount of the MOS transistor is . By analogy, the adjustment amount set of the adjustment process is ;
[0035] Based on the above discussion, in an optional embodiment of this application, the sampling frequency of the Hall current transformer is 10k, the sliding window size = 8000, and the sliding step = 500. The agent selects a deep convolutional neural network.
[0036] S12: The reinforcement learning agent selects a multi-scale jump separable convolutional neural network model.
[0037] It should be noted that an electronic DC load is an instrument. The agent model established in this application needs to run embedded in the ARM chip of the electronic load controller. The traditional deep neural network has a large amount of calculation and consumes a large amount of memory. This application proposes a deep network model as shown in Figure 4 , which combines the 1*1 convolution and depth separable convolution strategies, greatly reducing the amount of calculation. At the same time, the input information is extracted by multi-core multi-scale convolution to ensure the integrity and availability of the extracted information. The model modeling process mainly includes the following key steps.
[0038] Step 21: For the convenience of two-dimensional convolution, the long sequence signal of the collected pull current is segmented by connecting the head and tail. The first segment of the time series signal is used as the first row of the two-dimensional grayscale image, and the second segmented time series signal is used as the second row of the grayscale image, and so on, to establish the two-dimensional grayscale image data of the pull current;
[0039] Step 22: The agent proposed in this application runs embedded in the ARM chip of the electronic load. It is necessary to compress the computational amount of the convolution process, but the ability of the model to extract features cannot be reduced. For this purpose, a new model architecture of 1*1 convolution followed by multi-scale convolution is proposed: ① Use a 1*1 convolution kernel to perform convolution on the input grayscale image to form the first feature block. ② Use a 1*1 convolution kernel to perform convolution on the input grayscale image, and then perform 3*3, 5*5, and 7*7 convolutions on the feature map after 1*1 convolution to form the second feature block. ③ Do not perform convolution on the original grayscale image, but directly perform 2*2 max pooling processing to form the third feature block. Finally, splice the three feature blocks to form the first-layer convolution feature map;
[0040] Step 23: Use the first-layer convolution feature map as the input of the second convolution layer. First, use a conventional 1*1 pointwise convolution kernel to fuse the input feature map. Then, to reduce the convolution computational amount, use the depthwise separable convolution method. Each convolution kernel only convolves the feature map of its own channel without performing multi-feature map fusion. There are as many output feature maps as there are input feature maps after convolution. Then, use a conventional 1*1 pointwise convolution again to fuse and convolve the multi-feature maps to form the output feature map;
[0041] Step 24: A skip connection method is used between the convolution feature map of the first layer and the output feature map of the second layer to achieve information intercommunication, eliminate useless information, and reduce the parameters to be optimized. The output feature layer after the skip connection enters the channel attention mechanism module to optimize and allocate the weights of the feature maps of each channel, select important information, and further reduce the computational amount of the next layer.
[0042] Based on the above discussion, in an optional embodiment of this application, the first feature block of the first convolution layer selects 32 channels, and each of the 3*3, 5*5, and 7*7 convolutions of the second feature block has 64 channels.
[0043] S13: Establish an adjustment amount reward value according to the difference between the actual current and the target current, and establish a training data set through multi-step adjustment by the agent.
[0044] It should be noted that the reinforcement learning training data set has no artificial labels. The model optimizes and updates the parameters based on the reward cumulative value of the adjustment amount sequence in each state during the state transition process, and finally obtains the voltage adjustment amount sequence that ensures the highest reward. Therefore, the design of the reward function is a key step in reinforcement learning. And during the state migration process, in each state, a new state migration should start from this state until the final state, and the training data set should be resampled in each state. Based on the newly sampled data set, the model parameters are updated in this state. The reward function design and training data sampling specifically include the following steps.
[0045] Step 31: In the initial state , input the environmental observation data into the intelligent agent, and the intelligent agent gives the MOS transistor gate voltage adjustment amount based on the existing model parameters , the applied to the electronic load, and the pull-down current feedback value is , Calculate the target current The difference from the feedback current , if after adjustment is close to the target value, then give Multiply by the ratio + , otherwise multiply by the ratio - , to obtain the reward value of the first adjustment amount , enter the state ;
[0046] Step 32: Starting from the state Repeat step 1 until the final target state , this process is called a sampling of a training sample, and the state sequence set of the first sampling is obtained , the adjustment amount sequence set , the reward sequence set after each adjustment , accumulate all the rewards of the first sampling to obtain the accumulated reward ;
[0047] Step 33: Repeat steps 31 - 32 times, and training sample sequence sets can be sampled: ① state sequence sets , , ... , , ② adjustment amount sequence sets, , , ... , , ③ accumulated reward sequence sets ;
[0048] Step 34: The above steps 31 - 33 are the sampling process from to . During the state transition process based on the MOS transistor gate voltage adjustment, in each state, steps 1 - 3 need to be executed, that is, from the state , ,..., each state transition process needs to sample and obtain A sequence sample set as described in step 3 has different numbers of sample sequences in different states.
[0049] Based on the above discussion, in an alternative embodiment of the present application, the number of sampling times for each state is a hyperparameter that needs to be optimized according to real data, and the number of states , and the proportional value are also hyperparameters that need to be optimized according to the statistics of the actual adjustment process.
[0050] S14: Apply weights to the reward value sequence based on the influence degree on the pull current to establish a reward objective function, and optimize the agent parameters with the goal of maximizing the reward objective function value based on the gradient descent algorithm.
[0051] It should be noted that the initial parameters of the agent model are randomly given. Therefore, the MOS transistor voltage adjustment amount output after inputting the state observation amount cannot guarantee the maximum reward value. Only by continuously repeating sampling to obtain a large number of training samples and using the gradient descent algorithm to continuously optimize the agent model parameters based on the objective function with the maximum reward can the goal of maximizing the cumulative reward value in the state transition process be achieved. The agent parameter optimization process includes the following steps:
[0052] Step 41: Design a double-loop reward accumulation structure. The outer loop is current states, and the inner loop is regulation amount sequences. At each state, training sample sequence sets need to be sampled;
[0053] Step 42: Design a reinforcement learning reward function. When optimizing the model parameters in the t state, the reward values before the time do not need to be considered t while considering the reward values after the time. As the influence of the
[0054] (1)
[0055] t represents the state occurrence time, with a total of states, represents the current state, represents the influence weight of the operation amount on the reward value;
[0056] Step 43: Adopt the gradient descent calculation formula:
[0057] (2)
[0058] Calculate the reward gradient value in the state, based on the parameter update formula:
[0059] (3)
[0060] Update the model parameters in the state, where i represents a sample, and there are a total of sampling samples, Prevent all rewards from being positive. \(\alpha\) is the influence factor and represents the step size of each optimization.
[0061] Based on the above discussion, in an optional embodiment of the present application, the gradient descent algorithm selects the stochastic gradient descent method. Take the average value of historical rewards, and the decay factor takes 0.6. The influence factor needs to be optimized according to the specific data used.
[0062] S15: Use the trained agent as a controller, take the environmental state as input, and give a discrete sequence of the optimal adjustment amount of the MOS transistor gate voltage to achieve fast and high-precision adjustment of the pull load current.
[0063] It should be noted that the trained agent gives a sequence of MOS transistor gate voltage adjustment amounts based on the current state of the electronic load. After applying this sequence to the electronic load circuit, the target value tracking control of the current can be quickly and accurately realized. The specific process includes:
[0064] Take the initial state pull load current mean value and high-frequency waveform data collected as the observed quantities and input them into the trained agent model. The model outputs an adjustment amount of the MOS transistor gate voltage. After the adjustment, the pull load current changes and enters the state. Then, collect the observed quantities in the new state and return them to the agent. This process is repeated continuously until the error range between the pull load current and the target value is reached, and the adjustment process ends. The operation sequence formed during the adjustment process can ensure the shortest adjustment time, the highest precision, and the smallest overshoot.
[0065] Based on the above discussion, in an optional embodiment of the present application, the static error requirement for the target value is within ±5%, and the adjustment time is calculated from the initial state to the final state.
[0066] In this application, specific examples are used to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.
Claims
1. An electronic load pulling current regulation method based on deep reinforcement learning, characterized in that Including: The regulation process of the electronic load's pulling current is simulated as a Markov decision process. Using the high-frequency timing signal of the pulling current as the environmental state, the agent solves the optimal MOS transistor gate voltage regulation amount sequence. The agent of the reinforcement learning architecture selects a multi-scale jump separable convolutional neural network model. After converting the high-frequency current sampling timing signal into a grayscale image, it enters the multi-core convolutional layer. The multi-core convolution selects four different sizes of convolutional kernels: 1*1 convolution followed by 3*3 convolution, 1*1 convolution followed by 5*5 convolution, and 1*1 convolution followed by 7*7 convolution. In the second layer of the model, first use 1*1 pointwise convolution to adjust the number of features, then select depthwise separable convolution to extract information, and then use 1*1 pointwise convolution to fuse information to obtain feature information. A jump connection method is used between the input and the feature information to achieve information intercommunication. The feature information enters a channel attention module to optimize the weights of each feature information vector in the second layer. Based on the difference between the actual current and the target current, a regulation amount reward value is established, and a training data set is established through multi-step regulation by the agent. Based on the influence degree on the pulling current, weights are applied to the reward value sequence to establish a reward objective function, and based on the gradient descent algorithm, the agent parameters are optimized with the maximum reward objective function value as the goal. Use the trained agent as the controller, and use the timing signals of the starting, current, and target load currents as the initial state s1, the current state s t and the final state s n respectively. Taking the environmental state as the input, give the discrete sequence of the optimal adjustment amount of the MOS transistor gate voltage to achieve fast and high-precision adjustment of the load current.
2. The method for adjusting the pulling current of an electronic load based on deep reinforcement learning according to claim 1, wherein, The regulation process of the electronic load's pulling current is simulated as a Markov decision process. Using the high-frequency timing signal of the pulling current as the environmental state, the agent solves the optimal MOS transistor gate voltage regulation amount sequence, including: Taking the high-frequency sampling timing signal of the pulling current as the observable quantity of the environmental state, with observable quantities s1, s2, …, s n as the input, the agent gives a discrete sequence a1, a2, …, a of the MOS transistor gate voltage adjustment amounts for controlling the magnitude of the pulling current n-1 , and under the control of the discrete sequence, a state transition set s1, s2, …, s of the pulling current is formed n , thereby simulating the adjustment process of the pulling current as a Markov decision process.
3. The method for adjusting the pulling current of an electronic load based on deep reinforcement learning according to claim 1, wherein Based on the difference between the actual current and the target current, a regulation amount reward value is established, and a training data set is established through multi-step regulation by the agent, including: To achieve the state transition of the pulling current from s1 to s n , the agent needs to complete n - 1 times of MOS transistor gate voltage regulation, forming a discrete sequence of regulation amounts a1, a2, …, a n-1 . After m times of cyclic regulation by the agent, m discrete sequences are formed. After each regulation amount is applied to the circuit, the pulling current feedback value is I, and the target current I * is calculated. The difference ΔI t between the target current and the feedback current is * ΔI = I t - I. If I is close to the target value after regulation, then multiply ΔI t by the ratio +δ, otherwise multiply by the ratio -δ, to obtain the reward value r n for each regulation amount. The reward values r1, r2, …, r j of a regulation amount sequence are weighted and accumulated to obtain the reward value G. For m regulation amount sequences, m reward values are obtained, and the reward values and the regulation amount sequences form the training data set for reinforcement learning.
4. The method for regulating the pulling current of an electronic load based on deep reinforcement learning according to claim 1, characterized in that, Based on the influence degree on the pulling current, weights are applied to the reward value sequence to establish a reward objective function, and based on the gradient descent algorithm, the agent parameters are optimized with the maximum reward objective function value as the goal, including: An optimization method for the parameters of an agent model using a double-loop structure, where the outer loop has n current states and the inner loop has m adjustment amount sequences. At each state, m training sample sequence sets need to be sampled, s t When optimizing the model parameters in the state, the rewards r1, r2, …, r before time t do not need to be considered t-1 , and at the same time, the reward values r after time t are considered t+1 , r t+2 , …, r n For s t The influence of the model parameters in the state gradually weakens. According to the influence degree on the model parameters, the reward value sequence r1, r2, …, r n Add weights to obtain the reinforcement learning reward objective function λ < 1 represents the influence weight of the operation amount on the reward value, t represents the moment when the state occurs, there are n states in total, k is a constant variable, k = t means k takes the current state, t - k represents that when entering the t state, only when the current state is k, the reward value weight is the largest at 1, and the reward weights of the subsequent states gradually decrease; then use Calculate the reward gradient value in the state of s t , where i represents the sample, and there are m sampling samples in total Represents the probability that the i-th sample appears for the adjustment amount a t in the state of s t . Based on the parameter update formula Update the model parameters in the state of s t , b represents the average value of the rewards to prevent all rewards from being positive 5. The method for regulating the pulling current of an electronic load based on deep reinforcement learning according to claim 1, wherein Use the trained agent as the controller, and use the timing signals of the starting, current, and target load currents as the initial state s1, current state s t and final state s n of the environment respectively. Taking the environmental state as the input, give the discrete sequence of the optimal adjustment amount of the MOS transistor gate voltage to achieve fast and high-precision adjustment of the load current, including: Input the high-frequency sampling timing signal of the pull load current in the s1 state into the trained intelligent agent model, and the model outputs the optimal MOS transistor gate voltage adjustment amount. After applying this adjustment amount, the current enters the s2 state. After continuous iteration, the model outputs a voltage adjustment sequence that ensures the shortest adjustment time, the least overshoot, and the smallest steady-state error.
Citation Information
Patent Citations
High-voltage direct-current electronic load switch protection circuit
CN117498262A
Method for evaluating performance of MOS (Metal Oxide Semiconductor) tube for electronic load based on multi-feature mode recognition
CN118643431A