A Wireless Power Transfer Efficiency Control Method Based on Deep Reinforcement Learning
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-23
- Publication Date
- 2026-08-14
AI Technical Summary
然而,Q-learning采用查表方式存储和更新价值函数,其动作与状态空间必须离散化,这使其仅适用于低维离散控制问题
Smart Images

Figure CN122577451A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of wireless power transmission system control technology, specifically a wireless power transmission efficiency control method based on deep reinforcement learning. Background Technology
[0002] Wireless power transfer (WPT) technology enables contactless power transfer through a spatial medium and has been widely applied in electric vehicles, medical devices, and drones. To ensure efficient energy transfer, the primary circuit typically requires an inverter and resonant network to generate a stable high-frequency sinusoidal current in the transmitting coil. However, in actual operation, the relative positional offset and angular changes of the transmitting and receiving coils cause significant fluctuations in coupling strength. Simultaneously, the battery's equivalent internal resistance dynamically changes with charging parameters and operating mode. These time-varying circuit parameters directly cause deviations in transmission power and efficiency, causing the system's operating point to deviate from its maximum efficiency or maximum power point, severely degrading the overall transmission performance of the system.
[0003] Existing research on the impact of mutual inductance and load parameter variations on energy transfer performance mainly focuses on impedance matching, frequency tracking, and model-based optimization control. From a control paradigm perspective, these works can be broadly divided into two categories: one is to reconstruct matching conditions through parameter identification or prediction; the other is to directly adjust the system's operating state to track the resonant point.
[0004] Regarding impedance matching, the invention application with publication number CN121906825A predicts the mutual inductance and load resistance in an LCC-S topology in real time using a pre-trained two-stage residual neural network. It then combines this with a dynamic optimal impedance matching three-dimensional surface to determine the optimal equivalent load resistance, thereby adjusting the duty cycle of the primary and secondary Buck-Boost circuits to achieve efficient constant power output under mutual inductance fluctuations and load changes. This method utilizes the nonlinear mapping capability of deep learning to avoid complex online solutions, but it is essentially still an open-loop feedforward compensation: the matching effect heavily relies on the accuracy of the neural network's prediction of system parameters. Offline training data is difficult to cover all actual operating conditions; once complex disturbances outside the training set occur, prediction biases accumulate and lead to matching inaccuracies. Real-time performance is also challenged. The invention application with publication number CN122026633A proposes a hybrid optimization strategy based on a BP neural network and an improved white whale optimization algorithm to perform online optimization of the adjustable T-type matching network parameters to minimize the reflection coefficient. This scheme combines the inference speed of neural networks with the global search capability of metaheuristic algorithms to achieve adaptive impedance matching in megahertz WPT systems. However, the convergence stability and speed of its search process depend on the parameter settings of the optimization algorithm, and in highly dynamic or strongly nonlinear impedance change scenarios, online iterative optimization may face the risk of slow convergence or even oscillation. At the same time, the prediction of impedance matching effect by the BP neural network is also limited by the distribution range of the training data, and the problem of insufficient generalization ability still exists.
[0005] In frequency tracking, the invention application with publication number CN121966039A employs a second-order generalized integrator phase-locked loop (SOGI-PLL) to perform orthogonal decomposition and phase error extraction on the inverter output voltage, thereby adjusting the inverter operating frequency to lock the system resonant point. This method requires only a single voltage signal, has a simple structure, and utilizes bandpass filtering characteristics to suppress harmonics and noise interference, exhibiting a certain degree of robustness under conditions of gradual parameter changes within a small range. However, its dynamic performance is limited by the inherent bandwidth constraint of the phase-locked loop: under extreme conditions such as sudden changes in coupling coefficient or load, there is an irreconcilable contradiction between the dynamic response speed and steady-state accuracy of frequency tracking, which can easily lead to transient detuning.
[0006] In summary, all the above solutions rely to varying degrees on accurate mathematical models or parameter identification information: whether it's deep learning prediction models, metaheuristic optimization solutions, or phase-locked loop frequency tracking, their decision quality is directly or indirectly constrained by model accuracy and the coverage of prior data. When the system faces multiple operating conditions, strong uncertainties, and complex disturbances, these model- or data-driven open-loop, quasi-online methods struggle to guarantee continuous adaptive optimal control.
[0007] In recent years, reinforcement learning has demonstrated unique advantages in decision-making problems under uncertain environments. As an end-to-end policy learning algorithm, reinforcement learning agents autonomously make decisions and continuously optimize policies based on observed states through repeated interactions with the environment, without the need to pre-identify system parameters or build precise models. This provides a self-optimizing control path for optimizing the transmission performance of WPT systems. Patent application CN117081276A explores this area, proposing a method for improving the efficiency of two-parameter perturbation WPT systems based on the Q-learning algorithm. This method maintains stable output power and improves efficiency by dynamically adjusting the equivalent capacitance value of a controllable switched capacitor. However, Q-learning uses a lookup table approach to store and update the value function, requiring its action and state spaces to be discretized. This limits its applicability to low-dimensional discrete control problems. For complex tasks like multi-condition anti-interference in WPT systems, which have continuous state and action spaces, Q-learning inevitably encounters the curse of dimensionality, leading to difficulties in policy convergence or even policy failure.
[0008] The above analysis shows that existing methods based on neural networks, optimization algorithms, or Q-learning are either limited by their dependence on accurate models and data generalization bottlenecks, or restricted to discrete space decision-making, making it difficult to achieve stable and efficient online control in continuous, variable, and highly uncertain real-world WPT (Warning-Purpose Transmission) conditions. Therefore, how to invent a deep reinforcement learning control method and system capable of end-to-end learning directly in a high-dimensional continuous state-action space, possessing both strong robustness and self-optimization capabilities, to meet the multi-condition adaptive control requirements of wireless power transmission systems, has become a key problem urgently needing to be solved in this technical field. Summary of the Invention
[0009] To address the shortcomings of existing technologies, the technical problem this invention aims to solve is to provide a wireless power transfer efficiency control method based on deep reinforcement learning.
[0010] The present invention solves the aforementioned technical problem by adopting the following technical solution: A wireless power transfer efficiency control method based on deep reinforcement learning, characterized by the following steps: Step 1: Construct the circuit topology of the magnetically coupled wireless power transmission system, including the DC power supply, high-frequency inverter module, primary-side compensation topology and transmitting coil module, receiving coil and secondary-side compensation topology module, controllable rectifier and filter module, and load connected in sequence; The high-frequency inverter module includes four switches connected in a full-bridge configuration. The primary-side compensation topology and transmitting coil module includes a primary-side compensation inductor, a primary-side compensation capacitor, a primary-side compensation capacitor, and a primary-side coil connected in an LCC compensation network configuration. The receiving coil and secondary-side compensation topology module includes a secondary-side coil, a secondary-side compensation inductor, a secondary-side compensation capacitor, and a secondary-side compensation capacitor. The controllable rectifier and filter module includes a controllable rectifier circuit composed of switches five through eight and filter capacitors. Step 2: Based on the circuit topology of the magnetically coupled wireless power transfer system, construct the magnetically coupled wireless power transfer system; given the range of mutual inductance variation, the range of load resistance variation, and the range of phase shift angle variation of the controllable rectifier circuit, collect the system input power and output power of the magnetically coupled wireless power transfer system under different states, and calculate the system transmission efficiency; at the same time, obtain the DC voltage, input DC current, load voltage and load current, and the change in phase shift angle relative to the previous moment. These parameters form a sample, and several samples form a dataset. Step 3: Model the anti-interference control problem of the WPT system as a Markov decision process, specifically including: State space: state of time Defined as: (1) In the formula, DC voltage The input DC current, This is the load voltage. For load current, For system input power, For system output power, For system transmission efficiency, For phase shift angle, This is the phase shift angle increment; Action space: Moment of action Defined as: (2) The action variable is the phase shift angle increment of the secondary-side controllable rectifier circuit, which satisfies: (3) in, For the amplitude limiting function, and These are the minimum and maximum phase shift angles allowed by the secondary-side controllable rectifier circuit, respectively. The reward function is expressed as: (4) (5) In the formula, for Instant rewards for each moment This is the efficiency gain coefficient. To smooth out the penalty coefficient, This represents the maximum value of the phase shift angle change. , These are the upper and lower limits of the system output power, respectively; Step 4: The PPO algorithm is used to train the policy network and the value network. During the training phase, the mutual inductance and load resistance of the wireless power transmission system are randomly set. The policy network outputs the probability distribution parameters of the phase shift angle increment based on the current state and obtains the phase shift angle increment from the probability distribution. The phase shift angle increment is applied to the wireless power transmission system to obtain the next state and the immediate reward, forming the trajectory data corresponding to the current policy. Step 5: Deploy the trained policy network in the wireless power transmission system controller. The policy network outputs the phase shift angle increment according to the current state to update the target phase shift angle. Step 6: Convert the target phase shift angle into complementary drive signals for each switch in the secondary-side controllable rectifier circuit. The drive duty cycle of each switch is kept at a preset duty cycle. The two switches in the same bridge arm are complementary and a dead time is set. By changing the relative phase shift angle between the drive signals of the two bridge arms, the equivalent load presented by the secondary-side controllable rectifier circuit to the resonant network is adjusted, thereby tracking the maximum efficiency operating point that meets the output power constraint when the mutual inductance and load change.
[0011] Compared with the prior art, the beneficial effects of the present invention are: This paper proposes a wireless power transfer efficiency control method based on deep reinforcement learning. It effectively overcomes the shortcomings of existing Q-learning methods, such as their limitation by discrete states and action spaces, resulting in insufficient control accuracy and susceptibility to system chattering. It also addresses the limitations of traditional model-based optimal impedance matching methods, which heavily rely on precise mathematical models and exhibit poor adaptability under time-varying parameters and disturbances. This invention indirectly outputs continuous phase shift angles through the PPO algorithm, achieving fine-tuning in a continuous action space with smooth, chatter-free control. Employing a model-free, data-driven learning approach, it can adaptively track the maximum efficiency point online even with large-scale dynamic changes in parameters such as mutual inductance and load, without requiring prior knowledge of the system's precise model, significantly enhancing robustness. Furthermore, leveraging the reward function reconstruction capability of PPO, it can flexibly integrate multiple optimization objectives such as efficiency and power, elevating single-objective optimization to multi-objective comprehensive optimal control, balancing system safety and performance. The algorithm possesses end-to-end learning potential, simplifying the system structure and improving integration and reliability. This invention can significantly improve the transmission efficiency, response speed, and operational stability of wireless power transfer systems under complex dynamic conditions. Attached Figure Description
[0012] Figure 1 The topology of a magnetically coupled wireless power transfer system; Figure 2 This is the overall control block diagram; Figure 3 The waveform diagram shows the drive signal of the controllable rectifier circuit. Figure 4 This is a graph showing the relationship between the phase shift angle and the equivalent impedance. Detailed Implementation
[0013] Specific embodiments are given below with reference to the accompanying drawings. These specific embodiments are only used to describe the technical solution of the present invention in detail and are not intended to limit the scope of protection of this application.
[0014] This invention provides a wireless power transfer efficiency control method based on deep reinforcement learning, comprising the following steps: Step 1: Construct the circuit topology of a magnetically coupled wireless power transfer system with an LCC-LCC resonant network, including a DC power supply, a high-frequency inverter module, a primary-side compensation topology and a transmitting coil module, a receiving coil and a secondary-side compensation topology module, a controllable rectifier and filter module, and a load connected in sequence. like Figure 1 As shown, the high-frequency inverter module includes a first switching transistor connected in a full-bridge configuration. Switch No. 4 The primary-side compensated topology and transmitting coil module includes primary-side compensated inductors connected in the form of an LCC compensated network. Primary-side compensation capacitor No. 1 Second primary-side compensation capacitor and primary coil The receiving coil and secondary compensation topology module include the secondary coil. Secondary compensation inductor Secondary-side compensation capacitor No. 1 Secondary-side compensation capacitor No. 2 The controllable rectifier and filter module includes a No. 5 switching transistor. Switch No. 8 The controllable rectifier circuit and filter capacitor are composed of By controlling switch tube number five Switch No. 8 The phase shift angle between the drive signals can continuously adjust the equivalent impedance presented to the primary side by the controllable rectifier circuit.
[0015] Step 2: Based on the circuit topology of the magnetically coupled wireless power transfer system, construct the magnetically coupled wireless power transfer system; given the range of mutual inductance variation, the range of load resistance variation, and the range of phase shift angle variation of the controllable rectifier circuit, collect the system input power of the magnetically coupled wireless power transfer system under different states (i.e., different mutual inductance, load, and phase shift angle). and output power Calculation system transmission efficiency Simultaneously, acquire DC voltage. Input DC current Load voltage and load current The change in phase angle relative to the previous moment These parameters form a sample, and several samples form a dataset.
[0016] Step 3: Model the anti-interference control problem of the WPT system as a Markov decision process, specifically including: State space: state of time Defined as: (1) Action space: Moment of action Defined as: (2) The action variable is the phase shift angle increment of the secondary-side controllable rectifier circuit, which satisfies: (3) in, For the amplitude limiting function, and These are the minimum and maximum phase shift angles allowed by the secondary-side controllable rectifier circuit, respectively. The reward function aims to guide the agent to maximize system transmission efficiency and suppress drastic phase angle fluctuations while ensuring that the output power does not fall below a set threshold. The reward function includes an efficiency term, a power penalty term, and a smoothing penalty term. The efficiency term is defined as a function of the system transmission efficiency. The power penalty term is defined as a quadratic penalty with respect to the system output power; a larger penalty is applied when the system output power exceeds the desired power range, with a higher penalty weight assigned to the side with insufficient power. The smoothing penalty term is defined as the square of the normalized phase angle change, used to suppress frequent and drastic phase adjustments. The reward function is expressed as: (4) (5) In the formula, for Instant rewards for each moment This is the efficiency gain coefficient. To smooth out the penalty coefficient, This represents the maximum value of the phase shift angle change. , These are the upper and lower limits of the system output power, respectively; A neural network is used to establish a nonlinear mapping relationship between the state space and the action space. The policy network outputs parameters (including mean and standard deviation) of the action distribution based on the state, realizing high-precision, smooth, and adaptive adjustment of the phase shift angle of the secondary rectifier. The value network outputs an estimate of the state value based on the state.
[0017] Step 4: The Proximal Policy Optimization (PPO) algorithm is used to train the policy network and the value network. During the training phase, the mutual inductance and load resistance of the wireless power transmission system are randomly set. The policy network outputs the probability distribution parameters of the phase shift angle increment based on the current state and obtains the phase shift angle increment from the probability distribution. The phase shift angle increment is applied to the wireless power transmission system to obtain the next state and the immediate reward, forming the trajectory data corresponding to the current policy. First, initialize the policy network parameters and value network parameters; set hyperparameters, including learning rate, pruning factor, discount factor, generalized advantage estimation (GAE) parameters, number of training epochs, and batch size.
[0018] Then, under the current policy, the policy network determines the state based on the current time step. Output Action The current is applied to the controllable rectifier circuit, and the state at the next moment is observed after the system has started operating. And calculate the immediate reward based on the reward function. ; transfer samples Store the data in the experience buffer pool; repeat this process until a sufficient amount of time's worth of trajectory data is collected.
[0019] The generalized dominance estimation algorithm is used to calculate the estimated value of the dominance function at each time step in the trajectory. The generalized dominance estimation is achieved by exponentially weighting the temporal difference error, thus achieving the optimal trade-off between estimation bias and variance. Based on the advantage function estimate, construct the pruning agent objective function: (6) In the formula, For the expectation, for The probability ratio of the strategies at time t. for The estimated value of the dominance function at time 1. This is a pruning function used to constrain the policy probability ratio within an interval. Within this framework, it is important to avoid excessively large step sizes in a single policy update, which could lead to training instability. The clipping factor; The strategy probability ratio represents a measure of similarity between the current strategy and the old strategy, and is defined as: (7) In the formula, As the current strategy, This is the old strategy; Constructing the value loss function: (8) In the formula, For value networks; for The cumulative discounted value over time equals the instant reward. The product of the discount factor; Construct the overall optimization objective function: (9) In the formula, This is the value loss weighting coefficient. The weighting coefficients for the entropy regularization term; This is the entropy regularization term for the policy, used to increase the exploratory nature of the policy and prevent the training process from converging to a local optimum too early. An adaptive moment estimation (Adam) optimizer is used to train the policy network and value network based on the overall optimization objective function until convergence, resulting in the trained policy network.
[0020] Step 5: Deploy the trained policy network in the wireless power transmission system controller. The policy network outputs the phase shift angle increment based on the current state to update the target phase shift angle; by adjusting the phase shift angle, the equivalent impedance can be continuously adjusted within a certain range, such as... Figure 4 As shown, as the phase shift angle increases, the equivalent impedance can be made to approach the optimal load downwards, thereby achieving maximum efficiency tracking.
[0021] Step 6: Convert the target phase shift angle into complementary drive signals for each switch in the secondary-side controllable rectifier circuit. The drive duty cycle of each switch is maintained at a preset duty cycle. Two switches in the same bridge arm are complementaryly turned on and a dead time is set. By changing the relative phase shift angle between the two bridge arm drive signals, the equivalent load presented by the secondary-side controllable rectifier circuit to the resonant network is adjusted. Figure 4 As shown, as the phase shift angle increases, the equivalent impedance can be made to approach the optimal load downwards, thereby achieving maximum efficiency tracking.
[0022] Any aspects not covered in this invention are applicable to existing technologies.
Claims
1. A wireless power transfer efficiency control method based on deep reinforcement learning, characterized in that, Includes the following steps: Step 1: Construct the circuit topology of the magnetically coupled wireless power transmission system, including the DC power supply, high-frequency inverter module, primary-side compensation topology and transmitting coil module, receiving coil and secondary-side compensation topology module, controllable rectifier and filter module, and load connected in sequence; The high-frequency inverter module includes four switches connected in a full-bridge configuration. The primary-side compensation topology and transmitting coil module includes a primary-side compensation inductor, a primary-side compensation capacitor, a primary-side compensation capacitor, and a primary-side coil connected in an LCC compensation network configuration. The receiving coil and secondary-side compensation topology module includes a secondary-side coil, a secondary-side compensation inductor, a secondary-side compensation capacitor, and a secondary-side compensation capacitor. The controllable rectifier and filter module includes a controllable rectifier circuit composed of switches five through eight and filter capacitors. Step 2: Based on the circuit topology of the magnetically coupled wireless power transfer system, construct the magnetically coupled wireless power transfer system; given the range of mutual inductance variation, the range of load resistance variation, and the range of phase shift angle variation of the controllable rectifier circuit, collect the system input power and output power of the magnetically coupled wireless power transfer system under different states, and calculate the system transmission efficiency; at the same time, obtain the DC voltage, input DC current, load voltage and load current, and the change in phase shift angle relative to the previous moment. These parameters form a sample, and several samples form a dataset. Step 3: Model the anti-interference control problem of the WPT system as a Markov decision process, specifically including: State space: state of time Defined as: (1) In the formula, DC voltage The input DC current, This is the load voltage. For load current, For system input power, For system output power, For system transmission efficiency, For phase shift angle, This is the phase shift angle increment; Action space: Moment of action Defined as: (2) The action variable is the phase shift angle increment of the secondary-side controllable rectifier circuit. And satisfy: (3) in, For the amplitude limiting function, and These are the minimum and maximum phase shift angles allowed by the secondary-side controllable rectifier circuit, respectively. The reward function is expressed as: (4) (5) In the formula, for Instant rewards for each moment This is the efficiency gain coefficient. To smooth out the penalty coefficient, This represents the maximum value of the phase shift angle change. , These are the upper and lower limits of the system output power, respectively; Step 4: The PPO algorithm is used to train the policy network and the value network. During the training phase, the mutual inductance and load resistance of the wireless power transmission system are randomly set. The policy network outputs the probability distribution parameters of the phase shift angle increment according to the current state and obtains the phase shift angle increment from the probability distribution. The phase shift angle increment is applied to the wireless power transmission system to obtain the next state and the immediate reward, forming the trajectory data corresponding to the current policy. Step 5: Deploy the trained policy network in the wireless power transmission system controller. The policy network outputs the phase shift angle increment according to the current state to update the target phase shift angle. Step 6: Convert the target phase shift angle into complementary drive signals for each switch in the secondary-side controllable rectifier circuit. The drive duty cycle of each switch is kept at a preset duty cycle. The two switches in the same bridge arm are complementary and a dead time is set. By changing the relative phase shift angle between the drive signals of the two bridge arms, the equivalent load presented by the secondary-side controllable rectifier circuit to the resonant network is adjusted, thereby tracking the maximum efficiency operating point that meets the output power constraint when the mutual inductance and load change.
Citation Information
Patent Citations
Method for improving efficacy of two-parameter disturbance wireless power transmission system based on Q-learning algorithm
CN117081276A
Wireless power transmission control method and device based on deep learning
CN121906825A
Wireless power transmission system frequency tracking method based on second-order generalized integrator phase-locked loop
CN121966039A
Self-adaptive impedance matching system based on hybrid optimization strategy and working method
CN122026633A