Wireless charging multi-coil rapid matching and positioning method based on Q-learning algorithm
By adopting a multi-coil fast matching and positioning method based on the Q-learning algorithm in wireless charging technology, the problem of complex calculation and insufficient response capabilities of position detection algorithms in the prior art is solved, efficient and low-cost energy transmission is achieved, and good performance is maintained under extreme conditions.
Patent Information
- Application Number
- CN202510510107.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-04-23
AI Technical Summary
In the existing wireless charging technology, the position detection algorithm relies on front-end auxiliary sensor equipment, has complex calculations and insufficient responsiveness, especially in extreme weather conditions, and the limitations of the data set limit the generalization ability of the algorithm.
The wireless charging multi-coil fast matching and positioning method based on the Q-learning algorithm is adopted. By establishing a multi-coil array transmission platform and a flexible power supply network, combining dynamic exploration rate adjustment, confidence weighted action selection and gradient perception enhancement mechanism, we adapt to the dynamic changing environment to achieve the minimum system cost and the shortest path to open the optimal energy transmission unit.
It greatly reduces the number of turns on the energy transmission unit, improves matching efficiency, reduces system costs, and maintains good performance under extreme weather conditions, enhancing the generalization ability and application depth of the algorithm.
Smart Images

Figure CN120049642A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to a method for fast matching and positioning of multiple coils in wired charging, and more specifically, to a method for fast matching and positioning of multiple coils in wireless charging based on Q-learning an algorithm. Background Art
[0002] Wireless charging technology realizes wireless power transmission, completely getting rid of the bondage of cables. This non-contact charging method avoids safety hazards such as electric shock, leakage, and electric sparks, ensuring the safety and reliability of the operation process. At the same time, it reduces the downtime caused by battery replacement and does not require manual intervention, realizing unmanned autonomous charging.
[0003] The reconfigurable array magnetic structure can effectively alleviate the problem of misalignment when the power receiving device lands, and improve the system's tolerance to misalignment. However, the prerequisite is that it can intelligently select and activate the transmitting coil that highly matches the position of the power receiving device according to the random landing position and direction of the power receiving device (receiving coil), so as to ensure the efficient operation and performance optimization of the entire charging process.
[0004] Although significant progress has been made in the current research on position detection algorithms, there are still multiple challenges. The primary problem is that algorithm design mostly relies on the information collected by front-end auxiliary sensor devices and does not fully integrate the intrinsic information in the specific background of wireless power transmission for power receiving devices such as drones. For example, the introduction of technologies such as infrared, laser, and image recognition has improved the positioning accuracy, but also increased the drone load, system cost, and complexity. Especially under extreme weather conditions, the effectiveness of image recognition algorithms drops significantly. Due to insufficient image clarity and contrast or label failure, it is difficult to extract features and the position detection ability is limited.
[0005] Regarding the computational complexity and real-time requirements of algorithms, the current position detection methods applied in the context of wireless power transmission mainly theoretically deduce the position by manually calculating voltage, current, and mutual inductance, mainly realizing detection from the theoretical calculation level. The calculation is cumbersome and complex, and the accompanying computational burden limits its fast response ability in practical applications. At the same time, the limitations of the algorithm dataset, especially the lack of methods for collecting and establishing dedicated datasets for the field of wireless power transmission, limit the generalization ability and application depth of the algorithm. Summary of the Invention
[0006] To solve the technical problems that existing algorithms mostly rely on front-end information collection, are computationally cumbersome, and have limited response capabilities, the technical solution adopted in the present application is: to provide a method for fast matching and positioning of multiple coils in wireless charging based on Q-learning an algorithm, including the following steps: Step S1: Establish a multi-coil array transmitting platform and a power supply network that realizes flexible switching of the corresponding energy transmission units of the platform; Step S2: Change the position of the receiving coil and the activation status of different energy transmission units, obtain the induction amount, and establish the charging environment information; Step S3: Establish an exploration environment under the algorithm model, model a discretized two-dimensional grid map, where each grid corresponds to an energy transmission unit in the array transmitting platform, define the state S of the grid as the corresponding induction amount, and initialize the Q values of all grids; Step S4: When the induction amount of the receiving coil changes and enters the coil matching process, the agent explores in an unknown environment, executes actions according to the action selection strategy, and iteratively updates the Q value table; the action selection strategy includes a dynamic exploration rate adjustment mechanism, a confidence-weighted action selection mechanism, and a gradient perception enhancement mechanism; Step S5: Calculate the cumulative reward according to the actions executed by the agent; Step S6: Evaluate the action strategy output by the agent according to the evaluation network, and determine whether the model training strategy meets the initialization requirements. If yes, execute Step S7; if not, execute Step S4; Step S7: Extract the Q value table and output the optimal energy transmission unit and the shortest activation path.
[0007] Preferably, the dynamic exploration rate adjustment mechanism dynamically adjusts the exploration rate through the state access density and the induction amount gradient sensitivity.
[0008] Preferably, the implementation process of the dynamic exploration rate adjustment mechanism is as follows: (1) Establish an exploration rate decay function : ; Among them, N (s,a) is the cumulative execution times of action a in state s, reflecting the familiarity of the agent's actions; C (s) is the total access times of state s, indicating the degree of exploration of the current agent in the state; β is the decay rate parameter, used to control the current exploration rate decline speed; is the curvature parameter used to adjust the nonlinear degree of the decay curve; ε 0 is the initial exploration rate, that is, the maximum exploration probability.
[0009] (2) Establish a state access density threshold to force exploration of low-access areas: ; Through the above constraint conditions, it is restricted to prompt the agent to forcibly maintain an exploration rate of more than 30% when a new state appears, that is, the access times are less than 5; (3) Introduce an induction gradient sensitivity triggering mechanism to temporarily increase the exploration rate of the agent when the induction amount is close to the maximum value, that is, when the change rate of the induction amount is less than 10%. Preferably, a confidence-weighted action selection mechanism is used to quantify the uncertainty of action value through a confidence index and guide the exploration of actions with low visit times.
[0010] Preferably, the specific implementation method of the confidence-weighted action selection mechanism is as follows: (1) Calculate the action confidence index Q c (s,a) , and the formula is as follows: ; Among them, Q(s,a) represents the traditional Q-learning action value function; κ is the confidence coefficient, which can control the weight of exploration and exploitation. The larger the value, the more inclined to explore; C(s) is the total number of visits to state s; N (s,a) is the cumulative execution times of action a in state s; j is a constant to prevent division by zero; the uncertainty term √(ln(1 + C(s)) / N(s,a)) characterizes the degree of insufficient exploration of the action. (2) Actions with higher confidence are given higher selection probabilities. The action probability distribution of the agent in state s satisfies: ; Among them, P(a ∣ s) is the probability that the agent selects action s when it is in state a ; Q c (s,a) is the corrected Q value after introducing the action confidence index; a ′ is all executable actions of the agent in state s; Q c (s,a ′ ) are the corrected Q values when the agent executes different actions in state s; is the temperature parameter dynamic adjustment factor related to state s ; Introduce the temperature parameter dynamic adjustment factor ,Through the degree of certainty of the temperature parameter adaptive ,adjustment strategy, the confidence index of the modified Q value is ,further converted into a probability distribution, and the probability ,distribution is changed from flat to sharp through the temperature adjustment factor as the ,agent’s exploration process proceeds, and eventually the agent tends to a ,purely greedy state, and the current optimal state is ,selected intelligently.
[0011] ; in , Status s Related temperature parameter dynamic adjustment factors; C (s) is the total number of visits to state s, which indicates the degree of exploration of the state by the current agent; η is a hyperparameter.
[0012] Preferably, the gradient perception enhancement mechanism uses the induction quantity in the wireless charging environment to encode the field strength distribution information into the action selection, calculates the induction quantity gradient vector field and the direction preference weight, superimposes the direction preference weight on the basis of the confidence-weighted Q value, and gives priority to actions that are both high-value and consistent with the gradient direction.
[0013] Preferably, iteratively updating the Q value table refers to iteratively updating the Q value table using the Bellman equation: ; in, Q t , Q t+1 They are the Q values of the agent at time t and the next time t+1 respectively; Q t+1 (s t ,a t ) For the agent at time t+1, In state s t , Execute actions a t When , the updated Q value indicates the expected return of the state-action pair under the current strategy; s t , s t+1 are the states of the agent at time t and the next time t+1 respectively; A is the set of all actions, that is, A includes a, and a has four actions: up, down, left, and right; a t is the action performed by the agent at time t; R t+1 (s t , at ) is the cumulative reward obtained by the agent at time t + 1 after executing the action s t and transferring to the next state a t from state s t at time t; α is the learning rate; γ is the discount factor.
[0014] Preferably, the cumulative reward R t has the following calculation formula: ; where R 1 is the immediate reward, R 2 is the step penalty, R 3 is the termination reward, R 4 is the boundary penalty, and step is the number of times the agent executes the action.
[0015] Preferably, the evaluation network refers to calculating the ratio m : m = k / d ; where k is the number of actions executed by the agent, d is the Manhattan distance between the initially randomly activated energy transfer unit and the optimal energy transfer unit of the agent; When m approaches 1, the initialization requirement is met.
[0016] The beneficial effects of the present invention are as follows. The present invention proposes a wireless charging multi - coil fast matching and positioning method based on Q-learning algorithm, which adaptively learns the dynamically changing environment, activates the optimal energy transfer unit with the minimum system cost, i.e., the shortest path, and greatly reduces the activation times compared with the traversal activation search. Moreover, as the scale of the array layout increases, the matching efficiency will be further improved.
[0017] At the hardware level, the multi - coil array transmitting platform with a reconfigurable array configuration cooperates with a flexible power supply network, can freely activate the required coils according to the algorithm, provides a hardware basis for realizing efficient charging, and its coil configuration and power supply network design are flexible, suitable for any reconfigurable form of array multi - coil matching strategy.
[0018] From an algorithmic perspective, the action selection strategy is optimized through multiple innovative mechanisms, eliminating the need for complex and cumbersome calculations and theoretical derivations. The dynamic exploration rate adjustment mechanism adjusts the exploration rate based on the state access density and the sensitivity of the induction gradient, balancing exploration and exploitation, enabling the model to converge rapidly and effectively avoiding being trapped in local optimal solutions. The confidence-weighted action selection mechanism quantifies the uncertainty of action values, encourages the exploration of actions with low access frequencies, optimizes the cold start process, and improves the intelligence of action selection. The gradient-aware enhancement mechanism incorporates physical characteristics into the model, guiding the agent to explore along the direction of increasing induction, accelerating local convergence. Meanwhile, the reward function design provides reasonable feedback to the agent, and the multi-objective optimization framework comprehensively considers aspects such as immediate rewards, step penalties, termination rewards, and boundary penalties to ensure the rapid convergence of the model and the generation of reasonable strategies. In addition, this application utilizes an evaluation network to evaluate the agent's action strategy by ratio, effectively monitoring the training process and the algorithm convergence state, ensuring the effectiveness of model training. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] To more clearly illustrate the technical solutions in the embodiments of this application, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0020] Figure 1 A flow diagram of a method for fast matching and positioning of multiple coils in wireless charging based on Q-learning an algorithm provided in an embodiment of this application; Figure 2 A schematic diagram of the pickup voltage distribution provided in an embodiment of this application; Figure 3 A schematic diagram of the strategy path of a method for fast matching and positioning of multiple coils in wireless charging based on Q-learning an algorithm provided in an embodiment of this application; Figure 4 A schematic diagram of the traversal strategy path provided in an embodiment of this application; Figure 5 A Q-learning value change diagram of a method for fast matching and positioning of multiple coils in wireless charging based on m an algorithm provided in an embodiment of this application; Figure 6 An array transmitting platform composed of 19 coils provided in an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0022] Please refer to Figure 1 , which is a schematic flowchart of a method for fast matching and positioning of multiple coils for wireless charging based on Q-learning algorithm provided by an embodiment of the present application. For the convenience of description, only the parts related to this embodiment are shown and are described in detail as follows: In one embodiment, a method for fast matching and positioning of multiple coils for wireless charging based on Q-learning algorithm includes: Step S1: Establish a multi-coil array transmitting platform and a power supply network for realizing flexible switching of corresponding coils; wherein, the power supply network in the array transmitting platform is connected to each coil, and can realize the free activation of the array multi-coils, and can ensure that the coils to be activated can be flexibly activated according to the algorithm output strategy.
[0023] Specifically, the reconfigurable array-type multi-coil array transmitting platform is composed of multiple transmitting coil arrays arranged in an array, and the coil configuration can be in various forms such as circular, square, hexagonal, etc.
[0024] The power supply network can be freely designed according to the wireless charging energy transfer requirements and output conditions. At the same time, the design is not limited to using single-phase, two-phase, three-phase, four-phase and other excitation modes. Different excitation modes of the array multi-coils result in different composition methods of each coil energy transfer unit. For example, every three adjacent transmitting coils can be excited to form a freely combinable energy transfer unit. This means that in the subsequent matching process between the transmitting coil and the receiving coil of the power receiving device, according to the algorithm output, in each search process of the system, the excitation mode of the array multi-coils is three adjacent transmitting coils, and every three adjacent transmitting coils form a new energy transfer unit to transfer energy to the power receiving device. Therefore, the power supply network only needs to meet the requirement that the transmitting coils can self-organize to turn on and excite the corresponding array coils.
[0025] Step S2: Change the position of the receiving coil and the activation of different energy transfer units, obtain the coil inductance, extract data information and establish charging environment information; Change the position of the receiving coil and the activation status of different energy transmission units to extract the induced quantity in any case. The set of these induced quantities is the charging environment information. For different position situations of the power receiving device (including the receiving coil) and different activation situations of different energy transmission units, the induced quantity is different. This induced quantity reflects the matching status between the currently activated energy transmission unit and the receiving coil of the power receiving device. According to the comparison between the set threshold value of the induced quantity for the optimal matching status and the current induced quantity, if the induced quantity reaches the set threshold value, it is considered that the current system has reached the optimal matching status, and the matching and positioning process is then achieved. At the same time, the induced quantity is related to the selection of subsequent state quantities because it serves as the state quantity value of the algorithm agent. Specifically, the selection of the induced quantity can include any data with intuitive measurability such as the pickup voltage of the receiving coil, the induced current of the receiving coil, or the excitation current of the transmitting coil. Among them, the induced quantity can be reflected on the receiving coil side or the transmitting coil side, which is specifically determined according to the selection of the induced quantity.
[0026] In particular, the selection of the induced quantity does not use the basic mutual inductance or the secondary derived theoretical index as the measurement index for the induction situation, avoiding cumbersome theoretical derivations and the establishment of precise theoretical models to improve the overall detection efficiency.
[0027] Step S3: According to the charging environment information extracted in Step S2, establish Q-learning the exploration environment under the algorithm model.
[0028] Specifically, model the charging environment information as a discretized two-dimensional grid map. Each grid represents a potential activation state of the energy transmission unit, and the coordinates (x, y) of the grid correspond to the position of the energy transmission unit in the array transmitting platform.
[0029] Regarding the grid network state representation, the state S of each grid is defined as the voltage value U t (x,y) picked up by the receiving coil when the energy transmission unit corresponding to the grid is activated. This voltage value is obtained through the sensor measurement in Step S2. At the same time, for the environmental initialization setting, initialize the Q value of all grids in the grid map to zero or a small random value. This means that the agent has no prior knowledge of the value of all actions in the initial state.
[0030] Furthermore, the size of the grid map is determined according to the physical size and the number of coils of the transmitting array. Specifically, for the gridification method, a uniform gridification network is adopted, dividing the area covered by the transmitting array into grids of equal size and performing normalization processing to eliminate the differences between different array coils. The selection of the grid size needs to balance the positioning accuracy and the computational complexity. For example, if the transmitting array consists of 8x8 energy transmission units, the grid map can be set to 8x8.
[0031] Step S4: The agent explores in an unknown environment, executes actions, and updates its state. Specifically, the implementation process is as follows: Step S4-1: Definition of state and action: Define the state s to represent the magnitude of the induction of the receiving coil when the agent activates the current energy transfer unit, reflecting the coupling state of the receiving coil.
[0032] In each state s, the agent can execute the action a to iterate the state. The action a is set to continue to activate the next energy transfer unit of the current energy transfer unit to the left, right, up, or down.
[0033] To elaborate on this, define the state s to represent the magnitude of the induction of the receiving coil when the agent activates the current energy transfer unit. The current state s can reflect the current coil coupling state, and at the same time, the current system coil matching situation can be further determined according to the state quantity, which is reflected as the agent moving at different grid network points. That is, define the state s to represent the coordinate position of each grid in the exploration map.
[0034] Based on the multi-voltage extreme value characteristics of the wireless charging environment with near-field coupling characteristics, for each action, set the reconstruction directions in the horizontal and vertical four directions of the current energy transfer unit to simplify the control drive form, and at the same time perform refined search to find the best matching coil. Specifically, when the agent is in the exploration environment under the algorithm model, after activating a certain energy transfer unit of the array emission platform in the current state and detecting a certain induction amount, the action a for updating the state is set to continue to activate the next energy transfer unit of the current energy transfer unit to the left, right, up, or down, that is, the agent moves up, down, left, and right at the current grid network point. The agent's action setting can ensure any reconstruction method of the densely paved array coils of the array emission platform, avoiding falling into a local optimal solution or missing the optimal solution due to a large step size. Therefore, the agent's state transition equation is: ; where, S t 、S t+1 are the states iterated by the agent at time t and the next time t + 1 respectively, and a represents the action that the agent can execute; Step S4-2: Action selection strategy: In the traditional ε - greedy strategy, the agent will select the action with the largest current Q value with a probability of 1 - ε , that is, utilize it, and with a probability of εRandomly select an action with a certain probability, which is exploration. The purpose of this strategy is to maximize the benefits by using known information while also maintaining a certain degree of exploration in the model. However, this strategy has the defect of a fixed exploration rate, resulting in slow convergence of the agent, increasing the training time and cost, and being unable to make decisions quickly.
[0035] Therefore, to address this problem, the present invention proposes an optimized and improved action selection strategy. The specific key mechanisms include dynamic exploration rate adjustment, confidence-weighted action selection, and gradient-aware enhanced fusion and collaboration mechanisms. The details are as follows: (1) Dynamic exploration rate adjustment mechanism. Specifically, the exploration rate is dynamically adjusted through the state visit density and the gradient sensitivity of the induction quantity to balance exploration and exploitation.
[0036] First, establish an exploration rate decay function : ; Among them, N (s,a) is the cumulative execution times of action a in state s, reflecting the familiarity of the agent's actions; C (s) is the total number of visits to state s, indicating the degree of exploration of the current state by the agent; β is the decay rate parameter, used to control the current exploration rate decline speed, ∂ is the curvature parameter used to adjust the non-linearity degree of the decay curve; ε 0 is the initial exploration rate, that is, the maximum exploration probability. The exploration rate decays as the number of state visits C ( s ) increases, and the decay rate is jointly controlled by the decay rate β and the curvature . β , , ε 0 The ranges are all 0 to 1 and can be dynamically adjusted according to the actual situation. In the specific embodiment, the optimized values are β = 0.2, ∂ = 0.5, ε 0 = 0.8. Dynamically balance exploration and exploitation through the state visit frequency, fully explore in the early stage, and gradually converge to the optimal strategy in the later stage. β Control the decay speed, the larger the value, the faster the decay. Adjust the bending degree of the decay curve, the larger the value, the slower the initial decay.
[0037] Furthermore, establish a state visit density threshold to force exploration of low-visit areas: ; Through the above constraint conditions, when a new state appears, that is, when the access count is less than 5, the agent is forced to maintain an exploration rate of more than 30% to prevent the model from converging prematurely and falling into a local optimum.
[0038] Furthermore, an induction quantity gradient sensitivity trigger mechanism is introduced to temporarily increase the exploration rate when the induction quantity is close to the maximum value, avoiding falling into a sub-optimal solution. Specifically, when approaching the maximum induction quantity (such as the maximum voltage region), that is, when the gradient is less than 10%, the exploration rate of the agent is forced to increase by 15% with probability. The formula is as follows: ; where, U(s) is the induction quantity corresponding to state s, U max is the maximum induction quantity, ΔU is the induction quantity in state s U(s) and the maximum induction quantity U max is the absolute value of the difference between them.
[0039] Through the setting of the above dynamic exploration rate adjustment mechanism, high exploration is maintained in the low-access area, and local search is enhanced in the high-gradient pressure difference area to achieve rapid convergence of the model, realizing a dynamic balance between the two. At the same time, the strategy is adaptively adjusted and corrected through gradient detection to match the non-uniform discrete and multi-voltage extreme value characteristics of the wireless charging environment.
[0040] (2) Confidence-weighted action selection mechanism, which quantifies the uncertainty of action value through a confidence index to guide the exploration of actions with low access counts.
[0041] Specifically combined with Q value guidance and uncertainty evaluation, the traditional greedy strategy is replaced by an improved confidence index. On the basis of the original Q value, a confidence term negatively correlated with the action access count N ([[]] s , a ) is added. When N ([[]] s , a ) → 0, the confidence term increases significantly, encouraging the attempt of actions that have not been fully explored.
[0042] Specifically, an action confidence index Q c (s,a) is proposed, and the calculation formula is as follows: ; where, Q(s,a) represents the traditional Q-learning action value function; κ is the confidence coefficient, which can control the weight of exploration and exploitation. The larger the value, the more inclined to exploration,κ = 1.2 is an empirically optimized value; C(s) is the total number of visits to state s; N (s,a) is the cumulative execution count of action a in state s; j = 10 −5 is a constant to prevent zero and ensure numerical stability; the uncertainty term √(ln(1 + C(s)) / N(s,a)) characterizes the degree of insufficient exploration of the action.
[0043] Furthermore, actions with higher confidence are given higher selection probabilities, and the action probability distribution of the agent in state s satisfies: ; where P(a ∣ s) is the probability that the agent selects action s when it is in state a ; Q c (s,a) is the corrected Q value after introducing the action confidence index; a ′ are all the executable actions of the agent in state s; Q c (s,a ′ ) are the corrected Q values when the agent executes different actions in state s; is the dynamic adjustment factor of the temperature parameter related to state s ;
[0044] Introducing the dynamic adjustment factor of the temperature parameter , it is realized that the more the agent visits the state, the sharper the probability distribution, that is, it biases towards high Q-value actions, while when a new state is visited, the distribution is flatter, that is, uniform exploration is encouraged. The specific calculation formula is: when C ( s ) → 0, ≈ 1.5, the probability distribution is more uniform, encouraging exploration; specifically, when C ( s ) → ∞, → 1, the probability distribution is more concentrated, biasing towards high Q-value actions, and the certainty degree of the temperature parameter adaptive adjustment strategy is adjusted.
[0045] ; where , is the dynamic adjustment factor of the temperature parameter related to state s ; C (s) is the total number of visits to state s, indicating the degree of exploration of the current agent for the state;η is a hyperparameter, with a value range of 0 to 1, which can be dynamically adjusted according to the actual situation. Here, an empirical value of 0.5 is taken.
[0046] Through the above confidence-weighted action selection mechanism, the selection probability of actions with fewer execution times is further automatically increased, and the action selection in the new state is smoother, avoiding the inefficiency of random exploration and realizing cold start optimization.
[0047] (3) Gradient-aware enhancement mechanism. By using the induction quantity in the wireless charging environment (such as the physical characteristics of voltage gradient), the field strength distribution information is encoded into the action selection to achieve gradient-aware exploration enhancement, strengthen the action movement direction along the direction of increasing induction quantity, realize gradient direction guidance, and at the same time integrate the actual physical laws into the model optimization strategy.
[0048] Specifically, taking the voltage gradient field calculation as an example, the voltage gradient vector field is calculated by the central difference method ∇U(s) , which reflects the direction with the fastest local voltage change. Its calculation formula satisfies: ; where U(x,y) is the induction quantity at the coil energy transfer unit with coordinates (x,y) ; Δx = Δy = 1, indicating the adjacent coil spacing after normalization processing.
[0049] Furthermore, on this basis, the direction preference weight w dir (a) is set as: ; where v a is the unit vector of the movement direction corresponding to the action a (such as [-1, 0] for left); λ is the direction reinforcement coefficient, used to control the intensity of the gradient influence and adjust the reward intensity of gradient alignment. The larger the value, the more obvious the direction preference. The empirical value is λ = 0.8; ∇U(s) · v a is the dot product of the voltage gradient vector field ∇U(s) and the unit vector of the action direction v a , used to measure the consistency between the action direction and the gradient direction. If the direction of the action a is consistent with the voltage gradient direction (the dot product is large), then the weight w dir (a) increases.
[0050] Furthermore, the action along the gradient direction is weighted 1.8 times. That is, when λ = 0.8, only 20% of the value of the reverse gradient action is retained. Q c This realizes the automatic tracking of the voltage rising path in the non-uniform field, and at the same time quickly crosses the low-voltage array platform area through the gradient information. Based on the confidence-weighted Q value, the direction preference is superimposed, and the action that is both of high value and consistent with the gradient direction is preferentially selected. The gradient direction indicates the path with the fastest voltage rise, and the weight adjustment enables the agent to explore along the "potential rising direction" and avoid random walks. The specific action selection is optimized as follows: ; where a selected is the action finally selected and executed by the agent; argmax a is the maximum value among all actions a ; Q c (s,a) is the action confidence index.
[0051] In summary, in the action selection strategy, the above three mechanisms work together. Specifically, the dynamic adjustment of the exploration rate mechanism provides the global search ability, with a high exploration rate (80%) in the initial stage and gradually decaying in the later stage. The induction gradient sensitivity trigger mechanism temporarily increases the exploration rate when the induction quantity approaches the maximum value (i.e., when a potentially optimal area is detected in the local environment). The confidence-weighted action selection mechanism ensures the intelligence of action selection, introduces an uncertainty measure in the Q value, solves the exploration problem of "actions with few visits but may be better", and gradually shifts from exploration to exploitation. The gradient perception enhancement mechanism injects prior knowledge of the physical electromagnetic field, accelerates local convergence, and combines with the confidence-weighted action selection mechanism to form a "physically heuristic" directional exploration strategy. Finally, it guides the agent to optimize action selection. In the initial stage of the algorithm, the high exploration rate and gradient guidance are used to quickly locate the high-voltage area, and in the later stage, it stabilizes to the optimal strategy through the attenuation mechanism, and finally converges to a pure greedy strategy.
[0052] Step S4-3: Update and iterate the Q-value table, and use the Bellman equation to iteratively update the Q-value table to gradually approach the optimal strategy. The Bellman equation is as follows: ; where Q t , Q t+1 are the Q-value sizes of the agent at time t and the next time t + 1 respectively; Q t+1 (s t ,at ) For the agent at time t+1, being in state s t and executing action a t The updated Q - value represents the expected return of this state - action pair under the current policy; s t and s t+1 are the states iterated by the agent at time t and the next time t + 1 respectively; A is the set of all actions, that is, A includes a, and a has four actions: up, down, left, and right; a t is the action executed by the agent at time t; R t+1 (s t , a t ) For the agent at time t, from state s t executing action a t and transferring to the next state s t+1 The cumulative reward obtained at time t + 1.
[0053] α is the learning rate, which represents the magnitude of the weight when updating the Q - value function each time, and determines the degree to which new information, that is, the currently obtained reward and the estimated future reward, covers the old information. The old information is set to 0.05; γ is the discount factor, which reflects the agent's degree of emphasis on future returns and its time preference. It is set to 0.9, which is beneficial for faster convergence to the optimal solution. The settings of the two balance the performance and stability of the algorithm model.
[0054] According to the Bellman equation, the Q - values of different states of the agent can be iteratively updated. During the training process, the Q - value table is further updated by continuously updating the Q - values, and finally the Q - value table converges and the policy is extracted. By querying the Q - value table, at time t, the agent selects the action corresponding to the maximum Q - value to execute to achieve the optimal policy.
[0055] Step S5: Calculate the cumulative reward according to the action executed by the agent; The reward function is designed as a feedback mechanism to provide feedback to the agent about the rationality of its behavior in a given environment or the quality of the algorithm strategy. For this task model, the reward function framework is formalized as a multi - objective optimization framework. Specifically, the reward function R t is defined as a composite function including four components: immediate rewardR 1 , step penalty R 2 , termination reward R 3 and boundary penalty R 4 .
[0056] (1) Immediate reward R 1 : Evaluate the immediate feedback received by the agent after executing an action, and encourage the agent to select actions that can increase the voltage of the receiving coil.
[0057] Immediate reward R 1 is proportional to the magnitude of the state variable, especially proportional to the pickup voltage of the receiving coil. A higher pickup voltage results in a greater direct return, guiding the agent to select energy transfer units that gradually approach the optimal receiver position.
[0058] (2) Step penalty R 2 : Reflects the system cost generated by activating each energy transfer unit, penalizes redundant actions, and encourages reaching the goal along the shortest path.
[0059] Step penalty R 2 Encourages the agent to determine the optimal energy transfer unit with the fewest activation steps, thereby minimizing the activation cost of the system. In the map of the algorithm grid, this is equivalent to reaching the destination along the shortest path. Therefore, the penalty increases proportionally with the number of activation steps.
[0060] (3) Termination reward R 3 : When the optimal matching state is reached, that is, when the agent reaches the position of the optimal energy transfer unit on the algorithm map, a relatively large termination reward is assigned. This signals to the agent that it has reached the destination.
[0061] (4) Boundary penalty R 4 : Restricts the agent's ineffective exploration in low-voltage areas.
[0062] In theory, the array emission system can infinitely expand its array scale according to the number of UAVs and the charging requirements, resulting in a potentially infinite search environment map for the model. However, due to the limited spatial range of the magnetic field in near-field coupling, energy transfer is effectively achieved only when the energy transfer unit near the excitation receiving coil is excited. Therefore, for the agent, it is unnecessary and wasteful to explore and activate the energy transfer unit outside the area where the pickup voltage is almost zero. To solve this problem, a boundary penalty is imposed on the agent, restricting the search environment map of the model to the area where the pickup voltage is almost zero and preventing over-exploration beyond this area.
[0063] The above reward settings are designed to ensure fast model convergence and generate reasonable algorithm strategies. Therefore, the reward function accumulates rewards R t The calculation formula is: ; where step is the number of times the agent executes an action.
[0064] Meanwhile, after model training, the specific parameters of the reward function in the algorithm model are finally set as: ; Step S6: Evaluate the action strategy output by the agent according to the evaluation network, and determine whether the model training strategy meets the initialization requirements. If yes, execute step S7; if not, execute step S4.
[0065] The evaluation network refers to calculating the ratio m : m = k / d ; where k is the number of actions executed by the agent, and d is the Manhattan distance between the initially randomly activated energy transfer unit and the optimal energy transfer unit of the agent.
[0066] When m approaches 1, it meets the initialization requirements and step S7 is executed; otherwise, step S4 is executed.
[0067] Given the diversity in the reconstruction types and sizes of array emission platforms, to ensure that this method can be generally applicable to array multi-coil systems in various reconstruction forms, grid map points are used to represent individual energy emission units. In this process, the distance between grid points does not directly depend on the physical spacing of the actual multi-coils, but a normalization design strategy is adopted, such that the distance between any adjacent grids is uniformly set to 1. Additionally, for a specific landing scenario during each training process, the agent starts exploring from a randomly selected initial grid point, and through continuous exploration, finally locates the optimal energy transmission unit, thereby achieving the generalization ability of training. Also, given that the agent starts each training from a random point as the initial exploration position, therefore, simply counting the number of actions performed by the agent, i.e., the system startup cost, cannot comprehensively evaluate the effectiveness of its strategy. To address the complexity of balancing exploration difficulty and strategy effectiveness, the proposed algorithm model first defines the Manhattan distance between two points ( X 1 , Y 1 ) and ( X 2 , Y 2 ) as: d t ; ; When m value asymptotically approaches 1 during model training, it indicates that the agent starting from a random initial point on the current grid map has achieved an action count equal to the Manhattan distance between the two points. This observation directly verifies the achievement of the optimization goal, confirming that the agent has completed the search task through the shortest path, i.e., achieving the optimal matching of the coils.
[0068] This metric of the evaluation network proposed by the present invention is used to monitor the training process and evaluate the convergence state of the algorithm. When it is observed during the training process that m value gradually approaches 1 with a small fluctuation amplitude, this indicates that the agent has been able to effectively locate the optimal energy transmission unit through the shortest path when randomly starting a certain emission unit, achieve the precise matching of the coils, and minimize the system startup cost. At this time, it can be considered that the training of the model has reached the convergence state and can make intelligent decisions based on the current state information. Therefore, at this stage, the algorithm strategy can be directly extracted and the Q-value table generated during the learning process can be stored for subsequent direct query and use. At the same time, the position coordinates of the array units activated by the current system can be output, and based on this position, the position of the receiving coil of the power receiving device can be located, thereby completing the coil positioning and coil matching. Conversely, if mIf the value does not meet this convergence criterion, it indicates that the agent is still in the disordered exploration and learning stage and needs to perform more actions to find the optimal energy transfer unit. Therefore, the agent will continue to explore the environment.
[0069] Step S7: Extract the Q-value table and output the coil positioning and matching strategy, that is, the optimal energy transfer unit and the shortest activation path.
[0070] This application proposes a wireless charging multi-coil fast matching and positioning method based on Q-learning the algorithm, which adaptively learns the dynamically changing environment, activates the optimal energy transfer unit with the minimum system cost, that is, the shortest path, and greatly reduces the activation times compared with traversing and activating to search. Moreover, as the scale of the array layout increases, the matching efficiency will be further improved.
[0071] At the hardware level, the multi-coil array transmitting platform with a reconfigurable array configuration cooperates with a flexible power supply network, can freely activate the required coils according to the algorithm, provides a hardware basis for realizing efficient charging, and its coil configuration and power supply network design are flexible, suitable for any reconfigurable form of array multi-coil matching strategy.
[0072] From the perspective of the algorithm, the action selection strategy is optimized through a variety of innovative mechanisms, without complex and cumbersome calculations and theoretical derivations. The dynamic exploration rate adjustment mechanism adjusts the exploration rate according to the state access density and the sensitivity of the induction quantity gradient, balances exploration and exploitation, enables the model to converge quickly, and effectively avoids falling into local optimal solutions. The confidence-weighted action selection mechanism quantifies the uncertainty of the action value, encourages the exploration of actions with low access times, optimizes the cold start process, and improves the intelligence of action selection. The gradient perception enhancement mechanism integrates physical characteristics into the model, guides the agent to explore along the direction of increasing induction quantity, and accelerates local convergence. At the same time, the reward function design provides reasonable feedback for the agent, and the multi-objective optimization framework comprehensively considers from aspects such as immediate reward, step penalty, termination reward, and boundary penalty to ensure that the model converges quickly and generates reasonable strategies. In addition, this application uses an evaluation network to evaluate the agent's action strategy by ratio, effectively monitors the training process and the algorithm convergence state, and ensures the effectiveness of model training.
[0073] In summary, the present invention autonomously learns the charging environment information collected by the algorithm, conducts intelligent derivation, reduces the tediousness of theoretical calculations, and improves the decision-making efficiency and accuracy. This transformation from data-driven to intelligent decision-making will provide a more flexible, efficient, and reliable coil positioning and matching strategy for wireless power transmission, so as to activate the optimal energy transfer unit with the shortest path, and then realize efficient and stable wireless power transmission. Specific embodiments The Q - value table records the expected return, i.e., the Q - value, in the form of state - action pairs for performing different actions (i.e., turning on the next energy transfer unit in any of the up, down, left, or right directions) under different system states. For the selection of the induction quantity, the current receiving coil pickup voltage is defined as the state detection quantity as follows: ; where, S t is the state of the agent's pickup voltage at the current time t; U t represents the pickup voltage of the receiving coil at the current time t, and different subscripts represent the magnitudes of the pickup voltages of the receiving coils in different states.
[0075] In addition, the actions that the agent can choose each time are defined as turning on the energy transfer units adjacent to the current three - phase array energy transfer unit in the four directions of up, down, left, and right. Then the actions are: .
[0076] By continuously exploring the voltage distribution under the random landing positions and attitudes of the drone, after the algorithm model converges, it can quickly and accurately determine the startup sequence of the three - phase units in the array according to the current receiving coil pickup voltage. Taking one case as an example, Figure 2 shows the voltage distribution situation explored by the model. Different colors represent different intervals of voltage values. Extract the Q - value table. According to the output strategy, the agent can randomly turn on a certain energy transfer unit, and output the specified coil opening path and direction according to the current receiving coil pickup voltage state, and finally find the optimal output position coordinates of the energy - emitting unit.
[0077] Please refer to Figure 3 , in the 8x8 state - space configuration, using the traversal search method shown in the figure, the arrow represents the execution direction, and 64 switch operations are required to find the optimal energy transfer unit (asterisk position).
[0078] In contrast, please refer to Figure 4 , the coil matching and positioning method proposed by the present invention significantly optimizes this process, more intelligently selects the execution direction (arrow direction), and the right side of the figure shows the relationship between the pickup voltage and the color. In the action - selection strategy, in the initial stage of the agent's exploration (the total number of visits to state s( C ( s ) < 5): Try different actions with an exploration rate of 80%, and at the same time, gradient guidance makes it prefer to move in the direction of increasing voltage. In the middle stage ( C ( s ) ≈ 20): The exploration rate drops to about 50%, and the confidence - weighted strategy focuses on high - Q - value actions, but still retains a certain exploration ability. In the later stage ( C ( s)>100): The exploration rate is close to 0, fully relying on high-confidence actions with gradient alignment, and stably maintaining the maximum voltage. It requires a minimum of only 2 searches and a maximum of no more than 7 switching actions, reducing the average number of system switches to 5 times. Figure 5 Shows the algorithm convergence evaluation metrics of the algorithm model of the present invention during the training process m The change process. After training and iteration, the proposed algorithm model can stably converge to around 1, which means that the algorithm model has the ability to reach the optimal energy transmission unit along the shortest path at this time, completing the coil matching process.
[0079] The method proposed by the present invention has achieved a 92.2% improvement in efficiency compared with traditional traversal search, greatly accelerating the process of positioning and coil matching. In addition, as the state space is further expanded, the efficiency of this method shows a continuous increasing trend, effectively reducing the overall cost of system switching operations.
[0080] It can be seen that the method proposed by the present invention not only realizes the high efficiency of energy transmission, but also brings significant cost benefits to the system by reducing the operation frequency, demonstrating its superior performance in the field of energy transmission.
[0081] In addition, under the array emission platform composed of 19 coils shown in Figure 6 , 5 drone positions were randomly selected, corresponding to the five positions P1, P2, P3, P4, and P5 in the figure, for generalization verification. The experimental results are shown in the following table. It can be seen from the experimental results that if traversal is used to achieve coil matching, 24 energy transmission units need to be traversed at these five positions to find the optimal energy transmission unit. In the method of the present invention, the number of actions of the coils is much less than the number of times of searching for the best energy transmission unit by traversing and opening the coils. At the same time, the algorithm model output by the model can already intelligently output the energy transmission unit according to the current state, that is, the finally matched unit is obtained accordingly, and the position estimation error is less than 4.2 cm. It should be noted that for the measurement of the relative error size, the ultimate goal of the method of the present invention is the best matching of the array coils, that is, the currently opened energy transmission unit is best matched with the drone receiving coil to achieve the function of efficient energy transmission. Therefore, the model can finally output the central coordinates of the optimal energy transmission unit that is finally opened, rather than outputting the actual position of the drone receiving coil, but still has a certain position estimation ability, and the position estimation error is not greater than 4.2 cm. Therefore, it can be considered that the proposed model takes into account the functions of drone positioning and coil matching.
[0082] Table 1 Coil matching results
[0083] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in the present invention can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0084] The above-described embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of each embodiment of this application, and should all be included within the protection scope of this application.
Claims
1. A method based on Q-learning The wireless charging multi-coil fast matching and positioning method based on the algorithm is characterized by: The following steps are involved: Step S1: Establish a multi-coil array transmitting platform and a power supply network that enables flexible switching of corresponding energy transmission units of the platform; Step S2: changing the position of the receiving coil and the opening status of different energy transmission units, obtaining the induction quantity, and establishing the charging environment information; Step S3: Establish an exploration environment under the algorithm model, model a discretized two-dimensional grid map, each grid corresponds to the energy transmission unit in the array launch platform, define the state S of the grid as the corresponding induction quantity, and initialize the Q value of all grids; Step S4: The receiving coil induction changes and enters the coil matching process. The intelligent agent explores in the unknown environment, performs actions according to the action selection strategy, and iteratively updates the Q value table; the action selection strategy includes a dynamic exploration rate adjustment mechanism, a confidence-weighted action selection mechanism, and a gradient perception enhancement mechanism; Step S5: Calculate the cumulative reward based on the actions performed by the agent; Step S6: Evaluate the action strategy output by the agent based on the evaluation network to determine whether the model training strategy meets the initialization requirements. If yes, execute step S7; If no, proceed to step S4; Step S7: extract the Q value table to complete the coil matching process, and output the optimal energy transfer unit and the shortest opening path.
2. The method according to claim 1 Q-learning The wireless charging multi-coil fast matching and positioning method based on the algorithm is characterized by: The dynamic exploration rate adjustment mechanism dynamically adjusts the exploration rate through the state access density and the sensitivity of the sensing gradient.
3. The method according to claim 2 Q-learning The wireless charging multi-coil fast matching and positioning method based on the algorithm is characterized by: The dynamic exploration rate adjustment mechanism is implemented as follows: (1) Establishing the exploration rate decay function : ; in, N (s,a) is the cumulative number of executions of action a in state s, reflecting the familiarity of the agent’s actions; C (s) is the total number of visits to state s, which indicates the degree of exploration of the state by the current agent; β is the decay rate parameter, which is used to control the speed at which the current exploration rate decreases; The curvature parameter is used to adjust the nonlinearity of the attenuation curve; ε 0 is the initial exploration rate, i.e., the maximum exploration probability; (2) Establish a state access density threshold to force exploration of low-access areas: ; The above constraints force the agent to maintain an exploration rate of more than 30% when the agent appears in a new state, that is, the number of visits is less than 5; (3) Introduce a sensitivity trigger mechanism for the induction gradient, which temporarily increases the agent's exploration rate when the induction is close to the maximum value, that is, when the induction gradient change is less than 10%.
4. The method according to claim 1 or 2 Q-learning The wireless charging multi-coil fast matching and positioning method based on the algorithm is characterized by: The confidence-weighted action selection mechanism quantifies the uncertainty of action value through confidence indicators, and guides the exploration of actions with low access times.
5. The method according to claim 4 Q-learning The wireless charging multi-coil fast matching and positioning method based on the algorithm is characterized by: The confidence-weighted action selection mechanism is specifically implemented as follows: (1) Calculate action confidence index Q c (s,a) , the formula is as follows: ; in, Q(s,a) Representing tradition Q-learning The action-value function under the algorithm; κ is the confidence coefficient, which can control the weight of exploration and utilization. The larger the value, the more inclined to exploration; C(s) is the total number of visits to state s; N (s,a) is the cumulative number of executions of action a in state s; j is a zero-proof constant; and the degree of underexploration of the action is represented by the uncertainty term √(ln(1+C(s)) / N(s,a)); (2) Actions with high confidence are given higher selection probabilities, and the action probability distribution of the agent in state s satisfies: ; in, P(a ∣ s) The agent is in state s When you select Action a The probability of Q c (s,a) Correction after introducing action confidence index Q value; a ′ are all the actions that the agent can perform in state s; Q c (s,a ′ ) Correction for the agent to perform different actions in state s Q value; Status s Related temperature parameter dynamic adjustment factors; Introducing dynamic adjustment factors for temperature parameters ,Through the degree of certainty of the temperature parameter adaptive adjustment strategy, the confidence index of the modified Q value is further converted into a probability distribution, and the temperature adjustment factor changes the probability distribution from flat to sharp as the agent exploration process proceeds, the agent eventually tends to a purely greedy state, and the current optimal state is intelligently selected; ; in, is the dynamic adjustment factor of the temperature parameter related to state s; C (s) is the total number of visits to state s, which indicates the degree of exploration of the state by the current agent; η is a hyperparameter.
6. The method according to claim 4 Q-learning The wireless charging multi-coil fast matching and positioning method based on the algorithm is characterized by: The gradient perception enhancement mechanism uses the induction quantity in the wireless charging environment to encode the field strength distribution information into the action selection, calculates the induction quantity gradient vector field and the direction preference weight, superimposes the direction preference weight on the basis of the confidence-weighted Q value, and gives priority to actions that are both high-value and consistent with the gradient direction.
7. The method according to claim 1 Q-learning The wireless charging multi-coil fast matching and positioning method based on the algorithm is characterized by: The iterative updating of the Q value table refers to iteratively updating the Q value table using the Bellman equation: ; in, Q t , Q t+1 They are the Q values of the agent at time t and the next time t+1 respectively; Q t+1 (s t ,a t ) is the agent at time t+1 , In state s t , Execute actions a t When , the updated Q value indicates the expected value return of the state-action pair under the current strategy; s t , s t+1 are the states of the agent at time t and the next time t+1 respectively; A is the set of all actions, that is, A includes a, and a has four actions: up, down, left, and right; a t is the action performed by the agent at time t; R t+1 (s t , a t ) The agent at time t, from state s t Execute an action a t Transition to next state s t+1 After that, the cumulative reward obtained at time t+1; α is the learning rate; γ is the discount factor.
8. The method according to claim 1 Q-learning The wireless charging multi-coil fast matching and positioning method based on the algorithm is characterized by: The cumulative reward R t The calculation formula is: ; in, R 1 is an immediate reward, R 2 is the step penalty, R 3 is the termination reward, R 4 is the boundary penalty, and step is the number of times the agent performs an action.
9. The method according to claim 1 Q-learning The wireless charging multi-coil fast matching and positioning method based on the algorithm is characterized by: The evaluation network refers to the calculation ratio m : m = k / d ; in, k is the number of actions performed by the agent, d The Manhattan distance between the initial energy transfer unit randomly opened by the agent and the optimal energy transfer unit; when m When it approaches 1, the initialization requirement is met.
Citation Information
Patent Citations
Intelligent power plant autonomous polling robot polling system and method
CN109599945A
Mobile blind guiding robot system and blind guiding method
CN111609851A
Optimal decision-making method based on improved Q-learning
CN112598137A
Master-slave water surface robot recovery guiding method based on environment driving
CN115328143A
Unmanned aerial vehicle control method and system based on path planning
CN115421517A
Cited By
Battery replacement cabinet distributed energy management control system and method based on edge calculation
CN121172914A