A Wireless Charging Multi-Coil Fast Matching and Positioning Method Based on Q-Learning Algorithm
Through the fast matching and positioning method of wireless charging multi-coil based on the Q-learning algorithm, the problem of complex calculation and insufficient response capabilities of position detection algorithms in the prior art is solved, and efficient matching and positioning of wireless charging systems is achieved, which is suitable for extreme weather conditions.
Patent Information
- Application Number
- CN202510510107.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-04-23
AI Technical Summary
In the existing wireless charging technology, the position detection algorithm relies on front-end auxiliary sensor equipment, has complex calculations and insufficient responsiveness, especially in extreme weather conditions.
The wireless charging multi-coil fast matching and positioning method based on the Q-learning algorithm is adopted. By establishing a multi-coil array transmission platform and a flexible power supply network, combining dynamic exploration rate adjustment, confidence weighted action selection and gradient perception enhancement mechanism, adaptively learn the dynamic changing environment to achieve rapid matching and positioning of the optimal energy transmission unit.
It significantly reduces the number of system openings, improves matching efficiency, and can maintain efficient performance under extreme weather conditions without requiring complex calculations and theoretical derivation.
Smart Images

Figure CN120049642B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to a method for fast matching and positioning of multiple coils for wire charging, and more specifically, to a method for fast matching and positioning of multiple coils for wireless charging based on Q-learning an algorithm. Background Art
[0002] Wireless charging technology realizes wireless power transmission, completely getting rid of the bondage of cables. This non-contact charging method avoids safety hazards such as electric shock, leakage, and electric sparks, ensuring the safety and reliability of the operation process. At the same time, it reduces the downtime caused by battery replacement and does not require manual intervention, realizing autonomous unmanned charging.
[0003] The reconfigurable array magnetic structure can effectively alleviate the problem of misalignment when the power receiving device lands, and improve the system's tolerance to misalignment. However, the prerequisite is that it can intelligently select and activate the transmitting coil that highly matches the position of the power receiving device according to the random landing position and direction of the power receiving device (receiving coil), so as to ensure the efficient operation and performance optimization of the entire charging process.
[0004] Although significant progress has been made in the current research on position detection algorithms, there are still multiple challenges. The primary problem is that algorithm design mostly relies on the information collected by front-end auxiliary sensor devices, and does not fully integrate the intrinsic information in the specific background of wireless power transmission for power receiving devices such as drones. For example, the introduction of technologies such as infrared, laser, and image recognition has improved the positioning accuracy, but also increased the drone load, system cost, and complexity. Especially in extreme weather conditions, the effectiveness of image recognition algorithms drops significantly. Due to insufficient image clarity and contrast or label failure, it is difficult to extract features and the position detection ability is limited.
[0005] Regarding the computational complexity and real-time requirements of the algorithm, the current position detection methods applied in the context of wireless power transmission mainly theoretically deduce the position by manually calculating voltage, current, and mutual inductance, mainly realizing detection from the theoretical calculation level. The calculation is rather cumbersome and complex, and the accompanying computational burden limits its fast response ability in practical applications. At the same time, the limitations of the algorithm dataset, especially the lack of methods for collecting and establishing dedicated datasets for the field of wireless power transmission, limit the generalization ability and application depth of the algorithm. Summary of the Invention
[0006] To solve the technical problems that existing algorithms mostly rely on front-end information collection, are computationally cumbersome, and have limited response capabilities, the technical solution adopted in the present application is: to provide a method for fast matching and positioning of multiple coils for wireless charging based on Q-learning an algorithm, including the following steps:
[0007] Step S1: Establish a multi-coil array emission platform and a power supply network that enables flexible switching of the corresponding energy transmission units of the platform;
[0008] Step S2: Change the position of the receiving coil and the activation status of different energy transmission units, obtain the induction quantity, and establish charging environment information;
[0009] Step S3: Establish an exploration environment under an algorithm model, model a discretized two-dimensional grid map, where each grid corresponds to an energy transmission unit in the array emission platform, define the state S of the grid as the corresponding induction quantity, and initialize the Q values of all grids;
[0010] Step S4: When the induction quantity of the receiving coil changes and enters the coil matching process, the agent explores in an unknown environment, executes actions according to the action selection strategy, and iteratively updates the Q value table; the action selection strategy includes a dynamic exploration rate adjustment mechanism, a confidence-weighted action selection mechanism, and a gradient perception enhancement mechanism;
[0011] Step S5: Calculate the cumulative reward according to the actions executed by the agent;
[0012] Step S6: Evaluate the action strategy output by the agent based on the evaluation network, and determine whether the model training strategy meets the initialization requirements. If yes, execute Step S7; if not, execute Step S4;
[0013] Step S7: Extract the Q value table and output the optimal energy transmission unit and the shortest activation path.
[0014] Preferably, the dynamic exploration rate adjustment mechanism dynamically adjusts the exploration rate through the state access density and the induction quantity gradient sensitivity.
[0015] Preferably, the implementation process of the dynamic exploration rate adjustment mechanism is as follows:
[0016] (1) Establish an exploration rate decay function :
[0017] ;
[0018] Wherein, N (s,a) is the cumulative execution times of action a in state s, reflecting the familiarity of the agent's actions; C (s) is the total access times of state s, indicating the degree of exploration of the current agent for the state; β is the decay rate parameter, used to control the current exploration rate decline speed; is the curvature parameter used to adjust the nonlinear degree of the decay curve; ε 0 is the initial exploration rate, that is, the maximum exploration probability.
[0019] (2) Establish a state access density threshold to force exploration of low-access areas:
[0020] ;
[0021] Through the above constraint conditions, when a new state appears, that is, when the access count is less than 5, the agent is forced to maintain an exploration rate of more than 30%;
[0022] (3) Introduce an induction quantity gradient sensitivity trigger mechanism to temporarily increase the agent's exploration rate when the induction quantity is close to the maximum value, that is, when the change in the induction quantity gradient is less than 10%;
[0023] Preferably, a confidence-weighted action selection mechanism quantifies the uncertainty of action value through a confidence index to guide the exploration of actions with low access counts.
[0024] Preferably, the specific implementation method of the confidence-weighted action selection mechanism:
[0025] (1) Calculate the action confidence index Q c (s, a) , the formula is as follows:
[0026] ;
[0027] Among them, Q(s, a) represents the action value function of the traditional Q-learning ; κ is the confidence coefficient, which can control the weight of exploration and exploitation. The larger the value, the more inclined to explore; C(s) is the total number of accesses to state s; N (s,a) is the cumulative execution count of action a in state s; j is a constant to prevent division by zero; the uncertainty term √(ln(1 + C(s)) / N(s,a)) represents the degree of insufficient exploration of the action;
[0028] (2) Actions with higher confidence are given a higher selection probability. The action probability distribution of the agent in state s satisfies:
[0029] ;
[0030] Among them, P(a ∣ s) is the probability that the agent selects action s when it is in state a ; Q c (s, a) is the introduced action confidence index; a ′ are all executable actions of the agent in state s; Q c(s, a ′ ) For performing an action a ′ The action confidence index in a case; For the state s The dynamic adjustment factor of the temperature parameter related thereto;
[0031] Introduce the dynamic adjustment factor of the temperature parameter , through the certainty degree of the temperature parameter adaptive adjustment strategy, further convert the confidence index of the corrected Q value into a probability distribution. As the temperature adjustment factor changes during the agent's exploration process, the probability distribution changes from flat to sharp, and finally the agent tends to a pure greedy state and intelligently selects the current optimal state.
[0032] ;
[0033] Wherein , For the state s The dynamic adjustment factor of the temperature parameter related thereto; C N(s) is the total number of visits to state s, indicating the degree of exploration of the current agent for the state; η Is a hyperparameter.
[0034] Preferably, the gradient perception enhancement mechanism encodes the field strength distribution information into the action selection by using the induction amount in the wireless charging environment, calculates the induction amount gradient vector field and the direction preference weight, and superimposes the direction preference weight on the confidence-weighted Q value to preferentially select an action that is both of high value and consistent with the gradient direction.
[0035] Preferably, the Q-value table is iteratively updated, which means that the Q-value table is iteratively updated using the Bellman equation:
[0036] ;
[0037] Wherein, Q t 、 Q t+1 Are the Q-value sizes of the agent at time t and the next time t + 1 respectively; Q t+1 (s t ,a t ) Is the agent at time t+1, In the state s t 、Performing the action a t When, the updated Q-value size, indicating the expected return of the state-action pair under the current policy;s t and s t+1 are the iterative states of the agent at time t and the next time t+1 respectively; A is the set of all actions, that is, A includes a, and a has four actions: up, down, left, and right; a t is the action executed by the agent at time t; R t+1 (s t , a t ) is the cumulative reward obtained by the agent at time t+1 after transferring from state s t by executing action a t to the next state s t ; α is the learning rate; γ is the discount factor.
[0038] Preferably, the cumulative reward R t has the following calculation formula:
[0039] ;
[0040] wherein, R 1 is the immediate reward, R 2 is the step penalty, R 3 is the termination reward, R 4 is the boundary penalty, and step is the number of times the agent executes the action.
[0041] Preferably, the evaluation network refers to calculating the ratio m :
[0042] m = k / d ;
[0043] wherein, k is the number of actions executed by the agent, d is the Manhattan distance between the initial energy transfer unit randomly opened by the agent and the optimal energy transfer unit;
[0044] When m approaches 1, the initialization requirement is satisfied.
[0045] The beneficial effects of the present invention are as follows. The present invention proposes a method based on Q-learningWireless charging multi-coil fast matching and positioning method of the algorithm, adaptively learning the dynamically changing environment, enabling the optimal energy transmission unit to be turned on with the minimum system cost, i.e., the shortest path, greatly reducing the number of turn-on times compared with traversing and searching, and as the scale of the array layout increases, the matching efficiency will be further improved.
[0046] At the hardware level, the multi-coil array transmitting platform with a reconfigurable array configuration cooperates with a flexible power supply network, can freely turn on the required coils according to the algorithm, provides a hardware foundation for realizing efficient charging, and its coil configuration and power supply network design are flexible, suitable for any reconfigurable form of array multi-coil matching strategy.
[0047] From the perspective of the algorithm, the action selection strategy is optimized through various innovative mechanisms, without complex and cumbersome calculations and theoretical derivations. The dynamic exploration rate adjustment mechanism adjusts the exploration rate according to the state access density and the sensitivity of the induction quantity gradient, balancing exploration and exploitation, enabling the model to converge quickly and effectively avoiding falling into local optimal solutions. The confidence-weighted action selection mechanism quantifies the uncertainty of the action value, encourages the exploration of actions with low access times, optimizes the cold start process, and improves the intelligence of action selection. The gradient perception enhancement mechanism integrates physical characteristics into the model, guiding the intelligent agent to explore along the direction of increasing induction quantity and accelerating local convergence. At the same time, the reward function design provides reasonable feedback to the intelligent agent, and the multi-objective optimization framework comprehensively considers aspects such as immediate reward, step penalty, termination reward, and boundary penalty to ensure that the model converges quickly and generates reasonable strategies. In addition, this application uses an evaluation network to evaluate the intelligent agent's action strategy by ratio, effectively monitoring the training process and the algorithm convergence state, and ensuring the effectiveness of model training. Brief Description of the Drawings
[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0049] Figure 1 A flow schematic diagram of a wireless charging multi-coil fast matching and positioning method based on Q-learning the algorithm provided in an embodiment of the present application;
[0050] Figure 2 A schematic diagram of the pickup voltage distribution provided in an embodiment of the present application;
[0051] Figure 3 A schematic diagram of the strategy path of a wireless charging multi-coil fast matching and positioning method based on Q-learning the algorithm provided in an embodiment of the present application;
[0052] Figure 4 Schematic diagram of the traversal strategy path provided by an embodiment of the present application;
[0053] Figure 5 A kind provided by an embodiment of the present application based on Q-learning Algorithm for wireless charging multi - coil fast matching and positioning method of m Value change diagram;
[0054] Figure 6 Array transmitting platform composed of 19 coils provided by an embodiment of the present application. Detailed implementation manners
[0055] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0056] Please refer to Figure 1 , A flow chart of a wireless charging multi - coil fast matching and positioning method based on Q-learning Algorithm provided by an embodiment of the present application. For the convenience of description, only the parts related to this embodiment are shown and are described in detail as follows:
[0057] In one of the embodiments, a wireless charging multi - coil fast matching and positioning method based on Q-learning Algorithm includes:
[0058] Step S1: Establish a multi - coil array transmitting platform and a power supply network for realizing flexible switching of corresponding coils; wherein, the power supply network in the array transmitting platform is connected to each coil, and can realize the free activation of the array multi - coils, and can ensure flexible activation of the coils to be activated according to the algorithm output strategy.
[0059] Specifically, the reconfigurable array - type multi - coil array transmitting platform is composed of multiple transmitting coil arrays arranged, and the coil configuration can be in various forms such as circular, square, hexagonal, etc.
[0060] The power supply network can be freely designed according to the wireless charging energy transfer requirements and output conditions. At the same time, the design is not limited to using single-phase, two-phase, three-phase, four-phase and other excitation modes. Different array multi-coil excitation modes result in different composition methods of each coil energy transfer unit. For example, every three adjacent transmitting coils can be excited to form a freely combinable energy transfer unit. This means that in the subsequent matching process between the transmitting coil and the receiving coil of the power receiving device, according to the algorithm output, in each search process of the system, the excitation mode of the array multi-coil is three adjacent transmitting coils, and each three adjacent transmitting coils form a new energy transfer unit to transfer energy to the power receiving device. Therefore, the power supply network only needs to meet the requirement that the transmitting coil can self-organize to turn on and excite the corresponding array coils.
[0061] Step S2: Change the position of the receiving coil and the activation of different energy transfer units, obtain the coil induction, and extract data information to establish the charging environment information;
[0062] Change the position of the receiving coil and the activation of different energy transfer units to extract the induction in any case. The set of these inductions is the charging environment information. For different position conditions of the power receiving device (including the receiving coil) and different activation conditions of the energy transfer units, the induction is different. This induction reflects the matching state between the currently activated energy transfer unit and the receiving coil of the power receiving device. According to the comparison between the set threshold of the induction for the best matching state and the current induction, if the induction reaches the set threshold, it is considered that the current system has reached the best matching state, and the matching and positioning process is achieved. At the same time, the induction is related to the selection of subsequent state variables because it is used as the state variable value of the algorithm agent. Specifically, the selection of the induction can include any data with intuitive measurability such as the pickup voltage of the receiving coil, the induction current of the receiving coil, or the excitation current of the transmitting coil. Among them, the induction can be reflected on the receiving coil side or the transmitting coil side, which is specifically set according to the selection of the induction.
[0063] In particular, the selection of the induction does not use the basic mutual inductance or the secondary derived theoretical index as the measurement index for the induction situation, so as to avoid cumbersome theoretical derivations and the establishment of precise theoretical models to improve the overall detection efficiency.
[0064] Step S3: According to the charging environment information extracted in Step S2, establish Q-learning the exploration environment under the algorithm model.
[0065] Specifically, the charging environment information is modeled as a discretized two-dimensional grid map. Each grid represents a potential energy transfer unit activation state, and the coordinates (x, y) of the grid correspond to the position of the energy transfer unit in the array transmitting platform.
[0066] For the representation of the grid network state, the state S of each grid is defined as the voltage value U picked up by the receiving coil when the energy transfer unit corresponding to the grid is activated. t (x, y). This voltage value is obtained by measuring with the sensor in step S2. At the same time, for the environmental initialization setting, the Q values of all grids in the grid map are initialized to zero or a small random value. This means that the agent has no prior knowledge of the value of all actions in the initial state.
[0067] Furthermore, the size of the grid map is determined according to the physical size of the transmitting array and the number of coils. Specifically, for the gridification method, a uniform gridification network is adopted, and the area covered by the transmitting array is divided into grids of equal size and normalized to eliminate the differences between different array coils. The choice of grid size needs to balance the positioning accuracy and the computational complexity. For example, if the transmitting array consists of 8x8 energy transfer units, the grid map can be set to 8x8.
[0068] Step S4: The agent explores in the unknown environment, executes actions and updates the state; specifically, the implementation process is as follows:
[0069] Step S4-1: Definition of state and action:
[0070] Define the state s to represent the induction amount of the receiving coil when the agent activates the current energy transfer unit, reflecting the coupling state of the receiving coil.
[0071] In each state s, the agent can execute the action a to iterate the state. The action a is set to turn on the next energy transfer unit of the current energy transfer unit to the left or right or up or down.
[0072] For a detailed description of this, define the state s to represent the induction amount of the receiving coil when the agent activates the current energy transfer unit. The current state s can reflect the current coil coupling state, and at the same time, the matching situation of the current system coils can be further determined according to the state quantity, which is reflected as the agent moving at different grid network points. That is, define the state s to represent the coordinate position of each grid in the exploration map.
[0073] Due to the multi-voltage extreme value characteristics in the wireless charging environment based on near-field coupling characteristics, for each action, the reconstruction directions in the horizontal and vertical directions of the current energy transmission unit are set to be enabled to simplify the control drive form, and at the same time, a refined search is performed to find the best matching coil. Specifically, when the agent is in the exploration environment under the algorithm model, in the current state, that is, after a certain energy transmission unit of the array emission platform is enabled, on the basis of detecting a certain induction amount, the action a for updating the state is set to continue to enable the next energy transmission unit of the current energy transmission unit to the left or right or up or down, that is, the agent moves up, down, left, and right at the current grid network point. The agent's action setting can ensure any reconstruction method of the densely paved array coils of the array emission platform, avoiding getting stuck in local optimal solutions or missing the optimal solution due to a large step size. Therefore, the agent state transition equation is:
[0074] ;
[0075] Among them, S t 、S t+1 are the states iterated by the agent at time t and the next time t + 1 respectively, and a represents the action that the agent can execute;
[0076] Step S4-2: Action selection strategy:
[0077] In the traditional ε - greedy strategy, the agent will select the action with the largest current Q value with a probability of 1- ε , that is, exploitation, and randomly select an action with a probability of ε , that is, exploration. The purpose of this strategy is to maximize the benefit by using known information while also keeping the model having a certain degree of exploration. However, this strategy has the defect of a fixed exploration rate, resulting in slow convergence of the agent, thereby increasing the training time and cost and being unable to make decisions quickly.
[0078] Therefore, to address this problem, the present invention proposes an optimized and improved action selection strategy. The specific key mechanisms include dynamic exploration rate adjustment, confidence-weighted action selection, and gradient-aware enhanced fusion and collaboration mechanisms. The details are as follows:
[0079] (1) Dynamic exploration rate adjustment mechanism, specifically, the exploration rate is dynamically adjusted through the state access density and the sensitivity of the induction amount gradient to balance exploration and exploitation.
[0080] First, establish an exploration rate decay function :
[0081] ;
[0082] Among them, N(s, a) is the cumulative execution count of action a in state s, reflecting the familiarity of the agent's actions; C (s) is the total number of visits to state s, indicating the degree of exploration of the current state by the agent; β is the decay rate parameter, used to control the rate of decline of the current exploration rate, ∂ is the curvature parameter used to adjust the non-linearity of the decay curve; ε 0 is the initial exploration rate, i.e., the maximum exploration probability. The exploration rate decays as the number of state visits C ( s ) increases, and the decay rate is determined by the decay rate β and the curvature jointly controlled, β 、 、 ε 0 all range from 0 to 1 and can be dynamically adjusted according to the actual situation. The optimal values are obtained in specific embodiments β = 0.2, ∂ = 0.5, ε 0 = 0.8. By dynamically balancing exploration and exploitation through the state visit frequency, sufficient exploration is carried out in the early stage, and gradually converges to the optimal strategy in the later stage. β Controls the decay speed. The larger the value, the faster the decay. Adjusts the bending degree of the decay curve. The larger the value, the slower the initial decay.
[0083] Further, a state visit density threshold is established to force exploration of low-visit areas:
[0084] ;
[0085] Through the above constraint conditions, it is ensured that when a new state appears, that is, when the number of visits is less than 5, the exploration rate is forced to remain above 30%, preventing the model from converging prematurely and falling into a local optimum.
[0086] Even further, an induction quantity gradient sensitivity trigger mechanism is introduced to temporarily increase the exploration rate when the induction quantity is close to the maximum value, avoiding falling into a sub-optimal solution. Specifically, when approaching the maximum induction quantity (such as the maximum voltage region), that is, when the gradient is less than 10%, the exploration rate of the agent is forced to increase by 15% probability. The formula is as follows:
[0087] ;
[0088] Among them, U(s) is the induction quantity corresponding to state s, U max is the maximum induction quantity, ΔU is the induction quantity in state s U(s) and the maximum induction quantity U max The absolute value of the difference between them.
[0089] Through the setting of the above dynamic exploration rate adjustment mechanism, high explorability is maintained in low-access areas, and local search is enhanced in high-gradient pressure difference areas to achieve rapid convergence of the model, realizing a dynamic balance between the two. At the same time, through gradient detection, an adaptive adjustment and correction strategy is used to match the non-uniform discrete and multi-voltage extreme value characteristics of the wireless charging environment.
[0090] (2) Confidence-weighted action selection mechanism, which quantifies the uncertainty of action value through a confidence index to guide the exploration of actions with low access times.
[0091] Specifically combined with Q value guidance and uncertainty evaluation, replacing the traditional greedy strategy through an improved confidence index. On the basis of the original Q value, a confidence term negatively correlated with the action access times N ([[]] s , a ) is added. When N ([[]] s , a ) → 0, the confidence term increases significantly, encouraging the attempt of actions that have not been fully explored.
[0092] Specifically, an action confidence index Q c (s, a) is proposed, and the calculation formula is as follows:
[0093] ;
[0094] Among them, Q(s, a) represents the traditional Q-learning action value function; κ is the confidence coefficient, which can control the weight of exploration and exploitation. The larger the value, the more inclined to explore. κ = 1.2 is the empirically optimized value; C(s) is the total number of visits to state s; N (s, a) is the cumulative execution times of action a in state s; j = 10 −5 is a constant to prevent zero to ensure numerical stability; the uncertainty term √(ln(1 + C(s)) / N(s, a)) characterizes the degree of insufficient exploration of the action.
[0095] Furthermore, actions with high confidence are given a higher selection probability, and the action probability distribution of the agent in state s satisfies:
[0096] ;
[0097] Among them, P(a ∣ s) is when the agent is in state s , and selects actiona probability; Q c (s, a) is the action confidence index; a ′ are all executable actions of the agent in state s; Q c (s, a ′ ) is the action a ′ under the case of executing action is the state s related temperature parameter dynamic adjustment factor.
[0098] Introduce the temperature parameter dynamic adjustment factor , so that the more the agent's state is accessed, the sharper the probability distribution is, that is, it biases towards high-Q value actions, and when a new state is accessed, the distribution is smoother, that is, uniform exploration is encouraged. The specific calculation formula is: when C ( s ) → 0, ≈1.5, the probability distribution is more uniform, encouraging exploration; specifically, when C ( s ) → ∞, →1, the probability distribution is more concentrated, biasing towards high-Q value actions, and the certainty degree of the temperature parameter adaptive adjustment strategy is adjusted.
[0099] ;
[0100] where , is the state s related temperature parameter dynamic adjustment factor; C (s) is the total number of visits to state s, indicating the degree of exploration of the current agent for the state; η is a hyperparameter, and its value range is 0~1, which can be dynamically adjusted according to the actual situation. Here, an empirical value of 0.5 is taken.
[0101] Through the above confidence-weighted action selection mechanism, the selection probability of actions with fewer execution times is automatically increased, and at the same time, the action selection in the new state is smoother, avoiding the inefficiency of random exploration and realizing cold start optimization.
[0102] (3) Gradient perception enhancement mechanism. Utilize the induction quantity in the wireless charging environment (such as the physical characteristics of voltage gradient), encode the field strength distribution information into the action selection to achieve gradient perception exploration enhancement, strengthen the action movement direction along the direction of increasing induction quantity, realize gradient direction guidance, and at the same time integrate the actual physical laws into the model optimization strategy.
[0103] Specifically, taking the calculation of the voltage gradient field as an example, the voltage gradient vector field is calculated by the central difference method ∇U(s) , reflecting the direction in which the local voltage changes fastest. Its calculation formula satisfies:
[0104] ;
[0105] where U(x, y) is the induction at the coil energy transfer unit of coordinate (x, y) ; Δx = Δy = 1, indicating the adjacent coil spacing after normalization processing.
[0106] Furthermore, on this basis, the direction preference weight w dir (a) is set as:
[0107] ;
[0108] where v a is the unit vector of the movement direction corresponding to the action a (e.g., [-1, 0] for left); λ is the direction strengthening coefficient, used to control the intensity of the gradient influence and adjust the reward intensity of gradient alignment. The larger the value, the more obvious the direction preference. The empirical value λ = 0.8; ∇U(s) · v a is the dot product of the voltage gradient vector field ∇U(s) and the unit vector of the action direction v a , used to measure the consistency between the action direction and the gradient direction. If the direction of the action a is consistent with the voltage gradient direction (the dot product is large), then the weight w dir (a) increases.
[0109] Even further, the action along the gradient direction is weighted 1.8 times. That is, when λ = 0.8, only 20% of the value of the reverse gradient action is retained Q c . It realizes automatically tracking the voltage rising path in the non-uniform field, and at the same time quickly crosses the low-voltage array platform area through the gradient information. Based on the confidence-weighted Q value, the direction preference is superimposed, and the actions that are both of high value and consistent with the gradient direction are preferentially selected. The gradient direction indicates the path with the fastest voltage rise, and the weight adjustment enables the agent to explore along the "electric potential rising direction" and avoid random walks. The specific action selection is optimized as:
[0110] ;
[0111] Among them, a selected is the action finally selected and executed by the agent; argmax a is the maximum value among all actions a ; Q c (s, a) is the action confidence index.
[0112] In summary, in the action selection strategy, through the synergistic effect of the above three mechanisms, specifically, the dynamic adjustment of the exploration rate mechanism provides global search ability, with a high exploration rate (80%) in the initial stage and gradually decaying in the later stage. The induction quantity gradient sensitivity trigger mechanism temporarily increases the exploration rate when the induction quantity approaches the maximum value (i.e., when a potentially optimal area is detected in the local environment). The confidence-weighted action selection mechanism ensures the intelligence of action selection, introduces an uncertainty measure in the Q value, solves the exploration problem of "actions with few visits but may be better", and gradually shifts from exploration to exploitation. The gradient perception enhancement mechanism injects prior knowledge of physical electromagnetic fields, accelerates local convergence, and combines with the confidence-weighted action selection mechanism to form a "physically heuristic" directional exploration strategy. Finally, it guides the agent to optimize action selection. In the initial stage of the algorithm, it quickly locates the high-voltage area through high exploration rate and gradient guidance, and in the later stage, it stabilizes to the optimal strategy through the attenuation mechanism, and finally converges to a pure greedy strategy.
[0113] Step S4-3: Update and iterate the Q-value table, and use the Bellman equation to iteratively update the Q-value table to gradually approximate the optimal strategy. The Bellman equation is as follows:
[0114] ;
[0115] Among them, Q t , Q t+1 are the Q-value sizes of the agent at time t and the next time t+1, respectively; Q t+1 (s t ,a t ) is the agent at time t+1, in state s t , executes action a t , the updated Q-value size, representing the expected return of this state-action pair under the current strategy; s t , s t+1are the states of the agent at time t and the next time t+1 during iteration respectively; A is the set of all actions, that is, A includes a, and a has four actions: up, down, left, and right; a t is the action executed by the agent at time t; R t+1 (s t , a t ) is the agent at time t, from state s t executes the action a t transfers to the next state s t+1 and the cumulative reward obtained at time t+1.
[0116] α is the learning rate, which represents the magnitude of the weight when updating the Q-value function each time, determines the degree to which new information, that is, the current obtained reward and the estimated future reward, covers the old information, and the old information is set to 0.05; γ is the discount factor, which reflects the agent's emphasis on future rewards and its time preference, is set to 0.9, which is conducive to converging to the optimal solution faster, and the settings of the two balance the performance and stability of the algorithm model.
[0117] According to the Bellman equation, the Q-values of the agent in different states can be iteratively updated, and the Q-value table is further updated by continuously updating the Q-values during the training process, and finally the Q-value table converges and the policy is extracted. By querying the Q-value table, it can be obtained that at time t, the agent selects the action corresponding to the maximum Q-value to execute to achieve the optimal policy.
[0118] Step S5: Calculate the cumulative reward according to the action executed by the agent;
[0119] The reward function is designed as a feedback mechanism to provide feedback to the agent on the rationality of its behavior in a given environment or the quality of the algorithm policy. For this task model, the reward function framework is formalized as a multi-objective optimization framework. Specifically, the reward function R t is defined as a composite function including four components: immediate reward R 1. Step penalty R 2. Termination reward R 3 and boundary penalty R 4.
[0120] (1) Immediate reward R 1: Evaluate the immediate feedback received by the agent after executing an action, and encourage the agent to select actions that can increase the voltage of the receiving coil.
[0121] Immediate reward R 1 is proportional to the magnitude of the state variable, especially the pick-up voltage of the receiving coil. A higher pick-up voltage results in a greater direct return, guiding the agent to select an energy transfer unit that gradually approaches the optimal receiver position.
[0122] (2) Step penalty R 2 reflects the system cost incurred by activating each energy transfer unit, penalizes redundant actions, and encourages reaching the target via the shortest path.
[0123] Step penalty R 2 encourages the agent to determine the optimal energy transfer unit with the fewest activation steps, thereby minimizing the system's activation cost. In the map of the algorithm grid, this is equivalent to reaching the destination via the shortest path. Therefore, the penalty increases proportionally with the number of activation steps.
[0124] (3) Termination reward R 3: When the optimal matching state is reached, i.e., when the agent reaches the position of the optimal energy transfer unit on the algorithm map, a relatively large termination reward is assigned. This signals to the agent that it has reached the destination.
[0125] (4) Boundary penalty R 4 limits the agent's ineffective exploration in low-voltage areas.
[0126] Theoretically, the array emission system can infinitely expand its array scale according to the number of drones and charging requirements, resulting in a potentially infinite search environment map for the model. However, due to the limited spatial range of the magnetic field in near-field coupling, energy transfer is effectively achieved only when the energy transfer unit near the excitation receiving coil is activated. Therefore, for the agent, it is unnecessary and wasteful to explore and activate energy transfer units outside the area where the pick-up voltage is almost zero. To solve this problem, a boundary penalty is imposed on the agent, restricting the model's search environment map to the area where the pick-up voltage is almost zero and preventing excessive exploration beyond this area.
[0127] The above reward settings are designed to ensure fast model convergence and generate reasonable algorithmic strategies. Therefore, the reward function cumulative reward R t is calculated as follows:
[0128] ;
[0129] where step is the number of actions performed by the agent.
[0130] Meanwhile, after model training, the specific parameters of the reward function in the algorithm model are finally set as:
[0131] ;
[0132] Step S6: Evaluate the action strategy output by the agent according to the evaluation network, and determine whether the model training strategy meets the initialization requirements. If yes, execute Step S7; if not, execute Step S4.
[0133] The evaluation network refers to calculating the ratio m :
[0134] m = k / d ;
[0135] Where k is the number of actions executed by the agent, d is the Manhattan distance between the initially randomly activated energy transfer unit and the optimal energy transfer unit of the agent.
[0136] When m approaches 1, the initialization requirements are met, and Step S7 is executed; otherwise, Step S4 is executed.
[0137] Given the diversity in the reconstruction types and sizes of the array launch platforms, to ensure that this method can be generally applicable to array multi-coil systems with various reconstruction configurations, grid map points are used to represent individual energy emission units. In this process, the distance between grid points is not directly based on the physical spacing of the actual multi-coils, but a normalized design strategy is adopted, such that the distance between any adjacent grids is uniformly set to 1. Additionally, for a specific landing scenario during each training process, the agent starts exploring from a randomly selected initial grid point and, through a continuous exploration process, finally locates the optimal energy transfer unit, thereby achieving the generalization ability of the training. Also, considering that the agent starts each training from a random point as the initial exploration position, therefore, simply counting the number of actions executed by the agent, i.e., the system activation cost, cannot comprehensively evaluate the effectiveness of its strategy. To address the complexity of balancing the exploration difficulty and the strategy effectiveness, the proposed algorithm model first defines the Manhattan distance between two points ( X 1, Y 1) and ( X 2, Y 2) as: d t Defined as:
[0138] ;
[0139] When mWhen the value asymptotically approaches 1 during model training, it indicates that starting from a random initial point on the current grid map, the agent has achieved an action count equal to the Manhattan distance between two points. This observation directly verifies the achievement of the optimization objective, confirming that the agent has completed the search task through the shortest path, that is, achieving the optimal matching of the coil.
[0140] This metric of the evaluation network proposed by the present invention is used to monitor the training process and evaluate the convergence state of the algorithm. When observed during the training process m When the value gradually approaches 1 with a small fluctuation amplitude, it indicates that the agent has been able to effectively locate the optimal energy transmission unit through the shortest path when randomly starting a certain transmitting unit, achieve the precise matching of the coil, and minimize the system startup cost. At this time, it can be considered that the training of the model has reached the convergence state and can make intelligent decisions based on the current state information. Therefore, at this stage, the algorithm strategy can be directly extracted, and the Q-value table generated during the learning process can be stored for subsequent direct query and use. At the same time, the position coordinates of the array units currently activated by the system can be output, and based on this position, the position of the receiving coil of the power receiving device can be located, thus completing the coil positioning and coil matching. On the contrary, if m the value does not meet this convergence criterion, it indicates that the agent is still in the stage of disorderly exploration and learning and needs to perform more actions to find the optimal energy transmission unit. Therefore, the agent will continue to explore the environment.
[0141] Step S7: Extract the Q-value table and output the coil positioning and matching strategy, that is, the optimal energy transmission unit and the shortest startup path.
[0142] This application proposes a wireless charging multi-coil fast matching and positioning method based on Q-learning the algorithm, which adaptively learns the dynamically changing environment, activates the optimal energy transmission unit with the minimum system cost, that is, the shortest path, greatly reducing the number of startups compared with traversing and searching, and as the scale of the array layout increases, the matching efficiency will be further improved.
[0143] At the hardware level, the multi-coil array transmitting platform with a reconfigurable array configuration cooperates with a flexible power supply network, can freely activate the required coils according to the algorithm, provides a hardware foundation for achieving efficient charging, and its coil configuration and power supply network design are flexible, suitable for any reconfigurable array multi-coil matching strategy.
[0144] From an algorithmic perspective, the action selection strategy is optimized through various innovative mechanisms without the need for complex and cumbersome calculations and theoretical derivations. The dynamic exploration rate adjustment mechanism adjusts the exploration rate according to the state access density and the sensitivity of the induction gradient, balancing exploration and exploitation, enabling the model to converge quickly and effectively avoiding getting stuck in local optimal solutions. The confidence-weighted action selection mechanism quantifies the uncertainty of action values, encourages exploration of actions with low access counts, optimizes the cold start process, and improves the intelligence of action selection. The gradient perception enhancement mechanism incorporates physical characteristics into the model, guiding the agent to explore along the direction of increasing induction, accelerating local convergence. At the same time, the reward function design provides reasonable feedback to the agent, and the multi-objective optimization framework comprehensively considers aspects such as immediate reward, step penalty, termination reward, and boundary penalty to ensure that the model converges quickly and generates reasonable strategies. In addition, this application uses an evaluation network to evaluate the agent's action strategy by ratio, effectively monitoring the training process and the algorithm convergence state, ensuring the effectiveness of model training.
[0145] In summary, the present invention autonomously learns the charging environment information collected by the algorithm, conducts intelligent derivation, reduces the cumbersome nature of theoretical calculations, and improves the decision-making efficiency and accuracy. This transformation from data-driven to intelligent decision-making will provide a more flexible, efficient, and reliable coil positioning and matching strategy for wireless power transmission, so as to activate the optimal energy transmission unit with the shortest path and then achieve efficient and stable wireless power transmission. Specific embodiments
[0147] The Q-value table records the expected return, i.e., the Q-value, of performing different actions (i.e., activating the next energy transmission unit in any of the up, down, left, or right directions) in different system states in the form of state-action pairs. For the selection of the induction quantity, the current receiving coil pickup voltage is defined as the state detection quantity as:
[0148] ;
[0149] where S t is the state of the agent's pickup voltage at the current time t; U t represents the pickup voltage of the receiving coil at the current time t, and different subscripts represent the magnitudes of the pickup voltages of the receiving coil in different states.
[0150] In addition, the actions that the agent can select each time are defined as activating the energy transmission units adjacent to the current three-phase array power transmission unit in the up, down, left, and right directions, so the actions are: .
[0151] By continuously exploring the voltage distribution under the random landing positions and postures of the drone, after the algorithm model converges, it can quickly and accurately determine the startup sequence of the three-phase units in the array according to the current receiving coil pickup voltage. Taking one case as an example,Figure 2 It shows the voltage distribution explored by the model, where different colors represent different intervals of voltage values. The Q-value table is extracted. According to the output strategy, the intelligent agent can randomly turn on a certain energy transmission unit, and output the specified coil turning-on path and direction based on the voltage state picked up by the current receiving coil. Finally, the optimal output position coordinates of the energy transmitting unit are found.
[0152] Please refer to Figure 3 , in the 8x8 state space configuration, the traversal search method shown in the figure is adopted. The arrow represents the execution direction, and 64 switching operations need to be performed to find the optimal energy transmission unit (the asterisk position).
[0153] In contrast, please refer to Figure 4 , the coil matching and positioning method proposed by the present invention significantly optimizes this process, more intelligently selects the execution direction (arrow direction), and the right side of the figure shows the relationship between the picked-up voltage and the color. In the action selection strategy, in the initial stage of the agent's exploration (the total number of visits to state s ( C ( s ) < 5): Try different actions with an exploration rate of 80%, and at the same time, the gradient guidance makes it preferentially move in the direction of increasing voltage. In the middle stage ( C ( s ) ≈ 20): The exploration rate drops to about 50%, and the confidence-weighted strategy focuses on high-Q-value actions, but still retains a certain exploration ability. In the later stage ( C ( s ) > 100): The exploration rate approaches 0, completely relying on high-confidence actions aligned with the gradient to stably maintain the maximum voltage. It only needs to perform 2 searches at least and no more than 7 switching operations at most, and the average number of system switchings is reduced to 5 times. Figure 5 It shows the change process of the algorithm convergence evaluation index m of the algorithm model of the present invention during the training process. After the proposed algorithm model undergoes training iterations, it can stably converge to around 1, that is, it can be considered that the algorithm model already has the ability to reach the optimal energy transmission unit along the shortest path and complete the coil matching process.
[0154] The method proposed by the present invention has achieved a 92.2% improvement in efficiency compared with the traditional traversal search, greatly accelerating the process of positioning and coil matching. In addition, as the state space is further expanded, the efficiency of this method shows a continuous increasing trend, effectively reducing the overall cost of system switching operations.
[0155] It can be seen that the method proposed by the present invention not only realizes the high efficiency of energy transmission, but also brings significant cost benefits to the system by reducing the operation frequency, demonstrating its superior performance in the field of energy transmission.
[0156] In addition, in Figure 6Under the array emission platform composed of 19 coils shown in the figure, 5 UAV positions were randomly selected, namely the five positions P1, P2, P3, P4, and P5 in the figure, for generalization verification. The experimental results are shown in the following table. It can be seen from the experimental results that if traversal is used to achieve coil matching, 24 energy transmission units need to be traversed at these five positions to find the optimal energy transmission unit. In the method of the present invention, the number of coil actions is much smaller than the number of times of searching for the optimal energy transmission unit by traversing the opened coils. At the same time, the algorithm model output by the model can already intelligently output the energy transmission unit according to the current state, that is, the finally matched unit is obtained accordingly. At the same time, the position estimation error is less than 4.2 cm. It should be noted that for the measurement of the relative error size, the ultimate goal of the method of the present invention is the optimal matching of the array coils, that is, the currently opened energy transmission unit is optimally matched with the UAV receiving coil to achieve the function of efficient energy transmission. Therefore, the model can finally output the central coordinates of the optimal energy transmission unit finally opened, rather than output the actual position of the UAV receiving coil, but still has a certain position estimation ability, and the position estimation error is not greater than 4.2 cm. Therefore, it can be considered that the proposed model takes into account the functions of UAV positioning and coil matching.
[0157] Table 1 Coil Matching Results
[0158]
[0159] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in the present invention can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but this implementation should not be considered to exceed the scope of this application.
[0160] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of each embodiment of this application, and should all be included in the protection scope of this application.
Claims
1. A method based on Q-learning The wireless charging multi-coil fast matching and positioning method based on the algorithm is characterized by: The following steps are involved: Step S1: Establish a multi-coil array transmitting platform and a power supply network that enables flexible switching of corresponding energy transmission units of the platform; Step S2: changing the position of the receiving coil and the opening status of different energy transmission units, obtaining the induction quantity, and establishing the charging environment information; Step S3: Establish an exploration environment under the algorithm model, model a discretized two-dimensional grid map, each grid corresponds to the energy transmission unit in the array launch platform, define the state S of the grid as the corresponding induction quantity, and initialize the Q value of all grids; Step S4: The receiving coil induction changes and enters the coil matching process. The intelligent agent explores in the unknown environment, performs actions according to the action selection strategy, and iteratively updates the Q value table; the action selection strategy includes a dynamic exploration rate adjustment mechanism, a confidence-weighted action selection mechanism, and a gradient perception enhancement mechanism; Step S5: Calculate the cumulative reward based on the actions performed by the agent; Step S6: Evaluate the action strategy output by the agent based on the evaluation network to determine whether the model training strategy meets the initialization requirements. If yes, execute step S7; If no, proceed to step S4; Step S7: extract the Q value table to complete the coil matching process, and output the optimal energy transfer unit and the shortest opening path; The dynamic exploration rate adjustment mechanism dynamically adjusts the exploration rate through the state access density and the sensitivity of the sensing gradient; The confidence-weighted action selection mechanism quantifies the uncertainty of action value through confidence indicators, guiding the exploration of actions with low access counts; The gradient perception enhancement mechanism uses the induction quantity in the wireless charging environment to encode the field strength distribution information into the action selection, calculates the induction quantity gradient vector field and the direction preference weight, superimposes the direction preference weight on the basis of the confidence-weighted Q value, and selects an action that is both high-value and consistent with the gradient direction.
2. The method according to claim 1 Q-learning The wireless charging multi-coil fast matching and positioning method based on the algorithm is characterized by: The dynamic exploration rate adjustment mechanism is implemented as follows: (1) Establishing the exploration rate decay function : ; in, N (s,a) is the cumulative number of executions of action a in state s, reflecting the familiarity of the agent’s actions; C (s) is the total number of visits to state s, which indicates the degree of exploration of the state by the current agent; β is the decay rate parameter, which is used to control the speed at which the current exploration rate decreases; The curvature parameter is used to adjust the nonlinearity of the attenuation curve; ε 0 is the initial exploration rate, i.e., the maximum exploration probability; (2) Establish a state access density threshold to force exploration of low-access areas: ; The above constraints force the agent to maintain an exploration rate of more than 30% when the agent appears in a new state, that is, the number of visits is less than 5; (3) Introduce a sensitivity trigger mechanism for the induction gradient, which temporarily increases the agent's exploration rate when the induction is close to the maximum value, that is, when the induction gradient change is less than 10%.
3. The method according to claim 1 Q-learning The wireless charging multi-coil fast matching and positioning method based on the algorithm is characterized by: The confidence-weighted action selection mechanism is specifically implemented as follows: (1) Calculate action confidence index Q c (s,a) , the formula is as follows: ; in, Q(s,a) Representing tradition Q-learning The action-value function under the algorithm; κ is the confidence coefficient, which can control the weight of exploration and utilization. The larger the value, the more inclined to exploration; C(s) is the total number of visits to state s; N (s,a) is the cumulative number of executions of action a in state s; j is a zero-proof constant; and the degree of underexploration of the action is represented by the uncertainty term √(ln(1+C(s)) / N(s,a)); (2) Actions with high confidence are given higher selection probabilities, and the action probability distribution of the agent in state s satisfies: ; in, P(a ∣ s) The agent is in state s When you select Action a probability; Q c (s,a) is the action confidence index; a ′ are all the actions that the agent can perform in state s; Q c (s,a ′ ) To perform actions a ′ Action confidence indicators in the case; Status s Related temperature parameter dynamic adjustment factors; Introducing dynamic adjustment factors for temperature parameters ,Through the degree of certainty of the temperature parameter adaptive adjustment strategy, the confidence index of the modified Q value is converted into a probability distribution, and the temperature adjustment factor changes the probability distribution from flat to sharp as the agent exploration process proceeds, the agent eventually tends to a purely greedy state, and the current optimal ; in, is the dynamic adjustment factor of the temperature parameter related to state s; C (s) is the total number of visits to state s, which indicates the degree of exploration of the state by the current agent; η is a hyperparameter.
4. The method according to claim 1 Q-learning The wireless charging multi-coil fast matching and positioning method based on the algorithm is characterized by: The iterative updating of the Q value table refers to iteratively updating the Q value table using the Bellman equation: ; in, Q t , Q t+1 They are the Q values of the agent at time t and the next time t+1 respectively; Q t+1 (s t ,a t ) is the agent at time t+1 , In state s t , Execute actions a t When , the updated Q value indicates the expected value return of the state-action pair under the current strategy; s t , s t+1 are the states of the agent at time t and the next time t+1 respectively; A is the set of all actions, that is, A includes a, and a has four actions: up, down, left, and right; a t is the action performed by the agent at time t; R t+1 (s t , a t ) The agent at time t, from state s t Execute an action a t Transition to next state s t+1 After that, the cumulative reward obtained at time t+1; α is the learning rate; γ is the discount factor.
5. The method according to claim 1 Q-learning The wireless charging multi-coil fast matching and positioning method based on the algorithm is characterized by: The cumulative reward R t The calculation formula is: ; in, R 1 is an immediate reward, R 2 is the step penalty, R 3 is the termination reward, R 4 is the boundary penalty, and step is the number of times the agent performs an action.
6. The method according to claim 1 Q-learning The wireless charging multi-coil fast matching and positioning method based on the algorithm is characterized by: The evaluation network refers to the calculation ratio m : m = k / d ; in, k is the number of actions performed by the agent, d The Manhattan distance between the initial energy transfer unit randomly opened by the agent and the optimal energy transfer unit; when m When it approaches 1, the initialization requirement is met.
Citation Information
Patent Citations
Optimal decision-making method based on improved Q-learning
CN112598137A
Multi-agent target collaborative search method and system
CN115952736A