Wireless charging positioning method, device and equipment and storage medium

By using multi-sensor data fusion and a deep reinforcement learning decision model, the accuracy problem of wireless charging positioning in complex environments has been solved, achieving an efficient and stable charging process that can adapt to various scenario requirements.

CN121173013APending Publication Date: 2025-12-19CHONGQING CHANGAN AUTOMOBILE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511328428.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-17
Publication Date
2025-12-19

AI Technical Summary

Technical Problem

Existing wireless charging positioning methods suffer from significant positioning errors in complex environments, especially in scenarios involving metal interference, magnetic core differences, and high temperatures. Traditional UWB and pure magnetic coupling schemes struggle to achieve high-precision alignment, resulting in low charging efficiency.

Method used

Employing multi-sensor data fusion technology, combined with attention mechanisms and deep reinforcement learning decision models, data is collected through UWB base station arrays, magnetic induction coil arrays, inertial measurement units, and temperature sensors. Kalman filtering correction and normalization are performed, attention weights are dynamically adjusted, and actor critic networks are used to optimize charging locations, achieving multi-objective optimization.

Benefits of technology

It improves the alignment accuracy of the wireless charging coil, reduces charging efficiency loss, enhances the versatility and robustness of the system, adapts to various wireless charging scenarios, and ensures a stable and efficient charging process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121173013A_ABST
    Figure CN121173013A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of wireless charging, and discloses a wireless charging positioning method, device and equipment and a storage medium, and the method comprises the steps: collecting current sensing data corresponding to a plurality of sensors which are used for collecting the information of a wireless charging coil; fusing the current sensing data to obtain an initial state space; determining an attention weight corresponding to each sensor; calibrating the attention weight to the initial state space to obtain a target state space; inputting the target state space into a decision-making model based on deep reinforcement learning, and outputting a current action space through the decision-making model; and adjusting the position of the to-be-charged equipment according to the current action space. According to the invention, the accuracy of wireless charging positioning based on multi-sensor data is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of wireless charging technology, specifically to a wireless charging positioning method, device, equipment, and storage medium. Background Technology

[0002] Wireless charging technology transmits electrical energy through electromagnetic induction or magnetic resonance, but its transmission efficiency is highly dependent on the alignment accuracy of the transmitting and receiving coils. Traditional positioning methods mainly include UWB (Ultra Wide Band) positioning and pure magnetic coupling schemes.

[0003] UWB positioning suffers from multipath effects and electromagnetic noise in metallic environments, causing positioning errors to increase to over 20cm. Magnetic coupling positioning relies on sensing the strength and direction of the magnetic field to determine its relative position to the transmitting coil. It doesn't depend on external signals and is unaffected by metallic environments. However, pure magnetic coupling can only determine relative position and cannot provide absolute coordinates like UWB positioning, resulting in lower positioning accuracy. Some technologies consider combining UWB positioning and pure magnetic coupling to improve wireless charging positioning accuracy. However, research has shown that interference factors in the environment are often dynamically changing. Sometimes UWB positioning is more accurate, and sometimes pure magnetic coupling is more accurate. Current multi-sensor data fusion solutions still suffer from some positioning errors due to poor fusion algorithms. Summary of the Invention

[0004] In view of this, the present invention provides a wireless charging positioning method, apparatus, device and storage medium to solve the problem of large errors in wireless charging positioning based on multi-sensor data.

[0005] In a first aspect, the present invention provides a wireless charging positioning method, the method comprising: collecting current sensing data corresponding to multiple sensors, the sensors being used to collect information of a wireless charging coil; fusing the current sensing data to obtain an initial state space; determining an attention weight corresponding to each sensor; calibrating the attention weight to the initial state space to obtain a target state space; inputting the target state space into a decision model based on deep reinforcement learning to output a current action space through the decision model; and adjusting the position of the device to be charged according to the current action space.

[0006] Based on the aforementioned technical methods, wireless charging positioning is achieved through multi-sensor data acquisition and fusion, combined with attention mechanisms and deep reinforcement learning decision models. This allows for dynamic adaptation to the reliability of sensors in different scenarios, avoiding the positioning limitations of a single sensor in complex environments (such as metal interference or magnetic core differences), and accurately outputting device position adjustment commands. It effectively improves the alignment accuracy of the wireless charging coil, reduces charging efficiency losses due to positioning deviations, and adapts to various wireless charging scenarios, enhancing the versatility and robustness of the positioning system and ensuring a stable and efficient charging process.

[0007] In some optional implementations, the acquisition of current sensing data corresponding to multiple sensors includes: acquiring first sensing data from a UWB base station array, wherein the UWB base station array is disposed at the transmitting or receiving end of the wireless charging coil for UWB positioning of the wireless charging coil; acquiring second sensing data from a magnetic induction coil array, wherein the magnetic induction coil array is disposed at the transmitting or receiving end of the wireless charging coil for magnetic coupling positioning of the wireless charging coil; acquiring third sensing data from an inertial measurement unit, wherein the inertial measurement unit is disposed at the device to be charged for measuring the pose of the device to be charged; and acquiring fourth sensing data from a temperature sensor, wherein the temperature sensor is disposed at the transmitting end of the wireless charging coil for monitoring the operating temperature of the magnetic core in the transmitting end of the wireless charging coil.

[0008] Based on the aforementioned technical means, four types of sensors are provided to achieve full coverage of multi-dimensional sensing data. The UWB base station array provides absolute coordinates, the magnetic induction coil array captures magnetic field information, the inertial measurement unit monitors the device's pose to determine its positional deviation, and the temperature sensor controls the magnetic core temperature to determine the degree of length measurement deviation. This multi-source data complementarity avoids the shortcomings of single sensors and provides comprehensive, high-quality data support for subsequent data fusion and precise positioning, further improving positioning accuracy and system anti-interference capabilities.

[0009] In some optional implementations, fusing the current sensing data to obtain an initial state space includes: correcting the current sensing data corresponding to each sensor using Kalman filtering; normalizing the corrected current sensing data; and concatenating the normalized current sensing data into a vector to obtain the initial state space.

[0010] Based on the aforementioned technical methods, Kalman filtering is used to correct the perceived data, filtering out noise interference and ensuring data authenticity and stability. Normalization is then used to eliminate differences in data dimensions, preventing any type of data from dominating the fusion process due to its numerical scale. This results in a more accurate and regular initial state space, laying a reliable foundation for subsequent attention weight calculations and decision model inputs, reducing the impact of data errors on the positioning results, indirectly improving positioning accuracy, and ensuring the efficient progress of subsequent positioning steps.

[0011] In some optional implementations, determining the attention weight for each sensor includes: obtaining an attention extraction matrix; calculating the query vector and key vector corresponding to each sensor data based on the attention extraction matrix and the initial state space; calculating vector similarity using the query vector and key vector corresponding to each sensor data; scaling and normalizing the vector similarity of each sensor data; and assigning weights based on the processed vector similarity of each sensor data to obtain the attention weight for each sensor.

[0012] Based on the aforementioned technical methods, a query vector and a key vector are generated by multiplying a learnable attention extraction matrix with the initial state space. Vector similarity is then calculated and processed to assign weights. This allows for the dynamic identification of the importance of different sensor data in the current scenario; for example, it increases the weight of magnetic coupling data when there is metal interference, and emphasizes horizontal magnetic field data when there are significant differences in magnetic cores. This avoids the obscuring of key information caused by equal-weighted fusion, making the subsequent state space more closely match actual positioning needs, improving the targeting and accuracy of positioning decisions, and enhancing the system's adaptability to complex environments.

[0013] In some optional implementations, the step of training the decision model includes: obtaining a reward function, the reward function including a charging efficiency term, a motion amplitude term, a positioning deviation term, and a temperature term, wherein the charging efficiency term is adjusted by multiplying it by a positive reward coefficient, and the motion amplitude term, positioning deviation term, and temperature term are all adjusted by multiplying them by a negative penalty coefficient; inputting the target state space into the actor network of the decision model, and outputting the current motion space through the actor network; calculating a reward value based on the reward function through the critic network of the decision model; and adjusting the model parameters of the actor network and the critic network according to the reward value.

[0014] Based on the aforementioned technical means, this invention designs a multi-dimensional reward function that balances charging efficiency, motion stability, positioning accuracy, and device safety. A decision-making model is trained collaboratively through an actor-critic network. The actor network outputs a precise motion space, while the critic network rationally evaluates the reward value; the two work together to optimize model parameters. This enables the model to learn a positioning strategy that balances efficiency and safety, reducing invalid actions and positioning errors, avoiding problems such as low charging efficiency and device damage caused by inappropriate strategies, and improving the intelligence and practicality of the positioning system.

[0015] In some optional implementations, the reward function is:

[0016]

[0017] In the formula, r t Represents the reward value, η charge This refers to the charging efficiency term. Indicates the amplitude of the action, Π misalign The positioning deviation term is represented by max(0,TT). core ) represents the temperature term, T represents the current temperature, T core The target temperature is represented by λ1, λ2, and λ3, which are the corresponding penalty coefficients.

[0018] Based on the aforementioned technical means, this invention, through the design of a quantitative and multi-dimensionally coupled reward function formula, overcomes the limitations of traditional positioning algorithms' "single-objective optimization" and constructs a dynamic optimization mechanism integrating "efficiency, stability, and safety." The reward function couples charging efficiency with three types of penalty terms: motion amplitude, positioning deviation, and temperature, through weighted coefficients. This avoids focusing solely on positioning error or charging efficiency as a single objective, and avoids the contradiction of sacrificing device safety for efficiency or reducing charging speed to maintain stability. The reward function achieves a dynamic balance of multiple objectives through adjustable penalty coefficients. Furthermore, the temperature term uses max(0, TT)... core The nonlinear function of the magnetic core loss is precisely matched to its physical characteristics. When the temperature does not exceed the threshold, the penalty is 0, which does not affect normal positioning optimization. Once the temperature exceeds the threshold, the penalty increases synchronously with the temperature difference, forcing the decision model to actively adjust the device position to reduce magnetic coupling loss, thus achieving "preventive thermal control" and ensuring device safety and positioning stability in high-power wireless charging scenarios. The motion amplitude term suppresses large motions, which can guide the model to output smoother position adjustment commands, balancing positioning efficiency and device lifespan.

[0019] In some optional implementations, the method further includes adjusting the matrix parameters of the attention extraction matrix based on the reward value.

[0020] Based on the aforementioned technical means, this invention configures a learnable attention extraction matrix and synchronously adjusts the parameters of the attention extraction matrix according to the reward value, thus linking attention weight allocation with decision model optimization. When the model discovers through the reward value that a certain type of sensor data is more critical for localization, the matrix parameters can be optimized simultaneously, making subsequent weight calculations more accurate. This forms a closed loop of data fusion, decision output, and parameter feedback, continuously improving the effectiveness of the attention mechanism, thereby optimizing the quality of the target state space, ensuring a steady improvement in localization accuracy, and enhancing the system's adaptability and iterative capabilities.

[0021] In some alternative implementations, the vertical and horizontal coils in the wireless charging coil are arranged orthogonally.

[0022] Based on the aforementioned technical methods, the wireless charging coil adopts an orthogonal layout of vertical and horizontal coils, which can utilize the complementary vertical and horizontal magnetic field components to cancel out magnetic field distortion caused by differences in the magnetic core. Even when the properties of the magnetic core material change, the total mutual inductance remains relatively stable, reducing the impact of magnetic field distortion on positioning accuracy. Improving the system's resistance to magnetic core differences at the hardware structure level, combined with algorithm optimization, further reduces positioning errors, ensures stable charging efficiency, and enhances the versatility and reliability of the wireless charging system under different magnetic core scenarios.

[0023] Secondly, the present invention provides a wireless charging positioning device, the device comprising: a data acquisition module for acquiring current sensing data corresponding to multiple sensors, wherein the sensors are used to acquire information of a wireless charging coil; a state space generation module for fusing the current sensing data to obtain an initial state space; an attention weight calculation module for determining the attention weight corresponding to each sensor; a state space adjustment module for calibrating the attention weight to the initial state space to obtain a target state space; a decision module for inputting the target state space into a decision model based on deep reinforcement learning to output a current action space through the decision model; and a position adjustment module for adjusting the position of the device to be charged according to the current action space.

[0024] Thirdly, the present invention provides a computer device, comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the computer instructions to perform the method described in the first aspect or any corresponding embodiment thereof.

[0025] Fourthly, the present invention provides a computer-readable storage medium storing computer instructions for causing a computer to perform the method described in the first aspect or any corresponding embodiment thereof.

[0026] Fifthly, the present invention provides a computer program product, including computer instructions for causing a computer to perform the method described in the first aspect or any corresponding embodiment thereof.

[0027] (1) Based on the above technical means, wireless charging positioning is achieved through multi-sensor data acquisition and fusion, combined with attention mechanism and deep reinforcement learning decision model. It can dynamically adapt to the reliability of sensors in different scenarios, avoid the positioning limitations of a single sensor in complex environments (such as metal interference, magnetic core differences), and accurately output device position adjustment commands. It effectively improves the alignment accuracy of wireless charging coil, reduces charging efficiency loss caused by positioning deviation, and adapts to various wireless charging scenarios, enhancing the versatility and robustness of the positioning system, and ensuring a stable and efficient charging process.

[0028] (2) Based on the above technical means, four types of sensors are provided to achieve full coverage of multi-dimensional sensing data. The UWB base station array provides absolute coordinates, the magnetic induction coil array captures magnetic field information, the inertial measurement unit monitors the device's pose to determine the device's pose deviation, and the temperature sensor controls the magnetic core temperature to determine the degree of deviation in length measurement. The multi-source data complements each other. This not only avoids the shortcomings of a single sensor, but also provides comprehensive and high-quality data support for subsequent data fusion and accurate positioning, further improving positioning accuracy and system anti-interference capability.

[0029] (3) Based on the above technical means, the perceived data is corrected by Kalman filtering to remove noise interference and ensure the authenticity and stability of the data; then, normalization is used to eliminate differences in data dimensions and avoid the fusion process being dominated by the numerical scale of a certain type of data. The resulting initial state space is more accurate and regular, laying a reliable foundation for subsequent attention weight calculation and decision model input, reducing the impact of data errors on the positioning results, indirectly improving positioning accuracy, and ensuring the efficient progress of subsequent positioning steps.

[0030] (4) Based on the above technical means, a query vector and a key vector are generated by multiplying a learnable attention extraction matrix with the initial state space. Vector similarity is calculated and processed to assign weights. This can dynamically identify the importance of different sensor data in the current scenario. For example, when there is metal interference, the weight of magnetic coupling data is increased, and when there are large differences in the magnetic core, the weight of horizontal magnetic field data is emphasized. This avoids the obscuring of key information caused by equal-weighted fusion, making the subsequent state space more consistent with the actual positioning needs, improving the pertinence and accuracy of positioning decisions, and enhancing the system's adaptability to complex environments.

[0031] (5) Based on the above technical means, this invention designs a multi-dimensional reward function that takes into account charging efficiency, motion stability, positioning accuracy, and equipment safety. The decision-making model is trained collaboratively through an actor-critic network. The actor network outputs a precise motion space, while the critic network rationally evaluates the reward value; the two work together to optimize the model parameters. This enables the model to learn a positioning strategy that balances efficiency and safety, reducing invalid actions and positioning deviations, avoiding problems such as low charging efficiency and equipment damage caused by inappropriate strategies, and improving the intelligence and practicality of the positioning system.

[0032] (6) Based on the above technical means, this invention, through the design of a quantitative and multi-dimensionally coupled reward function formula, breaks through the limitations of the traditional positioning algorithm's "single-objective optimization" and constructs a dynamic optimization mechanism integrating "efficiency-stability-safety". The reward function couples the charging efficiency term with three types of penalty terms: motion amplitude, positioning deviation, and temperature, through coefficient weighting. This avoids using positioning error or charging efficiency as a single objective, and avoids the contradiction of sacrificing device safety for efficiency or reducing charging speed to maintain stability. The reward function achieves a dynamic balance of multiple objectives through adjustable penalty coefficients. In addition, the temperature term adopts max(0,TT) coreThe nonlinear function of the magnetic core loss is precisely matched to its physical characteristics. When the temperature does not exceed the threshold, the penalty is 0, which does not affect normal positioning optimization. Once the temperature exceeds the threshold, the penalty increases synchronously with the temperature difference, forcing the decision model to actively adjust the device position to reduce magnetic coupling loss, thus achieving "preventive thermal control" and ensuring device safety and positioning stability in high-power wireless charging scenarios. The motion amplitude term suppresses large motions, which can guide the model to output smoother position adjustment commands, balancing positioning efficiency and device lifespan.

[0033] (7) Based on the above technical means, this invention configures a learnable attention extraction matrix and adjusts the parameters of the attention extraction matrix synchronously according to the reward value, so that the attention weight allocation and decision model optimization are linked. When the model finds that a certain type of sensor data is more critical to positioning through the reward value, the matrix parameters can be optimized synchronously to make the subsequent weight calculation more accurate. A closed loop of data fusion, decision output, and parameter feedback is formed, which continuously improves the effectiveness of the attention mechanism, thereby optimizing the quality of the target state space, ensuring a steady improvement in positioning accuracy, and enhancing the system's adaptability and iterative capability.

[0034] (8) Based on the above technical means, the wireless charging coil adopts an orthogonal layout of vertical and horizontal coils, which can utilize the complementary vertical and horizontal magnetic field components to cancel out the magnetic field distortion caused by differences in the magnetic core. When the properties of the magnetic core material change (such as differences in permeability), the total mutual inductance can still remain relatively stable, reducing the impact of magnetic field distortion on positioning accuracy. By improving the system's resistance to magnetic core differences at the hardware structure level, and in conjunction with algorithm optimization, positioning errors are further reduced, charging efficiency is ensured to be stable, and the versatility and reliability of the wireless charging system under different magnetic core scenarios are enhanced. Attached Figure Description

[0035] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0036] Figure 1 This is a flowchart illustrating a wireless charging positioning method according to an embodiment of the present invention.

[0037] Figure 2 This is a schematic diagram of the structure of a wireless charging positioning device according to an embodiment of the present invention;

[0038] Figure 3 This is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. Detailed Implementation

[0039] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0040] According to embodiments of the present invention, a wireless charging positioning method, apparatus, device, storage medium, and program product embodiment are provided. It should be noted that the steps shown in the flowcharts in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowcharts, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0041] This embodiment provides a wireless charging positioning method. Figure 1 This is a flowchart of a wireless charging positioning method according to an embodiment of the present invention, which includes the following steps:

[0042] Step S101: Collect current sensing data corresponding to multiple sensors, which are used to collect information about the wireless charging coil.

[0043] Step S102: Fuse the current sensing data to obtain the initial state space;

[0044] Step S103: Determine the attention weight corresponding to each sensor;

[0045] Step S104: Label the attention weights to the initial state space to obtain the target state space;

[0046] Step S105: Input the target state space into the decision model based on deep reinforcement learning, so as to output the current action space through the decision model;

[0047] Step S106: Adjust the position of the device to be charged according to the current motion space.

[0048] Specifically, embodiments of the present invention collect relevant information about the wireless charging coil using various sensors. This collected information includes, but is not limited to, location and attribute information. The sensors are dedicated sensors deployed around the wireless charging coil to meet alignment requirements; any sensor capable of collecting the wireless charging coil's location and attribute information is sufficient. In one specific embodiment, a UWB base station array, a magnetic induction coil array, an inertial measurement unit, and a temperature sensor are included, forming a multimodal sensing system encompassing four dimensions: position, magnetic field, attitude, and temperature. The "current sensing data" refers to the dynamic data collected by each sensor in the real-time wireless charging positioning scenario, meeting millisecond-level update frequencies to match the position change speed of the device to be charged (such as new energy vehicles, automated guided vehicles, drones, mobile phones, and any other device supporting wireless charging), avoiding positioning errors caused by data lag. This includes the first sensing data from the UWB base station array, the second sensing data from the magnetic induction coil array, the third sensing data from the inertial measurement unit, and the fourth sensing data from the temperature sensor.

[0049] The core of wireless charging efficiency relies on the precise alignment of the transmitting and receiving coils. A single sensor cannot cope with interference in complex scenarios. UWB base station arrays can provide absolute coordinate information, but are susceptible to multipath effects in metallic environments. Magnetic induction coil arrays can sense magnetic field strength and direction, reflecting the relative position of the coils, but lack absolute position reference. Inertial measurement units can monitor attitude information such as pitch and roll angles of the device being charged through accelerometers and gyroscopes, but suffer from error accumulation. Temperature sensors can capture the operating temperature of the magnetic core at the transmitting end of the wireless charging coil in real time, preventing magnetic field distortion and equipment damage caused by overheating of the magnetic core. Through multi-sensor collaborative acquisition, data redundancy and complementarity can be achieved, providing comprehensive and reliable raw information for subsequent positioning calculations. Taking the wireless charging scenario of automated guided vehicles (AGVs) as an example, in a factory workshop environment, UWB base station arrays (deployed at the transmitting or receiving end of the wireless charging coil) collect three-dimensional coordinate data of the receiving coil on the AGV, while magnetic induction coil arrays (set at the transmitting or receiving end of the wireless charging coil) collect magnetic field strength data to determine the positional deviation between the receiving and transmitting coils based on changes in magnetic field strength. An inertial measurement unit (e.g., mounted on the chassis of an automated guided vehicle) collects attitude data of the automated guided vehicle (such as pitch angle and roll angle), and a temperature sensor (e.g., attached to the surface of the transmitting coil core) collects core temperature data. These four types of data are simultaneously uploaded to the positioning control system to complete the current sensing data acquisition.

[0050] This invention utilizes four types of sensors to achieve comprehensive multi-dimensional sensing data coverage. A UWB base station array provides absolute coordinates, a magnetic induction coil array captures magnetic field information, an inertial measurement unit monitors the device's pose to determine pose deviation, and a temperature sensor controls the magnetic core temperature to determine the degree of length measurement deviation. This multi-source data complementarity avoids the limitations of single sensors and provides comprehensive, high-quality data support for subsequent data fusion and precise positioning, further improving positioning accuracy and system anti-interference capabilities.

[0051] Next, the current sensing data are fused to obtain the initial state space. Data fusion refers to the preprocessing and integration of heterogeneous data collected from multiple sensors using algorithms. The initial state space is the fused set of vectors that comprehensively reflects the current state of the wireless charging positioning system. Its dimensions correspond to the sensor type and data characteristics, and the differences in dimensions and noise interference between data points must be eliminated. The initial state space clearly presents the position of the drone's receiving coil, its surrounding magnetic field environment, its own attitude, and the temperature of the transmitting coil's magnetic core, providing high-quality input for subsequent attention weight calculations.

[0052] Attention weight refers to the weight coefficient assigned to each sensor based on the reliability and importance of each sensor's data in the current positioning scenario, with a value ranging from [0,1]. The sum of the weights of all sensors is 1. Wireless charging positioning scenarios are subject to dynamic interference factors (such as reduced UWB reliability due to metallic environments, fluctuations in magnetic induction data accuracy due to differences in magnetic core materials, and increased importance of temperature data due to high temperatures). If all sensor data is directly superimposed and fused, positioning errors can easily occur due to insufficient weight for key data or excessive weight for redundant data. The attention mechanism dynamically allocates weights by calculating the correlation between sensor data and the positioning target, allowing the system to prioritize sensor data that is more reliable in the current scenario. Taking a car charging scenario near a metal shelf in a factory as an example, the metallic environment causes severe multipath effects in UWB data, resulting in a smaller attention weight; magnetic induction data is less affected by metal, resulting in a higher attention weight; IMU data is used to compensate for car attitude deviations, with a moderate weight; and temperature data currently shows no anomalies, so a smaller weight can be used. In this case, the system prioritizes magnetic induction data for positioning calculations, which meets the scenario requirements.

[0053] Next, the attention weights corresponding to each sensor are weighted and calculated (usually by direct multiplication) with the corresponding feature data in the initial state space to obtain the weighted target state space. The data in the target state space highlights the state vectors contributed by key sensor data. Its dimension is the same as the initial state space, but the importance of each feature data is distinguished by the weights. Taking a wireless charging scenario with excessively high magnetic core temperature as an example, the attention weight of the temperature sensor is increased to 0.3. The proportion of its corresponding feature data in the target state space is significantly increased after weighting, enabling the subsequent decision model to prioritize the thermal state of the magnetic core and avoid positioning errors and device damage caused by high temperature. The implicit importance of temperature data in the initial state space is fully reflected through weight calibration.

[0054] Next, the target state space is input into the deep reinforcement learning decision model. The decision model adopts an Actor-Critic architecture, where the actor network is responsible for outputting specific action instructions based on the input state space, and the critic network is responsible for evaluating the value of the current state and the rationality of the action. The action space refers to the specific set of parameters for adjusting the position of the device to be charged, which typically includes the displacement adjustment amounts on the horizontal X-axis and Y-axis, as well as the attitude angle adjustment amounts. The numerical range is preset according to the device's motion capabilities (e.g., displacement adjustment amount ±5cm, angle adjustment amount ±2°). This embodiment of the invention transforms wireless charging positioning into a sequential decision problem in a dynamic scenario. The device to be charged needs to continuously adjust its position according to the real-time state to achieve precise coil alignment, and the adjustment process must take into account charging efficiency, device stability, and safety (e.g., magnetic core temperature). Deep reinforcement learning guides the model to learn the optimal decision strategy through a "reward function." The actor network outputs actions based on the target state space. After the system executes the action, a new state is generated. The critic network calculates the reward value of the action based on the reward function, and then optimizes the network parameters in reverse through the reward value, so that the model gradually learns the action output logic that "maximizes long-term rewards." Finally, the control system adjusts the position of the device to be charged based on the current motion space output by the decision model.

[0055] This invention addresses the limitations of traditional positioning methods (such as pure UWB and pure magnetic coupling) in complex scenarios by employing multi-sensor collaborative perception, data fusion, dynamic attention weighting, deep reinforcement learning decision-making, and closed-loop position adjustment. From the perception layer perspective, the combination of four types of sensors covers position, magnetic field, attitude, and temperature dimensions, avoiding positioning failures caused by interference from a single sensor. The introduction of the attention mechanism enables the system to dynamically adapt to scene changes (such as metallic environments, magnetic core differences, and high temperatures), prioritizing reliable sensor data. The deep reinforcement learning decision model achieves multi-objective optimization of "efficiency-stability-safety" through a reward function, avoiding the contradiction of sacrificing device safety for efficiency or reducing accuracy for stability. Closed-loop position adjustment ensures continuous coil alignment during charging, controlling charging efficiency fluctuations. In practical applications, the method provided by this invention is adaptable to various devices such as guide vehicles, drones, and consumer electronics. Even in complex scenarios such as metallic interference, magnetic core material differences, and high temperatures, it can still control coil alignment deviation within centimeter levels, significantly improving charging efficiency. It exhibits good scene adaptability and robustness, reducing the deployment and maintenance costs of wireless charging systems.

[0056] In some optional implementations, step S102 above includes:

[0057] Step a1: Correct the current sensing data corresponding to each sensor by using Kalman filtering;

[0058] Step a2: Normalize the corrected current perception data, and concatenate the normalized current perception data into a vector to obtain the initial state space.

[0059] Specifically, Kalman filtering is a filtering algorithm based on the state equation of a linear system, recursively estimating the system state through observation data. Its core lies in filtering out random interference noise in sensor-acquired data through two iterative "prediction-update" stages, combined with the statistical characteristics of the data (such as process noise covariance and observation noise covariance), ensuring the data is closer to the true value. Current sensing data acquired by multiple sensors is subject to environmental interference (such as coordinate fluctuations in UWB data due to metal reflection, attitude errors in inertial measurement unit data due to equipment vibration, and intensity deviations in magnetic induction data due to magnetic field distortion). If directly used for subsequent fusion, this noise will be introduced into the positioning calculation, leading to a decrease in the accuracy of the initial state space. Kalman filtering, however, can dynamically adjust the confidence level between the "predicted value" and the "observed value" by utilizing the historical variation patterns of sensor data and the current observation value, thereby outputting more stable and accurate corrected data, adapting to the dual requirements of real-time data and accuracy for wireless charging positioning. In practical implementation, corresponding Kalman filter parameters need to be configured for different sensors first, and then the correction is performed according to the Kalman filtering process. Taking UWB location data as an example, if the corrected position at the previous moment was X = 5.21m, the predicted position at the current moment according to the device motion model is X_pred = 5.211m. Simultaneously, the current UWB observation value X_obs = 5.205m is obtained. Using the Kalman gain formula and correction formula, the corrected UWB location data X_corr = 5.208m is obtained. Compared to the original observation value of 5.205m, this effectively filters out fluctuation noise of ±0.015m. The specific calculation process of the Kalman filtering algorithm is existing technology and will not be elaborated in this embodiment.

[0060] Next, in this embodiment, the corrected current sensing data is normalized, and the normalized current sensing data are concatenated into a vector to obtain the initial state space. The preprocessed initial state space can be represented as s. t =[p uwb B mag ,θ imu ,T core ], where p uwb B represents the first sensing data after preprocessing. mag θ represents the preprocessed second sensing data. imu T represents the preprocessed third-sensory data. core This represents the preprocessed fourth-sensor data. After fusion, the localization update frequency can reach 200Hz, and the response time is shortened to 20ms.

[0061] "Normalization" refers to the preprocessing operation of mapping corrected data with different dimensions (such as meters, Teslas, degrees, and Celsius) and different numerical ranges to a unified interval (in this embodiment, the interval [0,1] is used as an example only and is not a limitation). The purpose is to eliminate numerical imbalances caused by differences in dimensions, preventing one type of data from dominating subsequent calculations due to its larger value, while another type of data is ignored due to its smaller value. The initial state space is a vector formed by concatenating all normalized sensor data in a preset order. Its dimension is equal to the sum of the number of features in all sensor data, and it can completely and regularly reflect the current state of the wireless charging positioning system. Normalization compresses all data into the [0,1] interval through linear transformation, ensuring that all types of data have equal "feature contribution potential" in subsequent calculations, while also accelerating the parameter convergence speed of the deep reinforcement learning decision model.

[0062] This invention optimizes multi-sensor data fusion by combining Kalman filtering and normalization, addressing both noise reduction and dimensional uniformity. Kalman filtering, with its customized parameters for different sensor noise characteristics, more accurately preserves dynamic data features compared to traditional mean filtering, while filtering out random interference, reducing the error rate of the corrected data by over 40%. Normalization, through explicit numerical mapping rules, avoids fusion bias caused by differences in data dimensions. Compared to unnormalized fusion schemes, it significantly improves the accuracy of subsequent attention weight calculations and accelerates the convergence speed of deep reinforcement learning decision models. In complex scenarios with varying magnetic core materials (such as a mixture of ferrite and nanocrystalline cores), the initial state space processed in this way more clearly reflects the impact of magnetic field distortion on magnetic induction data, providing high-quality feature input for the subsequent dynamic weight adjustment of the attention mechanism. Ultimately, this helps maintain high precision in wireless charging coil alignment and steadily improves charging efficiency, significantly outperforming traditional data fusion schemes.

[0063] In some optional implementations, step S103 above includes:

[0064] Step b1: Obtain the attention extraction matrix;

[0065] Step b2: Calculate the query vector and key vector corresponding to each sensor data based on the attention extraction matrix and the initial state space;

[0066] Step b3: Calculate vector similarity using the query vector and key vector corresponding to each sensor data.

[0067] Step b4: Scaling and normalizing the vector similarity of each sensor data.

[0068] Step b5 involves assigning weights based on the vector similarity of the processed data from each sensor, thus obtaining the attention weights for each sensor.

[0069] Specifically, to address the core variations and complex environments, the system should not equally trust all sensors. For example, when an excessively high core temperature or a high-permeability nanocrystalline material is detected, the algorithm should automatically reduce the weight of the vertical component of the magnetic field intensity (as it is susceptible to permeability variations) and increase the weight of the temperature sensor and the horizontal component of the magnetic field (as they are more stable). This attention mechanism enables this dynamic, adaptive sensor reliability assessment and selection.

[0070] The attention extraction matrix is ​​a trainable parameter matrix that is pre-trained using historical positioning data. Its core function is to explore the correlation between the initial state space and the wireless charging positioning target (coil alignment accuracy). Its dimensions need to match the feature dimensions of the initial state space and the number of sensors.

[0071] In this embodiment of the invention, the attention extraction matrix includes a query matrix and a key matrix. The query vector represents the core objective of the current positioning task (e.g., "prioritizing coil alignment accuracy while ensuring core temperature safety"), the key vector represents the core feature identifiers of each sensor data (e.g., UWB coordinate features, magnetic field strength features of the magnetic induction coil), and the value vector is the specific data value in the initial state space. The query vector is calculated by matrix multiplication of the query matrix and the initial state space vector, and the key vector is calculated by matrix multiplication of the key matrix and the initial state space vector. The query vector and key vector need to be split according to the sensor type to ensure that each sensor data has its own dedicated query vector and key vector.

[0072] Next, vector similarity is calculated using the query vector and key vector corresponding to each sensor data. Vector similarity is used to quantify the degree of matching between the features of each sensor data and the current positioning target. This embodiment of the invention uses the dot product similarity calculation method, which can efficiently reflect the linear correlation of vectors of the same dimension. Then, the vector similarity of each sensor data is scaled and normalized. According to the disclosure document, the core purpose of scaling is to avoid excessively high key vector dimensions leading to excessively large similarity values, which could cause the gradient vanishing in the subsequent Softmax function and affect the accuracy of weight calculation. Normalization uses the Softmax function to convert the scaled similarity into a probability distribution with a sum of 1, ensuring the rationality of weight allocation. Attention weights directly reflect the importance of sensor data in the current positioning scenario and need to be allocated based on the normalized similarity results. Simultaneously, fine-tuning is performed based on scene characteristics (such as metal interference and magnetic core temperature) to ensure that the weights truly reflect the reliability of the sensor data. In this embodiment, initial base weights are first assigned to each sensor. Then, during the positioning process, the attention weights are continuously fine-tuned according to the reward function of the decision model. After fine-tuning, the sum of all sensor weights remains 1, ultimately yielding the attention weights for each sensor.

[0073] The calculation formula is as follows:

[0074]

[0075] In the formula, α i q represents the attention weight corresponding to the i-th sensor. T Wh i q represents the calculated vector similarity. T h represents the query vector. i Let represent the key vector of the i-th sensor, and j represent the data from a total of j sensors.

[0076] For example, in a scenario application, in a high-temperature scenario (core temperature > 80°C, this is just an example and not a limitation), the weight of the temperature sensor and the magnetic induction coil is increased, and the weight of UWB is decreased; in a scenario with metal interference, the weight of magnetic coupling data is increased to 70%; when there are significant differences in the core (such as nanocrystals vs. ferrites), the horizontal magnetic field component data is weighted first (because it is less affected by changes in permeability).

[0077] This invention achieves accurate calculation of attention weights through a complete process of matrix acquisition, vector calculation, similarity matching, scaling and normalization, and weight allocation. Compared with traditional fixed-weight fusion schemes, the attention extraction matrix is ​​trained based on historical scene data, which can adapt to complex working conditions such as magnetic core differences and metal interference; secondly, vector similarity calculation and scaling and normalization processing ensure the quantitative rationality of weight allocation; in addition, the fine-tuning process combined with scene characteristics further improves the matching degree between weights and actual needs.

[0078] In some alternative implementations, the step of training the decision model includes:

[0079] Step c1: Obtain the reward function, which includes a charging efficiency term, a motion amplitude term, a positioning deviation term, and a temperature term. The charging efficiency term is adjusted by multiplying it by a positive reward coefficient, while the motion amplitude term, positioning deviation term, and temperature term are all adjusted by multiplying them by a negative penalty coefficient.

[0080] Step c2: Input the target state space into the actor network of the decision model, and output the current action space through the actor network;

[0081] Step c3: Calculate the reward value based on the reward function using the critic network of the decision model;

[0082] Step c4: Adjust the model parameters of the actor network and the critic network based on the reward value.

[0083] Specifically, in reinforcement learning, configuring the reward function is the most crucial step in guiding the improvement of the decision model's learning performance. In this embodiment of the invention, the charging efficiency term is adjusted by multiplying it by the positive reward coefficient, while the action amplitude term, positioning deviation term, and temperature term are all adjusted by multiplying them by the negative penalty coefficient.

[0084] The charging efficiency refers to the real-time power transfer efficiency of the wireless charging system. This is calculated by collecting data from a voltage sensor at the output of the transmitting coil and a current sensor at the input of the receiving coil, and then using the ratio of the power of the receiving coil to the power of the transmitting coil. The positive reward coefficient is set to 1 by default (no additional multiplication is needed; the efficiency percentage is directly included in the reward). The motion amplitude refers to the sum of the squares of the position adjustments of the device being charged, including X-axis and Y-axis displacement adjustments (in centimeters) and attitude angle adjustments (in degrees). The positioning deviation refers to the center alignment deviation between the wireless charging transmitting and receiving coils, which can be obtained by calculating the Euclidean distance using the center coordinates of the two coils collected by the UWB base station array. The temperature refers to the difference between the current core temperature and the safety threshold (a penalty is only applied when the temperature exceeds the safety threshold). In practice, the positioning control system of the device being charged calls the preset reward function formula and parameters from the local algorithm library to prepare for subsequent reward value calculations.

[0085] In one specific implementation, the reward function is:

[0086]

[0087] In the formula, r t Represents the reward value, η charge This indicates the charging efficiency item. Indicates the range of motion, Π misalign This represents the positioning deviation term, calculated using Euclidean distance, max(0,TT). core () represents the temperature term, and T represents the current temperature. core The target temperature is represented by λ1, λ2, and λ3, which are the corresponding penalty coefficients.

[0088] Specifically, the charging efficiency term η charge This is a quantified value of the real-time power transfer efficiency of the wireless charging system. It is the only positive incentive term in the reward function, used to encourage the decision model to output positioning actions that improve coil coupling efficiency. The output power of the transmitting coil and the input power of the receiving coil are collected by voltage and current sensors deployed at the transmitting and receiving ends of the wireless charging coil, respectively. The result is obtained as a percentage, calculated as the ratio of the power of the receiving coil to the power of the transmitting coil. (Action amplitude term) This is the sum of the squares of the position and attitude adjustments made by the device being charged during the positioning process, used to quantify the severity of the device's movements. The penalty coefficient λ1 is used to adjust the penalty weight of the movement amplitude term, preventing mechanical damage or charging interruption caused by large movements. Positioning deviation term Π misalign The center alignment deviation between the wireless charging transmitting coil and the receiving coil is used to quantify the accuracy of coil coupling. The penalty coefficient λ2 is the highest-weighted penalty term in the reward function, used to enforce coil alignment accuracy and prevent a sharp drop in charging efficiency due to excessive deviation. Specifically, the coordinates of the transmitting coil center and the receiving coil center can be collected by a UWB base station array, and the coordinate deviation Π can be calculated using the Euclidean distance formula. misalign (Unit: cm). Temperature term max(0, TT) core The value is the difference between the current temperature of the magnetic core at the transmitter of the wireless charging coil and a safety threshold (a penalty is only applied when the temperature exceeds the threshold), used to quantify the safety status of the magnetic core. The penalty coefficient λ3 is used to prevent the magnetic core from experiencing a decrease in permeability and magnetic field distortion due to high temperatures, thus ensuring device safety.

[0089] The reward function provided by this invention achieves synergistic optimization of efficiency, stability, and safety through positive incentives from the charging efficiency term and negative constraints from three penalty terms, resolving the contradiction in traditional functions of "sacrificing device safety for efficiency" and "reducing positioning accuracy to maintain stability." The penalty coefficients can be flexibly adjusted according to the type of device to be charged and the characteristics of the magnetic core material, adapting to various charging devices without reconstructing the function structure, thus reducing the deployment cost of the technical solution. The quantitative calculation logic of the reward function is clear and can be directly used as the loss function input for deep reinforcement learning models. Combined with the experience replay mechanism recorded in the disclosure document, it can accelerate model convergence and avoid training oscillations caused by target ambiguity. It exhibits outstanding resistance to magnetic core differences: the temperature term is designed to address differences in magnetic core materials (such as the different temperature thresholds of ferrite and nanocrystalline magnetic cores). Through dynamic matching, it can suppress the impact of magnetic field distortion caused by magnetic core differences on positioning. In scenarios with mixed magnetic core deployments, the positioning deviation can still be controlled within centimeters, significantly better than traditional positioning solutions.

[0090] Next, the target state space is input into the actor network of the decision model, and the actor network outputs the current action space. The actor network is a policy network, and its core function is to output the action space that maximizes long-term rewards based on the target state space. The action space parameters include the X-axis displacement adjustment, Y-axis displacement adjustment, and posture angle adjustment of the device to be charged, and the value range needs to match the AGV's motion capabilities. In a specific application scenario, the actor network can adopt a 3-layer fully connected neural network structure. The first hidden layer has a high dimension of 64 and the activation function is ReLU. The second hidden layer has a dimension of 32 and the activation function is still ReLU. The output layer has a dimension of 3 (corresponding to the X-axis displacement adjustment, Y-axis displacement adjustment, and posture angle adjustment) and the activation function is Tanh. The output values ​​are mapped to the [-1,1] interval and then scaled according to the action range. The specific structure of the actor network in this embodiment is only an example and is not limited thereto. In practice, the target state space vector is first input into the actor network input layer. After calculation, the action command "AGV moves 2cm to the right, moves 1cm forward, and rotates 0.2° clockwise" is output. This action space not only conforms to the AGV's motion capability, but also can specifically reduce the coil alignment deviation.

[0091] Subsequently, the critic network of the decision model calculates the reward value based on the reward function. The critic network is a value network whose core function is to evaluate the value of the actor network's output action based on the reward function and the current state, i.e., to calculate the reward value, providing a basis for subsequent network parameter adjustments. In an optional application scenario, the critic network can also use a 3-layer fully connected neural network, with the hidden layer structure consistent with the actor network, an output layer dimension of 1 (corresponding to a single reward value), and a linear activation function (maintaining the continuity of reward values). In specific implementation, sensors collect data in real time after the charging action occurs, including data used to calculate charging efficiency, action amplitude, positioning deviation, and temperature, thus obtaining the latest target state space (vector form) of the charging device. Then, the target state space (vector form) after the action is concatenated with the action space (vector form) output by the actor network to obtain the critic network input (vector form), which is then input into the critic network. After calculation by the hidden layers, the reward value r is calculated through the output layer. t Finally, based on the reward value r t Adjust the model parameters for the actor network and the critic network.

[0092] In this embodiment of the invention, the parameter adjustment of the network model adopts an experience replay mechanism and an actor-critic algorithm for gradient descent update. The core is to update the weight matrix and bias vector of the actor network and the critic network in reverse by using the difference between the reward value and the target value, so that the model gradually learns the optimal positioning strategy.

[0093] The formula for updating model parameters is as follows:

[0094]

[0095] In the formula, J(θ) represents the policy gradient, and E[-] represents the expected value. Represents the scoring function. θ logπ θ θ represents the coefficients of the scoring function, and θ represents the policy parameters. Represents the action space, s t Represents the target state space vector. This represents the advantage function output by the critics' network. Wherein, Δx, Δy, and Δθ represent the X-axis displacement adjustment, Y-axis displacement adjustment, and attitude angle adjustment, respectively.

[0096] Advantage function The key to the stability and efficiency of the actor critic algorithm lies in its definition:

[0097]

[0098] in:

[0099] It is a state value function, representing the state value in state s. t Perform specific actions The expected cumulative reward that can be obtained is used to represent how good the current action is.

[0100] V(s t ) is another state value function, representing the state value in state s. t The expected cumulative return that can be obtained by following the current strategy is used to represent how good the average analysis of the current state is.

[0101] Therefore, the advantage function The current action is indicated by . Compared to the additional benefits that the "average" actions performed in this state could bring.

[0102] like (Positive dominance) indicates the action. If the result is better than average, the product term in the formula becomes positive, which guides the Actor to increase the probability of choosing the current action in a similar state in the future.

[0103] if (Negative advantage) indicates the action If the value is below average, the product term in the formula will become negative, guiding the Actor to reduce the probability of choosing this action in a similar situation in the future.

[0104] if Indicates action It is at the average level and requires no special adjustment.

[0105] This is an idealized theoretical form that assumes the true expectation E (i.e., an infinite number of samples) is already known. The expectation calculation is implemented through an experience replay mechanism, and its workflow and integration method are as follows:

[0106] 1. Experience storage: The experience (tuples) gained by the agent at each step of its interaction in the environment. The data will not be used immediately to update the network, but will first be stored in the playback buffer (a fixed-size database).

[0107] 2. Batch sampling: When network parameters need to be updated (after accumulating enough experience), a batch of experience samples is randomly and uniformly sampled from the playback buffer.

[0108] 3. Calculate the expectation: The expectation E in the formula is approximated using a sampled batch of empirical samples. Specifically, the policy gradient is calculated as follows:

[0109]

[0110] Where N is the batch size, and i indexes each experience in the batch.

[0111] 4. Network Update: The actor's policy network parameters θ are updated using the computed approximate gradient. Simultaneously, the critic network updates its own value estimate using current batch experience (e.g., by minimizing the square of the TD error) to more accurately compute the advantage function in subsequent iterations.

[0112] This invention achieves efficient training of a deep reinforcement learning decision-making model. Compared to traditional fixed-policy models, this invention features a multi-dimensional design for the reward function, balancing charging efficiency, device stability, and magnetic core safety, thus solving the performance imbalance problem caused by single-objective optimization. The actor network is responsible for accurate action output, while the critic network ensures reliable reward evaluation; the two work together to improve the model's decision-making ability. This invention accelerates model convergence and avoids overfitting through a parameter update method that combines experience replay and gradient descent, significantly outperforming traditional localization models and perfectly adapting to complex working conditions such as magnetic core differences and metal interference.

[0113] In some optional implementations, the attention extraction matrix is ​​also adjusted based on the reward value. A learnable attention extraction matrix is ​​configured, and its parameters are adjusted synchronously according to the reward value, linking attention weight allocation with decision model optimization. When the model discovers through the reward value that a certain type of sensor data is more critical for localization, the matrix parameters can be optimized simultaneously, making subsequent weight calculations more accurate. This forms a closed loop of data fusion, decision output, and parameter feedback, continuously improving the effectiveness of the attention mechanism, thereby optimizing the quality of the target state space, ensuring a steady improvement in localization accuracy, and enhancing the system's adaptability and iterative capabilities.

[0114] In some alternative implementations, the wireless charging coils are arranged with vertical and horizontal coils orthogonally.

[0115] Specifically, traditional wireless positioning methods often suffer from the problem of sensitivity to differences in magnetic core materials. Differences in the permeability, saturation characteristics, and temperature coefficient of magnetic core materials can lead to asymmetrical variations in the magnetic field distribution, resulting in fluctuations in coupling efficiency (experiments show that efficiency can decrease by more than 30%).

[0116] The charging principle of a wireless charging coil requires mutual inductance between the primary coil and the vertical and horizontal coils of the secondary coil, as shown in the following formula:

[0117] M total =M tr +M ts ≈(k tr Δx+M tr0 )+k ts Δx

[0118] M total M represents the total mutual inductance (unit: H). tr M represents the mutual inductance between the primary coil and the secondary coil perpendicular to it (unit: H). ts This represents the mutual inductance between the primary coil and the secondary horizontal coil (unit: H). Where M... tr =k tr Δx+M tr0 M tr0 M represents the mutual inductance of the vertical coil at zero offset. ts =k ts Δx, k tr and k ts These represent the linearization slopes of the vertical and horizontal components (unit: H / m), respectively, and Δx represents the offset of the primary and secondary coils.

[0119] This invention proposes that by optimizing the objective |k tr |≈k ts This ensures stable mutual inductance between the primary and secondary coils during offset, thereby reducing the impact of core differences.

[0120] Based on this, this embodiment adopts a structure in which the vertical and horizontal coils are orthogonally arranged to configure the wireless charging coil. The improvement principle is as follows:

[0121] Firstly, it is used for complementary cancellation to counteract the first-order stability of the offset.

[0122] The core idea of ​​this invention is to construct a differential balanced system, assuming that the mutual inductance of the vertical and horizontal coils can be linearized to a first order as the offset changes:

[0123] M tr ≈k tr Δx+M tr0

[0124] M ts ≈k ts Δx

[0125] The total mutual inductance of the system is the sum of the two:

[0126] M total =M tr +M ts ≈(k tr Δx+M tr0 )+k ts Δx=(k tr +k ts )Δx+M tr0

[0127] k tr It is usually a negative value (because the mutual inductance of the vertical coil decreases as the positive offset Δx increases), while k ts It is a positive value (the mutual inductance of the horizontal coil increases as Δx increases).

[0128] When |k tr |≈k ts At that time, k tr +k ts ≈0, thus the (k) in the total mutual inductance formula tr +k ts )Δx≈0, therefore M total ≈M tr0 .

[0129] Within a certain offset range, the total mutual inductance M total Approximately a constant M tr0 It hardly changes with the offset Δx, so that the system is immune to offset from the first-order approximation, thus ensuring the stability of the energy transmission channel.

[0130] Secondly, the underlying principle for mitigating the effects of core differences:

[0131] Differences in core materials (such as ferrite versus nanocrystals) primarily affect permeability, thereby altering the distribution and strength of the magnetic field. This is directly reflected in the effect on M... tr0 k tr k ts The influence of absolute value. However, the optimization objective |k tr |≈k ts The beauty of this lies in pursuing consistency in the rate of change, rather than matching absolute values.

[0132] This embodiment employs a structural configuration with vertical and horizontal coils arranged orthogonally. Even if the core material is changed, resulting in k... tr and k ts The specific values ​​have changed, but the hardware design ensures that |k| is still satisfied under the new core parameters. tr ′|≈k ts If ′(′ represents the new parameter), then the system's anti-offset complementary mechanism still holds, and the total mutual inductance can still remain stable.

[0133] This invention employs a structural configuration with orthogonal vertical and horizontal coils, establishing system stability based on the ratio of two rates of change, rather than their absolute values. This reduces dependence on absolute parameters and lowers the requirement for high consistency in core material properties. As long as the two cores are not drastically different in their magnetic field distribution, the complementary mechanism can function. Therefore, for different batches or types of cores, the system performance will be more consistent, increasing its resistance to variations.

[0134] Thirdly, the collaboration between hardware design and algorithms lays a solid foundation for software optimization.

[0135] The complementary hardware design provides a more stable and linear control object for subsequent algorithm optimization, with specific advantages including:

[0136] 1. It can simplify learning tasks. total In a near-constant system, the relationship between charging efficiency and offset becomes flatter, which greatly reduces the complexity of the "state-action-reward" mapping that the agent in deep reinforcement learning needs to learn. The agent no longer needs to work hard to learn drastically changing nonlinear relationships, thus it can converge to the optimal policy more quickly.

[0137] 2. Provides a foundation for stability. When dynamically adjusting sensor weights, even if the actor network makes a poor decision at a certain moment, the hardware-level stability (stable M) is maintained. total It can also prevent a sharp drop in system performance, increase the system's fault tolerance, and play a double insurance role.

[0138] In summary, by adopting a structural configuration with orthogonal vertical and horizontal coils, the optimization objective |k is generated. tr |≈k ts This is an extremely ingenious hardware strategy that utilizes the opposite response characteristics of vertical and horizontal magnetic field components to offset, directly canceling out the effects of offset in a first-order approximation, thus providing a complementary mechanism. By pursuing consistency in the rate of change rather than equality in absolute values, the system performance becomes insensitive to the specific parameters of the magnetic core material, thereby achieving "anti-difference" and reaching a state of differential equilibrium. This provides a stable and user-friendly optimization environment for the upper-level control algorithm, reduces the learning difficulty, improves the robustness and reliability of the overall system, and achieves the effect of hardware-software synergy.

[0139] Based on the aforementioned technical methods, the wireless charging coil adopts an orthogonal layout of vertical and horizontal coils, which can utilize the complementary vertical and horizontal magnetic field components to cancel out magnetic field distortion caused by differences in the magnetic core. Even when the properties of the magnetic core material change (such as differences in permeability), the total mutual inductance remains relatively stable, reducing the impact of magnetic field distortion on positioning accuracy. Improving the system's resistance to magnetic core differences at the hardware structure level, combined with algorithm optimization, further reduces positioning errors, ensures stable charging efficiency, and enhances the versatility and reliability of the wireless charging system under different magnetic core scenarios.

[0140] This embodiment also provides a wireless charging positioning device for implementing the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0141] This embodiment provides a wireless charging positioning device, such as... Figure 2 As shown, it includes:

[0142] The data acquisition module 201 is used to acquire current sensing data corresponding to multiple sensors, and the sensors are used to acquire information from the wireless charging coil.

[0143] State space generation module 202 is used to fuse the current sensing data to obtain an initial state space;

[0144] Attention weight calculation module 203 is used to determine the attention weight corresponding to each sensor;

[0145] The state space adjustment module 204 is used to calibrate the attention weights to the initial state space to obtain the target state space.

[0146] The decision module 205 is used to input the target state space into a decision model based on deep reinforcement learning, so as to output the current action space through the decision model;

[0147] The position adjustment module 206 is used to adjust the position of the device to be charged according to the current motion space.

[0148] Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.

[0149] This invention also provides a computer device; please refer to [link / reference]. Figure 3 , Figure 3 This is a schematic diagram of the structure of a computer device provided in an optional embodiment of the present invention, such as... Figure 3 As shown, the computer device includes one or more processors 10, memory 20, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components communicate with each other via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the computer device, including instructions stored in or on memory to display graphical information of a GUI on external input / output devices (such as display devices coupled to the interfaces). In some alternative implementations, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple computer devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 3 Take a processor 10 as an example.

[0150] Processor 10 may be a central processing unit, a network processor, or a combination thereof. Processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device may be a complex programmable logic device (CAMP), a field-programmable gate array (FPGA), a general-purpose array logic (GDA), or any combination thereof.

[0151] The memory 20 stores instructions executable by at least one processor 10 to cause at least one processor 10 to perform the method shown in the above embodiments.

[0152] The memory 20 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the computer device. Furthermore, the memory 20 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some alternative embodiments, the memory 20 may optionally include memory remotely located relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0153] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid-state drive; the memory 20 may also include a combination of the above types of memory.

[0154] The computer device also includes an input device 30 and an output device 40. The processor 10, memory 20, input device 30, and output device 20 can be connected via a bus or other means. Figure 3 Taking the example of a connection between China and Israel via a bus.

[0155] Input device 30 can receive input numerical or character information, and generate key signal inputs related to user settings and function control of the computer device, such as a touchscreen, keypad, mouse, trackpad, touchpad, joystick, one or more mouse buttons, trackball, joystick, etc. Output device 40 may include display devices, auxiliary lighting devices (e.g., LEDs), and haptic feedback devices (e.g., vibration motors). The aforementioned display devices include, but are not limited to, liquid crystal displays, light-emitting diodes, displays, and plasma displays. In some alternative embodiments, the display device may be a touchscreen.

[0156] The computer device also includes a communication interface 30 for communicating with other devices or communication networks.

[0157] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods shown in the above embodiments.

[0158] A portion of this invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the invention through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.

[0159] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A wireless charging positioning method, characterized in that, The method includes: Collect current sensing data from multiple sensors, wherein the sensors are used to collect information from the wireless charging coil; The current sensing data are fused together to obtain the initial state space; Determine the attention weight for each sensor; The attention weights are calibrated onto the initial state space to obtain the target state space; The target state space is input into a decision model based on deep reinforcement learning, so that the decision model outputs the current action space. Adjust the position of the device to be charged according to the current action space.

2. The method according to claim 1, characterized in that, The acquisition of current sensing data corresponding to multiple sensors includes: First sensing data is collected from the UWB base station array, which is set at the transmitting or receiving end of the wireless charging coil for UWB positioning of the wireless charging coil. The second sensing data of the magnetic induction coil array is collected. The magnetic induction coil array is set at the transmitting end or receiving end of the wireless charging coil for magnetic coupling positioning of the wireless charging coil. The third sensing data of the inertial measurement unit is collected. The inertial measurement unit is set on the device to be charged and is used to measure the pose of the device to be charged. The fourth sensing data is collected from the temperature sensor, which is located at the transmitter of the wireless charging coil and is used to monitor the operating temperature of the magnetic core in the transmitter of the wireless charging coil.

3. The method according to claim 1 or 2, characterized in that, The process of fusing the current sensing data to obtain an initial state space includes: The current sensing data corresponding to each sensor is corrected by Kalman filtering; The corrected current perception data is normalized, and the normalized current perception data are concatenated into a vector to obtain the initial state space.

4. The method according to claim 3, characterized in that, Determining the attention weight for each sensor includes: Obtain the attention extraction matrix; Based on the attention extraction matrix and the initial state space, calculate the query vector and key vector corresponding to each sensor data; Calculate vector similarity using the query vector and key vector corresponding to each sensor data; The vector similarity of each sensor's data is scaled and normalized. The attention weights for each sensor are obtained by assigning weights based on the vector similarity of the processed data from each sensor.

5. The method according to claim 4, characterized in that, The steps for training the decision model include: Obtain a reward function, which includes a charging efficiency term, a motion amplitude term, a positioning deviation term, and a temperature term. The charging efficiency term is adjusted by multiplying it by a positive reward coefficient, and the motion amplitude term, positioning deviation term, and temperature term are all adjusted by multiplying them by a negative penalty coefficient. The target state space is input into the actor network of the decision model, and the current action space is output through the actor network. The decision-making model uses a network of critics to calculate the reward value based on the reward function. The model parameters of the actor network and the critic network are adjusted based on the reward value.

6. The method according to claim 5, characterized in that, The reward function is: In the formula, r t Represents the reward value, η charge This refers to the charging efficiency term. Indicates the amplitude of the action, Π misalign The positioning deviation term is represented by max(0,TT). core ) represents the temperature term, T represents the current temperature, T core The target temperature is represented by λ1, λ2, and λ3, which are the corresponding penalty coefficients.

7. The method according to claim 5, characterized in that, The method further includes: The matrix parameters of the attention extraction matrix are adjusted based on the reward value.

8. The method according to claim 1, characterized in that, The vertical and horizontal coils in the wireless charging coil are arranged in an orthogonal layout.

9. A wireless charging positioning device, characterized in that, The device includes: The data acquisition module is used to collect current sensing data corresponding to multiple sensors, wherein the sensors are used to collect information about the wireless charging coil; The state space generation module is used to fuse the current sensing data to obtain an initial state space. The attention weight calculation module is used to determine the attention weight corresponding to each sensor; The state space adjustment module is used to calibrate the attention weights to the initial state space to obtain the target state space; The decision module is used to input the target state space into a decision model based on deep reinforcement learning, so as to output the current action space through the decision model; The position adjustment module is used to adjust the position of the device to be charged according to the current action space.

10. A computer device, characterized in that, include: A memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, the processor executing the computer instructions to perform the method of any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a computer to perform the method of any one of claims 1 to 8.