Wireless power transmission method and system

By monitoring environmental parameters in real time and building prediction models, the optimal multi-frequency resonance frequency combination and power distribution scheme are calculated, and the inductor and capacitance component values ​​are dynamically adjusted, which solves the problem that the transmission efficiency of the wireless power transmission system is affected by environmental factors, and efficient and stable power transmission is achieved.

CN120074039AActive Publication Date: 2025-05-30LIAONING UNIVERSITY OF TECHNOLOGY

Patent Information

Application Number
CN202510143215.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-05-30
Estimated Expiration
2045-02-10

AI Technical Summary

Technical Problem

The transmission efficiency of existing wireless power transmission systems is greatly affected by distance, obstacles, and environmental factors, and the frequency and power configuration lacks intelligent adjustment.

Method used

By monitoring environmental parameters in real time, a prediction model is constructed to predict the expected transmission efficiency under different frequency channels, the optimal multi-frequency resonance frequency combination and corresponding power distribution scheme are calculated, and the inductor and capacitance component values ​​are dynamically adjusted using an adaptive tuning network.

Benefits of technology

It realizes the efficiency and stability of wireless power transmission, can automatically adapt to environmental changes, improves the flexibility and response speed of the system, and reduces energy loss.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120074039A_ABST
    Figure CN120074039A_ABST
Patent Text Reader

Abstract

The invention provides a wireless power transmission method and system, and belongs to the technical field of power. Firstly, environmental parameters in a transmission path are collected, and then transmission efficiency under different frequency channels is predicted by using a constructed prediction model; based on a prediction result, dynamically calculating an optimal multi-frequency resonant frequency combination and power distribution scheme by adopting a deep reinforcement learning algorithm so as to adapt to a current transmission environment; the transmitting end dynamically adjusts inductance and capacitance element values according to the calculated optimal configuration, ensures that the resonance condition is matched with the selected frequency, and transmits an alternating electromagnetic field according to a power distribution scheme; a receiving end captures and efficiently absorbs the energy through a resonance circuit, and converts the energy into direct-current electric energy to be used by a load; reasonably distributing the combined electric energy according to load requirements; according to the invention, transmission efficiency and environment change can be monitored in real time, parameters are dynamically adjusted through a closed-loop feedback mechanism, and high efficiency and stability of wireless power transmission are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of electric power, and particularly relates to a wireless power transmission method and system. Background Art

[0002] Traditional power transmission relies on physical wires to transmit electrical energy. This method has limitations in some application scenarios, such as crossing complex terrains, powering mobile devices, and environments where it is difficult to lay cables. With the development of technology, wireless power transmission technology has gradually become a research hotspot, aiming to overcome the inconvenience of wired transmission and achieve more flexible, safe, and efficient energy distribution. Early attempts at wireless power transmission, such as Tesla coils, demonstrated that electricity can be transmitted without contact, but the efficiency and practicality were limited. In recent years, with the progress of technologies such as electromagnetic induction, radio frequency, resonance, and microwave, the efficiency and feasibility of wireless power transmission have been significantly improved.

[0003] However, existing wireless power transmission systems still face challenges, including significant influence of transmission efficiency by distance, obstacles, and environmental factors, as well as lack of intelligent adjustment of frequency and power configuration. Therefore, it is particularly important to develop a wireless power transmission method and system that can adapt to environmental changes in real time, dynamically optimize transmission parameters, and be efficient and stable. Summary of the Invention

[0004] Based on the above technical problems, the present invention provides a wireless power transmission method and system, which optimize the transmission efficiency by real-time monitoring and adjusting environmental parameters. The frequency and power are dynamically adjusted through an intelligent scheduling algorithm to ensure the high efficiency and stability of power transmission.

[0005] The present invention provides a wireless power transmission method, which includes:

[0006] Step S1: Collect environmental parameters in the wireless power transmission path;

[0007] Step S2: Construct a prediction model to predict the expected transmission efficiency under different frequency channels;

[0008] Step S3: Calculate the optimal multi-frequency resonance frequency combination and the corresponding power distribution scheme;

[0009] Step S4: The transmitting end dynamically adjusts the values of inductance and capacitance elements according to the optimal operating frequency combination and an adaptive tuning network;

[0010] Step S5: The transmitting end emits an alternating electromagnetic field to the receiving end according to the power distribution scheme; the receiving end captures the alternating electromagnetic field emitted by the transmitting end and absorbs energy through a resonant circuit;

[0011] Step S6: The receiving end integrates the energy received from each frequency channel for combination, and intelligently distributes the combined electric energy to the target load according to the load demand.

[0012] Optionally, the construction of the prediction model predicts the expected transmission efficiency under different frequency channels, specifically including:

[0013] Define a feature extraction sub-network, which receives an input layer and a list representing the number of neurons in each layer, and then creates a series of fully connected layers, each layer using the ReLU activation function;

[0014] Define an attention mechanism layer, which converts the input into a probability distribution through a fully connected layer (using the Softmax activation); multiplies these probabilities with the original input through a dot product operation;

[0015] Construct a prediction model, create multiple input layers according to the input shape; create a feature extraction sub-network for each input layer; use a concatenation layer to combine the outputs of all sub-networks into a single vector; add an attention mechanism layer to dynamically emphasize the feature combinations that the model focuses on; use the concatenation layer again to combine the original combined features and the features processed by the attention mechanism; add a fully connected layer to further process these features; create an output layer for each frequency channel, using the Sigmoid activation function, and each output corresponds to the prediction of a frequency channel; compile the model, and the loss function is the mean squared error.

[0016] Optionally, the calculation of the optimal multi-frequency resonance frequency combination and the corresponding power distribution scheme specifically includes:

[0017] Initialize the deep reinforcement learning agent class for learning and making decisions on the optimal configuration of frequency and power; define the state space size; define the action space size; use a deque to set an experience replay buffer for storing each interaction data; set the discount factor, initial exploration rate, minimum exploration rate, exploration rate decay rate, learning rate; construct a deep Q-network model, implemented by a construction function;

[0018] Construct a neural network function for estimating the action value, with the input being the current environmental state (such as transmission distance, attenuation), and the output being the value of each action; the model consists of three fully connected layers, the first and second layers use the ReLU activation function to increase the non-linear processing ability, and the last layer uses the linear activation function to output the expected value of each action (i.e., the Q value);

[0019] Save the agent's experience (state, action, reward) to the memory replay, that is, store the data of one interaction in the experience replay buffer;

[0020] The action selected according to the current strategy (exploration or exploitation): when the random number is less than the current exploration rate, randomly select an action; otherwise, select the action corresponding to the maximum Q value predicted by the model as the expected optimal action;

[0021] Randomly sample a batch of experiences from memory for learning, and optimize the network weights by updating the target Q value to minimize the mean squared error between the predicted Q value and the target Q value, specifically including:

[0022] Randomly sample a certain number of samples from the experience replay buffer; the samples contain the states, actions taken, rewards obtained, next states, and information on whether it is the end during past exploration;

[0023] Process each sampled sample in a loop and unpack each element in the sample;

[0024] Calculate the target value. If this experience is a terminal state, the target value is the immediate reward obtained; if it is not a terminal state, the target value is the immediate reward plus the discounted value of the future expected reward;

[0025] Use the current model to predict the Q values of all actions in a given state;

[0026] Assign the calculated target value to the Q value corresponding to the action taken;

[0027] Use the updated target Q value as a label to train the model once, and optimize its weights to reduce the gap between the predicted Q value and the target Q value;

[0028] Update the initial exploration rate;

[0029] Output the instruction set, including the selected multi-frequency resonance frequency combination and the corresponding power distribution plan.

[0030] Optionally, the transmitting end dynamically adjusts the values of the inductor and capacitor elements according to the optimal operating frequency combination and the adaptive tuning network, specifically including:

[0031] After calculating the best frequency combination, send instructions to the controller of the tuning network through a digital interface. The instructions contain the target frequency f target ; The controller calculates the adjustment amount of the current capacitor C target or inductor L current according to the target frequency f current to match the target frequency;

[0032] According to the calculated adjustment amount, the controller dynamically adjusts the physical parameters of the variable capacitor or variable inductor through motor control or an electronic driver to achieve the calculated value of the capacitor C new or inductor L new ;

[0033] Monitor the tuned resonant frequency f using a high-frequency precision frequency sensor measured and compare it with the target frequency f target . If there is a deviation e, adjust the capacitance or inductance through a closed-loop control algorithm, expressed as:

[0034]

[0035] where u(t) is the control signal; K b , K i and K d are the proportional, integral, and derivative gains respectively.

[0036] Optionally, the transmitting end emits an alternating electromagnetic field to the receiving end according to a power distribution scheme; the receiving end captures the alternating electromagnetic field emitted by the transmitting end and absorbs energy through a resonant circuit, specifically including:

[0037] According to the calculated frequency and the corresponding power distribution, for each calculated frequency f α , the frequency synthesizer generates a corresponding alternating current signal, expressed as:

[0038]

[0039] where I α (t) is the current of the α-th frequency component; A α is the amplitude; φ α is the phase angle; the calculated frequencies are f 1 , f 2 ,..., f n ; the power distributions are P allocated,1 , P allocated,2 ,..., P allocated,n ;

[0040] According to the power distribution scheme, amplify the current signal of each frequency. Considering the amplification efficiency, the actual output power is expressed as:

[0041] P out,c = η PA,c ·P allocated,c

[0042] where P out,c is the actual output power amplified by the c-th channel; η PA,c is the power amplifier efficiency of the c-th channel;

[0043] Before transmission, use pre-distortion technology. The receiving end is equipped with multiple resonant circuits corresponding to the frequencies of the transmitting end. Each circuit is optimized for a specific frequency. By adjusting the values of the capacitors and inductors, ensure that the resonant frequency of the circuit matches the frequency of the received electromagnetic wave;

[0044] Adjust the Q value of the resonant circuit to improve the energy absorption efficiency, expressed as:

[0045]

[0046] In the formula, ∝ is proportional to; f i is the emission frequency; η abs,r is the absolute efficiency of the r-th resonant circuit; f 0,r is the center frequency of the r-th resonant circuit; Q r is the quality factor of the r-th resonant circuit; f BW,r is the bandwidth of the r-th resonant circuit;

[0047] Convert the alternating current absorbed in the resonant circuit into direct current, which is realized through a rectifier bridge circuit, expressed as:

[0048]

[0049] In the formula, η conv is the conversion efficiency; P DC is the converted DC power; P AC is the absorbed AC power.

[0050] The present invention also provides a wireless power transmission system, which includes:

[0051] An environmental parameter collection module for collecting environmental parameters in the wireless power transmission path;

[0052] A transmission efficiency prediction module for constructing a prediction model to predict the transmission efficiency of different frequency channels;

[0053] A power scheme allocation module for calculating the optimal multi-frequency resonance frequency combination and the corresponding power allocation scheme;

[0054] An electrical component adjustment module for the transmitting end to dynamically adjust the values of inductance and capacitance elements according to the optimal operating frequency combination and the adaptive tuning network;

[0055] A transmitting and receiving matching module for the transmitting end to transmit an alternating electromagnetic field to the receiving end according to the power allocation scheme; the receiving end captures the alternating electromagnetic field emitted by the transmitting end and absorbs energy through the resonant circuit;

[0056] An energy intelligent allocation module for the receiving end to integrate the energy received from each frequency channel for merging, and intelligently allocate the merged electric energy to the target load according to the load demand.

[0057] Optionally, the transmission efficiency prediction module specifically includes:

[0058] Define the feature extraction sub-network, which receives an input layer and a list representing the number of neurons in each layer, and then creates a series of fully connected layers, each using the ReLU activation function;

[0059] Define the attention mechanism layer, which converts the input into a probability distribution through a fully connected layer (using the Softmax activation); multiplies these probabilities with the original input through a dot product operation;

[0060] Build the prediction model. Create multiple input layers according to the input shape; create a feature extraction sub-network for each input layer; use a concatenation layer to merge the outputs of all sub-networks into a single vector; add an attention mechanism layer to dynamically emphasize the feature combinations that the model focuses on; use the concatenation layer again to combine the original merged features and the features processed by the attention mechanism; add a fully connected layer to further process these features; create an output layer for each frequency channel, using the Sigmoid activation function, and each output corresponds to the prediction of a frequency channel; compile the model, and the loss function is the mean squared error.

[0061] Optionally, the power scheme allocation module specifically includes:

[0062] Initialize the deep reinforcement learning agent class for learning and making decisions on the optimal configuration of frequency and power; define the state space size; define the action space size; use a deque to set up an experience replay buffer for storing each interaction data; set the discount factor, initial exploration rate, minimum exploration rate, exploration rate decay rate, and learning rate; build the deep Q-network model, implemented by a construction function;

[0063] Build a neural network function for estimating the action value. The input is the current environmental state (such as transmission distance, attenuation), and the output is the value of each action; the model consists of three fully connected layers. The first and second layers use the ReLU activation function to increase the non-linear processing ability, and the last layer uses the linear activation function to output the expected value of each action (i.e., the Q value);

[0064] Save the agent's experience (state, action, reward) to the memory replay, that is, store the data of one interaction in the experience replay buffer;

[0065] For the action selected according to the current policy (exploration or exploitation), when the random number is less than the current exploration rate, randomly select an action; otherwise, select the action corresponding to the maximum Q value predicted by the model as the expected optimal action;

[0066] Randomly sample a batch of experiences from the memory for learning, and optimize the network weights by updating the target Q value to minimize the mean squared error between the predicted Q value and the target Q value, specifically including:

[0067] Randomly sample a certain number of batches of samples from the experience replay buffer; the samples contain the states during past exploration, the actions taken, the rewards obtained, the next states, and information on whether it is the end.

[0068] Process each sampled sample in a loop and unpack each element in the sample.

[0069] Calculate the target value. If this experience is a terminal state, the target value is the immediate reward obtained; if it is not a terminal state, the target value is the immediate reward plus the discounted value of the future expected reward.

[0070] Use the current model to predict the Q-values of all actions in a given state.

[0071] Assign the calculated target value to the Q-value corresponding to the action taken.

[0072] Use the updated target Q-values as labels to train the model once and optimize its weights to reduce the gap between the predicted Q-values and the target Q-values.

[0073] Update the initial exploration rate.

[0074] Output the instruction set, including the selected multi-frequency resonance frequency combination and the corresponding power distribution plan.

[0075] Optionally, the electrical component adjustment module specifically includes:

[0076] After calculating the optimal frequency combination, send the instruction to the controller of the tuning network through the digital interface. The instruction contains the target frequency f target ; The controller calculates the current capacitance C target or inductance L current according to the target frequency f current adjustment amount to match the target frequency;

[0077] According to the calculated adjustment amount, the controller dynamically adjusts the physical parameters of the variable capacitor or variable inductor through motor control or electronic drive to achieve the calculated capacitance C new or inductance L new value;

[0078] Use a high-frequency precision frequency sensor to monitor the tuned resonance frequency f measured and compare it with the target frequency f target . If there is a deviation e, adjust the capacitance or inductance through a closed-loop control algorithm, expressed as:

[0079]

[0080] In the formula, u(t) is the control signal; K b , K i and Kd They are the proportional, integral, and derivative gains respectively.

[0081] Optionally, the transmitting and receiving matching module specifically includes:

[0082] According to the calculated frequency and the corresponding power distribution, for each calculated frequency f α , the frequency synthesizer generates a corresponding alternating current signal, expressed as:

[0083]

[0084] In the formula, I α (t) is the current of the α-th frequency component; A α is the amplitude; φ α is the phase angle; the calculated frequency is f 1 , f 2 ,..., f n ; the power distribution is P allocated,1 , P allocated,2 ,..., P allocated,n ;

[0085] According to the power distribution scheme, the current signal of each frequency is amplified. Considering the amplification efficiency, the actual output power is expressed as:

[0086] P out,c = η PA,c ·P allocated,c

[0087] In the formula, P out,c is the actual output power after amplification of the c-th channel; η PA,c is the power amplifier efficiency of the c-th channel;

[0088] Before transmission, pre-distortion technology is adopted. The receiving end is equipped with multiple resonant circuits corresponding to the frequencies of the transmitting end. Each circuit is optimized for a specific frequency. By adjusting the values of the capacitor and inductor, it is ensured that the resonant frequency of the circuit matches the frequency of the received electromagnetic wave;

[0089] Adjust the Q value of the resonant circuit to improve the energy absorption efficiency, expressed as:

[0090]

[0091] In the formula, ∝ is proportional to; f i is the transmission frequency; η abs,r is the absolute efficiency of the r-th resonant circuit; f 0,r is the center frequency of the r-th resonant circuit; Q r is the quality factor of the r-th resonant circuit; f BW,r is the bandwidth of the r-th resonant circuit;

[0092] Convert the alternating current absorbed in the resonant circuit into direct current, which is achieved through a rectifier bridge circuit, expressed as:

[0093]

[0094] In the formula, η conv is the conversion efficiency; P DC is the direct current power after conversion; P AC is the alternating current power absorbed.

[0095] Compared with the prior art, the present invention has the following beneficial effects:

[0096] Through the deep reinforcement learning algorithm, the system can autonomously learn and optimize the frequency and power configuration, automatically adapt to changes in the transmission environment, without manual intervention, improving the flexibility and response speed of the system; the prediction model and optimization algorithm ensure that the best frequency combination and power distribution scheme can be found in various environments, improving the energy transmission efficiency and reducing losses; the closed-loop feedback control mechanism continuously monitors and adjusts the system state to ensure efficient and stable power transmission even in the face of environmental interference or load changes, enhancing the reliability and continuous power supply ability of the system; the system design is applicable to a variety of application scenarios, including but not limited to homes, industries, mobile device charging, and power supply for remote monitoring devices, broadening the application scenarios of wireless power transmission. Brief Description of the Drawings

[0097] Figure 1 is a flowchart of a wireless power transmission method of the present invention;

[0098] Figure 2 is a structural diagram of a wireless power transmission system of the present invention. Detailed Embodiments

[0099] The present invention will be further described below in conjunction with specific implementation cases and the drawings, but the present invention is not limited to these embodiments.

[0100] Embodiment 1

[0101] As Figure 1 shown, the present invention discloses a wireless power transmission method, which includes:

[0102] Step S1: Collect environmental parameters in the wireless power transmission path.

[0103] Step S2: Build a prediction model to predict the expected transmission efficiency under different frequency channels.

[0104] Step S3: Calculate the optimized multi-frequency resonance frequency combination and the corresponding power distribution scheme.

[0105] Step S4: The transmitting end dynamically adjusts the values of inductance and capacitance components according to the optimal operating frequency combination and the adaptive tuning network.

[0106] Step S5: The transmitting end transmits an alternating electromagnetic field to the receiving end according to the power distribution scheme; the receiving end captures the alternating electromagnetic field emitted by the transmitting end and absorbs energy through the resonant circuit.

[0107] Step S6: The receiving end integrates the energy received from each frequency channel for combination, and intelligently distributes the combined electric energy to the target load according to the load demand.

[0108] The following elaborates on each step in detail.

[0109] Step S1: Collect environmental parameters in the wireless power transmission path.

[0110] Step S1 specifically includes:

[0111] Design and deployment of the environmental perception unit, specifically including:

[0112] The environmental perception unit usually integrates a variety of sensors and signal processing devices, including but not limited to:

[0113] Distance sensors (such as lidar, ultrasonic sensors, or infrared sensors), which measure the real-time distance between the transmitting end and the receiving end, and identify and track the positions of obstacles.

[0114] Substance identification sensors (such as near-infrared spectrometers or material identification radars), which analyze the types and materials of obstacles, because obstacles with different materials and shapes have different absorption, reflection, and diffraction effects on electromagnetic waves, affecting the transmission path and efficiency, and are crucial for frequency selection.

[0115] Load status monitors, which are connected to the receiving end or directly integrated into the receiving device, and monitor real-time power consumption, voltage, current and other parameters of the load to ensure that the power supply matches the demand.

[0116] Environmental condition sensors (such as temperature, humidity, and air pressure sensors), which monitor environmental factors because these factors may affect the propagation of electromagnetic waves and the performance of devices.

[0117] These perception units are installed at the transmitting end, the receiving end, and key positions in the transmission path to ensure comprehensive coverage and real-time monitoring of all relevant parameters.

[0118] Data acquisition and preliminary processing, specifically including:

[0119] The environmental perception unit continuously collects data from the above-mentioned various sensors to achieve continuous monitoring of the transmission environment.

[0120] The collected data is subjected to preliminary verification and filtering to remove outliers or noise and ensure data accuracy. The original data also needs to be normalized to adapt to the input format of subsequent algorithms.

[0121] Data transmission and synchronization, specifically including:

[0122] The processed environmental parameter data is packaged into a unified format for easy transmission and analysis.

[0123] Use low-power wireless communication technologies (such as Wi-Fi, Bluetooth, Zigbee, or a dedicated wireless protocol) to send the data to the central processing unit or the intelligent scheduling system.

[0124] Ensure the time synchronization of all sensing units to facilitate the system's accurate understanding of the spatio-temporal relationship of the data, which is crucial for dynamically adjusting transmission parameters.

[0125] Through the above steps, real-time environmental parameter monitoring provides real-time and accurate environmental information for the entire wireless power transmission system and is the basis for realizing intelligent and efficient energy transmission.

[0126] Step S2: Build a prediction model to predict the expected transmission efficiency under different frequency channels.

[0127] Step S2 specifically includes:

[0128] Collect information on the current transmission path through environmental perception technologies (such as sensors, cameras, RF signal feedback, etc.), which may include but are not limited to changes in transmission distance (due to movement or changes in obstacle positions), the degree of electromagnetic interference in the surrounding environment, the material of obstacles and their absorption or reflection characteristics of electromagnetic waves at specific frequencies, and the specific requirements of the receiving-end load (such as changes in power requirements).

[0129] Based on historical data, build a prediction model that combines machine learning algorithms to predict the expected transmission efficiency under different frequency channels according to the input environmental parameters, specifically including:

[0130] Import the TensorFlow library and the layers and models modules in Keras to build and compile a neural network model, expressed as import tensorflow as tf; from tensorflow.keras import layers, models.

[0131] Define a feature extraction subnetwork that takes an input layer and a list representing the number of neurons in each layer, and then creates a series of fully connected layers (Dense layers), each using the ReLU activation function; this process is performed for each input feature vector with the aim of extracting higher-level features from the original input, denoted as def create_feature_subnetwork(input_layer, layer_sizes):; x = input_layer; for size in layer_sizes: x = layers.Dense(size, activation='relu')(x); return x.

[0132] Define an attention mechanism layer that converts the input into a probability distribution through a fully connected layer (using the Softmax activation), where these probabilities represent the importance of each input feature; multiply these probabilities with the original input through a dot product operation to highlight the important features, denoted as def create_attention_layer(inputs, attention_size): attention_probs = layers.Dense(attention_size, activation='softmax', name='attention_vec')(inputs); attention_mul = layers.multiply([inputs, attention_probs]); return attention_mul.

[0133] Build a prediction model. Create multiple input layers according to the input shapes (corresponding to different input sources), expressed as input_layers = [layers.Input(shape=(shape,)) for shape in input_shapes]; create a feature extraction sub-network for each input layer, expressed as subnets = [create_feature_subnetwork(input_layer, [128, 128]) for input_layer in input_layers]; use the concatenation (concat) layer to combine the outputs of all sub-networks into a single vector, which helps the model understand the interaction between different inputs, expressed as merged = layers.concatenate(subnets); add an attention mechanism layer to dynamically emphasize the feature combinations that the model focuses on, expressed as attention_output = create_attention_layer(merged, 128); use the concatenation (concat) layer again to combine the original merged features and the features processed by the attention mechanism, expressed as merged_with_attention = layers.concatenate([merged, attention_output]); add a fully connected layer to further process these features, expressed as dense_output = layers.Dense(256, activation='relu')(merged_with_attention); create an output layer for each frequency channel, using the Sigmoid activation function, and each output corresponds to the prediction of a frequency channel, expressed as final_outputs = [layers.Dense(1, activation='sigmoid')(dense_output) for _ in range(5)]; compile the model, and the loss function is MSE (mean squared error), expressed as model = models.Model(inputs = input_layers, outputs = final_outputs).

[0134] In this embodiment, the prediction model includes multiple input layers, which respectively process different types of input data (such as distance, electromagnetic interference, obstacle material, etc.), a feature extraction sub-network, an attention mechanism layer, a fusion layer, and an output layer (generating prediction efficiency for each frequency channel).

[0135] Step S3: Calculate the optimized multi-frequency resonance frequency combination and the corresponding power allocation scheme.

[0136] Step S3 specifically includes:

[0137] Evaluate the suitability of each potential resonance frequency (from a preset series of frequencies) for energy transmission in the current environment, considering factors such as the penetration ability of the frequency, the attenuation degree affected by the environment, and the coupling efficiency with the receiving end; determine one or more groups of frequency combinations that can minimize energy loss and improve transmission efficiency to the greatest extent. For the selected frequency combinations, calculate the optimal power ratio to be allocated to each frequency channel, specifically including:

[0138] Introduce libraries and initialize, such as the deque in the NumPy, TensorFlow, and collections modules, as well as the random library, which are the basis for building and training the DQN model.

[0139] Initialize the deep reinforcement learning agent class for learning and making decisions on the optimal configuration of frequencies and powers, expressed as def__init__(self,state_size,action_size):; define the state space size (the number of environmental parameters), expressed as self.state_size = state_size; define the action space size (such as the number of possible frequency and power combinations), expressed as self.action_size = action_size; use a deque to set up an experience replay buffer for storing each interaction data, expressed as self.memory = deque(maxlen = 2000); set the discount factor (used to calculate the current value of future rewards), expressed as self.gamma = 0.95; set the initial exploration rate (determining whether to take random actions or frequencies predicted by the model), expressed as self.epsilon = 1.0; set the minimum exploration rate (to prevent completely stopping exploration), expressed as self.epsilon_min = 0.01; set the exploration rate decay rate (reducing the exploration rate after each learning), expressed as self.epsilon_decay = 0.995; set the learning rate (affecting the amplitude of model parameter updates), expressed as self.learning_rate = 0.001; build the deep Q-network model, expressed as self.model = self._build_model(), implemented by the build function.

[0140] Build a neural network function for estimating action values. The input is the current environmental state (such as transmission distance, attenuation, etc.), and the output is the value of each possible action (i.e., the expected return). The model consists of three fully connected layers. The ReLU activation function is used in the first and second layers to increase the non-linear processing ability. The linear activation function is used in the last layer to output the expected value of each action (i.e., the Q value), expressed as def_build_model(self): model = models.Sequential(); model.add(Dense(24, input_dim = self.state_size, activation='relu')); model.add(Dense(24, activation='relu')); model.add(Dense(self.action_size, activation='linear')); model.compile(loss='mse', optimizer = keras.optimizers.Adam(lr = self.learning_rate)); return model.

[0141] Save the agent's experience (state, action, reward, etc.) to the memory replay, that is, store the data of one interaction in the experience replay buffer for subsequent learning, expressed as def remember(self, state, action, reward, next_state, done): self.memory.append((state, action, reward, next_state, done)).

[0142] The action selected according to the current policy (exploration or exploitation). When the random number is less than the current exploration rate, randomly select an action; otherwise, select the action corresponding to the maximum Q value predicted by the model, that is, the expected optimal action, expressed as def act(self, state): if np.random.rand() <= self.epsilon: return random.randrange(self.action_size); act_values = self.model.predict(state); return np.argmax(act_values[0]).

[0143] Randomly sample a batch of experiences from memory and optimize the network weights by updating the target Q-value (considering future rewards) to minimize the mean squared error (MSE) between the predicted Q-value and the target Q-value, specifically including:

[0144] Randomly sample batch_size samples from the experience replay buffer self.memory. These samples contain information about the state, action, reward, next state, and whether it is done during past exploration, denoted as minibatch = random.sample(self.memory, batch_size).

[0145] Process each sampled sample in a loop and unpack each element in the sample, denoted as for state, action, reward, next_state, done in minibatch:.

[0146] Calculate the target value. If this experience is a terminal state (done is True), the target value is the immediate reward obtained; if it is not a terminal state, the target value, in addition to the immediate reward, also includes the discounted present value of future expected rewards, that is, the Q-value calculated using the Bellman equation. Here, self.gamma is used as the discount factor to predict the maximum Q-value of all possible actions in the next state, denoted as target = reward if done else reward + self.gamma * np.amax(self.model.predict(np.array([next_state]))[0]).

[0147] Use the current model to predict the Q-values of all actions for the given state, denoted as target_f = self.model.predict(state).

[0148] Assign the calculated target value target to the Q-value corresponding to the taken action for targeted update, that is, the TD error correction in Q-learning, denoted as target_f[0][action] = target.

[0149] Use the updated target Q-value target_f as a label to train the model once (single iteration, epochs = 1), optimizing its weights to reduce the gap between the predicted Q-value and the target Q-value. verbose = 0 means not to display the log information during training, which is expressed as self.model.fit(state, target_f, epochs = 1, verbose = 0).

[0150] Update the important parameter epsilon in the exploration and exploitation strategy. Epsilon usually controls the probability of randomly selecting an action, initially high to encourage exploration, and gradually decreasing as learning progresses to increase the exploitation of the known best strategy. Reduce its value by multiplying the current epsilon by the decay factor self.epsilon_decay and ensure that it does not fall below the preset minimum value self.epsilon_min to prevent the exploration probability from dropping to zero, which is expressed as self.epsilon = max(self.epsilon_min, self.epsilon × self.epsilon_decay).

[0151] Define the sizes of the state space and the action space, and instantiate an object of the reinforcement learning agent class to guide the frequency and power adjustment of the wireless power transfer system. By continuously interacting with the environment, accumulating experience, and performing batch learning at the appropriate time, the DQN agent gradually learns to select the frequency and power configuration strategy that is most beneficial to the energy transfer efficiency and system stability in different environmental states, especially suitable for solving problems such as wireless power transfer that require decision-making in complex and dynamic environments.

[0152] In this embodiment, the state space and the action space define the framework for the agent to interact with the environment. The state space represents how many state variables can be used to describe the system or the environment. These state parameters may include but are not limited to: transmission distance (distance from the transmitter to the receiver), obstacle information (such as the number, location, type of obstacles and their absorption or reflection characteristics of electromagnetic waves at specific frequencies), environmental attenuation degree (attenuation amount of electromagnetic waves in the air or when passing through obstacles), receiver coupling efficiency (receiver's ability to receive wireless power), load demand (current power level required by the receiver), system energy consumption (current energy consumption level), system temperature (to prevent overheating risk), and other environmental factors such as temperature, humidity, etc. that may affect the wireless transmission efficiency.

[0153] The action space refers to the number of different actions that an agent (in this case, a wireless power transfer system) can take in a given state; it indicates that different combinations of frequencies and power configurations are designed, and the agent can select one of these configurations as an action. Each action may represent a different frequency selection (each frequency may correspond to one or more specific wireless power transfer channels); the power allocation scheme (how to allocate the total output power at the selected frequency to achieve the best transmission efficiency and system stability); due to the complexity of actual operations and the limitations of computing resources, directly setting an extremely large action space may be impractical and not conducive to model convergence; a hierarchical or combined strategy is adopted to reduce the action space. The first is frequency band optimization, that is, dividing the available frequencies into several frequency bands and selecting an optimal frequency within each band, thus decomposing the problem into selecting a frequency band and adjusting the power within the selected band; for example, if there are three frequency bands and three power levels within each band, then the action space = 3 (frequency band selection) × 3 (power levels) = 9; the second is to pre-compute a series of efficient frequency-power combinations instead of exhausting all possibilities; if 10 groups of efficient combinations are determined after screening, then the action space = 10.

[0154] In this embodiment, deep reinforcement learning (DRL) is used to optimize frequency selection and power allocation. Deep reinforcement learning is a method that allows the model to learn the optimal strategy by interacting with the environment and is applicable to decision-making problems in dynamic environments, such as the frequency and power optimization of wireless power transfer systems; the above prediction model solves the problem of how to predict the transmission efficiency under different conditions, while deep reinforcement learning, after obtaining these prediction capabilities, solves the problem of how to make the best frequency and power adjustment strategies based on these predictions to achieve the highest efficiency.

[0155] Finally, the algorithm outputs a set of instructions, including the selected multi-frequency resonance frequency combination and the corresponding power allocation plan, and these instructions are used to adjust the transmission parameters of the transmitter and the receiving and processing settings of the receiver, specifically including:

[0156] After intelligent frequency and power optimization, the system algorithm (intelligent scheduling algorithm) will generate a set of specific execution instructions, which constitute the core content of the optimization scheme. It not only includes the most efficient multi-frequency resonance frequency combination in the current environment but also clarifies the power allocation details to be adopted on each frequency channel. Specifically:

[0157] Frequency combination, which lists one or more sets of optimal frequencies that have been proven to achieve the best electromagnetic wave coupling, reduce energy loss, and improve transmission efficiency under the current environmental conditions; each frequency corresponds to a specific wireless power transfer channel to ensure that energy can be efficiently transmitted through multiple paths.

[0158] The power distribution plan assigns the most suitable power output to each selected frequency channel, which is based on the evaluation of the efficiency of each channel, the total energy consumption limit, and the actual demand of the load. The aim is to ensure that the power demand of the load is maximally met without sacrificing the system stability.

[0159] The optimization scheme is output in the form of an instruction set, which is convenient for the system control unit to directly interpret and convert into actual control commands for the hardware, such as adjusting the frequency synthesizer and power amplifier settings at the transmitting end, as well as the matching network and energy converter configurations at the receiving end.

[0160] According to the instructions output by the algorithm, the hardware components make corresponding adjustments to implement the new frequency and power strategies. Meanwhile, the system continuously monitors the actual transmission effect. Through the closed-loop feedback mechanism, if a decrease in efficiency or environmental changes are detected, the algorithm will restart the optimization calculation, dynamically adjust the strategy, and ensure that the transmission efficiency and stability are always maintained at the best state, specifically including:

[0161] Hardware adjustment: According to the optimization scheme, the hardware components at the transmitting end (such as RF frequency source, power amplifier) and the hardware at the receiving end (such as resonant circuit, rectifier) will be automatically adjusted to the specified frequency and power settings; this step ensures that the physical layer of the system matches the optimal configuration calculated by the algorithm.

[0162] System configuration: The software layer will also perform corresponding configuration updates to ensure that the control system can accurately track and manage the current frequency and power settings, and prepare for subsequent feedback control.

[0163] Real-time monitoring: The sensors and monitoring modules built into the system continuously track the performance indicators of wireless power transmission, such as transmission efficiency, energy consumption, temperature, and changes in environmental parameters (new obstacles or distance changes).

[0164] Performance evaluation: Through the algorithm, the collected real-time data is quickly analyzed to evaluate whether the transmission efficiency of the current configuration reaches the expectation and whether there are any changes that may affect the efficiency or stability.

[0165] Closed-loop feedback: Once a performance decrease or environmental change beyond the threshold is detected, the algorithm will automatically trigger the re-optimization process, perform environmental parameter monitoring and intelligent scheduling algorithm calculation again, generate a new frequency and power optimization scheme, and instruct the hardware to adjust.

[0166] Step S4: After receiving the instruction, the transmitting end dynamically adjusts the values of the inductor and capacitor elements according to the optimal operating frequency combination and the adaptive tuning network.

[0167] Step S4 specifically includes:

[0168] After calculating the optimal frequency combination, the instructions are sent to the controller of the tuning network through a digital interface (such as a serial communication protocol like I 2 C, SPI or network interface), and the instructions contain the target frequency f target .

[0169] Based on the target frequency f target , the controller calculates the adjustment amount of the current capacitance C current or inductance L current to match the target frequency, specifically including:

[0170] Assume the capacitance and inductance before adjustment are C current and L current , respectively, then:

[0171]

[0172] Read the current capacitance C current or inductance L current . If f current (current frequency) is different from f target , adjust it to the target frequency f target . Solve one of the following equations to find the adjusted C new or L new , and keep the other parameter unchanged;

[0173]

[0174] If the capacitance C current is selected for adjustment, then:

[0175]

[0176] If the inductance L current is selected for adjustment, then:

[0177]

[0178] Based on the calculated adjustment amount, the controller dynamically adjusts the physical parameters of the variable capacitor or variable inductor through motor control (such as a stepper motor or servo motor) or an electronic driver (such as a piezoelectric driver) to achieve the calculated capacitance C new or inductance L new value.

[0179] The variable capacitor adjusts the capacitance value by changing the distance or area between the two plates. The capacitance formula is:

[0180] C = εA / d

[0181] Where ε is the permittivity; A is the plate area; d is the distance between the plates; dynamic adjustment is achieved through motor drive or microelectromechanical systems (MEMS).

[0182] The variable inductor adjusts the inductance value by changing the number of turns of the coil, the coil diameter, or the permeability of the magnetic core material. The inductance formula is:

[0183] L = μN 2 S / l

[0184] Where μ is the permeability; N is the number of turns; S is the cross-sectional area of the coil; l is the length of the coil; dynamic adjustment is achieved using technologies such as magnetorheological fluids or shape memory alloys.

[0185] Use a high-frequency precision frequency sensor (such as a PLL or FLL circuit) to monitor the tuned resonant frequency f measured , and compare it with the target frequency f target . If there is a deviation e, then adjust the capacitance or inductance through a closed-loop control algorithm (such as PID control), expressed as:

[0186]

[0187] e = f measured -f target

[0188] Where u(t) is the control signal; K b , K i and K d are the proportional, integral, and derivative gains respectively; the role of the proportional part is to provide an immediate corrective action, the magnitude of which is directly proportional to the magnitude of the current error, aiming to quickly respond and start correcting the system deviation; the integral part integrates the error signal e(t) to accumulate past errors to eliminate the steady-state error; the derivative part is the derivative of the error signal e(t), that is, calculates the error change rate, and this part is used to predict and suppress rapid changes in the system.

[0189] In this embodiment, the transmitting end receives the instructions from the intelligent scheduling controller, adjusts the working states of each resonant coil through the multi-frequency transmitting module according to the optimal working frequency combination, and simultaneously uses the adaptive tuning network to adjust the impedance characteristics of the transmitting end in real time to ensure that the resonance condition matches the selected frequency and distributes the corresponding power to each frequency channel.

[0190] In the implementation process of step S4, after the intelligent scheduling algorithm determines the frequency combination, it transmits instructions to the controller of the tuning network through a digital interface; based on the target frequency, the controller calculates the amount by which the current capacitance or inductance value needs to be adjusted to match the target frequency; by controlling the motor or electronic driver, the physical parameters of the variable capacitor and variable inductor are adjusted to achieve the calculated capacitance / inductance value; a sensor is used to monitor the tuned resonance frequency and compare it with the target frequency; if there is a deviation, further fine-tuning is performed through a closed-loop control algorithm until the desired frequency matching is achieved; after the system enters the stable state, the transmission efficiency and stability are continuously monitored to ensure that the tuning effect meets the performance indicators.

[0191] Step S5: The transmitting end transmits an alternating electromagnetic field to the receiving end according to the power distribution scheme; the receiving end captures the alternating electromagnetic field emitted by the transmitting end, absorbs energy through the resonant circuit, and converts the received electromagnetic energy into electrical energy.

[0192] Step S5 specifically includes:

[0193] According to the calculated frequency and the corresponding power distribution, for each calculated frequency f α , the frequency synthesizer generates a corresponding alternating current signal, expressed as:

[0194]

[0195] In the formula, I α (t) is the current of the α-th frequency component; A α is the amplitude; φ α is the phase angle; the calculated frequency is f 1 , f 2 ,..., f n ; the power distribution is P allocated,1 , P allocated,2 ,..., P allocated,n .

[0196] According to the power distribution scheme, the current signal of each frequency is amplified. Considering the amplification efficiency, the actual output power is expressed as:

[0197] P out,c = η PA,c ·P allocated,c

[0198] In the formula, P out,c is the actual output power after amplification of the c-th channel; η PA,c is the power amplifier efficiency of the c-th channel; the frequency signal is implicit in the power amplification process. After synthesizing the multi-frequency signal, the current signal of each frequency component needs to be amplified according to the power distribution scheme.

[0199] In the actual signal amplification process, the current signal I α (t) represents the time-domain expression of each frequency component, while P calculated in the power allocation scheme allocated,c defines the power level expected to be output for each frequency component; the working objective of the amplifier is to adjust its gain or working state according to P allocated,c so that the power of the amplified signal is as close as possible to this allocated value; this formula directly focuses on the relationship between the output power, the theoretical allocated power, and the amplifier efficiency. It implies the process of the amplifier processing the current signal I α (t), that is, by adjusting the gain to make the output power reach the expectation without directly changing the current amplitude expression I α (t) of the signal.

[0200] Before transmission, pre-distortion technology is adopted to ensure the correct waveform of the non-linearly amplified signal. The amplified signal needs to be precisely synchronized for transmission to ensure the superposition effect of multi-frequency signals. The purpose of waveform control and synchronous transmission is to ensure that all amplified frequency components can be perfectly superimposed in time and phase, based on the carefully designed frequency and power configuration in the previous steps, as well as the characteristics of the amplified signal.

[0201] The receiving end is equipped with multiple resonant circuits corresponding to the frequencies of the transmitting end. Each circuit is optimized for a specific frequency (i.e., f 1 , f 2 ,..., f n ). By adjusting the values of capacitors and inductors, it is ensured that the resonant frequency of the circuit matches the frequency of the received electromagnetic wave, thereby achieving maximum energy transfer.

[0202] Adjust the Q value of the resonant circuit to improve the energy absorption efficiency, expressed as:

[0203]

[0204] In the formula, ∝ is proportional to; f i is the transmission frequency; η abs,r is the absolute efficiency of the rth resonant circuit, an index measuring the energy conversion efficiency, that is, how much of the input energy is effectively utilized rather than lost; f 0,r is the center frequency of the rth resonant circuit, that is, the frequency point at which the resonant circuit is designed and optimized to work; Q r is the quality factor of the rth resonant circuit; the higher the quality factor, the better the selectivity of the circuit, the more effective the amplification of signals near the center frequency, and the better the suppression of signals far from the center frequency; a high Q value usually means higher working efficiency; f BW,r is the bandwidth of the rth resonant circuit, referring to the width of the frequency range within which the circuit can effectively respond; in a resonant system, the bandwidth is usually related to the degree of deviation from the center frequency. A narrow bandwidth means higher selectivity.

[0205] The above content describes the absolute efficiency η abs,r and the frequency offset f i -f o,r as well as the quality factor Q r and the bandwidth f BW,r are closely related; specifically, when the actual operating frequency f i is close to the resonance center frequency f 0,r , and the bandwidth f BW,r is relatively small (i.e., good selectivity), and at the same time the quality factor Q r is high, the expression in the denominator will be very small, resulting in the value of the entire fraction approaching 1, thereby making the efficiency η abs,r high; on the contrary, if the frequency deviates far from the center frequency, the bandwidth is wide, or the quality factor is low, the efficiency will decrease.

[0206] When the resonant circuit resonates with the electromagnetic wave, it can effectively absorb electromagnetic energy and convert it into current in the circuit; these currents are then sent to an energy converter (such as a rectifier bridge, DC-DC converter) to be converted into DC electrical energy to supply the load or for storage.

[0207] Converting the alternating current absorbed in the resonant circuit into direct current is achieved through a rectifier bridge circuit, expressed as:

[0208]

[0209] where η conv is the conversion efficiency; P DC is the converted DC power; P AC is the absorbed AC power.

[0210] The above content describes the energy conversion efficiency from AC to DC. The resonant circuit improves the energy absorption efficiency η r through an optimized quality factor Q abs,r , and the absorbed alternating current is then converted into direct current through circuits such as a rectifier bridge, directly depending on the quality of the absorbed AC electrical energy in the previous step, reflecting an effective conversion process from high-frequency electromagnetic energy to available DC electrical energy.

[0211] Monitor the reception efficiency and environmental changes through sensors, adjust the resonant frequency and circuit parameters in real time to maintain the best reception state; establish a communication link with the transmitter to exchange reception status information and further optimize the transmission parameters; the DC power after energy conversion becomes the basis for evaluating the system performance. Through sensor monitoring and feedback mechanisms, the receiver can adjust the resonant frequency and circuit parameters according to the actual conversion efficiency, and even communicate with the transmitter to optimize the transmission parameters; this feedback loop ensures that the system can dynamically adapt to environmental changes and continuously optimize the energy transmission efficiency.

[0212] Step S6: The receiver integrates the energy received from each frequency channel for combination and intelligently distributes the combined electric energy to the target load according to the load demand to ensure stable power supply.

[0213] Step S6 specifically includes:

[0214] The DC powers obtained after energy conversion for each channel are respectively P DC,1 , P DC,2 ,..., P DC,n . Power combination aggregates these independent DC powers into a single output power P total , expressed as:

[0215]

[0216] Power combination can be achieved in various ways, such as using a DC-DC converter (e.g., parallel Buck or Boost converters) or a simple parallel circuit, maintaining voltage matching and current balance, and avoiding power losses or safety issues caused by mismatches.

[0217] Evaluate the instantaneous power demand of the target load by monitoring the voltage V load and current I load of the load, and calculate the required power P Load = V load ×I Load .

[0218] According to P Load and the currently available P total , perform proportional distribution, that is, the power ratio contributed by each channel is in line with the total power demand, or consider factors such as maximizing efficiency, battery charge status (e.g., in energy storage applications), and load priority for distribution.

[0219] Monitor the difference between the power P actual actually allocated to the load and the demand P Load , and adjust the output of each channel accordingly to ensure stable power supply to the load.

[0220] Subsequently, the system continuously monitors the actual transmission efficiency and environmental changes, feeds the real-time data back to the intelligent scheduling algorithm, and dynamically adjusts parameters such as the operating frequency and power distribution as needed to form a closed-loop control to maintain the best transmission efficiency, specifically including:

[0221] Use sensors (such as power meters, efficiency measurement modules) to monitor the actual wireless power transmission efficiency in real time and evaluate the conversion and utilization efficiency of energy from the transmitter to the receiver; including but not limited to distance changes, the appearance or removal of obstacles, temperature and humidity changes, etc., which may affect the efficiency and stability of wireless power transmission.

[0222] The collected real-time data includes transmission efficiency, environmental parameter changes, load status (such as power demand fluctuations), etc., and is summarized to the central processor or cloud platform through the data acquisition system.

[0223] After receiving new data, the intelligent scheduling algorithm runs again. Based on the current environmental conditions, previous transmission efficiency data, and new load demands, it recalculates parameters such as the optimal frequency combination and power distribution. It involves complex optimization algorithms, such as genetic algorithms, particle swarm optimization, or machine learning models, to find the optimal solution under the current conditions.

[0224] According to the results recalculated by the algorithm, the system dynamically adjusts the operating frequency, power distribution scheme, and even the resonant circuit parameters (if necessary) of the transmitter; this includes adjusting the output of the transmit frequency synthesizer, power amplifier, and notifying the receiver to adjust the resonant circuit, etc.

[0225] After adjustment, the system enters the monitoring stage again, forming a continuous closed-loop feedback cycle. New efficiency and environmental data will be collected and compared with the adjusted system performance to evaluate the adjustment effect; if the expected performance is not achieved, the system will trigger the adjustment process again until the transmission efficiency is maximized.

[0226] Embodiment 2

[0227] As Figure 2 shown, the present invention discloses a wireless power transmission system, which includes:

[0228] An environmental parameter collection module 10 for collecting environmental parameters in the wireless power transmission path.

[0229] A transmission efficiency prediction module 20 for constructing a prediction model to predict the transmission efficiency of different frequency channels.

[0230] A power scheme allocation module 30 for calculating the optimal multi-frequency resonance frequency combination and the corresponding power distribution scheme.

[0231] An electrical component adjustment module 40 for the transmitter to dynamically adjust the inductance and capacitance element values according to the optimal operating frequency combination and the adaptive tuning network.

[0232] A transmitting-receiving matching module 50 for the transmitter to transmit an alternating electromagnetic field to the receiver according to the power distribution scheme; the receiver captures the alternating electromagnetic field emitted by the transmitter and absorbs energy through the resonant circuit.

[0233] An energy intelligent allocation module 60 for the receiver to integrate the energy received from each frequency channel for merging, and intelligently allocate the merged electric energy to the target load according to the load demand.

[0234] As an alternative implementation, the transmission efficiency prediction module 20 of the present invention specifically includes:

[0235] Define a feature extraction sub-network that receives an input layer and a list representing the number of neurons in each layer, and then creates a series of fully connected layers, each layer using the ReLU activation function.

[0236] Define an attention mechanism layer that converts the input into a probability distribution through a fully connected layer (using the Softmax activation); multiply these probabilities with the original input through a dot product operation.

[0237] Build a prediction model, create multiple input layers according to the input shape; create a feature extraction sub-network for each input layer; use a concatenation layer to merge the outputs of all sub-networks into a single vector; add an attention mechanism layer to dynamically emphasize the feature combinations that the model focuses on; use the concatenation layer again to combine the original merged features and the features processed by the attention mechanism; add a fully connected layer to further process these features; create an output layer for each frequency channel, using the Sigmoid activation function, and each output corresponds to the prediction of a frequency channel; compile the model, and the loss function is the mean squared error.

[0238] As an alternative implementation, the power scheme allocation module 30 of the present invention specifically includes:

[0239] Initialize a deep reinforcement learning agent class for learning and making decisions on the optimal configuration of frequency and power; define the state space size; define the action space size; use a deque to set an experience replay buffer for storing each interaction data; set the discount factor, initial exploration rate, minimum exploration rate, exploration rate decay rate, learning rate; build a deep Q-network model, implemented by a build function.

[0240] Build a neural network function for estimating the action value, the input is the current environmental state (such as transmission distance, attenuation), and the output is the value of each action; the model consists of three fully connected layers, the first and second layers use the ReLU activation function to increase the non-linear processing ability, and the last layer uses the linear activation function to output the expected value of each action (i.e., Q value).

[0241] Save the agent's experience (state, action, reward) to the memory replay, that is, store the data of one interaction in the experience replay buffer.

[0242] According to the action selected by the current policy (exploration or exploitation), when the random number is less than the current exploration rate, randomly select an action; otherwise, select the action corresponding to the maximum Q value predicted by the model as the expected optimal action.

[0243] Randomly extract a batch of experiences from memory for learning, optimize the network weights by updating the target Q-value, and minimize the mean squared error between the predicted Q-value and the target Q-value, specifically including:

[0244] Randomly extract a certain number of samples from the experience replay buffer; the samples contain the state during past exploration, the actions taken, the rewards obtained, the next state, and information on whether it is the end.

[0245] Process each extracted sample in a loop and unpack each element in the sample.

[0246] Calculate the target value. If this experience is a terminal state, the target value is the immediate reward obtained; if it is not a terminal state, the target value is the immediate reward plus the discounted value of the future expected reward.

[0247] Use the current model to predict the Q-values of all actions in a given state.

[0248] Assign the calculated target value to the Q-value corresponding to the action taken.

[0249] Use the updated target Q-value as a label to train the model once, optimize its weights to reduce the gap between the predicted Q-value and the target Q-value.

[0250] Update the initial exploration rate.

[0251] Output the instruction set, including the selected multi-frequency resonance frequency combination and the corresponding power distribution plan.

[0252] As an optional implementation manner, the electrical component adjustment module 40 of the present invention specifically includes:

[0253] After calculating the optimal frequency combination, send the instruction to the controller of the tuning network through the digital interface, and the instruction contains the target frequency f target ; The controller calculates the current capacitance C target according to the target frequency f current or the inductance L current adjustment amount to match the target frequency.

[0254] According to the calculated adjustment amount, the controller dynamically adjusts the physical parameters of the variable capacitor or variable inductor through motor control or electronic drive to reach the calculated capacitance C new or the inductance L new value.

[0255] Use a high-frequency precision frequency sensor to monitor the tuned resonance frequency f measured , and compare it with the target frequency f target . If there is a deviation e, adjust the capacitance or inductance through a closed-loop control algorithm, expressed as:

[0256]

[0257] Wherein, u(t) is the control signal; K b , K i and K d are the proportional, integral and differential gains respectively.

[0258] As an alternative implementation, the transmitting and receiving matching module 50 of the present invention specifically includes:

[0259] According to the calculated frequency and the corresponding power distribution, for each calculated frequency f α , the frequency synthesizer generates a corresponding alternating current signal, expressed as:

[0260]

[0261] Wherein, I α (t) is the current of the α-th frequency component; A α is the amplitude; φ α is the phase angle; the calculated frequencies are f 1 , f 2 ,..., f n ; the power distributions are P allocated,1 , P allocated,2 ,..., P allocated,n .

[0262] According to the power distribution scheme, amplify the current signal of each frequency. Considering the amplification efficiency, the actual output power is expressed as:

[0263] P out,c =η PA,c ·P allocated,c

[0264] Wherein, P out,c is the actual output power after amplification of the c-th channel; η PA,c is the power amplifier efficiency of the c-th channel.

[0265] Before transmission, pre-distortion technology is adopted. The receiving end is equipped with a plurality of resonant circuits corresponding to the frequencies of the transmitting end. Each circuit is optimized for a specific frequency. By adjusting the values of the capacitor and inductor, it is ensured that the resonant frequency of the circuit matches the frequency of the received electromagnetic wave.

[0266] Adjust the Q value of the resonant circuit to improve the energy absorption efficiency, expressed as:

[0267]

[0268] Wherein, ∝ is proportional to; f i is the transmission frequency; η abs,ris the absolute efficiency of the r-th resonant circuit; f 0,r is the center frequency of the r-th resonant circuit; Q r is the quality factor of the r-th resonant circuit; f BW,r is the bandwidth of the r-th resonant circuit.

[0269] Converting the alternating current absorbed in the resonant circuit into direct current is achieved through a rectifier bridge circuit, expressed as:

[0270]

[0271] In the formula, η conv is the conversion efficiency; P DC is the converted DC power; P AC is the absorbed AC power.

[0272] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A wireless power transmission method, characterized in that: The method comprises: Step S1: collecting environmental parameters in the wireless power transmission path; Step S2: construct a prediction model to predict the expected transmission efficiency under different frequency channels; Step S3: Calculate the optimal multi-frequency resonance frequency combination and the corresponding power allocation scheme; Step S4: The transmitter dynamically adjusts the values ​​of the inductor and capacitor components according to the optimal operating frequency combination and the adaptive tuning network; Step S5: the transmitting end transmits an alternating electromagnetic field to the receiving end according to the power allocation scheme; the receiving end captures the alternating electromagnetic field emitted by the transmitting end and absorbs energy through the resonant circuit; Step S6: The receiving end integrates the energy received from each frequency channel and combines it, and intelligently distributes the combined electric energy to the target load according to the load demand.

2. The wireless power transmission method according to claim 1, characterized in that: The constructing of a prediction model to predict the expected transmission efficiency under different frequency channels specifically includes: Define a feature extraction subnetwork that takes an input layer and a list of the number of neurons in each layer, and then creates a series of fully connected layers, each using a ReLU activation function. Define the attention mechanism layer, convert the input into a probability distribution through a fully connected layer (using Softmax activation); multiply these probabilities with the original input through a dot product operation; Build a prediction model and create multiple input layers based on the input shape; create a feature extraction subnetwork for each input layer; use a concatenation layer to merge the outputs of all subnetworks into a single vector; add an attention mechanism layer to dynamically emphasize the feature combinations that the model focuses on; use a concatenation layer again to combine the original merged features with the features processed by the attention mechanism; add a fully connected layer to further process these features; create an output layer for each frequency channel, use the Sigmoid activation function, and each output corresponds to the prediction of a frequency channel; compile the model, and the loss function is the mean squared error.

3. The wireless power transmission method according to claim 1, characterized in that: The calculation of the optimized multi-frequency resonance frequency combination and the corresponding power allocation scheme specifically includes: Initialize the deep reinforcement learning agent class for learning and decision-making on the optimal configuration of frequency and power; define the size of the state space; define the size of the action space; use a double-ended queue to set up the experience replay buffer to store each interaction data; set the discount factor, initial exploration rate, minimum exploration rate, exploration rate decay rate, and learning rate; build a deep Q network model, which is implemented by the construction function; Construct a neural network function for estimating the value of an action. The input is the current environment state (such as transmission distance, attenuation), and the output is the value of each action. The model consists of three fully connected layers. The first and second layers use the ReLU activation function to increase nonlinear processing capabilities. The last layer uses a linear activation function to output the expected value of each action (i.e., Q value). Save the agent's experience (state, action, reward) to memory replay, that is, store the data of an interaction in the experience replay buffer; The action selected according to the current strategy (exploration or exploitation) is selected randomly when the random number is less than the current exploration rate; otherwise, the action corresponding to the maximum Q value predicted by the model is selected as the expected optimal action; A batch of experiences are randomly extracted from memory for learning, and the network weights are optimized by updating the target Q value to minimize the mean square error between the predicted Q value and the target Q value, including: Randomly extract a certain batch of samples from the experience replay buffer; the samples contain the state of the past exploration, the actions performed, the rewards obtained, the next state, and whether it has ended; Loop through each extracted sample and unpack each element in the sample; Calculate the target value. If this experience is a terminal state, the target value is the immediate reward obtained; if it is not a terminal state, the target value is the discounted value of the immediate reward plus the expected future reward; Use the current model to predict the Q-values ​​of all actions in a given state; Assign the calculated target value to the Q value corresponding to the action taken; Use the updated target Q value as the label, train the model once, and optimize its weights to reduce the gap between the predicted Q value and the target Q value; Update the initial exploration rate; Outputting an instruction set including a selected multi-frequency resonant frequency combination and a corresponding power allocation plan.

4. The wireless power transmission method according to claim 1, wherein: The transmitting end dynamically adjusts the values ​​of the inductor and capacitor components according to the optimal operating frequency combination and the adaptive tuning network, specifically including: After calculating the best frequency combination, the command is sent to the controller of the tuning network through the digital interface. The command contains the target frequency f target ; The controller is based on the target frequency f target , calculate the current capacitance C current or inductance L current The amount of adjustment to match the target frequency; Based on the calculated adjustment amount, the controller dynamically adjusts the physical parameters of the variable capacitor or variable inductor through motor control or electronic drive to achieve the calculated capacitance C new or inductance L new value; Use a high-frequency precision frequency sensor to monitor the tuned resonant frequency f measured and with the target frequency f target By comparison, if there is a deviation e, the capacitance or inductance is adjusted through a closed-loop control algorithm, expressed as: Where u(t) is the control signal; K b , K i and K d are the proportional, integral and derivative gains respectively.

5. The wireless power transmission method according to claim 1, characterized in that: The transmitting end transmits an alternating electromagnetic field to the receiving end according to the power allocation scheme; the receiving end captures the alternating electromagnetic field emitted by the transmitting end and absorbs energy through a resonant circuit, specifically including: According to the calculated frequency and the corresponding power allocation, for each calculated frequency f α , the frequency synthesizer generates the corresponding alternating current signal, which is expressed as: In the formula, I α (t) is the current of the αth frequency component; A α is the amplitude; φ α is the phase angle; the calculated frequencies are f1, f2, ..., f n ; Power distribution is P allocated,1 , P allocated,2 , ..., P allocated,n ; According to the power allocation scheme, the current signal of each frequency is amplified. Considering the amplification efficiency, the actual output power is expressed as: P.S out,c η PA,c ·P allocated,c Where P out,c is the actual output power of the cth channel after amplification; η PA,c is the power amplifier efficiency of the cth channel; Pre-distortion technology is used before transmission. The receiving end is equipped with multiple resonant circuits corresponding to the transmitting end frequency. Each circuit is optimized for a specific frequency. By adjusting the values ​​of capacitors and inductors, the circuit resonant frequency is ensured to match the frequency of the received electromagnetic wave. Adjust the Q value of the resonant circuit to improve the energy absorption efficiency, expressed as: Where ∝ is proportional to; f i is the transmitting frequency; η abs,r is the absolute efficiency of the rth resonant circuit; f 0,r is the center frequency of the rth resonant circuit; Q r is the quality factor of the rth resonant circuit; f BW,r is the bandwidth of the rth resonant circuit; The AC power absorbed in the resonant circuit is converted into DC power through a rectifier bridge circuit, which can be expressed as: Where η conv is the conversion efficiency; P DC is the converted DC power; P AC is the absorbed AC power.

6. A wireless power transmission system, characterized in that: The system comprises: An environmental parameter collection module, used to collect environmental parameters in the wireless power transmission path; The transmission efficiency prediction module is used to build a prediction model to predict the transmission efficiency of different frequency channels; A power scheme allocation module is used to calculate the optimal multi-frequency resonance frequency combination and the corresponding power allocation scheme; An electrical component adjustment module is used for dynamically adjusting the values ​​of inductance and capacitance components at the transmitter according to the optimal operating frequency combination and the adaptive tuning network; The transmitting and receiving matching module is used for the transmitting end to transmit the alternating electromagnetic field to the receiving end according to the power allocation scheme; the receiving end captures the alternating electromagnetic field emitted by the transmitting end and absorbs energy through the resonant circuit; The energy intelligent distribution module is used to integrate the energy received from each frequency channel at the receiving end, and intelligently distribute the combined electric energy to the target load according to the load demand.

7. The wireless power transmission system according to claim 6, characterized in that: The transmission efficiency prediction module specifically includes: Define a feature extraction subnetwork that takes an input layer and a list of the number of neurons in each layer, and then creates a series of fully connected layers, each using a ReLU activation function. Define the attention mechanism layer, convert the input into a probability distribution through a fully connected layer (using Softmax activation); multiply these probabilities with the original input through a dot product operation; Build a prediction model and create multiple input layers based on the input shape; create a feature extraction subnetwork for each input layer; use a concatenation layer to merge the outputs of all subnetworks into a single vector; add an attention mechanism layer to dynamically emphasize the feature combinations that the model focuses on; use a concatenation layer again to combine the original merged features with the features processed by the attention mechanism; add a fully connected layer to further process these features; create an output layer for each frequency channel, use the Sigmoid activation function, and each output corresponds to the prediction of a frequency channel; compile the model, and the loss function is the mean squared error.

8. The wireless power transmission system according to claim 6, characterized in that: The power scheme allocation module specifically includes: Initialize the deep reinforcement learning agent class for learning and decision-making on the optimal configuration of frequency and power; define the size of the state space; define the size of the action space; use a double-ended queue to set up the experience replay buffer to store each interaction data; set the discount factor, initial exploration rate, minimum exploration rate, exploration rate decay rate, and learning rate; build a deep Q network model, which is implemented by the construction function; Construct a neural network function for estimating the value of an action. The input is the current environment state (such as transmission distance, attenuation), and the output is the value of each action. The model consists of three fully connected layers. The first and second layers use the ReLU activation function to increase nonlinear processing capabilities. The last layer uses a linear activation function to output the expected value of each action (i.e., Q value). Save the agent's experience (state, action, reward) to memory replay, that is, store the data of an interaction in the experience replay buffer; The action selected according to the current strategy (exploration or exploitation) is selected randomly when the random number is less than the current exploration rate; otherwise, the action corresponding to the maximum Q value predicted by the model is selected as the expected optimal action; A batch of experiences are randomly extracted from memory for learning, and the network weights are optimized by updating the target Q value to minimize the mean square error between the predicted Q value and the target Q value, including: Randomly extract a certain batch of samples from the experience replay buffer; the samples contain the state of the past exploration, the actions performed, the rewards obtained, the next state, and whether it has ended; Loop through each extracted sample to unpack each element in the sample; Calculate the target value. If this experience is a terminal state, the target value is the immediate reward obtained; if it is not a terminal state, the target value is the discounted value of the immediate reward plus the expected future reward; Use the current model to predict the Q-values ​​of all actions in a given state; Assign the calculated target value to the Q value corresponding to the action taken; Use the updated target Q value as the label, train the model once, and optimize its weights to reduce the gap between the predicted Q value and the target Q value; Update the initial exploration rate; Outputting an instruction set including a selected multi-frequency resonant frequency combination and a corresponding power allocation plan.

9. The wireless power transmission system according to claim 6, characterized in that: The electrical component adjustment module specifically includes: After calculating the best frequency combination, the command is sent to the controller of the tuning network through the digital interface. The command contains the target frequency f target ; The controller is based on the target frequency f target , calculate the current capacitance C current or inductance L current The amount of adjustment to match the target frequency; Based on the calculated adjustment amount, the controller dynamically adjusts the physical parameters of the variable capacitor or variable inductor through motor control or electronic drive to achieve the calculated capacitance C new or inductance L new value; Use a high-frequency precision frequency sensor to monitor the tuned resonant frequency f measured and with the target frequency f target By comparison, if there is a deviation e, the capacitance or inductance is adjusted through a closed-loop control algorithm, expressed as: Where u(t) is the control signal; K b , K i and K d are the proportional, integral and derivative gains respectively.

10. The wireless power transmission system according to claim 6, characterized in that: The transmitting and receiving matching module specifically includes: According to the calculated frequency and the corresponding power allocation, for each calculated frequency f α , the frequency synthesizer generates the corresponding alternating current signal, which is expressed as: In the formula, I α (t) is the current of the αth frequency component; A α is the amplitude; φ α is the phase angle; the calculated frequencies are f1, f2, ..., f n ; Power distribution is P allocated,1 , P allocated,2 , ..., P allocated,n ; According to the power allocation scheme, the current signal of each frequency is amplified. Considering the amplification efficiency, the actual output power is expressed as: P.S out,c η PA,c ·P allocated,c Where P out,c is the actual output power of the cth channel after amplification; η PA,c is the power amplifier efficiency of the cth channel; Pre-distortion technology is used before transmission. The receiving end is equipped with multiple resonant circuits corresponding to the transmitting end frequency. Each circuit is optimized for a specific frequency. By adjusting the values ​​of capacitors and inductors, the circuit resonant frequency is ensured to match the frequency of the received electromagnetic wave. Adjust the Q value of the resonant circuit to improve the energy absorption efficiency, expressed as: Where ∝ is proportional to; f i is the transmitting frequency; η abs,r is the absolute efficiency of the rth resonant circuit; f 0,r is the center frequency of the rth resonant circuit; Q r is the quality factor of the rth resonant circuit; f BW,r is the bandwidth of the rth resonant circuit; The AC power absorbed in the resonant circuit is converted into DC power through a rectifier bridge circuit, which can be expressed as: Where η conv is the conversion efficiency; P DC is the converted DC power; P AC is the absorbed AC power.

Citation Information

Patent Citations

  • Wireless charger transmission efficiency optimization method and system based on data analysis

    CN117992779A

  • Techniques For Delivering Pulsed Wireless Power

    US20240356376A1

Cited By

  • Wireless charging power adaptive adjustment method and system based on machine learning

    CN120824874A