A secure charging method and device for resisting attacks in a wireless charging network

By constructing charging and electromagnetic radiation models and combining them with reinforcement learning algorithms to generate defense strategies, the electromagnetic radiation safety and charging efficiency issues of wireless charging networks after being attacked are solved, achieving efficient and safe charging in attack scenarios.

CN119231776BActive Publication Date: 2025-11-04NANJING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411295186.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-16
Publication Date
2025-11-04
Estimated Expiration
2044-09-16

AI Technical Summary

Technical Problem

While wireless charging networks are vulnerable to attacks and cannot guarantee the safety of electromagnetic radiation and the effectiveness of charging, existing methods have failed to effectively address the risk of electromagnetic radiation exposure caused by malicious attackers manipulating wireless chargers.

Method used

By establishing a charging and electromagnetic radiation model that considers wave interference, and using a piecewise constant function to approximate nonlinear electromagnetic radiation, the two-dimensional region is divided into multiple equivalent electromagnetic radiation sub-regions. Combining primal dual optimization and a multi-agent deep reinforcement learning framework, a constrained reinforcement learning algorithm is proposed to generate a defense strategy to achieve Stackelberg equilibrium, ensuring electromagnetic radiation safety and maximizing charging utility.

Benefits of technology

After being attacked, it can effectively improve the overall charging efficiency and ensure electromagnetic radiation safety, which is significantly better than traditional methods, ensuring the efficiency and security of wireless charging networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119231776B_ABST
    Figure CN119231776B_ABST
Patent Text Reader

Abstract

The application discloses a kind of security charging method and device in wireless charging network to resist attack.The method comprises: constructing a charging and electromagnetic radiation model considering wave interference, to accurately calculate charging power and electromagnetic radiation distribution;Using piecewise constant function approximation nonlinear electromagnetic radiation, two-dimensional region is divided into multiple equivalent sub-regions, the maximum radiation point calculation method based on interference characteristics is used to simplify the electromagnetic radiation safety constraints of all points in the region into the finite constraints of maximum point, and the security charging problem of the point of attack in wireless charging network is changed into a double-layer optimization problem with finite constraints;Adopting dual learning method combined with reinforcement learning method realizes the defense charging strategy of optimal attack, and the strategy can realize Stackelberg equilibrium.The application effectively defends the attack on electromagnetic radiation safety, and the charging utility is improved by 40.48% on average compared with the comparative algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of electric energy transmission of wireless charging network, and more particularly to a method for ensuring the efficiency and safety of electric energy transmission in a wireless charging network with an adversarial scenario. BACKGROUND

[0002] Wireless Power Transmission (WPT) technology in wireless charging network refers to the technology that wireless chargers radiate electromagnetic waves to transmit energy to chargeable devices through air gap. With its reliability, non-contact and low maintenance, wireless power transmission technology has become a commercially viable charging method. However, during the process of electric energy transmission, the surrounding environment is inevitably exposed to a certain degree of electromagnetic radiation. High electromagnetic radiation exposure is identified as a threat to human health, which may cause risks such as tissue damage, cardiovascular disease and brain tumor. Therefore, a qualified wireless charging method must comply with electromagnetic radiation safety standards to ensure that the electromagnetic radiation intensity at any point does not exceed the safety threshold.

[0003] Currently, in order to provide stable energy supply, achieve full coverage of devices, and ensure charging continuity when charging devices fail, the deployment density of wireless chargers in wireless charging network is usually higher than the theoretical minimum value. Although this redundant design enhances the robustness of the network, it inadvertently provides an opportunity for attackers. Since wireless chargers lack physical tamper-proof measures, they are usually unattended, remotely placed and easily accessible. Malicious attackers can capture and manipulate these devices to disrupt the electromagnetic radiation safety constraints of the network and endanger biological safety by activating additional chargers. Existing wireless charging methods do not consider this adversarial charging scenario and are difficult to ensure safety after being attacked. SUMMARY

[0004] The present application aims to provide a safe charging method and device against attacks in wireless charging network to ensure the safety and actual effectiveness of electromagnetic radiation after the network is attacked.

[0005] In order to achieve the above application purposes, the technical solutions of the present application are as follows:

[0006] In the first aspect, a safe charging method against attacks in wireless charging network comprises the following steps:

[0007] (1) Establish a charging model and an electromagnetic radiation model considering wave interference according to the distribution of wireless charging network devices, wherein the charging model considers the phase change of electromagnetic waves radiated by different chargers when reaching the chargeable device due to the difference in path length to calculate the cumulative charging power provided by multiple chargers;

[0008] (2) based on the established charging model and electromagnetic radiation model, a mathematical model of the security charging problem of the wireless charging network under the attack is constructed, and the optimization goal of the security charging problem of the wireless charging network under the attack is to maximize the sum of the charging efficiency of all chargeable sensors, while ensuring the electromagnetic radiation safety constraint, that is, the total electromagnetic radiation at any point in the network does not exceed the predetermined safety threshold after suffering an attack on the electromagnetic radiation safety;

[0009] (3) the two-dimensional plane area of the wireless charging network is divided into multiple equivalent electromagnetic radiation sub-areas by using a piecewise constant function to approximate the nonlinear electromagnetic radiation, the maximum radiation value point of each equivalent electromagnetic radiation sub-area is calculated according to the interference characteristics, and the security charging problem of the wireless charging network under the attack is converted into a finite constraint double-layer optimization problem based on the radiation constraint of the maximum radiation value point;

[0010] (4) for the converted double-layer optimization problem, a constraint reinforcement learning algorithm is proposed to generate a defense strategy in combination with the original dual optimization and the multi-agent deep reinforcement learning framework, so as to realize the Stackelberg equilibrium to resist the optimal attack, in the constraint reinforcement learning algorithm, the defender and the attacker are regarded as agents executing the deep deterministic policy gradient, have a network for execution and a double evaluation network, use the Lagrange relaxation technology to convert the constraint optimization problem into an unconstrained problem, and use the primal-dual optimization method to iteratively update the strategy and the dual variable.

[0011] In a second aspect, a security charging strategy decision device for a wireless charging network under attack is provided, comprising:

[0012] The charging model and electromagnetic radiation model considering wave interference are established according to the distribution of the wireless charging network equipment, and the charging model considers the phase change of the electromagnetic waves radiated by different chargers when they reach the chargeable equipment due to the difference in path length to calculate the cumulative charging power provided by multiple chargers;

[0013] The security charging problem mathematical model construction module is used to construct a mathematical model of the security charging problem of the wireless charging network under the attack based on the established charging model and electromagnetic radiation model, and the optimization goal of the security charging problem of the wireless charging network under the attack is to maximize the sum of the charging efficiency of all chargeable sensors, while ensuring the electromagnetic radiation safety constraint, that is, the total electromagnetic radiation at any point in the network does not exceed the predetermined safety threshold after suffering an attack on the electromagnetic radiation safety;

[0014] The security charging problem mathematical model conversion module is configured to approximate the nonlinear electromagnetic radiation by using a piecewise constant function, divide a two-dimensional planar area of the wireless charging network into a plurality of equivalent electromagnetic radiation sub-areas, calculate a maximum radiation value point of each equivalent electromagnetic radiation sub-area according to interference characteristics, and convert the security charging problem of the wireless charging network against an attack into a double-layer optimization problem with limited constraints based on radiation constraints of the maximum radiation value point.

[0015] The defense strategy generation module based on constraint reinforcement learning is configured to propose a constraint reinforcement learning algorithm for generating a defense strategy against the converted double-layer optimization problem by combining original dual optimization and a multi-agent deep reinforcement learning framework, and realize a Stackelberg equilibrium to resist an optimal attack.

[0016] In a third aspect, the present application further provides a computer device, comprising: one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the programs realize the steps of the security charging method against an attack in a wireless charging network when executed by the processors.

[0017] In a fourth aspect, the present application further provides a computer readable storage medium having a computer program stored thereon, and the computer program realizes the steps of the security charging method against an attack in a wireless charging network when executed by a processor.

[0018] Advantages:

[0019] (1) The present application proposes a security charging method against an attack in a wireless charging network, which firstly constructs a charging and electromagnetic radiation model considering wave interference to accurately calculate charging power and electromagnetic radiation distribution, then approximates nonlinear electromagnetic radiation by using a piecewise constant function, divides a two-dimensional area into a plurality of equivalent sub-areas, simplifies electromagnetic radiation safety constraints of all points in the area into limited constraints of maximum value points by using a maximum radiation point calculation method based on interference characteristics, converts the security charging problem of the wireless charging network against an attack into a double-layer optimization problem with limited constraints, and finally realizes a defense charging strategy against an optimal attack by using a dual learning method combined with a reinforcement learning method, which can realize a Stackelberg equilibrium. The method determines the on-off state of a wireless charger to maximize overall charging utility and ensure electromagnetic radiation safety while coping with an attack.

[0020] (2) This invention combines primal dual optimization with a multi-agent deep reinforcement learning framework to propose a constrained reinforcement learning algorithm. The primal dual training method can effectively manage key constraints, optimize charging utility while ensuring electromagnetic radiation safety. The multi-agent deep reinforcement learning framework can capture the complex interaction between the defender and the attacker and formulate a robust defense strategy. This integration enables the algorithm to strive to achieve Stackelberg equilibrium while continuously optimizing performance, and finally determine the best defense strategy in the adversarial charging scenario. Attached Figure Description

[0021] Figure 1 A schematic diagram of a counter-charging scenario considering wave interference.

[0022] Figure 2 This is a diagram showing the electromagnetic radiation distribution of three wireless chargers.

[0023] Figure 3 A schematic diagram for calculating the maximum radiation value.

[0024] Figure 4 A framework diagram for constrained reinforcement learning algorithms.

[0025] Figure 5 A comparison of the charging efficiency of the method of the present invention with that of Bi-AC, PDO, Greedy, and Random methods under different attack counts k and electromagnetic radiation threshold Rt. Detailed Implementation

[0026] Due to the interference of electromagnetic waves, in the overlapping charging area of ​​multiple wireless chargers, the electromagnetic radiation received by the rechargeable device not only weakens with increasing distance, but also fluctuates significantly due to wave interference. Figure 1 This paper presents a wireless charging network in an adversarial charging scenario considering wave interference. It can be seen that due to the wave interference between the two chargers, constructive and destructive interference alternate, with constructive interference significantly enhancing electromagnetic radiation intensity and easily leading to high electromagnetic radiation exposure. Existing empirical formulas for calculating electromagnetic radiation deviate significantly from actual values, with errors reaching up to twice the actual values. Therefore, without considering interference effects, the calculated electromagnetic radiation intensity may be far lower than the actual exposure level, failing to ensure safety. To ensure electromagnetic radiation safety and practical effectiveness after an attack, this invention studies the safe charging problem in an adversarial charging scenario. The goal is to meet electromagnetic radiation safety constraints while maximizing overall charging efficiency in an adversarial charging scenario, activating the wireless charger to provide efficient and safe charging services for rechargeable devices in the network. By determining the on / off state of the wireless charger, it is possible to maximize overall charging efficiency and ensure electromagnetic radiation safety while responding to attacks.

[0027] There are three technical challenges for secure charging against adversarial attacks. The first challenge is that the secure charging problem is essentially an NP-hard problem, because of the infinite number of constraints on electromagnetic radiation at each point on the plane, and the non-convexity of the objective function hinders the direct application of classical optimization methods. Even if these constraints are simplified into a finite set, the problem is still a multi-dimensional 0 / 1 knapsack problem. The second challenge is that wave interference makes the secure charging problem more complex. While optimizing the charging utility and the safety constraints of electromagnetic radiation, the influence of wave interference on both must be considered. The cumulative charging utility and electromagnetic radiation of multiple chargers are not simply additive, but are affected by the complex interaction of their spatial arrangement. The third challenge is that the introduction of adversarial attacks turns the secure charging problem into a bi-level optimization problem, which is also an NP-hard problem. In addition, the problem of adversarial attacks and wave interference is tightly coupled and cannot be solved independently. Therefore, solving the secure charging problem against adversarial attacks requires dealing with multiple complex and interrelated NP-hard problems.

[0028] To this end, in terms of problem modeling and formalization, the present application establishes a charging and electromagnetic radiation model considering wave interference, accurately calculates the charging power and electromagnetic radiation distribution of the network, considers attacks against electromagnetic radiation safety, and formalizes the problem as a bi-level optimization problem. In terms of the infinite number of electromagnetic radiation safety constraints, the present application reduces the infinite electromagnetic radiation safety constraints to a finite number by discretizing the region and calculating the maximum radiation value point corresponding to each attack and defense strategy. In terms of defense strategy determination, the present application adopts a primal-dual optimization and multi-agent deep reinforcement learning framework, proposes a constrained reinforcement learning algorithm for generating defense strategies, and realizes a Stackelberg equilibrium to resist optimal attacks.

[0029] According to the secure charging method against attacks in the wireless charging network of the present application, the context of the wireless charging network is that there are a plurality of chargeable devices for monitoring the network in the plane, and there are also a plurality of wireless chargers with determined positions but undetermined switch states. There is also a malicious attacker in the network who can capture and manipulate a certain number of wireless chargers to change their switch states to violate the electromagnetic radiation safety constraints. The problem is how to determine the switch states of the wireless chargers to still satisfy the electromagnetic radiation safety constraints and maximize the overall charging utility after being attacked.

[0030] The secure charging method against adversarial attacks of the present application aims to satisfy the electromagnetic radiation safety constraints while maximizing the overall charging utility in the adversarial charging scenario, activates the wireless chargers, and provides efficient and safe charging services for the chargeable devices in the network, including the following steps:

[0031] (1) Establish a charging model and an electromagnetic radiation model considering wave interference, and then construct a mathematical model of the secure charging problem against adversarial attacks in the wireless charging network based on the established charging model and electromagnetic radiation model;

[0032] (2) Discretize the two-dimensional region into equivalent sub-regions by electromagnetic radiation approximation, and then calculate the maximum radiation value point according to the interference characteristics, so as to convert the charging problem into a double-layer optimization problem with limited constraints.

[0033] (3) Combining the original dual optimization and the multi-agent deep reinforcement learning framework, a constrained reinforcement learning algorithm is proposed to generate a defense strategy and realize the Stackelberg equilibrium to resist the optimal attack.

[0034] In step (1), the problem modeling and formal representation are performed. In the embodiment of the present application, the wave interference considering charging scenario is as follows: assuming that n identical wireless chargers and m identical chargeable sensors are located in a two-dimensional plane Omega, the wireless chargers uniformly radiate electromagnetic waves in all directions to charge the chargeable sensors, and s i and o j represent the positions of the wireless chargers s i and the chargeable sensors o j . There is also a malicious attacker in the network, who tries to cause high electromagnetic radiation exposure in some areas by turning on additional k wireless chargers.

[0035] The interference of waves will cause the electromagnetic radiation distribution in the network to present alternating enhanced intensity (constructive interference) and weakened intensity (destructive interference) regions. The shape and size of each region are different, please refer to the attached Figure 2 , which shows a complex electromagnetic radiation distribution radiated by three wireless chargers. The specific model construction method is as follows:

[0036] According to the widely accepted experience charging model, the charging model of a single wireless charger is established, assuming that the distance between the wireless charger s i and the chargeable sensor o j is d, and the charging power from the wireless charger s i to the chargeable sensor o j is given by the following formula:

[0037]

[0038] Wherein, A0 is the amplitude of the electromagnetic wave radiated by the wireless charger s i , is the path loss, a and b are two predetermined constants determined by the hardware of the charger or device and the surrounding environment, and D is the maximum charging distance of the wireless charger s i .

[0039] When a device is charged by multiple wireless chargers, the cumulative charging power received by the chargeable device is not the sum of the received power from multiple chargers, but is jointly affected by the amplitudes and phases of all electromagnetic waves reaching the device. The phase difference, i.e., the phase change of the electromagnetic waves radiated by different chargers when they reach the chargeable device due to the difference in path length: Whenever the path length differs by one wavelength λ w , the phase change goes through a period of 2π. The cumulative charging power from multiple chargers S i to the chargeable sensor o j is expressed as:

[0040]

[0041] For a device located within the charging area of multiple chargers, the charging power it receives is not only affected by path loss, but also by the interference effect of multiple chargers. As the distance increases, the charging power gradually decreases, but due to the existence of phase difference, the power may fluctuate up and down.

[0042] Since the charging utility and the electromagnetic radiation intensity are both proportional to the power received by the device, the charging utility model is established as follows:

[0043]

[0044] where x is the charger activation strategy, x i indicates whether the wireless charger s i is activated, C1 is a predetermined constant in the charging utility model, P ij and P kj represent the charging power transmitted by charger s i to sensor o j and o k , respectively, and U(x) is the charging utility obtained by charger activation strategy x.

[0045] Similarly, the electromagnetic radiation model at any position in the plane is established as follows:

[0046]

[0047] where C2 is a predetermined constant in the electromagnetic radiation model, p is any position in the two-dimensional plane, P ip and P kp represent the charging power transmitted by chargers s i and s k to point p, and e p (x) is the electromagnetic radiation intensity at point p.

[0048] The problem is formalized as follows. First, based on the established charging model and electromagnetic radiation model, the secure charging problem against attacks in a wireless charging network is defined. This invention models the secure charging problem against attacks as a Stackelberg game process, where the defender first determines the on / off state of the charger (Stage I), and then the attacker, after observing the defense strategy, activates an additional charger (Stage II) to compromise electromagnetic radiation security.

[0049] (P1)StageI(Defense):

[0050] Stage II (Attack):

[0051]

[0052] Where x and y represent n binary variables x i and y i Composed of defensive and offensive strategies, x i and y i These represent wireless chargers s i Whether it is activated / attacked. Therefore, the set of activated chargers S′={s i ∈S|x i ∨y i The network, i = 1, 2, ..., n, is jointly controlled by attack and defense strategies x and y. In a wireless charging network, the attack and defense state can be described as a two-phase Stackelberg game process, where the defender first determines the on / off state of the wireless chargers. The attacker, observing this strategy, intelligently turns on k additional wireless chargers, aiming to maximize the electromagnetic radiation peak and thus disrupt the network's electromagnetic radiation security constraints. k is the maximum number of wireless chargers the attacker can capture, and R... t This is a given electromagnetic radiation safety threshold. The objective of this invention is to determine the electromagnetic radiation safety threshold for each wireless charger using this two-phase Stackelberg game. i The on / off state ensures that, after a network attack, the electromagnetic radiation intensity at any location in the two-dimensional plane does not exceed a given safety threshold R. t In other words, the defense strategy must ensure that the network still meets electromagnetic radiation safety constraints after being attacked, while maximizing the overall charging efficiency of the devices.

[0053] In step (2), this invention utilizes region discretization technology and interference characteristics to obtain the maximum radiation value point corresponding to each attack and defense strategy. Please refer to the appendix. Figure 3 This is a schematic diagram illustrating the calculation of a maximum radiation value point as disclosed in an embodiment of the present invention. The specific steps of the method for calculating a maximum radiation value point disclosed in an embodiment of the present invention are as follows:

[0054] First, approximate the electromagnetic radiation intensity. Let e(d) denote the electromagnetic radiation when the charging distance is d, and use piecewise constant function Approximate the electromagnetic radiation:

[0055]

[0056] wherein is the approximation of e(d). l(0) = 0, l(J) = D, D is the maximum charging distance, l(q) = β((1+∈) q / 2 -1), (q = 1, 2,..., Q-1), β and ∈ are constants, and s i is the center of the circle, and concentric circles with radii l(1), l(2),..., l(Q) are drawn, wherein l(q) represents the radius of the qth concentric circle of the charger, and Q represents the total number of concentric circles. Thus, the two-dimensional plane is divided into several equivalent electromagnetic radiation sub-regions, as shown in (a) of FIG. 1. Figure 3

[0057] Second, draw the constructive interference curve corresponding to each attack and defense strategy by using the interference characteristics. Specifically, the electromagnetic waves radiated by the wireless chargers will affect the electromagnetic radiation intensity in the network due to the interference effect. When the phase difference of two or more chargers is an even multiple of half the wavelength, constructive interference is formed; when the phase difference is an odd multiple of half the wavelength, destructive interference is formed. According to the attack and defense strategy, the on-off state of the charger is determined, and the constructive interference curve of all the open chargers determined by the attack and defense strategy is drawn to show the points in the network where the electromagnetic radiation is significantly enhanced due to the interference effect, as shown in (b) of FIG. 1. Figure 3

[0058] Third, in each equivalent electromagnetic radiation sub-region, the intersection points of the constructive interference curves are calculated to identify several points that are significantly enhanced in the path attenuation process due to the interference effect, and the electromagnetic radiation intensity of these points is calculated to finally determine the point with the highest intensity as the maximum radiation value point, as shown by the red star in (c) of FIG. 1. Figure 3

[0059] Thus, the safety constraints of the electromagnetic radiation of all points in the two-dimensional plane of problem P1 are converted into safety constraints of the electromagnetic radiation of the maximum radiation value point, i.e. Thus, P1 with an infinite number of constraint conditions is equivalently converted into P2 with a limited number of constraint conditions:

[0060]

[0061] In step (3), the present application determines the defense strategy against the optimal attack based on the original dual optimization and reinforcement learning algorithm. Please refer to the attached Figure 4 ​​​A framework of a constraint reinforcement learning algorithm against attacks is disclosed in the embodiments of the present application. The constraint reinforcement learning algorithm against attacks disclosed in the embodiments of the present application has the following specific steps:

[0062] In the wireless charging network, the decision of the defender and the attacker strategy is structured as a Markov game model Wherein, is the state space, is the discrete action space of the agent i, is the joint action space of the agents, is the probability transition mechanism. The agent i obtains the reward and generates the cost (defined by the reward function and the cost function respectively) according to its state and behavior. The goal of the agent i is to maximize its total expected return and minimize its total expected cost where γ∈[0,1] is the discount factor, used to balance long-term and short-term returns, and T is the time horizon.

[0063] In the established constraint reinforcement learning algorithm, the defender and the attacker are respectively agent 1 and agent 2, with the execution network and the double evaluation network. The agent adopts the approximate deterministic strategy π i approximated by the neural network parameter θ i . The double evaluation network includes the reward evaluation network Q i (s,a|φ i ) and the cost review network C i (s,a|ξ i ). Wherein, a is the joint action of the defender and the attacker agent.

[0064] According to the characteristics of the Stackelberg game, the agent decides and collects experience in the following order:

[0065]

[0066] s.t.C1(s′,a1,a2)≤b1.

[0067] a′2←π2(s′,a1′;θ2),

[0068] s.t.C2(s′,a1′,a2)≤b2,

[0069] Wherein, b1 and b2 are the maximum threshold values of the respective constraints of the defender and the attacker. Specifically, the strategy of the defender aims to select the action that maximizes the charging utility increment, while ensuring that the electromagnetic radiation intensity of the maximum radiation value point does not exceed the specified threshold R tThe attacker's goal is to maximize the peak electromagnetic radiation of the network within a finite number of attacks, k. To ensure that subsequent attack and defense strategies reach a Tucklberg equilibrium, the defender considers all possible combinations of defense and attack when making decisions.

[0070] The constrained optimization problem is transformed into the following unconstrained problem using the Lagrange relaxation technique:

[0071]

[0072] Among them, L i (π,λ) is the Lagrange penalty function, which combines the cost term with the original long-run reward. The Lagrange multiplier λ dynamically balances the trade-off between reward and cost, ensuring that the constraints are met without affecting the overall objective.

[0073] The agent's policy π and dual variable λ are iteratively updated using a primal-dual optimization method:

[0074]

[0075] Where, α i and β i The learning rates for the policy and dual variables, θ, are respectively. i For policy network parameters, For the objective function L i (π,λ) for policy parameter θ i The gradient guides the optimization direction of the policy. Furthermore, b is the threshold of the cost constraint, [x] + =max{0,x} is the projection onto the dual space.

[0076] Agent from experience replay buffer Randomly select a batch of empirical samples Based on these samples, the policy gradient can be estimated by calculating the expected value, as shown in the following formula:

[0077]

[0078] At the same time, the agent uses the reward r obtained from the experience samples. i and cost c i Update the evaluation network parameters, evaluate each action, and update the parameters of the reward evaluation network by minimizing the temporal difference error:

[0079]

[0080] Where the target value y i for:

[0081] The parameters of the cost evaluation network are also updated by minimizing the temporal difference error:

[0082]

[0083] where the target value g i is:

[0084] To improve the stability of learning, the parameters of the target network (including the policy network π i ', the reward value network Q i ', and the cost value network C i ') are soft-updated using a delay parameter: θ' <- τθ + (1 - τ)θ i ', φ' <- τφ + (1 - τ)φ i ', ξ' <- τξ + (1 - τ)ξ i ', where τ is a small constant that controls the update rate of the target network parameters.

[0085] The constrained reinforcement learning algorithm adopts a centralized training and decentralized execution framework. During execution, the defender agent maximizes the Lagrangian penalty function based on local observations, and then the attacker agent responds optimally after observing the defender's strategy. The agents collect experiences through interactions with the environment and store them in an experience replay buffer When the environment reaches the termination state (i.e., the defender and attacker strategies are determined), it will be reset. During training, the agents optimize the strategies by sharing experiences. The evaluation network calculates the temporal difference error based on the reward, cost, and observation values and updates the parameters. The agents update the policy parameters using sampled policy gradients and adjust the Lagrange multipliers to balance reward maximization and cost minimization. The algorithm stabilizes the training process by iteratively updating the target network parameters. These steps enable the agents to learn the optimal charger activation strategy, effectively defend against attacks, and maximize charging utility. The pseudo-code of the algorithm is as follows.

[0086] Algorithm 1: Constrained reinforcement learning algorithm for solving the secure charging problem against adversarial attacks

[0087] Initialize: policy network parameters θ = [θ i ] i∈{1,2} , reward and cost evaluation network parameters φ = [φ i ] i∈{1,2} , and ξ = [ξ i ] i∈{1,2} ;

[0088] Initialize the target network parameters to be equal to the main network parameters: θ' <- θ, φ' <- φ, ξ' <- ξ;

[0089] Initialize the Lagrange multipliers: λ = [λ i ];

[0090] 1 For each episode e = 1,..., M: % Experience collection phase

[0091] 2 Each agent i e {1, 2} observes state s, takes action a i , and gets reward r i and cost c i ;

[0092] 3 Execute action a and observe next state s';

[0093] 4 Store (s, a i , r i , c i , s') to replay buffer

[0094] 5 If s' is a terminal state, reset environment state;

[0095] 6 For each step t = 1,..., H: % Parameter update phase

[0096] 7 Randomly sample a batch of transition tuples from

[0097] 8 Compute target reward and cost: y i = r i + γQ i '(s', a1', a'2), g i = c i + γC i '(s', a1', a'2);

[0098] 9 Update reward evaluation network by minimizing loss:

[0099] 10 Update cost evaluation network by minimizing loss:

[0100] 11 Update policy using sampled policy gradient:

[0101] 12 Update Lagrange multipliers:

[0102] 13 Update target networks: θ' i ← τθ i + (1 - τ)θ' i , φ' i ← τφ i + (1 - τ)φ' i , ξ' i ← τξ i+(1-τ)ξ′ i .

[0103] To verify the performance of the method of the present invention, comparative experiments were conducted. Please refer to the experimental results. Figure 5 The present invention's method (CARL) and the two-layer actor-critic method (Bi-AC), primal dual optimization method (PDO), greedy method (Greedy), and random method (Random) respectively achieve the following results in the number of attacks k ( Figure 5 (a) and electromagnetic radiation safety threshold R t ( Figure 5 Experimental comparison under changes in (b)). Except for the changed parameters, the default parameters were set as follows: n = 20, m = 50, D = 4m, a = 10, b = 10, C1 = 0.01, C2 = 1, H = 1000, γ = 0.9. Experimental results show that the safe charging method against attacks in this invention is significantly better than the Bi-AC, PDO, Greedy, and Random methods. On average, it is superior to the Bi-AC, PDO, Greedy, and Random methods in terms of the number of attacks k and the electromagnetic radiation safety threshold R. t In terms of performance, this invention outperforms the other four comparative methods by 34.66% and 46.3%, respectively. The invention maintains high charging efficiency in adversarial scenarios, demonstrating excellent robustness.

[0104] Based on the same technical concept as the method embodiment, another embodiment of the present invention also provides a secure charging strategy decision-making device for combating attacks in a wireless charging network, comprising:

[0105] A module for constructing charging and electromagnetic radiation models considering wave interference is used to establish charging and electromagnetic radiation models considering wave interference based on the distribution of wireless charging network devices. The charging model considers the phase change of electromagnetic waves radiated by different chargers when they reach the rechargeable device due to differences in path length, and calculates the cumulative charging power provided by multiple chargers.

[0106] The mathematical model construction module for the safe charging problem is used to construct a mathematical model for the safe charging problem against attacks in a wireless charging network based on the established charging model and electromagnetic radiation model. The optimization objective of the safe charging problem against attacks in the wireless charging network is to maximize the sum of the charging utility of all rechargeable sensors while ensuring electromagnetic radiation safety constraints. The electromagnetic radiation safety constraints refer to the fact that after being attacked against electromagnetic radiation safety, the total electromagnetic radiation at any point in the network does not exceed a predetermined safety threshold.

[0107] The security charging problem mathematical model conversion module is configured to approximate the nonlinear electromagnetic radiation by using a piecewise constant function, divide a two-dimensional planar region of the wireless charging network into a plurality of equivalent electromagnetic radiation sub-regions, calculate a maximum radiation value point of each equivalent electromagnetic radiation sub-region according to interference characteristics, and convert the security charging problem of the wireless charging network against the attack into a double-layer optimization problem with limited constraints based on radiation constraints of the maximum radiation value point.

[0108] The defense strategy generation module based on the constraint reinforcement learning is configured to propose a constraint reinforcement learning algorithm for generating a defense strategy for the converted double-layer optimization problem by combining the original dual optimization and a multi-agent deep reinforcement learning framework, and realize a Stackelberg equilibrium to resist the optimal attack, in which the defender and the attacker are regarded as agents executing a deep deterministic policy gradient, have a network execution and a double evaluation network, the constraint optimization problem is converted into an unconstrained problem by using a Lagrange relaxation technology, and the policy and the dual variable are iteratively updated by using an original-dual optimization method.

[0109] It should be understood that the security charging strategy decision device against the attack in the wireless charging network in the embodiments of the present application can realize all the technical solutions in the above method embodiments, and the functions of each functional module can be specifically realized according to the methods in the above method embodiments, and the specific implementation process can be referred to the related description in the above embodiments, which will not be described here.

[0110] The present application also provides a computer device, comprising: one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the program is executed by the processor to realize the steps of the security charging method against the attack in the wireless charging network as described above.

[0111] The present application also provides a computer readable storage medium having a computer program stored thereon, wherein the computer program is executed by the processor to realize the steps of the security charging method against the attack in the wireless charging network as described above.

[0112] Those skilled in the art will understand that the embodiments of the present application can be provided as a method, device (system), computer device or computer program product. Therefore, the present application can adopt a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer usable program code.

[0113] The application is described with reference to flowcharts of methods according to embodiments of the application. It will be understood that each block of the flowchart, and combinations of blocks in the flowchart, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flowchart block or blocks. Figure 1 The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks.

[0114] These computer program instructions can also be stored in a computer readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer readable memory produce an article of manufacture including instructions which implement the function specified in the flowchart block or blocks. Figure 1 The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks.

[0115] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks.

[0116] Finally, it should be noted that the above-described embodiments are merely intended to illustrate the technical solutions of the present application, but not to limit the same. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that the specific embodiments of the present application can be modified or equivalent replacements without departing from the spirit and scope of the present application, and any modifications or equivalent replacements shall be included in the protection scope of the claims of the present application.

Claims

1. A secure charging method against attacks in a wireless charging network, characterized in that, Includes the following steps: (1) Based on the distribution of wireless charging network devices, establish a charging model and an electromagnetic radiation model that consider wave interference. The charging model considers the phase change of electromagnetic waves radiated by different chargers when they reach the rechargeable device due to differences in path length, and calculates the cumulative charging power provided by multiple chargers. Assume there are n identical wireless chargers and m identical chargeable sensors located in a two-dimensional plane Ω, the charging model is represented as follows: where x is the charger activation strategy, x i represents whether the wireless charger s i is activated, C1 is a constant predetermined in the charging model; P ij and P kj respectively represent the charging power transmitted by the charger s i to the sensor o j and o k , and U(x) is the charging utility obtained by the charger activation strategy x; Δφ is the phase change of the electromagnetic waves radiated by different chargers when reaching the chargeable device due to the difference in path length. The electromagnetic radiation model at any position p in the plane is represented as follows: where C2 is a predetermined constant in the electromagnetic radiation model, P ip and P kp respectively represent the charging power transmitted to the point p by the charger s i and s k respectively, e p (x) is the electromagnetic radiation intensity at the point p; (2) Based on the established charging model and electromagnetic radiation model, a mathematical model for the secure charging problem against attacks in a wireless charging network is constructed. The optimization objective of the secure charging problem against attacks in the wireless charging network is to maximize the sum of the charging utilities of all rechargeable sensors while ensuring electromagnetic radiation safety constraints. The electromagnetic radiation safety constraints refer to the fact that after being attacked against electromagnetic radiation safety, the total electromagnetic radiation at any point in the network does not exceed a predetermined safety threshold. The secure charging problem against attacks in the wireless charging network is as follows: where x and y are charger activation strategy and attack strategy, the charger activation strategy is taken as the defense strategy, x and y are composed of n binary variables x i and y i , respectively, representing whether the wireless chargers s i are activated / attacked, the set of activated chargers S' = {s i ∈ S | x i ∨ y i , i = 1, 2, …, n} is jointly controlled by x and y; U(x, y) is the charging utility under the x and y strategy, e p (x, y) is the electromagnetic radiation intensity at point p under the x and y strategy; problem P1 describes the attack-defense state in the wireless charging network as a Stackelberg game process, the defender Defense determines the activation strategy of the chargers in the first stage Stage I, and the attacker Attack intelligently turns on additional k wireless chargers after observing the strategy in the second stage Stage II, aiming to maximize the electromagnetic radiation peak to break the electromagnetic radiation safety constraint of the network, R t is a given electromagnetic radiation safety threshold; (3) By approximating nonlinear electromagnetic radiation using a piecewise constant function, the two-dimensional planar region of the wireless charging network is divided into multiple equivalent electromagnetic radiation sub-regions. The maximum radiation value point of each equivalent electromagnetic radiation sub-region is determined based on interference characteristics. Based on the radiation constraint of the maximum radiation value point, the secure charging problem against attacks in the wireless charging network is transformed into a finite-constrained bi-level optimization problem, expressed as: wherein, represents the maximum radiant point electromagnetic radiation; (4) For the transformed bi-layer optimization problem, a constrained reinforcement learning algorithm is proposed to generate defense strategies and realize Stackelberg equilibrium to counter the optimal attack by combining the original dual optimization and the multi-agent deep reinforcement learning framework. In the constrained reinforcement learning algorithm, the defender and the attacker are regarded as agents that execute the gradient of the deep deterministic policy. They have an execution network and a dual evaluation network. The constrained optimization problem is transformed into an unconstrained problem by using the Lagrange relaxation technique. The original-dual optimization method is used to iteratively update the policy and dual variables.

2. The method according to claim 1, characterized in that, Methods for determining the point of maximum radiation include: The first step is to approximate the electromagnetic radiation intensity, using e(d) to represent the electromagnetic radiation at a charging distance of d, and employing a piecewise constant function to approximate the electromagnetic radiation intensity: in For an approximation of e(d), with each wireless charger s i Draw concentric circles with center β and radii l(1), l(2), ..., l(Q), where l(q) represents the radius of the q-th concentric circle of the charger, Q represents the total number of concentric circles, and l(q) = β((1+ω)). q / 2 -1), β and ∈ are constants; The second step is to use the interference characteristics to plot the constructive interference curve corresponding to each attack and defense strategy. Specifically, the switching state of the charger is determined according to the attack and defense strategy, and all constructive interference curves of the chargers are plotted to show the points in the network where electromagnetic radiation is significantly enhanced due to the interference effect. The third step is to calculate the intersection points of the constructive interference curves in each equivalent electromagnetic radiation sub-region to identify several points that are significantly enhanced due to the interference effect during the path attenuation process, and to calculate the electromagnetic radiation intensity of these points, and determine the point with the highest intensity as the maximum radiation value point.

3. The method according to claim 1, characterized in that, The constrained reinforcement learning algorithm adopts a centralized training and distributed execution framework. During execution, the defender agent maximizes the Lagrange penalty function based on local observations, while the attacker agent responds after observation. The agent collects experience by interacting with the environment and stores it in the experience replay buffer. The algorithm is reset when the environment reaches the termination state, i.e., the policy is determined. During training, the agent shares experience to optimize the policy, and the evaluation network updates parameters through temporal difference error; the agent uses policy gradient to update policy parameters and adjusts Lagrange multipliers to balance rewards and costs. The training process iteratively updates the target network parameters, enabling the agent to learn the optimal charger activation strategy, effectively defend against optimal attacks, and maximize overall charging efficiency.

4. The method according to claim 1, characterized in that, The constraint reinforcement learning algorithm includes: constructing the attack and defense strategy decision as a Markov game model. <S,A i ,P,r i ,c i ,γ>, where S is the state space, A i Let i be the discrete action space of agent i. Let P be the joint action space of the agents, where S×A i →[0,1] represents the probability transition mechanism, where agent i receives rewards and incurs costs based on its state and behavior; agent i's goal is to maximize its total expected return. and minimizing its total expected cost in γ∈[0,1] is the discount factor, and T is the time range; The agent employs a neural network parameterized by θ i An approximate deterministic policy π i ; the double evaluation network includes a reward evaluation network Q i (s, a | φ i ) and a cost review network C i (s, a | ξ i ), wherein a is the joint action of the agent; the goal of the agent is to maximize the expected return under a limited cost budget b i , and its decision-making process is as follows: stC1(s′,a1,a2)≤b1. a′2←π2(s′,a′1;θ2), stC2(s′,a1′,a2)≤b2. The constrained optimization problem is transformed into the following unconstrained problem using the Lagrange relaxation technique: where L i (π,λ) is a Lagrangian penalty function that combines the cost term with the original long-term return, and the Lagrangian multiplier λ dynamically balances the trade-off between the reward and the cost, ensuring that the constraints are satisfied without affecting the overall objective. The primal-dual optimization method is used to iteratively update the policy π and the dual variable λ: Where, α i and β i The learning rates are the policy and dual variables, respectively. For the objective function L i (π,λ) for policy parameter θ i The gradient of [x], where b is the threshold of the cost constraint, and [x] is the gradient of [x]. + =max{0,x} is the projection onto the dual space; The policy gradient of an agent is defined as: Where B represents an empirical data point extracted from the playback buffer D, and D records the empirical samples of all agents; The parameters of the reward evaluation network are updated by minimizing the temporal difference error: Where the target value y i for: The parameters of the cost evaluation network are updated as follows: Where the target value g i for: The parameters of the target network are soft-updated using a delay parameter: θ' <- τθ + (1 - τ)θ' i φ' <- τφ + (1 - τ)φ' i ξ' <- τξ + (1 - τ)ξ' i where τ is a constant that controls the rate of update of the target network parameters.

5. A secure charging strategy decision-making device for countering attacks in a wireless charging network, characterized in that, Its application in the secure charging method against attacks in a wireless charging network as described in claim 1 includes: A module for constructing charging and electromagnetic radiation models considering wave interference is used to establish charging and electromagnetic radiation models considering wave interference based on the distribution of wireless charging network devices. The charging model considers the phase change of electromagnetic waves radiated by different chargers when they reach the rechargeable device due to differences in path length, and calculates the cumulative charging power provided by multiple chargers. The mathematical model construction module for the safe charging problem is used to construct a mathematical model for the safe charging problem against attacks in a wireless charging network based on the established charging model and electromagnetic radiation model. The optimization objective of the safe charging problem against attacks in the wireless charging network is to maximize the sum of the charging utility of all rechargeable sensors while ensuring electromagnetic radiation safety constraints. The electromagnetic radiation safety constraints refer to the fact that after being attacked against electromagnetic radiation safety, the total electromagnetic radiation at any point in the network does not exceed a predetermined safety threshold. The mathematical model conversion module for the safe charging problem is used to approximate nonlinear electromagnetic radiation using a piecewise constant function, divide the two-dimensional planar region of the wireless charging network into multiple equivalent electromagnetic radiation sub-regions, calculate the maximum radiation value point of each equivalent electromagnetic radiation sub-region based on the interference characteristics, and transform the safe charging problem against attacks in the wireless charging network into a two-layer optimization problem with finite constraints based on the radiation constraints of the maximum radiation value point. A constraint-based reinforcement learning-based defense policy generation module is proposed to address the transformed bi-layer optimization problem. Combining primal-dual optimization and a multi-agent deep reinforcement learning framework, a constraint-based reinforcement learning algorithm is developed to generate defense policies and achieve Stackelberg equilibrium to counter optimal attacks. In this algorithm, the defender and attacker are treated as agents executing deep deterministic policy gradients, and the algorithm includes an execution network and a dual evaluation network. Lagrange relaxation is used to transform the constrained optimization problem into an unconstrained problem, and the primal-dual optimization method is used to iteratively update the policy and dual variables.

6. A computer device, characterized in that, include: One or more processors; Memory; And one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, wherein when the programs are executed by the processors, they implement the steps of the secure charging method against attacks in a wireless charging network as claimed in any one of claims 1-4.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the secure charging method against attacks in a wireless charging network as described in any one of claims 1-4.

Citation Information

Patent Citations

  • Robustly safe recharging scheduling method in wireless rechargeable sensor network

    CN108509742A

  • Adversarial sample detection method and electronic equipment

    CN110321790A