An energy efficiency optimization method for resisting intelligent full-duplex attack in cooperation of IRS and UAV

By employing an energy efficiency optimization method that integrates IRS and UAV, and utilizing a Stackelberg game model and a multi-agent algorithm to optimize the collaborative strategy of UAV and IRS, the complex threat of intelligent full-duplex attacks is addressed, thereby improving the security and energy efficiency of the communication system.

CN121125184BActive Publication Date: 2026-05-15ZHEJIANG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHEJIANG UNIV OF TECH
Filing Date
2025-08-21
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing technologies fail to fully exploit the collaborative potential of multiple nodes in UAV-IRS-ground jammers, neglecting the dynamic game relationship between legitimate parties and intelligent unauthorized attackers. This makes it difficult for traditional security mechanisms to cope with the complex threats of intelligent full-duplex attacks, and makes it difficult to solve multi-dimensional strongly coupled variables in real time and efficiently.

Method used

A collaborative energy efficiency optimization method for IRS and UAV is constructed. The dynamic attack and defense interaction between the legitimate and illegitimate parties is simulated through the Stackelberg game model. The MATD3-PER algorithm and AO algorithm are used to optimize the flight trajectory of the safe UAV, the user transmission power, the IRS phase shift and the ground jammer beam. The projected gradient descent method is combined to optimize the flight trajectory and jamming power of the UAV attacker, so as to achieve the optimal global energy efficiency.

Benefits of technology

It effectively resists intelligent full-duplex attacks, significantly enhances the physical layer security of communication systems, achieves adaptive adaptation and energy efficiency optimization of attack and defense strategies, and solves the optimization problem of strong coupling of multi-dimensional variables.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121125184B_ABST
    Figure CN121125184B_ABST
Patent Text Reader

Abstract

The application discloses an energy efficiency optimization method for resisting intelligent full-duplex attacks in IRS and UAV cooperation, which is applied to a communication system containing legal parties and illegal parties, and the legal parties include a safe UAV, an IRS, a user and a ground jammer. The energy efficiency optimization method for resisting intelligent full-duplex attacks in IRS and UAV cooperation fully releases the multi-device cooperation potential by constructing a safe UAV-IRS-ground jammer multi-node cooperation architecture, can effectively resist the eavesdropping and jamming double threats of intelligent full-duplex attackers, and significantly enhances the physical layer security of the communication system. A Stackelberg game model is introduced to simulate the dynamic attack and defense interaction of the legal parties and the illegal parties, the strategy iteration of both parties converges to an equilibrium state, the adaptive adaptation of attack and defense strategies and the global energy efficiency optimization in the confrontation scene are realized, and on the algorithm, the legal parties adopt the MATD3-PER algorithm and the AO algorithm, the illegal parties adopt the MATD3-PER algorithm and the projection gradient descent method, and the optimization problem of multi-dimensional variable strong coupling is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of physical layer security technology, specifically relating to an energy efficiency optimization method for resisting smart full-duplex attacks through the collaboration of IRS and UAV. Background Technology

[0002] Unmanned Aerial Vehicles (UAVs), with their high mobility and rapid deployment capabilities, have become a core component of air-to-ground wireless communication networks, playing a crucial role in emergency communications and coverage of remote areas. Intelligent Reconfigurable Surfaces (IRS), as an emerging technology, can reconstruct the wireless signal propagation environment by dynamically adjusting the phase of a large number of passive reflective elements, effectively overcoming channel fading and improving communication quality. Co-deploying IRS with UAVs can build an intelligent and efficient communication architecture, providing important support for next-generation communication systems such as 6G.

[0003] However, open airspace communication faces severe security challenges. Full-duplex attackers can simultaneously execute eavesdropping and jamming attacks. On the one hand, they exploit channel reciprocity to approach legitimate communication links and illegally intercept sensitive information; on the other hand, they actively transmit jamming signals to legitimate receivers, undermining communication reliability. Meanwhile, complex ground environments, such as building obstructions and multipath effects, cause drastic fluctuations in the quality of legitimate link channels, further weakening the system's anti-jamming capabilities and making traditional security mechanisms inadequate to cope with such complex threats.

[0004] Existing technical solutions have significant shortcomings. Most studies optimize single elements in isolation, failing to fully explore the collaborative potential of multiple nodes in the "UAV-IRS-Ground Jammer" system. They also neglect the dynamic game relationship between legitimate and intelligent unauthorized attackers, lacking a systematic attack and defense optimization framework. Furthermore, the multi-dimensional, strongly coupled variables make it difficult for traditional algorithms to achieve real-time and efficient solutions. Summary of the Invention

[0005] The purpose of this invention is to address the problems raised in the background art by proposing an energy efficiency optimization method for resisting intelligent full-duplex attacks through the collaboration of IRS and UAV.

[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0007] This invention proposes an energy efficiency optimization method for resisting intelligent full-duplex attacks through IRS and UAV collaboration. This method is applied to a communication system containing both legitimate and illegitimate parties. The legitimate parties include a secure UAV, an IRS, users, and a ground-based jammer. The illegitimate parties include a full-duplex UAV attacker. Users communicate with the secure UAV via the IRS, the UAV attacker eavesdrops on user information and interferes with the secure UAV via the IRS, and the ground-based jammer simultaneously interferes with the UAV attacker. The energy efficiency optimization method for resisting intelligent full-duplex attacks through IRS and UAV collaboration includes:

[0008] S1. Calculate the secure communication rate of each user, and calculate the energy efficiency of the legitimate party and the energy efficiency of the illegitimate party based on the secure communication rate of the users.

[0009] S2. Construct the objective function for maximizing the average energy efficiency of the legal side, and construct the objective function for maximizing the average energy efficiency of the non-legal side;

[0010] S3. The objective function for maximizing the average energy efficiency of the legitimate party is decomposed into four first sub-problems that optimize the flight trajectory of the safe UAV, the user transmission power, the IRS phase shift, and the beam of the ground jammer, respectively. The objective function for maximizing the average energy efficiency of the illegitimate party is decomposed into two second sub-problems that optimize the flight trajectory of the UAV attacker and the jamming target, and optimize the jamming power of the UAV attacker, respectively.

[0011] S4. Establish a Stackelberg game model. Solve the four first subproblems for the legitimate side first, and then solve the two second subproblems for the illegitimate side. The legitimate side uses the MATD3-PER algorithm to find the safe UAV flight trajectory, and the AO algorithm to solve the subproblems of optimizing user transmission power, optimizing IRS phase shift, and optimizing the beam of ground jammers. The illegitimate side uses the MATD3-PER algorithm to solve the subproblems of optimizing the flight trajectory and jamming target of the UAV attacker, and the projected gradient descent algorithm to solve the subproblem of optimizing the interference power of the UAV attacker. Finally, apply the optimal solutions of the six subproblems to the communication system to achieve UAV energy consumption optimization.

[0012] Preferably, each user is represented as , The total number of users, and each security UAV is represented as , The total number of secure UAVs, and the attacker of each UAV is represented as , The total number of UAV attackers, represented by each IRS. , The total number of IRSs, and the number of antennas of ground jammers are: The preset time period of the communication system is divided into There are 1 time slot, and each time slot is represented as The reflection elements of each IRS are represented as follows: , Let be the number of reflection elements in each IRS. In time slot t, the reflection coefficient matrix of each IRS is expressed as: , It is a complex number. The imaginary unit, For the first The first IRS The phase shift of each reflecting element in response to the incident signal. The dimension is Complex matrix;

[0013] In time slot t, the channel gains from each user to the secure UAV, from each user to the IRS, and from each IRS to the secure UAV are expressed as follows: , and Then the composite channel gain from each user to the secure UAV is , Indicates conjugate transpose;

[0014] In time slot t, the channel gains from each user to the UAV attacker and from each IRS to the UAV attacker are expressed as follows: and The composite channel gain from each user to the UAV attacker is expressed as: ;

[0015] In time slot t, the channel gains from each UAV attacker to the secure UAV and from the ground jammer to the UAV attacker are expressed as follows: and .

[0016] Preferably, the step of calculating the secure communication rate of each user and calculating the energy efficiency of the legitimate party and the energy efficiency of the non-legitimate party based on the user's secure communication rate includes:

[0017] The formula for calculating the secure communication rate for each user is as follows:

[0018] ;

[0019] in,

[0020] ;

[0021] ;

[0022] ;

[0023] in, This represents the secure communication rate of each user in time slot t. This represents the communication rate between each user and the secure UAV in time slot t. This represents the rate at which each UAV attacker eavesdrops on each user in time slot t. Indicates user and IRS The correlation coefficient, , Indicates user via IRS Communicate with secure UAV, otherwise This means that the user Not through IRS For secure UAV communication, each user is associated with its nearest IRS, and each user is associated with only one IRS. , Indicates the user in time slot t With security UAV The correlation coefficient, , Indicates user With security UAV Communication, otherwise This means that the user Not related to safe UAV Communication involves each user communicating with the nearest secure UAV covering their area. , Indicates the user in time slot t Transmission power, This indicates that the attacker was in time slot tUAV. For safe UAV Interference factors, This indicates a UAV attacker. Towards secure UAV Send interference signals, otherwise This means UAV attackers Not to secure UAV Sending interference signals, subject to constraints , This indicates that the attacker was in time slot tUAV. For safe UAV Interference power, Indicates a safe UAV noise power, This indicates the self-interference efficiency of a UAV attacker. This indicates ground-based jammers targeting UAV attackers. Interference beams, Indicates UAV attacker The noise power;

[0024] The formula for calculating the energy efficiency of a legal entity is as follows:

[0025] ;

[0026] in, This represents the energy efficiency of the legal side in time slot t. This represents the power balance factor for each user's transmission. This represents the flight power consumption balance factor for each safety UAV. This represents the flight power consumption of each safe UAV in time slot t. The power balance factor representing the ground-based interferator. The norm of a vector;

[0027] The formula for calculating the energy efficiency of non-legal entities is as follows:

[0028] ;

[0029] in, This represents the energy efficiency of the non-legal party in time slot t. This represents the interference power balance factor for each UAV attacker. This represents the flight power consumption balance factor for each UAV attacker. This represents the flight power consumption of each UAV attacker in time slot t.

[0030] Preferably, the construction of the objective function for maximizing the average energy efficiency of the legal side and the construction of the objective function for maximizing the average energy efficiency of the non-legal side include:

[0031] The formula for the objective function that maximizes the average energy efficiency of the legal side is as follows:

[0032] P1: ;

[0033] st C1.1: , ;

[0034] C1.2: , ;

[0035] C1.3: , ;

[0036] C1.4: , ;

[0037] C1.5: , ;

[0038] C1.6: , ;

[0039] C1.7: , , ;

[0040] in,

[0041] ;

[0042] ;

[0043] ;

[0044] ;

[0045] Where C1.1 is the constraint on the IRS phase shift, and C1.2 is the constraint on the user's transmission power. For users The maximum transmission power, C1.3 is the total power constraint of the ground jammer. C1.4 is the maximum power of the ground-based jammer and the flight power consumption constraint for a safe UAV. For safe UAV The total flight power, C1.5 and C1.6 are the safe flight speeds of the UAV in time slot t, respectively. and angle Restrictions, For safe UAV The maximum flight speed, C1.7 is the safe flight range limit for UAVs. and These are secure UAVs The horizontal and vertical axes of the flight range in time slot t, and These are secure UAVs The x-coordinate and y-coordinate of the maximum flight range. This represents the set of flight trajectories of all safe UAVs across all time slots. Indicates a safe UAV In the three-dimensional coordinates of time slot t, and , Indicates a safe UAV Flight altitude This represents the set of transmission power for all users across all time slots. This represents the set of phase shifts for all IRSs across all time slots. This represents the set of interference beams from ground-based jammers across all time slots.

[0046] The formula for maximizing the average energy efficiency of non-legal parties is as follows:

[0047] P2: ;

[0048] st C2.1: ;

[0049] C2.2: ;

[0050] C2.3: , ;

[0051] C2.4: ;

[0052] C2.5: ;

[0053] C2.6: ;

[0054] C2.7: ;

[0055] in,

[0056] ;

[0057] ;

[0058] ;

[0059] Among them, C2.1 and C2.2 are constraints on the interference factors of the UAV attacker, and C2.3 is a constraint on the flight power consumption of the UAV attacker. For UAV attackers The total flight power, C2.4 is the interference power constraint for UAV attackers. and C1 and C2.6 represent the minimum and maximum interference power of the UAV attacker, respectively, and C2.5 and C2.6 represent the flight speed of the UAV attacker in time slot t, respectively. and angle Restrictions, For UAV attackers The maximum flight speed, C2.7, is the flight range limit for UAV attackers. and These are UAV attackers The horizontal and vertical axes of the flight range in time slot t, and These are UAV attackers The x and y coordinates of the maximum flight range in time slot t. This represents the set of flight trajectories of all UAV attackers across all time slots. Indicates UAV attacker In the three-dimensional coordinates of time slot t, and , Indicates UAV attacker Flight altitude This represents the set of interference power from all UAV attackers across all time slots. This represents the set of interference factors from all UAV attackers across all time slots.

[0060] Preferably, the objective function for maximizing the average energy efficiency of the legitimate party is decomposed into four first sub-problems that optimize the safe UAV flight trajectory, user transmission power, IRS phase shift, and ground jammer beam, respectively. The objective function for maximizing the average energy efficiency of the illegitimate party is decomposed into two second sub-problems that optimize the UAV attacker's flight trajectory and the jamming target, and optimize the UAV attacker's jamming power, respectively.

[0061] When optimizing the flight trajectory of a safe UAV, the three optimization quantities in the other three first sub-problems are fixed, thus obtaining the sub-problems for optimizing the flight trajectory of a safe UAV:

[0062] P1.1: ;

[0063] st C1.4-C1.7;

[0064] The three first subproblems of optimizing user transmit power, IRS phase shift, and beamwidth for ground interference are transformed into single-slot optimization problems:

[0065] For each time slot, when optimizing user transmission power, the three optimization quantities in the other three first sub-problems are fixed, thus obtaining the sub-problems for optimizing user transmission power:

[0066] P1.2: ;

[0067] st C1.2;

[0068] For each time slot, when optimizing the IRS phase shift, the three optimization quantities in the other three first sub-problems are fixed, thus obtaining the sub-problems for optimizing the IRS phase shift:

[0069] P1.3: ;

[0070] st C1.1;

[0071] For each time slot, when optimizing the beam for ground jammers, the three optimization quantities in the other three first sub-problems are fixed, thus obtaining the sub-problems for optimizing the beam for ground jammers:

[0072] P1.4: ;

[0073] st C1.3;

[0074] When optimizing the flight trajectory of the UAV attacker and the interference target, the optimization amount of another second sub-problem is fixed, thus obtaining the sub-problem for optimizing the flight trajectory of the UAV attacker and the interference target:

[0075] P2.1: ;

[0076] st C2.1-C2.3, C2.5-C2.7;

[0077] For each time slot, when optimizing the interference power of a UAV attacker, the optimization amount of another second sub-problem is fixed, thus obtaining the sub-problem for optimizing the interference power of a UAV attacker:

[0078] P2.2: ;

[0079] st C2.4.

[0080] Preferably, when establishing the Stackelberg game model, the secure UAV and the UAV attacker constitute the Stackelberg game model, with the secure UAV acting as the leader and the UAV attacker acting as the leader's follower.

[0081] To solve the four first subproblems, we first solve the subproblem of optimizing the flight trajectory of the safe UAV. Each safe UAV is defined as an agent, referred to as the first agent, and the state space of each first agent in time slot t is: , Indicates a safe UAV In the time slot The three-dimensional coordinates , Indicates a safe UAV In the time slot The remaining flight power, defining the action space of each first agent in time slot t as... Define that all first agents have the same reward function in time slot t, and that is... ,in The penalty for not satisfying constraint C1.7, ,in The penalty factor representing the safe UAV flying out of the range limit boundary;

[0082] A first Actor-Critic neural network is introduced for each first agent, and each first Actor-Critic neural network includes a first Actor current network, a first Actor target network, two first Critic current networks, and two first Critic target networks.

[0083] Preferably, for solving the two second sub-problems, the first sub-problem of optimizing the flight trajectory of the UAV attacker and the interference target is solved. Each UAV attacker is defined as an agent, referred to as the second agent, and the state space of each second agent in time slot t is as follows: ,in Indicates UAV attacker In the time slot The three-dimensional coordinates, and , Indicates UAV attacker In the time slot The remaining flight power, This indicates that the UAV attacker was in the time slot. The interference target is defined as the action space of each second agent in time slot t. ,Bundle Relaxation is the key to having Constrained continuous variables, interval Divided into equal parts There are several sub-intervals, each corresponding to a safe UAV except for the first interval. The reward function for all second agents in time slot t is defined to be the same, and is... ,in The penalty for not satisfying constraint C2.7, ,in The penalty factor representing the safe UAV flying out of the range limit boundary;

[0084] A second Actor-Critic neural network is introduced for each second agent, and each second Actor-Critic neural network includes a second Actor current network, a second Actor target network, two second Critic current networks, and two second Critic target networks.

[0085] Preferably, from time slot 1 to The MATD3-PER algorithm trains the first Actor-Critic neural network of the first agent:

[0086] S4.1 Initialize the initial state space of each first agent, and set the time slots... ;

[0087] S4.2. For each first agent, obtain the current state space. And the set of current state spaces of all first agents is The current state space of each first agent is used as the input to the current network of the corresponding first actor, and the current network of each first actor outputs the action space of the first agent. ;

[0088] S4.3 Each first agent executes its own action space, and the set of action spaces of all first agents is: ;

[0089] S4.4, Start executing the AO algorithm;

[0090] S4.4.1. With fixed IRS phase shift and ground interference beam, a subproblem for optimizing user transmission power is solved using a convex optimization method.

[0091] S4.4.2. With fixed user transmission power and ground interference beam, solve the subproblem of optimizing IRS phase shift using SDR and SCA methods;

[0092] S4.4.3. With fixed user transmission power and IRS phase shift, the SCA method is used to solve the subproblem of beam optimization for ground interference.

[0093] S4.5. After each first agent performs an action, the corresponding reward is represented as follows: The set of rewards for all first agents is And observe the state space of each first agent in the next time slot, and the set of the state spaces of all first agents in the next time slot is: ;

[0094] S4.6, The first experiences of each first agent Store it in the first cache, and each time it is stored, set the priority of the first experience stored at the current time to the highest priority in the first cache;

[0095] S4.7 When the first buffer is full, the first Actor-Critic neural network is updated, i.e., S4.7.1 is executed; if it is not full, S4.8 is executed.

[0096] S4.7.1, from the first buffer according to A batch of first-hand experience samples, among which Sample the first intelligent agent u to the first The probability of an experience. For the first intelligent agent u, the first Prioritizing first experience, among which Let `x` be the summation index variable for the first buffer, representing the iteration through all the first experiences stored in the first buffer. [0,1] is used to control the adjustment coefficients for random sampling and greedy sampling, and is based on probability. Calculate the first agent u's first... The importance weight of each first experience is: ,in, The size of the first buffer. To offset the impact of the priority experience replay method on the convergence results;

[0097] S4.7.2 For each sampled first experience, calculate the first target action, the first target Q value, the current first Q value, and the first Critic loss respectively. The first Critic loss is calculated using the importance weight of the first experience. Use gradient descent to update the two first Critic current networks.

[0098] S4.7.3 When the preset update time is reached, update the current network of the first Actor: first calculate the gradient, and then use gradient ascent to update the current network of the first Actor;

[0099] S4.7.4 For each first agent, update the parameters of its own first Actor target network and the two first Critic target networks through soft updates;

[0100] S4.7.5 Calculate the TD error for each first agent corresponding to the sampled first experience. and according to The size of the first empirical sampled priority is updated, and The larger the value, the higher the priority.

[0101] S4.8, Order Repeat steps S4.2-S4.8 until... Then, one round of training of the first Actor-Critic neural network for each first agent is completed.

[0102] After each pair of first agent's first Actor-Critic neural networks completes one round of training, the second agent's second Actor-Critic neural network is trained for one round.

[0103] Preferably, from time slot 1 to The MATD3-PER algorithm trains the second Actor-Critic neural network for the second agent:

[0104] S4.10. Initialize the initial state space of each second agent, and set the time slots... ;

[0105] S4.11. For each second agent, obtain the current state space. And the set of current state spaces of all second agents is The current state space of each second agent is used as the input to the current network of the corresponding second actor, and the current network of each second actor outputs the action space of the second agent. ;

[0106] S4.12, Each second agent executes its own action space, and the set of action spaces of all second agents is: ;

[0107] S4.13. The subproblem of optimizing the interference power of UAV attackers is solved using the projected gradient descent algorithm.

[0108] S4.14. After each second agent completes its action, the corresponding reward is represented as follows: The reward set for all second agents is And observe the state space of each second agent in the next time slot, and the set of the state spaces of all second agents in the next time slot is: ;

[0109] S4.15, the second experience of the second intelligent agent Store it in the second cache, and each time it is stored, set the priority of the currently stored second experience to the highest priority in the second cache;

[0110] S4.16 When the second buffer is full, the second Actor-Critic neural network is updated, i.e., S4.16.1 is executed; if it is not full, S4.17 is executed.

[0111] S4.16.1, from the second buffer according to A second set of samples was taken, among which Two intelligent agents Sampling to the The probability of an experience. As a second intelligent agent The The priority of the second experience, among which Let `x` be the summation index variable for the second buffer, representing the summation of all second experiences stored in the second buffer according to probability. Computational Second Agent The The importance weight of each second experience is: ,in, This is the size of the second buffer.

[0112] S4.16.2 For each sampled second experience, calculate the second target action, the second target Q value, the current second Q value, and the second Critic loss respectively. The second Critic loss is calculated using the importance weight of the second experience. Use gradient descent to update the current network of the two second Critics.

[0113] S4.16.3 When the preset update time is reached, update the current network of the second Actor: first calculate the gradient, and then use gradient ascent to update the current network of the second Actor;

[0114] S4.16.4 For each second agent, update the parameters of its respective second actor target network and the two second critic target networks through soft updates;

[0115] S4.16.5 Calculate the TD error for each sampled second agent corresponding to the second experience. and according to The size of the sampled second empirical priority is updated, and The larger the value, the higher the priority.

[0116] S4.17, Order Repeat steps S4.11-S4.17 until... Then, one round of training of the second Actor-Critic neural network for each second agent is completed;

[0117] S4.18. Enter the next round of training for the first Actor-Critic neural network and the second Actor-Critic neural network. Repeat S4.1-S4.18 continuously until the preset number of rounds is reached, then the training of the first Actor-Critic neural network and the second Actor-Critic neural network is completed.

[0118] Preferably, after training the first Actor-Critic neural network and the second Actor-Critic neural network is completed, for each first agent:

[0119] S4.19, will The current state space of the time slot is used as the input of the first actor's current network in the first actor-critic neural network, to obtain the action space of the first agent in the current time slot. Based on the action space of the first agent in the current time slot, the three-dimensional coordinates of the safe UAV in the current time slot are calculated. Then, the AO algorithm is used to obtain the optimal user transmission power, optimal IRS phase shift and optimal ground jammer beam in the current time slot.

[0120] S4.20, Order Repeat steps S4.19-S4.20 until... Then we get 1 to The three-dimensional coordinates of the time-slotted safe UAV are used to obtain the flight trajectory of the safe UAV, and this flight trajectory is from 1 to... The optimal flight trajectory of a safe UAV within a time slot, and the results obtained from 1 to... Within each time slot, the optimal user transmission power, optimal IRS phase shift, and optimal ground interfering beam are achieved.

[0121] For each second agent:

[0122] S4.21, will The current state space of the time slot is used as the input of the current network of the second actor in the trained second actor-critic neural network to obtain the action space of the second agent in the current time slot. The three-dimensional coordinates of the UAV attacker in the current time slot are calculated based on the action space of the second agent in the current time slot. The optimal interference target of the UAV attacker in the current time slot is calculated based on the action space of the second agent in the current time slot. The optimal interference power of the current time slot is obtained by using the projection gradient descent algorithm.

[0123] S4.22, Order Repeat steps S4.21-S4.22 until... Then we get 1 to The three-dimensional coordinates of the time-slot UAV attacker are obtained, thus revealing the UAV attacker's flight trajectory, which is defined as 1 to... The optimal flight trajectory of a UAV attacker within the time slot, and the results of 1 to... Within each time slot, the optimal target for interference by a UAV attacker, and the result of 1 to... Within a time slot, the optimal interference power for a UAV attacker in each time slot.

[0124] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0125] This energy efficiency optimization method for resisting intelligent full-duplex attacks through IRS and UAV collaboration fully unleashes the potential of multi-device collaboration by constructing a multi-node collaborative architecture of secure UAV-IRS-ground jammer. It can effectively resist the dual threats of eavesdropping and interference from intelligent full-duplex attackers, significantly enhancing the physical layer security of the communication system. By introducing a Stackelberg game model to simulate the dynamic attack and defense interaction between the legitimate and illegitimate parties, the strategies of both parties converge to an equilibrium state through iterative convergence, achieving adaptive adaptation of attack and defense strategies and global energy efficiency optimization in adversarial scenarios. In terms of algorithms, the legitimate party adopts the MATD3-PER algorithm and the AO algorithm, while the illegitimate party adopts the MATD3-PER algorithm and the projective gradient descent method, solving the optimization problem of strong coupling of multi-dimensional variables. Attached Figure Description

[0126] Figure 1 This is a flowchart illustrating the energy efficiency optimization method for resisting intelligent full-duplex attacks using IRS and UAV collaboration according to the present invention.

[0127] Figure 2 This is a structural block diagram of the communication system of the present invention;

[0128] Figure 3 This is a convergence graph of the cumulative reward values ​​of the legitimate and illegitimate parties during the training process of this invention. Detailed Implementation

[0129] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0130] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to limit the invention.

[0131] In one embodiment, such as Figures 1-2 As shown, an energy efficiency optimization method for resisting smart full-duplex attacks through IRS and UAV collaboration is provided. It is applied to a communication system containing legitimate and illegitimate parties. The legitimate parties include a secure UAV, an IRS, a user, and a ground jammer. The illegitimate parties include a full-duplex UAV attacker. In this method, the user communicates with the secure UAV through the IRS, the UAV attacker eavesdrops on user information and interferes with the secure UAV through the IRS, and the ground jammer simultaneously interferes with the UAV attacker.

[0132] First, it should be noted that each user is represented as , The total number of users, and each security UAV is represented as , The total number of secure UAVs, and the attacker of each UAV is represented as , The total number of UAV attackers, represented by each IRS. , The total number of IRSs, and the number of antennas of ground jammers are: The preset time period of the communication system is divided into There are 1 time slot, and each time slot is represented as The reflection elements of each IRS are represented as follows: , Let be the number of reflection elements in each IRS. In time slot t, the reflection coefficient matrix of each IRS is expressed as: , It is a complex number. The imaginary unit, For the first The first IRS The phase shift of each reflecting element in response to the incident signal. The dimension is Complex matrix;

[0133] In time slot t, the channel gains from each user to the secure UAV, from each user to the IRS, and from each IRS to the secure UAV are expressed as follows: , and Then the composite channel gain from each user to the secure UAV is , Indicates conjugate transpose;

[0134] In time slot t, the channel gains from each user to the UAV attacker and from each IRS to the UAV attacker are expressed as follows: and The composite channel gain from each user to the UAV attacker is expressed as: ;

[0135] In time slot t, the channel gains from each UAV attacker to the secure UAV and from the ground jammer to the UAV attacker are expressed as follows: and .

[0136] The energy efficiency optimization method for resisting smart full-duplex attacks through the collaboration of IRS and UAV includes:

[0137] S1. Calculate the secure communication rate for each user, and calculate the energy efficiency of the legitimate party and the energy efficiency of the non-legitimate party based on the user's secure communication rate, including:

[0138] The formula for calculating the secure communication rate for each user is as follows:

[0139] ;

[0140] in,

[0141] ;

[0142] ;

[0143] ;

[0144] in, This represents the secure communication rate of each user in time slot t. This represents the communication rate between each user and the secure UAV in time slot t. This represents the rate at which each UAV attacker eavesdrops on each user in time slot t. Indicates user and IRS The correlation coefficient, , Indicates user via IRS Communicate with secure UAV, otherwise This means that the user Not through IRS For secure UAV communication, each user is associated with its nearest IRS, and each user is associated with only one IRS. , Indicates the user in time slot t With security UAV The correlation coefficient, , Indicates user With security UAV Communication, otherwise This means that the user Not related to safe UAV Communication involves each user communicating with the nearest secure UAV covering their area. , Indicates the user in time slot t The transmission power (where the transmission power from the user to the secure UAV is the same as the transmission power from the user to the UAV attacker). This indicates that the attacker was in time slot tUAV. For safe UAV Interference factors, This indicates a UAV attacker. Towards secure UAV Send interference signals, otherwise This means UAV attackers Not to secure UAV Sending interference signals, subject to constraints , This indicates that the attacker was in time slot tUAV. For safe UAV Interference power, Indicates a safe UAV noise power, This indicates the self-interference efficiency of a UAV attacker. This indicates ground-based jammers targeting UAV attackers. Interference beams, Indicates UAV attacker noise power, , For intermediate parameters;

[0145] The formula for calculating the energy efficiency of a legal entity is as follows:

[0146] ;

[0147] in, This represents the energy efficiency of the legal side in time slot t. This represents the power balance factor for each user's transmission. This represents the flight power consumption balance factor for each safety UAV. This represents the flight power consumption of each safe UAV in time slot t. The power balance factor representing the ground-based interferator. The norm of a vector;

[0148] The formula for calculating the energy efficiency of non-legal entities is as follows:

[0149] ;

[0150] in, This represents the energy efficiency of the non-legal party in time slot t. This represents the interference power balance factor for each UAV attacker. This represents the flight power consumption balance factor for each UAV attacker. This represents the flight power consumption of each UAV attacker in time slot t.

[0151] S2. Construct the objective function for maximizing the average energy efficiency of the legal sides, and construct the objective function for maximizing the average energy efficiency of the non-legal sides, including:

[0152] The formula for the objective function that maximizes the average energy efficiency of the legal side is as follows:

[0153] P1: ;

[0154] st C1.1: , ;

[0155] C1.2: , ;

[0156] C1.3: , ;

[0157] C1.4: , ;

[0158] C1.5: , ;

[0159] C1.6: , ;

[0160] C1.7: , , ;

[0161] in,

[0162] ;

[0163] ;

[0164] ;

[0165] ;

[0166] Where C1.1 is the constraint on the IRS phase shift, and C1.2 is the constraint on the user's transmission power. For users The maximum transmission power, C1.3 is the total power constraint of the ground jammer. C1.4 is the maximum power of the ground-based jammer and the flight power consumption constraint for a safe UAV. For safe UAV The total flight power, C1.5 and C1.6 are the safe flight speeds of the UAV in time slot t, respectively. and angle Restrictions, For safe UAV The maximum flight speed, C1.7 is the safe flight range limit for UAVs. and These are secure UAVs The horizontal and vertical axes of the flight range in time slot t, and These are secure UAVs The x-coordinate and y-coordinate of the maximum flight range. This represents the set of flight trajectories of all safe UAVs across all time slots. Indicates a safe UAV In the three-dimensional coordinates of time slot t, and , Indicates a safe UAV The flight altitude (the same for any time slot safe UAV). This represents the set of transmission power for all users across all time slots. This represents the set of phase shifts for all IRSs across all time slots. This represents the set of interference beams from ground-based jammers across all time slots.

[0167] The formula for maximizing the average energy efficiency of non-legal parties is as follows:

[0168] P2: ;

[0169] st C2.1: ;

[0170] C2.2: ;

[0171] C2.3: , ;

[0172] C2.4: ;

[0173] C2.5: ;

[0174] C2.6: ;

[0175] C2.7: ;

[0176] in,

[0177] ;

[0178] ;

[0179] ;

[0180] Among them, C2.1 and C2.2 are constraints on the interference factors of the UAV attacker, and C2.3 is a constraint on the flight power consumption of the UAV attacker. For UAV attackers The total flight power, C2.4 is the interference power constraint for UAV attackers. and C1 and C2.6 represent the minimum and maximum interference power of the UAV attacker, respectively, and C2.5 and C2.6 represent the flight speed of the UAV attacker in time slot t, respectively. and angle Restrictions, For UAV attackers The maximum flight speed, C2.7, is the flight range limit for UAV attackers. and These are UAV attackers The horizontal and vertical axes of the flight range in time slot t, and These are UAV attackers The x and y coordinates of the maximum flight range in time slot t. This represents the set of flight trajectories of all UAV attackers across all time slots. Indicates UAV attacker In the three-dimensional coordinates of time slot t, and , Indicates UAV attacker The flight altitude (the same for a UAV attacker in any time slot). This represents the set of interference power from all UAV attackers across all time slots. This represents the set of interference factors from all UAV attackers across all time slots.

[0181] S3. The objective function for maximizing the average energy efficiency of the legitimate party is decomposed into four first sub-problems that optimize the safe UAV flight trajectory, user transmission power, IRS phase shift, and ground jammer beam, respectively. The objective function for maximizing the average energy efficiency of the illegitimate party is decomposed into two second sub-problems that optimize the UAV attacker's flight trajectory and jamming target, and optimize the UAV attacker's jamming power, respectively.

[0182] When optimizing the flight trajectory of a safe UAV, the three optimization quantities in the other three first sub-problems are fixed (i.e., the user transmission power, IRS phase shift, and beamwidth of ground interference are fixed), thus obtaining the sub-problems for optimizing the flight trajectory of a safe UAV:

[0183] P1.1: ;

[0184] st C1.4-C1.7;

[0185] The three first subproblems of optimizing user transmit power, IRS phase shift, and beamwidth for ground interference are transformed into single-slot optimization problems:

[0186] For each time slot, when optimizing user transmission power, the three optimization quantities in the other three first sub-problems are fixed (i.e., the beams of the safe UAV flight trajectory, IRS phase shift, and ground interference are fixed), thus obtaining the sub-problems for optimizing user transmission power:

[0187] P1.2: ;

[0188] st C1.2;

[0189] For each time slot, when optimizing the IRS phase shift, the three optimization quantities in the other three first sub-problems are fixed (i.e., the safe UAV flight trajectory, user transmission power, and ground jammer beam are fixed), thus obtaining the sub-problems for optimizing the IRS phase shift:

[0190] P1.3: ;

[0191] st C1.1;

[0192] For each time slot, when optimizing the beam for ground-based jammers, three optimization quantities from the other three first sub-problems are fixed (i.e., the safe UAV flight trajectory, user transmission power, and IRS phase shift are fixed), thus obtaining the sub-problem for optimizing the beam for ground-based jammers:

[0193] P1.4: ;

[0194] st C1.3;

[0195] When optimizing the flight trajectory of the UAV attacker and the target of interference, the optimization amount of another second sub-problem is fixed (i.e., the interference power of the UAV attacker is fixed), thus obtaining the sub-problem for optimizing the flight trajectory of the UAV attacker and the target of interference:

[0196] P2.1: ;

[0197] st C2.1-C2.3, C2.5-C2.7;

[0198] For each time slot, when optimizing the interference power of the UAV attacker, the optimization amount of another second sub-problem is fixed (i.e., the flight trajectory of the UAV attacker and the interference target are fixed), thus obtaining the sub-problem for optimizing the interference power of the UAV attacker:

[0199] P2.2: ;

[0200] st C2.4.

[0201] S4. Establish a Stackelberg game model. Solve the four first sub-problems for the legitimate side first, then solve the two second sub-problems for the illegitimate side. The legitimate side uses the MATD3-PER algorithm (Multi-Agent Twin Delayed Deep Deterministic policy gradient-Prioritized Experience Replay) to find the safe UAV flight trajectory, and the AO algorithm (Alternating Optimization) to solve the sub-problems optimizing user transmission power, IRS phase shift, and ground jammer beam. The illegitimate side uses the MATD3-PER algorithm to solve the sub-problems optimizing the UAV attacker's flight trajectory and jamming target, and the projected gradient descent algorithm to solve the sub-problem optimizing the UAV attacker's jamming power. Finally, apply the optimal solutions to the six sub-problems to the communication system to achieve UAV energy consumption optimization, including:

[0202] When establishing the Stackelberg game model, the secure UAV and the UAV attacker constitute the Stackelberg game model, with the secure UAV acting as the leader and the UAV attacker acting as the leader's follower.

[0203] To solve the four first subproblems, we first solve the subproblem of optimizing the flight trajectory of the safe UAV. Each safe UAV is defined as an agent, referred to as the first agent, and the state space of each first agent in time slot t is: , Indicates a safe UAV In the time slot The three-dimensional coordinates , Indicates a safe UAV In the time slot The remaining flight power (where the remaining flight power is the total flight power minus the flight power consumption) is defined as the action space of each first agent in time slot t. Define that all first agents have the same reward function in time slot t, and that is... ,in The penalty for not satisfying constraint C1.7, ,in The penalty factor representing the safe UAV flying out of the range limit boundary;

[0204] The MATD3-PER algorithm outputs the optimal flight trajectory for all safe UAVs, and the AO algorithm solves the sub-problems of optimizing user transmission power, IRS phase shift, and ground jammer beams.

[0205] A first Actor-Critic neural network is introduced for each first agent, and each first Actor-Critic neural network includes a first Actor current network, a first Actor target network, two first Critic current networks, and two first Critic target networks.

[0206] To solve the two second subproblems, the first subproblem is to optimize the flight trajectory of the UAV attacker and the interference target. Each UAV attacker is defined as an agent, referred to as the second agent, and the state space of each second agent in time slot t is as follows: ,in Indicates UAV attacker In the time slot The three-dimensional coordinates, and , Indicates UAV attacker In the time slot The remaining flight power, This indicates that the UAV attacker was in the time slot. The interference target (and the interference target is one of all safe UAVs or is -1 (indicating no interference)) is defined as the action space of each second agent in time slot t. ,Bundle Relaxation is the key to having Constrained continuous variables, interval Divided into equal parts There are several sub-intervals, each corresponding to a safe UAV except for the first interval. The reward function for all second agents in time slot t is defined to be the same, and is... ,in The penalty for not satisfying constraint C2.7, ,in The penalty factor representing the safe UAV flying out of the range limit boundary;

[0207] A second Actor-Critic neural network is introduced for each second agent, and each second Actor-Critic neural network includes a second Actor current network, a second Actor target network, two second Critic current networks, and two second Critic target networks.

[0208] From time slot 1 to The MATD3-PER algorithm trains the first Actor-Critic neural network of the first agent:

[0209] S4.1 Initialize the initial state space of each first agent, and set the time slots... ;

[0210] S4.2. For each first agent, obtain the current state space. That is, the first intelligent agent (Safe UAV) In time slots The current state space of all first agents is and the set of their current state spaces is . The current state space of each first agent is used as the input to the current network of the corresponding first actor, and the current network of each first actor outputs the action space of the first agent. That is, the first intelligent agent In the time slot Action space and state space;

[0211] S4.3 Each first agent executes its own action space, and the set of action spaces of all first agents is: ;

[0212] S4.4, Start executing the AO algorithm;

[0213] S4.4.1. With fixed IRS phase shift and ground interference beam, a convex optimization method is used to solve the subproblem of optimizing user transmission power (this solution process belongs to the prior art): First, auxiliary variables are introduced to handle non-convex max operations, and then the optimization problem is transformed into a standard convex optimization problem. A solver is used to solve the problem, and finally the user transmission power of the current time slot is obtained.

[0214] S4.4.2. With fixed user transmission power and ground interference beams, the subproblem of optimizing IRS phase shift is solved using SDR (Semidefinite Relaxation) and SCA (Successive Convex Approximation) methods (this solution process is existing technology): For each IRS, the problem is transformed by constructing a positive semidefinite matrix through SDR. After the transformation, the objective function is still non-convex. Then, the safe rate expression is processed using the SCA method. Finally, Gaussian randomization is used to recover the feasible solution, and the phase shift of all IRSs in the current time slot is obtained.

[0215] S4.4.3. With fixed user transmission power and IRS phase shift, the SCA method is used to solve the subproblem of beamforming against ground jammers (this solution process is existing technology): First, auxiliary variables are introduced to handle the max operation, a first-order Taylor expansion is performed on the eavesdropping rate to construct a convex optimization problem, and then the solution is iteratively solved until convergence, and finally the jamming beamforming vector of the current time slot is output.

[0216] S4.5. After each first agent performs an action, the corresponding reward is represented as follows: That is, the first intelligent agent In the time slot The reward, the set of rewards for all first agents is And observe the state space of each first agent in the next time slot. That is, the first intelligent agent In the next time slot The state space of all first agents in the next time slot is and the set of state spaces of all first agents in the next time slot is . ;

[0217] S4.6, The first experiences of each first agent Stored in the first cache (according to the number of the first agent) (Sequential storage), and each time storage occurs, the priority of the first experience currently being stored is set to the highest priority in the first cache area;

[0218] S4.7 When the first buffer is full, the first Actor-Critic neural network is updated, i.e., S4.7.1-S4.7.5 are executed. If it is not full, S4.8 is executed.

[0219] S4.7.1, from the first buffer according to A batch of first-hand experience samples, among which Sample the first intelligent agent u to the first The probability of an experience. For the first intelligent agent u, the first The priority of the first experience (i.e., the priority set for S4.6), among which Let `x` be the summation index variable for the first buffer, representing the iteration through all the first experiences stored in the first buffer. [0,1], typically set to 0.6, is used to control the adjustment coefficients for random sampling and greedy sampling, and is determined based on probability. Calculate the first agent u's first... The importance weight of each first experience is: ,in, The size of the first buffer. To offset the influence of the priority experience replay method on the convergence results, a value of 0.4 is generally used;

[0220] S4.7.2 For each sampled first experience, calculate the first target action, the first target Q value, the current first Q value, and the first Critic loss respectively. The first Critic loss is calculated using the importance weight of the first experience. Use gradient descent to update the two first Critic current networks.

[0221] S4.7.3 When the preset update time is reached, update the current network of the first Actor: first calculate the gradient, and then use gradient ascent to update the current network of the first Actor;

[0222] S4.7.4 For each first agent, update the parameters of its own first Actor target network and the two first Critic target networks through soft updates;

[0223] S4.7.5 Calculate the TD error for each first agent corresponding to the sampled first experience. and according to The size of the first empirical sampled priority is updated, and The larger the value, the higher the priority.

[0224] S4.8, Order Repeat steps S4.2-S4.8 until... Then, one round of training of the first Actor-Critic neural network for each first agent is completed.

[0225] After each pair of first agent's first Actor-Critic neural networks completes one round of training, the second agent's second Actor-Critic neural network is trained for one round.

[0226] From time slot 1 to The training process of the second Actor-Critic neural network for the second agent using the MATD3-PER algorithm is as follows:

[0227] S4.10. Initialize the initial state space of each second agent, and set the time slots... ;

[0228] S4.11. For each second agent, obtain the current state space. That is, the second intelligent agent (UAV attacker) In time slots The current state space of the second agent is and the set of the current state spaces of all second agents is . The current state space of each second agent is used as the input to the current network of the corresponding second actor, and the current network of each second actor outputs the action space of the second agent. That is, the second intelligent agent In the time slot Action space and state space;

[0229] S4.12, Each second agent executes its own action space, and the set of action spaces of all second agents is: ;

[0230] S4.13. The subproblem of optimizing the interference power of UAV attackers is solved by using the projection gradient descent algorithm (this solution process belongs to the prior art): Under the goal of maximizing the energy efficiency of the illegal party, that is, under the goal of minimizing the safe rate and interference power consumption, the interference power is iteratively updated along the gradient direction, and after each update, it is projected into the feasible region until convergence.

[0231] S4.14. After each second agent completes its action, the corresponding reward is represented as follows: That is, the second intelligent agent In the time slot The reward, the set of rewards for all second agents is And observe the state space of each second agent in the next time slot. That is, the second intelligent agent In the next time slot The state space of all second agents in the next time slot is and the set of state spaces of all second agents in the next time slot is . ;

[0232] S4.15, the second experience of the second intelligent agent Stored in the second cache (according to the number of the second agent) (Sequential storage), and each time storage occurs, the priority of the currently stored second experience is set to the highest priority in the second cache area;

[0233] S4.16 When the second buffer is full, the second Actor-Critic neural network is updated, i.e., S4.16.1-S4.16.5 are executed. If it is not full, S4.17 is executed.

[0234] S4.16.1, from the second buffer according to A second set of samples was taken, among which Two intelligent agents Sampling to the The probability of an experience. As a second intelligent agent The The priority of the second experience (i.e., the one set for S4.15), where Let `x` be the summation index variable for the second buffer, representing the summation of all second experiences stored in the second buffer according to probability. Computational Second Agent The The importance weight of each second experience is: ,in, This is the size of the second buffer.

[0235] S4.16.2 For each sampled second experience, calculate the second target action, the second target Q value, the current second Q value, and the second Critic loss respectively. The second Critic loss is calculated using the importance weight of the second experience. Use gradient descent to update the current network of the two second Critics.

[0236] S4.16.3 When the preset update time is reached, update the current network of the second Actor: first calculate the gradient, and then use gradient ascent to update the current network of the second Actor;

[0237] S4.16.4 For each second agent, update the parameters of its respective second actor target network and the two second critic target networks through soft updates;

[0238] S4.16.5 Calculate the TD error for each sampled second agent corresponding to the second experience. and according to The size of the sampled second empirical priority is updated, and The larger the value, the higher the priority.

[0239] S4.17, Order Repeat steps S4.11-S4.17 until... Then, one round of training of the second Actor-Critic neural network for each second agent is completed;

[0240] S4.18. Enter the next round of training for the first Actor-Critic neural network and the second Actor-Critic neural network. Repeat S4.1-S4.18 continuously until the preset number of rounds is reached, then the training of the first Actor-Critic neural network and the second Actor-Critic neural network is completed.

[0241] After training the first and second Actor-Critic neural networks, for each first agent:

[0242] S4.19, will The current state space of the time slot is used as the input of the first actor's current network in the first actor-critic neural network, to obtain the action space of the first agent in the current time slot. Based on the action space of the first agent in the current time slot, the three-dimensional coordinates of the safe UAV in the current time slot are calculated. Then, the AO algorithm is used to obtain the optimal user transmission power, optimal IRS phase shift and optimal ground jammer beam in the current time slot.

[0243] S4.20, Order Repeat steps S4.19-S4.20 until... Then we get 1 to The three-dimensional coordinates of the time-slotted safe UAV are used to obtain the flight trajectory of the safe UAV, and this flight trajectory is from 1 to... The optimal flight trajectory of a safe UAV within a time slot, and the results obtained from 1 to... Within each time slot, the optimal user transmission power, optimal IRS phase shift, and optimal ground interfering beam are achieved.

[0244] For each second agent:

[0245] S4.21, will The current state space of the time slot is used as the input of the current network of the second actor in the trained second actor-critic neural network to obtain the action space of the second agent in the current time slot. The three-dimensional coordinates of the UAV attacker in the current time slot are calculated based on the action space of the second agent in the current time slot. The optimal interference target of the UAV attacker in the current time slot is calculated based on the action space of the second agent in the current time slot. The optimal interference power of the current time slot is obtained by using the projection gradient descent algorithm.

[0246] S4.22, Order Repeat steps S4.21-S4.22 until... Then we get 1 to The three-dimensional coordinates of the time-slot UAV attacker are obtained, thus revealing the UAV attacker's flight trajectory, which is defined as 1 to... The optimal flight trajectory of a UAV attacker within the time slot, and the results of 1 to... Within each time slot, the optimal target for interference by a UAV attacker, and the result of 1 to... Within a time slot, the optimal interference power for a UAV attacker in each time slot.

[0247] Finally, the optimal solutions to the six sub-problems are applied to the communication system to achieve UAV energy consumption optimization.

[0248] In another embodiment, the method proposed in this invention is further verified by combining specific experimental results. The experiment is implemented in Python 3.7 and runs on a computer with an Intel(R) Core(TM) i7-10510U CPU and 12G of memory.

[0249] Figure 3 The diagram shows the convergence plots of our proposed method for solving the valid and invalid sides.

[0250] Experimental results show that in the first 500 rounds, during the strategy exploration and intense competition phase, both sides continuously iterate their initial strategies, and the reward value fluctuates dramatically due to significant changes in strategies, demonstrating the significant impact of strategy iteration on energy efficiency optimization. From round 500 to 1000, the legitimate player, through the algorithm's learning of the environment and the illegal player's strategies, gradually iterates out more targeted optimization strategies. The illegal player's reward value initially decreases, then briefly increases due to trial-and-error strategy adjustments, but decreases again after failing to break through the legitimate player's defense. After 1000 rounds, both sides' strategies tend to reach a steady state, conforming to the equilibrium characteristics of a Stackelberg game, and both sides' reward values ​​are higher than in the initial exploration phase, proving that the algorithm can effectively drive the game towards equilibrium. By solving the game problem of energy efficiency optimization for both sides through dynamic strategy iteration, the practicality and scientific validity of the algorithm are verified.

[0251] This energy efficiency optimization method for resisting intelligent full-duplex attacks through IRS and UAV collaboration fully unleashes the potential of multi-device collaboration by constructing a multi-node collaborative architecture of secure UAV-IRS-ground jammer. It can effectively resist the dual threats of eavesdropping and interference from intelligent full-duplex attackers, significantly enhancing the physical layer security of the communication system. By introducing a Stackelberg game model to simulate the dynamic attack and defense interaction between the legitimate and illegitimate parties, the strategies of both parties converge to an equilibrium state through iterative convergence, achieving adaptive adaptation of attack and defense strategies and global energy efficiency optimization in adversarial scenarios. In terms of algorithms, the legitimate party adopts the MATD3-PER algorithm and the AO algorithm, while the illegitimate party adopts the MATD3-PER algorithm and the projective gradient descent method, solving the optimization problem of strong coupling of multi-dimensional variables.

[0252] It should be understood that, although Figure 1 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order in which these steps are executed, and they can be performed in other orders. Figure 1 At least some of the steps in the process may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least some of the sub-steps or stages of other steps.

[0253] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0254] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the appended claims.

Claims

1. An energy efficiency optimization method for resisting intelligent full-duplex attacks through IRS and UAV collaboration, characterized in that: The energy efficiency optimization method for resisting smart full-duplex attacks through IRS and UAV collaboration is applied to a communication system containing both legitimate and illegitimate parties. The legitimate parties include a secure UAV, an IRS, users, and a ground-based jammer. The illegitimate parties include a full-duplex UAV attacker. Users communicate with the secure UAV through the IRS, the UAV attacker eavesdrops on user information and interferes with the secure UAV through the IRS, and the ground-based jammer simultaneously interferes with the UAV attacker. The energy efficiency optimization method for resisting smart full-duplex attacks through IRS and UAV collaboration includes: S1. Calculate the secure communication rate of each user, and calculate the energy efficiency of the legitimate party and the energy efficiency of the illegitimate party based on the secure communication rate of the users. S2. Construct the objective function for maximizing the average energy efficiency of the legal side, and construct the objective function for maximizing the average energy efficiency of the non-legal side; S3. The objective function for maximizing the average energy efficiency of the legitimate party is decomposed into four first sub-problems that optimize the flight trajectory of the safe UAV, the user transmission power, the IRS phase shift, and the beam of the ground jammer, respectively. The objective function for maximizing the average energy efficiency of the illegitimate party is decomposed into two second sub-problems that optimize the flight trajectory of the UAV attacker and the jamming target, and optimize the jamming power of the UAV attacker, respectively. S4. Establish a Stackelberg game model. Solve the four first subproblems for the legitimate side first, and then solve the two second subproblems for the illegitimate side. The legitimate side uses the MATD3-PER algorithm to find the safe UAV flight trajectory, and the AO algorithm to solve the subproblems of optimizing user transmission power, optimizing IRS phase shift, and optimizing the beam of ground jammers. The illegitimate side uses the MATD3-PER algorithm to solve the subproblems of optimizing the flight trajectory and jamming target of the UAV attacker, and the projected gradient descent algorithm to solve the subproblem of optimizing the interference power of the UAV attacker. Finally, apply the optimal solutions of the six subproblems to the communication system to achieve UAV energy consumption optimization.

2. The energy efficiency optimization method for resisting intelligent full-duplex attacks through IRS and UAV collaboration as described in claim 1, characterized in that: Each user is represented as , The total number of users, and each security UAV is represented as , The total number of secure UAVs, and the attacker of each UAV is represented as , The total number of UAV attackers, represented by each IRS. , The total number of IRSs, and the number of antennas of ground jammers are: The preset time period of the communication system is divided into There are 1 time slot, and each time slot is represented as The reflection elements of each IRS are represented as follows: , Let be the number of reflection elements in each IRS. In time slot t, the reflection coefficient matrix of each IRS is expressed as: , It is a complex number. The imaginary unit, For the first The first IRS The phase shift of each reflecting element in response to the incident signal. The dimension is Complex matrix; In time slot t, the channel gains from each user to the secure UAV, from each user to the IRS, and from each IRS to the secure UAV are expressed as follows: , and Then the composite channel gain from each user to the secure UAV is , Indicates conjugate transpose; In time slot t, the channel gains from each user to the UAV attacker and from each IRS to the UAV attacker are expressed as follows: and The composite channel gain from each user to the UAV attacker is expressed as: ; In time slot t, the channel gains from each UAV attacker to the secure UAV and from the ground jammer to the UAV attacker are expressed as follows: and .

3. The energy efficiency optimization method for resisting smart full-duplex attacks through IRS and UAV collaboration as described in claim 1, characterized in that: The calculation of the secure communication rate for each user, and the calculation of the energy efficiency of the legitimate party and the energy efficiency of the non-legitimate party based on the user's secure communication rate, include: The formula for calculating the secure communication rate for each user is as follows: ; in, ; ; ; in, This represents the secure communication rate of each user in time slot t. This represents the communication rate between each user and the secure UAV in time slot t. This represents the rate at which each UAV attacker eavesdrops on each user in time slot t. Indicates user and IRS The correlation coefficient, , Indicates user via IRS Communicate with secure UAV, otherwise This means that the user Not through IRS For secure UAV communication, each user is associated with its nearest IRS, and each user is associated with only one IRS. , Indicates the user in time slot t With security UAV The correlation coefficient, , Indicates user With security UAV Communication, otherwise This means that the user Not related to safe UAV Communication involves each user communicating with the nearest secure UAV covering their area. , Indicates the user in time slot t Transmission power, This indicates that the attacker was in time slot tUAV. For safe UAV Interference factors, This indicates a UAV attacker. Towards secure UAV Send interference signals, otherwise This means UAV attackers Not to secure UAV Sending interference signals, subject to constraints , This indicates that the attacker was in time slot tUAV. For safe UAV Interference power, Indicates a safe UAV noise power, This indicates the self-interference efficiency of a UAV attacker. This indicates ground-based jammers targeting UAV attackers. Interference beams, Indicates UAV attacker The noise power; The formula for calculating the energy efficiency of a legal entity is as follows: ; in, This represents the energy efficiency of the legal side in time slot t. This represents the power balance factor for each user's transmission. This represents the flight power consumption balance factor for each safety UAV. This represents the flight power consumption of each safe UAV in time slot t. The power balance factor representing the ground-based interferator. The norm of a vector; The formula for calculating the energy efficiency of non-legal entities is as follows: ; in, This represents the energy efficiency of the non-legal party in time slot t. This represents the interference power balance factor for each UAV attacker. This represents the flight power consumption balance factor for each UAV attacker. This represents the flight power consumption of each UAV attacker in time slot t.

4. The energy efficiency optimization method for resisting intelligent full-duplex attacks through IRS and UAV collaboration as described in claim 3, characterized in that: The construction of the objective function for maximizing the average energy efficiency of the legal side and the objective function for maximizing the average energy efficiency of the non-legal side include: The formula for the objective function that maximizes the average energy efficiency of the legal side is as follows: P1: ; s.t. C1.1: , ; C1.2: , ; C1.3: , ; C1.4: , ; C1.5: , ; C1.6: , ; C1.7: , , ; in, ; ; ; ; Where C1.1 is the constraint on the IRS phase shift, and C1.2 is the constraint on the user's transmission power. For users The maximum transmission power, C1.3 is the total power constraint of the ground jammer. C1.4 is the maximum power of the ground-based jammer and the flight power consumption constraint for a safe UAV. For safe UAV The total flight power, C1.5 and C1.6 are the safe flight speeds of the UAV in time slot t, respectively. and angle Restrictions, For safe UAV The maximum flight speed, C1.7 is the safe flight range limit for UAVs. and These are secure UAVs The horizontal and vertical axes of the flight range in time slot t, and These are secure UAVs The x-coordinate and y-coordinate of the maximum flight range. This represents the set of flight trajectories of all safe UAVs across all time slots. Indicates a safe UAV In the three-dimensional coordinates of time slot t, and , Indicates a safe UAV Flight altitude This represents the set of transmission power for all users across all time slots. This represents the set of phase shifts for all IRSs across all time slots. This represents the set of interference beams from ground-based jammers across all time slots. The formula for maximizing the average energy efficiency of non-legal parties is as follows: P2: ; s.t. C2.1: ; C2.2: ; C2.3: , ; C2.4: ; C2.5: ; C2.6: ; C2.7: ; in, ; ; ; Among them, C2.1 and C2.2 are constraints on the interference factors of the UAV attacker, and C2.3 is a constraint on the flight power consumption of the UAV attacker. For UAV attackers The total flight power, C2.4 is the interference power constraint for UAV attackers. and C1 and C2.6 represent the minimum and maximum interference power of the UAV attacker, respectively, and C2.5 and C2.6 represent the flight speed of the UAV attacker in time slot t, respectively. and angle Restrictions, For UAV attackers The maximum flight speed, C2.7, is the flight range limit for UAV attackers. and These are UAV attackers The horizontal and vertical axes of the flight range in time slot t, and These are UAV attackers The x and y coordinates of the flight range at time slot t are the largest. This represents the set of flight trajectories of all UAV attackers across all time slots. Indicates UAV attacker In the three-dimensional coordinates of time slot t, and , Indicates UAV attacker Flight altitude This represents the set of interference power from all UAV attackers across all time slots. This represents the set of interference factors from all UAV attackers across all time slots.

5. The energy efficiency optimization method for resisting intelligent full-duplex attacks through IRS and UAV collaboration as described in claim 4, characterized in that: The objective function for maximizing the average energy efficiency of the legitimate party is decomposed into four first sub-problems that optimize the safe UAV flight trajectory, user transmission power, IRS phase shift, and ground jammer beam, respectively. The objective function for maximizing the average energy efficiency of the illegitimate party is decomposed into two second sub-problems that optimize the UAV attacker's flight trajectory and the jamming target, and optimize the UAV attacker's jamming power, respectively. These include: When optimizing the flight trajectory of a safe UAV, the three optimization quantities in the other three first sub-problems are fixed, thus obtaining the sub-problems for optimizing the flight trajectory of a safe UAV: P1.1: ; st C1.4-C1.7; The three first subproblems of optimizing user transmit power, IRS phase shift, and beamwidth for ground interference are transformed into single-slot optimization problems: For each time slot, when optimizing user transmission power, the three optimization quantities in the other three first sub-problems are fixed, thus obtaining the sub-problems for optimizing user transmission power: P1.2: ; st C1.2; For each time slot, when optimizing the IRS phase shift, the three optimization quantities in the other three first sub-problems are fixed, thus obtaining the sub-problems for optimizing the IRS phase shift: P1.3: ; st C1.1; For each time slot, when optimizing the beam for ground jammers, the three optimization quantities in the other three first sub-problems are fixed, thus obtaining the sub-problems for optimizing the beam for ground jammers: P1.4: ; st C1.3; When optimizing the flight trajectory of the UAV attacker and the interference target, the optimization amount of another second sub-problem is fixed, thus obtaining the sub-problem for optimizing the flight trajectory of the UAV attacker and the interference target: P2.1: ; st C2.1-C2.3, C2.5-C2.7; For each time slot, when optimizing the interference power of a UAV attacker, the optimization amount of another second sub-problem is fixed, thus obtaining the sub-problem for optimizing the interference power of a UAV attacker: P2.2: ; st C2.

4.

6. The energy efficiency optimization method for resisting smart full-duplex attacks through IRS and UAV collaboration as described in claim 5, characterized in that: When establishing the Stackelberg game model, the secure UAV and the UAV attacker constitute the Stackelberg game model, with the secure UAV acting as the leader and the UAV attacker acting as the leader's follower. To solve the four first subproblems, we first solve the subproblem of optimizing the flight trajectory of the safe UAV. Each safe UAV is defined as an agent, referred to as the first agent, and the state space of each first agent in time slot t is: , Indicates a safe UAV In the time slot The three-dimensional coordinates , Indicates a safe UAV In the time slot The remaining flight power, defining the action space of each first agent in time slot t as... Define that all first agents have the same reward function in time slot t, and that is... ,in The penalty for not satisfying constraint C1.7, ,in The penalty factor representing the safe UAV flying out of the range limit boundary; A first Actor-Critic neural network is introduced for each first agent, and each first Actor-Critic neural network includes a first Actor current network, a first Actor target network, two first Critic current networks, and two first Critic target networks.

7. The energy efficiency optimization method for resisting intelligent full-duplex attacks through IRS and UAV collaboration as described in claim 6, characterized in that: To solve the two second subproblems, the first subproblem is to optimize the flight trajectory of the UAV attacker and the interference target. Each UAV attacker is defined as an agent, referred to as the second agent, and the state space of each second agent in time slot t is as follows: ,in Indicates UAV attacker In the time slot The three-dimensional coordinates, and , Indicates UAV attacker In the time slot The remaining flight power, This indicates that the UAV attacker was in the time slot. The interference target is defined as the action space of each second agent in time slot t. ,Bundle Relaxation is the key to having Constrained continuous variables, interval Divided into equal parts There are several sub-intervals, each corresponding to a safe UAV except for the first interval. The reward function for all second agents in time slot t is defined to be the same, and is... ,in The penalty for not satisfying constraint C2.7 ,in The penalty factor representing the safe UAV flying out of the range limit boundary; A second Actor-Critic neural network is introduced for each second agent, and each second Actor-Critic neural network includes a second Actor current network, a second Actor target network, two second Critic current networks, and two second Critic target networks.

8. The energy efficiency optimization method for resisting smart full-duplex attacks through IRS and UAV collaboration as described in claim 7, characterized in that: From time slot 1 to The MATD3-PER algorithm trains the first Actor-Critic neural network of the first agent: S4.1 Initialize the initial state space of each first agent, and set the time slots... ; S4.

2. For each first agent, obtain the current state space. And the set of current state spaces of all first agents is The current state space of each first agent is used as the input to the current network of the corresponding first actor, and the current network of each first actor outputs the action space of the first agent. ; S4.3 Each first agent executes its own action space, and the set of action spaces of all first agents is: ; S4.4, Start executing the AO algorithm; S4.4.

1. With fixed IRS phase shift and ground interference beam, a subproblem for optimizing user transmission power is solved using a convex optimization method. S4.4.

2. With fixed user transmission power and ground interference beam, solve the subproblem of optimizing IRS phase shift using SDR and SCA methods; S4.4.

3. With fixed user transmission power and IRS phase shift, the SCA method is used to solve the subproblem of beam optimization for ground interference. S4.

5. After each first agent performs an action, the corresponding reward is represented as follows: The set of rewards for all first agents is And observe the state space of each first agent in the next time slot, and the set of the state spaces of all first agents in the next time slot is: ; S4.6, The first experiences of each first agent Store it in the first cache, and each time it is stored, set the priority of the first experience stored at the current time to the highest priority in the first cache; S4.7 When the first buffer is full, the first Actor-Critic neural network is updated, i.e., S4.7.1 is executed; if it is not full, S4.8 is executed. S4.7.1, from the first buffer according to A batch of first-hand experience samples were collected, among which Sample the first intelligent agent u to the first The probability of an experience. For the first intelligent agent u, the first Prioritizing first experience, among which Let `x` be the summation index variable for the first buffer, representing the iteration through all the first experiences stored in the first buffer. [0,1] is used to control the adjustment coefficients for random sampling and greedy sampling, and is determined according to probability. Calculate the first agent u's first... The importance weight of each first experience is: ,in, The size of the first buffer. To offset the impact of the priority experience replay method on the convergence results; S4.7.2 For each sampled first experience, calculate the first target action, the first target Q value, the current first Q value, and the first Critic loss respectively. The first Critic loss is calculated using the importance weight of the first experience. Use gradient descent to update the two first Critic current networks. S4.7.3 When the preset update time is reached, update the current network of the first Actor: first calculate the gradient, and then use gradient ascent to update the current network of the first Actor; S4.7.4 For each first agent, update the parameters of its own first Actor target network and the two first Critic target networks through soft updates; S4.7.5 Calculate the TD error for each first agent corresponding to the sampled first experience. and according to The size of the first empirical sampled priority is updated, and The larger the value, the higher the priority. S4.8, Order Repeat steps S4.2-S4.8 until... Then, one round of training of the first Actor-Critic neural network for each first agent is completed. After each pair of first agent's first Actor-Critic neural networks completes one round of training, the second agent's second Actor-Critic neural network is trained for one round.

9. The energy efficiency optimization method for resisting intelligent full-duplex attacks through IRS and UAV collaboration as described in claim 8, characterized in that: From time slot 1 to The MATD3-PER algorithm trains the second Actor-Critic neural network for the second agent: S4.

10. Initialize the initial state space of each second agent, and set the time slots... ; S4.

11. For each second agent, obtain the current state space. And the set of current state spaces of all second agents is The current state space of each second agent is used as the input to the current network of the corresponding second actor, and the current network of each second actor outputs the action space of the second agent. ; S4.12, Each second agent executes its own action space, and the set of action spaces of all second agents is: ; S4.

13. The subproblem of optimizing the interference power of UAV attackers is solved using the projected gradient descent algorithm. S4.

14. After each second agent completes its action, the corresponding reward is represented as follows: The reward set for all second agents is And observe the state space of each second agent in the next time slot, and the set of the state spaces of all second agents in the next time slot is: ; S4.15, the second experience of the second intelligent agent Store it in the second cache, and each time it is stored, set the priority of the currently stored second experience to the highest priority in the second cache; S4.16 When the second buffer is full, the second Actor-Critic neural network is updated, i.e., S4.16.1 is executed; if it is not full, S4.17 is executed. S4.16.1, from the second buffer according to A second set of samples was taken, among which Two intelligent agents Sampling to the The probability of an experience. As a second intelligent agent The The priority of the second experience, among which Let `x` be the summation index variable for the second buffer, representing the summation of all second experiences stored in the second buffer according to probability. Computational Second Agent The The importance weight of each second experience is: ,in, This is the size of the second buffer. S4.16.2 For each sampled second experience, calculate the second target action, the second target Q value, the current second Q value, and the second Critic loss respectively. The second Critic loss is calculated using the importance weight of the second experience. Use gradient descent to update the current network of the two second Critics. S4.16.3 When the preset update time is reached, update the current network of the second Actor: first calculate the gradient, and then use gradient ascent to update the current network of the second Actor; S4.16.4 For each second agent, update the parameters of its respective second actor target network and the two second critic target networks through soft updates; S4.16.5 Calculate the TD error for each sampled second agent corresponding to the second experience. and according to The size of the sampled second empirical priority is updated, and The larger the value, the higher the priority. S4.17, Order Repeat steps S4.11-S4.17 until... Then, one round of training of the second Actor-Critic neural network for each second agent is completed; S4.

18. Enter the next round of training for the first Actor-Critic neural network and the second Actor-Critic neural network. Repeat S4.1-S4.18 continuously until the preset number of rounds is reached, then the training of the first Actor-Critic neural network and the second Actor-Critic neural network is completed.

10. The energy efficiency optimization method for resisting intelligent full-duplex attacks through IRS and UAV collaboration as described in claim 9, characterized in that: After training the first and second Actor-Critic neural networks, for each first agent: S4.19, will The current state space of the time slot is used as the input of the first actor's current network in the first actor-critic neural network, to obtain the action space of the first agent in the current time slot. Based on the action space of the first agent in the current time slot, the three-dimensional coordinates of the safe UAV in the current time slot are calculated. Then, the AO algorithm is used to obtain the optimal user transmission power, optimal IRS phase shift and optimal ground jammer beam in the current time slot. S4.20, Order Repeat steps S4.19-S4.20 until... Then we get 1 to The three-dimensional coordinates of the time-slotted safe UAV are used to obtain the flight trajectory of the safe UAV, and this flight trajectory is from 1 to... The optimal flight trajectory of a safe UAV within a time slot, and the results obtained from 1 to... Within each time slot, the optimal user transmission power, optimal IRS phase shift, and optimal ground interfering beam are achieved. For each second agent: S4.21, will The current state space of the time slot is used as the input of the current network of the second actor in the trained second actor-critic neural network to obtain the action space of the second agent in the current time slot. The three-dimensional coordinates of the UAV attacker in the current time slot are calculated based on the action space of the second agent in the current time slot. The optimal interference target of the UAV attacker in the current time slot is calculated based on the action space of the second agent in the current time slot. The optimal interference power of the current time slot is obtained by using the projection gradient descent algorithm. S4.22, Order Repeat steps S4.21-S4.22 until... Then we get 1 to The three-dimensional coordinates of the time-slot UAV attacker are obtained, thus revealing the UAV attacker's flight trajectory, which is defined as 1 to... The optimal flight trajectory of a UAV attacker within the time slot, and the results of 1 to... Within each time slot, the optimal target for interference by a UAV attacker, and the result of 1 to... Within a time slot, the optimal interference power for a UAV attacker in each time slot.