Unmanned aerial vehicle cluster communication perception integrated network physical layer secure transmission method and device, medium and equipment
Patent Information
- Application Number
- CN202510177236.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-05-23
AI Technical Summary
The physical layer security of UAV clusters in ISAC scenarios has not been fully studied, and there is non-convexity in the coordinated optimization of base station beamforming and UAV cluster trajectory, which is difficult to solve by traditional methods.
By obtaining the initialization data in the integrated data transmission scenario of communication and perception, the problem of maximizing the average confidentiality rate is constructed, and an improved near-end strategy optimization algorithm is used to solve the problem, and a collaborative control strategy is obtained to optimize the base station beamforming and the flight trajectory of the UAV cluster to achieve secure transmission of the physical layer.
Under the condition of ensuring perceptual performance, the physical layer security performance of the transmission of the ISAC communication network in the UAV cluster is improved, and the quality differences between legal channels and eavesdropping channels are optimized.
Smart Images

Figure CN120034856A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of drone communications, and in particular to a method, device, medium and equipment for secure transmission of the physical layer of a drone cluster communication perception integrated network. Background Art
[0002] Integrated Sensing and Communication (ISAC), as one of the key technologies of the Sixth Generation (6G) communication network, is expected to play an important role in applications such as high-precision positioning, low-altitude security, and smart transportation, improve spectrum efficiency, and reduce hardware costs. In addition, unmanned aerial vehicles (UAVs) are expected to undertake diverse tasks in communication networks due to their high flexibility and on-demand deployment capabilities. ISAC technology can be used to locate and identify UAVs in airspace. The combination of ISAC and UAVs can also bring a larger perception range, lower deployment costs, and additional design freedom to communication networks, and increase the environmental perception capabilities of UAVs. Furthermore, UAV clusters have the advantages of information fusion and resource complementarity compared to single UAVs, which can improve the accuracy and reliability of UAVs in performing various tasks.
[0003] However, the high line of sight (LoS) property of the communication link between UAV and other nodes in the communication network makes its communication link and control link vulnerable to eavesdropping. At the same time, the security of transmission has not been fully studied in scenarios involving the combination of UAV and ISAC technology. Physical layer security technology achieves secure transmission through modulation and coding of communication signals, and can be used to ensure the secure transmission of sensitive information of UAV clusters in ISAC scenarios. To improve the physical layer security performance of UAV clusters, the key lies in the coordinated optimization of network parameters such as base station beamforming and UAV cluster trajectory, but this optimization problem has a non-convex form and is difficult to solve through methods such as semidefinite relaxation (SDR) and successive convex approximation (SCA). Summary of the invention
[0004] The main purpose of this application is to provide a method, device, medium and equipment for secure transmission of the physical layer of an integrated network for communication perception of a UAV cluster, aiming to ensure the secure transmission of the ISAC physical layer of a UAV cluster with an efficient and flexible network parameter optimization method.
[0005] To achieve the above-mentioned objectives, the present application provides a method for secure transmission of the physical layer of a drone cluster communication and perception integrated network, including: obtaining initialization data in a communication and perception integrated data transmission scenario, wherein the initialization data includes the initial position of each drone in the drone cluster, the initial position of each eavesdropping drone, the initial position of each perception device and the transmission power of the ISAC base station; constructing an average confidentiality rate maximization problem, wherein the average confidentiality rate maximization problem is a joint optimization problem of base station beamforming and drone cluster trajectory under constraints; based on the initialization data, an improved proximal strategy optimization algorithm is used to solve the average confidentiality rate maximization problem to obtain a collaborative control strategy, wherein the collaborative control strategy is used to simultaneously optimize the beamforming of the base station and the flight trajectory of the drone cluster to securely transmit the physical layer data between the perception device and the drone cluster.
[0006] Optionally, the process of constructing the average confidentiality rate maximization problem includes: determining a first transmission rate at which the ISAC base station transmits confidential data to each of the drones based on a received signal from each of the drones, and determining a second transmission rate at which each of the eavesdropping drones eavesdrops on the confidential data from each of the drones based on a received signal from each of the eavesdropping drones; determining the confidentiality transmission rate of each of the drones based on the first transmission rate and the second transmission rate, and determining a maximum value expression for the average confidentiality rate of the cluster drones based on the confidentiality transmission rate of each of the drones; constructing constraints corresponding to the maximum value expression of the average confidentiality rate; and constructing the average confidentiality rate maximization problem based on the maximum value expression and the constraints.
[0007] Optionally, the constraints corresponding to the construction of the maximum expression of the average confidentiality rate include: determining the speed constraints of each of the UAVs based on the relative positions of each of the UAVs; determining the safety distance constraints of each of the UAVs based on the displacement of each of the UAVs; determining the transmission power constraints of the ISAC base station based on the norm of the beam matrix; determining the signal strength constraints of each of the sensing devices based on the transmission radiation patterns of the ISAC base station in azimuth and elevation; determining the constraints corresponding to the maximum expression of the average confidentiality rate based on the speed constraints, the safety distance constraints, the transmission power constraints and the signal strength constraints.
[0008] Optionally, before determining the first transmission rate at which the ISAC base station transmits confidential data to each of the drones based on the received signals of each of the drones, and determining the second transmission rate at which each of the eavesdropping drones eavesdrops on the confidential data of each of the drones based on the received signals of each of the eavesdropping drones, the drone cluster communication and perception integrated network physical layer security transmission method also includes: obtaining the confidential communication signal vector and perception signal vector of the ISAC base station in any time slot, as well as the communication beamforming matrix and the perception beamforming matrix; obtaining a first sub-transmit signal based on the communication beamforming matrix and the confidential communication signal vector; obtaining a second sub-transmit signal based on the perception beamforming matrix and the perception signal vector. ; Based on the first sub-transmit signal and the second sub-transmit signal, the transmit signal of the ISAC base station is obtained; based on the channel gain of each of the UAVs, the transmit signal of the ISAC base station and the additive Gaussian white noise of each of the UAVs, the receive signal of each of the UAVs is correspondingly determined, and based on the channel gain of each of the eavesdropping UAVs, the transmit signal of the ISAC base station and the additive Gaussian white noise of each of the eavesdropping UAVs, the receive signal of each of the eavesdropping UAVs is correspondingly determined, wherein the channel gain of each of the UAVs and the channel gain of each of the eavesdropping UAVs are determined based on path loss, small-scale signal gain and transmit array steering vector, and the array element spacing of the azimuth and pitch angles of the transmit array steering vector is equal to half a wavelength.
[0009] Optionally, the improved proximal strategy optimization algorithm is used to solve the average confidentiality rate maximization problem to obtain a collaborative control strategy, including: determining a Markov decision process based on the average confidentiality rate maximization problem; optimizing the Markov decision process based on the proximal strategy optimization algorithm obtained after interaction and updating with the environment to obtain the collaborative control strategy.
[0010] Optionally, determining the Markov decision process based on the average confidentiality rate maximization problem includes: in the current time slot, based on the channel state information and the position information of each of the eavesdropping drones and each of the drones in the drone cluster, determining the action space of the Markov decision process based on the beamforming matrix of the base station and the trajectory of the drone cluster; and determining the total return of the Markov decision process based on the sum of the total confidentiality rate between the drones, the perception performance return of each of the perception devices, and the anti-collision return between the drones.
[0011] Optionally, the proximal policy optimization algorithm is determined based on a policy network and a value network, and the proximal policy optimization algorithm obtained after the update based on the interaction with the environment optimizes the Markov decision process to obtain the collaborative control strategy, including: determining the probability distribution of selecting action a and the corresponding reward probability in the current state s based on the state space, the action space and the total reward, and using a policy network to characterize the probability distribution, wherein the policy network refers to a parameterized neural network; based on the expectation of the reward probability obtained by the agent when performing action a in the current state s, obtaining the action value function, based on the expectation of the reward probability obtained by the agent when performing action a according to the optimal strategy in the current state s. , obtain the state value function, and obtain the advantage function based on the difference between the action value function and the state value function, wherein the state value function is obtained by approximating the state value corresponding to the current state based on the parameterized value network; the advantage function is estimated by using a generalized advantage estimation algorithm to obtain a generalized advantage estimation function; the objective function of the policy network is determined by using the generalized advantage estimation function weighted by the strategy ratio and the generalized advantage estimation function weighted by the clipping function; the objective function of the value network is determined based on the generalized advantage estimation function; the collaborative control strategy of the agent is determined based on the objective function of the policy network, the objective function of the value network and the agent-environment interaction data.
[0012] In addition, to achieve the above-mentioned purpose, the second aspect of the present application also provides a physical layer security transmission device for integrated communication and perception of a drone cluster, including: an acquisition module, used to obtain initialization data in a communication and perception integrated data transmission scenario, wherein the initialization data includes the initial position of each drone in the drone cluster, the initial position of each eavesdropping drone, the initial position of each perception device and the transmission power of the ISAC base station; a problem construction module, used to construct an average confidentiality rate maximization problem, wherein the average confidentiality rate maximization problem is a joint optimization problem of base station beamforming and drone cluster trajectory under constraints; a transmission control module, used to solve the average confidentiality rate maximization problem based on the initialization data, using an improved proximal strategy optimization algorithm to obtain a collaborative control strategy, and based on the collaborative control strategy, simultaneously optimize the beamforming of the base station and the flight trajectory of the drone cluster to securely transmit the physical layer data between the perception device and the drone cluster.
[0013] To achieve the above-mentioned objectives, the present application also provides a computer-readable storage medium, which includes instructions, which, when executed on a computer, enables the computer to execute the drone cluster communication perception integrated network physical layer security transmission method provided in the first aspect.
[0014] To achieve the above-mentioned purpose, the present application also provides an electronic device, which includes: at least one processor, a memory and an input-output unit; wherein the memory is used to store a computer program, and the processor is used to call the computer program stored in the memory to execute the drone cluster communication perception integrated network physical layer security transmission method provided in the first aspect.
[0015] The embodiments of the present application propose a method, device, medium and equipment for secure transmission of the physical layer of a drone cluster communication and perception integrated network. The method obtains initialization data in a communication and perception integrated data transmission scenario, wherein the initialization data includes the initial position of each drone in the drone cluster, the initial position of each eavesdropping drone, the initial position of each perception device and the transmission power of the ISAC base station; constructs an average confidentiality rate maximization problem, wherein the average confidentiality rate maximization problem is a joint optimization problem of base station beamforming and drone cluster trajectory under constraints; based on the initialization data, an improved proximal strategy optimization algorithm is used to solve the average confidentiality rate maximization problem to obtain a collaborative control strategy, wherein the collaborative control strategy is used to simultaneously optimize the beamforming of the base station and the flight trajectory of the drone cluster to securely transmit the physical layer data between the perception device and the drone cluster. That is to say, the present application maximizes the quality difference between the legitimate channel and the eavesdropping channel through the joint optimization of the transmit beamforming of the ISAC base station and the drone cluster trajectory, thereby improving the transmission security performance while ensuring the perception performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 A flowchart diagram of an embodiment of a method for secure transmission of physical layer of integrated network for communication and perception of drone clusters provided in this application;
[0017] Figure 2 A UAV cluster ISAC communication network model diagram provided for an embodiment of a method for secure transmission of physical layer of an integrated network of UAV cluster communication perception of this application;
[0018] Figure 3 The three-dimensional transmission beam pattern of the ISAC base station at the 200th time slot provided in the first embodiment of the UAV cluster communication perception integrated network physical layer security transmission method of the present application;
[0019] Figure 4 A schematic diagram of the trajectory of the UAV cluster and a simulation result diagram of the positions of the ISAC base station, the eavesdropping UAV, and the sensing target after the optimization is completed in the first embodiment of the UAV cluster communication and perception integrated network physical layer security transmission method of the present application;
[0020] Figure 5 A structural block diagram is provided for an embodiment of a physical layer secure transmission device with integrated communication perception for drone cluster communication of the present application.
[0021] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0022] It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0023] In the prior art, the Proximal Policy Optimization (PPO) algorithm is an algorithm for optimizing policy networks, which belongs to the category of model-free reinforcement learning algorithms. It is mainly used to solve the problems of low sample efficiency and unstable training in the policy gradient-based algorithm during training. The goal of reinforcement learning is to allow the agent to take a series of actions in the environment to maximize the cumulative reward. In this process, the optimization of the policy network (the neural network used to generate actions) is the key.
[0024] The basic principle of PPO. PPO stabilizes the training process by limiting the difference between the new strategy and the old strategy. It uses an important concept - the probability ratio. The strategy ratio refers to the ratio of the probability of taking a certain action under the new strategy to the probability of taking the same action under the old strategy. By controlling this ratio, the strategy update will not be too drastic. For example, assuming the old strategy is and the new strategy is, for a certain state and action, the strategy ratio.
[0025] The objective function of PPO. The objective function of PPO contains a clipped surrogate objective. The basic idea is to clip the policy ratio to prevent the policy from updating too quickly. The specific form is, where is the estimated advantage function, which measures the advantage of taking a certain action in a certain state relative to the average action, and is a small positive number used to control the scope of clipping.
[0026] Advantage function estimation for PPO. The advantage function plays a key role in PPO. It is used to evaluate how much better it is to take a certain action in a given state than to act according to the average strategy. A common way to estimate the advantage function is to use Generalized Advantage Estimation (GAE). GAE estimates the advantage function more accurately by combining multi-step returns and reducing the variance of the estimate.
[0027] Advantages of the PPO algorithm. Compared with traditional policy gradient algorithms, PPO can use collected samples for learning more effectively. It can achieve better training results with a smaller number of samples. For example, in some complex robot control tasks or game environments, collecting samples may be difficult or costly, and the high sample efficiency of PPO is particularly important. Due to the use of mechanisms such as policy ratio clipping, PPO can avoid large fluctuations in the policy update process. This makes the training process more stable and less prone to training failures or performance degradation due to too fast policy updates.
[0028] The state space of a Markov decision process is defined as: The state space is the set of all possible states in a Markov decision process (MDP). It is represented by symbols. A state contains all the necessary information to describe the current situation of the system. This information is the basis for the agent to make decisions.
[0029] Action Space is defined as: Action Space is the set of all actions that the agent can take in each state, represented by symbols. For each state, the agent selects an action from the action space to perform.
[0030] Reward is defined as: Reward is the immediate feedback obtained by the agent from the environment after taking an action in a certain state. It is usually represented by , which is a scalar value. The reward function is used to measure whether the agent's actions are beneficial to achieving the goal.
[0031] The action-value function, also known as the Q-function, is used in the Markov decision process (MDP) to measure the expectation of the cumulative reward that the agent can obtain after taking a specific action in a specific state. It is represented by . Essentially, it is a prediction of future returns, taking into account the sum of all the rewards that the agent can obtain from executing the action in the current state and following a certain strategy.
[0032] Reward is the feedback signal that an agent receives from the environment immediately after taking an action in a certain state.
[0033] The return is the cumulative value of all rewards obtained by the agent along a trajectory starting from a certain point in time, and a discount factor is usually taken into account.
[0034] The present invention proposes a method for secure transmission of the physical layer of an integrated network of UAV cluster communication perception based on deep reinforcement learning, aiming to solve the problems mentioned in the above background technology.
[0035] Reference Figure 1 The first embodiment of the present application provides a method for secure transmission of a physical layer of an integrated network of drone cluster communication perception. The method for secure transmission of a physical layer of an integrated network of drone cluster communication perception may include:
[0036] S10, obtaining initialization data in a communication perception integrated data transmission scenario, wherein the initialization data includes an initial position of each drone in the drone cluster, an initial position of each eavesdropping drone, an initial position of each perception device, and a transmission power of an ISAC base station;
[0037] It is worth noting that, assuming that the ground ISAC base station with communication and perception integration function transmits confidential information to the UAV cluster, the perception targets are all located on the ground, and there are multiple malicious eavesdroppers in the communication environment. In order to improve the security performance of confidential information transmission, the transmit beamforming of the ground ISAC base station and the trajectory of the UAV cluster are jointly optimized to maximize the average confidentiality rate during the communication and perception integration process T, while meeting the perception performance of multiple perception targets.
[0038] First, the communication-sensing integration process is discretized to facilitate subsequent optimization, that is, T = NΔ t , where N is the total number of time slots after discretization, Δ t is a sufficiently short single time slot length. Next, the average confidentiality rate maximization problem is reduced to a joint optimization problem of base station beamforming and UAV cluster trajectory under constraints such as perception performance, UAV speed, inter-UAV safety distance, and ISAC base station transmission power. This optimization problem is difficult to solve using traditional convex optimization algorithms because the objective function is non-convex and the optimization variables are tightly coupled.
[0039] Specifically, refer to Figure 2 The network includes ground-based ISAC base stations (equipped with N t =N x ×N y Element array antenna), U legal UAV individuals (position c at time slot n u [n]=[c x [n],c y [n],c z [n]] T ), E eavesdropping UAVs (position c e ) and K sensing targets located on the ground (position c k ).
[0040] S20, constructing an average confidentiality rate maximization problem, wherein the average confidentiality rate maximization problem is a joint optimization problem of base station beamforming and drone cluster trajectory under constraints;
[0041] S30. Based on the initialization data, an improved proximal strategy optimization algorithm is used to solve the average confidentiality rate maximization problem and obtain a collaborative control strategy, wherein the collaborative control strategy is used to simultaneously optimize the beamforming of the base station and the flight trajectory of the drone cluster to securely transmit the physical layer data between the sensing device and the drone cluster.
[0042] This application uses the joint optimization of the transmit beamforming of the ISAC base station and the trajectory of the UAV cluster to artificially design the channel state information of the legitimate channel and the eavesdropping channel, thereby maximizing the quality difference between the legitimate channel and the eavesdropping channel, and improving the security performance of the transmission while ensuring the perception performance. Simulation experiments show that the present invention can effectively optimize the transmit beamforming of the base station and the trajectory of the UAV cluster, and improve the physical layer security performance of the UAV cluster ISAC communication network transmission
[0043] In addition, this application designs an optimization algorithm based on Deep Reinforcement Learning (DRL) to solve the problem of maximizing the average confidentiality rate. DRL is essentially derived from dynamic programming and is good at dealing with sequence optimization problems. It should be noted that the use of an improved proximal policy optimization algorithm in this application to solve the problem of maximizing the average confidentiality rate does not require approximation of the optimization objectives and constraints, and the algorithm complexity is lower, the optimization time is shorter, and it is suitable for online deployment. Furthermore, the policy network after training can also be used as an initial model for solving similar problems, thereby providing migration application capabilities.
[0044] The following is a further description of the physical layer secure transmission method for the integrated network of drone cluster communication perception based on deep reinforcement learning proposed in this application through the accompanying drawings and examples.
[0045] In an embodiment of the present application, the process of constructing the average confidentiality rate maximization problem includes:
[0046] Determine a first transmission rate at which the ISAC base station transmits confidential data to each drone based on the received signal of each drone, and determine a second transmission rate at which each eavesdropping drone eavesdrops confidential data on each drone based on the received signal of each eavesdropping drone;
[0047] Among them, based on the receiving signal model of the UAV, the transmission rate (first transmission rate) of the u-th legal UAV at time slot n is obtained as:
[0048]
[0049] In the formula, f u [n] is the u-th column of the beamforming matrix F[n], corresponding to the beamforming vector of the ISAC base station to the u-th legal UAV. Accordingly, the transmission rate (second transmission rate) of the e-th eavesdropping UAV to eavesdrop on the confidential signal of the u-th legal UAV is:
[0050]
[0051] Determine the confidentiality transmission rate of each drone based on the first transmission rate and the second transmission rate, and determine the maximum value expression of the average confidentiality rate of the cluster drones based on the confidentiality transmission rate of each drone;
[0052] Construct the constraint conditions corresponding to the maximum expression of average confidentiality rate;
[0053] In the embodiment of the present application, the constraint conditions corresponding to the expression of the maximum value of the average confidentiality rate are constructed, including:
[0054] Determine the speed constraint of each UAV based on the relative position of each UAV;
[0055] Based on the displacement of each UAV, determine the safety distance constraint conditions of each UAV;
[0056] Based on the norm of the beam matrix, determine the transmit power constraint of the ISAC base station;
[0057] Based on the transmission pattern of the ISAC base station in azimuth and elevation, determine the signal strength constraints of each sensing device;
[0058] The constraints corresponding to the expression of the maximum value of the average confidentiality rate are determined based on the speed constraints, the safety distance constraints, the transmission power constraints and the signal strength constraints.
[0059] The average confidentiality rate maximization problem is constructed based on the maximum value expression and constraints.
[0060] For example, assuming the worst case that all eavesdropping UAVs can eavesdrop on the confidential signals of each legitimate UAV, the confidential transmission rate of the uth legitimate UAV at time slot n can be expressed as:
[0061]
[0062] Where, [x] + It means to take the maximum value between x and 0.
[0063] Based on the confidentiality transmission rate, we can obtain the maximum expression of the average confidentiality rate and the corresponding constraints:
[0064]
[0065] Where V max is the maximum flight speed of a legal UAV, D min represents the safe distance between any two legal UAVs, c u,0 Specify the initial position of the uth legal UAV, P t is the transmission power of the ground base station, To meet the threshold of target k perception performance.
[0066] In an embodiment of the present application, before determining the first transmission rate at which the ISAC base station transmits confidential data to each drone based on the received signal of each drone, and determining the second transmission rate at which each eavesdropping drone eavesdrops on confidential data of each drone based on the received signal of each eavesdropping drone, the drone cluster communication perception integrated network physical layer security transmission method further includes:
[0067] Obtain the confidential communication signal vector and the sensing signal vector of the ISAC base station in any time slot, as well as the communication beamforming matrix and the sensing beamforming matrix;
[0068] Obtaining a first sub-transmit signal based on a communication beamforming matrix and a secure communication signal vector;
[0069] Obtaining a second sub-transmit signal based on the sensing beamforming matrix and the sensing signal vector;
[0070] Obtaining a transmission signal of the ISAC base station based on the first sub-transmission signal and the second sub-transmission signal;
[0071] Based on the channel gain of each UAV, the transmission signal of the ISAC base station and the additive Gaussian white noise of each UAV, the receiving signal of each UAV is determined accordingly, and based on the channel gain of each eavesdropping UAV, the transmission signal of the ISAC base station and the additive Gaussian white noise of each eavesdropping UAV, the receiving signal of each eavesdropping UAV is determined accordingly, wherein the channel gain of each UAV and the channel gain of each eavesdropping UAV are determined based on the path loss, the small-scale signal gain and the transmitting array steering vector, and the array element spacing of the azimuth and pitch angles of the transmitting array steering vector is equal to half the wavelength.
[0072] The processor in the server or terminal obtains the receiving signal models of the drone cluster and the eavesdropping drone respectively by executing the above steps.
[0073] The ground base station transmits a signal x[n]=F at time slot n. c [n]s c [n]+F s [n]s s [n], where and They represent the communication beamforming matrix and the sensing beamforming matrix at time slot n, respectively, and s c [n]∈C U×1 and are the secure communication signal vector and the perception signal vector at time slot n, respectively, and satisfy In this way, the received signals of the u-th legitimate UAV and the e-th eavesdropping UAV at time slot n can be expressed as:
[0074]
[0075] in are the channel gain vectors from the base station to the u-th legitimate UAV and the e-th eavesdropping UAV at time slot n, respectively. are the additive white Gaussian noise (AWGN) at the uth legitimate UAV and the e-th eavesdropping UAV at time slot n, respectively. The legitimate channel and the eavesdropping channel are modeled as millimeter wave channels h i [n] = β 0 α 0 a(φ i [n],θ i [n]), where i∈{U,E} represents a legitimate UAV or an eavesdropping UAV, β 0 is the LoS path loss, is the small-scale channel gain, and a(φ,θ) is the transmitting array steering vector corresponding to the array element spacing equal to half the wavelength at azimuth angle φ and elevation angle θ:
[0076]
[0077] On this basis, the perception model of the perception device is established. For the perception target with the orientation (θ, φ), the transmission pattern of the ground base station at time slot n can be expressed as:
[0078] P φ,θ [n] = a H (φ,θ)(F[n]F H [n])a(φ,θ),
[0079] When constraining the perceived performance, it is required that P φ,θ [n] is greater than a certain threshold, that is, the constraint on perception performance is transformed into a constraint on the signal strength in each perceived target direction.
[0080] The processor in the server or terminal can determine the average confidentiality rate maximization problem and the corresponding constraints by obtaining the receiving signal models of the drone cluster and the eavesdropping drone and the perception model of the perception device.
[0081] Since the average confidentiality rate maximization problem involves non-convex objective functions and constraints, the optimization problem is non-convex, and the beamforming vector and trajectory of the UAV are tightly coupled, so it is difficult to solve it through traditional optimization methods such as CVX and alternating optimization. Therefore, we propose a joint optimization method for ground base station transmit beamforming and UAV cluster trajectory based on PPO. First, the optimization problem is transformed into a Markov decision process (MDP), and then the corresponding optimization method is designed based on the transformed MDP. The transformed MDP is described in detail below.
[0082] In the embodiment of the present application, an improved proximal strategy optimization algorithm is used to solve the average confidentiality rate maximization problem, and a collaborative control strategy is obtained, including:
[0083] Determine the Markov decision process based on the average confidentiality rate maximization problem;
[0084] The Markov decision process is optimized based on the proximal policy optimization algorithm obtained after interaction and updating with the environment to obtain a collaborative control strategy.
[0085] Specifically, a Markov decision process is determined based on the average confidentiality rate maximization problem, including:
[0086] At the current time slot, the state space of the Markov decision process is determined based on the channel state information and position information of each eavesdropping UAV and each UAV in the UAV cluster;
[0087] The state space S of the MDP should contain information that can guide the agent to plan its own actions. The state space of the nth step can be expressed as s n ={h u [n],h e [n],c 1 [n-1],…,c U [n-1],c 1 ,…,c E ,c 1 ,…,c K}, whose dimension is 2(U+E)N t +3(U+E+K).
[0088] The action space of the Markov decision process is determined based on the beamforming matrix of the base station and the trajectory of the drone cluster;
[0089] The action space of the MDP should contain the optimization variables of the problem under consideration, namely the beamforming matrix F[n] and the UAV cluster trajectory c u [n]. Therefore, the action of step n can be expressed as a n={F[n],c 1 [n],…,c U [n]}. F[n] is a complex number, so this part of the action space is divided into the real part and the imaginary part In addition, we replace the trajectory of the UAV with its speed along the x-axis and y-axis. After the above transformation, the final action is Thus, the dimension of the MDP action space is
[0090] And the total reward of the Markov decision process is determined based on the sum of the total confidentiality rate between the UAVs, the perception performance return of each perception device and the anti-collision return between the UAVs.
[0091] Considering the characteristics of maximizing the total reward in the MDP and DRL optimization process, we take the total confidentiality rate between legal UAVs at each step As the total return R n In addition, considering the perceived performance constraints We introduce perceptual performance rewards in the form of a penalty function Incorporate the perceived performance corresponding to target k into the reward:
[0092]
[0093] This means that if the perception performance of the kth perception target is not met, a penalty of -η is added to the total reward. Finally, define the anti-collision reward R col [n] Prevent legal UAVs from colliding with each other:
[0094]
[0095] The total return can be expressed as
[0096] In an embodiment of the present application, the proximal policy optimization algorithm is determined based on the policy network and the value network, and the proximal policy optimization algorithm obtained after interaction and updating with the environment optimizes the Markov decision process to obtain a collaborative control strategy, including:
[0097] Based on the state space, action space and total reward, determine the probability distribution of selecting action a and the corresponding reward probability in the current state s, and use the policy network to characterize the probability distribution, where the policy network refers to a parameterized neural network;
[0098] Based on the expected probability of reward obtained by the agent when performing action a in the current state s, the action value function is obtained; based on the expected probability of reward obtained by the agent when performing action a according to the optimal strategy in the current state s, the state value function is obtained, and the advantage function is obtained based on the difference between the action value function and the state value function, where the state value function is obtained by approximating the state value corresponding to the current state based on the parameterized value network;
[0099] The advantage function is estimated by using the generalized advantage estimation algorithm to obtain the generalized advantage estimation function;
[0100] The objective function of the policy network is determined by using the generalized advantage estimation function weighted by the policy ratio and the generalized advantage estimation function weighted by the clipping function;
[0101] Determine the objective function of the value network based on the generalized advantage estimation function;
[0102] The collaborative control strategy of the intelligent agent is determined based on the objective function of the policy network, the objective function of the value network and the interaction data between the intelligent agent and the environment.
[0103] Specifically, after completing the transformation of MDP, we give a joint optimization algorithm based on PPO. The PPO algorithm is essentially a policy-based DRL algorithm, that is, the agent optimizes the action execution strategy π(a|s) based on its experience gained from exploring the environment. π(a|s) represents the probability distribution of selecting action a in state s. The goal of the agent is to maximize the value of state s by optimizing the strategy π(a|s), that is, the action value function Q π The mathematical expectation of (s,a). The parameterized neural network represents the policy π(a|s) as The objective function of the strategy optimization can be expressed as Calculate its The gradient of is:
[0104]
[0105] In order to improve the efficiency of updating policy network parameters, the PPO algorithm introduces the old strategy The gradient is further processed as follows:
[0106]
[0107] At the same time, in order to improve learning efficiency and increase the stability of the learning process, the generalized advantage estimation (GAE) is used to estimate the advantage function, and GAE is used with the state s n The difference in value function replaces the action value function in the above formula
[0108]
[0109] in Indicates the use of The parameterized value network for state s n approximate the value of .
[0110] The objective function of strategy optimization is:
[0111]
[0112] in is the strategy ratio, that is, the objective function for updating the strategy network parameters, and the objective function for updating the value network parameters is:
[0113]
[0114] After the above improvements, the PPO algorithm has good robustness and convergence performance compared with other policy-based DRL algorithms, and is suitable for solving optimization problems with high-dimensional parameter space.
[0115] In this implementation, the present application first reduces the average confidentiality rate maximization problem to a Markov decision process, where the state s at time slot n is n is the channel state information of the UAV cluster and the eavesdropper in the current time slot n and the location of all nodes in the network. Action a n is the beamforming matrix of the base station and the location of the UAV cluster, reporting r n It can be divided into two parts, namely the total confidentiality rate that the UAV cluster can obtain in the current time slot and the perception performance constraint introduced in the form of a penalty function. Through such a setting, our agent can optimize the action space with the goal of maximizing the cumulative return on the basis of satisfying all constraints, that is, maximizing the average confidentiality rate between UAV clusters. The algorithm adopted is improved on the proximal policy optimization (PPO) to adapt it to this optimization problem. The policy network is used to give the execution action for the current state, and the value network is used to evaluate the pros and cons of the action given by the policy network. The objective function for updating the policy network parameters is:
[0116]
[0117] In the formula, represents the difference between the old and new strategies, is the generalized advantage estimate, clip is the clipping function, and ò is the clipping hyperparameter. The objective function for updating the value network parameters is:
[0118]
[0119] In the formula, Indicates state s n Through the continuous interaction between the agent and the environment, data is collected and the policy network and value network are updated to obtain the optimal strategy, and finally a physical layer secure transmission method for UAV cluster communication perception integration is given.
[0120] The specific effects of the present invention can be further illustrated by the following simulation:
[0121] Assume that the number of legal UAV individuals is U=4, and the initial position of each individual is c 1,0 =[30,30,40] T , c 2,0 =[45,20,35] T , c 3,0 =[35,25,40] T and c 4,0 =[40,30,35] T The number of eavesdropping UAVs is E = 3, and their positions are c 1 =[40,50,30] T , c 2 =[45,45,50] T and c 3 =[30,45,55] T The number of perceived targets is K = 4, and their positions are c 1 =[30,10,0] T , c 2 =[50,40,0] T , c 3 =[45,60,0] T and c 4 =[35,20,0] T The ISAC base station is equipped with a 4×4 element array antenna. Assume that the perception threshold of each perception target is equal, and the variance of AWGN at each legitimate UAV and each eavesdropping UAV is also equal. Experiments were conducted in a simulation environment, and the three-dimensional transmission beam pattern of the ISAC base station and the trajectory diagram of the UAV cluster at the 200th time slot obtained by the method proposed in the present invention after the optimization process is completed are obtained, as shown in Figure 2. Figure 3 and Figure 4 shown.
[0122] refer to Figure 5 Based on the above embodiments, the present application also provides a physical layer security transmission device for integrated drone cluster communication perception. The physical layer security transmission device 100 includes an acquisition module 101, a question construction module 102 and a transmission control module 103:
[0123] The acquisition module 101 is used to acquire initialization data in the communication perception integrated data transmission scenario, wherein the initialization data includes the initial position of each drone in the drone cluster, the initial position of each eavesdropping drone, the initial position of each perception device, and the transmission power of the ISAC base station;
[0124] The problem construction module 102 is used to construct an average confidentiality rate maximization problem, wherein the average confidentiality rate maximization problem is a joint optimization problem of base station beamforming and drone cluster trajectory under constraints;
[0125] The transmission control module 103 is used to solve the average confidentiality rate maximization problem based on the initialization data and adopt an improved proximal strategy optimization algorithm to obtain a collaborative control strategy. Based on the collaborative control strategy, the beamforming of the base station and the flight trajectory of the drone cluster are simultaneously optimized to securely transmit the physical layer data between the sensing device and the drone cluster.
[0126] Based on the above embodiments, the present application also provides a computer-readable storage medium, which includes instructions. When the instructions are executed on a computer, the computer executes the drone cluster communication perception integrated network physical layer security transmission method provided in any of the above embodiments.
[0127] Based on the above embodiments, the present application further provides an electronic device, the electronic device comprising:
[0128] at least one processor, memory, and input-output unit;
[0129] Among them, the memory is used to store computer programs, and the processor is used to call the computer programs stored in the memory to execute the drone cluster communication perception integrated network physical layer security transmission method provided in the above embodiment.
[0130] The above are only preferred embodiments of the present application, and are not intended to limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A method for secure transmission of physical layer of integrated network of UAV cluster communication perception, characterized in that: include: Acquire initialization data in a communication-sensing integrated data transmission scenario, wherein the initialization data includes the initial position of each drone in the drone cluster, the initial position of each eavesdropping drone, the initial position of each sensing device, and the transmission power of the ISAC base station; Constructing an average confidentiality rate maximization problem, wherein the average confidentiality rate maximization problem is a joint optimization problem of base station beamforming and drone cluster trajectory under constraints; Based on the initialization data, an improved proximal strategy optimization algorithm is used to solve the average confidentiality rate maximization problem to obtain a collaborative control strategy, wherein the collaborative control strategy is used to simultaneously optimize the beamforming of the base station and the flight trajectory of the drone cluster to securely transmit the physical layer data between the sensing device and the drone cluster.
2. The method for secure transmission of the physical layer of the integrated network of drone cluster communication perception as claimed in claim 1, characterized in that: The process of constructing the average confidentiality rate maximization problem includes: Determine a first transmission rate at which the ISAC base station transmits confidential data to each of the drones based on the received signals of each of the drones, and determine a second transmission rate at which each of the eavesdropping drones eavesdrops on the confidential data of each of the drones based on the received signals of each of the eavesdropping drones; Determine the confidentiality transmission rate of each of the drones based on the first transmission rate and the second transmission rate, and determine the maximum value expression of the average confidentiality rate of the cluster drones based on the confidentiality transmission rate of each of the drones; Constructing the constraint condition corresponding to the maximum value expression of the average confidentiality rate; The average confidentiality rate maximization problem is constructed based on the maximum value expression and the constraint conditions.
3. The method for secure transmission of the physical layer of the integrated network of drone cluster communication perception as claimed in claim 2, characterized in that: The constraint conditions corresponding to constructing the expression of the maximum value of the average confidentiality rate include: Determining a speed constraint condition for each of the drones based on the relative positions of each of the drones; Determining a safety distance constraint condition for each of the drones based on the displacement of each of the drones; Determining a transmit power constraint of the ISAC base station based on a norm of the beam matrix; Determining the signal strength constraint condition of each of the sensing devices based on the transmission pattern of the ISAC base station at the azimuth and elevation angles; The constraint condition corresponding to the average confidentiality rate maximum value expression is determined based on the speed constraint condition, the safety distance constraint condition, the transmission power constraint condition and the signal strength constraint condition.
4. The method for secure transmission of the physical layer of the integrated network of drone cluster communication perception as claimed in claim 2, characterized in that: Before determining the first transmission rate at which the ISAC base station transmits confidential data to each of the drones based on the received signals of each of the drones, and determining the second transmission rate at which each of the eavesdropping drones eavesdrops on the confidential data of each of the drones based on the received signals of each of the eavesdropping drones, the drone cluster communication perception integrated network physical layer security transmission method further includes: Obtaining a confidential communication signal vector and a perception signal vector, as well as a communication beamforming matrix and a perception beamforming matrix of the ISAC base station in any time slot; Obtaining a first sub-transmit signal based on the communication beamforming matrix and the secure communication signal vector; Obtaining a second sub-transmit signal based on the sensing beamforming matrix and the sensing signal vector; Obtaining a transmission signal of the ISAC base station based on the first sub-transmission signal and the second sub-transmission signal; Based on the channel gain of each of the UAVs, the transmission signal of the ISAC base station and the additive Gaussian white noise of each of the UAVs, the received signal of each of the UAVs is determined accordingly, and based on the channel gain of each of the eavesdropping UAVs, the transmission signal of the ISAC base station and the additive Gaussian white noise of each of the eavesdropping UAVs, the received signal of each of the eavesdropping UAVs is determined accordingly, wherein the channel gain of each of the UAVs and the channel gain of each of the eavesdropping UAVs are both determined based on path loss, small-scale signal gain and a transmitting array steering vector, and the array element spacing of the azimuth and pitch angles of the transmitting array steering vector is equal to half a wavelength.
5. The method for secure transmission of the physical layer of the integrated network of drone cluster communication perception as claimed in claim 1, characterized in that: The improved proximal strategy optimization algorithm is used to solve the average confidentiality rate maximization problem to obtain a collaborative control strategy, including: Determining a Markov decision process based on the average confidentiality rate maximization problem; The Markov decision process is optimized based on the proximal strategy optimization algorithm obtained after interaction and updating with the environment to obtain the collaborative control strategy.
6. The method for secure transmission of the physical layer of the drone cluster communication perception integrated network as claimed in claim 5, characterized in that: Determining the Markov decision process based on the average confidentiality rate maximization problem includes: In the current time slot, based on the channel state information and the position information of each of the eavesdropping drones and each of the drones in the drone cluster, determine the state space of the Markov decision process; and determining an action space of the Markov decision process based on a beamforming matrix of the base station and a trajectory of the drone cluster; And the total reward of the Markov decision process is determined based on the sum of the total confidentiality rate between the drones, the perception performance return of each perception device and the anti-collision return between the drones.
7. The method for secure transmission of the physical layer of the integrated network of drone cluster communication perception as claimed in claim 6, characterized in that: The proximal policy optimization algorithm is determined based on the policy network and the value network, and the proximal policy optimization algorithm obtained after interaction and updating with the environment optimizes the Markov decision process to obtain the collaborative control strategy, including: Determine a probability distribution of selecting action a and a corresponding reward probability in the current state s based on the state space, the action space and the total reward, and characterize the probability distribution using a policy network, wherein the policy network refers to a parameterized neural network; Based on the expectation of the probability of the reward obtained by the agent when performing action a in the current state s, an action value function is obtained; based on the expectation of the probability of the reward obtained by the agent of the proximal policy optimization algorithm when performing action a according to the optimal strategy in the current state s, a state value function is obtained; and based on the difference between the action value function and the state value function, an advantage function is obtained, wherein the state value function is obtained by approximating the state value corresponding to the current state based on the parameterized value network; The advantage function is estimated by using a generalized advantage estimation algorithm to obtain a generalized advantage estimation function; Determine the objective function of the policy network using a generalized advantage estimation function weighted by the policy ratio and a generalized advantage estimation function weighted by the clipping function; Determining an objective function of the value network based on the generalized advantage estimation function; The collaborative control strategy of the agent is determined based on the objective function of the policy network, the objective function of the value network and the interaction data between the agent and the environment.
8. A physical layer security transmission device for integrated UAV cluster communication perception, characterized in that: include: An acquisition module is used to acquire initialization data in a communication perception integrated data transmission scenario, wherein the initialization data includes an initial position of each drone in the drone cluster, an initial position of each eavesdropping drone, an initial position of each perception device, and a transmission power of an ISAC base station; A problem construction module, used to construct an average confidentiality rate maximization problem, wherein the average confidentiality rate maximization problem is a joint optimization problem of base station beamforming and drone cluster trajectory under constraints; The transmission control module is used to solve the average confidentiality rate maximization problem based on the initialization data by using an improved proximal strategy optimization algorithm to obtain a collaborative control strategy, and to simultaneously optimize the beamforming of the base station and the flight trajectory of the drone cluster based on the collaborative control strategy to securely transmit the physical layer data between the sensing device and the drone cluster.
9. A computer-readable storage medium, characterized in that: It includes instructions, which, when running on a computer, enable the computer to execute the physical layer security transmission method for integrated network of drone cluster communication perception as described in any one of claims 1 to 7.
10. An electronic device, characterized in that: The electronic device comprises: at least one processor, memory, and input-output unit; The memory is used to store computer programs, and the processor is used to call the computer programs stored in the memory to execute the physical layer security transmission method for integrated network of drone cluster communication perception according to any one of claims 1 to 7.
Citation Information
Cited By
ISAC system secure transmission method based on joint beamforming and trajectory optimization
CN121077515A
ISAC system beam forming and trajectory joint optimization method based on three-dimensional space secure transmission
CN121770572A