Unmanned aerial vehicle secrecy energy efficiency optimization method and system based on multi-agent deep reinforcement learning
Through deep reinforcement learning of multiple agents, the reflection coefficient matrix of drone trajectory and base station beamforming is optimized, and the confidential energy efficiency problem in dynamic environments in drone network communication is solved, and efficient confidential energy efficiency optimization and timely response are achieved.
Patent Information
- Application Number
- CN202510529975.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-08-08
AI Technical Summary
In drone network communication, the prior art is difficult to respond in a timely manner and optimize the drone trajectory and multi-functional reconstructible intelligent surface assisted confidential energy efficiency in a dynamic environment, resulting in communication delay and security issues.
A method based on deep reinforcement learning of multiple agents is adopted to build a confidential energy efficiency optimization model, and the reflection coefficient matrix of intelligent surface-assisted by joint optimization of the drone trajectory, base station transmission beamforming and multi-functional reconstructible intelligent surface-assisted reflection coefficient matrix is achieved, combining Markov decision-making process model and soft actor-critician algorithm to achieve optimal confidential energy efficiency.
It improves the confidentiality energy efficiency of legitimate users, can respond to dynamic environmental changes in a timely manner, reduce communication delays, and improves system flexibility and security.
Smart Images

Figure CN120454815A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of drone network communication energy efficiency technology, and in particular to a drone confidentiality energy efficiency optimization method and system based on multi-agent deep reinforcement learning. Background Art
[0002] Reconfigurable smart surfaces (RIS) can intelligently adjust the phase and amplitude of the input signal to improve the signal strength received by the receiver. Due to the open nature of communication links, physical layer security has become a must-solve issue. Wireless communications can be protected by increasing the capacity of legitimate channels while reducing the decoding performance of eavesdroppers.
[0003] However, obstructions often degrade communication quality. Deploying drones carrying RIS can improve system flexibility and performance. Power splitting (PS), a key technology in simultaneous wireless information and power transfer (SWIPT), splits received signals into an information decoding stream and an energy harvesting stream, enabling efficient deployment of energy-constrained devices. Unlike traditional single-function RIS (SF-RIS) limited to reflection / refraction, self-sustaining RIS integrating PS technology enhances flexibility in signal resource allocation.
[0004] Deep reinforcement learning (DRL) can adapt well to the time-varying wireless environment of drones through the interaction between the agent and the environment. For example, a DRL algorithm called FlyReflect was proposed in the previous technology, which can jointly optimize the drone trajectory and RIS phase shift to maximize the system's total efficiency. However, for multi-device collaborative optimization systems involving drones, drones need to make decisions within milliseconds when performing tasks. Deploying neural networks directly on drones can reduce communication delays with other devices, but if the neural network is deployed on a base station (BS), the communication delay will affect the real-time performance of the system, especially in dynamic environments. Summary of the Invention
[0005] In order to solve the above technical problems, the purpose of the present invention is to provide a drone confidentiality energy efficiency optimization method and system based on multi-agent deep reinforcement learning, which can improve confidentiality energy efficiency and respond to the needs of dynamic environments in a timely manner.
[0006] To achieve the above objectives, one aspect of the present application proposes a method for optimizing drone security energy efficiency based on multi-agent deep reinforcement learning, comprising the following steps:
[0007] Based on the UAV trajectory, base station transmit beamforming, and the reflection coefficient matrix of the multifunctional reconfigurable intelligent surface-assisted UAV, a confidentiality energy efficiency optimization model for the multifunctional reconfigurable intelligent surface-assisted UAV secure communication system is constructed;
[0008] Constructing a Markov decision process model based on the confidentiality energy efficiency optimization model, wherein the Markov decision process model includes a state space, a base station action space, a multifunctional reconfigurable intelligent surface assisted UAV action space, and a reward function;
[0009] Based on a multi-agent deep reinforcement learning algorithm, the Markov decision process model is iteratively optimized to obtain the optimal confidentiality energy efficiency.
[0010] In some embodiments, the construction of a confidentiality energy efficiency optimization model for a multifunctional reconfigurable intelligent surface-assisted UAV secure communication system based on the UAV trajectory, base station transmit beamforming, and the reflection coefficient matrix of the multifunctional reconfigurable intelligent surface-assisted UAV specifically includes:
[0011] The multifunctional reconfigurable intelligent surface-assisted UAV secure communication system is defined to include a base station, a UAV, an eavesdropper, and several legitimate users, wherein the UAV is the multifunctional reconfigurable intelligent surface-assisted UAV;
[0012] defining the trajectory of the drone and the flight speed of the drone in each time slot, and determining the propulsion energy of the drone in each time slot based on the flight speed;
[0013] defining a transmit beamforming vector of the legal user in each time slot;
[0014] Defining the reflection coefficient matrix according to the phase shift and amplification coefficient of the smart surface of the reflection unit;
[0015] Defining a channel gain matrix between the base station and the drone and a reflection channel gain vector between the drone and the legitimate user;
[0016] determining, based on the amplification factor, the transmit beamforming vector, and the channel gain matrix, that an element acquires energy;
[0017] Determining an achievable rate for the legitimate user based on the reflection channel gain vector, the reflection coefficient matrix, the channel gain matrix, and the transmit beamforming vector;
[0018] Determining an eavesdropping rate of the legitimate user by the eavesdropper according to the achievable rate;
[0019] determining confidentiality energy efficiency according to the propulsion energy, the achievable rate, and the eavesdropping rate;
[0020] The confidentiality energy efficiency, the UAV trajectory, the transmit beamforming vector, the reflection coefficient matrix, the component acquisition energy, and the channel gain matrix are integrated to obtain the confidentiality energy efficiency optimization model.
[0021] In some embodiments, the confidentiality energy efficiency optimization model is:
[0022]
[0023] Among them, SEE n represents the confidentiality efficiency, L and H represent the trajectory of the drone and represents the position of the UAV in time slot n and n represents the nth time slot and W represents the base station transmit beamforming and represents the transmit beamforming vector of the kth legal user in time slot n, k represents the kth legal user and k∈K, P BS represents the base station transmit beamforming, Θ represents the reflection coefficient matrix and and Indicates that the component obtains energy, represents the amplification factor, represents the i-th row in the channel gain matrix, P RIS represents the maximum transmission power, ξ represents the energy consumption conversion efficiency and ξ∈(0,1].
[0024] In some embodiments, constructing a Markov decision process model based on the confidentiality energy efficiency optimization model specifically includes:
[0025] Determining the state space according to the UAV trajectory and the channel gain matrix;
[0026] determining the base station action space according to the transmit beamforming vector;
[0027] Determining an action space of the multifunctional reconfigurable intelligent surface-assisted UAV according to the flight distance, elevation angle, azimuth angle of the UAV and the reflection coefficient matrix;
[0028] defining a penalty for the drone flying out of a legal area, and determining the reward function according to the confidentiality energy efficiency and the penalty for flying out of the legal area;
[0029] The state space, the base station action space, the multifunctional reconfigurable intelligent surface assisted UAV action space and the reward function are integrated to obtain the Markov decision process model.
[0030] In some embodiments, the drone confidentiality energy efficiency optimization method further includes:
[0031] The multi-agent deep reinforcement learning algorithm is defined to include a base station agent, a multifunctional reconfigurable intelligent surface assisted UAV agent, and a MIX network, wherein the base station agent includes a first actor network and a first critic network, and the multifunctional reconfigurable intelligent surface assisted UAV agent includes a second actor network and a second critic network;
[0032] The first critic network and the second critic network are combined through the MIX network to obtain a QMIX network.
[0033] In some embodiments, the multi-agent deep reinforcement learning algorithm is used to iteratively optimize the Markov decision process model to obtain optimal confidentiality efficiency, specifically including:
[0034] Adjusting the Markov decision process model to obtain a Markov decision process optimization model;
[0035] outputting a policy distribution through the first actor network and the second actor network based on the Markov decision process optimization model;
[0036] Outputting a first local Q value according to the policy distribution through the first critic network;
[0037] Outputting a second local Q value according to the policy distribution through the second critic network;
[0038] Combining the first local Q value and the second local Q value through the QMIX network to obtain a joint Q value;
[0039] updating the first actor network, the first critic network, the second actor network, and the second critic network according to the joint Q value;
[0040] When the preset policy update condition is reached, the update is stopped and the optimal confidentiality energy efficiency is output.
[0041] In some embodiments, adjusting the Markov decision process model to obtain a Markov decision process optimization model specifically includes:
[0042] Normalizing the transmit beamforming vector in the base station action space according to the confidentiality energy efficiency optimization model;
[0043] Adjusting the reflection coefficient matrix in the multifunctional reconfigurable intelligent surface-assisted UAV action space according to the confidential energy efficiency optimization model;
[0044] The Markov decision process optimization model is obtained according to the normalized base station action space and the adjusted multifunctional reconfigurable intelligent surface assisted UAV action space.
[0045] To achieve the above objectives, another aspect of the present application provides a drone security energy efficiency optimization system based on multi-agent deep reinforcement learning, comprising:
[0046] The first module is used to build a confidentiality energy efficiency optimization model for the multifunctional reconfigurable intelligent surface-assisted UAV secure communication system based on the UAV trajectory, base station transmit beamforming, and the reflection coefficient matrix of the multifunctional reconfigurable intelligent surface-assisted UAV;
[0047] The second module is used to construct a Markov decision process model based on the confidentiality energy efficiency optimization model, wherein the Markov decision process model includes a state space, a base station action space, a multifunctional reconfigurable intelligent surface assisted UAV action space, and a reward function;
[0048] The third module is used to iteratively optimize the Markov decision process model based on a multi-agent deep reinforcement learning algorithm to obtain the optimal confidentiality energy efficiency.
[0049] To achieve the above-mentioned purpose, another aspect of an embodiment of the present application proposes an electronic device, which includes a memory, a processor, a program stored on the memory and executable on the processor, and a data bus for realizing connection and communication between the processor and the memory. When the program is executed by the processor, the method for optimizing the confidentiality and energy efficiency of drones based on multi-agent deep reinforcement learning as described above is realized.
[0050] To achieve the above-mentioned purpose, another aspect of an embodiment of the present application proposes a storage medium, which is a computer-readable storage medium for computer-readable storage. The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the drone confidentiality energy efficiency optimization method based on multi-agent deep reinforcement learning as described above.
[0051] The beneficial effects of the present invention are as follows: the present invention's method and system for optimizing the confidentiality energy efficiency of drones based on multi-agent deep reinforcement learning first constructs a confidentiality energy efficiency optimization model for a multi-functional reconfigurable intelligent surface-assisted drone secure communication system based on the drone trajectory, base station transmit beamforming, and the multi-functional reconfigurable intelligent surface-assisted drone's reflection coefficient matrix. Then, based on the confidentiality energy efficiency optimization model, a Markov decision process model is constructed. The Markov decision process model includes a state space, a base station action space, a multi-functional reconfigurable intelligent surface-assisted drone action space, and a reward function. Finally, based on a multi-agent deep reinforcement learning algorithm, the Markov decision process model is solved to obtain the optimal confidentiality energy efficiency. On the one hand, the present invention maximizes the confidentiality energy efficiency of legitimate users by jointly optimizing the drone's flight trajectory, base station transmit beamforming, and the multi-functional reconfigurable intelligent surface-assisted drone's reflection coefficient matrix. On the other hand, considering the need to respond promptly to dynamic environments and the fact that the problem is multivariable coupled and non-convex optimization, a fully cooperative multi-agent deep reinforcement learning algorithm (MADRL) based on a soft actor-critic (SAC) is proposed to find the optimal learning strategy, which can significantly improve confidentiality energy efficiency while responding promptly to the needs of dynamic environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following introduction is made to the drawings required for use in the embodiments of the present invention. It should be understood that the drawings introduced below are only for the convenience of clearly describing some embodiments of the technical solutions of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative work.
[0053] Figure 1 A flowchart of the steps of a method for optimizing the confidentiality and energy efficiency of drones based on multi-agent deep reinforcement learning according to an embodiment of the present invention;
[0054] Figure 2 A schematic diagram of the structure of a multifunctional reconfigurable intelligent surface-assisted UAV secure communication system provided by an embodiment of the present invention;
[0055] Figure 3 A schematic diagram of the framework of a multi-agent deep reinforcement learning algorithm provided by one embodiment of the present invention;
[0056] Figure 4 A comparison chart of rewards for different algorithms provided in one embodiment of the present invention;
[0057] Figure 5 A two-dimensional trajectory diagram of a drone provided by one embodiment of the present invention;
[0058] Figure 6A three-dimensional trajectory diagram of a drone provided by one embodiment of the present invention;
[0059] Figure 7 A comparison chart of confidentiality energy efficiency at different transmission powers provided by an embodiment of the present invention;
[0060] Figure 8 A schematic diagram of the structure of a drone security energy efficiency optimization system based on multi-agent deep reinforcement learning provided by an embodiment of the present invention;
[0061] Figure 9 A schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0062] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present application. They are merely examples of devices and methods consistent with some aspects of the embodiments of the present application as detailed in the appended claims.
[0063] It will be understood that the terms "first", "second", etc. used in this application may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0064] The terms "at least one", "plurality", "each", "any", etc. used in this application include "at least one", "two" or more, "plurality" or "each", "any" or "any one", "each" or "any one" as used herein.
[0065] Before explaining the embodiments of the present application in detail, some of the nouns and terms involved in the embodiments of the present application are first explained. The nouns and terms involved in the embodiments of the present application are subject to the following explanations.
[0066] Multi-Frequency Reconfigurable Intelligent Surface (MF-RIS): A multifunctional reconfigurable intelligent surface is an intelligent material composed of programmable units that can dynamically control the reflection, refraction, and absorption properties of electromagnetic waves (such as wireless signals) while supporting multiple functions (such as communication, sensing, and energy harvesting). It is suitable for scenarios such as 6G and the Internet of Things.
[0067] Simultaneous Wireless Information and Power Transfer (SWIPT): Simultaneous wireless information and power transfer is a technology that transmits data and energy simultaneously through the same wireless signal. It is mainly used in scenarios such as the Internet of Things (IoT), sensor networks, and low-power devices.
[0068] Reconfigurable smart surfaces (RIS) can intelligently adjust the phase and amplitude of the input signal to improve the signal strength received by the receiver. Due to the open nature of communication links, physical layer security has become a must-solve issue. Wireless communications can be protected by increasing the capacity of legitimate channels while reducing the decoding performance of eavesdroppers.
[0069] However, obstructions often degrade communication quality. Deploying drones carrying RIS can improve system flexibility and performance. Power splitting (PS), a key technology in simultaneous wireless information and power transfer (SWIPT), splits received signals into an information decoding stream and an energy harvesting stream, enabling efficient deployment of energy-constrained devices. Unlike traditional single-function RIS (SF-RIS) limited to reflection / refraction, self-sustaining RIS integrating PS technology enhances flexibility in signal resource allocation.
[0070] Deep reinforcement learning (DRL) can adapt well to the time-varying wireless environment of drones through the interaction between the agent and the environment. For example, a DRL algorithm called FlyReflect was proposed in the previous technology, which can jointly optimize the drone trajectory and RIS phase shift to maximize the system's total efficiency. However, for multi-device collaborative optimization systems involving drones, drones need to make decisions within milliseconds when performing tasks. Deploying neural networks directly on drones can reduce communication delays with other devices, but if the neural network is deployed on a base station (BS), the communication delay will affect the real-time performance of the system, especially in dynamic environments.
[0071] To this end, an embodiment of the present invention proposes a method for optimizing the confidentiality energy efficiency of drones based on multi-agent deep reinforcement learning. First, a confidentiality energy efficiency optimization model for a multi-functional reconfigurable intelligent surface-assisted drone secure communication system is constructed based on the drone trajectory, base station transmit beamforming, and the multi-functional reconfigurable intelligent surface-assisted drone reflection coefficient matrix. Next, a Markov decision process model is constructed based on the confidentiality energy efficiency optimization model. The Markov decision process model includes a state space, a base station action space, a multi-functional reconfigurable intelligent surface-assisted drone action space, and a reward function. Finally, the Markov decision process model is solved based on a multi-agent deep reinforcement learning algorithm to obtain the optimal confidentiality energy efficiency. On the one hand, the present invention maximizes the confidentiality energy efficiency of legitimate users by jointly optimizing the drone flight trajectory, base station transmit beamforming, and the multi-functional reconfigurable intelligent surface-assisted drone reflection coefficient matrix. On the other hand, considering the need for timely response to dynamic environments and the fact that the problem involves multivariable coupling and non-convex optimization, a fully cooperative multi-agent deep reinforcement learning algorithm (MADRL) based on a soft actor-critic (SAC) algorithm is proposed to find the optimal learning strategy, which can significantly improve confidentiality energy efficiency while timely responding to the needs of dynamic environments.
[0072] Reference Figure 1 , Figure 1 This is a flowchart of a method for optimizing the confidentiality and energy efficiency of a drone based on multi-agent deep reinforcement learning according to an embodiment of the present invention. The method includes steps S101 to S103:
[0073] S101. Based on the UAV trajectory, base station transmit beamforming, and the reflection coefficient matrix of the multifunctional reconfigurable intelligent surface-assisted UAV, a confidentiality energy efficiency optimization model for the multifunctional reconfigurable intelligent surface-assisted UAV secure communication system is constructed.
[0074] Specifically, embodiments of the present invention propose a drone-carried MF-RIS-assisted secure and energy-efficient SWIPT system (i.e., a multifunctional reconfigurable intelligent surface-assisted drone secure communication system). In this system, a base station communicates with multiple users via cascaded channels in the presence of eavesdroppers. By jointly optimizing the drone trajectory, the base station's transmit beamforming, and the MF-RIS reflection coefficient matrix, the system maximizes security energy efficiency.
[0075] As an optional implementation, step S101 may be specifically divided into the following steps S1011 to S10110:
[0076] S1011. Define a multifunctional reconfigurable intelligent surface-assisted UAV secure communication system, comprising a base station, a UAV, an eavesdropper, and several legitimate users, wherein the UAV is a multifunctional reconfigurable intelligent surface-assisted UAV;
[0077] Specifically, if Figure 2 The figure shows a schematic diagram of a secure communication system for a multifunctional reconfigurable intelligent surface-assisted unmanned aerial vehicle (UAV). The present invention provides a secure communication system for a multifunctional reconfigurable intelligent surface-assisted UAV (hereinafter referred to as F-RIS). The system consists of a base station (BS), a multifunctional reconfigurable intelligent surface (MF-RIS), K single-antenna legitimate users, and a single-antenna eavesdropper. Represented as a user set.
[0078] S1012. Define the trajectory of the drone and the flight speed of the drone in each time slot, and determine the propulsion energy of the drone in each time slot based on the flight speed;
[0079] Specifically, the position of the UAV at time slot n is expressed as in, Indicates a legal flight area.
[0080] At time slot n, the horizontal flight speed of the UAV It can be expressed as the following formula:
[0081]
[0082] Vertical flight speed of the drone It can be expressed as the following formula:
[0083]
[0084] Among them, t n represents the duration of time slot n. Then the propulsion energy of the rotorcraft in each time slot n is It can be expressed as the following formula:
[0085]
[0086] Among them, P0 and P1 are the blade profile power and the induced power in the hovering state respectively; P2 is the descending or ascending power; U tip is the rotor tip speed; v0 is the average induced speed of the rotor in the hovering state; d0 and s are the fuselage drag ratio and rotor stability, respectively; ρ and G represent the air density and rotor disk area, respectively.
[0087] S1013. Define a transmit beamforming vector for a legitimate user in each time slot.
[0088] Specifically, the transmission signal x of the base station (BS) n It is defined as follows:
[0089]
[0090] in, is the transmit beamforming vector of legitimate user k in time slot n, Indicates the transmitted data symbol.
[0091] S1014. Define a reflection coefficient matrix according to the phase shift and amplification coefficient of the smart surface of the reflection unit;
[0092] Specifically, is defined as the F-RIS reflection coefficient matrix, where and i∈{1,2,…,N} are the RIS phase shift and amplification factor of the i-th reflection unit respectively.
[0093] S1015. Define the channel gain matrix between the base station and the drone, and the reflection channel gain vector between the drone and the legitimate user;
[0094] Specifically, it is assumed that the channel gain matrix between the base station (BS) and the multifunctional reconfigurable intelligent surface assisted unmanned aerial vehicle (F-RIS), the reflection channel gain vector between the multifunctional reconfigurable intelligent surface assisted unmanned aerial vehicle (F-RIS) and the legitimate user, and the reflection channel gain vector between the multifunctional reconfigurable intelligent surface assisted unmanned aerial vehicle (F-RIS) and the eavesdropper are respectively expressed as as well as = . Due to the presence of obstacles, the direct link between the base station (BS) and the legitimate user will be ignored. The signal received by the i-th element in the multifunctional reconfigurable intelligent surface assisted drone (F-RIS) It can be expressed as the following formula:
[0095]
[0096] in, represents the equivalent channel from the base station (BS) to the i-th F-RIS element, that is, the channel gain matrix H r The i-th row of .
[0097] S1016. Determine the energy acquired by the element according to the amplification factor, the transmit beamforming vector, and the channel gain matrix;
[0098] Specifically, the controller adjusts the magnification factor according to Control F-RIS components to work in different modes. When , the component can obtain energy from the incident signal through power distribution technology, and the calculation formula is defined as:
[0099]
[0100] in, Indicates that all The total energy collected by the elements; η∈(0,1] represents the energy collection conversion efficiency. When , the signal received by F-RIS needs to be amplified, and the calculation formula is defined as:
[0101]
[0102] in, Indicates that all The sum of the energy consumption of the elements; ξ∈(0,1] represents the energy consumption conversion efficiency.
[0103] S1017. Determine an achievable rate for a legitimate user based on the reflection channel gain vector, the reflection coefficient matrix, the channel gain matrix, and the transmit beamforming vector;
[0104] Specifically, the signal received by the legitimate user k is It can be expressed as the following formula:
[0105]
[0106] in, Expressed as thermal noise on RIS; The signal-to-interference-plus-noise ratio (SINR) of the kth legal user is The calculation formula is:
[0107]
[0108] Then the rate that can be achieved by legitimate user k is for
[0109] S1018. Determine the eavesdropping rate of the eavesdropper on the legitimate user based on the achievable rate;
[0110] Specifically, the eavesdropping rate of k legitimate users is It can be expressed as the following formula:
[0111]
[0112] S1019. Determine confidentiality efficiency based on propulsion energy, achievable rate, and eavesdropping rate;
[0113] Specifically, according to the above formula, the definition of confidentiality energy efficiency SEE is as follows:
[0114]
[0115] Where [x] + =max(0,x).
[0116] S10110, integrate confidentiality energy efficiency, drone trajectory, transmit beamforming vector, reflection coefficient matrix, component acquisition energy and channel gain matrix to obtain a confidentiality energy efficiency optimization model.
[0117] As an optional implementation method, the confidentiality energy efficiency optimization model is:
[0118]
[0119] Among them, SEE n represents confidentiality efficiency, L and H represent the trajectory of the drone and represents the position of the UAV in time slot n and n represents the nth time slot and W represents the base station transmit beamforming and represents the transmit beamforming vector of the kth legal user in time slot n, k represents the kth legal user and k∈K, P BS represents the base station transmit beamforming, Θ represents the reflection coefficient matrix and and Indicates that the component obtains energy, represents the amplification factor, represents the i-th row in the channel gain matrix, P RIS represents the maximum transmission power, ξ represents the energy consumption conversion efficiency and ξ∈(0,1].
[0120] Specifically, the goal of the embodiments of the present invention is to jointly optimize the drone trajectory, base station transmit beamforming, and F-RIS reflection coefficient matrix to maximize the confidentiality energy efficiency of the F-RIS equipped drone. The above confidential energy efficiency optimization model is obtained. Formula (a) expresses the objective function, formula (b) represents the base station transmit power constraint, formula (c) ensures that the acquired energy can replenish the energy consumed in each time slot, formula (d) represents the maximum transmit power constraint of the active RIS unit, and formula (e) represents the UAV flight area constraint.
[0121] S102. Construct a Markov decision process model based on the confidential energy efficiency optimization model. The Markov decision process model includes a state space, a base station action space, a multifunctional reconfigurable intelligent surface-assisted UAV action space, and a reward function.
[0122] Specifically, due to the presence of multivariable coupling and non-convex constraints, traditional optimization algorithms face significant challenges. Therefore, an embodiment of the present invention employs deep reinforcement learning (DRL) and proposes a multi-agent algorithm framework to address non-convex problems. A fully cooperative multi-agent deep reinforcement learning (MADRL) framework is employed, in which a base station (BS) and a multifunctional reconfigurable intelligent surface-assisted unmanned aerial vehicle (F-RIS) serve as independent agents. The problem of maximizing the confidentiality energy rate is formulated as a Markov decision process (MDP).
[0123] As an optional implementation, step S102 may be further divided into the following steps S1021 to S1025:
[0124] S1021. Determine the state space according to the UAV trajectory and the channel gain matrix;
[0125] S1022. Determine a base station action space based on a transmit beamforming vector;
[0126] S1023. Determine an action space for the multifunctional reconfigurable intelligent surface-assisted UAV based on the flight distance, elevation angle, azimuth angle, and reflection coefficient matrix of the UAV.
[0127] S1024. Define a penalty for the drone flying outside the legal area, and determine a reward function based on the confidentiality energy efficiency and the penalty for flying outside the legal area;
[0128] S1025. Integrate the state space, base station action space, multifunctional reconfigurable intelligent surface assisted UAV action space and reward function to obtain the Markov decision process model.
[0129] In some optional embodiments, deep reinforcement learning (DRL) is based on a Markov decision process (MDP), which consists of a state space Action Space Reward Function And the state transition function In the multi-agent algorithm of the embodiment of the present invention, different agents have different action spaces. The Markov decision process model can be represented as 5 tuples, such as
[0130] Specifically, the state s n : For the base station agent and the multifunctional reconfigurable intelligent surface assisted UAV agent, it is assumed that they observe the same state. The state consists of the position of the UAV and the associated CSI, which is defined as:
[0131]
[0132] Base station (BS) action space The BS action includes the BS's transmit beamforming. Then, the BS action is defined as:
[0133]
[0134] Action space of multifunctional reconfigurable intelligent surface-assisted UAV (F-RIS) The F-RIS action consists of the UAV flight action and the MF-RIS reflection coefficient matrix. The flight action is defined in polar coordinates, including the flight distance d n , elevation angle φ n and azimuth Then, the action of F-RIS is defined as:
[0135]
[0136] Reward n To enhance algorithm stability and avoid local optimality, an exponential reward function is designed in this embodiment of the present invention to widen the SEE performance gap. The reward is defined as:
[0137]
[0138] Among them, p o ={0,1} indicates the penalty for whether the drone flies out of the legal area.
[0139] S103. Based on the multi-agent deep reinforcement learning algorithm, the Markov decision process model is iteratively optimized to obtain the optimal confidentiality energy efficiency;
[0140] As an optional implementation, the drone confidentiality energy efficiency optimization method further includes the following steps S201 and S202:
[0141] S201. Define a multi-agent deep reinforcement learning algorithm including a base station agent, a multifunctional reconfigurable intelligent surface assisted UAV agent, and a MIX network. The base station agent includes a first actor network and a first critic network, and the multifunctional reconfigurable intelligent surface assisted UAV agent includes a second actor network and a second critic network.
[0142] S202: Combine the first critic network and the second critic network through the MIX network to obtain a QMIX network.
[0143] Specifically, if Figure 3The figure shows the framework of the multi-agent deep reinforcement learning algorithm. F-RIS and BS each have their own dedicated actor network, while their local dual Q critics (Q1, Q2 select the minimum value) and a MIX network together form a centralized QMIX critic. During the training process, each agent's agent network generates an action distribution The actions taken result in the next state and reward. Experience is stored in a replay buffer and can be sampled in random batches to train the neural network.
[0144] As an optional implementation, step S103 may be specifically divided into the following steps S1031 to S1037:
[0145] S1031. Adjust the Markov decision process model to obtain a Markov decision process optimization model;
[0146] As an optional implementation, step S1031 may be further divided into the following steps S10311 to S10313:
[0147] S10311. Normalize the transmit beamforming vector in the base station action space according to the confidentiality energy efficiency optimization model;
[0148] Specifically, in order to meet the base station transmission power constraint, it is necessary to The action vector is given by Normalization is performed to generate a new transmit beamforming Right now:
[0149]
[0150] S10312. Adjust the reflection coefficient matrix in the action space of the multifunctional reconfigurable intelligent surface-assisted UAV based on a confidential energy efficiency optimization model;
[0151] Specifically, the reflection coefficient also needs to be modified to ensure that the energy obtained can supplement the energy consumed in each time slot. The new reflection coefficient matrix It can be expressed as:
[0152]
[0153] in, Represents the sum of all reflection element coefficients greater than 1 at time slot n.
[0154] if Violates the maximum transmit power constraint, then let
[0155] S10313. Based on the normalized base station action space and the adjusted multifunctional reconfigurable intelligent surface assisted UAV action space, a Markov decision process optimization model is obtained.
[0156] S1032. Based on the Markov decision process optimization model, outputting the strategy distribution through the first actor network and the second actor network;
[0157] S1033. Output a first local Q value according to the strategy distribution through the first critic network;
[0158] S1034. Output a second local Q value according to the strategy distribution through the second critic network;
[0159] S1035 , combining the first local Q value and the second local Q value through a QMIX network to obtain a joint Q value;
[0160] S1036. Update the first actor network, the first critic network, the second actor network, and the second critic network according to the joint Q value;
[0161] S1037. When the preset policy update conditions are met, stop updating and output the optimal confidentiality energy efficiency.
[0162] Specifically, the maximum entropy-based deep reinforcement learning algorithm (SAC) enhances the standard DRL by simultaneously maximizing the cumulative reward and policy entropy within the maximum entropy framework. This approach enhances the agent's exploration ability, thereby efficiently discovering the optimal solution in complex environments, which can be formulated as:
[0163]
[0164] in, Representative strategy π i , i∈{BS,F-RIS}, the state-action marginal distribution of the trajectory caused by. H(π i (·|s n )) indicates the adoption of the current strategy π i Entropy, α i represents the temperature parameter, which is automatically updated by optimizing the following loss:
[0165]
[0166] in, represents the entropy of the expected target, Represents the policy distribution output by the neural network at the nth time slot. The critic network evaluates the actions taken by the agent by learning a value function and provides feedback for these actions, whose value is:
[0167]
[0168] in, γ∈[0,1] represents the discount factor, which reflects the importance of future and immediate rewards; p represents the probability of entering the next state; V(s n ) is a function that evaluates the expected return of taking a specific strategy in a given state. The critic networks of multiple agents are combined into a QMIX network through the MIX network. The QMIX network inputs the state of the nth time slot and the local Q value function of each agent, and outputs the joint Q value Q tot , expressed as follows:
[0169]
[0170] Among them, q i represents the local critic network, q mix Represents a MIX network, which is based on the state s n Given weights W and bias b, multiple local Q values are combined.
[0171] By ψ Q Parameterized Q function It is defined by a neural network and trained by minimizing the mean squared error between the estimated and target soft Q values.
[0172]
[0173] in, Represents the experience replay pool, represents the target network, which reduces the overestimation of the Q-value function while enhancing the convergence stability.
[0174] Theoretical guarantees confirm the convergence of soft policy iteration to the optimal policy, and the neural network parameterization ψ π The policy update target (policy update condition) is defined as:
[0175]
[0176] The above describes an embodiment of the present invention's method for optimizing drone security energy efficiency based on multi-agent deep reinforcement learning. It can be appreciated that the present invention maximizes the security energy efficiency of legitimate users by jointly optimizing the drone's flight trajectory, base station transmit beamforming, and the multifunctional reconfigurable intelligent surface-assisted drone's reflection coefficient matrix. Taking into account the need for timely response to dynamic environments and the fact that the problem involves multivariable coupling and non-convex optimization, two agents are defined: a base station agent and a multifunctional reconfigurable intelligent surface-assisted drone agent. A fully cooperative multi-agent deep reinforcement learning algorithm (MADRL) based on a soft actor-critic (SAC) is proposed to find the optimal learning strategy, significantly improving security energy efficiency while also promptly responding to the needs of dynamic environments.
[0177] To further verify the accuracy of the embodiment of the present invention, the system model and algorithm performance are verified in combination with simulation experiments to further illustrate the effects of the embodiment of the present invention.
[0178] The base station (BS) is located at (250m, 250m) with a height of 10m. Users are randomly distributed in a circle with a radius of 50m and a center of (350m, 450m). Eavesdroppers are randomly distributed in the legal area. The height of the eavesdroppers and users is 2m. Channel H r 、h k and h e The path loss follows the Rician fading model, and the path loss varies with distance, which can be expressed as:
[0179]
[0180] Where ∈, d, and C0 are the path loss exponent, channel distance, and reference path loss when the reference distance D0 = 1m. β is the Rice factor, h LoS and h NLoS They represent the line-of-sight (LoS) and non-line-of-sight (NLoS) components, respectively.
[0181] First, the reward convergence performance of the fully cooperative multi-agent deep reinforcement learning algorithm based on soft actor-critic (SAC) (hereinafter referred to as mSAC algorithm) proposed in the embodiment of the present invention is compared with the currently widely used DRL algorithms (such as SAC, PPO, TD3). The models of the proposed mSAC algorithm and other DRL algorithms are trained using the same parameters, and the results obtained in 10,000 training cycles are shown. Figure 4 The following is a comparison of rewards for different algorithms, where the blurred part is the reward for each episode and the other part is the average reward. Figure 4 It can be seen that the mSAC algorithm proposed in the embodiment of the present invention is superior to other algorithms in terms of convergence speed and stability, and obtains the maximum reward.
[0182] like Figure 5 The following is a two-dimensional trajectory diagram of the UAV: Figure 6 The figure shows the three-dimensional trajectory of the UAV. Figure 5 and Figure 6 The two-dimensional and three-dimensional trajectories of the UAV at the initial positions of (50, 50, 250) meters, (50, 650, 250) meters, and (650, 650, 250) meters are shown in FIG. The UAV always descends to the lowest altitude to improve the channel gain and confidentiality rate by reducing large-scale fading. In addition, the embodiment of the present invention also compares different path loss coefficients in the UAV trajectory. Figure 5 and Figure 6 As can be seen in Figure 2, when the path loss exponent ∈1 from the base station to the F-RIS and the path loss exponent ∈2 from the F-RIS to the user are both 2.8, drones at different initial locations tend to fly toward the center between the base station and the user group. However, when ∈1 is 3.5 and ∈2 is 2.8, drones tend to fly closer to the base station. This is because the increased reward from F-RIS near the base station outweighs the reduced reward from moving away from the user. Similarly, when ∈1 is 2.8 and ∈2 is 3.5, drones tend to fly closer to the user.
[0183] like Figure 7 Shown are different transmission powers P BS The embodiment of the present invention also compares the performance of passive RIS (PRIS) on drones of different sizes in the system. It can be seen that the confidentiality energy efficiency of all solutions increases with the increase of transmission power. In addition, compared with the PRIS solution on drones, the confidentiality energy efficiency of the F-RIS solution is significantly improved. For example, when P BS =40dBm, compared with PRIS using N=16, the proposed MFRIS using N=16 can improve the confidentiality energy efficiency by 4.1%.
[0184] Reference Figure 8 , an embodiment of the present invention further provides a drone confidentiality energy efficiency optimization system based on multi-agent deep reinforcement learning, comprising:
[0185] The first module is used to build a confidentiality energy efficiency optimization model for the multifunctional reconfigurable intelligent surface-assisted UAV secure communication system based on the UAV trajectory, base station transmit beamforming, and the reflection coefficient matrix of the multifunctional reconfigurable intelligent surface-assisted UAV;
[0186] The second module is used to construct a Markov decision process model based on the confidential energy efficiency optimization model. The Markov decision process model includes the state space, the base station action space, the multifunctional reconfigurable intelligent surface assisted UAV action space and the reward function;
[0187] The third module is used to iteratively optimize the Markov decision process model based on a multi-agent deep reinforcement learning algorithm to obtain the optimal confidentiality energy efficiency.
[0188] The contents of the above-mentioned embodiments of the drone confidentiality energy efficiency optimization method based on multi-agent deep reinforcement learning are all applicable to the embodiments of the present drone confidentiality energy efficiency optimization system based on multi-agent deep reinforcement learning. The functions specifically implemented by the embodiments of the present drone confidentiality energy efficiency optimization system based on multi-agent deep reinforcement learning are the same as those of the above-mentioned embodiments of the drone confidentiality energy efficiency optimization method based on multi-agent deep reinforcement learning, and the beneficial effects achieved are also the same as the beneficial effects achieved by the above-mentioned embodiments of the drone confidentiality energy efficiency optimization method based on multi-agent deep reinforcement learning.
[0189] An embodiment of the present invention further provides an electronic device comprising: a memory, a processor, a program stored in the memory and executable on the processor, and a data bus for communication between the processor and the memory. When the program is executed by the processor, the aforementioned method for optimizing drone security and energy efficiency based on multi-agent deep reinforcement learning is implemented. The electronic device can be any intelligent terminal, including a tablet computer and an in-vehicle computer.
[0190] like Figure 9 FIG2 is a schematic diagram of the hardware structure of the electronic device provided by the embodiment of the present invention, referring to FIG2 Figure 9 , an embodiment of the present invention provides an electronic device, including:
[0191] The processor 1001 may be implemented as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is configured to execute relevant programs to implement the technical solutions provided by the embodiments of the present invention.
[0192] The memory 1002 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1002 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1002 and is called by the processor 1001 to execute the drone confidentiality energy efficiency optimization method based on multi-agent deep reinforcement learning according to the embodiment of the present invention.
[0193] Input / output interface 1003, used to implement information input and output;
[0194] Communication interface 1004, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.);
[0195] Bus 1005 , which transmits information between various components of the device (e.g., processor 1001 , memory 1002 , input / output interface 1003 , and communication interface 1004 );
[0196] The processor 1001 , the memory 1002 , the input / output interface 1003 and the communication interface 1004 are connected to each other in communication within the device via the bus 1005 .
[0197] An embodiment of the present invention also provides a storage medium, which is a computer-readable storage medium used for computer-readable storage. The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the above-mentioned drone confidentiality energy efficiency optimization method based on multi-agent deep reinforcement learning.
[0198] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0199] The embodiment of the present invention also discloses a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs Figure 1 The method shown.
[0200] In some optional embodiments, the function / operation mentioned in the block diagram may not occur in the order mentioned in the operation diagram. For example, depending on the function / operation involved, the two boxes shown in succession can actually be executed substantially simultaneously or the above-mentioned boxes can sometimes be executed in reverse order. In addition, the embodiment presented and described in the flow chart of the present invention is provided in an exemplary manner for the purpose of providing a more comprehensive understanding of the technology. The disclosed method is not limited to the operation and logic flow presented herein. Optional embodiments are contemplated in which the order of the various operations is changed and the sub-operations described as a part of a larger operation are performed independently.
[0201] In addition, although the present invention is described in the context of functional modules, it should be understood that, unless otherwise stated, one or more of the above-mentioned functions and / or features can be integrated into a single physical device and / or software module, or one or more functions and / or features can be implemented in separate physical devices or software modules. It is also understood that a detailed discussion of the actual implementation of each module is not necessary for understanding the present invention. More specifically, given the properties, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the module will be understood within the routine skills of an engineer. Therefore, a person skilled in the art can implement the present invention set forth in the claims using ordinary skills without undue experimentation. It is also understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.
[0202] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the above methods of each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program code.
[0203] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0204] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic devices), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable media on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, deciphering, or processing in another suitable manner as necessary, and then stored in a computer memory.
[0205] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0206] In the above description of this specification, reference to the terms "one embodiment / example," "another embodiment / example," or "certain embodiments / examples" means that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0207] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to the embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the claims and their equivalents.
[0208] The above is a specific description of the preferred implementation of the present invention, but the present invention is not limited to the embodiments. Those skilled in the art can make various equivalent modifications or substitutions without violating the spirit of the present invention. These equivalent modifications or substitutions are all included in the scope defined by the claims of this application.
Claims
1. A method for optimizing the confidentiality and energy efficiency of drones based on multi-agent deep reinforcement learning, characterized in that: The following steps are involved: Based on the UAV trajectory, base station transmit beamforming, and the reflection coefficient matrix of the multifunctional reconfigurable intelligent surface-assisted UAV, a confidentiality energy efficiency optimization model for the multifunctional reconfigurable intelligent surface-assisted UAV secure communication system is constructed; Constructing a Markov decision process model based on the confidentiality energy efficiency optimization model, wherein the Markov decision process model includes a state space, a base station action space, a multifunctional reconfigurable intelligent surface assisted UAV action space, and a reward function; Based on a multi-agent deep reinforcement learning algorithm, the Markov decision process model is iteratively optimized to obtain the optimal confidentiality energy efficiency.
2. The method for optimizing the confidentiality and energy efficiency of drones based on multi-agent deep reinforcement learning according to claim 1 is characterized in that: The method constructs a confidentiality energy efficiency optimization model for a multifunctional reconfigurable intelligent surface-assisted UAV secure communication system based on the UAV trajectory, base station transmit beamforming, and the reflection coefficient matrix of the multifunctional reconfigurable intelligent surface-assisted UAV, specifically including: The multifunctional reconfigurable intelligent surface-assisted UAV secure communication system is defined to include a base station, a UAV, an eavesdropper, and several legitimate users, wherein the UAV is the multifunctional reconfigurable intelligent surface-assisted UAV; defining the trajectory of the drone and the flight speed of the drone in each time slot, and determining the propulsion energy of the drone in each time slot based on the flight speed; defining a transmit beamforming vector of the legal user in each time slot; Defining the reflection coefficient matrix according to the phase shift and amplification coefficient of the smart surface of the reflection unit; Defining a channel gain matrix between the base station and the drone and a reflection channel gain vector between the drone and the legitimate user; determining, based on the amplification factor, the transmit beamforming vector, and the channel gain matrix, that an element acquires energy; Determining an achievable rate for the legitimate user based on the reflection channel gain vector, the reflection coefficient matrix, the channel gain matrix, and the transmit beamforming vector; Determining an eavesdropping rate of the legitimate user by the eavesdropper according to the achievable rate; determining confidentiality energy efficiency according to the propulsion energy, the achievable rate, and the eavesdropping rate; The confidentiality energy efficiency, the UAV trajectory, the transmit beamforming vector, the reflection coefficient matrix, the component acquisition energy, and the channel gain matrix are integrated to obtain the confidentiality energy efficiency optimization model.
3. The method for optimizing the confidentiality and energy efficiency of drones based on multi-agent deep reinforcement learning according to claim 1 is characterized in that: The confidentiality energy efficiency optimization model is: Among them, SEE n represents the confidentiality efficiency, L and H represent the trajectory of the drone and represents the position of the UAV in time slot n and n represents the nth time slot and W represents the base station transmit beamforming and represents the transmit beamforming vector of the kth legal user in time slot n, k represents the kth legal user and k∈K, P BS represents the base station transmit beamforming, Θ represents the reflection coefficient matrix and and Indicates that the component obtains energy, represents the amplification factor, represents the i-th row in the channel gain matrix, P RIS represents the maximum transmission power, ξ represents the energy consumption conversion efficiency and ξ∈(0,1].
4. The method for optimizing the confidentiality and energy efficiency of drones based on multi-agent deep reinforcement learning according to claim 2 is characterized in that: The Markov decision process model is constructed according to the confidentiality energy efficiency optimization model, specifically including: Determining the state space according to the UAV trajectory and the channel gain matrix; determining the base station action space according to the transmit beamforming vector; Determining an action space of the multifunctional reconfigurable intelligent surface-assisted UAV according to the flight distance, elevation angle, azimuth angle of the UAV and the reflection coefficient matrix; defining a penalty for the drone flying out of a legal area, and determining the reward function according to the confidentiality energy efficiency and the penalty for flying out of the legal area; The state space, the base station action space, the multifunctional reconfigurable intelligent surface assisted UAV action space and the reward function are integrated to obtain the Markov decision process model.
5. The method for optimizing the confidentiality and energy efficiency of drones based on multi-agent deep reinforcement learning according to claim 1, characterized in that: The UAV confidentiality energy efficiency optimization method further includes: The multi-agent deep reinforcement learning algorithm is defined to include a base station agent, a multifunctional reconfigurable intelligent surface assisted UAV agent, and a MIX network, wherein the base station agent includes a first actor network and a first critic network, and the multifunctional reconfigurable intelligent surface assisted UAV agent includes a second actor network and a second critic network; The first critic network and the second critic network are combined through the MIX network to obtain a QMIX network.
6. The method for optimizing the confidentiality and energy efficiency of drones based on multi-agent deep reinforcement learning according to claim 5 is characterized in that: The multi-agent deep reinforcement learning algorithm is used to iteratively optimize the Markov decision process model to obtain the optimal confidentiality energy efficiency, specifically including: Adjusting the Markov decision process model to obtain a Markov decision process optimization model; outputting a policy distribution through the first actor network and the second actor network based on the Markov decision process optimization model; Outputting a first local Q value according to the policy distribution through the first critic network; Outputting a second local Q value according to the policy distribution through the second critic network; Combining the first local Q value and the second local Q value through the QMIX network to obtain a joint Q value; updating the first actor network, the first critic network, the second actor network, and the second critic network according to the joint Q value; When the preset policy update condition is reached, the update is stopped and the optimal confidentiality energy efficiency is output.
7. The method for optimizing the confidentiality and energy efficiency of drones based on multi-agent deep reinforcement learning according to claim 6, characterized in that: The Markov decision process model is adjusted to obtain a Markov decision process optimization model, specifically comprising: Normalizing the transmit beamforming vector in the base station action space according to the confidentiality energy efficiency optimization model; Adjusting the reflection coefficient matrix in the multifunctional reconfigurable intelligent surface-assisted UAV action space according to the confidential energy efficiency optimization model; The Markov decision process optimization model is obtained according to the normalized base station action space and the adjusted multifunctional reconfigurable intelligent surface assisted UAV action space.
8. A drone confidentiality energy efficiency optimization system based on multi-agent deep reinforcement learning, characterized by: include: The first module is used to build a confidentiality energy efficiency optimization model for the multifunctional reconfigurable intelligent surface-assisted UAV secure communication system based on the UAV trajectory, base station transmit beamforming, and the reflection coefficient matrix of the multifunctional reconfigurable intelligent surface-assisted UAV; The second module is used to construct a Markov decision process model based on the confidentiality energy efficiency optimization model, wherein the Markov decision process model includes a state space, a base station action space, a multifunctional reconfigurable intelligent surface assisted UAV action space, and a reward function; The third module is used to iteratively optimize the Markov decision process model based on a multi-agent deep reinforcement learning algorithm to obtain the optimal confidentiality energy efficiency.
9. An electronic device, characterized in that: The electronic device includes a memory, a processor, a program stored in the memory and executable on the processor, and a data bus for realizing connection and communication between the processor and the memory. When the program is executed by the processor, the steps of the method for optimizing the confidentiality and energy efficiency of a drone based on multi-agent deep reinforcement learning are implemented as described in any one of claims 1 to 7.
10. A storage medium, which is a computer-readable storage medium and is used for computer-readable storage, characterized in that: The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the drone confidentiality energy efficiency optimization method based on multi-agent deep reinforcement learning as described in any one of claims 1 to 7.