Communication method and system based on multi-agent near-end strategy optimization algorithm
By using a multi-agent near-end strategy optimization algorithm, real-time data on the status of aerial intelligent metasurfaces and UAVs are collected, and spatiotemporal dual-domain resource allocation is performed. This solves the problem of coupling between beamforming parameters and UAV flight trajectories, achieving long-term, efficient, and secure communication.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGXI POWER GRID CO LTD NANNING POWER SUPPLY BUREAU
- Filing Date
- 2025-12-02
- Publication Date
- 2026-04-21
AI Technical Summary
In existing technologies, the beamforming parameters of aerial intelligent metasurfaces are highly coupled with the flight trajectory of drones in multi-user communication scenarios, resulting in limited endurance. Furthermore, traditional encryption technologies are easily cracked by quantum computing, making it difficult to achieve long-term and efficient secure communication.
A multi-agent near-end policy optimization algorithm is adopted to collect real-time state data of aerial intelligent metasurfaces and UAVs. Through centralized value assessment and decentralized policy execution, action parameters of UAVs and intelligent metasurfaces are generated to perform spatiotemporal dual-domain resource allocation and optimize communication tasks.
It achieves long-term, efficient, secure and reliable communication in complex communication environments, enhances the system's multi-objective collaborative optimization capabilities, avoids policy oscillations, and dynamically adjusts the weight allocation of communication throughput, energy efficiency and security performance.
Smart Images

Figure CN121908296A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of wireless communication technology, and in particular to a communication method and system based on a multi-agent near-end policy optimization algorithm. Background Technology
[0002] With the evolution of mobile communication technology and the demand for the Internet of Things, the openness of wireless channels poses a serious threat of eavesdropping in multi-user communication scenarios. Traditional encryption technologies are at risk of being cracked by quantum computing, and key management in high-level protocols faces increasing complexity. To overcome this problem, the Airborne Intelligent Metasurface (ARIS) has become a key 6G technology due to its ability to overcome ground obstacles through aerial deployment, flexibly constructing line-of-sight communication links to extend network coverage and ensure communication security. However, the beamforming parameters of the ARIS are highly coupled with the flight trajectory of its onboard platform (UAV), and the limited onboard energy of the UAV severely restricts its endurance. This necessitates multi-variable collaborative optimization of its trajectory, beamforming, and energy resources to balance system security and energy efficiency. To address these challenges, deep reinforcement learning, with its advantages of not requiring precise mathematical models and being able to dynamically adapt to the time-varying characteristics of channels, has become a core approach for solving such problems.
[0003] Currently, existing technologies such as early deep Q-networks, deep deterministic policy gradient algorithms, double-delay deep deterministic policy gradient algorithms, and near-end policy optimization algorithms have all been adopted.
[0004] However, since most of the above algorithms still require a pre-set weight to transform a multi-objective problem into a single-objective problem, the multi-objective collaborative optimization capability is insufficient. Consequently, the generated flight strategies and signal reflection strategies are difficult to achieve the long-term and efficient secure communication requirements. Summary of the Invention
[0005] In view of this, this application provides a communication method and system based on a multi-agent proximal policy optimization algorithm, the main purpose of which is to address the problem that existing methods cannot truly achieve the long-term and efficient secure communication requirements.
[0006] According to one aspect of this application, a communication method based on a multi-agent proximal policy optimization algorithm is provided, comprising: During the communication mission performed by the target aerial intelligent metasurface, the current time slot state data of the target aerial intelligent metasurface is collected in real time. The aerial intelligent metasurface is used to characterize the product of the combination of intelligent metasurface and UAV. The current time slot state data includes the current time slot UAV state data and the current time slot intelligent metasurface state data. Based on the UAV action policy generation network that has completed model training, the current time slot UAV action parameters are generated according to the current time slot UAV state data. Based on the intelligent metasurface action policy generation network that has completed model training, the current time slot intelligent metasurface action parameters are generated according to the current time slot intelligent metasurface state data. The UAV action policy generation network and the intelligent metasurface action policy generation network are obtained by joint training based on the multi-agent proximal policy optimization algorithm through centralized value evaluation and decentralized policy execution. The UAV is controlled to perform flight missions based on the current time slot UAV action parameters, and the smart metasurface is controlled to perform communication missions based on spatiotemporal dual-domain resource allocation based on the current time slot smart metasurface action parameters.
[0007] Preferably, the current time-slot smart metasurface action parameters include the current time-slot smart metasurface phase matrix, the current time-slot time allocation factor, and the current time-slot cell reflection control parameters. Controlling the smart metasurface to perform a communication task based on spatiotemporal dual-domain resource allocation according to the current time-slot smart metasurface action parameters includes: Based on the current time slot allocation factor, the current time slot is divided into stages to obtain the energy harvesting stage and the information transmission stage. During the energy harvesting phase, all metasurface units contained in the intelligent metasurface are controlled to perform energy harvesting tasks; During the information transmission phase, based on the current time slot unit reflection control parameters, all metasurface units are divided into multiple reflection units and multiple energy harvesting units. The multiple energy harvesting units are controlled to perform energy harvesting tasks, and the multiple reflection units are controlled to configure corresponding phase offset parameters based on the current time slot intelligent metasurface phase matrix to perform communication tasks, so as to carry out communication operations based on spatiotemporal dual-domain resource allocation.
[0008] Preferably, before the UAV action policy generation network based on the completed model training generates the current time slot UAV action parameters according to the current time slot UAV state data, and before the intelligent metasurface action policy generation network based on the completed model training generates the current time slot intelligent metasurface action parameters according to the current time slot intelligent metasurface state data, the method further includes: Construct a drone motion strategy generation network, an intelligent metasurface motion strategy generation network, and a centralized motion value estimation network; In the simulation environment, the current simulation time slot UAV state data is acquired. Based on the current simulation time slot UAV state data and the UAV action strategy generation network, UAV action parameters for the current simulation time slot are generated. Based on the current simulation time slot UAV action parameters, simulated UAV flight is performed to obtain the next simulation time slot UAV state data, and the propulsion energy consumption of the current simulation time slot UAV is calculated. Additionally, the current simulation time slot intelligent metasurface state data is acquired. Based on the current simulation time slot intelligent metasurface state data and the intelligent metasurface action strategy generation network, intelligent metasurface action parameters for the current simulation time slot are generated. Based on the current simulation time slot intelligent metasurface action parameters, simulated spatiotemporal dual-domain resource allocation communication is performed to obtain the next simulation time slot intelligent metasurface state data, and the effective throughput of the current simulation time slot communication system, the intelligent metasurface communication energy consumption during the current simulation time slot information transmission phase, and the current simulation time slot's... The system collects energy using an intelligent metasurface, and calculates the global reward value for the current simulation time slot based on the current simulation time slot UAV propulsion energy consumption, the current simulation time slot information transmission stage intelligent metasurface communication energy consumption, the current simulation time slot intelligent metasurface collected energy, and the current simulation time slot communication system effective throughput. It also combines the current simulation time slot UAV state data, the current simulation time slot intelligent metasurface state data, the current simulation time slot UAV action parameters, the current simulation time slot intelligent metasurface action parameters, the current simulation time slot global reward value, the next simulation time slot UAV state data, and the next simulation time slot intelligent metasurface state data to obtain the current simulation time slot communication system vector. Finally, it constructs the next simulation time slot communication system vector based on the next simulation time slot UAV state data and the next simulation time slot intelligent metasurface state data, and iteratively generates multiple communication system vectors. Based on the centralized action value estimation network, the estimated action value for the next simulation time slot is generated according to the UAV state data and the smart metasurface state data of the next simulation time slot. Based on the Bellman equation, the expected action value for the current simulation time slot is calculated according to the estimated action value for the next simulation time slot and the global reward value of the current simulation time slot. Furthermore, based on the centralized action value estimation network, the estimated action value for the current simulation time slot is generated according to the UAV state data and the smart metasurface state data of the current simulation time slot. Finally, based on the mean squared error loss between the estimated action value and the expected action value of the current simulation time slot, the parameters of the centralized action value estimation network are updated using a gradient descent algorithm. Based on the generalized dominance estimation algorithm, the dominance value of the current simulation time slot is calculated according to the estimated value of the current simulation time slot action, the estimated value of the next simulation time slot action, and the global reward value of the current simulation time slot. Based on the near-end policy optimization algorithm, a UAV action policy gradient loss function is constructed according to the current simulation time slot advantage value and the current simulation time slot UAV action parameters. Based on the UAV action policy gradient loss function, the parameters of the UAV action policy generation network are updated through the gradient ascent algorithm. Based on the near-end policy optimization algorithm, a gradient loss function for intelligent metasurface action policy is constructed according to the current simulation time slot advantage value and the intelligent metasurface action parameters of the current simulation time slot. Based on the gradient loss function for intelligent metasurface action policy, the parameters of the intelligent metasurface action policy generation network are updated through the gradient ascent algorithm. The centralized action value estimation network, the UAV action policy generation network, and the intelligent metasurface action policy generation network are alternately updated until the network converges, resulting in a centralized action value estimation network, a UAV action policy generation network, and an intelligent metasurface action policy generation network that have completed model training.
[0009] Preferably, before constructing the UAV motion strategy generation network, the intelligent metasurface motion strategy generation network, and the centralized motion value estimation network, the method further includes: The communication channel model is constructed and expressed by the following formula. , , in, Represents the first in base stations and smart metasurfaces The LoS probability of each metasurface unit, where A and B both represent communication environment constants. Indicates the base station and the first The elevation angle between each metasurface unit Indicates the first At the height of a metasurface unit in time slot t, ( , ) indicates the first The horizontal coordinates of a metasurface unit in time slot t; A reflective model of an aerial intelligent metasurface is constructed, expressed by the following formula. , in, This represents the diagonal matrix of reflection coefficients. Indicates the first The amplitude reflection coefficient of each metasurface unit. Indicates the first The complex form of the phase shift of each metasurface unit. Indicates the number of metasurface units; The base station transmit signal model is constructed and expressed by the following formula. , Where X represents the signal vector actually transmitted by the base station. Represents the set of legitimate users. This indicates that the eavesdroppers have gathered. Indicates user The precoded vector, Indicates user A circularly symmetric complex Gaussian signal with zero mean and unit variance; A model for the propulsion energy consumption of unmanned aerial vehicles (UAVs) is constructed, expressed by the following formula. , in, This represents the propulsion energy consumption of the UAV in time slot t. Indicates the duration of the time slot. Indicates the hovering blade power. This represents the speed of the drone in time slot t. Indicates the rotor tip speed. Indicates the fuselage drag ratio. Indicates air density, Indicates rotor solidity, Indicates the rotor disk area, Indicates hovering induced power. This indicates the average induced speed during hovering; An aerial intelligent metasurface energy harvesting model is constructed, expressed by the following formula. , in, This represents the energy collected by the airborne intelligent metasurface in time slot t. Indicates the duration of the energy harvesting phase. Indicates the number of rows of metasurface cells. Indicates the number of columns for the metasurface unit. Indicates energy harvesting efficiency. Indicates the base station and the first The conjugate transpose of the channel vectors of each metasurface unit This represents the total set of users and eavesdroppers. Represents binary control variables; A communication energy consumption model for the information transmission phase of an aerial intelligent metasurface is constructed, expressed by the following formula: , in, G represents the communication energy consumption of the airborne intelligent metasurface during the information transmission phase in time slot t, and G represents the base station precoding matrix. Based on the communication channel model, the airborne intelligent metasurface reflection model, the base station transmission signal model, the UAV propulsion energy consumption model, the airborne intelligent metasurface energy harvesting model, and the airborne intelligent metasurface communication energy consumption model during the information transmission phase, a simulation environment is constructed to build a communication system vector within the simulation environment.
[0010] Preferably, the formula for calculating the effective throughput of a communication system is expressed as follows: , , in, Indicates the effective throughput of the communication system. Indicates user Effective throughput in time slot t Indicates user In the time slot signal gain, Indicates user In the time slot Interference and noise.
[0011] Preferably, the global reward value calculation formula is expressed as follows: , in, This represents the global reward value for time slot t.
[0012] The preferred mean squared error loss function is expressed by the following formula: , in, E represents the mean squared error loss value, and E represents the expected value. Indicates the value of the current simulated time slot action estimation. Compared with the expected value of current simulated time slot actions The square of the difference; The gradient loss function for the action policy is expressed by the following formula.
[0013] in, This represents the gradient loss value of the drone's action policy or the gradient loss value of the intelligent metasurface's action policy. Indicates the strategy ratio, This indicates the current simulation slot dominance value. This indicates the PPO shearing threshold, and clip indicates the shearing mechanism.
[0014] According to another aspect of this application, a communication system based on a multi-agent proximal policy optimization algorithm is provided, comprising: The status data acquisition module is used to collect the current time slot status data of the target airborne intelligent metasurface in real time during the communication mission of the target airborne intelligent metasurface. The airborne intelligent metasurface is used to characterize the product of the combination of intelligent metasurface and UAV. The current time slot status data includes the current time slot UAV status data and the current time slot intelligent metasurface status data. The motion parameter generation module is used to generate UAV motion parameters for the current time slot based on the UAV state data of the current time slot, and to generate intelligent metasurface motion parameters for the current time slot based on the intelligent metasurface state data of the current time slot, using a UAV motion policy generation network that has completed model training. The UAV motion policy generation network and the intelligent metasurface motion policy generation network are obtained by joint training based on a multi-agent proximal policy optimization algorithm through centralized value evaluation and decentralized policy execution. The control module is used to control the UAV to perform flight missions according to the current time slot UAV action parameters, and to control the smart metasurface to perform communication missions based on spatiotemporal dual-domain resource allocation according to the current time slot smart metasurface action parameters.
[0015] Preferably, the current time-slot smart metasurface action parameters include the current time-slot smart metasurface phase matrix, the current time-slot time allocation factor, and the current time-slot cell reflection control parameters. The control module is used for: Based on the current time slot allocation factor, the current time slot is divided into stages to obtain the energy harvesting stage and the information transmission stage. During the energy harvesting phase, all metasurface units contained in the intelligent metasurface are controlled to perform energy harvesting tasks; During the information transmission phase, based on the current time slot unit reflection control parameters, all metasurface units are divided into multiple reflection units and multiple energy harvesting units. The multiple energy harvesting units are controlled to perform energy harvesting tasks, and the multiple reflection units are controlled to configure corresponding phase offset parameters based on the current time slot intelligent metasurface phase matrix to perform communication tasks, so as to carry out communication operations based on spatiotemporal dual-domain resource allocation.
[0016] Preferably, before the action parameter generation module, the system further includes a model training module, including: Network building units are used to construct UAV motion strategy generation networks, intelligent metasurface motion strategy generation networks, and centralized motion value estimation networks. The communication system vector generation unit is used to acquire the current simulation time slot UAV state data in a simulation environment, generate UAV action parameters for the current simulation time slot based on the UAV action strategy generation network, simulate UAV flight based on the current simulation time slot action parameters, obtain the next simulation time slot UAV state data, calculate the UAV propulsion energy consumption for the current simulation time slot, and acquire the current simulation time slot intelligent metasurface state data. Based on the current simulation time slot intelligent metasurface state data, generate intelligent metasurface action parameters for the current simulation time slot based on the intelligent metasurface action strategy generation network, simulate spatiotemporal dual-domain resource allocation communication based on the current simulation time slot intelligent metasurface action parameters, obtain the next simulation time slot intelligent metasurface state data, and calculate the effective throughput of the current simulation time slot communication system and the intelligent metasurface communication energy consumption during the current simulation time slot information transmission phase. The system collects energy from the intelligent metasurface in the current simulation time slot, and calculates the global reward value for the current simulation time slot based on the UAV propulsion energy consumption, the intelligent metasurface communication energy consumption during the information transmission phase of the current simulation time slot, the collected energy from the intelligent metasurface in the current simulation time slot, and the effective throughput of the communication system in the current simulation time slot. It then combines the UAV state data, the intelligent metasurface state data, the UAV action parameters, the intelligent metasurface action parameters, the global reward value, the UAV state data, and the intelligent metasurface state data of the next simulation time slot to obtain the current simulation time slot communication system vector. Finally, it constructs the next simulation time slot communication system vector based on the UAV state data and the intelligent metasurface state data of the next simulation time slot, and iteratively generates multiple communication system vectors. A centralized action value estimation network training unit is used to generate the action estimation value for the next simulation time slot based on the UAV state data and smart metasurface state data in the next simulation time slot, and to calculate the expected action value for the current simulation time slot based on the Bellman equation, the action estimation value for the next simulation time slot, and the global reward value for the current simulation time slot. It also generates the action estimation value for the current simulation time slot based on the UAV state data and smart metasurface state data in the current simulation time slot, and updates the parameters of the centralized action value estimation network using a gradient descent algorithm based on the mean squared error loss between the action estimation value and the expected action value for the current simulation time slot. The dominance value calculation unit is used to calculate the dominance value of the current simulation time slot based on the generalized dominance estimation algorithm, according to the estimated value of the current simulation time slot action, the estimated value of the next simulation time slot action, and the global reward value of the current simulation time slot. The UAV action policy generation network training unit is used to construct a UAV action policy gradient loss function based on the near-end policy optimization algorithm, according to the current simulation time slot advantage value and the UAV action parameters of the current simulation time slot, and update the parameters of the UAV action policy generation network based on the UAV action policy gradient loss function through the gradient ascent algorithm. The training unit of the intelligent metasurface action policy generation network is used to construct the gradient loss function of the intelligent metasurface action policy based on the proximal policy optimization algorithm, according to the current simulation time slot advantage value and the intelligent metasurface action parameters of the current simulation time slot, and update the parameters of the intelligent metasurface action policy generation network based on the gradient ascent algorithm. The training termination unit is used to alternately update the centralized action value estimation network, the UAV action policy generation network, and the intelligent metasurface action policy generation network until the network converges, resulting in the centralized action value estimation network, the UAV action policy generation network, and the intelligent metasurface action policy generation network that have completed model training.
[0017] Preferably, before the network construction unit, the model training module further includes a simulation environment construction unit, used for: The communication channel model is constructed and expressed by the following formula. , , in, Represents the first in base stations and smart metasurfaces The LoS probability of each metasurface unit, where A and B both represent communication environment constants. Indicates the base station and the first The elevation angle between each metasurface unit Indicates the first At the height of a metasurface unit in time slot t, ( , ) indicates the first The horizontal coordinates of a metasurface unit in time slot t; A reflective model of an aerial intelligent metasurface is constructed, expressed by the following formula. , in, This represents the diagonal matrix of reflection coefficients. Indicates the first The amplitude reflection coefficient of each metasurface unit. Indicates the first The complex form of the phase shift of each metasurface unit. Indicates the number of metasurface units; The base station transmit signal model is constructed and expressed by the following formula. , Where X represents the signal vector actually transmitted by the base station. Represents the set of legitimate users. This indicates that the eavesdroppers have gathered. Indicates user The precoded vector, Indicates user A circularly symmetric complex Gaussian signal with zero mean and unit variance; A model for the propulsion energy consumption of unmanned aerial vehicles (UAVs) is constructed, expressed by the following formula. , in, This represents the propulsion energy consumption of the UAV in time slot t. Indicates the duration of the time slot. Indicates the hovering blade power. This represents the speed of the drone in time slot t. Indicates the rotor tip speed. Indicates the fuselage drag ratio. Indicates air density, Indicates rotor solidity, Indicates the rotor disk area, Indicates hovering induced power. This indicates the average induced speed during hovering; An aerial intelligent metasurface energy harvesting model is constructed, expressed by the following formula. , in, This represents the energy collected by the airborne intelligent metasurface in time slot t. Indicates the duration of the energy harvesting phase. Indicates the number of rows of metasurface cells. Indicates the number of columns for the metasurface unit. Indicates energy harvesting efficiency. Indicates the base station and the first The conjugate transpose of the channel vectors of each metasurface unit This represents the total set of users and eavesdroppers. Represents binary control variables; A communication energy consumption model for the information transmission phase of an aerial intelligent metasurface is constructed, expressed by the following formula: , in, G represents the communication energy consumption of the airborne intelligent metasurface during the information transmission phase in time slot t, and G represents the base station precoding matrix. Based on the communication channel model, the airborne intelligent metasurface reflection model, the base station transmission signal model, the UAV propulsion energy consumption model, the airborne intelligent metasurface energy harvesting model, and the airborne intelligent metasurface communication energy consumption model during the information transmission phase, a simulation environment is constructed to build a communication system vector within the simulation environment.
[0018] Preferably, the formula for calculating the effective throughput of a communication system is expressed as follows: , , in, Indicates the effective throughput of the communication system. Indicates user Effective throughput in time slot t Indicates user In the time slot signal gain, Indicates user In the time slot Interference and noise.
[0019] Preferably, the global reward value calculation formula is expressed as follows: , in, This represents the global reward value for time slot t.
[0020] The preferred mean squared error loss function is expressed by the following formula: , in, E represents the mean squared error loss value, and E represents the expected value. Indicates the value of the current simulated time slot action estimation. Compared with the expected value of current simulated time slot actions The square of the difference; The gradient loss function for the action policy is expressed by the following formula.
[0021] in, This represents the gradient loss value of the drone's action policy or the gradient loss value of the intelligent metasurface's action policy. Indicates the strategy ratio, This indicates the current simulation slot dominance value. This indicates the PPO shearing threshold, and clip indicates the shearing mechanism.
[0022] According to another aspect of this application, a storage medium is provided, wherein at least one executable instruction is stored therein, the executable instruction causing a processor to perform an operation corresponding to the communication method based on the multi-agent proximal policy optimization algorithm described above.
[0023] According to another aspect of this application, a terminal is provided, comprising: a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other through the communication bus; The memory is used to store at least one executable instruction, which causes the processor to perform the operation corresponding to the communication method based on the multi-agent proximal policy optimization algorithm.
[0024] By employing the above technical solutions, the technical solutions provided in the embodiments of this application have at least the following advantages: This application provides a communication method and system based on a multi-agent near-end policy optimization algorithm. First, during the execution of a communication task by a target aerial intelligent metasurface, the current time-slot state data of the target aerial intelligent metasurface is collected in real time. The aerial intelligent metasurface represents the product of the combination of the intelligent metasurface and the UAV. The current time-slot state data includes the current time-slot UAV state data and the current time-slot intelligent metasurface state data. Second, based on a UAV action policy generation network that has completed model training, action parameters for the current time-slot UAV are generated according to the current time-slot UAV state data. Similarly, based on a trained intelligent metasurface action policy generation network, action parameters for the current time-slot intelligent metasurface are generated according to the current time-slot intelligent metasurface state data. Both the UAV action policy generation network and the intelligent metasurface action policy generation network are obtained through joint training using a multi-agent near-end policy optimization algorithm, through centralized value evaluation and decentralized policy execution. Finally, the UAV is controlled to execute a flight task based on the current time-slot UAV action parameters, and the intelligent metasurface is controlled to execute a communication task based on spatiotemporal dual-domain resource allocation based on the current time-slot intelligent metasurface action parameters. Compared with existing technologies, the centralized value assessment mechanism adopted in this application can perform unified value estimation of the joint actions of UAV and intelligent metasurface based on global system state information. This adaptively adjusts the weight distribution among communication throughput, energy efficiency and safety performance during multi-objective optimization, significantly enhancing the system's multi-objective collaborative optimization capability. Simultaneously, it provides globally consistent and dynamically adaptable optimization guidance for the training of the UAV action strategy generation network and the intelligent metasurface action strategy generation network, effectively avoiding policy oscillations caused by objective conflicts. Furthermore, through a decentralized policy execution mechanism, each agent, guided by the global value signal, performs policy optimization for local tasks such as UAV trajectory control and metasurface beamforming. This achieves optimal performance in its respective control dimension and realizes dynamic complementarity and joint improvement of the overall system efficiency through global collaboration, thereby achieving long-term, efficient, and secure communication in complex communication environments.
[0025] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description
[0026] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings: Figure 1 A flowchart of a communication method based on a multi-agent proximal policy optimization algorithm provided in an embodiment of this application is shown. Figure 2 This paper shows a schematic diagram of a system model of an aerial intelligent metasurface performing a communication task according to an embodiment of this application. Figure 3 A flowchart of a communication task execution method based on spatiotemporal dual-domain resource allocation provided in an embodiment of this application is shown; Figure 4 This illustration shows a schematic diagram of intelligent metasurface control based on spatiotemporal dual-domain resource allocation provided in an embodiment of this application; Figure 5 A flowchart illustrating the model training process provided in an embodiment of this application is shown; Figure 6 This paper illustrates a block diagram of a communication system based on a multi-agent proximal policy optimization algorithm provided in an embodiment of this application. Figure 7 A schematic diagram of the structure of a terminal provided in an embodiment of this application is shown. Detailed Implementation
[0027] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0028] At the same time, it should be understood that, for ease of description, the dimensions of the various parts shown in the accompanying drawings are not drawn according to actual scale.
[0029] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit the scope of this application and its application or use.
[0030] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and equipment should be considered part of the specification.
[0031] It should be noted that similar labels and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be discussed further in subsequent figures.
[0032] The embodiments of this application can be applied to computer systems / servers that can operate with a wide range of other general-purpose or special-purpose computing system environments or configurations. Examples of well-known computing systems, environments, and / or configurations suitable for use with computer systems / servers include, but are not limited to: personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and distributed cloud computing environments that include any of the above systems, etc.
[0033] Computer systems / servers can be described in the general context of computer system executable instructions (such as program modules) executed by the computer system. Typically, program modules can include routines, programs, object programs, components, logic, data structures, etc., which perform specific tasks or implement specific abstract data types. Computer systems / servers can be implemented in distributed cloud computing environments, where tasks are performed by remote processing devices linked through a communication network. In distributed cloud computing environments, program modules can reside on local or remote computing system storage media, including storage devices.
[0034] This application provides a communication method based on a multi-agent proximal policy optimization algorithm, such as... Figure 1 As shown, the method includes: 101. During the communication mission performed on the target airborne intelligent metasurface, collect the current time slot status data of the target airborne intelligent metasurface in real time.
[0035] The aerial smart metasurface is used to characterize the product of combining a smart metasurface with a drone, i.e., a drone equipped with a smart metasurface for performing communication tasks. The current time slot state data includes the current time slot drone state data and the current time slot smart metasurface state data. The current time slot drone state data includes the current time slot aerial smart metasurface location information, the current time slot set of legitimate user location information, and the current time slot set of eavesdropper location information. The current time slot smart metasurface state data includes the current time slot base station precoding matrix and the current time slot smart metasurface receiver channel. In this embodiment, the current execution end can be the aerial smart metasurface control module.
[0036] It should be noted that the schematic diagram of the system model for the aerial intelligent metasurface performing communication tasks is as follows: Figure 2As shown, the target includes an aerial smart metasurface, a base station, several legitimate users, and several eavesdroppers.
[0037] 102. Based on the UAV action policy generation network that has completed model training, generate the UAV action parameters for the current time slot according to the UAV state data in the current time slot. Also, based on the intelligent metasurface action policy generation network that has completed model training, generate the intelligent metasurface action parameters for the current time slot according to the intelligent metasurface state data in the current time slot.
[0038] Among them, the UAV action policy generation network and the intelligent metasurface action policy generation network are obtained by joint training based on the multi-agent proximal policy optimization algorithm through centralized value evaluation and decentralized policy execution; the UAV action policy generation network is used to generate UAV action policies, which are used to control UAVs to perform flight missions; the intelligent metasurface action policy generation network is used to generate intelligent metasurface action policies, which are used to control intelligent metasurfaces to perform communication missions.
[0039] 103. Control the UAV to perform flight missions based on the current time slot UAV motion parameters, and control the intelligent metasurface to perform communication missions based on spatiotemporal dual-domain resource allocation based on the current time slot intelligent metasurface motion parameters.
[0040] In this embodiment, the current time slot UAV action parameters can be directly used to control UAV flight. To improve the mission execution duration of the aerial intelligent metasurface and optimize energy efficiency, this embodiment designs a communication mission execution model based on spatiotemporal dual-domain resource allocation. Specifically, each time slot is divided into an energy harvesting phase and an information transmission phase. In the energy harvesting phase, all metasurface units contained in the intelligent metasurface perform energy harvesting. In the information transmission phase, some metasurface units are used to reflect signals to perform communication tasks, while the remaining metasurface units continue to harvest energy.
[0041] Compared with existing technologies, the centralized value assessment mechanism adopted in this application can perform unified value estimation of the joint actions of UAV and intelligent metasurface based on global system state information. This adaptively adjusts the weight distribution among communication throughput, energy efficiency and safety performance during multi-objective optimization, significantly enhancing the system's multi-objective collaborative optimization capability. Simultaneously, it provides globally consistent and dynamically adaptable optimization guidance for the training of the UAV action strategy generation network and the intelligent metasurface action strategy generation network, effectively avoiding policy oscillations caused by objective conflicts. Furthermore, through a decentralized policy execution mechanism, each agent, guided by the global value signal, performs policy optimization for local tasks such as UAV trajectory control and metasurface beamforming. This achieves optimal performance in its respective control dimension and realizes dynamic complementarity and joint improvement of the overall system efficiency through global collaboration, thereby achieving long-term, efficient, and secure communication in complex communication environments.
[0042] In one embodiment of this application, for further definition and explanation, such as Figure 3 As shown, step 103 of the embodiment involves controlling the intelligent metasurface to perform a communication task based on spatiotemporal dual-domain resource allocation according to the current time slot intelligent metasurface action parameters, including: 201. Based on the current time slot allocation factor, the current time slot is divided into stages to obtain the energy harvesting stage and the information transmission stage.
[0043] It should be noted that the current time-slot smart metasurface action parameters include the current time-slot smart metasurface phase matrix, the current time-slot time allocation factor, and the current time-slot cell reflection control parameters. Among them, the current time-slot time allocation factor is used to characterize the duration of the energy harvesting phase and the duration of the information transmission phase, and its proportion in the current time slot, and can take values (0,1); the current time-slot cell reflection control parameters are used to determine which metasurface cells are used to perform communication tasks and which metasurface cells are used to continue performing energy harvesting tasks during the information transmission phase; the current time-slot smart metasurface phase matrix is used to indicate how each metasurface cell used to perform communication tasks should act.
[0044] 202. During the energy harvesting phase, control all metasurface units contained in the intelligent metasurface to perform energy harvesting tasks.
[0045] 203. During the information transmission phase, based on the current time slot unit reflection control parameters, all metasurface units are divided into multiple reflection units and multiple energy harvesting units. Multiple energy harvesting units are controlled to perform energy harvesting tasks, and multiple reflection units are controlled to configure corresponding phase offset parameters based on the current time slot intelligent metasurface phase matrix to perform communication tasks, so as to carry out communication operations based on spatiotemporal dual-domain resource allocation.
[0046] In the embodiments of this application, the communication task execution method based on spatiotemporal dual-domain resource allocation, such as... Figure 4 As shown in the figure, The current time slot time allocation factor represents the duration of the energy harvesting phase, where T represents the duration of the current time slot. Each square in the diagram represents a metasurface unit. Specifically, this is first determined based on the current time slot time allocation factor. The current time slot T is divided into an energy harvesting phase and an information transmission phase, with the energy harvesting phase lasting for [duration missing]. The duration of the information transmission phase is Furthermore, as shown in the left half of the diagram, during the energy harvesting phase, all metasurface units are controlled to perform energy harvesting tasks; as shown in the right half of the diagram, during the information transmission phase, the reflection control parameters of the current time slot unit are used. To determine which metasurface units are used for communication tasks and which are used to continue energy harvesting tasks, it is understandable that... For binary control variables, when When, it indicates that the metasurface unit is used to perform communication tasks, when At that time, it was explained that the metasurface unit was used to continue performing the energy harvesting task. No. Each metasurface unit represents the reflected signal from the k-th receiver.
[0047] In one embodiment of this application, for further definition and explanation, such as Figure 5 As shown, in embodiment step 102, based on the UAV action policy generation network that has completed model training, the current time slot UAV action parameters are generated according to the current time slot UAV state data. Before the intelligent metasurface action policy generation network that has completed model training generates the current time slot intelligent metasurface action parameters according to the current time slot intelligent metasurface state data, the embodiment method further includes: 301. Construct a simulation environment.
[0048] Accordingly, step 301 of the embodiment specifically includes: The communication channel model is constructed and expressed by the following formula. , , in, Represents the first in base stations and smart metasurfaces The LoS probability of each metasurface unit, where A and B both represent communication environment constants. Indicates the base station and the first The elevation angle between each metasurface unit Indicates the first At the height of a metasurface unit in time slot t, ( , ) indicates the first The horizontal coordinates of a metasurface unit in time slot t; A reflective model of an aerial intelligent metasurface is constructed, expressed by the following formula. , in, This represents the diagonal matrix of reflection coefficients. Indicates the first The amplitude reflection coefficient of each metasurface unit. Indicates the first The complex form of the phase shift of each metasurface unit. Indicates the number of metasurface units; The base station transmit signal model is constructed and expressed by the following formula. , Where X represents the signal vector actually transmitted by the base station. Represents the set of legitimate users. This indicates that the eavesdroppers have gathered. Indicates user The precoded vector, Indicates user A circularly symmetric complex Gaussian signal with zero mean and unit variance; A model for the propulsion energy consumption of unmanned aerial vehicles (UAVs) is constructed, expressed by the following formula. , in, This represents the propulsion energy consumption of the UAV in time slot t. Indicates the duration of the time slot. Indicates the hovering blade power. This represents the speed of the drone in time slot t. Indicates the rotor tip speed. Indicates the fuselage drag ratio. Indicates air density, Indicates rotor solidity, Indicates the rotor disk area, Indicates hovering induced power. This indicates the average induced speed during hovering; An aerial intelligent metasurface energy harvesting model is constructed, expressed by the following formula. , in, This represents the energy collected by the airborne intelligent metasurface in time slot t. Indicates the duration of the energy harvesting phase. Indicates the number of rows of metasurface cells. Indicates the number of columns for the metasurface unit. Indicates energy harvesting efficiency. Indicates the base station and the first The conjugate transpose of the channel vectors of each metasurface unit This represents the total set of users and eavesdroppers. Represents binary control variables; A communication energy consumption model for the information transmission phase of an aerial intelligent metasurface is constructed, expressed by the following formula: , in, G represents the communication energy consumption of the airborne intelligent metasurface during the information transmission phase in time slot t, and G represents the base station precoding matrix. Based on the communication channel model, the airborne intelligent metasurface reflection model, the base station transmission signal model, the UAV propulsion energy consumption model, the airborne intelligent metasurface energy harvesting model, and the airborne intelligent metasurface communication energy consumption model in the information transmission stage, a simulation environment is constructed to build the communication system vector in the simulation environment.
[0049] 302. Construct a UAV motion strategy generation network, an intelligent metasurface motion strategy generation network, and a centralized motion value estimation network.
[0050] Among them, the UAV action strategy generation network is used to generate UAV action strategies, the intelligent metasurface action strategy generation network is used to generate intelligent metasurface action strategies, and the centralized action value estimation network is used to evaluate the action value of the target aerial intelligent metasurface, that is, the joint action of intelligent metasurface and UAV.
[0051] 303. Generate multiple communication system vectors in a loop to construct an experience pool.
[0052] Accordingly, step 303 of the embodiment specifically includes: in the simulation environment, acquiring the current simulation time slot UAV state data; based on the current simulation time slot UAV state data, generating the current simulation time slot UAV action parameters based on the UAV action strategy generation network; simulating UAV flight based on the current simulation time slot UAV action parameters to obtain the next simulation time slot UAV state data, and calculating the current simulation time slot UAV propulsion energy consumption; acquiring the current simulation time slot intelligent metasurface state data; based on the current simulation time slot intelligent metasurface state data, generating the current simulation time slot intelligent metasurface action parameters based on the intelligent metasurface action strategy generation network; simulating spatiotemporal dual-domain resource allocation communication based on the current simulation time slot intelligent metasurface action parameters to obtain the next simulation time slot intelligent metasurface state data, and calculating the current simulation time slot communication system effective throughput and the current simulation time slot information transmission stage intelligence. The system calculates the global reward value for the current simulation time slot based on the energy consumption of the intelligent metasurface for communication and the energy collected by the intelligent metasurface during the current simulation time slot's information transmission phase, the energy collected by the intelligent metasurface during the current simulation time slot's information transmission phase, and the effective throughput of the communication system during the current simulation time slot. It then combines the current simulation time slot's UAV state data, intelligent metasurface state data, UAV action parameters, intelligent metasurface action parameters, global reward value, UAV state data for the next simulation time slot, and intelligent metasurface state data for the next simulation time slot to obtain the current simulation time slot's communication system vector. Finally, it constructs the next simulation time slot's communication system vector based on the next simulation time slot's UAV state data and intelligent metasurface state data, thus iteratively generating multiple communication system vectors.
[0053] In this embodiment, within the simulation environment constructed in step 301 of the embodiment, firstly, for the UAV action strategy generation network, the current simulation time slot UAV state data is input into the UAV action strategy generation network to generate the current simulation time slot UAV action parameters. Based on these generated parameters, UAV flight is simulated to obtain the next simulation time slot UAV state data. Then, based on the aforementioned UAV propulsion energy consumption model, the propulsion energy consumption of the current simulation time slot UAV is calculated. Simultaneously, for the intelligent metasurface action strategy generation network, the current simulation time slot intelligent metasurface state data is input into the network to generate the current simulation time slot intelligent metasurface action parameters. Based on these parameters, simulated spatiotemporal dual-domain resource allocation communication is performed to obtain the next simulation time slot intelligent metasurface state data. Finally, based on the effective throughput calculation formula for the communication system, the effective throughput of the current simulation time slot communication system is calculated. , , in, Indicates the effective throughput of the communication system. Indicates user Effective throughput in time slot t Indicates user In the time slot signal gain, Indicates user In the time slot Interference plus noise, Based on the aforementioned communication energy consumption model of the aerial intelligent metasurface during the information transmission phase, the communication energy consumption of the intelligent metasurface during the current simulation time slot's information transmission phase is calculated. Furthermore, based on the aforementioned energy harvesting model of the aerial intelligent metasurface, the harvested energy of the intelligent metasurface during the current simulation time slot is calculated. Further, based on the global reward value calculation formula, and according to the current simulation time slot's UAV propulsion energy consumption, intelligent metasurface communication energy consumption during the current simulation time slot's information transmission phase, intelligent metasurface harvested energy during the current simulation time slot, and the effective throughput of the current simulation time slot's communication system, the global reward value for the current simulation time slot is calculated. , in, This represents the global reward value for time slot t; Finally, the current simulation time slot UAV state data, the current simulation time slot intelligent metasurface state data, the current simulation time slot UAV action parameters, the current simulation time slot intelligent metasurface action parameters, the current simulation time slot global reward value, the next simulation time slot UAV state data, and the next simulation time slot intelligent metasurface state data are combined to obtain the current simulation time slot communication system vector. Based on the above steps, the next simulation time slot communication system vector is constructed according to the next simulation time slot UAV state data and the next simulation time slot intelligent metasurface state data, and multiple communication system vectors are generated iteratively to build an experience pool.
[0054] 304. Training Concentrated Motion Value Estimation Network.
[0055] Accordingly, step 304 of the embodiment specifically includes: generating the estimated value of the action in the next simulation time slot based on the UAV state data and the smart metasurface state data in the next simulation time slot, and calculating the expected value of the action in the current simulation time slot based on the Bellman equation, the estimated value of the action in the next simulation time slot and the global reward value in the current simulation time slot; generating the estimated value of the action in the current simulation time slot based on the UAV state data and the smart metasurface state data in the current simulation time slot, and updating the parameters of the centralized action value estimation network using the gradient descent algorithm based on the mean square error loss value between the estimated value of the action in the current simulation time slot and the expected value of the action in the current simulation time slot.
[0056] In this embodiment, firstly, the UAV state data and the smart metasurface state data for the next simulation time slot are input into a centralized action value estimation network to generate the action estimation value for the next simulation time slot. Then, based on the Bellman equation, the expected action value for the current simulation time slot is calculated according to the action estimation value for the next simulation time slot and the global reward value for the current simulation time slot. The Bellman equation is expressed as the following formula. , in, This represents the expected value of the current simulated time slot action. This represents the combination of the current simulation time slot UAV state data and the current simulation time slot smart metasurface state data. This represents the global reward value for the current simulation time slot. Indicates the discount factor. Indicates the estimated value of the next simulation time slot action. Simultaneously, the current simulation time slot UAV state data and the current simulation time slot smart metasurface state data are input into the centralized action value estimation network to generate the action estimate value for the current simulation time slot. Finally, based on the mean squared error loss value between the current simulation time slot action estimate value and the current simulation time slot action expectation value, the parameters of the centralized action value estimation network are updated using the gradient descent algorithm. , in, E represents the mean squared error loss value, and E represents the expected value. Indicates the value of the current simulated time slot action estimation. Compared with the expected value of current simulated time slot actions The square of the difference.
[0057] 305. Based on the generalized dominance estimation algorithm, calculate the dominance value of the current simulation time slot according to the estimated value of the action in the current simulation time slot, the estimated value of the action in the next simulation time slot, and the global reward value of the current simulation time slot.
[0058] The generalized dominance estimation algorithm is expressed by the following formula. , This indicates the current simulation slot dominance value. This represents the attenuation factor for the advantage estimate. This represents the sum of the global reward value of the current simulation time slot and the estimated value of the action in the next simulation time slot, minus the estimated value of the action in the current simulation time slot.
[0059] 306. Training the drone motion strategy generation network.
[0060] Accordingly, step 306 of the embodiment specifically includes: based on the near-end policy optimization algorithm, constructing a UAV action policy gradient loss function according to the current simulation time slot advantage value and the current simulation time slot UAV action parameters; and updating the parameters of the UAV action policy generation network based on the UAV action policy gradient loss function through the gradient ascent algorithm. 307. Training the intelligent metasurface action strategy generation network.
[0061] Accordingly, step 307 of the embodiment specifically includes: based on the near-end policy optimization algorithm, constructing a gradient loss function for the intelligent metasurface action policy according to the current simulation time slot advantage value and the intelligent metasurface action parameters of the current simulation time slot; and updating the parameters of the intelligent metasurface action policy generation network based on the gradient ascent algorithm using the gradient loss function for the intelligent metasurface action policy.
[0062] The action policy gradient loss function is expressed by the following formula. , in, This represents the gradient loss value of the drone's action policy or the gradient loss value of the intelligent metasurface's action policy. The strategy ratio is represented (which can be obtained from the UAV action parameters or the smart metasurface action parameters in the current simulation time slot). This indicates the current simulation slot dominance value. This indicates the PPO shearing threshold, and clip indicates the shearing mechanism.
[0063] 308. Alternately update the centralized action value estimation network, the UAV action policy generation network, and the intelligent metasurface action policy generation network until the network converges, resulting in the centralized action value estimation network, the UAV action policy generation network, and the intelligent metasurface action policy generation network that have completed model training.
[0064] This application provides a communication method based on a multi-agent proximal policy optimization algorithm. First, during the execution of a communication task by a target aerial intelligent metasurface, the current time-slot state data of the target aerial intelligent metasurface is collected in real time. The aerial intelligent metasurface represents the product of the combination of the intelligent metasurface and the UAV. The current time-slot state data includes the current time-slot UAV state data and the current time-slot intelligent metasurface state data. Second, based on a UAV action policy generation network that has completed model training, action parameters for the current time-slot UAV are generated according to the current time-slot UAV state data. Similarly, based on a trained intelligent metasurface action policy generation network, action parameters for the current time-slot intelligent metasurface are generated according to the current time-slot intelligent metasurface state data. Both the UAV action policy generation network and the intelligent metasurface action policy generation network are obtained through joint training using a multi-agent proximal policy optimization algorithm, employing centralized value evaluation and decentralized policy execution. Finally, the UAV is controlled to execute a flight task based on the current time-slot UAV action parameters, and the intelligent metasurface is controlled to execute a communication task based on spatiotemporal dual-domain resource allocation based on the current time-slot intelligent metasurface action parameters. Compared with existing technologies, the centralized value assessment mechanism adopted in this application can perform unified value estimation of the joint actions of UAV and intelligent metasurface based on global system state information. This adaptively adjusts the weight distribution among communication throughput, energy efficiency and safety performance during multi-objective optimization, significantly enhancing the system's multi-objective collaborative optimization capability. Simultaneously, it provides globally consistent and dynamically adaptable optimization guidance for the training of the UAV action strategy generation network and the intelligent metasurface action strategy generation network, effectively avoiding policy oscillations caused by objective conflicts. Furthermore, through a decentralized policy execution mechanism, each agent, guided by the global value signal, performs policy optimization for local tasks such as UAV trajectory control and metasurface beamforming. This achieves optimal performance in its respective control dimension and realizes dynamic complementarity and joint improvement of the overall system efficiency through global collaboration, thereby achieving long-term, efficient, and secure communication in complex communication environments.
[0065] Furthermore, as a response to the above Figure 1 The implementation of the method shown in this application provides a communication system based on a multi-agent proximal policy optimization algorithm, such as... Figure 6 As shown, the system includes: Status data acquisition module 41, action parameter generation module 42, control module 43; The status data acquisition module 41 is used to collect the current time slot status data of the target airborne intelligent metasurface in real time during the communication task performed by the target airborne intelligent metasurface. The airborne intelligent metasurface is used to characterize the product of the combination of intelligent metasurface and UAV. The current time slot status data includes the current time slot UAV status data and the current time slot intelligent metasurface status data. The motion parameter generation module 42 is used to generate the current time slot UAV motion parameters based on the UAV state data of the current time slot, based on the UAV motion policy generation network that has completed model training, and to generate the current time slot intelligent metasurface motion parameters based on the intelligent metasurface state data of the current time slot, based on the intelligent metasurface motion policy generation network that has completed model training. The UAV motion policy generation network and the intelligent metasurface motion policy generation network are obtained by joint training based on the multi-agent proximal policy optimization algorithm through centralized value evaluation and decentralized policy execution. The control module 43 is used to control the UAV to perform flight missions according to the current time slot UAV action parameters, and to control the smart metasurface to perform communication missions based on spatiotemporal dual-domain resource allocation according to the current time slot smart metasurface action parameters.
[0066] In specific application scenarios, the current time slot intelligent metasurface action parameters include the current time slot intelligent metasurface phase matrix, the current time slot time allocation factor, and the current time slot cell reflection control parameters. The control module is used for: Based on the current time slot allocation factor, the current time slot is divided into stages to obtain the energy harvesting stage and the information transmission stage. During the energy harvesting phase, all metasurface units contained in the intelligent metasurface are controlled to perform energy harvesting tasks; During the information transmission phase, based on the current time slot unit reflection control parameters, all metasurface units are divided into multiple reflection units and multiple energy harvesting units. The multiple energy harvesting units are controlled to perform energy harvesting tasks, and the multiple reflection units are controlled to configure corresponding phase offset parameters based on the current time slot intelligent metasurface phase matrix to perform communication tasks, so as to carry out communication operations based on spatiotemporal dual-domain resource allocation.
[0067] In specific application scenarios, prior to the action parameter generation module, the system further includes a model training module, including: Network building units are used to construct UAV motion strategy generation networks, intelligent metasurface motion strategy generation networks, and centralized motion value estimation networks. The communication system vector generation unit is used to acquire the current simulation time slot UAV state data in a simulation environment, generate UAV action parameters for the current simulation time slot based on the UAV action strategy generation network, simulate UAV flight based on the current simulation time slot action parameters, obtain the next simulation time slot UAV state data, calculate the UAV propulsion energy consumption for the current simulation time slot, and acquire the current simulation time slot intelligent metasurface state data. Based on the current simulation time slot intelligent metasurface state data, generate intelligent metasurface action parameters for the current simulation time slot based on the intelligent metasurface action strategy generation network, simulate spatiotemporal dual-domain resource allocation communication based on the current simulation time slot intelligent metasurface action parameters, obtain the next simulation time slot intelligent metasurface state data, and calculate the effective throughput of the current simulation time slot communication system and the intelligent metasurface communication energy consumption during the current simulation time slot information transmission phase. The system collects energy from the intelligent metasurface in the current simulation time slot, and calculates the global reward value for the current simulation time slot based on the UAV propulsion energy consumption, the intelligent metasurface communication energy consumption during the information transmission phase of the current simulation time slot, the collected energy from the intelligent metasurface in the current simulation time slot, and the effective throughput of the communication system in the current simulation time slot. It then combines the UAV state data, the intelligent metasurface state data, the UAV action parameters, the intelligent metasurface action parameters, the global reward value, the UAV state data, and the intelligent metasurface state data of the next simulation time slot to obtain the current simulation time slot communication system vector. Finally, it constructs the next simulation time slot communication system vector based on the UAV state data and the intelligent metasurface state data of the next simulation time slot, and iteratively generates multiple communication system vectors. A centralized action value estimation network training unit is used to generate the action estimation value for the next simulation time slot based on the UAV state data and smart metasurface state data in the next simulation time slot, and to calculate the expected action value for the current simulation time slot based on the Bellman equation, the action estimation value for the next simulation time slot, and the global reward value for the current simulation time slot. It also generates the action estimation value for the current simulation time slot based on the UAV state data and smart metasurface state data in the current simulation time slot, and updates the parameters of the centralized action value estimation network using a gradient descent algorithm based on the mean squared error loss between the action estimation value and the expected action value for the current simulation time slot. The dominance value calculation unit is used to calculate the dominance value of the current simulation time slot based on the generalized dominance estimation algorithm, according to the estimated value of the current simulation time slot action, the estimated value of the next simulation time slot action, and the global reward value of the current simulation time slot. The UAV action policy generation network training unit is used to construct a UAV action policy gradient loss function based on the near-end policy optimization algorithm, according to the current simulation time slot advantage value and the UAV action parameters of the current simulation time slot, and update the parameters of the UAV action policy generation network based on the UAV action policy gradient loss function through the gradient ascent algorithm. The training unit of the intelligent metasurface action policy generation network is used to construct the gradient loss function of the intelligent metasurface action policy based on the proximal policy optimization algorithm, according to the current simulation time slot advantage value and the intelligent metasurface action parameters of the current simulation time slot, and update the parameters of the intelligent metasurface action policy generation network based on the gradient ascent algorithm. The training termination unit is used to alternately update the centralized action value estimation network, the UAV action policy generation network, and the intelligent metasurface action policy generation network until the network converges, resulting in the centralized action value estimation network, the UAV action policy generation network, and the intelligent metasurface action policy generation network that have completed model training.
[0068] In specific application scenarios, prior to the network construction unit, the model training module further includes a simulation environment construction unit, used for: The communication channel model is constructed and expressed by the following formula. , , in, Represents the first in base stations and smart metasurfaces The LoS probability of each metasurface unit, where A and B both represent communication environment constants. Indicates the base station and the first The elevation angle between each metasurface unit Indicates the first At the height of a metasurface unit in time slot t, ( , ) indicates the first The horizontal coordinates of a metasurface unit in time slot t; A reflective model of an aerial intelligent metasurface is constructed, expressed by the following formula. , in, This represents the diagonal matrix of reflection coefficients. Indicates the first The amplitude reflection coefficient of each metasurface unit. Indicates the first The complex form of the phase shift of each metasurface unit. Indicates the number of metasurface units; The base station transmit signal model is constructed and expressed by the following formula. , Where X represents the signal vector actually transmitted by the base station. Represents the set of legitimate users. This indicates that the eavesdroppers have gathered. Indicates user The precoded vector, Indicates user A circularly symmetric complex Gaussian signal with zero mean and unit variance; A model for the propulsion energy consumption of unmanned aerial vehicles (UAVs) is constructed, expressed by the following formula. , in, This represents the propulsion energy consumption of the UAV in time slot t. Indicates the duration of the time slot. Indicates the hovering blade power. This represents the speed of the drone in time slot t. Indicates the rotor tip speed. Indicates the fuselage drag ratio. Indicates air density, Indicates rotor solidity, Indicates the rotor disk area, Indicates hovering induced power. This indicates the average induced speed during hovering; An aerial intelligent metasurface energy harvesting model is constructed, expressed by the following formula. , in, This represents the energy collected by the airborne intelligent metasurface in time slot t. Indicates the duration of the energy harvesting phase. Indicates the number of rows of metasurface cells. Indicates the number of columns for the metasurface unit. Indicates energy harvesting efficiency. Indicates the base station and the first The conjugate transpose of the channel vectors of each metasurface unit This represents the total set of users and eavesdroppers. Represents binary control variables; A communication energy consumption model for the information transmission phase of an aerial intelligent metasurface is constructed, expressed by the following formula: , in, G represents the communication energy consumption of the airborne intelligent metasurface during the information transmission phase in time slot t, and G represents the base station precoding matrix. Based on the communication channel model, the airborne intelligent metasurface reflection model, the base station transmission signal model, the UAV propulsion energy consumption model, the airborne intelligent metasurface energy harvesting model, and the airborne intelligent metasurface communication energy consumption model during the information transmission phase, a simulation environment is constructed to build a communication system vector within the simulation environment.
[0069] In specific application scenarios, the formula for calculating the effective throughput of a communication system is expressed as follows: , , in, Indicates the effective throughput of the communication system. Indicates user Effective throughput in time slot t Indicates user In the time slot signal gain, Indicates user In the time slot Interference and noise.
[0070] In specific application scenarios, the formula for calculating the global reward value is expressed as follows: , in, This represents the global reward value for time slot t.
[0071] In specific application scenarios, the mean squared error loss function is expressed by the following formula. , in, E represents the mean squared error loss value, and E represents the expected value. Indicates the value of the current simulated time slot action estimation. Compared with the expected value of current simulated time slot actions The square of the difference; The gradient loss function for the action policy is expressed by the following formula.
[0072] in, This represents the gradient loss value of the drone's action policy or the gradient loss value of the intelligent metasurface's action policy. Indicates the strategy ratio, This indicates the current simulation slot dominance value. This indicates the PPO shearing threshold, and clip indicates the shearing mechanism.
[0073] This application provides a communication system based on a multi-agent near-end policy optimization algorithm. First, during the execution of a communication task by a target aerial intelligent metasurface, the system collects real-time time-slot state data of the target aerial intelligent metasurface. The aerial intelligent metasurface represents the product of the combination of the intelligent metasurface and the UAV. The current time-slot state data includes the current time-slot UAV state data and the current time-slot intelligent metasurface state data. Second, based on a UAV action policy generation network that has completed model training, action parameters for the current time-slot UAV are generated according to the current time-slot UAV state data. Similarly, based on a trained intelligent metasurface action policy generation network, action parameters for the current time-slot intelligent metasurface are generated according to the current time-slot intelligent metasurface state data. Both the UAV action policy generation network and the intelligent metasurface action policy generation network are obtained through joint training using a multi-agent near-end policy optimization algorithm, employing centralized value evaluation and decentralized policy execution. Finally, the system controls the UAV to execute flight tasks based on the current time-slot UAV action parameters and controls the intelligent metasurface to execute communication tasks based on spatiotemporal dual-domain resource allocation based on the current time-slot intelligent metasurface action parameters. Compared with existing technologies, the embodiments of this application utilize centralized value assessment, which enables adaptive adjustment of the focus among multiple optimization objectives based on global system state information, thereby improving the multi-objective collaborative optimization capability. This provides a stable and consistent optimization guide for the training of the UAV action strategy generation network and the intelligent metasurface action strategy generation network. Furthermore, by utilizing decentralized policy execution, the UAV action strategy generation network and the intelligent metasurface action strategy generation network focus on their respective performance indicators according to the centralized value assessment. This allows the generated flight strategies and signal reflection strategies to achieve optimal performance in their respective control dimensions at the local level, and to achieve efficient collaboration and dynamic complementarity at the global level, thus fulfilling the long-term and efficient secure communication requirements.
[0074] According to one embodiment of this application, a storage medium is provided, the storage medium storing at least one executable instruction, which can execute the communication method based on the multi-agent proximal policy optimization algorithm in any of the above method embodiments.
[0075] Based on this understanding, the technical solution of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, or portable hard drive) and includes several instructions to cause a computer device (such as a personal computer, server, or network device) to execute the methods described in the various implementation scenarios of this application.
[0076] Figure 7 The diagram shows a structural schematic of a terminal according to one embodiment of the present application. The specific embodiments of the present application do not limit the specific implementation of the terminal.
[0077] like Figure 7 As shown, the terminal may include: a processor 502, a communications interface 504, a memory 506, and a communications bus 508.
[0078] The processor 502, communication interface 504, and memory 506 communicate with each other via communication bus 508.
[0079] Communication interface 504 is used to communicate with other network elements such as clients or other servers.
[0080] The processor 502 is used to execute program 510, specifically to execute the relevant steps in the above-described communication method embodiment based on the multi-agent proximal policy optimization algorithm.
[0081] Specifically, program 510 may include program code that includes computer operation instructions.
[0082] Processor 502 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application. The computer device includes one or more processors, which may be processors of the same type, such as one or more CPUs; or they may be processors of different types, such as one or more CPUs and one or more ASICs.
[0083] Memory 506 is used to store program 510. Memory 506 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.
[0084] Specifically, program 510 can be used to cause processor 502 to perform the following operations: During the communication mission performed by the target aerial intelligent metasurface, the current time slot state data of the target aerial intelligent metasurface is collected in real time. The aerial intelligent metasurface is used to characterize the product of the combination of intelligent metasurface and UAV. The current time slot state data includes the current time slot UAV state data and the current time slot intelligent metasurface state data. Based on the UAV action policy generation network that has completed model training, the current time slot UAV action parameters are generated according to the current time slot UAV state data. Based on the intelligent metasurface action policy generation network that has completed model training, the current time slot intelligent metasurface action parameters are generated according to the current time slot intelligent metasurface state data. The UAV action policy generation network and the intelligent metasurface action policy generation network are obtained by joint training based on the multi-agent proximal policy optimization algorithm through centralized value evaluation and decentralized policy execution. The UAV is controlled to perform flight missions based on the current time slot UAV action parameters, and the smart metasurface is controlled to perform communication missions based on spatiotemporal dual-domain resource allocation based on the current time slot smart metasurface action parameters.
[0085] The storage medium may also include an operating system and a network communication module. The operating system is a program that manages the hardware and software resources of the physical device using the aforementioned communication method based on the multi-agent proximal policy optimization algorithm, supporting the operation of information processing programs and other software and / or programs. The network communication module is used to enable communication between the various components within the storage medium, as well as communication with other hardware and software in the information processing physical device.
[0086] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For system embodiments, since they largely correspond to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0087] The methods and systems of this application may be implemented in many ways. For example, they may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above-described order of steps for the methods is for illustrative purposes only, and the steps of the methods of this application are not limited to the order specifically described above, unless otherwise specifically stated. Furthermore, in some embodiments, this application may also be implemented as a program recorded on a recording medium, the program including machine-readable instructions for implementing the methods according to this application. Thus, this application also covers recording media storing programs for performing the methods according to this application.
[0088] Obviously, those skilled in the art should understand that the modules or steps of this application described above can be implemented using general-purpose computing systems. They can be centralized on a single computing system or distributed across a network of multiple computing systems. Optionally, they can be implemented using program code executable by a computing system, thereby storing them in a storage system for execution by the computing system. In some cases, the steps shown or described can be performed in a different order than those presented here, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, this application is not limited to any particular combination of hardware and software.
[0089] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A communication method based on a multi-agent proximal policy optimization algorithm, characterized in that, include: During the communication mission performed by the target airborne intelligent metasurface, the current time slot status data of the target airborne intelligent metasurface is collected in real time. The airborne intelligent metasurface (ARIS) is used to characterize the product of the combination of intelligent metasurface (RIS) and unmanned aerial vehicle (UAV). The current time slot status data includes the current time slot UAV status data and the current time slot intelligent metasurface status data. Based on the UAV action policy generation network that has completed model training, the current time slot UAV action parameters are generated according to the current time slot UAV state data. Based on the intelligent metasurface action policy generation network that has completed model training, the current time slot intelligent metasurface action parameters are generated according to the current time slot intelligent metasurface state data. The UAV action policy generation network and the intelligent metasurface action policy generation network are obtained by joint training based on the multi-agent proximal policy optimization algorithm through centralized value evaluation and decentralized policy execution. The UAV is controlled to perform flight missions based on the current time slot UAV action parameters, and the smart metasurface is controlled to perform communication missions based on spatiotemporal dual-domain resource allocation based on the current time slot smart metasurface action parameters.
2. The method according to claim 1, characterized in that, The current time-slot smart metasurface action parameters include the current time-slot smart metasurface phase matrix, the current time-slot time allocation factor, and the current time-slot cell reflection control parameters. Controlling the smart metasurface to perform a communication task based on spatiotemporal dual-domain resource allocation according to the current time-slot smart metasurface action parameters includes: Based on the current time slot allocation factor, the current time slot is divided into stages to obtain the energy harvesting stage and the information transmission stage. During the energy harvesting phase, all metasurface units contained in the intelligent metasurface are controlled to perform energy harvesting tasks; During the information transmission phase, based on the current time slot unit reflection control parameters, all metasurface units are divided into multiple reflection units and multiple energy harvesting units. The multiple energy harvesting units are controlled to perform energy harvesting tasks, and the multiple reflection units are controlled to configure corresponding phase offset parameters based on the current time slot intelligent metasurface phase matrix to perform communication tasks, so as to carry out communication operations based on spatiotemporal dual-domain resource allocation.
3. The method according to claim 1, characterized in that, The method further includes, before the UAV action policy generation network based on the completed model training generates the current time slot UAV action parameters according to the current time slot UAV state data, and before the intelligent metasurface action policy generation network based on the completed model training generates the current time slot intelligent metasurface action parameters according to the current time slot intelligent metasurface state data, the method further includes: Construct a drone motion strategy generation network, an intelligent metasurface motion strategy generation network, and a centralized motion value estimation network; In the simulation environment, the current simulation time slot UAV state data is acquired. Based on the current simulation time slot UAV state data and the UAV action strategy generation network, UAV action parameters for the current simulation time slot are generated. Based on the current simulation time slot UAV action parameters, simulated UAV flight is performed to obtain the next simulation time slot UAV state data, and the propulsion energy consumption of the current simulation time slot UAV is calculated. Additionally, the current simulation time slot intelligent metasurface state data is acquired. Based on the current simulation time slot intelligent metasurface state data and the intelligent metasurface action strategy generation network, intelligent metasurface action parameters for the current simulation time slot are generated. Based on the current simulation time slot intelligent metasurface action parameters, simulated spatiotemporal dual-domain resource allocation communication is performed to obtain the next simulation time slot intelligent metasurface state data, and the effective throughput of the current simulation time slot communication system, the intelligent metasurface communication energy consumption during the current simulation time slot information transmission phase, and the current simulation time slot's... The system collects energy using an intelligent metasurface, and calculates the global reward value for the current simulation time slot based on the current simulation time slot UAV propulsion energy consumption, the current simulation time slot information transmission stage intelligent metasurface communication energy consumption, the current simulation time slot intelligent metasurface collected energy, and the current simulation time slot communication system effective throughput. It also combines the current simulation time slot UAV state data, the current simulation time slot intelligent metasurface state data, the current simulation time slot UAV action parameters, the current simulation time slot intelligent metasurface action parameters, the current simulation time slot global reward value, the next simulation time slot UAV state data, and the next simulation time slot intelligent metasurface state data to obtain the current simulation time slot communication system vector. Finally, it constructs the next simulation time slot communication system vector based on the next simulation time slot UAV state data and the next simulation time slot intelligent metasurface state data, and iteratively generates multiple communication system vectors. Based on the centralized action value estimation network, the estimated action value for the next simulation time slot is generated according to the UAV state data and the smart metasurface state data of the next simulation time slot. Based on the Bellman equation, the expected action value for the current simulation time slot is calculated according to the estimated action value for the next simulation time slot and the global reward value of the current simulation time slot. Furthermore, based on the centralized action value estimation network, the estimated action value for the current simulation time slot is generated according to the UAV state data and the smart metasurface state data of the current simulation time slot. Finally, based on the mean squared error loss between the estimated action value and the expected action value of the current simulation time slot, the parameters of the centralized action value estimation network are updated using a gradient descent algorithm. Based on the generalized dominance estimation algorithm, the dominance value of the current simulation time slot is calculated according to the estimated value of the current simulation time slot action, the estimated value of the next simulation time slot action, and the global reward value of the current simulation time slot. Based on the near-end policy optimization algorithm, a UAV action policy gradient loss function is constructed according to the current simulation time slot advantage value and the current simulation time slot UAV action parameters. Based on the UAV action policy gradient loss function, the parameters of the UAV action policy generation network are updated through the gradient ascent algorithm. Based on the near-end policy optimization algorithm, a gradient loss function for intelligent metasurface action policy is constructed according to the current simulation time slot advantage value and the intelligent metasurface action parameters of the current simulation time slot. Based on the gradient loss function for intelligent metasurface action policy, the parameters of the intelligent metasurface action policy generation network are updated through the gradient ascent algorithm. The centralized action value estimation network, the UAV action policy generation network, and the intelligent metasurface action policy generation network are alternately updated until the network converges, resulting in a centralized action value estimation network, a UAV action policy generation network, and an intelligent metasurface action policy generation network that have completed model training.
4. The method according to claim 2, characterized in that, Before constructing the UAV motion policy generation network, the intelligent metasurface motion policy generation network, and the centralized motion value estimation network, the method further includes: The communication channel model is constructed and expressed by the following formula. , , in, Represents the first in base stations and smart metasurfaces The LoS probability of each metasurface unit, where A and B both represent communication environment constants. Indicates the base station and the first The elevation angle between each metasurface unit Indicates the first At the height of a metasurface unit in time slot t, ( , ) indicates the first The horizontal coordinates of a metasurface unit in time slot t; A reflective model of an aerial intelligent metasurface is constructed, expressed by the following formula. , in, This represents the diagonal matrix of reflection coefficients. Indicates the first The amplitude reflection coefficient of each metasurface unit. Indicates the first The complex form of the phase shift of each metasurface unit. Indicates the number of metasurface units; The base station transmit signal model is constructed and expressed by the following formula. , Where X represents the signal vector actually transmitted by the base station. Represents the set of legitimate users. This indicates that the eavesdroppers have gathered. Indicates user The precoded vector, Indicates user A circularly symmetric complex Gaussian signal with zero mean and unit variance; A model for the propulsion energy consumption of unmanned aerial vehicles (UAVs) is constructed, expressed by the following formula. , in, This represents the propulsion energy consumption of the UAV in time slot t. Indicates the duration of the time slot. Indicates the hovering blade power. This represents the speed of the drone in time slot t. Indicates the rotor tip speed. Indicates the fuselage drag ratio. Indicates air density, Indicates rotor solidity, Indicates the rotor disk area, Indicates hovering induced power. This indicates the average induced speed during hovering; An aerial intelligent metasurface energy harvesting model is constructed, expressed by the following formula. , in, This represents the energy collected by the airborne intelligent metasurface in time slot t. Indicates the duration of the energy harvesting phase. Indicates the number of rows of metasurface cells. Indicates the number of columns for the metasurface unit. Indicates energy harvesting efficiency. Indicates the base station and the first The conjugate transpose of the channel vectors of each metasurface unit This represents the total set of users and eavesdroppers. Represents binary control variables; A communication energy consumption model for the information transmission phase of an aerial intelligent metasurface is constructed, expressed by the following formula: , in, G represents the communication energy consumption of the airborne intelligent metasurface during the information transmission phase in time slot t, and G represents the base station precoding matrix. Based on the communication channel model, the airborne intelligent metasurface reflection model, the base station transmission signal model, the UAV propulsion energy consumption model, the airborne intelligent metasurface energy harvesting model, and the airborne intelligent metasurface communication energy consumption model during the information transmission phase, a simulation environment is constructed to build a communication system vector within the simulation environment.
5. The method according to claim 3, characterized in that, The formula for calculating the effective throughput of a communication system is expressed as follows: , , in, Indicates the effective throughput of the communication system. Indicates user Effective throughput in time slot t Indicates user In the time slot signal gain, Indicates user In the time slot Interference and noise.
6. The method according to claim 3, characterized in that, The formula for calculating the global reward value is expressed as follows: , in, This represents the global reward value for time slot t.
7. The method according to claim 3, characterized in that, The mean squared error loss function is expressed by the following formula. , in, E represents the mean squared error loss value, and E represents the expected value. Indicates the value of the current simulated time slot action estimation. Compared with the expected value of current simulated time slot actions The square of the difference; The gradient loss function for the action policy is expressed by the following formula. in, This represents the gradient loss value of the drone's action policy or the gradient loss value of the intelligent metasurface's action policy. Indicates the strategy ratio, This indicates the current simulation slot dominance value. This indicates the PPO shearing threshold, and clip indicates the shearing mechanism.
8. A communication system based on a multi-agent proximal policy optimization algorithm, characterized in that, include: The status data acquisition module is used to collect the current time slot status data of the target airborne intelligent metasurface in real time during the communication mission of the target airborne intelligent metasurface. The airborne intelligent metasurface is used to characterize the product of the combination of intelligent metasurface and UAV. The current time slot status data includes the current time slot UAV status data and the current time slot intelligent metasurface status data. The motion parameter generation module is used to generate UAV motion parameters for the current time slot based on the UAV state data of the current time slot, and to generate intelligent metasurface motion parameters for the current time slot based on the intelligent metasurface state data of the current time slot, using a UAV motion policy generation network that has completed model training. The UAV motion policy generation network and the intelligent metasurface motion policy generation network are obtained by joint training based on a multi-agent proximal policy optimization algorithm through centralized value evaluation and decentralized policy execution. The control module is used to control the UAV to perform flight missions according to the current time slot UAV action parameters, and to control the smart metasurface to perform communication missions based on spatiotemporal dual-domain resource allocation according to the current time slot smart metasurface action parameters.
9. A storage medium storing at least one executable instruction, characterized in that, The executable instructions cause the processor to perform the operations corresponding to the communication method based on the multi-agent proximal policy optimization algorithm as described in any one of claims 1-7.
10. A terminal, comprising: The processor, memory, communication interface, and communication bus are provided, wherein the processor, memory, and communication interface communicate with each other via the communication bus. The memory is used to store at least one executable instruction, characterized in that the executable instruction causes the processor to perform the operation corresponding to the communication method based on the multi-agent proximal policy optimization algorithm as described in any one of claims 1-7.