A High-Efficiency User Association and Trajectory Optimization Method for Multi-Intelligent Reflector-Assisted Unmanned Aerial Vehicle Networks

By using multi-intelligent reflector-assisted UAV network scenarios, combined with phase alignment and multi-agent optimization algorithms, the problems of coverage blind spots and resource coupling in UAV networks are solved, the task offloading is increased and energy consumption is reduced, and efficient user association and trajectory optimization are achieved.

CN122138179APending Publication Date: 2026-06-02QINGHAI UNIV FOR NATITIES
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
QINGHAI UNIV FOR NATITIES
Filing Date
2026-01-23
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

In existing technologies, multi-intelligent reflector-assisted UAV networks suffer from problems such as coverage blind spots, resource heterogeneity, and strategy homogenization in dynamic environments in terms of user association, trajectory optimization, and resource allocation, resulting in high computational complexity, high energy consumption, and slow convergence speed.

Method used

We construct a UAV network scenario assisted by multiple intelligent reflectors, optimize channel gain through phase alignment principle, adopt an improved multi-agent near-end policy optimization algorithm and convex optimization method, decouple user association and computing resource allocation in a hierarchical manner, optimize UAV trajectory by using action masking and hybrid expert model, and perform task offloading by combining orthogonal frequency division multiple access technology.

Benefits of technology

Without increasing the cost of drone deployment, it significantly increased the total amount of task offloading, reduced energy consumption, improved the convergence speed and training stability of the algorithm in dynamic environments, and achieved precise supply of physical layer resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122138179A_ABST
    Figure CN122138179A_ABST
Patent Text Reader

Abstract

This invention discloses a high-energy-efficiency user association and trajectory optimization method for multi-intelligent reflector-assisted unmanned aerial vehicle (UAV) networks. First, by introducing multiple intelligent reflectors and employing a phase-alignment-based phase-shift control strategy, a high-quality virtual line-of-sight link is constructed, enabling effective service to users in coverage blind spots. Then, a hierarchical decoupling strategy is proposed, utilizing a channel-based low-complexity matching mechanism and a convex optimization method to solve user association and computational resource allocation respectively, effectively reducing computational complexity. Simultaneously, trajectory optimization is modeled as a partially observable Markov decision process and solved using an improved multi-agent near-end policy optimization algorithm. By introducing a hybrid expert model to enhance policy diversity and using an action masking mechanism to enforce constraints on flight boundaries, the system can autonomously achieve the optimal balance between energy efficiency and service quality in dynamic environments. Simulation experiments verify the effectiveness and robustness of this scheme in improving task offloading and reducing energy consumption.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of wireless communication and mobile edge computing technology, specifically to a high-energy-efficiency method for jointly optimizing UAV flight trajectory, intelligent reflector phase shift control, user association, and computing resource allocation in a multi-intelligent reflector-assisted UAV network. Background Technology

[0002] The statements in this section are merely background information relating to this disclosure, and these statements may constitute prior art. In the process of developing this invention, the inventors discovered at least the following problems in the prior art.

[0003] With the development of 6G networks, compute-intensive applications pose significant challenges to the computing power of ground users. Drones, with their high flexibility and scalability, are considered an effective solution for addressing resource shortages in hotspot areas as aerial edge servers. However, drones have limited coverage, making it difficult to serve users scattered outside the coverage area, and increasing the number of drones significantly increases system costs and energy consumption. Intelligent reflectors, as low-cost and easily deployed devices, can reconstruct the wireless environment and establish virtual line-of-sight links by adjusting the phase shift of reflective elements, thereby enhancing the coverage range of drones.

[0004] Despite the great potential of introducing smart reflective surfaces into drone networks, many challenges remain.

[0005] First, the mobility of UAVs leads to dynamic changes in the channel environment, rendering static user association strategies ineffective. Second, resource heterogeneity makes joint scheduling of communication and computing resources difficult. Finally, UAV trajectory design needs to consider overlapping coverage areas and flight boundary constraints. However, traditional optimization methods struggle to handle this highly coupled mixed-integer nonlinear programming problem, while existing single-agent reinforcement learning methods often face problems of excessively large action spaces and policy homogeneity in multi-UAV scenarios, resulting in slow convergence and poor performance.

[0006] The technical solution of patent application number 202311835980.2, entitled "User Association and Trajectory Optimization Method Based on Multi-Intelligent Reflector UAV Communication," first decomposes the system energy efficiency maximization problem into three sub-problems: user association decision-making, UAV trajectory optimization, and joint decoding scheduling and power allocation. Then, it utilizes inverse soft Q-learning to obtain a long-term optimized user association strategy, and uses successive convex approximation and Dinkelbach ensemble methods to simplify sub-problem 2 from a non-convex fractional optimization problem into a subtractive convex optimization problem. Finally, it uses a penalty-based successive convex approximation method to transform it into a convex optimization problem with a penalty term. While the above-mentioned solutions alleviate the shortcomings of introducing intelligent reflective surfaces into UAV networks to some extent, they still have the problem of UAV coverage blind spots. Furthermore, the challenges of mixed-integer nonlinear programming with deep coupling of user association, trajectory control, and resource allocation, as well as the problems of policy homogenization and boundary constraints in multi-UAV collaboration, have not been well resolved. They cannot increase the total amount of task offloading without increasing the cost of UAV deployment, nor can they reduce computational complexity and achieve precise supply of physical layer resources. They also cannot avoid ineffective exploration, resulting in unsatisfactory convergence speed and training stability of the algorithm in dynamic environments. Summary of the Invention

[0007] In view of the above problems, the purpose of this invention is to solve some of the problems in the prior art, or at least alleviate these problems.

[0008] A highly energy-efficient user association and trajectory optimization method for multi-intelligent reflector-assisted unmanned aerial vehicle (UAV) networks includes the following steps:

[0009] Scenario and Model Construction: Construct a UAV computing network scenario assisted by multiple intelligent reflectors, and establish computing models, communication models and energy consumption models;

[0010] Construct an optimization problem and decompose it into sub-problems: Construct a joint optimization problem with the goal of maximizing the total amount of unloading tasks and minimizing energy consumption, and decompose it into four sub-problems: intelligent reflector phase control, UAV trajectory optimization, user association and computing resource allocation;

[0011] Subproblem solving: Based on the phase alignment principle, an optimal closed-loop phase shift control strategy based on the UAV and user positions is derived to enhance channel gain; the UAV trajectory optimization subproblem is modeled as a partially observable Markov decision process and solved using an improved multi-agent near-end policy optimization algorithm that incorporates action masking and a hybrid expert model; a channel-based matching mechanism is used to solve the user association problem, and a convex optimization method is used to solve the computational resource allocation problem.

[0012] Construct a joint optimization problem with the objectives of maximizing the total number of unloaded tasks and minimizing energy consumption, specifically:

[0013] make Represents the phase shift variable. Indicates the coordinates of the drone. Representing user-related variables, and This represents the variables used for calculating resource allocation;

[0014] The joint optimization problem is formulated as follows:

[0015]

[0016] in, Indicates in time slot Smart reflective surface The reflection phase shift matrix, Indicates drone coordinates Indicates in time slot drones For users The number of CPU cycles allocated; User association indicator; Indicates a time slot; , They represent drones exist Minimum and maximum flight range of the axis, , They represent drones exist Minimum and maximum flight range of the axis; This represents the upper limit of the total computational allocation. For computational tasks Execution latency, Indicates the size of the task data. Indicates in time slot Time user Data transmission rate This represents the maximum latency tolerance. For phase shift variables; Indicates the speed of the drone; This indicates the time slot length, i.e., the duration of each time step;

[0017] Constraint C1 ensures that each drone maintains a fixed distance traveled in each time slot; constraint C2 defines the drone's flight range; user-related variables. The value is defined by constraint C3; constraint C4 stipulates that each user can offload a task to at most one drone; constraint C5 restricts the drones. The total computing resources allocated to its service users should be ensured to remain within their limits. Within this constraint; constraint C6 stipulates that the latency of the served user must meet its maximum latency tolerance; constraint C7 defines the phase shift variable. The range of values; all tasks are assumed to be completed within one time slot; for tasks with large amounts of data and high computational requirements, they can be preprocessed and executed in multiple time slots.

[0018] Furthermore, the multi-intelligent reflector-assisted UAV computing network scenario includes... Individual users A smart reflective surface and A drone; via drones and The collaboration of several intelligent reflective surfaces dynamically provides computing services to users within the service area; these intelligent reflective surfaces are deployed on the surfaces of high-rise buildings, and are... It consists of several reflective elements;

[0019] The transfer modes used for task unloading include the following two modes:

[0020] The first transmission mode directly offloads the computing task to the drone, suitable for users. Located in drone The situation within the coverage area; its channel is a direct channel;

[0021] The second transmission mode can utilize intelligent reflective surfaces. Transmit computing tasks to the drone Execution, applicable to users Situations outside the service range of any drone; its channels include the direct channel and the channel from the user. via smart reflective surface To drones A virtual line-of-sight channel; the virtual line-of-sight channel includes a virtual line-of-sight channel from the user To intelligent reflective surface The channel, and from the smart reflector To drones The channel; the direct channel and the virtual line-of-sight channel are combined to form a composite channel; the composite channel can adapt to the user's state, as follows:

[0022]

[0023] Able to adjust The value is automatically adapted to the two transmission modes; orthogonal frequency division multiple access technology is used to facilitate task offloading;

[0024] The energy consumption includes computational energy consumption and flight energy consumption.

[0025] Furthermore, the subproblem of phase control for the intelligent reflector is:

[0026] Based on the phase alignment principle and with the goal of maximizing channel gain, the optimal closed-loop phase shift control strategy formula is derived as follows:

[0027]

[0028] The sub-problem of optimizing the drone trajectory is:

[0029]

[0030] The user association sub-problem is:

[0031]

[0032] The computational resource allocation subproblem is:

[0033]

[0034] Solving the subproblem includes the following steps:

[0035] For the phase control subproblem of intelligent reflector, an optimal closed-loop phase shift control strategy derived based on the phase alignment principle is used to obtain phase shift control that is coupled with the UAV coordinates in real time.

[0036] To address the subproblems of UAV trajectory optimization, user association, and computational resource allocation, a hierarchical decoupling strategy is proposed. A low-complexity channel-based matching mechanism and a convex optimization method are used to solve the user association and computational resource allocation subproblems, respectively. Specifically: the UAV trajectory optimization subproblem is modeled as a partially observable Markov decision process, defining its state, observations, actions, and reward function; an improved MAPPO algorithm incorporating a hybrid expert model and action mask is employed to make UAV trajectory design decisions based on current observation information, iteratively solving for the optimal UAV flight trajectory; the low-complexity channel-based matching mechanism is used to solve the user association subproblem, prioritizing association of users to UAVs with the highest channel gain; the convexity of the resource allocation problem is proven under given association conditions, and the interior-point method is used to solve the computational resource allocation subproblem.

[0037] Furthermore, an optimal closed-loop phase shift control strategy derived based on the phase alignment principle is used to obtain phase shift control coupled with UAV coordinates in real time. Specifically, based on the phase alignment principle, a real-time phase control strategy based on UAV coordinates is adopted, and its optimality in the considered scenario is established through optimal phase control during the mission transmission phase. The optimal phase control during the mission transmission phase is expressed as:

[0038]

[0039] in, Indicates the carrier frequency. , These represent the row spacing and column spacing of the reflective element, respectively; Indicates from intelligent reflective surface To drones The vertical departure angle, Indicates from user To intelligent reflective surface The vertical angle of arrival; Indicates from intelligent reflective surface To drones Horizontal departure angle, Indicates from user To intelligent reflective surface Horizontal angle of arrival; This indicates the row index of the reflecting element in the array. Indicates the column index of the reflecting element in the array; Represents the speed of light;

[0040] After applying the optimal phase control strategy, problem P1 is obtained, which is expressed as follows:

[0041]

[0042] Furthermore, the UAV trajectory optimization subproblem is modeled as a partially observable Markov decision process. An improved MAPPO algorithm incorporating a hybrid expert model and action mask is used to make UAV trajectory design decisions based on current observation information, including:

[0043] The intelligent agent shares the location information of all drones;

[0044] Model the problem P1 as a tuple. Partially observable Markov decision processes; among which, and These represent the state transition probability and the discount factor, respectively; in each time slot, the agent determines the state transition probability based on its observations. Select Action Interact with the environment; then, perform actions. Current state Transferred to And use a reward function to calculate the single-step reward. ;

[0045] Intelligent agents based on hybrid expert models include Actor networks and Critic networks;

[0046] An action mask is introduced at the output layer of the Actor network to remove unselectable actions. The steps include: the action mask is generated based on the drone's position, with entries corresponding to selectable actions set to 1 and entries corresponding to unselectable actions set to 0; the action logits are masked by the action mask, and then the probability of invalid actions is effectively forced to approach 0 through the softmax function.

[0047] The paradigm of "centralized training and decentralized execution" is adopted: global state is utilized during the training phase; during the agent execution phase, a hybrid expert model structure containing multiple expert networks and a gating network is introduced into the Actor network, and actions are generated based on its local observations through a gating network and multiple experts.

[0048] Actor networks based on hybrid experts can be updated using a multi-agent proximal policy optimization mechanism: Let Represents intelligent agents The current strategy, the parameters of the Critic network are represented as follows: For Actor networks, the input is local observations, and the output is local actions; Critic networks take the global state and the local observations corresponding to all agents as input, and output... State value;

[0049] The Actor network is trained using the following loss function:

[0050]

[0051] in This represents the clipping function. This represents the probability ratio between the current strategy and the old strategy; It represents the mathematical expectation or expected value, and is an average performance index calculated from the sampled trajectory data; This represents the advantage function under the old strategy, used to measure the additional benefit of taking a certain action in the current state relative to the average level; This represents the pruning parameter, which limits the magnitude of updates between the old and new policies to prevent excessive policy updates from causing training instability.

[0052] The generalized advantage estimator is calculated as follows:

[0053]

[0054] in, and Let represent the state-value function and the discount factor of the generalized advantage estimator, respectively; let This represents the state value assessment of the Critic network. and These represent the sampled training data and the cumulative discount reward, respectively. This represents the time-domain step index, used for weighted summation of rewards over multiple future time steps;

[0055] The Critic network is trained using the following loss function:

[0056] .

[0057] Furthermore, the tuple is defined as follows:

[0058] State: To reduce redundancy in neural network inputs, the state only contains environmental parameters that change over time; the global state contains the position coordinates of all UAVs.

[0059] Observation: To reduce coverage overlap during training, each drone observes its own position and the positions of other drones; Agent The observations were made by Give;

[0060] Actions: The actions of each agent Defined as drone Move 5 meters in one of the four directions; It is a joint action of all intelligent agents;

[0061] Rewards: Single-step rewards are calculated as follows:

[0062]

[0063] in, This indicates a penalty for overlapping drone coverage. For drones In the time slot Total energy consumption.

[0064] Furthermore, the agent divides the neural network into multiple sub-networks, each representing an expert, and controls the activation of the experts and outputs mixed features through a gating network. The gating network selects and activates experts based on the observation results, and the activated experts independently generate feature representations based on the observations. Then, the gating network applies a softmax layer to generate weight matrices for the experts, which are subsequently used to aggregate the experts' outputs. The output of the gating network is represented as follows:

[0065]

[0066] in, This indicates that the input features correspond to the input features. The output characteristics, Experts In the input Features generated in time; This indicates that the gating network targets the input. Assigned to experts The weights are expressed as:

[0067]

[0068] in, Indicating expert in gating network The relevant learnable weights.

[0069] Furthermore, a low-complexity matching mechanism based on the channel and a convex optimization method are used to solve for user association and computational resource allocation, respectively, including the following steps:

[0070] Given a fixed drone location, a low-complexity matching algorithm based on the channel is used to calculate the channel gain from the user to each drone, prioritizing the association of the user with the drone with the best channel conditions to reduce transmission latency.

[0071] Based on user associations, the subproblem of computing resource allocation is expressed as:

[0072]

[0073] Since the numerator of problem P2 is determined by the UAV coordinates and user-related decisions, it is transformed into the following equivalent problem:

[0074]

[0075] Problem P3 is a convex optimization problem, and the optimal solution is obtained using the interior point method.

[0076] A computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the high-efficiency user association and trajectory optimization method for a multi-intelligent reflector-assisted unmanned aerial vehicle network.

[0077] The present invention has the following beneficial effects:

[0078] To effectively address the issues of limited coverage and low energy efficiency of unmanned aerial vehicles (UAVs), this invention proposes a multi-intelligent reflector-assisted UAV network architecture and a joint trajectory and resource optimization scheme. First, by introducing multiple intelligent reflectors and employing a phase-alignment-based phase-shift control strategy, this invention constructs a high-quality virtual line-of-sight link, effectively overcoming the UAV coverage blind spot problem and significantly increasing the total task offload without increasing UAV deployment costs. Second, addressing the complex problem of mixed-integer nonlinear programming involving deep coupling of user association, trajectory control, and resource allocation, a hierarchical decoupling strategy is proposed. A low-complexity matching mechanism based on the channel and a convex optimization method are used to solve user association and computational resource allocation respectively, effectively reducing computational complexity and achieving precise supply of physical layer resources. Simultaneously, to address the issues of policy homogenization and boundary constraints in multi-UAV cooperation, this invention models trajectory optimization as a partially observable Markov decision process and uses an improved multi-agent near-end policy optimization algorithm for solution. By introducing a hybrid expert model to enhance policy diversity and using an action mask mechanism to enforce constraints on flight boundaries, invalid exploration is avoided, significantly improving the algorithm's convergence speed and training stability in dynamic environments. Simulation results show that the proposed scheme can consistently achieve optimal overall system performance under different user densities and the number of reflective elements. It effectively reduces energy consumption while increasing task offloading, and is significantly better than existing benchmark algorithms. Attached Figure Description

[0079] Figure 1 This is a schematic diagram of the unmanned aerial vehicle (UAV) network system model of the present invention;

[0080] Figure 2 A schematic diagram illustrating the principle of introducing an action masking mechanism in an Actor network;

[0081] Figure 3 This is a diagram illustrating the overall framework of the joint optimization scheme proposed in this invention.

[0082] Figure 4 A comparison of the convergence of the JTUC scheme designed for this invention with four benchmark schemes;

[0083] Figure 5 The graph shows the changing trends of total system reward, total unloaded task volume, and total energy consumption as the number of intelligent reflective surfaces increases.

[0084] Figure 6 This diagram illustrates the system's total reward, total unloaded task volume, and total energy consumption as the number of intelligent reflective surface reflective units changes.

[0085] Figure 7 A diagram illustrating the impact of user scale on various system metrics;

[0086] Figure 8Performance comparison chart for different computational intensities required for different tasks;

[0087] Figure 9 This diagram illustrates the impact of the number of drones on various performance indicators of the system. Detailed Implementation

[0088] The present invention will be further described below with reference to the accompanying drawings. The embodiments of the present invention are only used to illustrate the present invention and not to limit the present invention. Various substitutions and modifications made based on ordinary technical knowledge and common practices in the art without departing from the technical concept of the present invention should be included within the scope of the present invention.

[0089] To address the aforementioned problems, this invention proposes a high-efficiency user association and trajectory optimization method for multi-intelligent reflector-assisted unmanned aerial vehicle (UAV) networks, comprising the following steps:

[0090] S1: Scenario and Model Construction: Construct a UAV computing network scenario assisted by multiple intelligent reflectors, and establish computing models, communication models and energy consumption models;

[0091] S2: Construct an optimization problem and decompose it into sub-problems: Construct a joint optimization problem with the goal of maximizing the total amount of unloading tasks and minimizing energy consumption, and decompose it into four sub-problems: intelligent reflector phase control, UAV trajectory optimization, user association and computing resource allocation;

[0092] S3: Subproblem Solving: Based on the phase alignment principle, an optimal closed-loop phase shift control strategy based on the UAV and user positions is derived to enhance channel gain; the UAV trajectory optimization subproblem is modeled as a partially observable Markov decision process, and an improved multi-agent near-end policy optimization algorithm incorporating action masking and hybrid expert models is used to solve it; a channel-based matching mechanism is used to solve the user association problem, and a convex optimization method is used to solve the computational resource allocation problem.

[0093] The specific steps are as follows:

[0094] S1: Constructing scenarios and models.

[0095] Construct a UAV network scenario assisted by multiple intelligent reflectors, and establish a computing model, a communication model, and an energy consumption model.

[0096] This invention designs a drone network with multiple intelligent reflective surfaces to provide dynamic coverage enhancement. The network includes... Individual users A smart reflective surface and A drone, that is , ,as well as .pass drones and The collaboration of multiple intelligent reflective surfaces can dynamically provide computing services to users within the service area. Since drone coverage is limited, intelligent reflective surfaces can allow users outside the drone's coverage area to offload tasks. Figure 1 As shown, the drone network system includes multiple drones acting as aerial edge servers, users distributed on the ground, and multiple smart reflective surfaces deployed on building surfaces, and demonstrates two communication modes: direct transmission and smart reflective surface-assisted reflective transmission.

[0097] This invention takes into account the limited coverage area of ​​each drone, allowing users within its service area to offload computationally intensive tasks. Intelligent reflective surfaces are deployed on the surfaces of high-rise buildings, by... It consists of several reflective elements. By adjusting the phase shift of these elements, the smart reflective surface allows users to offload computational tasks to the drone.

[0098] definition (in , , , ) is the user association indicator, where Indicates user In the time slot Through intelligent reflective surface Offload computing tasks to drones ,otherwise . Indicates user The computational tasks were transmitted to the drone without the aid of a smart reflector. .

[0099] user computational tasks It can be described by three main characteristics, namely , and These represent the task data size, the required number of CPU cycles, and the latency threshold, respectively. Due to the user's lack of computing resources, all tasks were offloaded to the drone. (Computational tasks) The execution latency is expressed as:

[0100]

[0101] in, Indicates in time slot drones For users The number of CPU cycles allocated.

[0102] Because drones have limited coverage, users may be outside the service area and unable to access computing services. Therefore, this invention considers two transmission modes for task offloading. The first transmission mode is suitable for users... Located in drone Within the coverage area, the computational tasks are directly offloaded to the drone. users With located drones The Euclidean distance between them can be expressed as Then, in the time slot ,user With drones The direct channel between them is represented as:

[0103]

[0104] in This represents the path loss at a reference distance of 1m. This represents the path loss index between the user and the drone.

[0105] The second transmission mode appears to users When outside the service range of any drone, this utilizes a smart reflective surface. Transmit computing tasks to the drone Execution. Specifically, the channel in this mode consists of two parts: the aforementioned direct channel, and the channel from the user... via smart reflective surface To drones The virtual line-of-sight channel, denoted as The virtual line-of-sight channel consists of two parts: from the user... To intelligent reflective surface The channel, and from the smart reflector To drones The channel. From the user To intelligent reflective surface The channel is represented as:

[0106]

[0107] in, and Representing users respectively With intelligent reflective surface The path loss exponent and Rice parameter of the channel between them. and These represent the carrier frequency and the speed of light, respectively. The row spacing and column spacing of the reflective element are represented by... and express. This represents the non-line-of-sight channel components. Since line-of-sight components are dominant, user behavior is not considered subsequently. With drones The non-line-of-sight channel component between the channels. , and They represent from the user To intelligent reflective surface The distance, vertical angle of arrival, and horizontal angle of arrival.

[0108] Smart reflective surface The coordinates are represented as ,and , ,as well as .

[0109] Smart reflective surface With drones The channel between them is modeled as a line-of-sight channel, and its corresponding channel vector is represented as follows:

[0110]

[0111] in , and These represent the intelligent reflective surface. To drones The distance, vertical departure angle, and horizontal departure angle, and , as well as .

[0112] make To represent the amplitude parameter, from the user via smart reflective surface To drones The virtual line-of-sight channel can be represented as follows:

[0113]

[0114] in , Indicates in time slot Smart reflective surface The reflection phase shift matrix is ​​used to assist users. Offload its computing tasks to the drone .

[0115] To enhance user experience With drones The communication link between them combines direct channels and virtual line-of-sight channels to form a composite channel. This is in conjunction with the user's communication link. The relevant composite channels can adapt to the user's state, as shown below:

[0116]

[0117] It is worth noting that adjustments can be made. The value is used to automatically adapt to the two transmission modes. To reduce interference between users, this invention employs orthogonal frequency division multiple access (OFDMA) technology to facilitate task offloading, and the bandwidth allocated to each user is expressed as... In the time slot ,user The data transmission rate is calculated as follows:

[0118]

[0119] in Indicates noise power. Indicates the transmission power.

[0120] Given the limited energy capacity of UAVs, this paper primarily focuses on their energy consumption, including computational energy consumption and flight energy consumption. Computational energy consumption can be expressed as... .in Indicates a switched capacitor. (UAV) In the time slot The flight energy consumption is expressed as ,in This indicates the drone's flight power. Therefore, the drone In the time slot Total energy consumption is .

[0121] The system's stability is affected by the drone's lifespan. Therefore, this invention proposes an optimization problem aimed at maximizing the total amount of unloading tasks while minimizing energy consumption. The optimization objective is defined as:

[0122]

[0123] in This represents the penalty parameter. The penalty term is introduced to prevent the optimization process from converging in a direction that reduces system throughput.

[0124] S2: Construct the optimization problem and decompose it into subproblems.

[0125] We construct a joint optimization problem with the goal of maximizing the total number of unloaded tasks and minimizing energy consumption, and decompose it into subproblems of phase control, trajectory optimization, user association, and resource allocation.

[0126] make Represents the phase shift variable. Indicates the coordinates of the drone. Representing user-related variables, and This represents the variables used for calculating resource allocation.

[0127] Therefore, the optimization problem is described as follows:

[0128]

[0129] in, Indicates in time slot Smart reflective surface The reflection phase shift matrix, Indicates drone coordinates Indicates in time slot drones For users The number of CPU cycles allocated; User association indicator; Indicates a time slot; , They represent drones exist Minimum and maximum flight range of the axis, , They represent drones exist Minimum and maximum flight range of the axis; This represents the upper limit of the total computational allocation. For computational tasks Execution latency, Indicates the size of the task data. Indicates in time slot Time user Data transmission rate This represents the maximum latency tolerance. For phase shift variables; This indicates the time slot length, which is the duration of each time step.

[0130] Constraint C1 ensures that each drone maintains a fixed distance of movement in each time slot, where This represents the drone's speed. Constraint C2 defines the drone's flight range. User-related variables. The value is defined by constraint C3, and constraint C4 stipulates that each user can offload a task to at most one drone. Constraint C5 restricts the drones. The total computing resources allocated to its service users should be ensured to remain within their limits. Within this range, constraint C6 stipulates that the latency of the served user must meet its maximum latency tolerance. Constraint C7 defines the phase shift variable. The range of values ​​is defined. All tasks are assumed to be completed within a single time slot. For tasks with large amounts of data and high computational requirements, preprocessing can be performed and the data executed across multiple time slots.

[0131] Problem P0 is a mixed-integer nonlinear programming problem. The conflicting objectives of maximizing the total number of unloaded tasks and minimizing the total energy consumption further increase the complexity of solving problem P0.

[0132] S3: Solving subproblems.

[0133] In the process of solving the sub-problems of this invention, phase shift control depends on the current position of the UAV, the UAV trajectory determines the user association, and the result of the user association determines the resource allocation.

[0134] For the phase control subproblem of intelligent reflector, an optimal closed-loop phase shift control strategy derived based on the phase alignment principle is used to obtain phase shift control that is coupled with the UAV coordinates in real time.

[0135] For the subproblems of UAV trajectory optimization, user association, and computational resource allocation, a hierarchical decoupling strategy is proposed. A low-complexity matching mechanism based on the channel and a convex optimization method are used to solve the user association and computational resource allocation problems respectively. Specifically:

[0136] The UAV trajectory optimization subproblem is modeled as a partially observable Markov decision process, defining its state, observations, actions, and reward function. An improved MAPPO algorithm incorporating a hybrid expert model and action mask is employed to make UAV trajectory design decisions based on current observation information, iteratively solving for the optimal UAV flight trajectory. A low-complexity matching mechanism based on the channel is used to solve the user association subproblem, prioritizing association of users to UAVs with the highest channel gain. Given the association, the convexity of the resource allocation problem is proven, and the interior-point method is used to solve the computational resource allocation subproblem.

[0137] Figure 3 This is a diagram illustrating the overall framework of the joint optimization scheme proposed in this invention. The framework comprises three main parts: a centralized training module, a distributed execution module, and a policy network based on a hybrid expert model. Through iterative collaboration, each module dynamically activates different experts using a gating network, achieving joint optimization of resources and trajectories under multi-agent collaboration.

[0138] S31: Solve the phase control subproblem of the intelligent reflector.

[0139] Based on the phase alignment principle, an optimal closed-loop phase shift control strategy based on the UAV and user locations is derived to enhance channel gain.

[0140] For the intelligent reflector phase shift optimization subproblem, since the phase shift variable is highly coupled with the trajectory, this invention does not directly generate the phase shift through reinforcement learning, but derives the optimal solution based on the phase alignment principle. To maximize the channel capacity between the user and the UAV, the optimal intelligent reflector phase shift must ensure that the reflected signal and the direct signal are superimposed in phase at the receiving end. Specifically:

[0141] A real-time phase control strategy based on UAV coordinates is adopted, and its optimality in the considered scenario is established through optimal phase control during the mission transmission phase. The optimal phase control during the mission transmission phase can be expressed as:

[0142]

[0143] in, Indicates the carrier frequency. , These represent the row spacing and column spacing of the reflective element, respectively; Indicates from intelligent reflective surface To drones The vertical departure angle, Indicates from user To intelligent reflective surface The vertical angle of arrival; Indicates from intelligent reflective surface To drones Horizontal departure angle, Indicates from user To intelligent reflective surface Horizontal angle of arrival; This indicates the row index of the reflecting element in the array. Indicates the column index of the reflecting element in the array; Represents the speed of light;

[0144] After applying the optimal phase control strategy, problem P1 can be obtained, which is expressed as follows:

[0145]

[0146] S32: Solve for UAV trajectory optimization.

[0147] The trajectory design is modeled as a partially observable Markov decision process, and an improved multi-agent proximal policy optimization algorithm incorporating action masking and hybrid expert models is used to solve it.

[0148] An improved multi-agent proximal policy optimization algorithm is employed to optimize UAV trajectories. The global state includes the positions of all UAVs; local observations include the positions of the UAV itself and other UAVs. To address UAV flight boundary constraints, an action mask is introduced at the Actor network output layer. Based on the UAV's current position, the probability of invalid actions leading to out-of-bounds flight is forcibly set to 0, replacing the traditional penalty term and improving training efficiency. A hybrid expert model structure is introduced into the Actor network, comprising multiple expert networks and a gating network. The gating network dynamically activates different expert networks to generate strategies based on observations, preventing policy homogenization caused by multi-agent parameter sharing and enhancing exploration capabilities. A reward function is designed to maximize system utility and penalize overlapping coverage. Specifically:

[0149] In the scenarios under consideration, Each agent makes independent decisions based on its own observations. To mitigate coverage overlap, the agents share the position information of all UAVs. The trajectory optimization subproblem (i.e., the formula obtained after solving the phase control subproblem) P1 can be modeled as a tuple. Partially observable Markov decision processes, in which and These represent the state transition probability and the discount factor, respectively. In each time slot, the agent, based on its observations... Select Action Interact with the environment. Then, perform actions. Current state Transferred to And use a reward function to calculate the single-step reward. .

[0150] Specifically, a tuple is defined as follows:

[0151] (1) State: In order to reduce the redundancy of neural network input, the state only contains environmental parameters that change over time. Therefore, the global state contains the position coordinates of all UAVs.

[0152] (2) Observation: To reduce coverage overlap during training, each UAV observes its own position and the positions of other UAVs. (Agent) The observations were made by Provided.

[0153] (3) Actions: The actions of each agent Defined as drone Move 5 meters in one of the four directions. It is a joint action of all intelligent agents.

[0154] To address the boundary constraint C2, this invention introduces an action mask into the Actor network.

[0155] Figure 2 This diagram illustrates the principle of introducing an action masking mechanism into the Actor network. It demonstrates how action masking works by generating a mask based on the drone's real-time position and applying it to the network's output layer. This forcibly blocks illegal actions that fly out of bounds, effectively avoiding ineffective exploration and improving training efficiency.

[0156] Specifically, at each time step, the UAV's observations are input into the network, and the output layer outputs corresponding action logits. When the UAV is located at the service area boundary, the corresponding agent is restricted from choosing actions that cross the boundary. Traditionally, a penalty term is applied to encourage the agent to learn the boundary. This method may lead to a large number of unnecessary actions, resulting in low learning efficiency or hindering convergence. Therefore, this invention introduces an action mask to remove unselectable actions. The action mask is generated based on the UAV's position, with entries corresponding to selectable actions set to 1 and entries corresponding to unselectable actions set to 0. Then, the action logits are masked by the action mask, and subsequently, a softmax function is used to effectively force the probability of invalid actions to approach 0.

[0157] (4) Reward: The reward function is designed to guide the agent to take actions consistent with the optimization direction of the objective function of problem P0. The single-step reward is calculated as follows:

[0158]

[0159] in This indicates a penalty for overlapping drone coverage. For drones In the time slot The total energy consumption. It is worth noting that drone coverage overlap is determined by the joint flight decisions of all drones, which makes it impractical to use action masks at each agent to prevent overlap.

[0160] In single-agent environments, state transitions are typically deterministic for actions, meaning the action chosen by the agent in a given state determines the next state. However, in multi-agent environments, the actions of all agents collectively influence the environment, leading to non-stationarity. Non-stationarity hinders the convergence of independently trained agents. To address this issue, this paper adopts a "centralized training, distributed execution" paradigm. Figure 3 illustrates that the JTUC scheme utilizes global states during the training phase, while during agent execution, actions are generated based on their local observations through a gating network and multiple experts. The JTUC scheme is a joint trajectory design, user association, and computational resource allocation optimization scheme. Specifically, J stands for Joint, for joint optimization; T stands for Trajectory, for UAV flight trajectory design; U stands for User Association, for user association; and C stands for Computation Resource Allocation, for computational resource allocation.

[0161] Unlike traditional agent networks, hybrid expert models divide the neural network into several subnetworks, each representing an expert. A gating network controls the activation of these experts and their outputs, resulting in a mixture of features. The gating network selects and activates experts based on observations. These experts independently generate feature representations based on the observations. Then, the gating network applies a softmax layer to generate weight matrices for the experts. This weight matrix is ​​subsequently used to aggregate the experts' outputs. The output of the gating network can be represented as:

[0162]

[0163] in This indicates that the input features correspond to the input features. The output characteristics. Experts In the input Features generated at that time. This indicates that the gating network targets the input. Assigned to experts The weights can be expressed as:

[0164]

[0165] in Indicating expert in gating network The relevant learnable weights.

[0166] Since all agents in the system of this invention are isomorphic, sharing network weights and bias parameters can accelerate the convergence of the learning process. However, parameter sharing may lead to homogenization of behavior among agents, thus limiting their potential to search the action space. This limits the diversity of policies within a multi-agent system, thereby affecting system performance. Therefore, this invention employs a hybrid expert model in the Actor network. The hybrid expert model utilizes a gating network to activate different experts based on different local observations, resulting in policies that enhance the diversity of policies among agents.

[0167] Actor networks based on hybrid experts can be updated using a multi-agent proximal policy optimization mechanism. Represents intelligent agents The current strategy. The parameters of the Critic network are represented as follows: For an Actor network, the input is local observations, and the output is local actions. A Critic network takes the global state and the local observations corresponding to all agents as input, and outputs... The state value. The Actor network is trained using the following loss function:

[0168]

[0169] in This represents the clipping function. This represents the probability ratio between the current policy and the old policy. It represents the mathematical expectation or expected value, and is an average performance index calculated from the sampled trajectory data; This represents the advantage function under the old strategy, used to measure the additional benefit of taking a certain action in the current state relative to the average level; This represents the pruning parameter, which limits the magnitude of updates between the old and new policies to prevent excessive policy updates from causing training instability. The pruning function is used to limit the magnitude of policy updates to ensure training stability.

[0170] Furthermore, the generalized advantage estimator is calculated as follows:

[0171]

[0172] in and Let represent the state value function and the discount factor of the generalized advantage estimator, respectively. This represents the state value assessment of the Critic network. and These represent the sampled training data and the cumulative discount reward, respectively. This represents the time-domain step index, used for weighted summation of rewards over multiple future time steps.

[0173] The Critic network is trained using the following loss function:

[0174]

[0175] Algorithm 1 demonstrates the pseudocode for the JTUC scheme training process.

[0176]

[0177] S33: Solve the subproblem of user association and computational resource allocation.

[0178] A channel-based matching mechanism is used to solve the user association problem, and a convex optimization method is used to solve the computational resource allocation problem.

[0179] Given a fixed drone location, a low-complexity channel-based matching algorithm is used to calculate the channel gain from the user to each drone, prioritizing the association of the user to the drone with the best channel conditions to reduce transmission latency. With the association determined, the original problem is transformed into a convex optimization problem. This invention proves that, given a trajectory and association, the resource allocation problem that minimizes energy consumption is convex, and directly solves for the optimal CPU frequency allocation using an interior-point method. Specifically:

[0180] In each time slot, once the UAV's location is determined, it is necessary to optimize local service association and resource allocation. Representing association variables as actions leads to an excessively large action space. Reducing transmission latency can lower the UAV's computational energy consumption while meeting the user's maximum latency tolerance. In particular, reducing transmission latency allows for longer computational latency without violating latency constraints. This invention designs a channel-based matching method to associate users with UAVs. By selecting users with stronger channel gain, task transmission latency can be effectively reduced. Specifically, users within the direct coverage area of ​​the UAV directly offload tasks to the UAV; while users outside the direct coverage area but with strong composite channel gain via the IRS reflection link are also associated with the UAV to provide services.

[0181] Algorithm 2 demonstrates the user association process.

[0182]

[0183] Given the user associations determined by Algorithm 1, the computational resource allocation subproblem can be formulated as follows:

[0184]

[0185] Since the numerator of problem P2 is determined by the UAV coordinates and user-related decisions, the problem can be transformed into the following equivalent problem:

[0186]

[0187] Problem P3 is a convex optimization problem, and the optimal solution can be obtained using the interior point method.

[0188] The applicant conducted numerical simulation experiments comparing the proposed solution with four benchmark solutions.

[0189] MTO (Multi-Agent Dual-Delay Deep Deterministic Policy Gradient Scheme): This scheme is an optimization framework based on the multi-agent TD3 algorithm, which maximizes system energy efficiency by jointly optimizing UAV trajectory, phase shift, and task allocation. The UAV trajectory is optimized by the MTO scheme, while the user association and computational resource allocation strategies retain the settings from the JTUC scheme, and constraint C2 is introduced as a penalty term into the reward function.

[0190] MBS (MAPPO-based Joint Optimization Scheme): This is a comparative scheme based on the MAPPO algorithm, aiming to minimize total energy consumption by jointly optimizing task offloading ratio, computational resource allocation, and UAV trajectory. The scheme uses MBS logic to optimize the UAV trajectory, while its remaining solutions for trajectory control and computational resource allocation follow the same path as the JTUC scheme.

[0191] RDA (Random User Association Scheme): In this scheme, user association decisions are performed randomly. Apart from this, the UAV flight trajectory and computational resource allocation schemes employ the same solution strategy as the JTUC scheme. This benchmark primarily serves to verify the contribution of the proposed channel-based matching association algorithm to improving system performance.

[0192] WIS (Smart Surface-Free Solution): This is an IRS-free solution based on MAPPO, focusing on jointly optimizing offloading decisions, resource allocation, and encryption configuration to minimize latency, energy consumption, and privacy risks. Under this solution, the drone network operates without an IRS, and each drone can only provide services to users within its direct coverage area; the rest of the solution is the same as the JTUC solution.

[0193] Figure 4 The convergence results of the JTUC scheme designed in this invention and four benchmark schemes are shown. It can be seen that, thanks to the introduction of action masking and hybrid expert models, the JTUC scheme outperforms the other comparative schemes in both convergence speed and final reward value, exhibiting stronger stability.

[0194] Figure 5 The data shows the trends in total system reward, total offloaded tasks, and total energy consumption as the number of intelligent reflectors increases. It can be seen that increasing the number of intelligent reflectors significantly expands the drone's coverage area, thereby greatly improving the system's overall task offload capability.

[0195] Figure 6 This study reflects the performance of the system's total reward, total offloaded tasks, and total energy consumption as the number of intelligent reflector units changes. The results show that increasing the number of reflector units enhances channel quality, thereby increasing system reward while effectively reducing total system energy consumption by decreasing transmission latency.

[0196] Figure 7 The impact of user scale on various system metrics was analyzed. As the number of users increases, both the total number of offload tasks and total energy consumption show an upward trend. The JTUC solution maintains superior performance compared to the comparison solutions under different user numbers, especially in high-load scenarios, demonstrating good scalability.

[0197] Figure 8 The performance comparison is shown under different computational intensities required for different tasks. It can be observed that the total energy consumption of all schemes increases with the increase of the required computational intensity, but the JTUC scheme maintains the highest system utility due to its efficient resource allocation strategy.

[0198] Figure 9 This demonstrates the impact of the number of drones on various system performance indicators. As the number of drones increases, the JTUC solution can still effectively coordinate multiple drone trajectories, avoid strategy homogenization, and maintain high system benefits while improving coverage, outperforming other benchmark solutions.

[0199] Through numerical simulation experiments, this invention verifies the effectiveness and robustness of the JTUC scheme in increasing task offloading and reducing energy consumption.

[0200] A computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the high-efficiency user association and trajectory optimization method for a multi-intelligent reflector-assisted unmanned aerial vehicle network.

[0201] This invention constructs a high-quality virtual line-of-sight link by introducing multiple intelligent reflectors and employing a phase-alignment-based phase-shift control strategy, enabling effective service to users in coverage blind spots. Addressing the complex problem of mixed-integer nonlinear programming involving deep coupling of user association, trajectory control, and resource allocation, a hierarchical decoupling strategy is proposed. A low-complexity matching mechanism based on the channel and a convex optimization method are used to solve user association and computational resource allocation respectively, effectively reducing computational complexity. Furthermore, addressing the issues of policy homogenization and boundary constraints in multi-UAV cooperation, this invention models trajectory optimization as a partially observable Markov decision process and employs an improved multi-agent near-end policy optimization algorithm for solution. By introducing a hybrid expert model to enhance policy diversity and using an action masking mechanism to enforce constraints on flight boundaries, the system can autonomously achieve the optimal balance between energy efficiency and service quality in dynamic environments.

[0202] This invention increases the total task offload without increasing the deployment cost of UAVs, while reducing computational complexity and achieving precise supply of physical layer resources. It also avoids ineffective exploration, significantly improving the convergence speed and training stability of the algorithm in dynamic environments. Simulation results show that the proposed scheme can consistently achieve optimal overall system performance under different user densities and the number of reflective elements, effectively reducing energy consumption while increasing task offload, significantly outperforming existing benchmark algorithms.

[0203] The above description describes specific embodiments of the present invention and the technical principles employed. Any changes made in accordance with the concept of the present invention that do not exceed the spirit of the specification and drawings should still fall within the protection scope of the present invention.

Claims

1. A high-efficiency user association and trajectory optimization method for multi-intelligent reflector-assisted unmanned aerial vehicle (UAV) networks, characterized in that, Includes the following steps: Scenario and Model Construction: Construct a UAV computing network scenario assisted by multiple intelligent reflectors, and establish computing models, communication models and energy consumption models; Construct an optimization problem and decompose it into sub-problems: Construct a joint optimization problem with the goal of maximizing the total amount of unloading tasks and minimizing energy consumption, and decompose it into four sub-problems: intelligent reflector phase control, UAV trajectory optimization, user association and computing resource allocation; Subproblem solving: Based on the phase alignment principle, the optimal closed phase shift control strategy based on the UAV and user positions is derived to enhance channel gain; the UAV trajectory optimization subproblem is modeled as a partially observable Markov decision process, and an improved multi-agent near-end policy optimization algorithm incorporating action masking and hybrid expert models is used to solve it; A channel-based matching mechanism is used to solve the user association problem, and a convex optimization method is used to solve the computational resource allocation problem.

2. The high-efficiency user association and trajectory optimization method for multi-intelligent reflector-assisted UAV networks according to claim 1, characterized in that, Construct a joint optimization problem with the objectives of maximizing the total number of unloaded tasks and minimizing energy consumption, specifically: make Represents the phase shift variable. Indicates the coordinates of the drone. Representing user-related variables, and This represents the variables used for calculating resource allocation; The joint optimization problem is formulated as follows: in, Indicates in time slot Smart reflective surface The reflection phase shift matrix, Indicates drone coordinates Indicates in time slot drones For users The number of CPU cycles allocated; User association indicator; Indicates a time slot; , They represent drones exist Minimum and maximum flight range of the axis, , They represent drones exist Minimum and maximum flight range of the axis; This represents the upper limit of the total computational allocation. For computational tasks Execution latency, Indicates the size of the task data. Indicates in time slot Time user The data transmission rate This represents the maximum latency tolerance. For phase shift variables; Indicates the speed of the drone; This indicates the time slot length, i.e., the duration of each time step; Constraint C1 ensures that each drone maintains a fixed distance traveled in each time slot; constraint C2 defines the drone's flight range; user-related variables. The value is defined by constraint C3; constraint C4 stipulates that each user can offload a task to at most one drone; constraint C5 restricts the drones. The total computing resources allocated to its service users should be ensured to remain within their limits. Within this constraint; constraint C6 stipulates that the latency of the served user must meet its maximum latency tolerance; constraint C7 defines the phase shift variable. The range of values; all tasks are assumed to be completed within one time slot; for tasks with large amounts of data and high computational requirements, they can be preprocessed and executed in multiple time slots.

3. The high-efficiency user association and trajectory optimization method for multi-intelligent reflector-assisted UAV networks according to claim 1, characterized in that, The multi-intelligent reflector-assisted UAV computing network scenario includes Individual users A smart reflective surface and A drone; via drones and The collaboration of several intelligent reflective surfaces dynamically provides computing services to users within the service area; these intelligent reflective surfaces are deployed on the surfaces of high-rise buildings, and are... It consists of several reflective elements; The transfer modes used for task unloading include the following two modes: The first transmission mode directly offloads the computing task to the drone, suitable for users. Located in drone The situation within the coverage area; its channel is a direct channel; The second transmission mode can utilize intelligent reflective surfaces. Transmit computing tasks to the drone Execution, applicable to users Situations outside the service range of any drone; Its channels include the direct channel and the channel from the user. via smart reflective surface To drones A virtual line-of-sight channel; the virtual line-of-sight channel includes a virtual line-of-sight channel from the user To intelligent reflective surface The channel, and from the smart reflector To drones The channel; the direct channel and the virtual line-of-sight channel are combined to form a composite channel; the composite channel can adapt to the user's state, as follows: Able to adjust The value is automatically adapted to the two transmission modes; orthogonal frequency division multiple access technology is used to facilitate task offloading; The energy consumption includes computational energy consumption and flight energy consumption.

4. The high-efficiency user association and trajectory optimization method for multi-intelligent reflector-assisted UAV networks according to claim 2, characterized in that, The subproblem of phase control for the intelligent reflector is: Based on the phase alignment principle and with the goal of maximizing channel gain, the optimal closed-loop phase shift control strategy formula is derived as follows: The sub-problem of optimizing the drone trajectory is: The user association sub-problem is: The computational resource allocation subproblem is: Solving the subproblem includes the following steps: For the subproblem of phase control of intelligent reflector, an optimal closed-loop phase shift control strategy derived based on the phase alignment principle is used to obtain phase shift control that is coupled with UAV coordinates in real time. To address the subproblems of UAV trajectory optimization, user association, and computational resource allocation, a hierarchical decoupling strategy is proposed. A low-complexity channel-based matching mechanism and a convex optimization method are used to solve the user association and computational resource allocation subproblems, respectively. Specifically: the UAV trajectory optimization subproblem is modeled as a partially observable Markov decision process, defining its state, observations, actions, and reward function; an improved MAPPO algorithm incorporating a hybrid expert model and action mask is employed to make UAV trajectory design decisions based on current observation information, iteratively solving for the optimal UAV flight trajectory; the low-complexity channel-based matching mechanism is used to solve the user association subproblem, prioritizing association of users to UAVs with the highest channel gain; the convexity of the resource allocation problem is proven under given association conditions, and the interior-point method is used to solve the computational resource allocation subproblem.

5. The high-efficiency user association and trajectory optimization method for multi-intelligent reflector-assisted UAV networks according to claim 4, characterized in that, An optimal closed-loop phase shift control strategy derived based on the phase alignment principle is used to obtain phase shift control coupled with UAV coordinates in real time. Specifically, based on the phase alignment principle, a real-time phase control strategy based on UAV coordinates is adopted, and its optimality in the considered scenario is established through optimal phase control during the mission transmission phase. The optimal phase control during the mission transmission phase is expressed as follows: in, Indicates the carrier frequency. , These represent the row spacing and column spacing of the reflective element, respectively; Indicates from intelligent reflective surface To drones The vertical departure angle, Indicates from user To intelligent reflective surface The vertical angle of arrival; Indicates from intelligent reflective surface To drones Horizontal departure angle, Indicates from user To intelligent reflective surface Horizontal angle of arrival; This indicates the row index of the reflecting element in the array. Indicates the column index of the reflecting element in the array; Represents the speed of light; After applying the optimal phase control strategy, problem P1 is obtained, which is expressed as follows: 。 6. The high-efficiency user association and trajectory optimization method for multi-intelligent reflector-assisted UAV networks according to claim 5, characterized in that, The UAV trajectory optimization subproblem is modeled as a partially observable Markov decision process. An improved MAPPO algorithm incorporating a hybrid expert model and action masking is used to make UAV trajectory design decisions based on current observation information, including: The intelligent agent shares the location information of all drones; Model the problem P1 as a tuple. Partially observable Markov decision processes; among which, and These represent the state transition probability and the discount factor, respectively; in each time slot, the agent determines the state transition probability based on its observations. Select Action Interact with the environment; then, perform actions. Current state Transferred to And use a reward function to calculate the single-step reward. ; Intelligent agents based on hybrid expert models include Actor networks and Critic networks; An action mask is introduced at the output layer of the Actor network to remove unselectable actions. The steps include: the action mask is generated based on the drone's position, with entries corresponding to selectable actions set to 1 and entries corresponding to unselectable actions set to 0; the action logits are masked by the action mask, and then the probability of invalid actions is effectively forced to approach 0 through the softmax function. The paradigm of "centralized training and decentralized execution" is adopted: global state is utilized during the training phase; during the agent execution phase, a hybrid expert model structure containing multiple expert networks and a gating network is introduced into the Actor network, and actions are generated based on its local observations through a gating network and multiple experts. Actor networks based on hybrid experts can be updated using a multi-agent proximal policy optimization mechanism: Let Represents intelligent agents The current strategy, the parameters of the Critic network are represented as follows: For Actor networks, the input is local observations, and the output is local actions; Critic networks take the global state and the local observations corresponding to all agents as input, and output... State value; The Actor network is trained using the following loss function: in This represents the clipping function. This represents the probability ratio between the current strategy and the old strategy; It represents the mathematical expectation or expected value, and is an average performance index calculated from the sampled trajectory data; This represents the advantage function under the old strategy, used to measure the additional benefit of taking a certain action in the current state relative to the average level; This represents the pruning parameter, which limits the magnitude of updates between the old and new policies to prevent excessive policy updates from causing training instability. The generalized advantage estimator is calculated as follows: in, and Let represent the state-value function and the discount factor of the generalized advantage estimator, respectively; let This represents the state value assessment of the Critic network. and These represent the sampled training data and the cumulative discount reward, respectively. This represents the time-domain step index, used for weighted summation of rewards over multiple future time steps; The Critic network is trained using the following loss function: 。 7. The high-efficiency user association and trajectory optimization method for multi-intelligent reflector-assisted UAV networks according to claim 6, characterized in that, The tuple is defined as follows: State: To reduce redundancy in neural network inputs, the state only contains environmental parameters that change over time; the global state contains the position coordinates of all UAVs. Observation: To reduce coverage overlap during training, each drone observes its own position and the positions of other drones; Agent The observations were made by Give; Actions: The actions of each agent Defined as drone Move 5 meters in one of the four directions; It is a joint action of all intelligent agents; Rewards: Single-step rewards are calculated as follows: in, This indicates a penalty for overlapping drone coverage. For drones In the time slot Total energy consumption.

8. The high-efficiency user association and trajectory optimization method for multi-intelligent reflector-assisted UAV networks according to claim 6, characterized in that, The agent divides the neural network into multiple subnetworks, each representing an expert, and controls the activation of the experts and outputs mixed features through a gating network; the gating network selects and activates experts based on the observation results, and the activated experts independently generate feature representations based on the observations; Then, the gated network applies a softmax layer to generate a weight matrix for the experts, which is subsequently used to aggregate the experts' outputs; the output of the gated network is represented as: in, This indicates that the input features correspond to the input features. The output characteristics, Experts In the input Features generated in time; This indicates that the gating network targets the input. Assigned to experts The weights are expressed as: in, Indicating expert in gating network The relevant learnable weights.

9. The high-efficiency user association and trajectory optimization method for multi-intelligent reflector-assisted UAV networks according to claim 4, characterized in that, The user association and computational resource allocation are solved using a channel-based low-complexity matching mechanism and a convex optimization method, respectively, including the following steps: Given a fixed drone location, a low-complexity matching algorithm based on the channel is used to calculate the channel gain from the user to each drone, prioritizing the association of the user with the drone with the best channel conditions to reduce transmission latency. Based on user associations, the computational resource allocation subproblem can be expressed as: Since the numerator of problem P2 is determined by the UAV coordinates and user-related decisions, it is transformed into the following equivalent problem: Problem P3 is a convex optimization problem, and the optimal solution is obtained using the interior point method.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the high-efficiency user association and trajectory optimization method for multi-intelligent reflector-assisted unmanned aerial vehicle networks as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • User association and trajectory optimization method based on multi-intelligent reflector unmanned aerial vehicle communication

    CN117793752A