Dynamic resource allocation method of NCO-NOMA-assisted UAV collaborative IRS-MEC system
The UAV collaborative IRS-MEC system assisted by NCO-NOMA adopts hybrid optimization and deep reinforcement learning algorithms to dynamically allocate resources, solves the problems of user computing urgency and resource fairness, and minimizes the total system latency and improves resource utilization efficiency.
Patent Information
- Application Number
- CN202510642400.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-09-05
AI Technical Summary
In the existing UAV-assisted IRS-MEC system, the fully overlapping NOMA technology is difficult to balance the computing urgency and resource fairness of different users, which limits the overall performance of the system, especially in scenarios with heterogeneous latency requirements of user tasks.
A UAV collaborative IRS-MEC system assisted by NCO-NOMA is adopted. Through a hybrid optimization method and deep reinforcement learning algorithm, resources are dynamically allocated, including user transmission power, transmission time, IRS reflection phase shift and UAV flight trajectory, to optimize the task offloading and calculation process of the user group, and a convex optimization and DRL model is constructed to minimize the total system delay.
It achieves the minimization of system total latency and improvement of resource utilization efficiency in scenarios with heterogeneous task requirements, adapts to diversified service requirements, and improves the system's spectrum efficiency and user experience.
Smart Images

Figure CN120602995A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of MEC technology, and in particular relates to a dynamic resource allocation method for an NCO-NOMA-assisted UAV collaborative IRS-MEC system. Background Art
[0002] With the rapid development of information and communication technologies, the proliferation of mobile devices has driven the expansion of wireless communication networks, generating massive amounts of data traffic. According to the International Telecommunication Union (ITU), global mobile data traffic is expected to double by 2030, posing significant challenges in processing, offloading, latency, throughput, and energy efficiency. Mobile Edge Computing (MEC), an emerging computing technology, significantly reduces latency and improves real-time data processing capabilities by offloading computing tasks to the edge of the network, close to users. MEC combines the advantages of cloud computing and the edge of the network, providing an efficient and scalable data processing solution. However, traditional ground-based MEC base stations are limited by fixed deployments and struggle to adapt to dynamically changing user needs. Against this backdrop, unmanned aerial vehicles (UAVs), with their exceptional flexibility and maneuverability, can provide efficient computing and communication support to ground users in diverse environments, meeting their demands for high bandwidth and low latency, and are becoming an increasingly important complement to MEC networks. By introducing UAVs, wireless service providers can provide users with mobile MEC servers, further improving the system's computing power and service quality. Therefore, collaborative computing between UAVs and MECs has become an effective way to improve system performance and meet user needs. In UAV-MEC collaborative systems, researchers are exploring how to optimize system resources through flexible deployment and collaborative computing by combining UAVs with base stations (BSs). In recent years, much research has focused on optimizing resource allocation by introducing non-orthogonal multiple access (NOMA) technology. NOMA technology can effectively improve spectrum efficiency and reduce interference between users, thereby enhancing overall system performance and efficiency. Through reasonable resource optimization, NOMA not only improves system spectrum utilization but also enhances user experience, becoming an important means of optimizing UAV-MEC collaborative computing systems.
[0003] Intelligent Reflecting Surfaces (IRS), an emerging technology, have been widely adopted in MEC networks, primarily to enhance communication system performance. IRSs consist of multiple intelligent reflective elements, each of which independently improves the quality of the initial received signal by adjusting the signal's phase and amplitude. Furthermore, IRSs can be flexibly deployed in a variety of scenarios, such as building exteriors, moving vehicles, and ceilings, to enhance base station signal transmission quality. In MEC networks, the combination of IRSs and UAV technology can significantly improve system coverage, spectral efficiency, and energy efficiency. UAVs provide line-of-sight (LoS)-dominated transmission links, while IRSs optimize signal propagation paths by adjusting their reflective elements to achieve passive beamforming.
[0004] Existing UAV-assisted IRS-MEC systems primarily utilize Completely Overlapping NOMA (CO-NOMA) technology, where users within the same NOMA group share the same frequency and time resources, requiring all users to complete task transmissions synchronously. This transmission approach works well in scenarios where user task data volumes and latency requirements are similar. However, in real-world MEC systems, user computing tasks often have heterogeneous latency requirements, limiting CO-NOMA's adaptability to diverse service requirements. This synchronous transmission mechanism struggles to balance the computational urgency and resource fairness of different users, limiting the overall performance of the MEC system.
[0005] To address the above issues, Non-Completely Overlapping NOMA (NCO-NOMA) was proposed as a more flexible MEC multiple access scheme. Unlike CO-NOMA's fixed transmission duration, NCO-NOMA allows users within the same NOMA group to have different transmission times, meaning that users can complete task transmission and offload on partially overlapping time-frequency resources. By flexibly configuring each user's transmission duration and power strategy, it can not only better adapt to heterogeneous task requirements, but also effectively mitigate intra-group interference and improve communication resource utilization efficiency. Therefore, studying the dynamic resource allocation problem of the UAV collaborative IRS-MEC system based on NCO-NOMA has extremely important theoretical significance and practical value. Summary of the Invention
[0006] The main purpose of the present invention is to provide a dynamic resource allocation method for an NCO-NOMA-assisted UAV collaborative IRS-MEC system to solve the problems existing in the prior art and minimize the task completion delay of all users in the system.
[0007] In order to achieve the above object, the solution of the present invention is: A dynamic resource allocation method for UAV cooperative IRS-MEC system assisted by NCO-NOMA, which is aimed at a The proposed system consists of an IRS with a reflective element, several users, and several UAVs. Each UAV has computing capabilities and acts as an aerial MEC server to provide computing services to users. There is no communication link between the user and the UAV, but rather tasks are offloaded with the help of the IRS. The system is characterized by the following steps: Step 1. Definition Represents a collection of user groups. Indicates the total number of user groups, Indicates the user groups, each containing two users , Indicates the user's serial number; definition represents a collection of UAVs, represents the total number of UAVs, Indicates the UAVs; system time is evenly divided into time slots, denoted as , the duration of each time slot is Assuming that in each time slot, each user has a computationally intensive and delay-sensitive task that needs to be offloaded to the UAV for computation, the goal of the dynamic resource allocation method is to minimize the task completion delay of all users in the UAV-cooperative IRS-MEC system based on non-completely overlapping NOMA. Its objective function is defined as ,in Indicates that all users are in the time slot Total task delay in ; Step 2: A hybrid optimization method based on theoretical derivation and learning algorithm is performed, which specifically includes the following sub-steps: Step 2.1: Decompose the problem into two subproblems: the first is the joint optimization of user transmit power, individual transmission time slots, and computational frequency allocation; the second is the optimization of IRS phase shift, shared transmission time slots, UAV selection strategy, and flight trajectory; Step 2.2: Given the IRS phase shift, shared transmission period, UAV selection strategy, and flight trajectory, we introduce auxiliary variables to verify that the transmit power, individual transmission period, and frequency allocation optimization subproblems for each user group are standard convex difference programming problems. We then transform these problems into convex optimization problems and solve them using the concave-convex process algorithm. Step 3: For the real-time optimization of IRS phase shift, shared transmission period, UAV selection strategy and flight trajectory, it is modeled as a deep reinforcement learning problem, and a penalty mechanism is introduced in the reward function to strictly meet the requirements of all users in the time slot. Total task delay in The specific steps are as follows: Step 3.1 Define the state space and divide the time slots Status in It is defined as the set of the UAV position and all effective channel gains in the current time slot, that is: ; in, to Indicates UAV to location, to Represents a user to The channel gain with IRS, to Represents a user to The channel gain with IRS, to Indicates IRS and UAV to The channel gain between Step 3.2 Define the action space and divide the time slots Actions in It is defined as the set of IRS phase shift, shared transmission period, UAV selection strategy and flight trajectory, namely: ; in, to Indicates IRS and 1st to The phase shift between the reflective elements, to Indicates user group to Shared transmission time, to Indicates user group to The selected UAV index value, to Indicates the serial number to The horizontal flight angle of the UAV, to Indicates the serial number to The displacement step of the UAV in the current time slot; Step 3.3 Define the reward function and convert the time slot Rewards within Defined as: ; in, represents a negative exponent; Indicates the total system delay; 、 Both represent penalty items used to constrain unreasonable decisions; if the current time slot Medium Action Make user groups There is no solution to the joint optimization problem of calculating the frequency allocation and user transmission power for the individual transmission period. ,otherwise ; When multiple user groups select the same UAV and cause resource conflicts, The value of is set to the number of the same index value of the currently selected UAV, otherwise ; If the decision made in the current time slot results in at least one of the penalty items being non-zero, then let ; Step 4: Based on the SAC (soft actor-critic) algorithm, obtain the approximately optimal real-time IRS phase shift, shared transmission period, UAV selection strategy, and flight trajectory.
[0008] The user group The task processing process consists of two stages: task offloading and task calculation. In the above step 1, all users in the time slot Total task delay in The following constraints need to be met: Constraint 1. In the time slot Each UAV can only process computing tasks for one user group at most; Constraint 2. In the time slot Within, each user group’s task is handled by one UAV; Constraint 3. In the time slot Inside, UAV Flight speed Do not exceed the preset maximum flight speed ; Constraint 4. The initial position of any UAV is the same; Constraint 5. In the time slot The total amount of mission data for all users is limited by the uplink bandwidth and power; Constraint 6. In the time slot In user group Users and Transmit power 、 and Do not exceed the maximum transmit power limit , where the superscript 、 Respectively represent shared transmission period and individual transmission period; Constraint 7. In the time slot Internal, user Unloading energy consumption Less than its maximum available energy ; Constraint 8. In the time slot In user group Users Transfer time during shared transfer period The sum of the UAV task calculation time and its maximum task tolerance delay does not exceed , its users Transfer time during shared transfer period , transmission time of a single transmission period The sum of the UAV task calculation time and the UAV task calculation time does not exceed its maximum task tolerance delay ; Constraint 9. In the time slot Inside, UAV Assign to user The calculation frequency does not exceed its maximum calculation frequency ; Constraint 10. In the time slot Inside, UAV Computational energy consumption and flight energy consumption The sum does not exceed its maximum allowable energy consumption ;in, Represents a binary variable, used to refer to UAV If the selection strategy of the user group Offloading its tasks to UAVs but ,otherwise ; Indicates time slot The UAV calculates the energy consumed by the user's task.
[0009] The specific calculation process of step 4 is: Step 4.1: Given a discount factor , soft copy factor , Maximum number of rounds and the number of time slots per round ; Step 4.2 Randomly initialize weight parameters 、 、 、 and , and clear the experience replay pool ; Step 4.3 The SAC agent enters the training process consisting of two nested loops. The total training rounds, each round contains Time slots - at the beginning of each training round, initialize the environment and obtain the initial state ; Then in each time slot In the process, the SAC agent is based on the current state Select actions by strategy , Represents the policy network; after executing the action, the environment will feedback the immediate reward , and transfer the state to the next state ; At this time, the experience sample obtained in the current time slot Stored in the replay pool ; Next, from the playback pool Randomly extract a preset number of experience samples for training and update the weight parameters 、 、 、 and ; Step 4.4 After training is completed, the SAC algorithm outputs the optimal weight parameters of the policy network, thereby obtaining a near-optimal real-time solution for IRS phase shift, shared transmission period, UAV selection strategy, and flight trajectory.
[0010] After adopting the above technical solution, the present invention has the following technical effects: (1) This paper proposes a hybrid optimization method based on the combination of theoretical derivation and learning algorithm to achieve efficient real-time optimization in complex systems; (2) For time-varying channels and multi-UAV collaboration scenarios, a delay minimization problem for the UAV collaborative IRS-MEC system based on NCO-NOMA is constructed. The optimization problem is decomposed into the sub-problems of optimizing the individual transmission period, calculating frequency allocation and user transmit power, and optimizing the IRS phase shift, sharing the transmission period, UAV selection strategy, and flight trajectory. (3) Given the IRS phase shift, shared transmission period, UAV selection strategy, and flight trajectory, by introducing auxiliary variables, we prove that the transmission power, individual transmission period, and computational frequency allocation optimization subproblems for each user group belong to standard DC programming problems. We then transform them into convex optimization problems and solve them using the CCCP algorithm. (4) The real-time optimization of IRS phase shift, shared transmission period, UAV selection strategy and flight trajectory is modeled as a DRL problem. A SAC-based DRL algorithm is introduced to obtain a near-optimal real-time decision strategy, and a penalty mechanism is introduced into the reward function to strictly satisfy the constraints. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 This is a system scenario diagram applicable to a specific embodiment of the present invention.
[0012] Figure 2 This is a schematic diagram of the task offloading and calculation process of a user group in a system scenario in a specific embodiment of the present invention.
[0013] Figure 3 Flowchart of a specific embodiment of the present invention. DETAILED DESCRIPTION
[0014] In order to further explain the technical solution of the present invention, the present invention is described in detail below through specific embodiments.
[0015] refer to Figure 1 As shown, the present invention discloses a dynamic resource allocation method for a UAV cooperative IRS-MEC system assisted by NCO-NOMA, which is aimed at a UAV having This paper proposes a MEC system that includes an IRS (intelligent reflective surface) with multiple reflective elements, several users, and several UAVs (unmanned aerial vehicles). Each UAV has computing capabilities and acts as an aerial MEC server to provide computing services to users. There is no communication link between the users and the UAVs, but rather tasks are offloaded with the help of the IRS. To save spectrum resources and improve offloading efficiency, the present invention groups users into groups, with each group containing two users and occupying an independent subchannel, and offloads tasks using non-completely overlapping NOMA technology. All users' devices are equipped with a single antenna to meet lightweight deployment requirements. The goal of this invention is to minimize the task completion delay for all users in the UAV-cooperative IRS-MEC system based on non-completely overlapping NOMA by jointly optimizing user transmit power, transmission time, and IRS reflection phase shift, while meeting the task delay constraints of each user and the energy consumption constraints of each user and UAV.
[0016] The present invention comprises the following steps: Step 1. Definition Represents a collection of user groups. Indicates the total number of user groups, Indicates the user groups, each containing two users , Indicates the user's serial number; definition represents a collection of UAVs, represents the total number of UAVs, Indicates the UAVs; system time is evenly divided into time slots, denoted as , the duration of each time slot is Assuming that in each time slot, each user has a computationally intensive and delay-sensitive task that needs to be offloaded to the UAV for computation, the goal of the dynamic resource allocation method is to minimize the task completion delay of all users in the UAV-cooperative IRS-MEC system based on non-completely overlapping NOMA. Its objective function is defined as ,in Indicates that all users are in the time slot The total task delay in .
[0017] Specifically, refer to Figure 2 As shown, the user group The task processing process consists of two stages: task offloading and task calculation. In the above step 1, all users in the time slot Total task delay in The following constraints need to be met: Constraint 1. In the time slot Each UAV can only process computing tasks for one user group at most; Constraint 2. In the time slot Within, each user group’s task is handled by one UAV; Constraint 3. In the time slot Inside, UAV Flight speed Do not exceed the preset maximum flight speed ; Constraint 4. The initial position of any UAV is the same; Constraint 5. In the time slot The total amount of mission data for all users is limited by the uplink bandwidth and power; Constraint 6. In the time slot In user group Users and Transmit power 、 and Do not exceed the maximum transmit power limit , where the superscript 、 Respectively represent shared transmission period and individual transmission period; Constraint 7. In the time slot Internal, user Unloading energy consumption Less than its maximum available energy ; Constraint 8. In the time slot In user group Users Transfer time during shared transfer period The sum of the UAV task calculation time and its maximum task tolerance delay does not exceed , its users Transfer time during shared transfer period , transmission time of a single transmission period The sum of the UAV task calculation time and the UAV task calculation time does not exceed its maximum task tolerance delay ; Constraint 9. In the time slot Inside, UAV Assign to user The calculation frequency does not exceed its maximum calculation frequency ; Constraint 10. In the time slot Inside, UAV Computational energy consumption and flight energy consumption The sum does not exceed its maximum allowable energy consumption ;in, Represents a binary variable, used to refer to UAV If the selection strategy of the user group Offloading its tasks to UAVs but ,otherwise ; Indicates time slot The UAV calculates the energy consumed by the user's task.
[0018] Step 2: A hybrid optimization method based on theoretical derivation and learning algorithm is performed, which specifically includes the following sub-steps: Step 2.1: Decompose the problem into two subproblems: the first is the joint optimization of user transmit power, individual transmission time slots, and computational frequency allocation; the second is the optimization of IRS phase shift, shared transmission time slots, UAV selection strategy, and flight trajectory; Step 2.2: Given the IRS phase shift, shared transmission period, UAV selection strategy, and flight trajectory, we introduce auxiliary variables to verify that the transmit power, individual transmission period, and frequency allocation optimization subproblems for each user group are standard convex difference of convex functions (DC) programming problems. We then transform these problems into convex optimization problems and solve them using the Concave-Convex Procedure (CCCP) algorithm.
[0019] Step 3: For the real-time optimization of IRS phase shift, shared transmission period, UAV selection strategy and flight trajectory, it is modeled as a deep reinforcement learning (DRL) problem, and a penalty mechanism is introduced in the reward function to strictly meet the requirements of all users in the time slot. Total task delay in The specific steps are as follows: Step 3.1 Define the state space and divide the time slots Status in It is defined as the set of the UAV position and all effective channel gains in the current time slot, that is: ; in, to Indicates UAV to location, to Represents a user to The channel gain with IRS, to Represents a user to The channel gain with IRS, to Indicates IRS and UAV to The channel gain between Step 3.2 Define the action space and divide the time slots Actions in It is defined as the set of IRS phase shift, shared transmission period, UAV selection strategy and flight trajectory, namely: ; in, to Indicates IRS and 1st to The phase shift between the reflective elements, to Indicates user group to Shared transmission time, to Indicates user group to The selected UAV index value, to Indicates the serial number to The horizontal flight angle of the UAV, to Indicates the serial number to The displacement step of the UAV in the current time slot; Step 3.3 Define the reward function and convert the time slot Rewards within Defined as: ; in, represents a negative exponent; Indicates the total system delay; 、 Both represent penalty items used to constrain unreasonable decisions; if the current time slot Medium Action Make user groups There is no solution to the joint optimization problem of calculating the frequency allocation and user transmission power for the individual transmission period. ,otherwise ; When multiple user groups select the same UAV and cause resource conflicts, The value of is set to the number of the same index value of the currently selected UAV, otherwise ; If the decision made in the current time slot results in at least one of the penalty items being non-zero, then let .
[0020] From the design of the reward function, it can be seen that the introduction of the penalty term encourages the agent to avoid unreasonable strategies in subsequent time slots, thereby effectively minimizing the total delay of the system.
[0021] Step 4: Based on the SAC (soft actor-critic) algorithm, obtain the approximately optimal real-time IRS phase shift, shared transmission period, UAV selection strategy, and flight trajectory.
[0022] Specifically, the calculation process of step 4 above is: Step 4.1: Given a discount factor , soft copy factor , Maximum number of rounds and the number of time slots per round ; Step 4.2 Randomly initialize weight parameters 、 、 、 and , and clear the experience replay pool ; Step 4.3 The SAC agent enters the training process consisting of two nested loops. The total training rounds, each round contains Time slots - at the beginning of each training round, initialize the environment and obtain the initial state ; Then in each time slot In the process, the SAC agent is based on the current state Select actions by strategy , Represents the policy network; after executing the action, the environment will feedback the immediate reward , and transfer the state to the next state ; At this time, the experience sample obtained in the current time slot Stored in the replay pool ; Next, from the playback pool Randomly extract a preset number of experience samples for training and update the weight parameters 、 、 、 and ; Step 4.4 After training is completed, the SAC algorithm outputs the optimal weight parameters of the policy network, thereby obtaining a near-optimal real-time solution for IRS phase shift, shared transmission period, UAV selection strategy, and flight trajectory.
[0023] Through the above scheme, the present invention proposes a hybrid optimization method based on the combination of theoretical derivation and learning algorithm to achieve efficient real-time optimization under complex systems; for time-varying channels and multi-UAV collaborative scenarios, a UAV collaborative IRS-MEC system delay minimization problem based on NCO-NOMA is constructed; the optimization problem is decomposed into individual transmission period, calculation frequency allocation and user transmission power optimization sub-problems and IRS phase shift, shared transmission period, UAV selection strategy and flight trajectory optimization sub-problems; given the IRS phase shift, shared transmission period, UAV selection strategy and flight trajectory, by introducing auxiliary variables, it is proved that the transmission power of each user group, individual transmission period and calculation frequency allocation optimization sub-problems belong to standard DC programming problems, and it is converted into a convex optimization problem and solved using the CCCP algorithm; for the real-time optimization of IRS phase shift, shared transmission period, UAV selection strategy and flight trajectory, it is modeled as a DRL problem, and a SAC-based DRL algorithm is introduced to obtain a near-optimal real-time decision strategy, and a penalty mechanism is introduced in the reward function to strictly meet the constraints.
[0024] The above embodiments and drawings do not limit the product form and style of the present invention. Any appropriate changes or modifications made by ordinary technicians in the relevant technical field should be deemed to be within the patent scope of the present invention.
Claims
1. A dynamic resource allocation method for UAV cooperative IRS-MEC system assisted by NCO-NOMA, targeting a The MEC system consists of an IRS with a reflective element, several users and several UAVs; each UAV has computing capabilities and acts as an aerial MEC server to provide computing services to users; there is no communication link between users and UAVs, but tasks are offloaded with the help of IRS; it is characterized by The following steps are involved: Step 1. Definition Represents a collection of user groups. Indicates the total number of user groups, Indicates the User groups, each containing two users , Indicates the user's serial number; definition represents a collection of UAVs, represents the total number of UAVs, Indicates the UAVs; system time is evenly divided into time slots, denoted as , the duration of each time slot is Assuming that in each time slot, each user has a computationally intensive and delay-sensitive task that needs to be offloaded to the UAV for computation, the goal of the dynamic resource allocation method is to minimize the task completion delay of all users in the UAV-cooperative IRS-MEC system based on non-completely overlapping NOMA. Its objective function is defined as ,in Indicates that all users are in the time slot Total task delay in ; Step 2: A hybrid optimization method based on theoretical derivation and learning algorithm is performed, which specifically includes the following sub-steps: Step 2.1: Decompose the problem into two subproblems: the first is the joint optimization of user transmit power, individual transmission time slots, and computational frequency allocation; the second is the optimization of IRS phase shift, shared transmission time slots, UAV selection strategy, and flight trajectory; Step 2.2: Given the IRS phase shift, shared transmission period, UAV selection strategy, and flight trajectory, we introduce auxiliary variables to verify that the transmit power, individual transmission period, and frequency allocation optimization subproblems for each user group are standard convex difference programming problems. We then transform these problems into convex optimization problems and solve them using the concave-convex process algorithm. Step 3: For the real-time optimization of IRS phase shift, shared transmission period, UAV selection strategy and flight trajectory, it is modeled as a deep reinforcement learning problem, and a penalty mechanism is introduced in the reward function to strictly meet the requirements of all users in the time slot. Total task delay in The specific steps are as follows: Step 3.1 Define the state space and divide the time slots Status in It is defined as the set of the UAV position and all effective channel gains in the current time slot, that is: ; in, to Indicates UAV to location, to Represents a user to The channel gain with IRS, to Represents a user to The channel gain with IRS, to Indicates IRS and UAV to The channel gain between Step 3.2 Define the action space and divide the time slots Actions in It is defined as the set of IRS phase shift, shared transmission period, UAV selection strategy and flight trajectory, namely: ; in, to Indicates IRS and 1st to The phase shift between the reflective elements, to Indicates user group to Shared transmission time, to Indicates user group to The selected UAV index value, to Indicates the serial number to The horizontal flight angle of the UAV, to Indicates the serial number to The displacement step of the UAV in the current time slot; Step 3.3 Define the reward function and convert the time slot Rewards within Defined as: ; in, represents a negative exponent; Indicates the total system delay; 、 Both represent penalty items used to constrain unreasonable decisions; if the current time slot Medium Action Make user groups There is no solution to the joint optimization problem of calculating the frequency allocation and user transmission power for the individual transmission period. ,otherwise ; When multiple user groups select the same UAV and cause resource conflicts, The value of is set to the number of the same index value of the currently selected UAV, otherwise ; If the decision made in the current time slot results in at least one of the penalty items being non-zero, then let ; Step 4: Based on the SAC (soft actor-critic) algorithm, obtain the approximately optimal real-time IRS phase shift, shared transmission period, UAV selection strategy, and flight trajectory.
2. The dynamic resource allocation method of the NCO-NOMA-assisted UAV collaborative IRS-MEC system as claimed in claim 1 is characterized in that The user group The task processing process consists of two stages: task offloading and task calculation. In the above step 1, all users in the time slot Total task delay in The following constraints need to be met: Constraint 1. In the time slot Each UAV can only process computing tasks for one user group at most; Constraint 2. In the time slot Within, each user group’s task is handled by one UAV; Constraint 3. In the time slot Inside, UAV Flight speed Do not exceed the preset maximum flight speed ; Constraint 4. The initial position of any UAV is the same; Constraint 5. In the time slot The total amount of mission data for all users is limited by the uplink bandwidth and power; Constraint 6. In the time slot In user group Users and Transmit power 、 and Do not exceed the maximum transmit power limit , where the superscript 、 Respectively represent shared transmission period and individual transmission period; Constraint 7. In the time slot Internal, user Unloading energy consumption Less than its maximum available energy ; Constraint 8. In the time slot In user group Users Transfer time during shared transfer period The sum of the UAV task calculation time and its maximum task tolerance delay does not exceed , its users Transfer time during shared transfer period , transmission time of a single transmission period The sum of the UAV task calculation time and the UAV task calculation time does not exceed its maximum task tolerance delay ; Constraint 9. In the time slot Inside, UAV Assign to user The calculation frequency does not exceed its maximum calculation frequency ; Constraint 10. In the time slot Inside, UAV Computational energy consumption and flight energy consumption The sum does not exceed its maximum allowable energy consumption ;in, Represents a binary variable, used to refer to UAV If the selection strategy of the user group Offloading its tasks to UAVs but ,otherwise ; Indicates time slot The UAV calculates the energy consumed by the user's task.
3. The dynamic resource allocation method of the NCO-NOMA-assisted UAV collaborative IRS-MEC system as claimed in claim 1 is characterized in that The specific calculation process of step 4 is: Step 4.1 Given a discount factor , soft copy factor , Maximum number of rounds and the number of time slots per round ; Step 4.2 Randomly initialize weight parameters 、 、 、 and , and clear the experience replay pool ; Step 4.3 The SAC agent enters the training process consisting of two nested loops. The total training rounds, each round contains Time slots - at the beginning of each training round, initialize the environment and obtain the initial state ; Then in each time slot In the process, the SAC agent is based on the current state Select actions by strategy , Represents the policy network; after executing the action, the environment will feedback the immediate reward , and transfer the state to the next state ; At this time, the experience sample obtained in the current time slot Stored in the replay pool ; Next, from the playback pool Randomly extract a preset number of experience samples for training and update the weight parameters 、 、 、 and ; Step 4.4 After training is completed, the SAC algorithm outputs the optimal weight parameters of the policy network, thereby obtaining an approximately optimal real-time solution for IRS phase shift, shared transmission period, UAV selection strategy, and flight trajectory.
Citation Information
Cited By
Downlink fairness transmission rate optimization method for 5G / 6G UAV-IRS assisted NOMA-MEC system
CN122160794A