Air-ground network optimization method and system based on fairness guarantee of MEC
Through deep reinforcement learning based on MEC, the problem of uneven resource allocation in multiple UAV systems is solved, energy consumption is minimized and fairness is improved, and the adaptability and service quality of the air-ground integrated network is enhanced.
Patent Information
- Application Number
- CN202510268093.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-07-01
AI Technical Summary
The prior art ignores resource balanced allocation in multi-drone-assisted MEC systems, resulting in some drones being overloaded, and lacks in-depth exploration of multi-target optimization problems, especially how to achieve fairness and coordination of resource allocation in a multi-drone environment.
Using the MEC-based air-ground network optimization method, the policy network and critic network in deep reinforcement learning (DRL) are used to optimize the drone trajectory and task scheduling through initialization, training, and updating agent parameters, so as to achieve fairness and multi-objective optimization of resource allocation.
It achieves the minimization of energy consumption of IoTDs and UAVs, improves network coverage and service quality, enhances the adaptability and flexibility of the system, and meets the needs of low-latency service.
Smart Images

Figure CN120238958A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to, but is not limited to, the field of communication technologies, and particularly relates to an air-ground network optimization method and system for fairness guarantee based on MEC. Background Art
[0002] The air-ground integrated network (AGIN) expands the network coverage by integrating air and ground resources, provides a solution for the explosion of Internet of Things devices (IoTDs), and solves the coverage and capacity limitations of traditional networks. Multi-access edge computing (MEC) technology reduces service latency and improves the user experience by offloading computing tasks to the network edge, especially in high-bandwidth application scenarios such as virtual reality, augmented reality, and high-definition video transmission. The combination of AGIN and MEC, together with the flexibility of unmanned aerial vehicles (UAVs), optimizes the computing load of IoT devices, improves data transmission efficiency, and enhances the adaptability and flexibility of the network. Artificial intelligence technology, especially deep reinforcement learning (DRL), shows great potential in AGIN to solve complex system problems, provides effective solutions for the coordination of multi-UAV systems and multi-objective optimization problems, and enables real-time decision-making, which is crucial for building an air-ground integrated information processing network that can adapt to complex environmental challenges and meet the requirements of low-latency services.
[0003] When combining MEC and UAVs into the air-ground integrated network architecture, some challenges still need to be addressed.
[0004] (1) Coordination and fairness of multi-UAV systems
[0005] Multi-UAV-assisted MEC systems show great potential in providing computing and communication services. However, in the pursuit of performance improvement, the issue of system fairness is often overlooked. Although the layout of multi-UAVs in three-dimensional space has been considered to adapt to the selection of ground devices, how to achieve effective coordination among these UAVs remains a problem that needs further exploration. In addition, the challenge of balanced resource allocation in a multi-UAV environment is often ignored, which may lead to overload of some UAVs and reduce efficiency. Therefore, how to ensure the fairness of resource allocation and achieve effective coordination among multi-UAVs while improving system performance to avoid uneven resource allocation and overload problems has become a key problem to be solved urgently.
[0006] (2) Multi-objective optimization
[0007] In a multi-UAV-assisted MEC system, optimizing computational efficiency, energy consumption, latency, and security is a complex issue. The current challenge is how to improve computational efficiency through the comprehensive optimization of computing resources, bandwidth allocation, and UAV trajectories. The complexity of these problems lies in that they not only involve the optimization of energy consumption but also need to consider the optimization of fairness. Therefore, designing a comprehensive framework that can optimize these objectives simultaneously to adapt to the dynamically changing environment and improve the adaptability and flexibility of the system is an important direction in the current technological development.
[0008] (3) Application of Deep Reinforcement Learning
[0009] Deep reinforcement learning technology can adapt to environments with dynamically changing tasks and resources without the need for manual setting of detailed scheduling rules, thereby optimizing performance. In practical applications, DRL is used to optimize communication resource allocation, improve computational efficiency and response speed in mobile edge computing, and solve complex multi-variable optimization problems. Although DRL has shown effectiveness in approximating Q-values and has been applied to online resource allocation and scheduling design in wireless networks, in the case of large-scale problems, the training of DRL models may become unstable, and the huge decision space is one of the challenges faced by current technologies. In addition, intelligently designing the flight trajectories of UAVs in mobile edge computing networks to serve a large number of devices, especially considering the dynamic mobility of devices and the dynamic association between UAVs and devices, is an area that needs further exploration and improvement in the current technological development.
[0010] Through the above analysis, the problems and defects of the existing technologies are as follows:
[0011] How to effectively manage and schedule the surging IoTDs to ensure efficient data processing and low communication latency. Although MEC technology significantly reduces service latency and improves the user experience by offloading computing tasks to the network edge, existing research often ignores the problem of achieving balanced resource allocation in a multi-UAV environment, resulting in some UAVs being overloaded and reducing efficiency. In addition, existing literature lacks in-depth exploration of the application of DRL in multi-UAV systems, especially in multi-objective optimization problems such as coordination and fairness, as well as comprehensive multi-objective optimization. The complexity of these problems lies in that they not only involve the optimization of energy consumption and fairness index but also need to consider the complexity brought by task volume and task queuing. Therefore, in a dynamically changing environment, how to design a comprehensive optimization framework that can optimize multiple objectives simultaneously to adapt to environmental changes and improve the adaptability and flexibility of the system is the key problem that needs to be solved in the current technological development. Summary of the Invention
[0012] Aiming at the problems existing in the prior art, the present invention provides an air-ground network optimization method and system based on fairness guarantee of MEC.
[0013] The present invention is realized through the following technical solutions, an air-ground network optimization method and system based on fairness guarantee of MEC: An air-space-ground integrated network performance optimization framework relying on MEC is proposed. The framework uses the policy network (Actor Network) and the critic network (Critic Network) in DRL to train the agent and make decisions. First, the parameters of the policy network and the critic network are initialized, and relevant training hyperparameters are configured. During the training process, the agent interacts with the environment according to the current policy, executes corresponding actions and observes the state transition. The interaction experience of the agent is accumulated through the experience replay buffer and updated when the buffer is full, so as to maintain the latest learning samples. The loss function is used to calculate the policy performance of the agent, and the gradient descent algorithm is used to optimize the parameters of the policy network and the critic network. At the same time, the parameters of the policy network and the value network in the target network are softly updated regularly to ensure the stability of the learning process. The training continues until the policy network converges, and finally the trained policy network is applied to task allocation and trajectory planning to optimize the network performance.
[0014] Furthermore, an air-ground network optimization method and system based on fairness guarantee of MEC:
[0015] S101. Initialize the parameters of the main policy network and the critic network of the agent and the parameters of the target policy network and the critic network and the number of episodes EP, the maximum number of training steps T max , initialize the learning rates α c and α a corresponding to the critic network and the policy network, the discount factor γ, initialize the size D of the replay buffer, the size M of the mini-batch, and the noise ∈ for action exploration; initialize the network layout parameters, such as the number I of IoTDs, the number J of unmanned aerial vehicles, etc.
[0016] S102. After initializing the state of the agent, the agent interacts with the environment, and the main policy network generates corresponding actions according to the current policy.
[0017] S103. The agent operates according to the actions generated by the main policy network, then receives the rewards feedback by the environment, and updates its state according to these feedbacks.
[0018] S104. Store the experience tuple into the experience replay buffer. When the buffer reaches the capacity limit, introduce the latest experience by replacing the oldest experience data.
[0019] S105. Update the parameters of the main policy network and the critic network.
[0020] S106. Calculate the loss function of the critic network according to the temporal difference (TD) target and the value function predicted by the critic network, sample from the experience replay buffer, and update the target policy network and the critic network using the gradient descent method.
[0021] S107. Update the parameters of the main policy network and the critic network using a small batch of experience samples.
[0022] S108. Update the parameters of the policy network and the value network in the target network through a soft update mechanism.
[0023] S109. After iterative training until the algorithm converges stably, apply this policy to the agent to achieve optimal task offloading and resource allocation.
[0024] Furthermore, in S102, initialize the state of the agent. The agent interacts with the environment, and the main policy network generates an action based on the current policy. The state of the agent is represented as:
[0025] o j (t) = {u j (t - 1), φ j (t)},
[0026] where u j (t - 1) = (X j (t - 1), Y j (t - 1), H j (t - 1)), represents the three-dimensional coordinates of the UAV j at the end of time slot t - 1, corresponding to the starting point of time slot t; φ j (t) = {φ j (t)}, represents the task arrival metric of the IoT devices covered by the UAV j at the beginning of time slot t.
[0027] Furthermore, in S103: The agent executes the action generated by the main policy network, obtains a reward, and updates the state. The calculation formula of the reward reward in the state update is as follows:
[0028]
[0029] In the above formula, O j(t) represents the research objective of the present invention, that is, to minimize the overall energy consumption of the IoTD, the energy consumption of the UAV, reduce the number of discarded tasks, and improve fairness among the IoTDs.
[0030] Further, in step S105: Update the parameters of the main policy network and the critic network; update the current policy network by gradient ascent as follows:
[0031]
[0032] Where represents the partial derivative of θ j .
[0033] Further, in step S106: Calculate the loss function of the critic network according to the TD target and the value function predicted by the critic network, extract samples from the experience replay buffer, and update the target policy network and the evaluation network using the gradient descent method.
[0034] Furthermore, update the parameter w of the main value network and the current value network by gradient descent j as follows:
[0035]
[0036] Where and q(s m(t) , a m(t) ; w j ) represent the time difference error of the j-th weight at time slot t and the action value function, respectively.
[0037] Further, in step S108: Implement parameter updates of the policy network and the value network in the target network using a soft update mechanism; the soft update formula is as follows:
[0038]
[0039] Where θ j represents the parameter of the current policy network, represents the parameter of the target policy network, w j represents the parameter of the current value network, represents the parameter of the target value network, and χ ∈ [0, 1] is the parameter used to update the target network.
[0040] Another object of the present invention is to provide an air-ground network optimization method and system for implementing the fairness guarantee based on MEC, including:
[0041] A system initialization module for initializing the parameters of the deep deterministic policy gradient algorithm, including setting the number of episodes EP and the maximum number of training steps T max , and initializing the learning rates α corresponding to the critic network and the policy networkc and α a , discount factor γ, initialize the size D of the replay buffer, and the size M of the mini-batch;
[0042] The network construction module is used for the noise ∈ of action exploration; initialize the network layout parameters, such as the number I of IoTDs, the number J of UAVs, and other parameters;
[0043] The agent module is used to generate actions based on the current network state at the beginning of each cycle, and has the function of adding exploratory noise to these behaviors in order to introduce a certain degree of randomness during execution;
[0044] The action execution module is used to execute the resource allocation and access control policies;
[0045] The reward acquisition module is used to execute actions and calculate immediate rewards, evaluate rewards according to the long-term average utility of all devices in the system, and the state transition module that transfers the system from the current state to the next state;
[0046] The experience replay module is used to store the experience tuples of the system state, executed actions, obtained rewards, and the next state each time;
[0047] The data sampling module is used to extract mini-batch experiences from the stored experience replay module for learning;
[0048] The network update module is used to update the main policy network and the main value network according to the data in the experience replay module, including a parameter optimization unit that uses the gradient ascent method and the gradient descent method to adjust the network parameters;
[0049] The parameter update module is used to synchronize the parameter updates of the main network to the target policy network and the target value network, adopting a soft update strategy, so that the parameters of the target network are the weighted average of the parameters of the main network, and this synchronization process is realized through the parameter synchronization unit.
[0050] Combined with the above technical solutions and the solved technical problems, the advantages and positive effects of the technical solution to be protected by the present invention are as follows:
[0051] First, the MEC-assisted AGIN architecture proposed by the present invention realizes the minimization of the energy consumption and task loss quantity of IoTDs and UAVs through refined access control, UAV trajectory planning, and task scheduling, while maximizing the fairness index. This strategy not only improves the network coverage and service quality, but also in the air-ground integrated network architecture, the UAV, as a key aerial node, can provide a wider network coverage for remote areas, making up for the deficiency of the ground base station coverage.
[0052] The multi-objective optimization method adopted by the present invention simultaneously considers multiple objectives such as computational efficiency, energy consumption, latency, and security. Such comprehensive performance considerations can more comprehensively improve system performance and meet the comprehensive requirements for system performance in practical applications. Especially in critical tasks such as emergency response and environmental monitoring, this multi-objective optimization method can ensure the efficient and stable operation of the system in the face of different scenarios.
[0053] The present invention uses DRL technology to solve complex system problems, especially the coordination and fairness problems in multi-UAV systems, as well as multi-objective optimization problems. The introduction of DRL makes it possible to make real-time decisions in a dynamically changing environment, which is crucial for building an air-ground integrated information processing network that can adapt to complex environmental challenges and meet the requirements of low-latency services, and plays an important role in improving the intelligence level and service quality of the network.
[0054] The present invention applies the improved Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm to handle continuous and discrete variables, and satisfies the coupling constraints between variables by discretizing continuous action outputs. This innovative improvement of the algorithm improves the robustness of real-time decision-making, especially the real-time decision-making ability in a dynamic environment, solves the limitations of traditional methods in dealing with such problems, and is of great significance for improving the adaptability and flexibility of the system.
[0055] Second, the significant technical progress specifically achieved by the present invention lies in realizing an air-ground network optimization method and system based on MEC fairness guarantee, and this method has made significant progress in the following key aspects:
[0056] (1) Resource allocation and energy consumption optimization:
[0057] The MEC-based AGIN architecture proposed by the present invention realizes the minimization of energy consumption of IoTDs and UAVs through refined access control and task scheduling. By optimizing computing resources, bandwidth allocation, and UAV trajectories, the system performance and resource utilization rate are significantly improved, and at the same time, the energy consumption is reduced, reflecting the comprehensive consideration of the system in terms of performance improvement and energy consumption reduction.
[0058] (2) Task partitioning and scheduling strategy:
[0059] By jointly optimizing 3-D UAV trajectory planning and task scheduling, the present invention realizes an intelligent task partitioning strategy. The system can dynamically adjust task processing according to the capabilities of UAVs and network conditions, and intelligently decide whether to process by UAVs or base stations, so as to adapt to the dynamic changes of the network, enhancing the stability of the network and the user experience.
[0060] (3)Integrated Application of DRL:
[0061] Through the adoption of DRL technology, especially the MADDPG algorithm, the present invention realizes the real-time data parsing of UAV-assisted AGIN and the adaptive ability to environmental changes. This technology integration enables the network to make optimal decisions under variable dynamic conditions, thereby enhancing the intelligence level and service quality of the system. Through the implementation of the MADDPG algorithm, the network can process continuous and discrete variables, optimize the decision-making process, improve the efficiency of real-time control strategies, and further enhance the decision-making quality and response speed of the system.
[0062] (4)System Utility and Fairness Optimization:
[0063] The objective function design proposed by the present invention takes into account the maximization of the fairness index and optimizes the total system utility through a reward mechanism. This not only improves energy efficiency but also pays attention to environmental friendliness and system fairness, reflecting a comprehensive consideration of the multi-objective optimization of the system.
[0064] (5)Improvement of System Stability and Adaptability:
[0065] Through the application of the MADDPG algorithm, the present invention realizes the accurate calculation of the immediate feedback after the UAV executes tasks and optimizes the state transition process, significantly enhancing the stability and reliability of the system in the face of large-scale data processing and high-density requests, and strengthening the adaptability and flexibility of the network.
[0066] The application of the mathematical model provided by the present invention not only improves the operation efficiency and decision-making quality of the system but also enhances the adaptability of the system to environmental changes and long-term stability. By integrating 3D trajectory optimization, resource allocation, and access control strategy optimization, it realizes the real-time monitoring and mapping of the network state. These technical effects are crucial for processing a large amount of data and high-frequency interactions in modern edge computing environments.
[0067] Thirdly, an air-ground network optimization method and system based on fairness guarantee of MEC provided by the present invention adopts deep reinforcement learning technology to optimize the performance of the network through the interaction between the agent and the environment.
[0068] Initialize the state of the agent. The state of the agent includes multiple variables, such as: the horizontal distance of the UAV, the flight altitude of the UAV, environment-related parameters, path loss parameters, and the task arrival metrics of the IoTDs covered by UAV j at the beginning of time slot t, etc. These variables jointly define the environmental state of the agent at a specific moment and thus affect the decision-making of the agent.
[0069] The agent executes tasks according to the actions determined by the main policy network. After executing the corresponding actions, the agent obtains corresponding rewards based on the results obtained, and updates and transfers the state according to the rewards and execution results. The calculation of the rewards is based on minimizing the overall energy consumption of the IoTD, the energy consumption of the UAV, while reducing the number of discarded tasks, and improving the fairness among the IoTDs.
[0070] Optimize the parameters of the policy network through gradient ascent, enabling it to generate better action selections, thereby enhancing the overall operating performance of the system.
[0071] Use the time difference (TD) error and the value prediction of the critic network to calculate the loss function, and update the parameters of the policy network and the critic network through gradient descent. This is the core link in the optimization of the value function in deep reinforcement learning. By improving the accuracy of value prediction, the policy is continuously optimized, thereby enhancing the system performance.
[0072] Adopt soft update technology to update the parameters of the target network. This method realizes the smooth transition of the target network parameters by calculating the weighted average of the target network and the current network parameters, which helps to reduce the drastic fluctuations in the learning process and ensure the stability of the learning process.
[0073] Fourth, as the creative auxiliary evidence of the claims of the present invention, it is also reflected in the following important aspects:
[0074] 1. Achievement of comprehensive optimization objectives:
[0075] The MEC-based AGIN architecture proposed by the present invention realizes multi-objective optimization through joint optimization of access control, UAV trajectory planning, and task scheduling, including minimizing the total energy consumption of IoTDs and UAVs, reducing the number of task losses, and maximizing the fairness index. This comprehensive optimization strategy fills the gap in multi-objective optimization in the prior art.
[0076] 2. Solving key technical challenges:
[0077] In response to challenges such as high latency sensitivity, energy-intensive applications, and scarce wireless resources in the air-ground network, the technical solution of the present invention proposes effective solutions through the edge computing ability of MEC and the flexible deployment of UAVs. By optimizing computing resources, bandwidth allocation, and UAV trajectories, the service latency is reduced, the user experience is improved, and the adaptability and flexibility of the network are enhanced.
[0078] 2. Decentralized decision optimization:
[0079] The technical solution of this paper realizes decentralized decision optimization through the MADDPG algorithm, overcoming the dependence on centralized processing in traditional network optimization methods. The algorithm can handle continuous and discrete variables, optimize system control and decision-making, improve the robustness of real-time decision-making, and achieve comprehensive optimization of network performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0080] Figure 1 is the flow chart of the air-ground network optimization method and system based on MEC fairness guarantee provided by the embodiments of the present invention;
[0081] Figure 2 is a scenario diagram applicable to the embodiments of the present invention;
[0082] Figure 3 is the structural diagram of the air-ground network optimization method and system based on MEC fairness guarantee provided by the embodiments of the present invention.
[0083] Figure 4 is the convergence performance comparison diagram provided by the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0084] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0085] The technical solution provided by the present invention proposes an air-ground network optimization method and system based on MEC fairness guarantee, mainly using the DRL framework. The solution mainly includes:
[0086] 1) Deep reinforcement learning framework: The DRL framework is adopted, which includes the main policy network and the critic network of the agent, and is used to perform learning and decision-making tasks.
[0087] 2) Network initialization: Before the start of training, the parameters of the main policy network, the target policy network and the critic network are initialized, and the hyperparameters related to training are configured.
[0088] 3) Environment interaction and state management: The agent interacts with the environment according to the current policy, executes corresponding actions and realizes state transition.
[0089] 4) Experience replay: The replay buffer is used to store the experience tuples of the agent's interaction with the environment. When the buffer reaches the capacity limit, it is updated to ensure that the latest experience data is used in the learning process.
[0090] 5) Network parameter update: The parameters of the main policy network and the critic network are updated by calculating the loss function and applying the gradient descent algorithm to optimize the network performance.
[0091] 6) Target network soft update: Regularly use the soft update method to update the policies and value network parameters in the target network, which helps maintain the stability of the learning process.
[0092] 7) Policy convergence and implementation: Repeat the training process until the policy converges, and then apply the trained policy to task segmentation and resource allocation to achieve efficient decision-making.
[0093] The following are two specific embodiments provided by the present invention and their implementation schemes:
[0094] Embodiment 1: Intelligent environmental monitoring
[0095] The present invention can be applied to the real scenario of environmental monitoring, especially in remote areas or inaccessible environments, using drones to perform tasks such as water quality monitoring and air quality detection. The specific implementation process is as follows:
[0096] 1) Network parameter initialization: According to the requirements of environmental monitoring, set key network parameters such as the number of drones and IoT devices to adapt to the environmental monitoring tasks in remote areas.
[0097] 2) Policy and environment interaction: The drones execute environmental monitoring tasks according to the instructions generated by the policy network, such as water quality sampling and air quality data collection, to achieve an automated data collection process.
[0098] 3) Data collection and processing: The environmental data collected by the drones is transmitted to the MEC server in real time and processed quickly using edge computing capabilities to analyze the environmental conditions in a timely manner.
[0099] 4) Task and resource allocation: Apply the DRL policy to intelligently optimize the management and resource allocation of environmental monitoring data, improving the efficiency and accuracy of environmental monitoring.
[0100] 5) Network update and optimization: Continuously update and optimize the policy according to the monitoring results and resource usage to adapt to the dynamic changes of the environment and ensure the effectiveness of environmental protection and ecological monitoring.
[0101] Embodiment 2: Intelligent city management
[0102] The present invention can be applied to the urban environment, using drones and edge computing technologies for real-time traffic monitoring, environmental monitoring, and public safety management, improving the efficiency and response speed of urban management.
[0103] 1) Network parameter initialization: According to the specific requirements of urban management, configure drones and IoT devices and set necessary network parameters to support real-time traffic monitoring, environmental monitoring, and public safety management.
[0104] 2) Task execution: The drone automatically executes urban monitoring tasks according to the preset policy network, including traffic flow monitoring, environmental quality detection, and public safety patrol.
[0105] 3) Data transmission and processing: The drone collects key data in the city, such as traffic conditions, pollution levels, and safety incident information, and quickly transmits it back to the MEC server for processing.
[0106] 4) Dynamic resource allocation: Use DRL strategies to dynamically optimize the allocation of urban management resources, such as adjusting traffic lights to relieve congestion, or quickly deploying police forces and emergency resources in case of emergencies.
[0107] 5) Policy iteration and update: Continuously adjust and optimize urban management policies based on real-time monitoring data and resource scheduling effects to improve the efficiency and response speed of urban management, and ensure the smooth and safe operation of the city.
[0108] Embodiment 3
[0109] Aiming at the problems existing in the prior art, this embodiment provides an air-ground network optimization method and system based on fairness guarantee of MEC. Figure 2 It is a scenario diagram to which the method of the present invention can be applied. It consists of a ground layer and an air layer. The air layer contains J MEC-supported drones that are constantly moving and providing computing services for nearby IoTDs. Let denote the set of drones. In the system, drones can directly communicate with each other through wireless links. At the bottom layer, I represents fixed IoTDs deployed in 3D space, and the IoTD set is represented as The IoTDs are scattered within a certain range, and after their positions are determined, they no longer move and have corresponding real-world scenarios. Such scenarios can be sensor nodes in remote areas, such as sensor nodes for forest fire perception, water quality monitoring, etc. Considering the need to analyze the dynamic system changes, the present invention places the system within the framework of a quasi-static network environment, divides the continuous time flow into a series of discrete segments of equal length, and calls them "time slots" or "time steps". The set of time slots is given by The duration of a time slot is Δt. The IoTDs continuously generate some tasks at the beginning of each time slot. Assuming that the IoTDs have no processing capabilities, all tasks need to be offloaded and processed using drones.
[0110] Embodiment 4
[0111] Given the dynamic changes in the network environment and the uncertainty of information acquisition, the handling of problems becomes relatively complex. To effectively address this challenge, the present invention uses a Markov Decision Process (MDP) to re-model the problem. Considering that the variables involved in the problem are continuous, the present invention selects the MADDPG algorithm, which is a deep reinforcement learning algorithm specifically designed to handle optimization problems with continuous action spaces and can support real-time decision-making.
[0112] As Figure 1 shown, the air-ground network optimization method and system for fairness guarantee based on MEC provided in this embodiment is characterized by including the following steps:
[0113] S101. Initialize the parameters of the main policy network and the critic network of the agent and the parameters of the target policy network and the critic network and the number of episodes EP, the maximum number of training steps T max , initialize the learning rates α c and α a corresponding to the critic network and the policy network, the discount factor γ, initialize the replay buffer size D, the size M of the mini-batch, and the noise ∈ for action exploration; initialize the network layout parameters, such as the number I of IoTDs, the number J of drones, and other parameters.
[0114] S102. After initializing the state of the agent, the agent interacts with the environment, and the main policy network generates corresponding actions according to the current policy.
[0115] S103. The agent operates according to the actions generated by the main policy network, then receives the rewards feedback by the environment, and updates its state according to these feedbacks.
[0116] S104. Store the experience tuples in the experience replay buffer. When the buffer reaches the capacity limit, introduce the latest experience by replacing the oldest experience data.
[0117] S105. Update the parameters of the main policy network and the critic network.
[0118] S106. Calculate the loss function of the critic network according to the TD target and the value function predicted by the critic network, extract samples from the experience replay buffer, and update the target policy network and the critic network using the gradient descent method.
[0119] S107. Use a small batch of experience samples to update the parameters of the main policy network and the critic network.
[0120] S108. Update the parameters of the policy network and the value network in the target network through the soft update mechanism.
[0121] S109. After iterative training until the algorithm converges stably, apply this policy to the agent to achieve optimal task offloading and resource allocation.
[0122] In step S102, at the beginning of each period, initialize the state of the agent. The agent interacts with the environment, and the main policy network generates actions based on the current policy. The state of the agent is represented as:
[0123] o j (t) = {u j (t - 1), φ j (t)},
[0124] where u j (t - 1) = (X j (t - 1), Y j (t - 1), H j (t - 1)), represents the three-dimensional coordinates of the UAV j at the end of time slot t - 1, corresponding to the starting point of time slot t; φ j (t) = {φ j (t)}, represents the task arrival metric of the IoT devices covered by the UAV j at the beginning of time slot t.
[0125] In step S103, the agent executes the action generated by the main policy network, obtains a reward, and updates the state. The calculation formula for the reward reward in the state update is as follows:
[0126]
[0127] In the above formula, O j (t) represents the research objective of the present invention, that is, to minimize the overall energy consumption of the IoT devices, the energy consumption of the UAVs, reduce the number of discarded tasks, and improve the fairness among the IoT devices.
[0128] In step S105, update the parameters of the main policy network and the critic network; update the current policy network through gradient ascent as follows:
[0129]
[0130] where represents the partial derivative of θ j .
[0131] In step S106: Calculate the loss function of the critic network according to the TD target and the value function predicted by the critic network, extract samples from the experience replay buffer, and update the target policy network and the evaluation network using the gradient descent method.
[0132] Furthermore, update the parameters w of the main value network and the current value network through gradient descent j as follows:
[0133]
[0134] where and q(s m(t) , a m(t) ; w j ) represent the temporal difference error and the action value function of the j-th weight at time slot t, respectively.
[0135] In step S108, a soft update mechanism is adopted to update the parameters of the policy network and the value network in the target network; the soft update formula is as follows:
[0136]
[0137] where θ j represents the parameters of the current policy network, represents the parameters of the target policy network, w j represents the parameters of the current value network, represents the parameters of the target value network, and χ ∈ [0, 1] is the parameter used to update the target network.
[0138] To elaborate on the air-ground network optimization method and system for fairness guarantee based on MEC, the present invention provides two specific application embodiments, including the key details of the implementation scheme.
[0139] To prove the creativity and technical value of the technical solution of the present invention, this part is an application embodiment of the technical solution of the claims on a specific product or related technology.
[0140] Application Embodiment 1: Intelligent Environmental Monitoring
[0141] 1) Network parameter initialization: Deploy the main policy network and the critic network in the environmental monitoring center of the city. Initialize the learning rates α c and α a corresponding to the critic network and the policy network, the discount factor γ, initialize the size D of the replay buffer, the size M of the mini-batch, and the noise ∈ for action exploration; deploy the necessary IoTD and drones to achieve environmental monitoring and data collection.
[0142] 2) Strategy and environment interaction: The drone executes environmental monitoring tasks according to the actions generated by the policy network. Noise is incorporated into the policy to explore better monitoring paths.
[0143] 3) Data collection and processing: The environmental data collected by the drone is transmitted to the MEC server for rapid data processing.
[0144] 4) Task and resource allocation: Based on the DRL policy, the processing of data and the allocation of resources are gradually optimized, such as the allocation of bandwidth resources among different devices.
[0145] 5) Network update and optimization: According to the monitoring effect and resource usage, the policy is updated regularly to ensure that the monitoring system continuously and efficiently manages and transmits data.
[0146] Application Example 2: Smart City Management
[0147] 1) Network parameter initialization: The main policy network and the critic network are quickly deployed in the urban environment, and key network parameters are initialized. The drone device is used to provide real-time urban monitoring images and videos to support urban management and emergency response.
[0148] 2) Task execution: The drone executes urban monitoring and emergency tasks according to the policy. After each task execution, immediate feedback is obtained according to the execution effect, and the system state is updated accordingly.
[0149] 3) Data transmission and processing: The urban monitoring information collected by the drone is quickly transmitted to the MEC system for processing to ensure the timely transmission and processing of information.
[0150] 4) Dynamic resource allocation: Through the DRL policy, the allocation of urban management resources is dynamically adjusted, such as the coordination of drones and ground emergency teams, to optimize the efficiency of urban management.
[0151] 5) Policy iteration and update: According to the progress and feedback of actual urban monitoring and emergency response, the policy is continuously adjusted and optimized to improve the response speed and efficiency of urban management, ensuring that urban management is more accurate and efficient.
[0152] In these two application examples, an efficient, reliable, and energy-saving joint optimization method is demonstrated for the air-ground network optimization method and system based on fairness guarantee of MEC. This method is not only applicable to smart city management but also to various scenarios such as environmental monitoring, which once again proves its wide application potential and technical advantages.
[0153] To compare and analyze the present invention with the current industry standard algorithms, we conducted a standardized performance benchmark test. This analysis revealed the significant advantages demonstrated by the present invention in key performance indicators and provided an in-depth understanding of possible performance limitations.
[0154] 1. The present invention: Represents the proposed MEC-based fairness guarantee air-ground network optimization method and system.
[0155] 2. Random access control: This algorithm differs in that the management and regulation of the communication channel between the IoTD and the drone for access are random. Other parameter settings are as before.
[0156] 3. Random task scheduling: In the context of the joint optimization and access control framework, this algorithm randomly assigns power levels to various tasks as needed. Other parameters are the same as before.
[0157] Figure 4 The performance and convergence of the three algorithms are illustrated. It is obvious that each of the three algorithms exhibits strong convergence, but the algorithm of the present invention performs better. Although the advantage in convergence speed is not very obvious, the rewards after convergence are significantly higher and more stable. In random access control, the assistance of the drone to the IoTD is randomly assigned instead of being calculated for proximity assignment, resulting in more time-consuming task offloading and thus reducing the reward value. In random task scheduling, since the task resource allocation is random compared with the other two adaptive task algorithms, the computing resources for the current task cannot be reasonably matched, thus affecting task completion and resource utilization and leading to a significant decline in rewards.
[0158] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any modification, equivalent replacement, and improvement made within the spirit and principle of the present invention by those skilled in the art within the technical scope disclosed by the present invention shall be covered by the protection scope of the present invention.
Claims
1. A fairness-guaranteed air-ground network optimization method based on MEC solves the fairness problem of existing technologies in task offloading and resource allocation by introducing an innovative soft update mechanism and dual network structure. The method is characterized by: The following steps are involved: Initialize the parameters of the main policy network and the critic network, as well as the parameters of the target policy network and the critic network. By introducing a dual network architecture, the stability and convergence of the training process are ensured, effectively avoiding the overfitting problem in the traditional single network architecture; Configure the training hyperparameters of the agent, including learning rate, discount factor, experience replay buffer size, mini-batch size, and action exploration noise. Adopt an adaptive exploration noise mechanism to dynamically adjust the action search range and improve the adaptability of the agent in complex environments; Initialize the agent state and generate actions through the main strategy network. The proposed algorithm improves the accuracy and execution efficiency of action generation. The agent interacts with the environment, updates the state according to the action and receives the reward, stores the experience tuple in the experience playback buffer, and prioritizes the storage and training of high-value experience to improve the training effect; Mini-batches of samples are extracted from the experience replay buffer and the parameters of the main policy network and the critic network are optimized using gradient descent. The parameters of the target policy network and the target value network are updated by calculating the loss function of the critic network based on the temporal difference target. A soft update mechanism is used to adjust the parameters of the policy network and value network in the target network. At the same time, a new parameter adjustment strategy is adopted in the soft update mechanism to ensure a smooth transition between the target network and the main network, thereby improving the stability and convergence speed of training; Continue iterating training until the strategy converges, and apply the training strategy to task offloading and resource allocation. Related experimental results show that the overall performance of this method in task offloading and resource allocation is improved by 20% to 50%.
2. The air-ground network optimization method based on MEC fairness guarantee according to claim 1 is characterized in that: The state of the agent at initialization includes the three-dimensional position coordinates of the drone, the task arrival index of the IoTD within the coverage area, and the time slot information.
3. The air-ground network optimization method based on MEC fairness guarantee according to claim 1 is characterized in that: After the agent executes the action generated by the main policy network, the reward calculation formula is: The reward is equal to the energy usage of the IoTD multiplied by weight parameter one, plus the energy usage of the drone multiplied by weight parameter two, minus the number of abandoned tasks multiplied by weight parameter three, plus the fairness between IoTDs multiplied by weight parameter four.
4. The air-ground network optimization method based on MEC fairness guarantee according to claim 1 is characterized in that: Update the parameters of the main policy network by: Add the partial derivative of the action value function with respect to the policy parameters to the current policy network parameters to obtain the updated policy network parameters.
5. The air-ground network optimization method based on MEC fairness guarantee according to claim 1 is characterized in that: The loss function of the critic network is calculated as: The difference between the predicted value and the time-difference target value is squared and the average is taken as the loss function value to update the network parameters.
6. The air-ground network optimization method based on MEC fairness guarantee according to claim 1 is characterized in that: The target network parameters are updated using a soft update mechanism, which includes the following methods: Multiply the parameters of the current policy network by the update coefficient, and add the parameters of the target policy network multiplied by (1 minus the update coefficient) to obtain the updated target policy network parameters. Multiply the parameters of the current value network by the update coefficient, and add the parameters of the target value network multiplied by (1 minus the update coefficient) to obtain the updated target value network parameters.
7. An air-ground network optimization system based on MEC fairness guarantee, characterized in that: include: The initialization module is used to initialize the parameters of the main policy network, critic network, target policy network, and target value network, and configure the training hyperparameters, including learning rate, discount factor, experience replay buffer size, mini-batch size, and action exploration noise; The interaction module is used for the agent to interact with the environment, generate actions through the main policy network, update the state according to the actions, and receive rewards from the environment; The experience replay module is used to store the experience tuples generated by the interaction between the agent and the environment, and adopts a fixed-capacity replay buffer management mechanism to save the latest interaction data; The optimization module extracts small batches of samples from the experience replay buffer and optimizes the parameters of the main policy network and the critic network using gradient descent. The soft update module is used to adjust the parameters of the policy network and the value network in the target network through the soft update mechanism.
8. The air-ground network optimization system with fairness guarantee based on MEC according to claim 7 is characterized in that: The interaction module further comprises: The state initialization unit is used to initialize the state of the intelligent agent, including the three-dimensional position coordinates of the drone, the task arrival index and time slot information of the IoTD within the coverage area; An action generation unit, used for generating actions based on the main strategy network; The reward calculation unit is used to calculate the reward based on the energy usage of the IoTD, the energy usage of the drone, the number of discarded tasks, and the fairness between the IoTDs.
9. The air-ground network optimization system with fairness guarantee based on MEC according to claim 7, characterized in that: The optimization module further comprises: A loss calculation unit for calculating the loss function based on the temporal difference target and the action value predicted by the critic network; A parameter updating unit, which is used to optimize the parameters of the main policy network and the critic network by gradient descent; The target network update unit is used to optimize the parameters of the target policy network and the target value network based on the gradient descent method.
10. The air-ground network optimization system with fairness guarantee based on MEC according to claim 7, characterized in that: The soft update module implements the parameter update of the target network in the following manner: Combine the parameters of the current policy network with the parameters of the target policy network in a certain ratio to obtain the updated parameters of the target policy network; The parameters of the current value network are combined with the parameters of the target value network in a certain ratio to obtain updated target value network parameters, wherein the combination ratio is determined by a preset update coefficient.