Task unloading, content caching and resource allocation optimization method in edge low-altitude system
By adopting a deep reinforcement learning method based on active reasoning in edge low-altitude systems, the free energy of the agent is calculated and its strategies are optimized, and the problems of weak generalization ability, imbalance in exploration and utilization, unreasonable resource allocation and insufficient adaptability in the existing technology are solved, and a more efficient, flexible and stable edge low-altitude system is achieved.
Patent Information
- Application Number
- CN202510074191.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art faces problems of weak generalization capabilities, imbalance in exploration and utilization, unreasonable resource allocation and insufficient algorithm adaptability in edge low-altitude systems, resulting in limited effects and scope in complex and dynamic application scenarios.
Deep reinforcement learning method based on active reasoning is adopted, and the internal generation model is closer to the real external environment by calculating the free energy of the agent, and the optimal action is obtained based on the obtained free energy to obtain the minimum value. The global network parameters are updated through the backpropagation algorithm and gradient descent method to realize the optimal strategy of the agent in the low-altitude edge system.
It significantly improves the overall efficiency and flexibility of the edge low-altitude system, optimizes the content acquisition delay and resource utilization rate, improves the system's response speed and resource allocation efficiency, and enhances the system's adaptability and stability.
Smart Images

Figure CN120075055A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of wireless communication, and particularly relates to an optimization method for joint task offloading, content caching, and resource allocation in an MEC-enabled edge low-altitude system. Background Art
[0002] Edge computing is a distributed computing architecture that moves data processing and computing tasks to edge devices closer to the data source, rather than concentrating them in a remote data center. This approach reduces data transmission latency, improves real-time performance and processing efficiency. Edge computing is particularly suitable for applications that require immediate response, such as smart cities, industrial automation, etc. By deploying computing resources at the network edge, edge computing can also relieve the burden on the central server and improve the overall performance and reliability of the system. The low-altitude economy is an emerging economic field that combines drone technology and various industrial applications. Its core concept is to expand and innovate economic activities by utilizing low-altitude space resources. This economic model can effectively solve the limitations of some traditional industries in terms of geographical space and operation efficiency, and improve the flexibility and efficiency of economic activities. At the same time, the low-altitude economy supports more diverse business scenarios and lower operating costs, which helps to enhance the economic value of application scenarios such as logistics distribution, agricultural plant protection, geographical mapping, and emergency rescue. By reasonably planning and utilizing low-altitude resources, the low-altitude economy can not only promote industrial upgrading, but also provide more extensive employment opportunities and economic growth points, meeting the needs of modern society for innovation and sustainable development.
[0003] When applying edge computing to the low-altitude economy, some challenges need to be addressed.
[0004] (1) Shortage of airspace resources: With the development of the aviation industry, airspace resources have become increasingly scarce. Low-altitude economic activities need to be carried out within a certain airspace range. For example, business operations such as drone delivery and low-altitude tourism rely on the use of low-altitude airspace. However, the existing airspace management system may have many restrictions on the use of low-altitude airspace by the low-altitude economy. For instance, in some busy urban areas or near airports, in order to ensure the safe flight of civil airliners, the opening degree of low-altitude airspace may be relatively low, which limits the scope of low-altitude economic activities.
[0005] There are also differences in the management policies for low-altitude airspace in different regions, which makes low-altitude economic enterprises face complex airspace approval processes and different regulatory requirements when operating across regions. For example, an enterprise engaged in drone logistics business in multiple cities may need to understand and apply for the airspace policies of each city in detail, which increases the operating costs and time costs and hinders the large-scale development of the low-altitude economy.
[0006] (2) Shortage of computing resources: Low-altitude economic activities generate a large amount of data, such as farmland image data collected by drones in agricultural plant protection and terrain data obtained in geographic surveying. These data need to be processed and analyzed in a timely manner to provide support for decision-making. However, the shortage of computing resources may lead to slow data processing. For example, an agricultural plant protection company uses drones to monitor pests and diseases on a large area of farmland. If the large amount of image data collected cannot be processed in a timely manner, it will not be possible to quickly and accurately determine the occurrence of pests and diseases, thereby affecting the timeliness and effectiveness of prevention and control measures. For some low-altitude economic applications with high real-time requirements, such as path planning and real-time monitoring in drone logistics, insufficient computing power may not be able to meet the needs of rapid decision-making. For example, in a complex urban environment, drones need to quickly plan the optimal path based on real-time traffic conditions and delivery task requirements. If the computing power is insufficient, the path planning time will be extended, resulting in reduced delivery efficiency and affecting the efficient operation of the low-altitude economy.
[0007] Secondly, in the scenarios involved in MEC Multi-access Edge Computing, there are often extremely complex decision-making requirements. This is because the MEC network itself is dynamic, and its network status, resource allocation, and user needs are constantly changing. This dynamicity makes it extremely difficult to make decisions in the MEC scenario, bringing many huge challenges. Deep reinforcement learning (DRL), as an emerging technical means, provides a very potential solution to solve such complex decision-making problems. The uniqueness of DRL is that it can interact directly with the environment, continuously obtain information from the environment, and learn the best offloading strategy based on this information. In practical applications, many complex problems often have huge state spaces and action spaces. For example, in an MEC scenario with many user devices, multiple application types, and different network conditions, there may be a large number of state combinations and a variety of action options. However, deep reinforcement learning, with its powerful learning ability and algorithmic mechanism, can fully approximate the optimal offloading strategy even in such a complex situation, thus providing an effective way to solve complex decision-making problems in MEC scenarios. Classic DRL algorithms include DQN, PPO, and DDPG. However,
[0008] Traditional DRL algorithms have certain limitations. In terms of generalization ability, since it relies on a fixed reward function, it is often insufficient when faced with agents with different preferences or characteristics. Active reasoning can introduce agent preferences and environmental information to enhance the generalization ability of the algorithm. On the issue of the balance between exploration and utilization, the DRL algorithm itself is difficult to achieve a good balance. Active reasoning can use related mechanisms such as the free energy principle to help the algorithm find a better balance between exploration and utilization, thereby improving the overall performance and efficiency of the algorithm.
[0009] In summary, the main technical problems faced by existing technologies in industrial applications are weak generalization ability, imbalance between exploration and utilization, unreasonable resource allocation, and insufficient algorithm adaptability. These problems limit the effectiveness and scope of existing technologies in practical applications, and new mechanisms and methods need to be introduced for improvement and optimization. Summary of the invention
[0010] In view of the problems existing in the prior art, the joint optimization of the edge low-altitude system provided by the present invention specifically involves the joint optimization of collaborative content caching, task offloading and resource allocation methods.
[0011] The present invention is implemented as follows: a method for optimizing task offloading, content caching and resource allocation in an edge low-altitude system, which utilizes a deep reinforcement learning method based on active inference to optimize the low-altitude edge network, calculates the free energy of the intelligent body to make the internal generation model of the intelligent body closer to the real external environment, and obtains the optimal action based on the minimum value of the obtained free energy, updates the global network parameters through the back propagation algorithm and the gradient descent method, and undergoes multiple iterative training until the algorithm converges, ultimately achieving the optimal strategy for the intelligent body's joint content caching, task offloading and resource allocation in the low-altitude edge system.
[0012] Further, the following steps are included:
[0013] S101, initializing agent global network parameters;
[0014] Initialize the agent's global network θ, number of rounds N ep , maximum number of iterations I, number of candidate strategies J, planning horizon H, number of users N, number of optimal candidate strategies k, initialization of the learning rate of the global network lr, discount factor γ, initialization of strategy distribution η(π), initialization of transition probability distribution δ(s t |s t-1 ,θ,π); initialize hyper parameters, such as QoE threshold length Δt, etc.;
[0015] S102, at the beginning of each round, the initial state s t is set, and J candidate strategies are obtained by random sampling of the strategy distribution;
[0016] S103. Randomly sample J actions from these policies respectively, then obtain J conditional transition probability distributions based on the J candidate policies, and calculate the corresponding current rewards from the corresponding J actions;
[0017] S104. According to the active inference and the free energy principle, use the cumulative rewards and the conditional transition probability distributions to obtain the free energies of the candidate policies;
[0018] S105. Average the first k smallest free energies, and obtain the current policy distribution based on the average value;
[0019] S106. Sample a policy based on the obtained policy distribution, and then sample an action as the action of the current agent, and interact with the environment to obtain the next state;
[0020] S107. Store the experience tuple in the replay buffer; if the replay buffer is full, overwrite the earliest stored experience with the new experience;
[0021] S108. The global network outputs the predicted value of the free energy, calculate the error function with the actual free energy, and update the global network parameters using the backpropagation algorithm and gradient descent;
[0022] S109. Repeat the training until the algorithm converges, and finally obtain the policy distribution, so that the policy at each step can be randomly sampled from the distribution, thereby controlling the actions of the agent to obtain the optimal joint content caching and resource offloading.
[0023] Furthermore, in S102, at the beginning of each episode, the initial state is set; the policy distribution is randomly sampled to obtain J candidate policies, and then J actions are randomly sampled from the J policies respectively; the state of the agent is represented as:
[0024]
[0025] Among them, is the set of local interaction information calculation amounts of the user set at time slot t; The set of local interaction information data amounts of the user set at time slot t; b(t) = {b i (t)}: The set of background environment requests of the user set at time slot t; The set of object requests of the user set in the background environment at time slot t. The action generated by sampling with the policy π t is represented as:
[0026] a(t) = {κ(t), ζ(t), B a (t), B d (t), P u,a (t), P u,d (t)},
[0027] where κ(t): the content of the background environment cached on the UAV at time slot t; ζ(t): the content of the objects in the background environment cached on the UAV at time slot t; B(t) = {B a (t), B d (t)}: the bandwidth set allocated by the UAV to the user set at time slot t, including the analog signal bandwidth and the digital signal bandwidth; P u (t) = {P u,a (t), P u,d (t)}: the transmission power set allocated by the UAV to the user set at time slot t, including the analog signal transmission power and the digital signal transmission power.
[0028] Furthermore, in step S103: based on J candidate policies, J conditional transition probability distributions are obtained, and the corresponding current rewards are calculated from the corresponding J actions. The calculation formula for the immediate reward r(t) is as follows:
[0029]
[0030] In the above formula, Q i (t) - Cost i (t) represents the reciprocal of the weighted sum of the energy consumption, delay, and virtual service operator cost of all devices in the system at time slot t. In other words, the denominator is the required objective function.
[0031] Furthermore, in step S104: according to the active inference and the free energy principle, using the cumulative reward and the conditional transition probability distribution δ(s t |s t-1 , θ, π), the free energy of J candidate policies can be obtained, and it can be obtained that:
[0032]
[0033] In step S105: the average of the first k smallest free energies is calculated, and the current policy distribution is obtained based on the average value. This process is equivalent to taking the first k maximum values of the opposite of the free energy, sorting the values from largest to smallest, and then calculating the average value:
[0034]
[0035] And the policy distribution is obtained:
[0036]
[0037] Among them, σ(·) represents a diagonal Gaussian distribution, which has the property of independent dimensions for each dimension and is determined by the mean vector and the variance vector on the diagonal. It can prevent overfitting and simplify the inference process; in reinforcement learning, it can describe state uncertainty and help the agent make reasonable decisions in an uncertain environment;
[0038] The said S106: Sample the policy π based on the obtained policy distribution δ(π) t , and then sample the action a t as the current behavior of the agent, and interact with the environment to obtain the next state s t+1 ;
[0039] Furthermore, the said S108: The global network outputs the predicted value Q(s t , a t ; θ), calculate the error function L(θ) with the target (i.e., the actual free energy) , and use the backpropagation algorithm and gradient descent to update the global network parameters; the loss function is given by the following formula:
[0040]
[0041] The gradient of the loss function is given by the following formula:
[0042]
[0043] Then update the parameters of the global network through gradient descent, as follows:
[0044]
[0045] Among them, θ is the internal parameter of the global model network, and lr is the learning rate.
[0046] Another object of the present invention is to provide a joint content caching and task offloading optimization system for an MEC-enabled edge low-altitude system, including:
[0047] A system initialization module for initializing the parameters of the deep deterministic policy gradient algorithm. This module contains two parts: one is a configuration module, whose function is to set the network learning rate, discount factor, and replay buffer size; the other is a network construction module, whose role is to define network layout parameters such as the number of environmental contents and the positions of drones;
[0048] An agent module. At the beginning of each cycle, the agent module generates actions based on the current network state. This agent module uses the active inference mechanism and the free energy principle. It fits the distribution of the policy by selecting the minimum value among the maximum of the first k free energies and calculating the mean of these minimum values;
[0049] An action execution module for executing actions of content caching and task offloading policies;
[0050] A reward acquisition module for executing actions and calculating immediate rewards. The reward acquisition module calculates rewards based on the QoE and overhead of all devices in the system;
[0051] A state transition module: After the agent executes an action, the system state is transferred from the current state to the next state;
[0052] An experience replay module for storing experience tuples of the system state, action, reward, cumulative reward, and next state for each time;
[0053] A sampling module for sampling from the distribution δ(π) of the policy and sampling from the policy π;
[0054] A global network update module for using the backpropagation algorithm. This module includes a parameter optimization unit for adjusting network parameters using the gradient descent method.
[0055] Another object of the present invention is to provide a computer device, which includes a memory and a processor. When the computer program stored in the memory is executed by the processor, the processor executes the steps of the optimization method based on active inference in the MEC-empowered edge low-altitude system.
[0056] Another object of the present invention is to provide a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the processor executes the steps of the optimization method based on active inference in the MEC-empowered edge low-altitude system.
[0057] Another object of the present invention is to provide an information data processing terminal for implementing the optimization system based on active inference in the MEC- and UAV-empowered edge system.
[0058] Combined with the above technical solutions and solved technical problems, the advantages and positive effects of the technical solution to be protected by the present invention are:
[0059] First, the present invention significantly improves the overall efficiency and flexibility of the edge low-altitude system. Under this innovative architecture, the unmanned aerial vehicle (UAV) not only serves as an edge server node of the low-altitude network but also makes up for the deficiencies of ground IoT devices in terms of coverage, especially in remote and inaccessible areas. The UAV can provide real-time caching and offloading support in these areas and cooperate with ground IoT devices to achieve data collection, preliminary processing, and rapid transmission. This edge low-altitude integration solution expands the network coverage range. Even in areas with limited network coverage, it can efficiently meet various computing requirements, optimize the data processing efficiency, and improve the overall system performance. The dynamic flexibility of the UAV enables it to adjust the deployment of computing resources according to real-time needs, achieve efficient data forwarding and task processing, reduce data transmission latency, and alleviate the bandwidth bottleneck. At the same time, the multi-access edge computing (MEC) technology makes the computing resources closer to the data source, greatly reducing the time for data to be transmitted to the remote cloud, significantly improving the data processing speed and the system response ability. This computing method close to the source not only enhances the real-time processing ability of the system but also improves the adaptability to different environmental conditions, enabling the system to maintain a high degree of stability and reliability in dynamic network demands and environments. By combining the mobility of the UAV with the efficient data processing ability of the MEC technology, the edge low-altitude network can achieve more efficient data processing, lower latency, and better resource utilization, providing a more reliable, flexible, and efficient solution for complex and dynamic application scenarios, promoting the progress of intelligent network technology, and opening up new possibilities for future network applications.
[0060] The significant technological progress specifically achieved by the present invention lies in realizing an optimization method based on active inference in an MEC-enabled edge low-altitude system, and this method has made remarkable progress in the following key aspects:
[0061] 1) Reasonable content caching strategy:
[0062] This method significantly reduces the content acquisition latency in the edge low-altitude system by optimizing the content pre-cached in the edge server. This optimization measure not only improves the overall performance of the system but also effectively reduces the content acquisition latency and the load of the remote server.
[0063] 2) Efficient task offloading and resource allocation strategy:
[0064] This method significantly improves the resource utilization rate of the edge low-altitude network by optimizing the resource allocation ratio of the UAV to ground devices and the offloading strategy of user local computing content. This optimization measure not only improves the overall performance of the system but also effectively reduces the overhead and latency.
[0065] 3) Integration of reinforcement learning:
[0066] The present invention combines the deep reinforcement learning (DRL) algorithm with the active reasoning method, so that the system no longer relies solely on a single reward when making decisions, but instead uses additional information from the environment in a comprehensive manner. The system can make optimal decisions without explicit instructions by virtue of its ability to autonomously learn and adapt based on real-time data. This adaptive ability plays a very critical role in dealing with complex and dynamically changing environments.
[0067] 4) User QoE and system overhead optimization:
[0068] The incentive mechanism included in this method focuses on increasing user QoE and reducing system overhead, which is not only energy efficient but also environmentally friendly, which is particularly important in the context of rising energy costs and increasing attention to environmental protection.
[0069] 5) Improvement of system reliability and adaptability:
[0070] By screening multiple candidate strategies and performing multiple iterations at each step, the predicted free energy value gradually approaches the true value. Therefore, the solution of using a global network to approximate this value enhances the reliability of the system. In addition, by focusing on factors other than system rewards instead of relying on a deterministic reward function, the agent's exploration space is expanded, which indirectly improves the adaptability of the system.
[0071] 6) Network’s autonomous learning and optimization capabilities:
[0072] Through continuous iterative training and experience-based adjustments to the network, the method is able to continuously improve its decision-making mechanism, thereby enhancing overall performance. The cumulative effect of these technological advances has brought significant performance improvements to edge low-altitude systems supported by MEC, and has achieved improvements in overhead, stability, and adaptability. These optimizations are critical to meeting the growing complex computing needs of the big data era.
[0073] The optimization method based on active reasoning in the MEC-enabled low-altitude edge system provided by the present invention is based on the use of mathematical models to guide the behavior and learning process of the intelligent agent. The technical effects brought by these mathematical models can be discussed based on their characteristics:
[0074] 1) Free Energy
[0075] The calculation of free energy not only focuses on user QoE and system overhead costs, but also has a factor about environmental information, called information gain, which pays extra attention to agent preferences and increases the subjective initiative of agents. By directly linking rewards to user QoE and system overhead, this method encourages agents to explore environments where weighted sums tend to be minimized, thereby achieving a win-win situation of improving user service quality and saving operator costs.
[0076] 2) Global network loss function and update
[0077] Using the approximate value of the free energy obtained through approximation calculation, the global network updates the global network by means of backpropagation and gradient descent.
[0078] Policy optimization: By continuously adjusting the global network parameters, the system can learn and adopt a more effective decision-making policy distribution, thereby sampling specific policies.
[0079] Learning stability: Based on the target calculated by the active inference mechanism, the learning process can be balanced to avoid instability caused by excessive prediction errors.
[0080] Performance optimization: By accurately calculating the loss function and updating the network, the accuracy and efficiency of the system's decision-making are improved.
[0081] 3) Parameter update formula
[0082] Describes the parameter update method of the global network.
[0083] Policy gradual approximation: By gradually updating the global model network parameters, the system can smoothly transition to a new policy, preventing performance fluctuations caused by sharp changes.
[0084] Continuous learning and adaptation: This continuous parameter update mechanism ensures that the system can adapt to long-term environmental changes.
[0085] The application of the mathematical model provided by the present invention not only improves the operation efficiency and decision-making quality of the MEC-enabled edge low-altitude system, but also enhances its adaptability to environmental changes and long-term stability. These technical effects are crucial for modern edge computing environments that handle large amounts of data and high-frequency interactions.
[0086] The optimization method based on active inference in the MEC-enabled edge low-altitude system provided by the present invention optimizes the performance of the network through the interaction between the agent and the environment.
[0087] Initialize the state of the agent. The state of the agent includes multiple variables, such as: the amount of drone cache content data and computing power, the bandwidth between the drone and the user equipment, the set of user-requested content, the drone transmission power, etc. These variables jointly define the environmental state of the agent at a specific moment, thereby affecting the agent's decision-making.
[0088] The agent initially simulates the execution of actions and obtains immediate rewards, generates actions through simulation, and thereby obtains rewards. Then, through a series of cumbersome screening processes, the optimal actions are obtained. Only at this time does the agent execute the real actions and perform state transitions. The reward calculation formula takes into account the long-term overhead of all devices in the system and the user QoE, which is the comprehensive goal of the system design, aiming to reduce the system overhead while improving the user experience.
[0089] Calculate the loss function and update the network. Use the gradient descent method to update the current global network, which can help improve the prediction accuracy of the global network for free energy, thereby optimizing the system performance. This part of the operation can be analogized to the value function update in deep reinforcement learning. The key lies in improving the prediction accuracy to guide policy improvement.
[0090] The application of these steps and mathematical models has brought significant technological progress:
[0091] Policy optimization: Through deep reinforcement learning, the system can self-learn and optimize policies to adapt to the changing network environment.
[0092] Efficient resource utilization: The optimization of content caching task offloading allocation ensures that the computing power of all devices is efficiently utilized, especially for drones with multiple functions.
[0093] Energy consumption minimization: By optimizing the long-term overhead, it helps to achieve green communication and reduce the impact on the environment.
[0094] System stability and adaptability: By adopting a unique active inference mechanism, paying attention to other additional information besides rewards, increasing the agent's attention to its own preferences, and then enhancing stability and universality.
[0095] These improvements demonstrate the potential of deep reinforcement learning based on active inference in complex network systems, especially in realizing intelligent and efficient low-altitude communication networks.
[0096] Second, the application of the present invention has achieved significant technological progress. On the one hand, through the learning and decision-making of the agent, this method can adapt to the dynamically changing environment and achieve efficient resource utilization and task execution. On the other hand, the present invention improves the stability and convergence speed of learning by introducing the experience replay mechanism and the gradient ascent and descent update strategies, thereby further enhancing the overall performance of the system. In addition, this method shows a low computational complexity when dealing with large-scale low-altitude networks, meeting the real-time requirements. Finally, by optimizing the relationship between user QoE and system overhead, the present invention realizes an effective balance in low-altitude network service deployment and resource allocation.
[0097] In summary, by introducing MEC and ADRL technologies, the present invention solves the problems existing in the traditional low-altitude network service deployment and resource allocation methods, and realizes the intelligent, efficient, and dynamic management of low-altitude network services. This method deploys edge servers on drones, considering both analog and digital signal transmissions. Considering cache decisions and drone resource allocation, it aims to maximize the QoE of users while minimizing the overhead of users. Its application not only improves resource utilization and task execution success rate, but also reduces computational complexity and meets real-time requirements, bringing significant technological progress and application value to the development of the low-altitude network industry.
[0098] Third, the MEC-empowered low-altitude edge system optimization method proposed by the present invention solves the technical problems in resource scheduling, network bandwidth, and content caching efficiency in the prior art through an innovative combination of active inference, free energy model, and joint content caching, task offloading, and resource allocation algorithms, bringing significant technological progress. Specifically, it is reflected in the following aspects:
[0099] 1. Efficient resource scheduling and optimization in a dynamic environment
[0100] Problems in the prior art: The resource scheduling of traditional MEC (Mobile Edge Computing) systems is mostly static or fixed strategies, which are difficult to meet the needs of multi-user and high-dynamic environments, resulting in low resource allocation efficiency and inability to achieve real-time response.
[0101] Technological progress: The present invention constructs a dynamic adjustment mechanism through the combination of active inference and free energy theory. The agent updates the policy distribution according to the free energy calculation in each time slot, and uses the transition probability distribution and cumulative reward for multiple iterative trainings, enabling the resource scheduling to efficiently adapt to the dynamic needs of different users. Through continuous training, the agent can dynamically and real-time complete task offloading and resource optimization in a multi-user complex environment, improving the response speed and resource allocation efficiency of the system.
[0102] 2. Precise content caching and task offloading strategies
[0103] Problems in the prior art: Existing content caching and task offloading technologies usually cannot optimize cache content and offloading strategies simultaneously, especially facing the dual challenges of limited drone resources and dynamic cache requirements in low-altitude edge computing systems.
[0104] Technological progress: In the present invention, by jointly optimizing the content caching and task offloading strategies, the agent can dynamically adjust the cached content and offloading strategy of the drone according to the user's requirements. Using random sampling of the policy distribution and the backpropagation algorithm, the system can accurately cache suitable content, maximize the user's QoE (Quality of Experience), and avoid wasting resources due to excessive caching or frequent offloading. Compared with traditional methods, this optimization strategy significantly improves the accuracy of caching and offloading and reduces the resource consumption of the system.
[0105] 3. Efficient Integration of Free Energy Model and Active Inference
[0106] Problems in the prior art: Traditional optimization methods rely on fixed or single reward functions and cannot fully consider the diversity of different user requirements and environments, resulting in the agent being difficult to adapt to complex and changing scenarios and lacking generalization ability.
[0107] Technological progress: In the present invention, by introducing the free energy model, according to the principle of active inference, combining rewards and conditional transition probability distributions, the free energy of candidate strategies is calculated. After quantifying various user requirements and environmental characteristics into free energy, the agent can select the strategy with the lowest free energy, thereby more effectively performing resource allocation and offloading control. This method enhances the generalization ability of the agent in diverse scenarios, enabling the system to achieve stable and efficient optimization under different task requirements and environments, meeting the requirements of low-altitude systems for high adaptability.
[0108] 4. Fast Convergence of the Algorithm and Precise Policy Generation
[0109] Problems in the prior art: Existing reinforcement learning algorithms have a slow convergence speed in low-altitude systems and are difficult to quickly generate efficient strategies, resulting in insufficient real-time performance of the system and being difficult to meet the requirements of dynamic environments.
[0110] Technological progress: In the present invention, by initializing the multi-dimensional parameters of the agent (such as the number of candidate strategies, planning horizon, learning rate, discount factor, etc.), combining the backpropagation algorithm and the gradient descent method, the policy distribution of the agent is repeatedly trained in the global network. As the number of training times increases, the algorithm gradually converges, and finally a stable optimal policy distribution is obtained, enabling the agent to real-time control the content caching and resource offloading at each step. By accelerating convergence, the system can quickly respond in a dynamic environment, enhancing real-time performance and reliability.
[0111] 5. Experience Replay Mechanism and Improvement of Data Utilization Efficiency
[0112] Problems in the prior art: Traditional low-altitude edge systems are difficult to make full use of experience data during the training process, resulting in low data utilization rate and low training efficiency.
[0113] Technological Progress: The present invention adopts an experience replay buffer. The agent stores the state, action, and reward data of each interaction in the buffer, avoiding the waste of old and new data. When the buffer reaches its capacity limit, the system overwrites the oldest data with the latest experience, ensuring that the stored information is the latest environmental information and interaction results. This method improves the utilization efficiency of data, makes training more efficient, and reduces data redundancy.
[0114] 6. Energy Consumption and Cost Optimization of the System
[0115] Problems in the Prior Art: The content caching and resource allocation methods of existing MEC systems mostly rely on fixed parameters and it is difficult to balance energy consumption and cost in a dynamic environment.
[0116] Technological Progress: The present invention integrates the energy consumption, latency, and service cost of the system into a weighted sum as the denominator of the immediate reward. By maximizing the reciprocal of this objective function, the agent can effectively balance resource allocation and cost consumption and achieve optimal energy efficiency. Combined with the dynamic threshold control of QoE, while meeting user requirements, this system greatly reduces resource consumption and improves the overall economic benefits of the system.
[0117] In summary, the present invention has made significant progress in aspects such as efficient resource allocation, precise content caching and task offloading, fast convergence of algorithms, and system energy consumption optimization in a multi-user dynamic environment, providing technical support for the popularization of low-altitude edge computing systems in application scenarios such as smart cities, drone swarm communication, and real-time monitoring. Description of the Drawings
[0118] Figure 1 is the flowchart of the content caching and resource allocation method in the edge low-altitude system empowered by edge computing provided by the embodiment of the present invention.
[0119] Figure 2 is a scenario diagram applicable to the embodiment of the present invention.
[0120] Figure 3 is the implementation flowchart of the joint optimization method provided by the embodiment of the present invention.
[0121] Figure 4 is the relationship diagram between the system QoE and the drone storage capacity after the convergence of the simulation algorithm provided by the embodiment of the present invention.
[0122] Figure 5 is the relationship diagram between the user energy consumption and the drone bandwidth after the convergence of the simulation algorithm provided by the embodiment of the present invention. Detailed Embodiments
[0123] To make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below in conjunction with embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0124] In order to effectively solve the problems of UAV caching decision-making and resource allocation in the low-altitude network, considering the dynamic characteristics of the network and the complexity of the QoE and overhead of users obtaining content, the problem becomes extremely challenging. The present invention models the problem as a partially observable Markov decision process. Since the variables involved (such as UAV caching decisions, resource allocation ratios, etc.) are continuous, a deep reinforcement learning algorithm based on active inference is used, which is an algorithm that combines active inference methods to solve complex decision-making problems and supports real-time online decision-making.
[0125] As an emerging technology, multi-access edge computing has been widely applied in multiple industrial fields. Especially in the low-altitude economy field, it provides effective technical support for solving UAV-related problems. The following are two specific industrial application embodiments:
[0126] The following are two specific application embodiments of the present invention in the "Optimization Method for Joint Content Caching, Task Offloading and Resource Allocation Based on Active Inference":
[0127] Application Embodiment 1: UAV-Assisted Smart Agriculture Monitoring System
[0128] In smart agriculture, UAV swarms are used to monitor the growth status of crops, soil humidity, pests and diseases in real time. The farmland has a wide coverage area and the monitoring requirements are time-sensitive. Therefore, an efficient edge computing system is needed to complete the processing, caching and resource allocation of large-scale data.
[0129] 1. System initialization: The UAVs deployed in the farmland area establish communication with the cloud server and the farmer's mobile devices. Set the QoE threshold, the transmission power of the UAV, the cache capacity and the resource allocation strategy.
[0130] 2. Policy generation and execution: The system generates several candidate policies through policy distribution sampling according to the dynamic environment of the farmland, and based on the computing resources and data cache capacity allocated by each policy to meet the real-time needs of farmers and the agricultural management system.
[0131] 3. Free energy calculation and update: The system continuously obtains crop monitoring data during operation, and converts the reward values of different tasks (such as latency, energy consumption and task completion) into reward scores through the free energy model, so as to update the optimal policy.
[0132] 4. Content Caching and Task Offloading: The drone dynamically adjusts the cached content and task offloading according to the optimal strategy, quickly uploading important monitoring data to the cloud server and caching unimportant data on the drone to ensure the effective utilization of resources.
[0133] This method effectively improves the accuracy and real-time performance of crop monitoring in smart agriculture, reduces the energy consumption of the drone, extends its endurance time, and simultaneously enhances the overall system response efficiency, providing accurate data support for farmland management.
[0134] Application Example 2: Drone-Assisted Scheduling in Intelligent Transportation Monitoring System
[0135] Application Scenario: In an intelligent transportation system, drones are used to monitor urban traffic flow, accidents, road congestion, etc. in real time. Due to the complex urban road environment and wide distribution of monitoring points, an edge computing system is required to perform reasonable content caching and resource offloading for tasks to ensure the real-time performance of traffic monitoring.
[0136] Implementation Steps:
[0137] 1. Initial Parameter Setting: The agent initializes the global network parameters, including the cache capacity, bandwidth allocation, and transmission power of the drone, and simultaneously sets the QoE threshold and key parameters for urban traffic monitoring.
[0138] 2. Candidate Strategy Generation and Selection: The drone generates several candidate strategies through random sampling in each time slot and calculates the immediate reward under each strategy (i.e., the weighted sum of traffic monitoring quality, delay, and energy consumption).
[0139] 3. Free Energy Calculation and Strategy Update: Based on active inference and the principle of free energy, the system calculates the free energy of each strategy in real time and selects the strategy with the minimum free energy to update as the current optimal strategy, continuously optimizing resource allocation.
[0140] 4. Task Offloading and Cache Optimization: According to the current strategy, the drone dynamically adjusts the cached content, uploading the monitoring videos and data of key areas to the cloud and caching ordinary monitoring data on the drone; the system flexibly allocates resources according to data requirements to reduce bandwidth and power consumption.
[0141] By jointly optimizing content caching, task offloading, and resource allocation, the task execution efficiency and monitoring data processing ability of drones in the intelligent transportation system are significantly improved, relieving the pressure on urban traffic monitoring resources, achieving fast response and efficient monitoring, and contributing to the real-time feedback and emergency scheduling of traffic accidents.
[0142] Example 1: Low-Altitude Logistics Distribution System
[0143] The low-altitude logistics distribution system utilizes multi-access edge computing and deep reinforcement learning technology based on active inference. By deploying drones and related edge computing facilities in the logistics distribution area, it realizes efficient low-altitude logistics distribution services, improving distribution efficiency and user experience.
[0144] 1. Data collection: The drones are equipped with various sensors to collect data such as their own position, flight status, environmental information, and cargo status in real time. At the same time, ground Internet of Things devices (such as intelligent shelves, package trackers, etc.) also collect relevant logistics data, such as cargo in and out information, package location information, etc.
[0145] 2. Edge processing: The edge server deployed on the drone performs preliminary processing on the collected data. For example, it analyzes real-time flight data to determine whether the flight path needs to be adjusted to avoid obstacles or bad weather; it monitors the cargo status data to ensure the safety of the cargo during transportation.
[0146] 3. Decision-making and response: The deep reinforcement learning algorithm based on active inference makes optimal decisions according to the results of edge processing and the current system state (including the positions and tasks of other drones, the usage of ground facilities, etc.). For example, it determines the best flight path of the drone, decides whether to make a temporary stop at a certain ground base station or unload some cargo, etc. At the same time, the system will timely adjust the flight parameters and task execution process of the drone according to the decision results.
[0147] 4. Cooperative optimization: Multiple drones and drones and ground facilities can cooperate with each other. Drones can share flight paths and environmental information to avoid collisions and optimize the overall distribution route; drones cooperate with ground facilities (such as the scheduling system of the logistics center, intelligent shelves, etc.) to ensure the efficient loading and unloading and accurate distribution of goods. Real-time performance: Through edge computing and the real-time data processing ability of drones, the system can quickly respond to various changes in the logistics distribution process, such as order changes, traffic congestion, etc. Low latency: Data is processed at the local drone or nearby edge computing nodes, reducing the latency of data transmission to the cloud and improving the system response speed. Privacy protection: Edge computing ensures that logistics data is processed locally, and only necessary information is shared under the premise of security, reducing the risk of data leakage and protecting user privacy and business secrets.
[0148] Example 2: Low-altitude environmental monitoring system
[0149] The low-altitude environmental monitoring system utilizes MEC and ADRL technologies. By deploying a monitoring network consisting of drones and ground monitoring stations, it realizes the efficient collection, processing, and analysis of environmental data to improve the accuracy and timeliness of environmental monitoring.
[0150] 1. Data collection: The drone is equipped with professional environmental monitoring equipment (such as gas sensors, dust sensors, meteorological instruments, etc.). When flying in the low-altitude area, it collects atmospheric environment data in real time, including air quality indicators, meteorological parameters, etc. The ground monitoring station collects relevant environmental data such as soil and water quality and transmits it to the nearby edge computing node.
[0151] 2. Edge processing: The edge computing node integrates and preliminarily processes the data collected by the drone and the ground monitoring station. For example, it performs real-time analysis on the atmospheric environment data to determine whether the air quality exceeds the standard and whether the meteorological conditions are abnormal; it conducts a preliminary assessment of the soil and water quality data to determine whether there are signs of pollution.
[0152] 3. Decision-making and response: Based on the ADRL algorithm, the system makes corresponding decisions according to the results of edge processing and the current environmental conditions. For example, if it is found that the air quality in a certain area seriously exceeds the standard, the system will decide to dispatch more drones to that area for detailed monitoring, send alarm information to the relevant environmental protection departments in a timely manner, and at the same time adjust the monitoring strategy and the flight path of the drones.
[0153] 4. Collaborative optimization: The drones collaborate with each other and with the ground monitoring station. The drones can optimize their flight paths according to the information provided by the ground monitoring station to improve the monitoring efficiency; the ground monitoring station can receive the high-altitude environmental data fed back by the drones and conduct comprehensive analysis with the data it collects itself to enhance the monitoring accuracy. Real-time performance: Through the edge computing and the real-time data processing capabilities of the drones, the system can obtain and process environmental data in a timely manner and quickly respond to environmental changes. Low latency: The data is processed at the local or nearby edge computing nodes, reducing the data transmission latency and improving the system response speed. Privacy protection: Edge computing protects the privacy of environmental monitoring data, preventing data from being leaked during transmission and processing, and ensuring the security and reliability of the monitoring data.
[0154] Example 3: Low-altitude emergency rescue system
[0155] The low-altitude emergency rescue system adopts MEC and ADRL technologies. By deploying drones and related edge computing devices in the rescue area, it realizes rapid monitoring of the rescue scene and effective rescue operations, improving the rescue efficiency and success rate.
[0156] 1. Data collection: The drone flies over the rescue scene and, through the devices such as cameras, thermal imagers, and life detectors carried by it, collects key data such as images, videos, personnel positions, and vital signs of the scene in real time. At the same time, the mobile devices carried by the ground rescue personnel also collect some on-site information and transmit it to the nearby edge computing node.
[0157] 2. Edge Processing: The edge computing node quickly processes the data collected by the drones and ground rescue personnel. For example, it analyzes image and video data to identify dangerous areas, the locations and conditions of trapped people at the rescue scene; it monitors vital sign data to judge the degree of life danger of the people.
[0158] 3. Decision-making and Response: According to the ADRL algorithm, the system makes decisions by combining the results of edge processing and the current rescue situation. For example, it determines the best rescue path for the drones and the locations for dropping rescue supplies; it commands the action directions and rescue strategies of the ground rescue personnel; it promptly sends a report on the on-site situation to the rescue command center.
[0159] 4. Collaborative Optimization: There is close collaboration between the drones and the ground rescue personnel. The drones can provide an aerial perspective and real-time information for the ground rescue personnel to help them better understand the rescue scene; the ground rescue personnel can adjust their actions based on the feedback information from the drones to improve the rescue efficiency. Real-time Performance: Through the edge computing and the real-time data processing capabilities of the drones, the system can quickly acquire and process the data at the rescue scene and rapidly respond to rescue needs. Low Latency: The data is processed at local or nearby edge computing nodes, reducing the latency of data transmission to the cloud and improving the system response speed. Privacy Protection: Edge computing ensures the privacy of rescue data, preventing the adverse effects of data leakage on rescue operations and trapped people.
[0160] These three embodiments demonstrate the specific applications of multi-access edge computing and active inference-based deep reinforcement learning technologies in the fields of low-altitude logistics distribution, low-altitude environmental monitoring, and low-altitude emergency rescue, as well as the technical advantages and values they bring. With the continuous development and maturity of technologies, these technologies will be widely applied and promoted in more low-altitude-related industrial fields.
[0161] As Figure 1 described, a method for edge computing-enabled air-ground network service deployment and resource allocation provided by an embodiment of the present invention includes the following steps:
[0162] Step 1, construct a system initialization module and an agent module, etc. The system initialization module includes a configuration module and a network construction module, and the agent module uses an active inference mechanism and the free energy principle to fit the policy distribution;
[0163] Step 2, initialize the parameters in the system initialization module, including network learning rate, discount factor, replay buffer size, and network layout parameters such as the number of background contents and the positions of the drones;
[0164] Step 3, in the agent module, generate actions based on the current network state at the beginning of each cycle;
[0165] Step 4: Execute the UAV path planning and task offloading strategy through the action execution module;
[0166] Step 5: Use the reward acquisition module to execute actions and calculate the immediate reward. Calculate the reward based on the reciprocal of the weighted average of the latency, energy consumption, and operator cost of all devices in the system. At the same time, transfer the system state from the current state to the next state;
[0167] Step 6: The experience replay module stores the experience tuples of the system state, action, reward, cumulative reward, and next state for each time;
[0168] Step 7: The sampling module samples from the distribution of the policy and samples from the policy;
[0169] Step 8: Use the parameter optimization unit in the global network update module and use the backpropagation algorithm to adjust the network parameters by the gradient descent method;
[0170] Step 9: Determine whether the network converges. If so, obtain the final optimal solution; otherwise, start from Step 3 again.
[0171] The present invention provides a method for content caching and resource allocation in an edge low-altitude system based on edge computing, which realizes the optimization of content caching and resource allocation through active inference and the free energy principle. Its working principle is as follows:
[0172] First, initialize the global network parameters of the agent. First, initialize the global network parameters of the agent, including the number of rounds, the maximum number of iteration steps, the number of candidate policies, the planning horizon, the number of users, the number of optimal candidate policies, the learning rate, and the discount factor, etc. In addition, initialize the policy distribution and the transition probability distribution, as well as hyperparameters (such as the QoE threshold length, etc.). The initialization of these parameters provides a basis for the training process of the entire system, ensuring that the agent can perform policy evaluation and optimization in subsequent operations.
[0173] Second, generate candidate policies and initialize the state. At the beginning of each round, the agent sets the initial state and randomly samples J candidate policies through the policy distribution. These candidate policies are the set of actions that the agent may take in the current environment. The candidate policies provide input for subsequent action sampling and the calculation of the conditional transition probability distribution, and at the same time enable the agent to have a certain degree of diversity and exploration in each iteration.
[0174] Third, sample actions based on the candidate policies and calculate the reward. Randomly sample the corresponding action set through the candidate policies, calculate the conditional transition probability distribution using these actions, and generate a reward value by interacting with the environment according to the current state and action. These reward values are important indicators for measuring the quality of the policy, directly reflecting the performance of the current policy in content caching and resource allocation.
[0175] Fourth, optimize the policy distribution through the free energy principle. After obtaining the cumulative rewards and conditional transition probability distributions of all candidate policies, calculate the free energy values of each candidate policy according to the free energy principle. Average the top k policies with the smallest free energy values to generate a new policy distribution. This step realizes the optimization of the policy through the principle of minimizing free energy, ensuring that the agent can gradually converge to the global optimal policy.
[0176] Fifth, environment interaction and experience storage. Based on the optimized policy distribution, the agent further samples actions to interact with the environment to obtain the next state, and at the same time stores the experience tuple in the replay buffer. If the buffer is full, the earliest stored experience is overwritten with the new experience. This experience replay mechanism can effectively avoid overfitting and improve the training efficiency and generalization ability of the model.
[0177] Sixth, update the global network parameters and iterate the training. The global network outputs the predicted value of the free energy and calculates the error function between it and the actual free energy. Update the global network parameters through the backpropagation algorithm and the gradient descent method. The entire training process is repeated in multiple episodes until the algorithm converges, and finally a policy distribution is obtained. Through this policy distribution, the agent can randomly extract the optimal policy at each step to achieve the optimization of joint content caching and resource offloading.
[0178] This method realizes the intelligent control of content caching and resource allocation in a complex environment through active inference and the free energy principle, greatly improving the quality of service and resource utilization efficiency of the edge low-altitude system.
[0179] Figure 2 is a scenario diagram of the system involved in the present invention. The system includes a drone, multiple mobile user devices, and a cloud server. An edge server is deployed on the drone to form an edge low-altitude system to provide services for user devices. Let the number and set of mobile user devices be represented as I and The present invention divides time into multiple time slots and regards each time slot as a discrete time unit for executing tasks and processing communication processes. At the beginning of each time slot, each user device creates a service request, including a content request and a local computing task. Due to the limited processing capacity of the user device, it may be necessary to offload the local computing task to the drone for processing. To optimize the QoE and overhead of the user device to obtain content, the present invention considers caching some content on the drone and dynamically adjusts the caching decision and resource allocation according to the user request and system state. In the relevant scenario, there are K different environments and objects, where the data set is represented by. Among them, C f (t + 1) is the analog signal data volume of the environment or object; D f(t + 1) is the digital signal data volume of the environment or object; it is the CPU cycles required for the processing corresponding to this data;
[0180] As Figure 3 shown, the method for collaborative content caching decision, task offloading decision and resource allocation of the present invention includes the following steps:
[0181] S101. Initialize the global network parameters of the agent;
[0182] Initialize θ of the global network of the agent, the number of rounds N ep , the maximum number of iteration steps I, the number of candidate strategies J, the planning horizon H, the number of users N, the number of optimal candidate strategies k, initialize the learning rate lr of the global network, the discount factor γ, initialize the policy distribution η(π), initialize the transition probability distribution δ(s t |s t-1 , θ, π); Initialize hyperparameters, such as the QoE threshold length Δt, etc.;
[0183] S102. At the beginning of each round, the initial state s t is set, and J candidate strategies are obtained by random sampling of the policy distribution;
[0184] S103. J actions are obtained by random sampling from these strategies respectively, and J conditional transition probability distributions are obtained based on the J candidate strategies, and the corresponding current rewards are calculated from the corresponding J actions;
[0185] S104. According to the active inference and free energy principle, the free energy of each candidate strategy is obtained by using the cumulative reward and the conditional transition probability distribution;
[0186] S105. Average the first k smallest free energies, and obtain the current policy distribution based on the average value;
[0187] S106. Sample a strategy based on the obtained policy distribution, and then sample an action as the action of the current agent, and interact with the environment to obtain the next state;
[0188] S107. Store the experience tuple in the replay buffer; if the replay buffer is full, overwrite the earliest stored experience with the new experience;
[0189] S108. The global network outputs the predicted value of the free energy, calculates the error function with the actual free energy, and updates the global network parameters using the backpropagation algorithm and gradient descent;
[0190] S109. Repeat the training until the algorithm converges, and finally obtain the policy distribution, so that the policy of each step can be randomly sampled from the distribution, thereby controlling the action of the agent and obtaining the optimal joint content caching and resource offloading.
[0191] Furthermore, at the beginning of each round in S102, the initial state is set; J candidate policies are obtained by random sampling of the policy distribution, and then J actions are obtained by random sampling of each of the J policies; the state of the agent is represented as:
[0192]
[0193] where is the set of local interaction information calculation amounts of the user set at time slot t; the set of local interaction information data amounts of the user set at time slot t; b(t) = {b i (t)}: the set of background environment requests of the user set at time slot t; the set of object requests in the background environment of the user set at time slot t. The action generated by sampling from the self-policy π t is represented as:
[0194] a(t) = {κ(t), ζ(t), B a (t), B d (t), P u,a (t), P u,d (t)},
[0195] where κ(t): the content of the background environment cached on the UAV at time slot t; ζ(t): the content of the object in the background environment cached on the UAV at time slot t; B(t) = {B a (t), B d (t)}: the set of bandwidths allocated by the UAV for the user set at time slot t, including analog signal bandwidth and digital signal bandwidth; P u (t) = {P u,a (t), P u,d (t)}: the set of transmission powers allocated by the UAV for the user set at time slot t, including analog signal transmission power and digital signal transmission power.
[0196] Furthermore, in S103: J conditional transition probability distributions are obtained based on the J candidate policies, and the corresponding current rewards are calculated from the corresponding J actions. The calculation formula of the immediate reward reward is as follows:
[0197]
[0198] In the above formula, Q i (t) - Cost i (t) represents the reciprocal of the weighted sum of the energy consumption, delay, and virtual service operator cost of all devices in the system at time slot t. In other words, the denominator is the required objective function.
[0199] Further, in step S104: According to the active inference and free energy principle, the free energy of J candidate policies is obtained by using the cumulative reward and the conditional transition probability distribution δ(s t |s t-1 , θ, π), where the calculation formula for the negative value of the free energy is:
[0200]
[0201] In step S105: The average of the first k smallest free energies is calculated, and the current policy distribution is obtained based on the average value. This process is equivalent to taking the first k maximum values of the negative value of the free energy, sorting the values from largest to smallest, and then calculating the average value:
[0202]
[0203] And the policy distribution is obtained:
[0204]
[0205] where σ(·) represents a diagonal Gaussian distribution, which has the property of independent dimensions in each dimension and is determined by the mean vector and the variance vector on the diagonal. It is commonly used as a prior or posterior distribution in machine learning to prevent overfitting and simplify the inference process; in reinforcement learning, it can describe state uncertainty and help the agent make reasonable decisions in an uncertain environment;
[0206] In step S106: Based on the obtained policy distribution δ(π), a policy π t is sampled, and then an action a t is sampled as the current behavior of the agent, and the agent interacts with the environment to obtain the next state s t+1 ;
[0207] Further, in step S108: The global network outputs the predicted value Q(s t , a t ; θ) of the free energy, and the error function L(θ) with respect to the target (i.e., the actual free energy) is calculated. The global network parameters are updated using the backpropagation algorithm and gradient descent; the loss function is given by the following formula:
[0208]
[0209] The gradient of the loss function is given by the following formula:
[0210]
[0211] Then, the parameters of the global network are updated through gradient descent as follows:
[0212]
[0213] Among them, θ is the internal parameter of the global model network, and lr is the learning rate.
[0214] In order to elaborate on the optimization method based on active reasoning in the edge low-altitude system enabled by MEC, the present invention provides three specific application embodiments, including key details of the implementation scheme.
[0215] In order to prove the creativity and technical value of the technical solution of the present invention, this section provides application examples of the claimed technical solution on specific products or related technologies.
[0216] Example 1: Low-altitude logistics distribution system
[0217] The low-altitude logistics distribution system uses multi-access edge computing and deep reinforcement learning technology based on active inference to deploy drones and related edge computing facilities in the logistics distribution area to achieve efficient low-altitude logistics distribution services and improve distribution efficiency and user experience.
[0218] 1) Network parameter initialization: Deploy the global model network in the logistics distribution center and initialize the network parameters, including learning rate, playback buffer size, etc. At the same time, determine the network layout parameters such as the number of drones and the location of ground base stations. Deploy multiple drones for logistics distribution tasks, and deploy edge servers on each drone.
[0219] 2) Strategy and environment interaction: At the beginning of each cycle, the drone generates actions based on the active reasoning mechanism and the current network status, such as determining the flight path, deciding whether to stop at a ground base station to unload cargo, etc. The drone's agent module uses the active reasoning mechanism and the free energy principle to fit the distribution of strategies by selecting the mean of the largest previous free energy minimum.
[0220] 3) Data collection and processing: Drones are equipped with various sensors to collect real-time data such as their own location, flight status, environmental information, and cargo status. The collected data is initially processed by the edge server on the drone, such as determining whether the cargo is safe and whether the flight path needs to be adjusted.
[0221] 4) Task offloading allocation: Optimize the task offloading allocation strategy based on the ADRL algorithm. Determine where to unload the goods and how to allocate computing resources based on the drone cache decision and resource allocation to maximize delivery efficiency and minimize overhead.
[0222] 5) Network update and optimization: Update network parameters based on delivery results and resource utilization. For example, adjust learning rate, cache strategy, etc. based on whether the goods are delivered on time and the energy consumption of drones, so that the delivery system can maintain an efficient operation state.
[0223] Example 2: Low-altitude environment monitoring system
[0224] The low-altitude environmental monitoring system utilizes MEC and ADRL technologies. By deploying unmanned aerial vehicles (UAVs) and ground monitoring stations to form a monitoring network, it realizes the efficient collection, processing, and analysis of environmental data, so as to improve the accuracy and timeliness of environmental monitoring.
[0225] 1) Network parameter initialization: Deploy a global model network in the environmental monitoring center, initialize relevant network parameters, and determine network layout parameters such as the number of UAVs and the locations of ground base stations within the monitoring area. An edge server is deployed on each UAV.
[0226] 2) Policy and environment interaction: The UAV generates actions based on the active inference mechanism and the current environmental state, such as determining the monitoring path and adjusting the flight altitude. Its agent module uses active inference and the free energy principle to fit the policy distribution to adapt to environmental changes.
[0227] 3) Data collection and processing: The UAV is equipped with professional environmental monitoring equipment to collect atmospheric environmental data (such as air quality indicators, meteorological parameters, etc.) and ground-related environmental data in real time (such as soil and water quality information collected by ground monitoring equipment and transmitted to nearby base stations). The edge server on the UAV performs preliminary processing on the collected data, such as determining whether the air quality exceeds the standard and whether further monitoring of a certain area is required.
[0228] 4) Task offloading and allocation: Use the ADRL algorithm to optimize the task offloading and allocation strategy. According to the environmental monitoring requirements and the resource situation of the UAVs, reasonably allocate monitoring tasks and computing resources, such as determining which areas need to be key monitored and how to allocate the computing resources of the UAVs for data processing, etc., to improve the monitoring efficiency and accuracy.
[0229] 5) Network update and optimization: Update network parameters based on the monitoring results and resource utilization. For example, if there are abnormal changes in the environmental data in a certain area, adjust the monitoring path, caching strategy, and learning rate of the UAVs, etc., so that the monitoring system can better adapt to environmental changes and maintain high-efficiency monitoring capabilities.
[0230] Example 3: Low-altitude emergency rescue system
[0231] The low-altitude emergency rescue system adopts MEC and ADRL technologies. By deploying UAVs and related edge computing devices in the rescue area, it realizes the rapid monitoring of the rescue site and effective rescue operations, improving the rescue efficiency and success rate.
[0232] 1) Network parameter initialization: Deploy a global model network in the emergency rescue command center, initialize network parameters, including the learning rate, replay buffer size, etc., and at the same time determine network layout parameters such as the number of UAVs participating in the rescue and the locations of ground base stations. An edge server is deployed on each UAV.
[0233] 2) Strategy and environment interaction: The drone generates actions according to the active inference mechanism and the current rescue scene status, such as determining the rescue path, the position of dropping rescue supplies, etc. Its agent module uses active inference and the free energy principle to fit the strategy distribution in order to make reasonable decisions in a complex rescue environment.
[0234] 3) Data collection and processing: The drone uses devices such as cameras, thermal imagers, and life detectors carried on it to collect key data such as images, videos, personnel positions, and vital signs of the rescue scene in real time. Mobile devices carried by ground rescue personnel also collect some on-site information and transmit it to nearby base stations. The edge server on the drone quickly processes the collected data, such as identifying the positions of trapped people and judging the degree of life danger of the people.
[0235] 4) Task offloading and allocation: Optimize the task offloading and allocation strategy based on the ADRL algorithm. According to the actual situation of the rescue scene and the drone resources, reasonably allocate rescue tasks and computing resources, such as determining which areas need to be rescued with priority and how to allocate the computing resources of the drone for data processing and rescue supply dropping, etc., to improve the rescue efficiency and success rate.
[0236] 5) Network update and optimization: Update network parameters according to the rescue effect and resource utilization. For example, if the rescue operation does not go smoothly, adjust the rescue path, caching strategy, and learning rate of the drone, etc., so that the rescue system can better adapt to the changes in the rescue scene and improve the rescue efficiency.
[0237] Figure 4 It is a comparison graph of the user QoE obtained by the present invention and the existing collaborative content caching and replacement decision-making, computing offloading decision-making, user resource allocation method, and transmission power control for different drone storage capacities provided by the embodiments of the present invention.
[0238] Figure 5 It is a comparison graph of the user overhead obtained by the present invention and the existing collaborative content caching and replacement decision-making, computing offloading decision-making, user resource allocation method, and transmission power control for different drone bandwidths provided by the embodiments of the present invention.
[0239] Figure 4 The storage capacity of the drone is respectively set to 0.001×10 9 bit, 1×10 9 bit, 2×10 9 bit, 3×10 9bit. When the storage capacity is very small, it is difficult to satisfy the C6 constraint in Problem P1, so the user obtains a very small QoE. After the algorithm converges, due to the increase in storage capacity, more and more environment and objects can be cached on the UAV. Therefore, the demand for downloading background information from the cloud server will decrease, which will lead to a reduction in the end-to-end total delay. Obviously, the QoE value will increase with the increase of the Π value. In addition, the performance of the proposed algorithm is significantly better than other benchmark algorithms.
[0240] Figure 5 Describes the relationship between the user overhead and the maximum bandwidth of the UAV after the algorithm converges. When the UAV bandwidth is extremely small, the bandwidth that the UAV can allocate to the user is also extremely small. As the total UAV bandwidth value increases, the bandwidth allocated to the user also increases, and the user obtains a better user experience, but at the same time has to pay a greater overhead. It can be seen from the figure that the proposed algorithm obtains relatively smaller overhead compared to the two benchmark algorithms.
[0241] It should be noted that the embodiments of the present invention can be implemented by hardware, software, or a combination of software and hardware. The hardware part can be implemented using dedicated logic; the software part can be stored in a memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated design hardware. Those of ordinary skill in the art can understand that the above devices and methods can be implemented using computer-executable instructions and / or included in processor control code, for example, such code is provided on a carrier medium such as a disk, CD, or DVD-ROM, a programmable memory such as a read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and modules of the present invention can be implemented by hardware circuits of programmable hardware devices such as very large scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, etc., or programmable logic devices such as field programmable gate arrays, or can be implemented by software executed by various types of processors, or can be implemented by a combination of the above hardware circuits and software, such as firmware.
[0242] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be covered by the protection scope of the present invention.
Claims
1. A method for optimizing task offloading, content caching and resource allocation in an edge low-altitude system, characterized in that: A deep reinforcement learning method based on active inference is used to optimize the low-altitude edge network. The free energy of the intelligent agent is calculated to make the internal generation model of the intelligent agent closer to the real external environment, and the optimal action is obtained by minimizing the free energy. The global network parameters are updated through the back propagation algorithm and the gradient descent method. After multiple iterative training until the algorithm converges, the optimal strategy of joint content caching, task offloading and resource allocation of the intelligent agent in the low-altitude edge system is finally achieved.
2. The method for optimizing task offloading, content caching and resource allocation in an edge low-altitude system as claimed in claim 1, characterized in that: The specific steps include: S101, initializing agent global network parameters; Initialize the agent's global network parameters and the number of rounds N ep , maximum number of iterations I, number of candidate strategies J, planning horizon H, number of users N, number of optimal candidate strategies k, initialization of the learning rate of the global network lr, discount factor γ, initialization of strategy distribution η(π), initialization of transition probability distribution δ(s t |s t-1 ,θ,π); initialize hyper parameters, such as QoE threshold length Δt, etc.; S102, at the beginning of each round, the initial state s t is set, and J candidate strategies are obtained by random sampling of the strategy distribution; S103, the strategies sampled by S102 are sampled to obtain J actions respectively, and then J conditional transition probability distributions are obtained, and the corresponding current rewards are calculated from the corresponding J actions; S104. Based on the free energy principle, the free energy of each candidate strategy is obtained by using the rewards and conditional transition probability distribution obtained by interacting with the environment; S105, averaging the smallest k free energies, and obtaining a strategy distribution based on the average value; S106, sampling a strategy based on the strategy distribution obtained in S105, sampling an action based on the strategy as the action of the current agent, and interacting with the environment to obtain the next state; S107, storing the experience tuple in the replay buffer; If the replay buffer is full, the new experience will overwrite the oldest experience; S108, the predicted value of the global network output free energy is calculated, the error function between the predicted value and the actual free energy is calculated, and the global network parameters are updated using the back propagation algorithm and gradient descent; S109, repeat the training until the algorithm converges, and finally obtain the strategy distribution, so that the strategy of each step can be randomly extracted from the distribution to control the action of the intelligent agent and obtain the optimal joint content caching and resource unloading.
3. The method for optimizing task offloading, content caching and resource allocation in an edge low-altitude system as claimed in claim 2, characterized in that: In S102, at the beginning of each round, the initial state is set; J candidate strategies are randomly sampled from the strategy distribution, and then J actions are randomly sampled from each of the J strategies; the state of the agent is represented as: in, is the set of local interaction information calculations of the user set in time slot t; The local interactive information data volume of the user set in time slot t; b(t) = {b i (t)}: the background environment request set of the user set at time slot t; The set of object requests from the user set in the background environment at time slot t. Self-strategy π t The sampled actions are represented as: a(t)={κ(t),ζ(t),B a (t),B d (t),P u,a (t),P u,d (t)}, Where, κ(t): the background environment content cached on the drone at time slot t; ζ(t): the object content in the background environment cached on the drone at time slot t; B(t) = {B a (t),B d P (t)}: The bandwidth set allocated by the drone to the user set in time slot t, including analog signal bandwidth and digital signal bandwidth; u (t) = {P u,a (t),P u,d (t)}: The transmission power set allocated by the drone to the user set in time slot t, including analog signal transmission power and digital signal transmission power.
4. The method for optimizing task offloading, content caching and resource allocation in an edge low-altitude system as claimed in claim 2, characterized in that: S103: Based on the J candidate strategies, J conditional transition probability distributions are obtained, and the corresponding current rewards are calculated from the corresponding J actions. The calculation formula of the instant reward is as follows: In the above formula, Q i (t)-Cost i (t) represents the inverse of the weighted sum of the energy consumption, delay and virtual service operator cost of all devices in the system at time slot t. In other words, the denominator is the required objective function.
5. The method for optimizing task offloading, content caching and resource allocation in an edge low-altitude system as claimed in claim 2, characterized in that: S104: Based on active reasoning and free energy principle, using cumulative rewards and conditional transition probability distribution δ(s t |s t-1 ,θ,π) to obtain the free energy of J candidate strategies, where the calculation formula for the opposite number of free energy is: S105: averaging the k smallest free energies, and obtaining the current strategy distribution based on the average value, is equivalent to taking the first k maximum values of the opposite number of the free energy, sorting the values from large to small, and then obtaining the average value: And get the policy distribution: σ(·) represents the diagonal Gaussian distribution, which has the property that each dimension is independent of each other. It is determined by the mean vector and the variance vector on the diagonal. It is often used as a priori or posterior distribution to prevent overfitting and simplify the inference process. In reinforcement learning, it can describe state uncertainty and help intelligent agents make reasonable decisions in uncertain environments. S106: Sampling a strategy π based on the obtained strategy distribution δ(π) t , and then sample action a t As the current agent's behavior, it interacts with the environment to get the next state; the extraction process is summarized as follows: p t ~δ(π),a t ~π t 。 6. The optimization method based on active reasoning in the MEC-enabled edge low-altitude system as claimed in claim 2, characterized in that: S108: predicted value Q(s) of global network output free energy t ,a t ;θ), and the target The error function L(θ) is used to update the global network parameters using the back propagation algorithm and gradient descent; the loss function is given by: The gradient of the loss function is given by: Then update the parameters of the global network through gradient descent as follows: Among them, θ is the internal parameter of the global model network, and lr is the learning rate.
7. A system for implementing task offloading, content caching and resource allocation optimization in an edge low-altitude system as described in any one of claims 1 to 6, characterized in that: The task offloading, content caching and resource allocation optimization system in the edge low-altitude system includes: The system initialization module is used to initialize the parameters of the deep deterministic policy gradient algorithm. This module consists of two parts: the first is the configuration module, which is used to set the network learning rate, discount factor, and playback buffer size; the second is the network construction module, which is used to define network layout parameters such as the number of environmental contents and the location of drones; Agent module: At the beginning of each cycle, the agent module generates actions based on the current network state. This agent module uses the active reasoning mechanism and the free energy principle. It selects the minimum value of the largest first k free energies and calculates the mean of these minimum values to fit the distribution of strategies. An action execution module, which is used to execute content caching and task offloading strategies; A reward acquisition module is used to perform actions and calculate instant rewards. The reward acquisition module calculates rewards based on the QoE and overhead of all devices in the system. State transfer module: After the agent performs an action, the system state is transferred from the current state to the next state; An experience replay module is used to store each system state, action, reward, cumulative reward and experience tuple of the next state; Sampling module, a module for sampling from the distribution δ(π) of the strategy and a module for sampling from the strategy π; A global network update module, a global network update module for using a back-propagation algorithm, the module comprising a parameter optimization unit for adjusting network parameters using a gradient descent method.
8. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the optimization method based on active reasoning in the MEC-enabled edge low-altitude system as described in claims 1-6.
9. A computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the processor executes the steps of the optimization method based on active reasoning in the MEC-enabled edge low-altitude system as described in claims 1-6.
10. An information data processing terminal, characterized in that: The information data processing terminal is used to implement the task offloading, content caching and resource allocation optimization system in the edge low-altitude system as described in claim 7.
Citation Information
Cited By
Wild animal image compression and encryption transmission method driven by edge calculation
CN120408327A
Dynamic environment adaptive control method and device, equipment, medium and program product
CN120803276A