Active reasoning-based optimization method and system in low-altitude system with integrated communication, inductance and calculation
By introducing an optimization method based on active reasoning in low-altitude systems, various technical challenges in low-altitude networks are solved, efficient allocation of drone resources and task optimization are achieved, and the efficiency and adaptability of the system are significantly improved.
Patent Information
- Application Number
- CN202510073933.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art faces high dynamic and weak connectivity problems, optimal service dynamic deployment problems, resource allocation and management challenges, security problems, and airspace management problems in low-altitude networks, and the DRL model has shortcomings in generalization capabilities, exploration and utilization balance, model and environment complexity, and drone coverage.
An optimization method based on active reasoning in a low-altitude system integrated synesthesia computing is proposed. By initializing the global network parameters of the intelligence, dynamically generate alternative strategies, using the active reasoning and free energy principles to calculate free energy, filter the optimal strategy distribution, and update the global network parameters through the backpropagation algorithm to achieve efficient allocation of resources and optimization of tasks.
It significantly improves the efficiency and adaptability of the system, optimizes the flight path and resource allocation of the drone, reduces energy consumption and delays, enhances the reliability and adaptability of the system, and achieves more efficient data processing and lower latency.
Smart Images

Figure CN120075054A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of wireless communication, and particularly relates to an optimization method and system based on active inference in an integrated communication, sensing and computing low-altitude system. Background Art
[0002] Edge computing (MEC) is a decentralized data processing mode that transfers data processing and analysis tasks from a centralized cloud computing center to devices at the network edge, which are closer to the data source. This mode significantly reduces latency by shortening the data transmission distance in the network, improves data processing speed and real-time performance, and is particularly suitable for application scenarios with strict requirements for response time, such as autonomous driving, industrial automation, and remote healthcare. In industry, the application of edge computing is particularly prominent. Industry emphasizes intelligent manufacturing and automation, which involve a large number of sensors and devices that generate massive amounts of data. Edge computing can perform real-time analysis at the source of data generation, quickly identify abnormal situations in the production process, and take corresponding control measures. This real-time performance is crucial for ensuring the stability and security of the production process. In industrial production, data often contains the core secrets and intellectual property rights of enterprises. By performing local processing and storage on edge devices and only uploading necessary summary data to the cloud, the risk of data exposure can be minimized, protecting the core interests of enterprises.
[0003] The integration of low-altitude networks and edge computing is an innovative network architecture that combines low-altitude communication networks and edge computing technologies to provide more efficient, secure, and intelligent services. In this architecture, low-altitude networks generally refer to communication networks within a certain height range above the ground, including aerial platforms such as unmanned aerial vehicles (UAVs) and low-earth orbit satellites, while edge computing refers to the technology of performing data processing at the network edge to reduce latency and bandwidth usage. This integration can provide more efficient, secure, and intelligent services for low-altitude intelligent networks. For example, through edge computing technology, real-time data processing and decision-making can be carried out at the network edge, improving the efficiency and real-time performance of data processing and decision-making. At the same time, Internet of Things (IoT) security technologies can provide security protection measures such as encryption, authentication, and access control for data, enhancing the security of low-altitude intelligent networks. When combining MEC with a low-altitude network architecture, several challenges need to be addressed:
[0004] (1) High dynamic and weak connection problems: Since nodes in low-altitude networks (such as UAVs) have high mobility, this leads to frequent updates of the network topology and frequent link switching, increasing the difficulty of maintaining service continuity.
[0005] (2) Optimal service dynamic deployment problem: In satellite networks with a large coverage area and diverse access users, how to achieve dynamic optimal deployment of services and optimal utilization of network resources is an important research issue.
[0006] (3) Resource allocation and management: In a drone-assisted MEC system, how to effectively ensure the privacy of terminal device data and how to improve the wireless transmission environment and enhance the quality of wireless link transmission are the current challenges.
[0007] (4) Security challenges: The MEC converged network needs to face the differences in defense strategies between different security domains and the complex network environment where edge devices are located. Traditional security defense algorithms are difficult to adapt to edge nodes with limited resources.
[0008] (5) Airspace challenges: The network topology of the space-air-ground integrated network has expanded from two dimensions to three dimensions, with time-varying and uncertain characteristics, posing new challenges to network management and operation.
[0009] Secondly, in a low-altitude communication system supporting MEC, Deep Reinforcement Learning (DRL) has been proven to be an effective optimization solution. DRL learns how to make optimal decisions in a complex network environment through the interaction between the agent and the environment. In such a system, DRL can be used to optimize the flight trajectory of drones and task offloading decisions to achieve low latency and fair computational service offloading. In addition, DRL is also used to solve non-convex optimization problems and improve network performance. Especially in a dynamically changing communication system, DRL can intelligently adjust strategies to adapt to environmental changes, improve spectrum efficiency, and meet the Quality of Service (QoS) requirements of different users. The applications of DRL in low-altitude networks also include technologies such as cooperative communication, cognitive radio, and non-orthogonal multiple access, which jointly promote the development of the space-air-ground integrated network. For example, through the multi-agent cooperative learning optimized by DRL, learning exploration can be carried out according to the relationship between the global goal and the individual goal, and high-challenge tasks can be executed, even performing excellently in a non-homogeneous system or settings without known target allocation. However, traditional DRL methods have some drawbacks:
[0010] (1) State space complexity: In the real world, the state space of many problems is very large, making it difficult for traditional reinforcement learning algorithms to handle.
[0011] (2) Data dependence: Reinforcement learning requires a large amount of trial and error to learn rules and strategies, which requires a large amount of data. In practical applications, it may be very difficult to obtain a large amount of high-quality data.
[0012] (3) Exploration-exploitation balance: Reinforcement learning needs to balance exploring the unknown environment and maximizing rewards in the known environment. The exploration strategy needs to effectively sample in the unknown environment to promote better generalization.
[0013] (4) Training stability: Due to the feedback loop characteristics of reinforcement learning, decision changes affect the next state, potentially leading to unknown results. This makes the training process vulnerable to minor changes, resulting in unstable training outcomes.
[0014] The technical problems existing in the prior art in industrial applications are mainly reflected in the following aspects:
[0015] 1. Weak generalization ability:
[0016] DRL models often perform well in the training environment but may experience a decline in performance in new environments. Improving the generalization ability of the model to enable it to adapt to different environments and tasks is one of the current research hotspots.
[0017] 2. Imbalance between exploration and exploitation:
[0018] The exploration mechanism in DRL is crucial for learning new strategies, but excessive exploration may lead to low learning efficiency. Researchers are trying to solve this problem by improving the exploration strategy, but the effects are not as expected, such as using methods like upper confidence bounds to balance exploration and exploitation.
[0019] 3. High complexity of the model and the environment:
[0020] As the complexity of the environment increases, DRL models require more complex network architectures to capture environmental features. This may make model training more difficult and increase the demand for computing resources.
[0021] 4. UAV coverage:
[0022] In the application environment of low-altitude networks, designing an efficient flight path for UAVs is crucial. However, existing technical means may not fully consider the problem of expanding the sensing range of UAVs, that is, may not fully consider how to expand the monitoring and sensing range of the surrounding environment as much as possible while ensuring that the UAV successfully completes the established tasks.
[0023] In summary, the main technical problems faced by the prior art in industrial applications are weak generalization ability, imbalance between exploration and exploitation, high complexity of the model and the environment, and UAV coverage. These problems limit the effectiveness and scope of the prior art in practical applications and new mechanisms and methods need to be introduced for improvement and optimization. Summary of the Invention
[0024] In view of the problems existing in the prior art, the present invention provides an optimization method based on active inference in a communication-sensing-computation integrated low-altitude system.
[0025] The present invention is implemented as follows. An optimization method based on active inference in a communication-sensing-computation integrated low-altitude system includes:
[0026] S101. Initialize the global network parameters of the agent;
[0027] Initialize θ of the global network of the agent, the number of rounds N, the maximum number of training steps T, the number of optimization iterations I per step, the number of alternative strategies J, the number of optimal candidate strategies k, initialize the learning rate α of the global network, the discount factor γ, initialize the size D of the replay buffer, and initialize the policy distribution Initialize the transition probability distribution Initialize hyperparameters, such as the number of drones M, the time slot length Δt, etc.; Initialize random parameters, such as the transmission power of drones The signal-to-noise ratio SNR(t) at the initial moment.
[0028] S102. Except for the need to initialize the state in the first round, at the beginning of each round, the initial state s t is determined due to state transition, and J alternative strategies are obtained by random sampling of the policy distribution;
[0029] S103. J actions are obtained by random sampling from these strategies respectively, and J conditional transition probability distributions are obtained based on the J alternative strategies, and the corresponding current rewards are calculated from the corresponding J actions;
[0030] S104. According to the active inference and the free energy principle, the free energy of each alternative strategy is obtained by using the cumulative reward and the conditional transition probability distribution;
[0031] S105. Average the free energies of the first k smallest ones, and obtain the current policy distribution based on the average value;
[0032] S106. Sample a strategy based on the obtained policy distribution, and then sample an action as the current behavior of the agent, and interact with the environment to obtain the next state;
[0033] S107. Store the experience tuple in the replay buffer; if the replay buffer is full, delete the oldest experience to store the latest experience;
[0034] S108. The global network outputs the predicted value of the free energy, calculates the error function with the actual free energy, and updates the global network parameters using the backpropagation algorithm and gradient descent;
[0035] S109. Repeat the training until the algorithm converges, and finally obtain the policy distribution, so that the policy for each step can be randomly sampled from the distribution, thereby controlling the actions of the agent and obtaining the optimal resource control allocation.
[0036] Furthermore, in S102, at the beginning of each round, the initial state is set; J alternative strategies are obtained by random sampling of the policy distribution, and J actions are obtained by random sampling from each of the J strategies; the state of the agent is represented as:
[0037] s(t) = {λ(t), W(t), Δ(t), f ECP (t), SNR(t)}
[0038] wherein, λ m (t) is the sensing rate of the UAV m at time slot t; W m (t) is the bandwidth resource occupied by the UAV m at time slot t; wherein represents the UAV trajectory; f ECP (t) is the available computing resource of the edge service platform in time slot t, and SNR(t) is the signal-to-noise ratio in time slot t; The action generated by sampling from the self-policy π t is expressed as:
[0039] a(t) = {t sens (t), P trans (t)},
[0040] wherein, represents the sensing time of the UAV cluster; represents the transmission power of the UAV cluster.
[0041] Furthermore, for the said S103: Based on J alternative policies, J conditional transition probability distributions are obtained, and the corresponding current rewards are calculated from the corresponding J actions. The calculation formula of the immediate reward r(t) is as follows:
[0042]
[0043] In the above formula, I(t) - C(t) represents the economic benefit of the low-altitude system, which is the maximization objective function.
[0044] Furthermore, for the said S104: According to the active inference and the free energy principle, using the cumulative reward and the conditional transition probability distribution the free energies of J alternative policies are obtained. The calculation formula of the opposite number of the free energy is:
[0045]
[0046] For the said S105: The average of the first k smallest free energies is taken, and the current policy distribution is obtained based on the average value. This process is equivalent to taking the first k maximum values of the opposite number of the free energy, sorting the values from largest to smallest, and then obtaining the average value:
[0047]
[0048] And the policy distribution is obtained:
[0049]
[0050] Among them, σ(·) represents a continuous distribution related to the exponent of the natural number e, such as the exponential distribution and the gamma distribution;
[0051] The step S106: Based on the obtained policy distribution sample the policy π t , and then sample the action a t as the current agent's line behavior, and interact with the environment to obtain the next state; the sampling process is summarized as follows:
[0052] π t ~q(π), a t ~π t .
[0053] Furthermore, the step S108: The global network outputs the predicted value Q(s t , a t ; θ) of the free energy, find the error function L(θ) with the target (i.e., the actual free energy) , and update the global network parameters using the backpropagation algorithm and gradient descent; the loss function is given by the following formula:
[0054]
[0055] The gradient of the loss function is given by the following formula:
[0056]
[0057] Then update the parameters of the global network through gradient descent, as follows:
[0058]
[0059] Among them, θ is the internal parameter of the global model network, and α is the learning rate.
[0060] Another object of the present invention is to propose a system, which adopts an optimization method based on active inference to support low-altitude communication of MEC (Multi-Access Edge Computing). The system consists of the following key modules:
[0061] System startup module: responsible for initiating the parameters required by the backpropagation algorithm, including the network learning rate, discount factor, and size of the replay buffer. In addition, this module is also responsible for constructing the network and defining network layout parameters such as the number of drones and the location of ground edge servers.
[0062] Agent Control Module: At the beginning of each cycle, generate corresponding actions according to the current network state. This module integrates the active inference mechanism and the free energy principle, and fits the policy distribution by selecting the mean of the first k bits with the minimum free energy.
[0063] Action Execution Module: The module responsible for implementing resource allocation.
[0064] Reward Calculation Module: After executing an action, calculate the immediate reward. This module calculates the reward based on the reciprocal of the weighted average of the latency, energy consumption, and operator cost of all devices in the system, and manages the process of the system state transitioning from the current state to the next state.
[0065] Experience Storage Module: Used to save the system state, executed actions, obtained rewards, cumulative rewards, and the next state of each operation, forming experience tuples for subsequent learning.
[0066] Sample Extraction Unit: The function of this unit is to randomly extract samples from the policy distribution and extract samples from the policy itself.
[0067] Global Network Training Module: This module is responsible for updating the global network using the backpropagation algorithm. It contains a parameter optimization component that uses the gradient descent method to fine-tune the parameters of the network.
[0068] Another object of the present invention is to provide a computer device, the computer device includes a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the optimization method based on active inference in the communication supporting MEC.
[0069] Another object of the present invention is to provide a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the processor executes the steps of the optimization method based on active inference in the communication supporting MEC.
[0070] Another object of the present invention is to provide an information data processing terminal, and the information data processing terminal is used to implement the optimization system based on active inference in the communication supporting MEC.
[0071] Combined with the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solution to be protected by the present invention are:
[0072] First, the present invention has developed an innovative system architecture that significantly enhances the system's efficiency and adaptability by integrating drone technology. In this architecture, drones act as mobile computing and communication nodes, complementing the coverage of ground IoT devices, especially in remote or hard-to-reach areas. These drones can provide immediate computing and communication support, working in tandem with ground IoT devices to achieve data collection, preliminary processing, and rapid transmission. By expanding the network coverage, this integration solution can effectively meet diverse computing needs even in areas with limited network coverage, thereby optimizing the efficiency of data processing and enhancing the overall performance of the system. The dynamic flexibility of drones allows them to adjust the configuration of computing resources according to real-time requirements to achieve efficient data transmission and task processing, reduce data transmission latency, and effectively alleviate bandwidth limitations. At the same time, MEC technology deploys computing resources closer to the data source, reducing the time for data to be transmitted to the remote cloud center, significantly improving the data processing speed and the system's response ability. This computing method close to the data source not only enhances the system's real-time processing ability but also improves the system's adaptability to different environmental conditions, ensuring the stability and reliability of the system in the face of dynamic network demands and environments. By combining the mobility of drones and the efficient data processing capabilities of MEC technology, the air-ground integrated network can achieve more efficient data processing, lower latency, and better resource utilization, providing a reliable, flexible, and efficient solution for complex and dynamic application scenarios. This integration not only promotes the development of intelligent network technology but also provides new ideas for the innovation of future network applications.
[0073] The present invention proposes an MEC-supported low-altitude communication optimization method based on active inference technology, achieving the following key technological advancements:
[0074] (1) Efficient task offloading and allocation strategy: This method significantly improves the resource utilization rate of the air-ground integrated metaverse network by optimizing the resource allocation ratio from drones to ground base stations and from base stations to edge service platforms. This strategy not only enhances the overall performance of the system but also effectively reduces energy consumption and latency.
[0075] (2) Integration of reinforcement learning: The present invention combines the DRL algorithm with active inference, enabling the system to rely not only on a single reward when making decisions but also to comprehensively utilize environmental information. The system can autonomously learn and adapt based on real-time data to achieve optimal decisions in the absence of clear instructions.
[0076] (3) Energy consumption, latency, and economic profit: The reward mechanism of this method particularly focuses on reducing the system's energy consumption and latency, maximizing economic profit, improving energy efficiency, and being environmentally friendly. This is particularly crucial in the context of rising energy costs and increased emphasis on environmental protection.
[0077] (4) Enhancement of system reliability and adaptability: By screening alternative strategies and performing multiple iterations, the predicted free energy value approaches the true value, enhancing the system's reliability. The agent expands the exploration space by focusing on factors beyond system rewards, improving the system's adaptability.
[0078] (5) Autonomous learning and optimization ability of the network: This method enables the system to continuously optimize the decision-making process and improve overall performance through iterative training and experience-based network updates. The combined effect of these technological advancements has led to a significant improvement in the performance of the MEC-supported low-altitude communication system, while also enhancing multiple aspects such as energy efficiency, stability, and adaptive ability. These improvements are crucial for meeting the complex computational requirements of the big data era.
[0079] The optimization method based on active inference in the integrated communication, sensing, and computing low-altitude system proposed by the present invention focuses on using mathematical models to guide the behavior and learning process of the system. The characteristics and technical effects of these mathematical models are as follows:
[0080] (1) Calculation of free energy: The calculation of free energy not only covers system energy consumption, latency, and operator costs but also introduces an information gain factor that takes into account environmental information and agent preferences, enhancing the agent's subjective initiative. By directly associating rewards, this method encourages the agent to explore environments where the weighted sum tends to be minimized, achieving the dual goals of improving user service quality and saving operator costs.
[0081] (2) Global network loss function and update: The global network is updated using backpropagation and gradient descent methods by approximating the free energy approximation calculated. This process continuously adjusts the global network parameters, enabling the system to learn and adopt a more effective decision strategy distribution, thereby sampling specific strategies. In addition, the objective calculation based on the active inference mechanism helps balance the learning process, avoid instability caused by excessive prediction errors, and update the network by accurately calculating the loss function, improving the accuracy and efficiency of system decision-making.
[0082] (3) Parameter update formula: Describes the parameter update method of the global network, enabling the system to smoothly transition to new strategies and prevent performance fluctuations caused by sharp changes. This continuous parameter update mechanism ensures that the system can adapt to long-term environmental changes.
[0083] The application of the mathematical model provided by the present invention not only improves the operation efficiency and decision-making quality of the MEC-supported low-altitude communication system but also enhances its adaptability to environmental changes and long-term stability. These technical effects are crucial for modern edge computing environments that handle large amounts of data and high-frequency interactions.
[0084] The present invention proposes an optimization algorithm based on active inference in low-altitude communication supporting MEC, which improves network performance through the interaction between the agent and the environment. The state initialization of the agent involves multiple parameters, including the sensing speed of the unmanned aerial vehicle, the distance from the edge server, the available computing resources, the two-dimensional position coordinates, and the transmission power, etc. These parameters together constitute the environmental state of the agent at a certain moment and affect the decision-making of the agent.
[0085] In the initial stage, the agent will simulate the execution of actions and obtain immediate rewards. After a series of screening processes to determine the optimal action, the agent will then execute the actual action and update its state. The calculation of the reward takes into account the long-term energy consumption, latency, and the cost of the operator of all devices in the system, with the aim of improving the user experience while reducing the operator's cost. The current global network is updated through the gradient descent method to improve the accuracy of the network's free energy prediction, thereby optimizing the performance of the entire system. The application of these steps and mathematical models has brought significant technological progress:
[0086] Automatic improvement of the strategy: Using deep reinforcement learning technology, the system can autonomously learn and adjust its strategy to adapt to the dynamic changes of the network environment.
[0087] Efficient allocation of resources: By optimizing the allocation strategy of task offloading, it ensures that all computing resources, especially multi-functional unmanned aerial vehicles, are utilized with the highest efficiency.
[0088] Reduction of energy consumption: Through the optimization of long-term energy consumption, it helps to promote the development of green communication and reduce the negative impact on the environment.
[0089] Stability and flexibility of the system: An innovative active inference mechanism is introduced, which not only focuses on the reward signal but also considers other relevant information. This mechanism improves the agent's awareness of its own preferences, thereby enhancing the stability of the system and its ability to adapt to different environments.
[0090] These improvements demonstrate the potential of deep reinforcement learning based on active inference in complex network systems, especially in realizing intelligent and efficient air-ground integrated communication networks.
[0091] Second, as the creative auxiliary evidence of the claims of the present invention, it is also reflected in the following important aspects:
[0092] (1) Commercial value and economic benefits:
[0093] The present invention has carefully considered the operational efficiency of metaverse network service providers and applied innovative technologies to reduce energy consumption. At the same time, the present invention also ensures excellent performance in meeting the high-standard requirements of users' sensitivity to latency. By implementing a series of carefully designed optimization measures, the present invention not only improves the system efficiency at the technical level but also effectively increases the economic benefits of operators from an economic perspective. This comprehensive improvement provides valuable reference for future patent implementation and commercial applications, indicating that operators can achieve higher economic benefits while enhancing the user experience.
[0094] (2) Breakthrough of technical problems:
[0095] Traditional DRL algorithms are usually limited to several fixed reward function design patterns, resulting in these algorithms being mainly applicable to specific scenarios. Once the requirements or scenarios change, the training results often fail to meet expectations, showing poor generalization ability. Although some studies have tried to incorporate theories such as Lyapunov optimization theory and attention mechanisms into DRL, these attempts have still failed to solve the problem of weak generalization ability. The present invention significantly improves the generalization performance of the algorithm by combining the active inference mechanism in neuroscience and using the free energy principle to guide the algorithm. By transforming the free energy into the form of the sum of cumulative rewards and indeterminates, the present invention not only retains the reward mechanism in traditional DRL but also reflects the personalized preferences and characteristics of different agents through the indeterminates, thus achieving better adaptability and generalization ability in different scenarios.
[0096] Third, the present invention proposes an innovative solution to several key problems existing in the industrial application of the prior art and achieves significant technological progress.
[0097] First, aiming at the problem of insufficient adaptability of DRL algorithms in the prior art under different environments and agent preferences, the present invention introduces the active inference mechanism and the free energy principle. This innovative method enables the algorithm to better handle agents with different characteristics, thus achieving more extensive generalization and better training effects during the training process.
[0098] Second, the present invention proposes a solution to the balance problem between exploring new strategies and exploiting existing states in traditional DRL algorithms. By adjusting the free energy of alternative strategies and selecting actions based on this free energy, the algorithm of the present invention can achieve a better balance between exploration and exploitation, thereby improving the overall performance and efficiency of the algorithm.
[0099] In addition, the present invention also optimizes the resource allocation problem in the air-ground integrated computing. By introducing the policy distribution and conditional transition probability distribution, the algorithm can allocate resources for drones and other Internet of Things devices more efficiently, which not only improves the utilization efficiency of resources but also enhances the stability of the system.
[0100] Finally, the present invention has also made significant technological progress at the industrial application level. By optimizing the flight path planning and resource offloading strategies of drones, the present invention has successfully reduced the energy consumption, latency, and cost of the system while improving the overall performance. These technological advancements provide solid technical support for the application of MEC and drone technologies in low-altitude communication systems, injecting new impetus into the development of related industries.
[0101] Fourth, the present invention effectively solves the technical problems in the optimization of drone trajectories and resource allocation in the prior art and has made significant technological progress by introducing parameterized agent strategies, active inference algorithms, and free energy mathematical models. The specific manifestations are as follows:
[0102] 1. Efficient resource allocation and policy generation in dynamic environments
[0103] Problems in the prior art: Traditional methods for optimizing drone trajectories and resource allocation usually rely on fixed strategies and cannot dynamically respond to environmental changes, resulting in poor system response speed and efficiency when network resources are limited or channel conditions fluctuate greatly.
[0104] Technological progress: By initializing the global network parameters and policy distribution of the agent, the present invention dynamically generates multiple alternative strategies and continuously adjusts the policy distribution based on the free energy optimization algorithm of active inference. Combined with the transition probability distribution, the system can quickly adapt to different environmental conditions, providing an efficient and flexible resource allocation method for drones and significantly improving the reliability of task completion.
[0105] 2. Precise drone trajectory planning and resource control
[0106] Problems in the prior art: Existing trajectory planning methods are difficult to achieve precise trajectory adjustment in high-dynamic and complex environments, and there are redundancies in resource allocation, affecting the endurance of drones and the utilization rate of system resources.
[0107] Technological progress: Through the calculation of free energy and the generation of the optimal policy distribution, the present invention can achieve precise drone trajectory planning. The agent uses the random sampling mechanism of the policy distribution to precisely control the flight actions and resource usage of the drone at each step, significantly improving the accuracy of trajectory control and the utilization rate of system resources, and enhancing the endurance time and task completion quality of the drone.
[0108] 3. The introduction of the free energy model solves various degradation fault problems
[0109] Existing technical problems: Traditional methods are difficult to balance multiple task objectives in complex scenarios and are hard to achieve stable fault simulation and performance maintenance under multi-variable optimization. Especially, they face great challenges in the degradation fault simulation of unmanned aerial vehicles and data stream transmission.
[0110] Technological progress: The free energy model of the present invention integrates the active inference algorithm and conditional transition probability, can comprehensively consider the complex dynamics of the environment and different task requirements, and realizes flexible simulation of degradation faults. In resource allocation and flight trajectory optimization, the free energy model provides higher reliability and adaptability, reducing the impact of system degradation on tasks.
[0111] 4. The parameterized intelligent agent strategy and algorithm convergence are significantly improved
[0112] Existing technical problems: Existing algorithms for low-altitude systems mostly have problems such as slow convergence speed and poor adaptability, and cannot meet the real-time requirements in parameter optimization, which affects the task execution efficiency.
[0113] Technological progress: By initializing the multi-dimensional parameters of the intelligent agent, such as policy distribution, transition probability distribution, discount factor, learning rate, etc., the intelligent agent conducts repeated training by combining the policy distribution with the conditional transition probability distribution. With the backpropagation algorithm and gradient descent method, the algorithm can converge quickly during training, achieving higher computational efficiency and ensuring the real-time performance and task adaptability of the low-altitude system in complex environments.
[0114] 5. Efficient resource offloading and dynamic interaction mechanism
[0115] Existing technical problems: Traditional resource offloading methods are difficult to dynamically adjust strategies according to task requirements and lack an efficient experience storage mechanism, resulting in resource waste and low task execution efficiency.
[0116] Technological progress: The present invention stores the interaction data of the intelligent agent through an experience replay buffer, and uses the free energy feedback and reward mechanism during the task, combined with a dynamic resource offloading allocation strategy, to achieve efficient management of unmanned aerial vehicle resources. The unmanned aerial vehicle can intelligently adjust resource consumption in different regions according to the current policy distribution and task requirements, reducing energy consumption while ensuring the task execution effect.
[0117] 6. Improvement of economic and environmental benefits
[0118] Existing technical problems: Existing low-altitude systems often have problems of high cost and low efficiency when implementing tasks. Especially in complex task scenarios, resource consumption is large and it is difficult to achieve the optimal economic benefits.
[0119] Technological progress: By introducing an objective function for maximizing economic benefits, the present invention takes into account the economy of the system in immediate rewards, and realizes the reasonable allocation of UAV resources through the optimal policy distribution, thus enhancing the economic benefits of task completion. At the same time, intelligent trajectory planning and resource control reduce unnecessary energy consumption, indirectly achieving the maximization of environmental benefits.
[0120] In summary, the present invention has made remarkable technological progress in terms of dynamic environment adaptability, fault simulation and control, resource allocation efficiency, and economic benefits, and has broad application value and prospects. Description of the Drawings
[0121] Figure 1 It is a flowchart of an optimization method based on active inference in a low-altitude system integrating communication, sensing, and computing provided by an embodiment of the present invention.
[0122] Figure 2 It is an applicable scenario diagram provided by an embodiment of the present invention.
[0123] Figure 3 It is an implementation flowchart of an optimization method based on active inference in a low-altitude system integrating communication, sensing, and computing provided by an embodiment of the present invention.
[0124] Figure 4 It is a relationship diagram provided by an embodiment of the present invention, simulating the relationship between the cumulative system reward and the operation energy consumption after the algorithm converges and the UAV sensing rate respectively.
[0125] Figure 5 It is a comparison diagram of the convergence performance with several baseline algorithms provided by an embodiment of the present invention.
[0126] Figure 6 It is a structural block diagram of an optimization system based on active inference in low-altitude communication supporting MEC provided by an embodiment of the present invention. Detailed Embodiments
[0127] In order to make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0128] The following are two specific application embodiments for the "optimization method based on active inference in a low-altitude system integrating communication, sensing, and computing":
[0129] Embodiment 1: Real-time Monitoring and Data Processing System for UAV Swarms in Smart Cities
[0130] Application scenario: In a smart city, a swarm of drones is used for real-time monitoring of urban traffic, environmental conditions, and emergencies. Due to limited network resources and complex environments in urban areas, the swarm of drones needs to achieve optimal resource allocation to maintain efficient data processing and transmission capabilities.
[0131] Implementation steps:
[0132] 1) Initialize drone parameters: Set basic parameters such as the number of drones, time slot length, transmission power, and signal-to-noise ratio.
[0133] 2) Policy and state update: In each monitoring cycle, several alternative policies are randomly generated through the policy distribution of the agent. The drones complete different monitoring tasks according to each policy, and collect data such as urban traffic, environment, and population density.
[0134] 3) Reward and free energy calculation: Calculate the immediate reward based on the resource consumption and monitoring accuracy of each policy, and perform free energy calculation in combination with the transition probability distribution to screen out the optimal policy distribution.
[0135] 4) Resource optimization: Through policy distribution sampling, control the flight trajectories and resource allocation of drones in specific areas, and achieve coordination and task division among drones to ensure the stability and timeliness of data transmission.
[0136] This method improves the monitoring efficiency of the swarm of drones in a smart city, optimizes the utilization of network bandwidth resources, and enables drones to achieve real-time and efficient data collection and transmission in complex urban environments.
[0137] Example 2: Intelligent scheduling and resource allocation of drones in disaster relief
[0138] Application scenario: After disasters such as earthquakes and floods occur, a swarm of drones is used for real-time search for survivors, delivery of relief supplies, and maintaining communication with the rescue command center. The network signal in the disaster area is unstable, and the bandwidth and computing resources are limited. Therefore, it is necessary to optimize the flight trajectories and communication resources of the drones.
[0139] Implementation steps:
[0140] 1) System initialization: Set the initial positions, transmission powers, sensing rates, and regional signal-to-noise ratios of the swarm of drones, and initialize the policy distribution and transition probability distribution.
[0141] 2) Generation of alternative policies: At the start of each flight mission, the drones randomly generate several policies and complete the reconnaissance tasks through different path and power configurations.
[0142] 3) Free energy and policy update: Based on the real-time feedback of the survivors' positions, environmental data, and network feedback, calculate the immediate rewards of each policy, and use free energy to screen out the policy distribution with the lowest energy consumption, optimizing the flight path and resource allocation of the drones.
[0143] 4) Network parameter update: Combining the feedback from the disaster area and environmental changes, the agent gradually adjusts the policy distribution to achieve efficient scheduling and resource optimization of the drone swarm in the disaster area.
[0144] This method effectively extends the endurance time of the drones in the disaster area, improves the reconnaissance and material transportation efficiency, realizes the intelligent rescue scheduling of the drone swarm in complex environments, and provides timely disaster situation data and resource scheduling support for the rescue command center.
[0145] The technical solution of the present invention proposes an optimization method for a low-altitude communication system supporting Multi-Access Edge Computing (MEC), which is based on the Active Inference and Deep Reinforcement Learning (DRL) framework. The key components of this solution include:
[0146] (1) Deep reinforcement learning architecture: Deploy a deep neural network to represent the global model, and use DRL enhanced by active inference for policy learning and decision-making.
[0147] (2) Network settings: Initialize the parameters of the global model network and configure relevant training parameters.
[0148] (3) Environmental interaction and state control: The agent first interacts with multiple candidate policies in a virtual environment, determines the actual executed action by selecting the average value of the policy with the lowest free energy, and then updates the system state.
[0149] (4) Experience storage mechanism: Use an experience replay buffer to save experience tuples, which contain cumulative rewards instead of immediate rewards, and update when the buffer is full to keep the latest learning data.
[0150] (5) Network parameter adjustment: Adjust the parameters of the global model network through the calculation of the loss function and the gradient descent method.
[0151] (6) Policy stabilization and execution: Continuously train until the policy is stable, and then apply this policy to guide the path planning and task offloading of the drones.
[0152] 1. System initialization and parameter setting
[0153] The working principle of this method starts from the initialization of the global network parameters of the agent. The system sets parameters such as the number of episodes, the maximum number of training steps, the number of optimization iterations, the policy distribution, and the transition probability distribution, laying a foundation for the optimization training of the low-altitude system. Hyperparameters and random parameters such as the transmission power of the UAV and the initial signal-to-noise ratio are set in the initialization stage to ensure that the agent can randomly explore different states and actions in the low-altitude metaverse environment, guaranteeing the diversity of the initial learning process.
[0154] 2. Generation of State and Policy Distribution
[0155] At the beginning of each episode, the initial state of the agent is determined by the state transition of the previous round. Several alternative policies are generated through random sampling of the policy distribution, and each policy corresponds to a set of actions. These policies and actions reflect state parameters such as the sensing rate of the UAV, the bandwidth resource occupancy, the trajectory, and the computing resources of the edge service platform, ensuring a high environmental adaptability at each step and providing possibilities for subsequent decision-making and resource allocation.
[0156] 3. Conditional Transition Probability and Reward Calculation
[0157] Based on each alternative policy, the agent randomly selects a set of actions, then calculates the corresponding conditional transition probability distribution according to these policies, and calculates the immediate reward using the current state information and actions. The immediate reward reflects the economic benefits of the low-altitude system. By solving the maximization of the objective function, the agent can evaluate the contribution of the current policy to the overall system benefit under different policies, thus providing a basis for the screening of better policies and the calculation of free energy.
[0158] 4. Calculation of Free Energy and Optimal Policy Distribution
[0159] The system combines the cumulative reward and the conditional transition probability distribution according to the active inference and free energy principle to calculate the free energy of each alternative policy. Subsequently, the obtained free energy values are sorted, and several policies with the smallest free energy are selected for averaging to form a new policy distribution. The update of this policy distribution is continuously optimized in subsequent episodes, enabling the agent to evolve towards the optimal resource allocation and trajectory planning direction while continuously reducing the free energy.
[0160] 5. Agent Behavior Selection and Experience Storage
[0161] The agent samples a policy based on the current policy distribution and randomly generates an action to interact with the metaverse environment to obtain the next state. At each step, the agent's experience tuple (state, action, reward, and next state) is stored in the replay buffer. The replay buffer ensures the formation of the memory mechanism. When the buffer reaches the capacity limit, the oldest experience is automatically deleted to ensure the update and effectiveness of the agent's memory.
[0162] 6. Network Parameter Update and Final Policy Convergence
[0163] During the entire training process, the global network continuously outputs the predicted value of the free energy, generates an error function by comparing it with the actually calculated free energy value, and updates the global network parameters using the backpropagation algorithm and gradient descent. Through repeated training, the network gradually converges, and finally obtains the optimal policy distribution. At this time, the agent can control its own actions based on this policy distribution to achieve the optimal trajectory planning and resource offloading allocation of the drone in the metaverse environment, improving the overall efficiency and performance of the low-altitude system.
[0164] The following are two specific embodiments provided by the present invention and their implementation schemes:
[0165] Example 1: Intelligent City Traffic Monitoring System
[0166] This system monitors urban traffic facilities through drones, collects and processes relevant data.
[0167] 1) Network configuration: Determine network settings such as the number of drones and the location of the nearest base station.
[0168] 2) Policy interaction: Use a computer or intelligent device to simulate monitoring tasks according to preset policies, select several policies with the lowest free energy for averaging to determine the actual monitoring actions of the drones, and execute corresponding movement and monitoring tasks.
[0169] 3) Data collection and analysis: The drones collect data on the surrounding environment and send it to the monitoring center for processing.
[0170] 4) Task optimization and resource management: Use deep reinforcement learning policies to optimize resource allocation and task offloading.
[0171] 5) Network adjustment and performance improvement: Adjust and optimize the policies according to the monitoring results and resource utilization efficiency to improve the overall effectiveness of the monitoring system.
[0172] Example 2: On-site Navigation System
[0173] In a tourism scenario, in order to help tourists understand the real-time road conditions and scenic tour routes, drones are used to provide auxiliary services to avoid traffic peaks and complex sections.
[0174] 1) Network settings: Develop a detailed plan for the deployment of drones in the scenic area.
[0175] 2) Task execution: The drones provide the required scenic road condition information and real-scene navigation inside the scenic area to users in real time according to preset policies.
[0176] 3) Data collection and processing: The data collected by the drone is transmitted to the base station and processed by the MEC to provide value-added services such as route recommendations.
[0177] 4) Dynamic resource management: Use deep reinforcement learning strategies to dynamically adjust resource allocation to improve the efficiency of data transmission.
[0178] 5) Continuous optimization of strategies: According to user feedback and requirements, continuously iterate and update strategies to accelerate response speed and enhance user experience.
[0179] Example 3
[0180] In response to the challenges existing in the current technology, this embodiment provides an optimization method based on active inference in a communication-sensing-computation integrated low-altitude system of the present invention. As Figure 2 shown, the method of the present invention is applicable to specific application scenarios. The air-ground integrated network (AGIN) supported by the concerned MEC consists of two main levels: air and ground. In this scenario, a series of drones are deployed by the virtual service provider to provide virtual services to users within its service scope. The drones at the air level are responsible for perceiving the surrounding environment and collecting relevant information. The ground level consists of an edge computing platform and user equipment. The edge computing platform provides services within the service scope of the ground base station, and the operator can provide economic incentives to the platform in exchange for computing resources. The drones can send the collected data to the edge computing platform through wireless communication to reduce their own computing burden. The symbol represents the set of time slots.
[0181] In the relevant scenario, at the beginning of each time slot, users everywhere send virtual service requests to the operator, and this request is forwarded to the drones. To complete the sensing task and achieve large-area coverage, the drones need to fly and move. Eventually, all the sensed data is processed by the edge computing platform, relieving the drones of one task.
[0182] Secondly, the present invention provides a low-altitude communication system supporting MEC. The system includes an air layer and a ground layer. The ground layer consists of an edge computing platform and user equipment, and the air layer is formed by a cluster of drones responsible for sensing the ground environment. In each fixed time slot, the user equipment generates tasks, which need to be processed by the edge computing platform within the limited time slot. By reasonably allocating the drone resources in the low-altitude system, the economic benefits of the entire system are maximized.
[0183] Example 4
[0184] Considering the dynamic characteristics of the network and the uncertainty of information acquisition, the problem becomes quite tricky. To effectively solve it, the problem is reconstructed as a Markov Decision Process (MDP). Due to different preferences selected by the agent, an improved reinforcement learning algorithm based on active inference is used, which supports both continuous action spaces and discrete action spaces and enables real-time online decision-making.
[0185] As Figure 1 shown, the optimization method based on active inference in the MEC-enabled low-altitude metaverse system provided in this embodiment is characterized by including the following steps:
[0186] S101. Initialize the global network parameters of the agent;
[0187] Initialize θ of the global network of the agent, the number of rounds N, the maximum number of training steps T, the number of optimization iterations I per step, the number of alternative strategies J, the number of optimal candidate strategies k, initialize the learning rate α of the global network, the discount factor γ, initialize the replay buffer size D, and initialize the policy distribution Initialize the transition probability distribution Initialize hyperparameters such as the number of drones M, the time slot length Δt, etc.; initialize random parameters such as the transmission power of the drones The signal-to-noise ratio SNR(t) at the initial moment.
[0188] S102. Except for the need to initialize the state in the first round, at the beginning of each round, the initial state s t is determined due to state transition, and J alternative strategies are randomly sampled from the policy distribution;
[0189] S103. Randomly sample J actions from these strategies respectively, then obtain J conditional transition probability distributions based on the J alternative strategies, and calculate the corresponding current rewards from the corresponding J actions;
[0190] S104. According to the active inference and the free energy principle, use the cumulative reward and the conditional transition probability distribution to obtain the free energy of each alternative strategy;
[0191] S105. Average the first k smallest free energies and obtain the current policy distribution based on the average value;
[0192] S106. Sample a strategy based on the obtained policy distribution, and then sample an action as the current behavior of the agent and interact with the environment to obtain the next state;
[0193] S107. Store the experience tuple in the replay buffer; if the replay buffer is full, delete the oldest experience to store the latest experience;
[0194] S108. Predict the value of the global network output free energy, calculate the error function with the actual free energy, and update the global network parameters using the backpropagation algorithm and gradient descent;
[0195] S109. Repeat the training until the algorithm converges, and finally obtain the policy distribution, so that the policy at each step can be randomly sampled from the distribution to control the actions of the agent and obtain the optimal resource control allocation.
[0196] S102. At the beginning of each episode, the initial state is set; the policy distribution is randomly sampled to obtain J alternative policies, and then J actions are randomly sampled from each of the J policies; the state of the agent is represented as:
[0197] s(t) = {λ(t), W(t), Δ(t), f ECP (t), SNR(t)}
[0198] where, λ m (t) is the sensing rate of UAV m at time slot t; W m (t) is the bandwidth resource occupied by UAV m at time slot t; where represents the UAV trajectory; f ECP (t) is the available computing resource of the edge service platform in time slot t, and SNR(t) is the signal-to-noise ratio in time slot t; the action generated by policy π t is represented as:
[0199] a(t) = {t sens (t), P trans (t)},
[0200] where, the sensing time of the UAV cluster ② the transmission power of the UAV cluster
[0201] S103: Based on the J alternative policies, obtain J conditional transition probability distributions, and calculate the corresponding current reward from the corresponding J actions. The calculation formula of the immediate reward r(t) is as follows:
[0202]
[0203] In the above formula, I(t) - C(t) represents the economic benefit of the low-altitude system, which is the maximization objective function.
[0204] In step S104: According to the active inference and free energy principle, use the cumulative reward and the conditional transition probability distribution The free energies of J alternative strategies are obtained, and the calculation formula for the opposite of the free energy is as follows:
[0205]
[0206] Furthermore, in S105: averaging the first k smallest free energies and obtaining the current policy distribution based on the average value, this process is equivalent to taking the first k maximum values of the opposite of the free energy, sorting the values from largest to smallest and then obtaining the average value:
[0207]
[0208] And obtain the policy distribution:
[0209]
[0210] Among them, σ(·) represents a continuous distribution related to the exponent of the natural number e, such as the exponential distribution and the gamma distribution.
[0211] In step S106: based on the obtained policy distribution sample the policy π t , and then sample the action a t as the current agent's line behavior, and interact with the environment to obtain the next state; the sampling process is summarized as follows:
[0212]
[0213] In step S108: the global network outputs the predicted value Q(s t ,a t ; θ) of the free energy, calculate the error function L(θ) with the target (i.e., the actual free energy) , and update the global network parameters using the backpropagation algorithm and gradient descent; the loss function is given by the following formula:
[0214]
[0215] The gradient of the loss function is given by the following formula:
[0216]
[0217] Then update the parameters of the global network through gradient descent, as follows:
[0218]
[0219] Among them, θ is the internal parameter of the global model network, and α is the learning rate.
[0220] To elaborate on the optimization method based on active inference in the synaesthetic computing-integrated low-altitude system in detail, the present invention provides two specific application embodiments, including the key details of the implementation scheme.
[0221] To prove the creativity and technical value of the technical solution of the present invention, this part is an application embodiment of the technical solution of the claims on a specific product or related technology.
[0222] Example Application 1: Intelligent City Traffic Monitoring Solution
[0223] 1) System startup and configuration: In the monitoring center of urban traffic management, deploy and start the global model network. Initialize the network parameters, which include setting key hyperparameters such as the learning rate α of the critic network, determining the capacity D of the replay buffer, and defining the time slot length Δt, etc. At the same time, deploy multiple drones to monitor traffic facilities and real-time situations, and use nearby base stations and edge service platforms to assist in handling monitoring tasks.
[0224] 2) Policy interaction and environmental adaptation: The drones execute monitoring and movement tasks generated based on the active inference mechanism. To improve the computational efficiency of active inference, the devices in the monitoring department need to support multi-threaded parallel processing and have powerful computing capabilities.
[0225] 3) Data collection and analysis: The drones collect urban traffic data and transmit it to nearby base stations for preliminary processing.
[0226] 4) Task allocation and offloading: Adopt a strategy based on deep reinforcement learning (DRL) to optimize data processing and task offloading allocation. Through continuous iteration of the reinforcement learning algorithm, gradually improve the efficiency and effect of task allocation.
[0227] 5) System iteration and performance improvement: Regularly update and optimize the policy according to the monitoring effect and resource utilization. This ensures that the traffic monitoring system can continuously and efficiently process and transmit data, while improving the accuracy and response speed of monitoring, and can adapt to changing traffic conditions, providing strong technical support for urban traffic management.
[0228] Implementation Application Case 2: On-site Navigation System
[0229] 1) System configuration and startup: Deploy the global model network inside the tourist scenic area and perform initialization settings for network parameters. At the same time, configure multiple drones, which can provide real-time aerial observations and 3D models to help visually display key information such as traffic conditions, complex terrains, and pedestrian flow densities inside the scenic area.
[0230] 2) Task Execution and Monitoring: The drone executes various monitoring tasks according to the preset strategy and seamlessly transitions to the next state after completing the tasks to continue the next round of tasks.
[0231] 3) Data Collection and Analysis: The perceptual data collected by the drone is transmitted to the ground base station for further processing and analysis.
[0232] 4) Dynamic Resource Management: Utilize the deep reinforcement learning (DRL) strategy to dynamically adjust resource allocation to optimize the efficiency and effectiveness of task execution.
[0233] 5) Policy Optimization and Iteration: Continuously adjust and update the policy according to the user's feedback and requirements. This helps to improve the system's response speed and further enhance the efficiency and accuracy of task execution, ensuring timely and accurate navigation services for users. Through this continuous policy iteration, the system can better adapt to the changes in user needs and scenic environment and provide more accurate and efficient on-site navigation services.
[0234] In these two embodiments, an efficient, reliable and energy-saving solution is provided for the communication system supporting MEC, which is applicable to different application scenarios, from urban real-time traffic condition monitoring to real-scene navigation, demonstrating its wide application potential and technical advantages.
[0235] To more clearly demonstrate the positive effects achieved in the research process of the embodiments of the present invention, the advantages of the embodiments of the present invention in research and development compared with the prior art will be described next.
[0236] In Figure 4 the present invention simulates the relationship between the average system reward and the drone sensing rate after the algorithm converges. It can be seen that when the drone sensing rate is extremely small, both the total system energy consumption and the reward are very small, and the corresponding economic benefits are low. As the sensing rate increases, according to the corresponding calculation formula, the corresponding amount of sensed data increases, and the energy consumption and reward of the system show an upward trend. At the same time, the proposed algorithm has the best performance in both indicators, with the largest reward and the smallest total energy consumption. It can be seen that the proposed algorithm is superior to the other three algorithms in terms of increasing the reward function. Compared with the other three algorithms, the proposed algorithm can show a higher reward and lower energy loss. It shows that the embodiments of the present invention have achieved some positive effects during the simulation use process and indeed have great advantages compared with the prior art.
[0237] In Figure 5shows the convergence performance of all algorithms with reward as the performance metric under default parameter settings. The proposed algorithm converges around 60 episodes, and both baseline algorithms converge after 100 episodes, indicating that the proposed algorithm exhibits fast convergence when solving complex optimization problems. In addition, it can be noted that the algorithm with random sensing time has the strongest volatility, which is due to the fact that this algorithm increases the random variables in the environment, enhancing the uncertainty of exploration. Although both baseline algorithms converge, their strong volatility and slow convergence characteristics result in poor comprehensive performance. Specifically, compared with several benchmark algorithms, our algorithm always achieves higher performance metrics and exhibits less volatility. It is worth noting that some comparison algorithms only modify the agent without imposing strict decision (i.e., action) constraints from different perspectives to cater to unique preferences. The stable convergence of several algorithms observed in Figure 5 emphasizes the ability of our algorithm to adapt to agents with different preferences and highlights its superior generalization ability. Therefore, the proposed algorithm has strong robustness and stability.
[0238] As Figure 6 shown, another object of the present invention is to provide an active inference-based optimization system for MEC-supported low-altitude communication, including:
[0239] System startup module: responsible for initiating the parameters required for the backpropagation algorithm, including the network learning rate, discount factor, and size of the replay buffer. In addition, this module is also responsible for constructing the network and defining network layout parameters such as the number of drones and the location of ground edge servers.
[0240] Agent control module: at the beginning of each cycle, generate corresponding actions according to the current network state. This module integrates the active inference mechanism and the free energy principle, and fits the policy distribution by selecting the mean of the top k values of the free energy minimum.
[0241] Action implementation module: responsible for the module that implements resource allocation.
[0242] Reward calculation module: calculate the immediate reward after executing the action. This module calculates the reward based on the reciprocal of the weighted average of the delay, energy consumption, and operator cost of all devices in the system, and manages the process of transferring the system state from the current state to the next state.
[0243] Experience storage module: used to save the system state, executed actions, obtained rewards, cumulative rewards, and the next state of each operation, forming experience tuples for subsequent learning.
[0244] Sample extraction unit: the function of this unit is to randomly extract samples from the policy distribution and extract samples from the policy itself.
[0245] Global network training module: This module is responsible for updating the global network using the backpropagation algorithm. It contains a parameter optimization component that utilizes the gradient descent method to fine-tune the network's parameters.
[0246] Another object of the present invention is to provide a computer device, the computer device includes a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the optimization method based on active inference in the MEC-enabled low-altitude metaverse system.
[0247] Another object of the present invention is to provide a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the processor executes the steps of the optimization method based on active inference in the integrated communication and sensing low-altitude system.
[0248] Another object of the present invention is to provide an information data processing terminal for implementing the optimization system based on active inference in the MEC-supported low-altitude communication.
[0249] Example Application 1: Intelligent City Traffic Monitoring Solution
[0250] 1) System startup and configuration: At the monitoring center of urban traffic management, deploy and start the global model network. Initialize the network parameters, which include setting key hyperparameters such as the learning rate α of the critic network, determining the capacity D of the replay buffer, and defining the time slot length Δt. At the same time, deploy multiple drones to monitor traffic facilities and real-time situations, and utilize nearby base stations and edge service platforms to assist in processing monitoring tasks.
[0251] 2) Policy interaction and environmental adaptation: The drones execute monitoring and movement tasks generated based on the active inference mechanism. To improve the computational efficiency of active inference, the devices in the monitoring department need to support multi-threaded parallel processing and have powerful computing capabilities.
[0252] 3) Data collection and analysis: The drones collect urban traffic data and transmit it to nearby base stations for preliminary processing.
[0253] 4) Task allocation and offloading: Adopt a strategy based on deep reinforcement learning (DRL) to optimize data processing and task offloading allocation. Through continuous iteration of the reinforcement learning algorithm, gradually improve the efficiency and effectiveness of task allocation.
[0254] 5) System Iteration and Performance Improvement: Regularly update and optimize the strategy based on the monitoring results and resource utilization. This ensures that the traffic monitoring system can continuously and efficiently process and transmit data, while improving the accuracy and response speed of monitoring, adapting to changing traffic conditions, and providing strong technical support for urban traffic management.
[0255] Implementation Application Case 2: On-site Navigation System
[0256] 1) System Configuration and Startup: Deploy the global model network within the tourist scenic area and perform initialization settings for network parameters. Meanwhile, configure multiple drones that can provide real-time aerial observations and 3D models to help visually display key information such as traffic conditions, complex terrains, and pedestrian flow densities within the scenic area.
[0257] 2) Task Execution and Monitoring: The drones execute various monitoring tasks according to the preset strategy and seamlessly transition to the next state after completing the tasks to continue executing the next round of tasks.
[0258] 3) Data Collection and Analysis: The perceptual data collected by the drones is transmitted to the ground base station for further processing and analysis.
[0259] 4) Dynamic Resource Management: Utilize the deep reinforcement learning (DRL) strategy to dynamically adjust resource allocation to optimize the efficiency and effectiveness of task execution.
[0260] 5) Strategy Optimization and Iteration: Continuously adjust and update the strategy according to user feedback and requirements. This helps improve the system's response speed and further enhance the efficiency and accuracy of task execution, ensuring timely and accurate navigation services for users. Through this continuous strategy iteration, the system can better adapt to changes in user needs and scenic area environments, providing more accurate and efficient on-site navigation services.
[0261] In these two embodiments, an efficient, reliable, and energy-saving solution is provided for the communication system supporting MEC, applicable to different application scenarios, from urban real-time traffic condition monitoring to on-site navigation, demonstrating its broad application potential and technical advantages.
[0262] In Figure 4In this invention, the relationship between the average system reward after algorithm convergence and the UAV sensing rate is simulated. It can be seen that when the UAV sensing rate is extremely small, both the total system energy consumption and the reward are very small, and the corresponding economic benefits are relatively low. As the sensing rate increases, according to the corresponding calculation formula, the corresponding amount of sensed data increases, and the energy consumption and reward of the system show an upward trend. At the same time, the proposed algorithm has the best performance in both metrics, with the largest reward and the smallest total energy consumption. It can be seen that the proposed algorithm is superior to the other three algorithms in terms of increasing the reward function. Compared with the other three algorithms, the proposed algorithm can show a higher reward and lower energy loss. This demonstrates that some positive effects have been achieved during the simulation of the embodiments of the present invention, and it indeed has great advantages compared with the prior art.
[0263] In Figure 5 it shows the convergence performance of all algorithms with the reward as the performance metric under the default parameter settings. The proposed algorithm converges in about 60 rounds, and both of the two baseline algorithms converge after 100 rounds. This indicates that when solving complex optimization problems, the proposed algorithm exhibits the characteristic of fast convergence. In addition, it can be noted that the random sensing time algorithm has the strongest volatility, which is due to the fact that this algorithm adds random variables to the environment, increasing the uncertainty of exploration. Although both of the two baseline algorithms converge, their strong volatility and slow convergence characteristics result in poor comprehensive performance. Specifically, compared with several benchmark algorithms, our algorithm always achieves higher performance metrics and shows less volatility. It is worth noting that some comparison algorithms only modify the agent without imposing strict decision (i.e., action) constraints from different perspectives to cater to unique preferences. In Figure 5 the stable convergence of several algorithms observed emphasizes the ability of our algorithm to adapt to agents with different preferences and highlights its superior generalization ability. Therefore, the proposed algorithm has strong robustness and stability.
[0264] It should be noted that the embodiments of the present invention can be implemented by hardware, software, or a combination of software and hardware. The hardware part can be implemented using dedicated logic; the software part can be stored in a memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated designed hardware. Those of ordinary skill in the art can understand that the above devices and methods can be implemented using computer-executable instructions and / or included in processor control code, for example, such code is provided on a carrier medium such as a disk, CD, or DVD-ROM, a programmable memory such as a read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and their modules of the present invention can be implemented by hardware circuits of programmable hardware devices such as very large scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, etc., or field programmable gate arrays, programmable logic devices, etc., can also be implemented by software executed by various types of processors, or can be implemented by a combination of the above hardware circuits and software such as firmware.
[0265] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention should be covered by the protection scope of the present invention.
Claims
1. An optimization method based on active reasoning in a low-altitude system integrating synaesthesia and computing, characterized in that: The method comprises the following steps: 1) Initialize global network parameters and hyperparameters: Initialize the agent's global network parameters and optimize hyperparameters to build a multi-dimensional action space suitable for low-altitude systems; 2) Generate initial policy distribution and trajectory: Based on the strategy distribution and transition probability distribution, the multi-dimensional action space is randomly sampled to generate the initial trajectory and resource allocation strategy; 3) Evaluate the cumulative reward value: Using deep reinforcement learning methods to evaluate the cumulative reward value of each candidate strategy, the cumulative reward is an economic benefit indicator of the low-altitude system; 4) Free energy value calculation and strategy update: Based on the free energy principle, the free energy value of each strategy is calculated; The free energy values are averaged to generate a new strategy distribution; The gradient of global network parameters is calculated using the back-propagation algorithm, and the global network parameters are updated using the gradient descent method with free energy reduction. 5) Iterative optimization: In a complex dynamic environment, the adaptability and performance of the strategy distribution are optimized through multiple iterative training until the free energy value converges to a stable state or the cumulative reward meets the optimization goal; 6) Dynamically adapt to environmental changes: During the training process, it dynamically adapts to scenario changes including changes in mission requirements, adjustments to drone node distribution, and fluctuations in computing resource allocation to ensure the stability of strategy distribution and optimize performance; 7) Output optimization strategy distribution: Based on the optimized strategy distribution, the optimal allocation of resources for mission efficiency, energy utilization and delay cost in the low-altitude system is achieved.
2. The optimization method based on active reasoning in the low-altitude system integrating synaesthesia and computing as claimed in claim 1 is characterized in that: The following steps are involved: S101, initializing agent global network parameters; Initialize the agent's global network θ, the number of rounds N, the maximum number of training steps T, the number of optimization iterations for each step I, the number of alternative strategies J, the number of optimal candidate strategies k, the initialization of the global network's learning rate α, the discount factor γ, the initialization of the replay buffer size D, the initialization of the strategy distribution q(π), the initialization of the transition probability distribution p(s τ |s τ-1 ,θ,π); initialize hyper parameters, such as the number of drones M, time slot length Δt, etc.; Initialize random parameters, such as the UAV transmission power Signal-to-noise ratio SNR(t) at the initial moment. S102, except for the first round, the initial state s t Determined by state transition, the strategy distribution is randomly sampled to obtain J alternative strategies; S103, randomly sampling J actions from these strategies respectively, and then obtaining J conditional transition probability distributions based on the J alternative strategies, and calculating the corresponding current rewards from the corresponding J actions; S104. According to active reasoning and free energy principle, the free energy of each alternative strategy is obtained using cumulative rewards and conditional transition probability distribution; S105, averaging the first k smallest free energies, and obtaining the current strategy distribution based on the average value; S106, sampling a strategy based on the obtained strategy distribution, and then sampling an action as the current agent's behavior, and interacting with the environment to obtain the next state; S107, storing the experience tuple in the replay buffer; If the replay buffer is full, the oldest experience is deleted to store the latest experience; S108, the predicted value of the global network output free energy is calculated, the error function between the predicted value and the actual free energy is calculated, and the global network parameters are updated using the back propagation algorithm and gradient descent; S109, repeat the training until the algorithm converges, and finally obtain the strategy distribution, so that the strategy of each step can be randomly extracted from the distribution to control the action of the intelligent agent and obtain the optimal resource control allocation.
3. The optimization method based on active reasoning in the low-altitude system integrating synaesthesia and computing as claimed in claim 2 is characterized in that: In S102, at the beginning of each round, the initial state is set; J candidate strategies are randomly sampled from the strategy distribution, and then J actions are randomly sampled from each of the J strategies; the state of the agent is represented as: s(t)={λ(t),W(t),Δ(t),f ECP (t),SNR(t)} in, is the sensing rate of UAV m in time slot t; is the bandwidth resource occupied by UAV m in time slot t; Where Δ(t) = {Δ m (t)}={ΔX m (t),ΔY m (t),ΔH m (t)}, represents the trajectory of the drone; f ECP (t) is the available computing resources of the edge service platform in time slot t, SNR(t) is the signal-to-noise ratio in time slot t; strategy π t The sampled actions are represented as: a(t)={t sens (t),P trans (t)}, Among them, drone cluster perception time ② UAV cluster transmission power 4. The optimization method based on active reasoning in the low-altitude system integrating synaesthesia and computing as claimed in claim 2 is characterized in that: S103: Based on the J candidate strategies, J conditional transition probability distributions are obtained, and the corresponding current rewards are calculated from the corresponding J actions. The calculation formula of the immediate reward r(t) is as follows: In the above formula, I(t)-C(t) represents the economic benefit of the low-altitude system, which is the maximization objective function.
5. The optimization method based on active reasoning in the low-altitude system integrating synaesthesia and computing as claimed in claim 2 is characterized in that: S104: Based on active reasoning and free energy principle, using cumulative rewards And the conditional transition probability distribution The free energy of J alternative strategies is obtained, where the calculation formula for the opposite number of free energy is: S105: averaging the first k smallest free energies, and obtaining the current strategy distribution based on the average value, this process is equivalent to taking the first k maximum values of the opposite number of the free energy, sorting the values from large to small, and obtaining the average value: And get the policy distribution: Here, σ(·) represents a continuous distribution related to the exponential of the natural number e, such as the exponential distribution and the gamma distribution; S106: Based on the obtained strategy distribution Sampling strategy π t , and then sample the action at as the current agent's behavior, and interact with the environment to get the next state; the extraction process is summarized as follows: p t ~q(π),a t ~π t 。 6. The optimization method based on active reasoning in the low-altitude system integrating synaesthesia and computing as claimed in claim 2 is characterized in that: S108: predicted value Q(s) of global network output free energy t ,a t ;θ), and the target The error function L(θ) is used to update the global network parameters using the back propagation algorithm and gradient descent; the loss function is given by: The gradient of the loss function is given by: Then update the parameters of the global network through gradient descent as follows: Among them, θ is the internal parameter of the global model network, and α is the learning rate.
7. A low-altitude system with integrated synergy and computing that implements the optimization method based on active reasoning in the low-altitude system with integrated synergy and computing as described in any one of claims 1 to 6, characterized in that: The optimization system based on active reasoning in the low-altitude system supporting the integration of synaesthesia and computing includes: System initialization module, which is used to initialize the parameters of the deep deterministic policy gradient algorithm. It includes a configuration module for setting the network learning rate, discount factor, and playback buffer size, and a network construction module for defining network layout parameters such as the number of drones and the location of ground base stations; The agent module is used to generate actions based on the current network state at the beginning of each cycle. The agent module uses the active reasoning mechanism and the free energy principle to fit the distribution of strategies by selecting the mean of the largest top k free energy minima; Action execution module, which is used to execute the path planning and task offloading strategies of UAV; A reward acquisition module for executing actions and calculating instant rewards, which calculates rewards based on the inverse of the weighted average of latency, energy consumption, and operator costs of all devices in the system, and a state transition module for transferring the system state from the current state to the next state; An experience replay module is used to store each system state, action, reward, cumulative reward and experience tuple of the next state; Sampling module, a module for sampling from the distribution q(π) of the policy and a module for sampling from the policy π; A global network update module, a global network update module for using a back-propagation algorithm, the module comprising a parameter optimization unit for adjusting network parameters using a gradient descent method.
8. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the optimization method based on active reasoning in the low-altitude system with integrated synaesthesia and computing as described in any one of claims 1-6.
9. A computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor executes the steps of the optimization method based on active reasoning in the low-altitude system with integrated synaesthesia and computing as described in any one of claims 1 to 6.
10. An information data processing terminal, characterized in that: The information data processing terminal is used to implement the optimization system based on active reasoning in low-altitude communication supporting MEC as described in claim 7.
Citation Information
Patent Citations
Unmanned aerial vehicle intelligent trajectory planning and communication resource allocation method based on reinforcement learning
CN116704823A
Optimization method based on active reasoning in MEC enabled low-altitude element cosmic system
CN119250197A
Air-ground network optimization method and system based on MEC and digital twinning
CN119255263A
Cited By
Unmanned aerial vehicle take-off and landing point intelligent site selection method and system for low-altitude logistics
CN120235367A