Block chain enabled low-altitude edge network task unloading and resource optimization method and system
Through deep reinforcement learning and blockchain consensus mechanism, task offloading and resource allocation in low-altitude fusion networks are optimized, solving the problems of resource shortage and dynamic scheduling, achieving efficient and safe task processing and resource utilization, and adapting to the changing environment of low-altitude networks.
Patent Information
- Application Number
- CN202510695754.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-09-19
AI Technical Summary
In low-altitude fusion networks, existing technologies have failed to effectively solve the joint scheduling problems of resource shortage, task offloading and resource allocation, and lack dynamic optimization capabilities, resulting in unverifiable task execution results and security risks, making it difficult to adapt to highly dynamic and high-density application scenarios.
This paper adopts deep reinforcement learning methods, utilizes the agent's policy actor and critic network for decision-making, combines the Proximal Policy Optimization (PPO) algorithm and the blockchain consensus mechanism to optimize task offloading and resource allocation. By constructing a joint state space and action space, it realizes offloading target selection and computing resource block allocation. It also introduces dynamic view switching and credibility mechanisms to optimize master node and client selection.
It significantly improves resource utilization efficiency, reduces consensus delay and failure rate, optimizes task processing costs, improves system robustness and stability, adapts to multi-target scheduling requirements in low-altitude scenarios, and realizes safe, efficient, and low-latency intelligent scheduling.
Smart Images

Figure CN120676411A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of wireless communication technology, and in particular relates to a blockchain-enabled low-altitude edge network task offloading and resource optimization method and system. Background Art
[0002] Multi-access Edge Computing (MEC) is currently considered an important complement to cloud computing and has garnered widespread attention in IoT and smart device scenarios in recent years. By deploying computing and caching resources at edge nodes, IoT devices can perform compute-intensive and data-intensive applications, effectively reducing task processing latency and device energy consumption, while improving system response speed and resource utilization efficiency.
[0003] At the same time, low-altitude converged networks, as an emerging network architecture, integrate perception, communication, and computing functions to build drone edge platforms with local processing capabilities. Leveraging blockchain technology, they enable trusted recording of mission results and multi-party collaboration, providing strong support for future scenarios such as unmanned delivery, disaster response, and smart cities. Such networks typically include aerial computing nodes (such as drones equipped with MEC capabilities), ground terminals (such as IoT sensors and edge control stations), and communication infrastructure (such as base stations and on-chain verification nodes). Drones offer the advantages of high flexibility, rapid deployment, and coverage of inaccessible areas, effectively filling the gaps in ground infrastructure. Combined with blockchain technology, they can provide reliable support for data exchange and result storage during task offloading.
[0004] However, there are still many challenges in integrating MEC, blockchain and low-altitude networks:
[0005] (1) Communication and computing resource shortages
[0006] In a dynamic low-altitude environment, the resources of drones and edge nodes are limited, especially when multiple terminals are accessed concurrently or tasks are densely distributed. Bandwidth and computing resources are tight. If they are not allocated reasonably, it will cause resource waste or system congestion, reducing mission success rate and energy efficiency.
[0007] (2) Joint Scheduling Problem of Task Offloading and Resource Allocation
[0008] Since the amount of task data, computing intensity and completion time limit vary, and the processing capabilities of each network node are heterogeneous, the tasks need to be reasonably divided and matched to the most suitable computing nodes (such as drones or ground servers) to achieve optimal offloading efficiency. Especially under the conditions of multi-layer device collaboration, the complexity is greatly increased.
[0009] (3) Highly dynamic environment and long-term optimization issues
[0010] The positions of drones in low-altitude networks are constantly changing, and link states and node loads are highly dynamic, making traditional static or short-term optimization algorithms difficult to adapt. Given that the goal is often to maximize the long-term benefits of the system, it is necessary to develop dynamic optimization algorithms that are adaptable, convergent, and low-complexity.
[0011] Through the above analysis, the problems and defects of the existing technology are as follows:
[0012] Current literature has yet to consider how to dynamically and jointly optimize offloaded tasks and resource blocks based on network status, given limited edge node resources and sudden changes in communication load. This approach also lacks mechanisms to consider result credibility and data consistency, which can easily lead to unverifiable task execution results and security risks. Most studies fail to balance the trade-offs between task offload efficiency, blockchain consensus latency, and system energy consumption, making them difficult to adapt to the highly dynamic, high-density, and multi-target application scenarios of low-altitude converged networks. Therefore, a task offload and resource optimization method that integrates blockchain and edge computing is urgently needed to achieve the goals of secure, efficient, and low-latency intelligent scheduling in low-altitude converged networks. Summary of the Invention
[0013] In response to the problems existing in the existing technology, the present invention provides a blockchain-enabled low-altitude edge network task offloading and resource optimization method.
[0014] The present invention is implemented as follows: a blockchain-enabled low-altitude edge network task offloading and resource optimization method includes:
[0015] The method applies the policy actor and critic network of the intelligent agent and adopts the Deep Reinforcement Learning (DRL) method to make decisions, including: initializing the parameters of the policy actor network and the critic network and setting the training-related hyperparameters; the intelligent agent interacts with the environment based on the current policy, performs actions and performs state transitions; samples the entire segment in the environment using the parameters of the sampling policy actor network and stores the trajectory in memory; calculates the discounted reward, advantage function and objective function; updates the parameters of the policy actor network and the critic network as well as the parameters of the sampling policy actor and the critic network; repeats the training until the policy converges, and uses the trained policy for computation offloading and resource allocation.
[0016] Further, the following steps are included:
[0017] S101. Initialize the parameters θ and ω of the agent's policy actor network and critic network, the parameters θ′ and ω′ of the sampled policy actor network and critic network, the number of episodes, initialize the learning rates μ and σ corresponding to the policy actor network and critic network, the discount factor γ, initialize the experience pool; initialize the network layout parameters, such as the data size D of the input task i (t), task workload C i (t) and other parameters;
[0018] S102: Initialize the state of the agent, the agent interacts with the environment, and the policy network generates actions based on the current policy;
[0019] S103, the agent executes the generated action, obtains an immediate reward based on the executed action, and transfers the environment state to the next state;
[0020] S104, sampling the entire segment in the environment according to the parameters of the sampling strategy actor network, and storing the trajectory in memory;
[0021] S105. Calculate the discount reward;
[0022] S106, calculating the advantage function, adding the clipping factor to constrain the update rate, and calculating the objective function at the same time;
[0023] S107, update the parameters of the policy actor and critic network;
[0024] S108. Update the sampling strategy actor and critic network parameters according to the updated strategy actor and critic network parameters;
[0025] S109: Repeat iterative training, select the optimal action according to each state, obtain the maximum benefit, and finally obtain the optimal computing resource allocation and computing offloading strategy.
[0026] Furthermore, in step S102, the state of the agent is initialized, the agent interacts with the environment, and the main policy network generates actions based on the current policy. The state of the agent is represented as:
[0027] in An indicator indicating the link availability between IoT devices and drones; Indicates the number of computing resource blocks owned by the drone; Indicates the credit value of the node at the end of the previous time slot; A flag indicating whether client k has misbehaved during the request phase; An indicator indicating whether the master node k has misbehaved during the pre-preparation process; An indicator indicating whether consensus node k has misbehaved when submitting; It represents the additional delay of client k in the request phase; represents the additional delay of master node k in pre-preparation; represents the additional delay of consensus node k in committing; represents the average consensus delay to time slot t-1; N fail (t-1) represents the number of failures before time slot t-1. The main policy network generates actions based on the current policy as follows:
[0028]
[0029] in, It is determined by the choice of sensing resolution; It is the access control policy for IoT devices; It is the computational resource block allocation of the UAV; is the transmission rate distribution of the UAV; It is the customer’s choice decision; It is the master node selection.
[0030] Furthermore, in S103, the agent executes the action generated by the main strategy network, obtains a reward, and performs a state transition. The calculation formula for the reward in the state transition is as follows:
[0031]
[0032] In the above formula, R fail (t) represents the objective function.
[0033] Furthermore, the S106: calculating the advantage function:
[0034]
[0035] At this time there are:
[0036] δ t =r(t)+γV(s t+1 ;w)-V(s t ;w).
[0037] Proximal Policy Optimization (PPO) introduces a J-based θ′ A further improvement of the actor objective function of (θ) is to constrain the update rate by adding a clipping factor and update the PPO actor by maximizing the objective function, which is formulated as:
[0038]
[0039] Where ∈ is a hyperparameter, the clip function converts (π(a t ∣s t ;θ)) / π(a t ∣s t ; the value of θ′) is restricted to the range [1-,1+]; ensuring that the two distributions remain relatively close after minimizing the clip function.
[0040] Furthermore, the S107: update the parameters of the strategy actor and the critic network; the actor parameters are updated by the following formula:
[0041]
[0042] Considering the mean square error function of the value estimation, the loss function of the critic network is given: L critic (w)=[V(s t+1 ;w)-V(s t ;w)] 2 ,
[0043] Updated by this formula:
[0044]
[0045] Among them, δ t represents the TD error.
[0046] Another object of the present invention is to provide a blockchain-enabled low-altitude edge network task offloading and resource optimization system comprising:
[0047] System initialization module, which initializes all necessary parameters, including initializing experience memory, policy actor network parameters and critic network parameters, sampling policy actor parameters and critic network parameters;
[0048] The configuration module is used to set application-specific parameters, including task parameters such as the data size of the input task;
[0049] The agent module is used to generate actions based on the current network state at the beginning of each cycle; it is used to collect data samples, calculate the advantage function, and update the policy and value network during the policy evaluation process;
[0050] Action execution module, which is used to perform task offloading, computing resource allocation in task processing, and selection of master nodes and non-master nodes in the consensus process;
[0051] The reward acquisition module is used to execute actions and calculate instant rewards. The reward acquisition module is designed based on whether the system's constraints are met. If all constraints are met, rewards are obtained, otherwise penalties are obtained;
[0052] The state transfer module is used to transfer the system state from the current state to the next state;
[0053] The experience replay module is used to store the experience tuples of each system state, action, reward, and next state;
[0054] The data sampling module is used to extract certain fragments from the stored experience for learning;
[0055] The network update module is used to update the actor network and the critic network based on the data of the experience replay module;
[0056] Parameter update module, used for parameter update of policy actor network and critic network, as well as parameter update of sampling policy actor and critic network.
[0057] Another object of the present invention is to provide a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the blockchain-enabled low-altitude edge network task offloading and resource optimization method.
[0058] Another object of the present invention is to provide a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to perform the steps of the blockchain-enabled low-altitude edge network task offloading and resource optimization method.
[0059] Another object of the present invention is to provide an information data processing terminal, which is used to implement the blockchain-enabled low-altitude edge network task offloading and resource optimization system.
[0060] In combination with the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solutions to be protected by the present invention are as follows:
[0061] First, this paper considers a hybrid cloud and edge computing network that supports blockchain, which includes two systems: an MEC system and a blockchain system. The MEC system is responsible for generating pending computing tasks, which can be strategically selected for local execution or offloaded to edge nodes. The drone platform, acting as a mobile edge computing node, deploys computing resources and handles offloaded tasks. The blockchain consensus system is used to record and verify task execution results. Using APBFT, a reputation ranking and dynamic view switching mechanism are used to select master nodes and client nodes to form a consensus committee, enabling a highly efficient and low-latency consensus process for offload results. Secondly, considering the dynamic nature of tasks, the volatility of network conditions, and the random behavior of nodes in consensus, we use PPO to learn the environment state and obtain the optimal joint decision-making strategy for offload, computing resource allocation, and consensus committee selection.
[0062] Taking into account the highly dynamic nature of the system's environmental state, the volatility of task loads, and the uncertainty of consensus node behavior, this paper further introduces a deep reinforcement learning mechanism based on PPO. This mechanism constructs a state perception, action generation, and policy update process to implement the selection of task offloads and offload targets; the dynamic allocation of computing resource blocks required for task execution; the selection of sensing resolution in imaging tasks; and the intelligent decision-making of joint actions between clients and master nodes during the blockchain consensus process.
[0063] This paper implements a method for optimizing secure task offloading and computing resource allocation in a multi-access edge computing system supported by a consortium blockchain. This method achieves significant progress in the following key aspects:
[0064] 1) Collaborative optimization of task offloading and resource allocation:
[0065] The present invention constructs a joint state space and action space to guide the system to simultaneously select offloading targets and allocate computing resource blocks, significantly improving resource utilization efficiency and load balancing capabilities of edge nodes.
[0066] 2) Efficient blockchain consensus mechanism:
[0067] Adopting APBFT, introducing dynamic view switching and credibility mechanisms, and optimizing the selection strategies of master nodes and clients, it significantly reduces consensus delays and failure rates, and improves system robustness.
[0068] 3) Minimize system task processing costs:
[0069] When designing the reinforcement learning reward function, the system focuses on terminal energy consumption, offloading failure rate and processing delay, and minimizes the overhead of IoT terminal processing tasks through long-term strategy training.
[0070] 4) Adaptive sensing resolution selection strategy:
[0071] In image perception tasks, a resolution selection strategy is introduced as a trainable action to effectively reduce the complexity of image processing and improve the response speed of edge nodes.
[0072] 5) Deep reinforcement learning algorithm integration:
[0073] The present invention integrates the PPO algorithm, supports the coordinated update of the policy network and the value network, and enables the system to self-learn and adapt to the optimal scheduling strategy in an uncertain environment.
[0074] 6) Enhanced system stability and reliability:
[0075] Through the time difference target and soft update mechanism, the drastic fluctuations in the learning process can be effectively alleviated, ensuring that the system maintains stable operation in high concurrency and task-intensive scenarios.
[0076] 7) Network's continuous learning and online optimization capabilities:
[0077] Through replay buffering and iterative training mechanisms, the system can continuously optimize the task decision-making process, dynamically adapt to network topology and load changes, and improve performance throughout the entire life cycle.
[0078] 8) Multi-objective joint optimization design:
[0079] This method not only jointly optimizes the offloading strategy, resource allocation and consensus node selection, but also takes into account latency, energy consumption and security, meeting the multi-objective scheduling requirements in low-altitude scenarios.
[0080] 9) Ability to adapt to environmental dynamics:
[0081] In view of the characteristics of high node mobility and frequent link status fluctuations in low-altitude fusion networks, the PPO algorithm can maintain stable training and output robust strategies to adapt to changing environments.
[0082] 10) Real-time and efficiency of scheduling decisions:
[0083] The policy network outputs real-time decisions, combined with edge deployment and an efficient consensus mechanism, to ensure that the system can still respond quickly when facing large-scale access and dynamic requests.
[0084] The core of the optimization method for low-altitude fusion network systems based on APBFT consensus and MEC provided by this invention is to use mathematical models to guide the system's behavior and learning process. The technical effects brought by these mathematical models can be explored based on their characteristics:
[0085] 1) Calculation of rewards
[0086] Even the computational consensus of rewards focuses on the overall network cost, consensus delay and long-term failure rate.
[0087] Cost savings: By directly linking rewards to costs, this approach encourages reducing costs across the network and effectively improves resource utilization efficiency.
[0088] 2) Update of the policy master actor and critic network
[0089] The policy network is updated by randomly sampling small batches of experience data and using the gradient ascent method to update the current policy network. Policy Optimization: By continuously adjusting the policy network parameters, the system can learn and adopt more effective decision-making strategies.
[0090] Improved responsiveness: Using small batches of data enables the network to quickly adapt to environmental changes, enhancing the system's dynamic adjustment capabilities.
[0091] 3) Parameter update formula
[0092] This paper describes how the policy actor and critic network parameters update the target network, involving the current network and target network parameters.
[0093] Policy gradually approaches: By introducing the clipping function, the system can smoothly transition to the new policy and prevent performance fluctuations caused by drastic changes.
[0094] Continuous learning and adaptation: This continuous parameter update mechanism ensures that the system can adapt to long-term environmental changes.
[0095] Second, by introducing the PPO algorithm, the present invention implements an optimization method for a low-altitude fusion network system based on APBFT consensus and MEC, significantly improving network performance and service quality. This not only meets the growing demand for high-performance computing, but also optimizes resource utilization and reduces network costs. In terms of commercial applications, the present invention can be widely used in application scenarios that require efficient processing of large amounts of data and ensure data security and trust, bringing significant economic and social benefits to related industries. Therefore, the present invention has broad market prospects and huge commercial value.
[0096] Task scheduling in low-altitude fusion networks has always been subject to multiple limitations of system heterogeneity, network volatility, and consensus delay. Traditional static optimization or phased scheduling methods cannot cope with the collaborative scheduling problem in this highly dynamic environment. Especially under multi-objective constraints (delay, energy consumption, consensus success rate), the linkage between subsystems is difficult to model and optimize. The present invention constructs a joint state-action space, trains a policy network with generalization capabilities, realizes end-to-end joint decision-making of the system, and successfully opens up the closed-loop optimization process between edge computing and blockchain mechanisms for the first time, significantly improving the overall system intelligence and stability. BRIEF DESCRIPTION OF THE DRAWINGS
[0097] Figure 1 This is a flow chart of the blockchain-enabled low-altitude edge network task offloading and resource optimization method provided by an embodiment of the present invention.
[0098] Figure 2 This is a structural block diagram of the blockchain-enabled low-altitude edge network task offloading and resource optimization system provided by an embodiment of the present invention.
[0099] Figure 3 This is a flowchart of the implementation method of the blockchain-enabled low-altitude edge network task offloading and resource optimization provided by an embodiment of the present invention.
[0100] Figure 4 This is a comparison diagram of comprehensive convergence performance provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0101] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0102] like Figure 1 As shown, a blockchain-enabled low-altitude edge network task offloading and resource optimization method provided by an embodiment of the present invention includes the following steps:
[0103] S101. Initialize the parameters θ and ω of the agent's policy actor network and critic network, the parameters θ′ and ω′ of the sampled policy actor network and critic network, the number of episodes, initialize the learning rates μ and σ corresponding to the policy actor network and critic network, the discount factor γ, initialize the experience pool; initialize the network layout parameters, such as the data size D of the input task i (t), task workload C i (t) and other parameters;
[0104] S102: Initialize the state of the agent, the agent interacts with the environment, and the policy network generates actions based on the current policy;
[0105] S103, the agent executes the generated action, obtains an immediate reward based on the executed action, and transfers the environment state to the next state;
[0106] S104, sampling the entire segment in the environment according to the parameters of the sampling strategy actor network, and storing the trajectory in memory;
[0107] S105. Calculate the discount reward;
[0108] S106, calculating the advantage function, adding the clipping factor to constrain the update rate, and calculating the objective function at the same time;
[0109] S107, update the parameters of the policy actor and critic network;
[0110] S108. Update the sampling strategy actor and critic network parameters according to the updated strategy actor and critic network parameters;
[0111] S109: Repeat iterative training, select the optimal action according to each state, obtain the maximum benefit, and finally obtain the optimal computing resource allocation and computing offloading strategy.
[0112] The blockchain-enabled low-altitude edge network task offloading and resource optimization method provided by the embodiment of the present invention applies the agent's policy actor and critic network and adopts the deep reinforcement learning (DRL) method for decision-making, including:
[0113] Initialize the parameters of the policy actor network and the critic network, and set the training-related hyperparameters; the agent interacts with the environment based on the current strategy, performs actions and performs state transitions; samples the entire segment in the environment using the parameters of the sampling policy actor network and stores the trajectory in memory; calculates the discounted reward, advantage function and objective function; updates the parameters of the policy actor network and the critic network as well as the parameters of the sampling policy actor and the critic network; repeats the training until the strategy converges, and uses the trained strategy for computational offloading and resource allocation.
[0114] In S102 provided by the embodiment of the present invention, the state of the agent is initialized. The agent interacts with the environment, and the main policy network generates actions based on the current policy. The state of the agent is represented as:
[0115]
[0116] in An indicator indicating the link availability between IoT devices and drones; Indicates the number of computing resource blocks owned by the drone; Indicates the credit value of the node at the end of the previous time slot; A flag indicating whether client k has misbehaved during the request phase; An indicator indicating whether the master node k has misbehaved during the pre-preparation process; An indicator indicating whether consensus node k has misbehaved when submitting; It represents the additional delay of client k in the request phase; represents the additional delay of master node k in pre-preparation; represents the additional delay of consensus node k in committing; represents the average consensus delay to time slot t-1; Nfail (t-1) represents the number of failures before time slot t-1; the main strategy network generates actions based on the current strategy as follows:
[0117]
[0118] in, It is determined by the choice of sensing resolution; It is the access control policy for IoT devices; It is the computational resource block allocation of the UAV; is the transmission rate distribution of the UAV; It is the customer’s choice decision; It is the master node selection.
[0119] In step S103 provided by the embodiment of the present invention, the agent executes the generated action, obtains an immediate reward based on the executed action, and transfers the environment state to the next state. The calculation consensus of the reward is as follows:
[0120]
[0121] In the above formula, R fail (t) represents the objective function.
[0122] S106 provided by the embodiment of the present invention: Calculating the advantage function:
[0123]
[0124] At this time there are:
[0125] δ t =r(t)+γV(s t+1 ;w)-V(s t ;w).
[0126] PPO introduces the J-based θ′ A further improvement of the actor objective function of (θ) is to constrain the update rate by adding a clipping factor and update the PPO actor by maximizing the objective function, which is formulated as:
[0127]
[0128] Where ∈ is a hyperparameter, the clip function converts (π(a t ∣s t ;θ)) / π(a t ∣s t ; the value of θ′) is restricted to the range [1-,1+]; ensuring that the two distributions remain relatively close after minimizing the clip function.
[0129] In an embodiment of the present invention, S107 is provided: updating the parameters of the strategy actor and the critic network; the parameters of the actor are updated by the following formula:
[0130]
[0131] Considering the mean square error function of the value estimation, the loss function of the critic network is given: L critic (w)=[V(s t+1 ;w)-V(s t ;w)] 2 ,
[0132] Updated by this formula:
[0133]
[0134] Among them, δ t represents the TD error.
[0135] like Figure 2 As shown, an embodiment of the present invention provides a blockchain-enabled low-altitude edge network task offloading and resource optimization system comprising:
[0136] System initialization module, which initializes all necessary parameters, including initializing experience memory, policy actor network parameters θ and critic network parameters w, sampling policy actor parameters θ' and critic network parameters w';
[0137] The configuration module is used to set application-specific parameters, including task parameters such as the data size of the input task;
[0138] The agent module is used to generate actions based on the current network state at the beginning of each cycle; it is used to collect data samples, calculate the advantage function, and update the policy and value network during the policy evaluation process;
[0139] Action execution module, which is used to perform task offloading, computing resource allocation in task processing, and selection of master nodes and non-master nodes in the consensus process;
[0140] The reward acquisition module is used to execute actions and calculate instant rewards. The reward acquisition module is designed based on whether the system's constraints are met. If all constraints are met, rewards are obtained, otherwise penalties are obtained;
[0141] The state transfer module is used to transfer the system state from the current state to the next state;
[0142] The experience replay module is used to store the experience tuples of each system state, action, reward, and next state;
[0143] The data sampling module is used to extract certain fragments from the stored experience for learning;
[0144] The network update module is used to update the actor network and the critic network based on the data of the experience replay module;
[0145] Parameter update module, used for parameter update of policy actor network and critic network, as well as parameter update of sampling policy actor and critic network.
[0146] Traditional low-altitude edge networks often face challenges in task offloading and resource allocation, including delayed processing decisions, inflexible scheduling, and a lack of transparent and trusted coordination mechanisms. In multi-node environments, task scheduling decisions are prone to local optima, while resource scheduling often lacks a global perspective and effective incentive mechanisms, leading to low overall resource utilization and uncontrollable task delays. To address this bottleneck, this paper introduces a blockchain consensus mechanism and deep reinforcement learning strategies to construct an adaptively evolving optimization decision-making system, aiming to achieve dynamic linkage between the three layers of computing, communication, and coordination.
[0147] During the system initialization phase, the initial parameters of the policy network and value network are loaded to establish a trainable multi-agent learning structure. An experience memory mechanism is introduced to provide a data foundation for decision trajectories in subsequent reinforcement learning. The policy network is responsible for action selection, while the value network is used to estimate the expected reward of state-action pairs. The two networks continuously correct each other in discrete time steps. This design emphasizes convergence and transferability, enabling the system to extract optimal behavior patterns from complex dynamic environments.
[0148] The configuration module allows users or system managers to input settings based on the specific low-altitude mission type, such as mission data volume, node computing power, communication bandwidth, and other key parameters. This setting not only determines the optimization objective function of the task offloading strategy but also provides contextual constraints for the design of the reward function. By accurately characterizing the mission, the system achieves parameter-driven scheduling based on scenario requirements.
[0149] The agent module is the core decision-making unit. It perceives the network state during each time period and then invokes the policy network to generate actions, such as whether to offload tasks, which edge node to assign them to, and whether to run for consensus master node. The agent then updates its policy based on this feedback and performs value assessment to refine its behavior. The introduction of an advantage function ensures that the policy update process favors efficient actions, thereby improving the robustness of the overall network scheduling strategy.
[0150] At the actual execution layer, task offloading corresponds to data flow and computational migration, while the dynamic selection of consensus nodes is related to the confirmation and synchronization efficiency of information within the blockchain structure. This module not only schedules node resources but also coordinates the interface logic between the on-chain consensus structure and edge computing, achieving the coordinated optimization of task execution and consensus behavior, ensuring the simultaneous advancement of resource allocation and trusted records.
[0151] Immediate rewards after system execution are generated by the reward acquisition module, which constructs an evaluation function based on multi-dimensional constraints such as offload latency, resource utilization, node energy consumption, and network stability. If the system state satisfies all constraints, a positive reward is generated; otherwise, a penalty signal is triggered. This feedback participates in the gradient backpropagation of the policy network, guiding subsequent policy convergence towards the global optimum. This entire process, supported by experience replay and batch sampling mechanisms, promotes stable iterative policy learning.
[0152] Example 1
[0153] To address the challenges of existing technologies, this implementation provides an optimization method for a low-altitude converged network system based on APBFT consensus and MEC. First, it comprises two systems: a multi-access edge computing system and a blockchain system. MEC systems are designed for mission execution, while blockchain systems play a crucial role in establishing a secure and reliable transaction platform for MEC systems.
[0154] In an MEC system, there are I IoT devices, K drones, and one macro base station. The set of IoT devices and drones is denoted by I = {1,...,i,...I} and κ = {1,...,k,...K}, respectively. Each drone is equipped with a MEC server, which provides coverage and computing power to specific IoT devices. In a blockchain system, each drone has a dual function, serving as a consensus node for blockchain consensus. A controller equipped with a DRL algorithm is deployed on the MBS to collect dynamic environmental information and make optimization decisions based on it. The system operates in discrete time slots, with the duration t of each time slot being represented by Δt.
[0155] Example 2
[0156] Given the dynamic nature of the network and the uncertainty of information acquisition, the problem becomes quite complex. To effectively address this challenge, we remodel the problem as a Markov decision process (MDP). Because the variables involved are discontinuous, we choose to use the PPO algorithm, a DRL algorithm designed for solving reinforcement learning problems in continuous action spaces and capable of supporting real-time online decision making.
[0157] like Figure 3 , an embodiment of the present invention provides an optimization method in a low-altitude fusion network system based on APBFT consensus and MEC, which specifically includes the following steps:
[0158] S101. Initialize the parameters θ and ω of the agent's policy actor network and critic network, the parameters θ′ and ω′ of the sampled policy actor network and critic network, the number of episodes, initialize the learning rates μ and σ corresponding to the critic network and policy network, the discount factor γ, initialize the experience pool; initialize the network layout parameters, such as the data size D of the input task i (t), task workload C i (t) and other parameters;
[0159] S102: Initialize the state of the agent, the agent interacts with the environment, and the main policy network generates actions based on the current policy;
[0160] S103, the agent executes the generated action, obtains an immediate reward based on the executed action, and transfers the environment state to the next state;
[0161] S104, sampling the entire segment in the environment according to the parameters of the sampling strategy actor network, and storing the trajectory in memory;
[0162] S105. Calculate the discount reward;
[0163] S106, calculating the advantage function, adding the clipping factor to constrain the update rate, and calculating the objective function at the same time;
[0164] S107, update the parameters of the policy actor and critic network;
[0165] S108. Update the sampling strategy actor and critic network parameters according to the updated strategy actor and critic network parameters;
[0166] S109: Repeat iterative training, select the optimal action according to each state, obtain the maximum benefit, and finally obtain the optimal computing resource allocation and computing offloading strategy.
[0167] S102. Initialize the state of the agent. The agent interacts with the environment. The main policy network generates actions based on the current policy. The state of the agent is expressed as:
[0168]
[0169] in An indicator indicating the link availability between IoT devices and drones; Indicates the number of computing resource blocks owned by the drone; Indicates the credit value of the node at the end of the previous time slot; A flag indicating whether client k has misbehaved during the request phase; An indicator indicating whether the master node k has misbehaved during the pre-preparation process; An indicator indicating whether consensus node k has misbehaved when submitting; It represents the additional delay of client k in the request phase; represents the additional delay of master node k in pre-preparation; represents the additional delay of consensus node k in committing; represents the average consensus delay to time slot t-1; N fail (t-1) represents the number of failures before time slot t-1. The main policy network generates actions based on the current policy as follows:
[0170]
[0171] in, It is determined by the choice of sensing resolution; It is the access control policy for IoT devices; It is the computational resource block allocation of the UAV; is the transmission rate distribution of the UAV; It is the customer’s choice decision; It is the master node selection.
[0172] S103: The agent executes the action generated by the main strategy network, obtains the reward, and performs the state transition. The reward calculation formula during the state transition is as follows:
[0173]
[0174] In the above formula, R fail (t) represents the objective function.
[0175] S106: Calculate advantage function:
[0176]
[0177] At this time there are:
[0178] δ t =r(t)+γV(s t+1 ;w)-V(s t ;w).
[0179] In order to improve the performance, PPO introduces a θ′ Further improvement of the actor objective function of (θ). By adding a clipping factor to constrain the update rate, the PPO actor can be updated by maximizing the objective function, as follows:
[0180]
[0181] Where ∈ is a hyperparameter, the clip function converts (π(a t ∣s t ;θ)) / π(a t ∣s t ; θ′) is restricted to the range [1-,1+]. This approach ensures that after minimizing the clip function, the two distributions remain relatively close and avoid significant differences.
[0182] S107: Update the parameters of the policy actor and critic network; the actor parameters are updated by the following formula:
[0183]
[0184] This algorithm considers the mean square error function of the value estimation and gives the loss function of the critic network:
[0185] L critic (w)=[V(s t+1 ;w)-V(s t ;w)] 2 ,
[0186] At the same time, it can be updated by this formula:
[0187]
[0188] Among them, δ t represents the TD error.
[0189] In order to elaborate on the optimization method in the low-altitude fusion network system based on APBFT consensus and MEC, the present invention provides two specific application embodiments, including key details of the implementation scheme.
[0190] Application Example 1: Smart City IoT Application
[0191] 1) System initialization
[0192] A central controller was deployed in the low-altitude economic demonstration zone, along with multiple drones equipped with edge computing servers and ground-based IoT sensing nodes, to establish a complete communication network architecture. The system uses a distributed deployment approach to ensure network coverage and fault tolerance.
[0193] 2) Interaction between the agent and the environment
[0194] As intelligent agents, drone nodes perceive network status in real time, including channel quality, computing resources, and node trust, and dynamically adjust their operational strategies. By continuously monitoring environmental changes, the system can adaptively adjust communication parameters and computing task allocation.
[0195] 3) Action execution and reward acquisition
[0196] After the system executes its optimization decision, it conducts a comprehensive evaluation based on indicators such as task completion quality, response speed, and energy efficiency, and then awards corresponding rewards. This reward mechanism uses a multi-dimensional evaluation system that considers both immediate performance and long-term operational stability.
[0197] 4) Optimization and update
[0198] We use intelligent algorithms to continuously optimize network strategies, focusing on improving the accuracy of masternode selection, rational resource allocation, and the ability to identify malicious behavior. The optimization process uses incremental updates to ensure the system remains stable throughout the optimization process.
[0199] 5) Energy consumption and performance optimization
[0200] This embodiment demonstrates an innovative low-altitude network drone cluster collaboration solution. The system builds a complete communication network architecture by deploying an intelligent central controller and drone nodes equipped with edge computing capabilities. Using advanced intelligent algorithms to implement state strategy adjustment, the system can perceive the network status in real time, including key parameters such as channel quality, resource distribution, and node trust value. By establishing a multi-dimensional evaluation and reward mechanism, the system continuously optimizes core functions such as master client and node selection, resource allocation, and anomaly detection. Practical applications have proved that this solution significantly improves the collaborative efficiency of drone clusters, and demonstrates excellent performance in terms of communication quality, mission reliability, and safety, providing reliable technical support for various application scenarios in the low-altitude economy.
[0201] Application Example 2: Disaster Response and Rescue System
[0202] 1) System initialization
[0203] In disaster response scenarios, a swarm of drones with edge computing capabilities and blockchain consensus nodes are deployed to initialize the parameters of the PPO reinforcement learning network. Ground-based IoT devices are also configured for environmental perception and information collection. The system establishes a low-altitude converged network architecture to ensure smooth communication and coordinated computing resources within and outside the disaster area.
[0204] 2) Interaction between the agent and the environment
[0205] Drones and blockchain nodes, acting as intelligent agents, continuously perceive the state of the disaster area, including on-site data, communication link quality, task priorities, and node load. Based on this state, the intelligent agents use the PPO strategy to make decisions about offloading tasks, allocating computing resources, selecting image data resolution, and selecting blockchain consensus nodes.
[0206] 3) Action execution and reward acquisition
[0207] The system performs task offloading and resource scheduling according to policy instructions. The drone processes the collected imagery and sensor data at the edge, uploading the results to the blockchain system for secure storage and verification. After completing the action, the system calculates a reward signal based on metrics such as response latency, data processing quality, and consensus success rate, and feeds it back to the reinforcement learning module.
[0208] 4) Optimization and update
[0209] The reinforcement learning module iteratively updates the policy and value network parameters based on reward feedback, enabling adaptation to dynamic changes in disaster environments and optimizing resource scheduling. Drone path adjustments, computing task migration, and blockchain consensus node selection are optimized in real time to ensure the efficiency and safety of disaster response missions.
[0210] 5) System intelligent decision-making ensures rescue efficiency
[0211] Through the joint optimization method of the present invention, the disaster response system can achieve real-time, efficient and safe task offloading and resource allocation, effectively improve the collaborative processing capabilities and data credibility of the drone swarm, ensure the timeliness and accuracy of rescue decisions, and improve disaster response efficiency and on-site safety.
[0212] In these two embodiments, the optimization method in the low-altitude fusion network system based on APBFT consensus and MEC provides an efficient, reliable and energy-saving solution, which is suitable for different application scenarios, from smart cities to disaster response and rescue systems, demonstrating its wide application potential and technical advantages.
[0213] To comprehensively evaluate the performance of our invention, we compared it with several representative benchmark algorithms. These benchmark algorithms each have their own unique characteristics and advantages, effectively handling similar problems. By comparing these algorithms under the same conditions, we can more clearly understand the advantages of our invention and the potential for improvement.
[0214] 1) Random Offloading: The offloading decision for each IoT device-generated computing task is determined randomly. Tasks can be randomly offloaded to drone nodes. However, a task can only be processed at one location. This optimizes the allocation of computing resources, the selection of sensing resolution, and the selection of blockchain clients and master nodes.
[0215] 2) Random Master Node: The blockchain consensus master node is randomly selected from all drone nodes, including master nodes and non-master nodes. This optimizes task offloading decisions and computing resource allocation decisions.
[0216] 3) The present invention: It represents the proposed PPO-based joint task offloading, computing resource allocation, and blockchain client and master node selection algorithm.
[0217] exist Figure 4 In this paper, we demonstrate the convergence performance of all the aforementioned algorithms using default parameter settings. Convergence performance is evaluated based on three performance metrics: convergence speed, convergence stability, and reward value. A careful observation of the three curves indicates that the proposed algorithm has the fastest convergence speed and outperforms the two baseline algorithms.
[0218] It should be noted that the embodiments of the present invention can be implemented by hardware, software, or a combination of software and hardware. The hardware portion can be implemented using dedicated logic; the software portion can be stored in a memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated design hardware. Those skilled in the art will appreciate that the above-mentioned devices and methods can be implemented using computer-executable instructions and / or contained in processor control code, for example, such as a carrier medium such as a disk, CD or DVD-ROM, a programmable memory such as a read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and modules of the present invention can be implemented by hardware circuits such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, or programmable hardware devices such as field programmable gate arrays, programmable logic devices, etc., can also be implemented by software executed by various types of processors, or can be implemented by a combination of the above-mentioned hardware circuits and software, such as firmware.
[0219] The above description is only a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions and improvements made by any technician familiar with this technical field within the technical scope disclosed by the present invention and within the spirit and principles of the present invention should be covered by the scope of protection of the present invention.
Claims
1. A blockchain-enabled low-altitude edge network task offloading and resource optimization method, characterized in that: The following steps are involved: S101. Initialize the parameters of the policy actor network and the critic network, sample the policy network parameters, set the number of rounds, learning rate, discount factor and experience pool, and set the network layout and task input parameters; S102, initialize the agent state, the agent interacts with the environment, and the main policy network generates actions; S103: Execute the action, calculate the instant reward and complete the state transfer; S104, sampling strategy network parameters, obtaining environment segment trajectories and storing them in the experience pool; S105. Calculate the discount reward; S106, calculating the advantage function, adding a clipping factor to limit the strategy update rate, and constructing the objective function; S107, update the policy actor and critic network parameters; S108, updating the sampling strategy network based on the latest strategy parameters; S109, loop iterative training, select the best action, and optimize computing resource allocation and task offloading strategy.
2. The method according to claim 1, wherein The agent status includes link availability indicators, node credit value, client and consensus node behavior flags, historical delay, average consensus delay and number of failures.
3. The method according to claim 1, wherein The actions output by the master strategy network include sensing resolution selection, access control strategy, computing resource block resource allocation, transmission rate configuration, client node selection and master node selection.
4. The method according to claim 1, wherein The immediate reward is set based on the objective function value and the system constraints. When all constraints are met, a positive reward is obtained, and when the constraints are violated, a negative reward is obtained.
5. The method according to claim 1, wherein The advantage function is calculated by the difference between the current value function and the discounted cumulative reward, and is used to measure the direction and magnitude of strategy improvement.
6. The method according to claim 1, wherein The objective function of the policy actor network uses a clipping probability ratio to constrain update changes, improve training stability and avoid policy degradation.
7. The method according to claim 1, wherein The critic network uses the mean square error loss function and performs parameter backpropagation and update through TD error.
8. The method according to claim 1, wherein The experience pool uses a sliding window method to cache states, actions, rewards, and transfer tuples, and regularly replaces historical samples.
9. The method according to claim 1, wherein The sampling strategy network is independent of the training strategy network and integrates experience with the existing model through a soft update strategy to enhance generalization performance.
10. A blockchain-enabled low-altitude edge network task offloading and resource optimization system implementing the method according to any one of claims 1 to 9, characterized in that: include: System initialization module, configuration module, agent module, action execution module, reward acquisition module, state transfer module, experience replay module, data sampling module, network update module and parameter update module.
Citation Information
Cited By
Mobile edge computing resource allocation method and system based on block chain
CN121455679A