Optimization method and system in ipbft consensus multi-access edge computing system
By introducing deep reinforcement learning and the DPoS-improved PBFT algorithm into the MEC system of IPBFT consensus, task offloading and resource allocation are optimized, network architecture integration, real-time, trust and security issues are solved, and efficient and stable computing resource allocation and consensus performance are achieved.
Patent Information
- Application Number
- CN202411213712.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-30
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2044-08-30
AI Technical Summary
Existing technologies in the MEC system with IPBFT consensus have problems such as insufficient network architecture integration, difficulty in meeting real-time and latency requirements, trust and security issues, and insufficient response to environmental dynamics, which affect the efficiency and stability of the system.
The deep reinforcement learning (DRL) method is adopted to optimize task offloading and computing resource allocation through the policy actor and critic network. Combined with the PBFT algorithm improved by DPoS, an efficient consensus mechanism is designed, and the PPO algorithm is introduced for self-learning and adjustment of network status.
It achieves precise task offloading and resource allocation, improves the system's resource utilization and consensus performance, reduces consensus delay and failure rate, enhances the system's stability and reliability, adapts to complex and dynamic environments, and meets real-time and security requirements.
Smart Images

Figure CN119277416B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of wireless communication technology, and in particular relates to an optimization method and system for a multi-access edge computing system with IPBFT consensus, specifically to security task offloading and computing resource allocation. Background Art
[0002] Currently, the popularity of the Internet of Things (IoT) has led to a surge in user equipment (UE) that affects our daily lives. However, traditional centralized architectures, especially in cloud computing, have difficulty meeting the mobility and rapid response requirements of these terminals. This challenge has stimulated the emergence of multi-access edge computing (MEC). MEC is particularly suitable for real-time applications, enabling rapid response and localized decision-making. The scalability of cloud services combined with the low latency of edge computing creates a powerful and efficient computing ecosystem, providing a comprehensive solution that meets the diverse needs of different industries. In edge computing systems, tasks are typically processed on edge servers or cloud servers, which are collectively referred to as compute power providers (CPPs).
[0003] In the above scenarios, consortium blockchain systems are expected to significantly improve the reliability of collaboration. Task processing, resource utilization, and payment information are recorded as transactions, and each transaction is digitally signed by the service provider. Each blockchain node retains a copy of all transactions as immutable evidence of service provision and payment to facilitate the resolution of potential disputes, thereby establishing a robust and trusted environment for decentralized data processing in edge computing systems. The operator alliance deploys smart contracts, specifies responsibilities, reward and punishment standards, and triggers, etc., to ensure automatic execution to prevent disputes over contract performance. Integrating blockchain and MEC networks in a unified system can help create a task processing service environment with trustworthy, transparent, secure, immutable, and automated characteristics. However, there are still some challenges and problems when applying deep reinforcement learning (DRL) to MEC systems with IPBFT consensus:
[0004] 1) Network architecture integration issues
[0005] First, a multi-system architecture, encompassing both edge computing and blockchain systems, must be designed to ensure efficient data and task transmission and processing between these systems. Secondly, the integration of consensus mechanisms and computing resources is crucial. The Practical Byzantine Fault Tolerance (PBFT) consensus algorithm requires frequent communication and information exchange between nodes, which consumes significant network bandwidth and computing resources. Therefore, consensus algorithms need to be optimized to reduce resource consumption and improve efficiency.
[0006] In particular, PBFT faces the following challenges: (1) The number of consensus nodes is fixed, and all blockchain nodes participate in the consensus process, resulting in low consensus efficiency; (2) Due to the single point full node and two full node broadcasts, PBFT brings high bandwidth overhead, limiting its application in large-scale networks; (3) The arbitrary selection of master nodes carries the risk of selecting malicious or faulty nodes, causing consensus failure. Although changing the protocol solves this problem, its high communication complexity leads to resource waste, extended latency, and reduced system stability; (4) PBFT encounters difficulties in dynamic scenarios and lacks an effective mechanism for seamless transition during node changes. Restarting the entire network is costly, and frequent replacements can seriously affect system availability.
[0007] 2) Real-time and latency requirements
[0008] In MEC systems based on IPBFT consensus, real-time and latency requirements are crucial. MEC systems aim to provide low-latency, high-real-time services, placing extremely high demands on data processing and task response speeds. However, the IPBFT consensus mechanism involves complex inter-node communication and multiple message passes, which increases network latency. This poses a significant challenge for application scenarios that require immediate feedback and processing, such as autonomous driving, real-time monitoring, and augmented reality. Therefore, optimizing consensus algorithms to reduce communication overhead, improve consensus efficiency, and design efficient data transmission and processing mechanisms are key to meeting real-time and latency requirements. Only by ensuring real-time system response and minimal latency can the advantages of MEC systems be fully realized in various applications.
[0009] 3) Trust and security issues
[0010] Due to the distributed and heterogeneous nature of MEC systems, reliable trust relationships must be established between nodes to ensure the security of data transmission and processing. The IPBFT consensus mechanism improves system reliability through fault-tolerant design, but it also faces the risk of selecting malicious or faulty nodes, which can lead to consensus failure or system attacks. To this end, it is necessary to combine encryption technology, access control, and node reputation evaluation mechanisms to ensure that only trusted nodes can participate in the consensus process. At the same time, efficient security protocols must be designed to prevent data tampering, unauthorized access, and network attacks, ensuring the overall security and stability of the system. Only when trust and security issues are effectively addressed can the MEC system provide reliable computing and services that meet the high security requirements of users and applications.
[0011] 4) Environmental dynamics
[0012] In an MEC system based on IPBFT consensus, the random nature of task and network states, coupled with the tight coupling of various constraints, necessitates a high degree of flexibility and adaptability. Node reputations are also constantly changing, requiring the system to dynamically adjust task allocation and resource scheduling strategies to ensure service continuity and stability. Therefore, designing algorithms and mechanisms that can cope with this high degree of dynamism, such as real-time load balancing, dynamic resource management, and flexible consensus protocols, is key to ensuring the efficient operation of MEC systems in dynamic environments. These mechanisms must be able to rapidly respond to environmental changes, maintain system stability and efficiency, and maximize available resources to meet the needs of various applications.
[0013] Through the above analysis, the problems and defects of the existing technology are as follows:
[0014] (1) Insufficient network architecture integration: Existing technologies still have deficiencies in designing multi-system architectures, making it difficult to efficiently integrate edge computing systems and blockchain systems, which limits the transmission and processing efficiency of data and tasks.
[0015] (2) Timeliness and delay requirements are difficult to meet: The existing IPBFT consensus mechanism involves complex inter-node communication and multiple message transmissions, which causes delay problems in application scenarios that require low latency and high real-time performance (such as autonomous driving and real-time monitoring), affecting the performance and application effect of the MEC system.
[0016] (3) Trust and security issues: The IPBFT consensus mechanism faces the risk of malicious nodes or faulty nodes being selected, which may lead to consensus failure or system security being threatened. Existing technologies have not yet formed an effective solution in terms of node reputation evaluation and trust establishment mechanism.
[0017] (4) Insufficient response to the dynamic nature of the environment: Existing technologies are difficult to effectively cope with the high dynamics of MEC systems. The system lacks sufficient flexibility and adaptability when dealing with random changes in task status and network status. Summary of the Invention
[0018] In response to the problems existing in the prior art, the present invention provides an optimization method and system for a multi-access edge computing system with IPBFT consensus.
[0019] The present invention is implemented as follows: an optimization method for a multi-access edge computing system with IPBFT consensus, which applies a policy actor and critic network of an intelligent agent and adopts a deep reinforcement learning (DRL) method for decision making, including:
[0020] Initialize the parameters of the policy actor network and the critic network, and set the training-related hyperparameters; the agent interacts with the environment based on the current strategy, performs actions and performs state transitions; samples the entire segment in the environment using the parameters of the sampling policy actor network and stores the trajectory in memory; calculates the discounted reward, advantage function and objective function; updates the parameters of the policy actor network and the critic network as well as the parameters of the sampling policy actor and the critic network; repeats the training until the strategy converges, and uses the trained strategy for computational offloading and resource allocation.
[0021] Further, the following steps are included:
[0022] S101. Initialize the parameters θ and ω of the agent's policy actor network and critic network, the parameters θ′ and ω′ of the sampled policy actor network and critic network, the number of episodes, initialize the learning rates μ and σ corresponding to the policy actor network and critic network, the discount factor γ, initialize the experience pool; initialize the network layout parameters, such as the data size D of the input task i (t), task workload C i (t) and other parameters;
[0023] S102: Initialize the state of the agent, the agent interacts with the environment, and the policy network generates actions based on the current policy;
[0024] S103, the agent executes the generated action, obtains an immediate reward based on the executed action, and transfers the environment state to the next state;
[0025] S104, sampling the entire segment in the environment according to the parameters of the sampling strategy actor network, and storing the trajectory in memory;
[0026] S105. Calculate the discount reward;
[0027] S106, calculating the advantage function, adding the clipping factor to constrain the update rate, and calculating the objective function at the same time;
[0028] S107, update the parameters of the policy actor and critic network;
[0029] S108. Update the sampling strategy actor and critic network parameters according to the updated strategy actor and critic network parameters;
[0030] S109: Repeat iterative training, select the optimal action according to each state, obtain the maximum benefit, and finally obtain the optimal computing resource allocation and computing offloading strategy.
[0031] Furthermore, in S102, the agent interacts with the environment; at the beginning of each round, the system state s(t) is initialized; wherein the state includes three parts, namely, the state of the task s task (t), the state of the network s net (t) and the state information s in the consensus con (t), as follows:
[0032] s(t)={s task (t),s net (t),s con (t)}.
[0033] The status of the task task (t) is given by the following formula:
[0034] s task (t)={D(t),C(t)},
[0035] Among them, the input data volume of the task The workload of the task
[0036] Network Status net (t) is given by:
[0037] s net (t) = {A(t), R(t), U(t), r m (t),f m (t)},
[0038] Among them, the link availability between UE and SBSs is The link connection rate between UE and SBSs is Number of computing resource blocks of SBSs The link connection rate between UE and MBS is r m (t); The computing resources allocated by MBS to each task are fm (t+1);
[0039] State information s during the consensus process con (t) is given by the following formula:
[0040]
[0041] Where W(t-1)={W k (t)} is the credit value of the node at time slot t-1, Malicious node indicators The number of malicious nodes is When the master node is a malicious node, the indicator of whether the master node k has reached consensus When the non-master node is a malicious node, the indicator of whether the non-master node k has reached consensus When the master node is a normal node, the malicious delay of master node k in the Pre-prepare stage is When the non-master node is a normal node, the malicious delay of the non-master node k in the Commit phase is The number of failures at the previous moment is N fail (t-1); the average consensus delay is
[0042] At each time slot t, the agent decides its action a(t) based on the state s(t) given by
[0043]
[0044] in, It is the task offloading decision of UE; Make CRB allocation decisions during task processing; Select decision indicators for consensus nodes, where Select the indicator for the master node, y k (t) Select decision indicators for non-primary nodes.
[0045] Furthermore, the consensus calculation of the reward in S103: the agent executes the generated action, obtains an immediate reward based on the executed action, and transfers the environment state to the next state is as follows:
[0046]
[0047] In the above formula, Object(t) represents our objective function.
[0048] Furthermore, the S106: calculating the advantage function:
[0049]
[0050] Among them, at this time there are:
[0051] δ t =r(t)+γV(s t+1 ;w)-V(s t ;w).
[0052] γ is the discount factor. PPO introduces the J-based θ′ A further improvement of the actor objective function of (θ) is to constrain the update rate by adding a clipping factor and update the PPO actor by maximizing the objective function, which is formulated as:
[0053]
[0054] Where ∈ is a hyperparameter, the clip function converts (π(a t ∣s t ;θ)) / π(a t ∣s t ; the value of θ′) is restricted to the range [1-,1+]; ensuring that the two distributions remain relatively close after minimizing the clip function.
[0055] Furthermore, the S107: update the parameters of the strategy actor and the critic network; the actor parameters are updated by the following formula:
[0056]
[0057] Considering the mean square error function of the value estimation, the loss function of the critic network is given: L critic (w)=[V(s t+1 ;w)-V(s t ;w)] 2 ,
[0058] Updated by this formula:
[0059]
[0060] Among them, δ t represents the TD error.
[0061] Another object of the present invention is to provide a multi-access edge computing system based on IPBFT consensus, the system comprising:
[0062] System initialization module, which initializes all necessary parameters, including initializing experience memory, policy actor network parameters θ and critic network parameters w, sampling policy actor parameters θ' and critic network parameters w';
[0063] The configuration module is used to set application-specific parameters, including task parameters such as the data size of the input task;
[0064] The agent module is used to generate actions based on the current network state at the beginning of each cycle; it is used to collect data samples, calculate the advantage function, and update the policy and value network during the policy evaluation process;
[0065] Action execution module, which is used to perform task offloading, computing resource allocation in task processing, and selection of master nodes and non-master nodes in the consensus process;
[0066] The reward acquisition module is used to execute actions and calculate instant rewards. The reward acquisition module is designed based on whether the system's constraints are met. If all constraints are met, rewards are obtained, otherwise penalties are obtained;
[0067] The state transfer module is used to transfer the system state from the current state to the next state;
[0068] The experience replay module is used to store the experience tuples of each system state, action, reward, and next state;
[0069] The data sampling module is used to extract certain fragments from the stored experience for learning;
[0070] The network update module is used to update the actor network and the critic network based on the data of the experience replay module;
[0071] Parameter update module, used for parameter update of policy actor network and critic network, as well as parameter update of sampling policy actor and critic network.
[0072] Another object of the present invention is to provide a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the optimization method in a multi-access edge computing system based on IPBFT consensus.
[0073] Another object of the present invention is to provide a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to perform the steps of an optimization method in a multi-access edge computing system based on IPBFT consensus.
[0074] In combination with the above technical solutions and the technical problems solved, please analyze the advantages and positive effects of the technical solutions to be protected by the present invention from the following aspects:
[0075] First, in view of the technical problems existing in the above-mentioned prior art and the difficulty of solving these problems, this paper closely combines the technical solutions to be protected by the present invention and the results and data during the research and development process, and analyzes in detail and in depth how the technical solutions of the present invention solve the technical problems and some creative technical effects brought about by solving the problems. The specific description is as follows:
[0076] The first is efficiency and security. The present invention considers a hybrid cloud and edge computing network that supports blockchain, which includes two systems: an MEC system and a blockchain system. The MEC system is designed for task execution, while the blockchain system plays a vital role in establishing a secure and trusted transaction platform for the MEC system. In the MEC system, macro base stations (MBS) are positioned at a considerable distance, providing a wide coverage area to ensure that all UEs can obtain uninterrupted wireless communication services. Each small base station (SBSs) demonstrates its proficiency in covering a specific UE, thereby establishing a localized coverage area. SBSs achieve seamless connectivity through high-speed wired connections. Integrated with the cloud server, the MBS enhances its processing capabilities, ensuring that powerful processing capabilities are provided to the terminal.
[0077] Secondly, considering the dynamic nature of Ren, the volatility of network conditions and the random behavior of nodes in consensus, we use PPO to learn the environment state and obtain the optimal joint decision-making strategy for offloading, computing resource allocation and consensus committee selection.
[0078] Second, considering the technical solution as a whole or from the perspective of the product, the technical effects and advantages of the technical solution to be protected by the present invention are described in detail as follows:
[0079] This paper implements a method for optimizing secure task offloading and computing resource allocation in a multi-access edge computing system supported by a consortium blockchain. This method achieves significant progress in the following key aspects:
[0080] 1) Accurate task offloading and computing resource allocation
[0081] This approach significantly improves edge computing resource utilization and system performance by introducing precise task offloading and resource allocation strategies. The MEC system provides computing power to user devices to process tasks, while the consortium blockchain system provides trust and security for user devices during the task offloading process.
[0082] 2) Efficient consensus performance
[0083] This approach uses Delegated Proof of Stake (DPoS) and integrates it into PBFT for block consensus, improving upon the Practical Byzantine Fault Tolerance algorithm. This integrated IPBFT approach balances scalability and complexity, addresses the limitations of the PBFT algorithm, and significantly improves consensus performance.
[0084] 3) Minimize UE task processing cost
[0085] Task processing will bring economic costs, so this method carefully designs the reward mechanism to minimize the cost of UE processing tasks.
[0086] 4) Minimize consensus delays and failure rates
[0087] Consensus latency refers to the time required to reach consensus, and consensus failure can be caused by a variety of reasons, such as unsatisfied constraints and malicious master nodes. This method minimizes consensus latency and failure rate by optimizing offloading strategies, thereby improving system stability and efficiency.
[0088] 5) Integration of Deep Reinforcement Learning
[0089] Combining deep reinforcement learning algorithms with edge computing systems enables the system to self-learn and adjust based on real-time data. By using the PPO algorithm, the system can still make optimal decisions even without explicit instructions.
[0090] 6) Improvement of system stability and reliability
[0091] By accurately calculating and optimizing the immediate rewards and state transitions after action execution, this method significantly enhances the stability and reliability of the system, especially when processing large amounts of data and high-concurrency requests.
[0092] 7) The network's autonomous learning and optimization capabilities
[0093] This method enables the system to continuously optimize its decision-making process through continuous iterative training and experience-based network updates, thereby improving overall performance.
[0094] 8) Joint optimization problem
[0095] This approach jointly optimizes task offloading decisions, computing resource allocation, and master-slave node selection to improve the overall performance and efficiency of the system. This comprehensive optimization strategy not only improves resource utilization but also enhances the system's adaptability and stability in complex network environments.
[0096] 9) Complexity and dynamism
[0097] The complexity of the network and the dynamic nature of the environment make it difficult for traditional optimization methods to effectively deal with these challenges. To address this problem, this method uses the PPO algorithm to improve the adaptability and performance of the system in the face of complex and dynamic environments.
[0098] 10) Real-time and efficiency
[0099] This method can improve the real-time response speed and resource utilization efficiency of the system when processing tasks.
[0100] Third, the optimization method for the MEC system based on IPBFT consensus provided by this invention is based on the use of mathematical models to guide the system's behavior and learning process. The technical effects brought by these mathematical models can be explored based on their characteristics:
[0101] 1) Calculation of rewards
[0102] Even the computational consensus of rewards focuses on the overall network cost, consensus delay and long-term failure rate.
[0103] Cost savings: By directly linking rewards to costs, this approach encourages reducing costs across the network and effectively improves resource utilization efficiency.
[0104] 2) Update of the policy master actor and critic network
[0105] Update the current policy network by randomly sampling small batches of experience data and using the gradient ascent method.
[0106] Policy optimization: By continuously adjusting the policy network parameters, the system can learn and adopt more effective decision-making strategies.
[0107] Improved responsiveness: Using small batches of data enables the network to quickly adapt to environmental changes, enhancing the system's dynamic adjustment capabilities.
[0108] 3) Parameter update formula
[0109] This paper describes how the policy actor and critic network parameters update the target network, involving the current network and target network parameters.
[0110] Policy gradually approaches: By introducing the clipping function, the system can smoothly transition to the new policy and prevent performance fluctuations caused by drastic changes.
[0111] Continuous learning and adaptation: This continuous parameter update mechanism ensures that the system can adapt to long-term environmental changes.
[0112] Fourth, as auxiliary evidence for the inventiveness of the claims of the present invention, it is also reflected in the following important aspects:
[0113] (1) The expected benefits and commercial value of the technical solution of the present invention after transformation are:
[0114] By introducing the PPO algorithm, this invention implements an optimization method for the MEC system based on IPBFT consensus, significantly improving network performance and service quality. This not only meets the growing demand for high-performance computing, but also optimizes resource utilization and reduces network costs. In terms of commercial applications, this invention can be widely used in application scenarios that require efficient processing of large amounts of data while ensuring data security and trustworthiness, bringing significant economic and social benefits to related industries. Therefore, this invention has broad market prospects and huge commercial value. BRIEF DESCRIPTION OF THE DRAWINGS
[0115] Figure 1 This is a flow chart of an optimization method in an MEC system based on IPBFT consensus provided by an embodiment of the present invention;
[0116] Figure 2 It is an applicable scene graph provided by an embodiment of the present invention;
[0117] Figure 3 This is a flowchart of an implementation of an optimization method in an MEC system based on IPBFT consensus provided by an embodiment of the present invention;
[0118] Figure 4 This is a comparison diagram of comprehensive convergence performance provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0119] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0120] 1. Explanatory Examples In order to enable those skilled in the art to fully understand how to implement the present invention, this section provides an illustrative example that expands upon the technical solutions of the claims.
[0121] Example 1
[0122] In response to the problems existing in the existing technology, this implementation provides an optimization method in the MEC system based on IPBFT consensus. Figure 2 This is a diagram of a scenario in which the method of the present invention can be applied. First, it includes two systems: a multi-access edge computing system and a blockchain system. The MEC system is designed for task execution, while the blockchain system plays a crucial role in establishing a secure and reliable transaction platform for the MEC system.
[0123] In the MEC system, there are I user equipments, K small base stations and 1 macro base station, wherein the set of UEs and SBSs are respectively denoted as and The MBS is located at a considerable distance, providing extensive coverage, ensuring that all UEs can obtain uninterrupted wireless communication services. Each SBS demonstrates proficiency in covering specific UEs, thereby establishing a local coverage area. SBSs achieve seamless connectivity through high-speed wired connections. Integrated with cloud servers, MBS enhances its processing power, ensuring strong processing power for terminals.
[0124] The system of the present application works in time slots, where the time range is divided into T time slots, and the set is denoted as Each time slot has a length of At. At the beginning of each time slot, each UE generates a computationally intensive task, which is processed within that time slot. Each UE has limited computing power, which may not be sufficient to process the task while meeting the quality of service constraint requirements. However, the task can be transferred to the edge server of the SBS or the cloud server of the MBS through wireless communication, thereby seeking assistance from MEC. In the following, the MEC server and the cloud server are referred to as computing power providers, and all CPPs are managed by the same operator. After the CPPs process a task, the operator will charge the UE a certain fee for using communication and computing resources in task processing.
[0125] After the task is processed, the CPPs package the information of the computing task, the processing result and the corresponding fee, etc. into a transaction as evidence of the CPPs processing the task and as evidence of the UE paying for the corresponding task processing service. Then, each CPP uploads the transaction it generates to the blockchain system, where the transaction will first be verified and then packaged into a new block. Once the network reaches consensus on the new block, it will be appended to the distributed ledger, i.e. the blockchain of each node. Only when the block is successfully added to the blockchain will the task processing result be returned to the user and the CPPs will obtain the corresponding service fee. Otherwise, the task processing will fail and the consensus nodes will be punished. When a block is successfully added to the blockchain, the transactions it contains will be irreversible, thereby effectively guaranteeing the privacy and security of the user. A controller is deployed in the MBS, responsible for collecting and maintaining lists of terminal information, edge server details, service details and application information, etc. to optimize network configuration on demand. The controller receives and trains data through the DRL algorithm, enabling it to make intelligent joint optimization decisions for MEC and the blockchain system.
[0126] Embodiment 2
[0127] Given the dynamic nature of the network and the uncertainty of information acquisition, the problem becomes quite complex. To effectively address this challenge, we remodel the problem as a Markov decision process (MDP). Because the variables involved are discontinuous, we choose to use the PPO algorithm, a DRL algorithm designed for solving reinforcement learning problems in continuous action spaces and capable of supporting real-time online decision making.
[0128] like Figure 1 As shown, the embodiment of the present invention provides an optimization method in a MEC system based on IPBFT consensus, which specifically includes the following steps:
[0129] S101. Initialize the parameters θ and ω of the agent's policy actor network and critic network, the parameters θ′ and ω′ of the sampled policy actor network and critic network, the number of episodes, initialize the learning rates μ and σ corresponding to the critic network and policy network, the discount factor γ, initialize the experience pool; initialize the network layout parameters, such as the data size D of the input task i (t), task workload C i (t) and other parameters;
[0130] S102: Initialize the state of the agent, the agent interacts with the environment, and the main policy network generates actions based on the current policy;
[0131] S103, the agent executes the generated action, obtains an immediate reward based on the executed action, and transfers the environment state to the next state;
[0132] S104, sampling the entire segment in the environment according to the parameters of the sampling strategy actor network, and storing the trajectory in memory;
[0133] S105. Calculate the discount reward;
[0134] S106, calculating the advantage function, adding the clipping factor to constrain the update rate, and calculating the objective function at the same time;
[0135] S107, update the parameters of the policy actor and critic network;
[0136] S108. Update the sampling strategy actor and critic network parameters according to the updated strategy actor and critic network parameters;
[0137] S109: Repeat iterative training, select the optimal action according to each state, obtain the maximum benefit, and finally obtain the optimal computing resource allocation and computing offloading strategy.
[0138] Furthermore, in S102, the agent interacts with the environment; at the beginning of each round, the system state s(t) is initialized. The state consists of three parts, namely, the state of the task s task (t), the state of the network s net (t) and the state information s in the consensus con (t), as follows:
[0139] s(t)={s task (t),s net (t),s con (t)}.
[0140] The status of the task task (t) is given by the following formula:
[0141] s task (t)={D(t),C(t)},
[0142] The input data volume of the task D(t) = {D i (t)},i∈I; the workload of the task C(t)={C i (t)},i∈I.
[0143] Network Status net (t) is given by:
[0144] s net (t) = {A(t), R(t), U(t), r m (t),f m (t)},
[0145] Among them, the link availability between UE and SBSs is The link connection rate between UE and SBSs is Number of computing resource blocks of SBSs The link connection rate between UE and MBS is r m (t); The computing resources allocated by MBS to each task are f m (t+1).
[0146] State information s during the consensus process con (t) is given by the following formula:
[0147]
[0148] Where W(t-1)={W k (t)} is the credit value of the node at time slot t-1, Malicious node indicators The number of malicious nodes is When the master node is a malicious node, the indicator of whether the master node k has reached consensus When the non-master node is a malicious node, the indicator of whether the non-master node k has reached consensus When the master node is a normal node, the malicious delay of master node k in the Pre-prepare stage is When the non-master node is a normal node, the malicious delay of the non-master node k in the Commit phase is The number of failures at the previous moment is N fail (t-1); the average consensus delay is
[0149] At each time slot t, the agent decides its action a(t) based on the state s(t) given by
[0150]
[0151] in, It is the task offloading decision of UE; Make CRB allocation decisions during task processing; Select decision indicators for consensus nodes, where For master node selection, y k (t) Select decision indicators for non-master nodes.
[0152] Furthermore, the consensus calculation of the reward in S103: the agent executes the generated action, obtains an immediate reward based on the executed action, and transfers the environment state to the next state is as follows:
[0153]
[0154] In the above formula, Object(t) represents our objective function.
[0155] Furthermore, the S106: calculating the advantage function:
[0156]
[0157] At this time there are:
[0158] δ t =r(t)+γV(s t+1 ;w)-V(s t ;w).
[0159] In order to improve the performance, PPO introduces a θ′ Further improvement of the actor objective function of (θ). By adding a clipping factor to constrain the update rate, the PPO actor can be updated by maximizing the objective function, as follows:
[0160]
[0161] Where ∈ is a hyperparameter, the clip function converts (π(a t ∣s t ;θ)) / π(a t ∣s t ; θ′) is restricted to the range [1-,1+]. This approach ensures that after minimizing the clip function, the two distributions remain relatively close and avoid significant differences.
[0162] Furthermore, the S107: update the parameters of the strategy actor and critic network; the actor parameters are updated by the following formula:
[0163]
[0164] This algorithm considers the mean square error function of the value estimation and gives the loss function of the critic network:
[0165] L critic (w)=[V(s t+1 ;w)-V(s t ;w)] 2 ,
[0166] At the same time, it can be updated by this formula:
[0167]
[0168] Among them, δ t represents the TD error.
[0169] In order to elaborate on the optimization method in the MEC system based on IPBFT consensus, the present invention provides two specific application embodiments, including key details of the implementation scheme.
[0170] Application Example 1: Smart City IoT Application
[0171] 1) System initialization
[0172] Deploy the policy actor and critic network in the smart city monitoring center. Initialize network parameters, including the number of rounds, training steps, and learning rate. Deploy multiple IoT devices (such as edge computing nodes and sensors) and base stations to collect and transmit data.
[0173] 2) Interaction between the agent and the environment
[0174] Each IoT device acts as an intelligent agent and generates actions based on the state of its monitored environment (such as adjusting the sensor's acquisition frequency, the edge node's data processing priority, or the frequency of data uploads). Actions are generated based on the current policy.
[0175] 3) Action execution and reward acquisition
[0176] After performing an action, the agent receives an immediate reward based on the effect of the action (e.g., data transmission rate, data accuracy). The system then moves to the next state, which may be due to changes in the environment or changes in the data stream at different time periods.
[0177] 4) Optimization and update
[0178] Based on the collected data and rewards, the policy actor and critic network are updated to optimize the sensor control strategy and data processing priority of edge nodes. Reinforcement learning algorithms are used to gradually improve resource allocation, such as data transmission rate and allocation of computing resources.
[0179] 5) Energy consumption and performance optimization
[0180] The performance of the PPO algorithm is evaluated through simulations or field tests. Key metrics such as communication network coverage, signal quality, and data transmission rate are analyzed and compared with traditional methods. Based on the evaluation results, the algorithm parameters are adjusted and optimized to further improve the performance and stability of the communication network. Through the above steps, the proposed IPBF consensus-based MEC network joint optimization method can monitor a large number of IoT devices in smart cities in real time, such as smart lighting control and environmental monitoring.
[0181] Application Example 2: Financial and Payment Systems
[0182] 1) System initialization
[0183] In the financial and payment systems, the center deploys the PPO network and initializes network parameters, while also deploying necessary infrastructure such as servers and databases. These facilities not only support real-time processing and data storage of payment transactions, but also ensure that the system can operate safely and efficiently, meeting the security, stability, and performance requirements of financial services.
[0184] 2) Interaction between the agent and the environment
[0185] Each server acts as an intelligent agent, continuously interacting with the environment and generating actions based on the current state of the environment. These actions may include resource allocation, data processing optimization, and security control to ensure the stability and performance of financial and payment systems.
[0186] 3) Action execution and reward acquisition
[0187] After executing the current action, the system receives immediate rewards based on key indicators such as the rationality of resource allocation, data transmission rate, and security. The system then moves to the next state based on changes in the environment or time period.
[0188] 4) Optimization and update
[0189] The collected data and rewards are used to update the parameters of the PPO network. Through continuous iteration and training, the algorithm gradually improves the stability and efficiency of the communication system, thereby ensuring the security and reliability of the payment system.
[0190] 5) Payment system reliability
[0191] The PPO algorithm of this invention enables intelligent management and optimization of financial and payment systems. This algorithm focuses on improving the reliability and stability of communication links, ensuring the security and integrity of data during transmission. Furthermore, the PPO algorithm optimizes the allocation and utilization efficiency of system resources, significantly improving the overall system's response speed and processing capabilities, thereby enhancing the user experience and system reliability. Through the above steps, this invention provides an efficient and intelligent solution for financial and payment system security.
[0192] In these two embodiments, the optimization method in the MEC system based on IPBFT consensus provides an efficient, reliable and energy-saving solution, which is suitable for different application scenarios, from smart cities to financial and payment systems, demonstrating its wide application potential and technical advantages.
[0193] To comprehensively evaluate the performance of our invention, we compared it with several representative benchmark algorithms. These benchmark algorithms each have their own unique characteristics and advantages, effectively handling similar problems. By comparing these algorithms under the same conditions, we can more clearly understand the advantages of our invention and the potential for improvement.
[0194] 1) Random Offloading: Each UE's offloading decision is randomly determined. Each UE's task can be independently processed on its own, the SBS, or the MBS. However, a task can only be processed at one location. This optimizes the allocation of computing resources and the selection of the blockchain committee.
[0195] 2) Random Committee: Blockchain committee nodes are randomly selected from all SBSs, including master nodes and non-master nodes. This optimizes task offloading decisions and computing resource allocation decisions.
[0196] 3) The present invention: It represents the proposed PPO-based joint task offloading, computational resource allocation, and blockchain committee selection algorithm.
[0197] exist Figure 4 In this paper, we demonstrate the convergence performance of all the above algorithms with default parameter settings. The convergence performance is evaluated based on three performance indicators: convergence speed, convergence stability, and reward value. The results show that the rewards of the algorithm proposed in this paper tend to stabilize around the 100th round, and the reward return thereafter is approximately 5.2×10 4 The random unloading algorithm converges around the 170th round, with a reward of 2×10 4 The random committee algorithm converges around the 150th round, and its reward is 1.2×10 4 By carefully observing the three curves, it can be inferred that the algorithm proposed by the present invention has the fastest convergence speed and outperforms the two benchmark algorithms in terms of performance.
[0198] 3. Evidence of the effects of the embodiments: The embodiments of the present invention have achieved some positive effects during the development or use process, and indeed have great advantages over the existing technology. The following content describes them with reference to the data, charts, etc. of the experimental process.
[0199] Figure 4 In this article, we provide a comprehensive performance evaluation of the proposed algorithm by comparing it with three baseline algorithms. The graph clearly shows that the return of the proposed algorithm stabilizes around episode 100, after which its return fluctuates around 5.2×10^4. The PPO-based-committee-selection algorithm converges around episode 270, with its return fluctuating around 4.2×10^4. Random-offloading converges around episode 170, with its return fluctuating around 2×10^4. A careful observation of the three curves shows that the proposed algorithm demonstrates the fastest convergence speed and outperforms the two baseline algorithms in terms of performance.
[0200] It should be noted that the embodiments of the present invention can be implemented by hardware, software, or a combination of software and hardware. The hardware portion can be implemented using dedicated logic; the software portion can be stored in a memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated design hardware. Those skilled in the art will appreciate that the above-mentioned devices and methods can be implemented using computer-executable instructions and / or contained in processor control code, for example, such as a carrier medium such as a disk, CD or DVD-ROM, a programmable memory such as a read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and modules of the present invention can be implemented by hardware circuits such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, or programmable hardware devices such as field programmable gate arrays, programmable logic devices, etc., can also be implemented by software executed by various types of processors, or can be implemented by a combination of the above-mentioned hardware circuits and software, such as firmware.
[0201] The above description is only a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions and improvements made by any technician familiar with this technical field within the technical scope disclosed by the present invention and within the spirit and principles of the present invention should be covered by the scope of protection of the present invention.
Claims
1. An optimization method in a multi-access edge computing system based on IPBFT consensus, characterized in that: The method is applied to the agent's policy actor network and value critic network, using deep reinforcement learning (DRL) methods for decision making, including: Initialize the parameters of the actor network and the critic network, and set the training-related hyperparameters. The agent interacts with the environment based on the current strategy, performs actions, and performs state transitions. The entire segment of the environment is sampled using the parameters of the sampling actor network, and the trajectory is stored in memory. The discounted reward, advantage function, and objective function are calculated. The parameters of the actor network and the critic network, as well as the parameters of the sampling actor network and the sampling critic network, are updated. Training is repeated until the strategy converges, and the trained strategy is used for computation offloading and resource allocation. The method comprises the following steps: S101. Initialize the parameters θ and ω of the actor network and the critic network of the agent, the parameters θ′ and ω′ of the sampling actor network and the sampling critic network, the number of episodes, initialize the learning rates μ and σ corresponding to the actor network and the critic network, the discount factor γ, initialize the experience pool; initialize the network layout parameters, including the data size D of the input task i (t), task workload C i (t); S102: Initialize the state of the agent, the agent interacts with the environment, and the policy network generates actions based on the current policy; S103, the agent executes the generated action, obtains an immediate reward based on the executed action, and transfers the environment state to the next state; S104, sampling the entire segment in the environment according to the parameters of the sampling actor network, and storing the trajectory in memory; S105. Calculate the discount reward; S106, calculating the advantage function, adding the clipping factor to constrain the update rate, and calculating the objective function at the same time; S107, update the parameters of the actor network and the critic network; S108. Update the sampling actor and sampling critic network parameters according to the updated actor network and critic network parameters; S109: Repeat iterative training, select the optimal action according to each state, obtain the maximum benefit, and finally obtain the optimal computing resource allocation and computing offloading strategy.
2. The optimization method in the multi-access edge computing system based on IPBFT consensus according to claim 1, characterized in that: In step S102, the agent interacts with the environment; at the beginning of each round, the system state s(t) is initialized; wherein the state consists of three parts, namely, the state of the task s task (t), the state of the network s net (t) and the state information s in the consensus con (t), as follows: s(t)={s task (t),s net (t),s con (t)}. The status of the task task (t) is given by the following formula: s task (t)={D(t),C(t)}, Among them, the input data volume of the task The workload of the task Network Status net (t) is given by: s net (t)={A(t),R(t),U(t),r m (t),f m (t)}, Among them, the link availability between the user equipment UE and the small base stations (SBSs) is The link connection rate between UE and SBSs is Number of computing resource blocks of SBSs The link connection rate between UE and macro base station (MBS) is r m (t); The computing resources allocated by MBS to each task are f m (t+1); State information s during the consensus process con (t) is given by the following formula: Where W(t-1)={W k (t)} is the credit value of the node at time slot t-1, Malicious node indicators The number of malicious nodes is When the master node is a malicious node, the indicator of whether the master node k has reached consensus When the non-master node is a malicious node, the indicator of whether the non-master node k has reached consensus When the master node is a normal node, the malicious delay of master node k in the Pre-prepare stage is When the non-master node is a normal node, the malicious delay of the non-master node k in the Commit phase is The number of failures at the previous moment is N fail (t-1); the average consensus delay is At each time slot t, the agent decides its action a(t) based on the state s(t). The action a(t) is: in, It is the task offloading decision of UE; Make CRB allocation decisions during task processing; Select decision indicators for consensus nodes, where Select the indicator for the master node, y k (t) Select decision indicators for non-primary nodes.
3. The optimization method in the multi-access edge computing system based on IPBFT consensus according to claim 1, characterized in that: In step S103, the agent executes the generated action, obtains an immediate reward based on the executed action, and transfers the environment state to the next state. The calculation consensus of the reward is as follows: In the above formula, Object(t) represents our objective function.
4. The optimization method in the multi-access edge computing system based on IPBFT consensus according to claim 1, characterized in that S106: Calculate the advantage function: At this time there are: δ t =r(t)+γV(s t+1 ;w)-V(s t ;w). PPO (proximal policy optimization) introduces a θ′ A further improvement of the actor objective function of (θ) is to constrain the update rate by adding a clipping factor and update the PPO actor network by maximizing the objective function, which is formulated as: The clip function converts (π(a t ∣s t ;θ)) / π(a t ∣s t ; the value of θ′) is restricted to the range [1-, 1+]; ensuring that the two distributions remain relatively close after minimizing the clip function.
5. The optimization method in the multi-access edge computing system based on IPBFT consensus according to claim 1, characterized in that: S107: Update the parameters of the strategy actor network and the critic network; the parameters of the actor network are updated by the following formula: Considering the mean square error function of the value estimation, the loss function of the critic network is given: L critic (w)=[V(s t+1 ;w)-V(s t ;w)] 2 , Updated by this formula: Among them, δ t Indicates the time difference TD error.
6. A multi-access edge computing system based on IPBFT consensus according to the method of claim 1, characterized in that: The system comprises: System initialization module, which initializes all necessary parameters, including initializing experience memory, actor network parameters θ and critic network parameters w, sampling actor network parameters θ' and sampling critic network parameters w'; Configuration module, used to set application-specific parameters, including the data size of the input task; The agent module is used to generate actions based on the current network state at the beginning of each cycle; it is used to collect data samples, calculate the advantage function, and update the policy and value network during the policy evaluation process; Action execution module, which is used to perform task offloading, computing resource allocation in task processing, and selection of master nodes and non-master nodes in the consensus process; The reward acquisition module is used to execute actions and calculate instant rewards. The reward acquisition module is designed based on whether the system's constraints are met. If all constraints are met, rewards are obtained, otherwise penalties are obtained; The state transfer module is used to transfer the system state from the current state to the next state; The experience replay module is used to store the experience tuples of each system state, action, reward, and next state; The data sampling module is used to extract certain fragments from the stored experience for learning; The network update module is used to update the actor network and critic network based on the data of the experience playback module; Parameter update module, used for parameter update of actor network and critic network, as well as parameter update of sampling actor network and sampling critic network.
7. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the optimization method in the multi-access edge computing system based on IPBFT consensus as claimed in claim 1.
8. A computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor executes the steps of the optimization method in the multi-access edge computing system based on IPBFT consensus as claimed in claim 1.