A method and system for edge computing collaborative task offloading based on dynamic reputation value
Through the edge computing collaborative task offload method based on dynamic reputation values, the node reputation is evaluated using beta distribution and blockchain technology, combined with deep reinforcement learning optimization task offload strategy, the problem of resource limitations in traditional edge computing architecture and single point failure of reputation system is solved, and efficient and flexible task success rate is maximized.
Patent Information
- Application Number
- CN202510448347.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2045-04-10
AI Technical Summary
Traditional edge computing architectures are difficult to meet resource limitations under high concurrency and low latency requirements, and the existing reputation mechanism relies on centralized systems to have a risk of single point of failure and data tampering, making it difficult to adapt to dynamically changing edge environments.
The edge computing collaborative task offload method based on dynamic reputation values is adopted, node reputation is evaluated through the beta distribution three-factor reputation evaluation method, and reputation storage is carried out in combination with blockchain technology, and deep reinforcement learning is used to optimize task offload strategy to build a decentralized part to observe Markov decision-making process, and training near-end strategy optimization algorithms to obtain the optimal offload strategy.
It improves adaptability and task success rate in dynamic edge environments, reduces the possibility of malicious nodes, avoids excessive dependence on reputation values, and realizes flexible utilization and efficient coordination of resources.
Smart Images

Figure CN120151947B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of edge computing technology, and specifically relates to an edge computing collaborative task offloading method and system based on dynamic reputation value. Background Art
[0002] With the increasing popularity of the Internet of Things (IoT) (e.g., smart homes, smart cities, and smart factories) and advances in wireless technology, a multitude of emerging services and applications (e.g., virtual reality, telemedicine, and autonomous driving) are placing stringent demands on edge computing for high concurrency and low latency. Traditional edge computing architectures, due to the resource limitations of a single node, struggle to meet these requirements. Edge computing can achieve resource sharing through multi-node collaboration, significantly improving resource utilization and user experience quality.
[0003] In edge computing, collaborative task offloading has attracted widespread attention, and numerous methods have been developed to optimize system performance. However, these approaches still face numerous shortcomings. Traditional approaches based on mathematical optimization or game theory are either computationally complex, difficult to adapt to dynamic environments, and prone to local optimal solutions. Alternatively, they rely on strict assumptions to compute static Nash equilibria, making it difficult to fully describe the real world. Most research based on deep reinforcement learning (DRL) attempts to maximize latency or energy consumption, but assumes that all computing resources in the system are trustworthy. In reality, since edge nodes may come from different manufacturers or service providers, some may disrupt collaboration due to selfish motives. To address this, many studies have considered introducing reputation mechanisms into offloading schemes to distinguish honest from malicious nodes, using reputation as a key factor in selecting collaborative edge nodes. However, such schemes rely on centralized reputation systems, which may present single points of failure or data tampering risks. Furthermore, they lack flexibility in complex and dynamically changing edge environments and cannot effectively optimize resource competition and utilization. Summary of the Invention
[0004] The purpose of the present invention is to provide an edge computing collaborative task offloading method and system based on dynamic reputation value.
[0005] In a first aspect, the present invention provides an edge computing collaborative task offloading method based on dynamic reputation value, which comprises the following steps:
[0006] A task model for collaborative task offloading is established; the task model includes a task generation model, a task transmission model, a task calculation model, and a task delay model, and a task offloading decision optimization problem is formulated with the goal of maximizing the task success rate; based on the constructed task model, the task offloading decision optimization problem is transformed into a decentralized partially observable Markov decision process; wherein, based on the observations of the edge node in each time slot, an observation space is constructed; based on all possible actions of the edge node in each time slot, an action space is constructed; by introducing a reputation value, the reward for the current edge node to offload the task to the target edge node is set; if the task is discarded due to exceeding the deadline or fails to be processed due to malicious behavior, a negative reward is imposed on the current edge node; if the task is correctly processed within the deadline, a positive reward is imposed on the current edge node, and the positive reward includes the reputation value of the target edge node; blockchain technology is used to store the reputation value of the edge node;
[0007] A neural network model for generating task offloading strategies is constructed, and experience fragments consisting of observations, actions, and rewards are constructed based on the interaction between edge nodes and the environment in a decentralized partially observable Markov decision process. The experience fragments at different times are stored in a rolling buffer, and the rolling buffer is used to train the neural network model. Based on the observations of edge nodes, the trained neural network model is used to obtain the optimal task offloading strategy. The task offloading strategy includes the actions performed by each edge node after receiving the task.
[0008] As a preference, add the task difficulty factor to the reputation value of the edge node , reputation value The method to obtain is as follows:
[0009]
[0010] in, and Respectively u After the credit value is updated j The number of positive and negative feedback of edge nodes; For the forgetting factor; is the historical task completion rate factor; is the action feedback. If the edge node processes the task correctly within the deadline, If the edge node performs malicious behavior or fails to process the task within the deadline, ; ; N is the number of edge nodes.
[0011] As a preference, the task difficulty factor The method to obtain is as follows:
[0012]
[0013] in, Indicates the task size; Indicates task processing density; Indicates the maximum size of the task; Indicates the maximum value of task processing density.
[0014] Preferably, the malicious behavior includes denial of service, refusal to pledge, delayed processing, and error processing.
[0015] Preferably, the observations of the edge nodes include task size, task processing density, task arrival probability, computing power, transmission rate between the edge node and other edge nodes, and historical reputation data of all edge nodes.
[0016] As a preference, the interaction process between the edge node and the environment is as follows:
[0017] When the task arrives at the current edge node, the current edge node selects an action based on the current observation and actor network; if the current edge node offloads the task to the target edge node for processing, the current edge node publishes a smart contract, and the target edge node pledges tokens; after the pledge is completed, the current edge node offloads the task to the target edge node for processing; if the target edge node refuses service or refuses to pledge, the task is processed locally by the current edge node; after the task processing is completed or the deadline is reached, the reputation value of the target edge node is dynamically updated.
[0018] Preferably, in the task transmission model, the tasks of each edge node are transmitted in a first-in-first-out order. The tasks that have completed the transmission are automatically placed in the processing queue of the target computing node and are no longer forwarded to other edge nodes. The transmission time of the transmission task between different edge nodes is expressed as the task size divided by the transmission rate between different edge nodes.
[0019] As an optimal choice, the task delay model constructs the total delay for task processing completion based on transmission delay, processing delay and waiting delay. , depending on the task processing method, the total delay The method to obtain is as follows:
[0020]
[0021] in, Indicates that the task is on the edge node Processing wait time in; Indicates that the task is from the edge node To the edge node Transmission waiting time; Indicates that the task is processed at the current edge node. Indicates that the task is offloaded to other edge nodes for processing;
[0022] If the total delay of a task is greater than the deadline, the task will be discarded from the network and the deadline will be used as the total delay of the task.
[0023] Preferably, the neural network model adopts a proximal policy optimization algorithm; during the training process, the parameters in the actor network are updated by a truncated proxy objective function and an entropy regularization term, and the parameters in the critic network are updated by a mean square error loss of a value function.
[0024] In the second aspect, the present invention provides an edge computing collaborative task offloading system based on dynamic reputation value, including an Internet of Things device, an edge node and a smart contract; the Internet of Things device is used to generate tasks and transmit the tasks to the edge node for processing; the edge node is used to process tasks from the Internet of Things device and tasks offloaded by other edge nodes; the edge computing collaborative task offloading system is used to execute the above-mentioned edge computing collaborative task offloading method; the smart contract is responsible for the execution and supervision of the task offloading process, and stores relevant information; the relevant information includes task description, task requirements, token staking rules, task result verification rules and reputation update rules.
[0025] The present invention has the following beneficial effects:
[0026] 1. The present invention uses a three-factor reputation evaluation method based on Beta distribution to efficiently and accurately evaluate the reputation value of edge nodes, and introduces the historical task completion rate factor and task difficulty factor into the reputation update algorithm, expanding the multidimensionality of the reputation update algorithm. It enables edge nodes to add trust indicators based on historical performance and task difficulty to enhance the evaluation of edge node reputation values, improve adaptability in dynamic edge environments, and improve the processing capability of high-dimensional state space.
[0027] 2. This invention optimizes task offloading strategies through deep reinforcement learning, integrating reputation into the agent training process. This guides edge nodes to learn trustworthy task offloading strategies, minimizing the likelihood of malicious nodes offloading tasks while maximizing task success rates. This approach avoids over-reliance on reputation, improves the flexibility of offloading decisions, and effectively coordinates resource competition and utilization. Furthermore, this invention achieves decentralized storage of reputation values through blockchain, avoiding single points of failure and data tampering. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 This is the overall flow chart of the edge computing collaborative task offloading method in the present invention.
[0029] Figure 2 Schematic diagram of the edge computing collaborative task offloading system in the present invention.
[0030] Figure 3 This is a diagram of the training framework of the proximal strategy optimization algorithm in the present invention. DETAILED DESCRIPTION
[0031] The present invention will be further described below with reference to the accompanying drawings.
[0032] like Figure 1 This paper presents a method for collaborative edge computing task offloading based on dynamic reputation. The proposed system comprises IoT devices, edge nodes, and smart contracts. Edge nodes play a dual role: computing resource providers and blockchain nodes. Edge nodes not only process tasks from local IoT devices and tasks offloaded from other edge nodes, but also participate in the blockchain network, jointly maintaining the reputation blockchain, including verifying transactions, generating blocks, and broadcasting blockchain status. Reputation plays a key role in this process, serving as a metric to quantify node behavior and reflect its trustworthiness by recording its historical performance. Furthermore, edge nodes can be categorized as honest or malicious. Malicious nodes may violate the system protocol through behaviors such as denial of service, delayed task processing, or submitting erroneous results. Smart contracts execute and oversee the task offloading process and store relevant information, including task descriptions, task requirements, token staking rules, task result verification rules, and reputation update rules.
[0033] The edge computing collaborative task offloading method includes the following steps:
[0034] Step S1: Establish a blockchain-supported edge computing collaborative task offloading model, task generation model, task transmission model, task calculation model, and task delay model, and formulate an optimization problem. The specific steps are as follows:
[0035] Step S1-1: Establishing a blockchain-supported edge computing collaborative task offloading model
[0036] like Figure 2 As shown, construct a connected graph To represent the edge computing collaborative task offloading model; where, is the set of edge nodes, ; e i For the i edge nodes; N is the number of edge nodes; is the set of links connecting different edge nodes, ; For the i The edge node and j The time is discretized into multiple links with constant duration. The time slot index is Indicates that ; T is the total time. In each time slot At the beginning, set the reputation value set of all edge nodes ;in, R i For the i The reputation value of each edge node. It can receive tasks from local IoT devices and then decide whether to process the tasks locally or offload them to other edge nodes.
[0037] Step S1-2: Establishing a task generation model
[0038] In the time slot From IoT devices to edge nodes Mission Use triples To describe; among them, Indicates the task size; Indicates the task processing density, that is, the number of CPU cycles required to process one data unit; Indicates the deadline for task processing. In addition, in the time slot Reaching the edge node Mission Following the Bernoulli distribution, recorded as the probability of task arrival .
[0039] Step S1-3: Establishing a task transfer model
[0040] Each edge node maintains A sending queue is used to transmit tasks to other edge nodes; at the same time, A receive queue is used to receive tasks from local IoT devices and other edge nodes. Tasks are transmitted in a first-in first-out (FIFO) order. If there is a task being transmitted in the current transmission queue, other tasks must wait. Tasks that have completed transmission are automatically placed in the processing queue of the target computing node, awaiting further processing and are no longer forwarded to other computing resources. The transmission rate of each link varies with time and is unknown a priori. Transmission Task From the edge node To the edge node Transmission time Expressed as:
[0041]
[0042] in, Indicates time slot Edge Node and The transmission rate between.
[0043] Step S1-4: Establishing a task calculation model
[0044] Whether it is a task processed locally or a task received from other edge nodes, it will be automatically placed in the processing queue. Tasks in the processing queue are processed in FIFO order. If other tasks occupy computing resources, the current task must wait in the processing queue. Task Processing time Expressed as:
[0045]
[0046] in, Represents an edge node computing power.
[0047] Step S1-5: Establishing task delay model
[0048] The total delay for task processing to complete includes transmission delay, processing delay, and waiting delay, but does not include the delay of task offloading decision because this delay is very short. Total delay to complete Depending on how the task is handled, it can be defined as:
[0049]
[0050] in, Indicates a task At the edge node Processing wait time in; Indicates a task From the edge node To the edge node Transmission waiting time; Indicates that the task is processed locally. Indicates that the task is offloaded to other edge nodes for processing.
[0051] In addition, the task Should be within its deadline Completed before, Otherwise, the task will expire and be discarded from the network. At this time, the total delay of the task is Will be recorded as its deadline .
[0052] Steps S1-6: Formulate the optimization problem
[0053] The overall goal of this invention is to maximize the mission success rate by minimizing the possibility of offloading tasks to malicious nodes and enhancing the collaboration between edge nodes. is defined as follows:
[0054]
[0055] in, Indicates the number of tasks successfully processed; Indicates the number of tasks that were discarded because they were not processed within the deadline; Indicates the number of tasks that failed due to malicious behavior of edge nodes.
[0056] Step S2: Introduce a three-factor reputation evaluation method based on beta distribution in the edge computing architecture supported by blockchain to evaluate the credibility of each edge node and provide a trustworthy platform for collaborative edge nodes. Each edge node is assigned a medium initial reputation value in the initial time slot; in this embodiment, the initial reputation value is 0.5. If the edge node At the deadline If the correct result is submitted to the smart contract, feedback will be obtained ; If the edge node Feedback will be given if malicious behavior (including denial of service, refusal to stake, delayed processing, incorrect processing) or failure to respond within the deadline causes the task to be discarded due to timeout. Based on this, u The number of positive feedbacks from edge nodes after the secondary reputation value update , negative feedback number and new reputation value The expression is:
[0057]
[0058] in, is the forgetting factor, which is used to introduce the time decay mechanism; The historical task completion rate factor is used to prevent malicious nodes from quickly restoring their reputation by occasionally processing tasks correctly. is the task difficulty factor, which is used to reflect the computational complexity of the task.
[0059] In a dynamically changing edge environment, the forgetting factor Ensure that the system can gradually dilute the impact of historical behavior, thereby enhancing the real-time nature of reputation. If an edge node can complete a complex task, then it can be given more reputation; if an edge node cannot even complete a simple task, then its reputation should be significantly reduced. Task Difficulty Factor Expressed as:
[0060]
[0061] in, Indicates the maximum size of the task; Indicates the maximum value of task processing density.
[0062] Historical task completion rate factor and task difficulty factor The existence of expands the multidimensionality of the reputation update algorithm, which is equivalent to adding a trust indicator based on historical performance and task difficulty to the node to enhance the evaluation of the node's reputation value.
[0063] Step S3: Convert the optimization problem formulated in step S1 into a decentralized partially observable Markov decision process (DPOMDP). The specific steps are as follows:
[0064] Step S3-1: Modeling DPOMDP
[0065] The optimization problem formulated in step S1-6 is formalized as a DPOMDP, which can be represented by the tuple Indicates. Among them, Indicates a A set of agents, each of which corresponds to an edge node. is the state space, Indicates time slot Environmental conditions at the time; is the action space, in the time slot , each agent Will choose an action , forming a joint action ; is the state transition function, Indicates the environmental status All agents perform joint actions After that, the environment transfers to the next gap environment state probability; is the reward function. When the environment state is transferred, each agent You will receive a reward ; For the observation space, in partially observable scenes, each agent Only local information of itself can be observed, and based on a decentralized strategy Select an action; Represents a given observation The probability distribution of all possible actions under is the initial environment state distribution, that is, the initial environment state By distribution to obtain; is a discount factor used to balance the importance of immediate rewards and future rewards.
[0066] Step S3-2: Define reputation-driven multi-dimensional observations
[0067] In the time slot , each edge node Observe its local information, edge nodes In the time slot Observation Defined as:
[0068]
[0069] in, is the task size; Processing density for tasks; is the task arrival probability; For computing power; For edge nodes Transmission rate with other edge nodes; H is the historical reputation data of all edge nodes, is a A matrix of size, representing the total from the past to the present The historical reputation data of the step.
[0070] In addition, due to the local information The elements in have different numerical scales, and each element is divided by its maximum value to ensure that each element is in the range of [0,1]. This reduces the impact on training stability and convergence.
[0071] Step S3-3: Define actions
[0072] When the task Reaching the edge node When the edge node Execute actions based on current observations, i.e. select appropriate edge nodes to process tasks. Therefore, the action space This should include all computing resources and can be defined as:
[0073]
[0074] in, To indicate that the edge node has not received the task; Indicates that the task is iEdge node processing; .
[0075] If the edge node Select Action , it means that the task will be processed locally. In addition, when the edge node does not need to make an offload decision because it has not received a task, it can select an action by blocking other available actions. .
[0076] Step S3-4: Define reputation-driven rewards
[0077] The immediate reward should take into account the credibility of the task offloading decision and whether the task processing is successful. If it is discarded due to exceeding the deadline or fails due to malicious behavior, a negative reward will be imposed, i.e. ; If the task If the correct result is successfully processed and submitted within the deadline, a positive reward will be obtained, i.e. ;in, is the reputation value of the target edge node. This reward design encourages the agent to strike a balance between high-confidence offloading decisions and a high success rate. If the agent overly pursues high-confidence offloading decisions, it may overload some high-reputation edge nodes, resulting in more tasks being dropped.
[0078] Step S4: Solve the DPOMDP in step S3 based on the proximal policy optimization algorithm to find the optimal task offloading strategy. The specific steps are as follows:
[0079] Step S4-1: Constructing a neural network model
[0080] like Figure 3 As shown, let the proximal strategy optimization algorithm be Parameterized Actor Networks To model each edge node The approximate offloading strategy of , that is, the action performed by each edge node after receiving the task; let the proximal strategy optimization algorithm be composed of Parameterized Critic Network To model each edge node The approximate value function of each edge node. Passed through the input layer to the actor network and critics network Then, a Long Short-Term Memory (LSTM) layer is used to classify the reputation value of H To predict the node reputation value of the next time slot Afterwards, two fully connected (FC) layers are used to learn the To action The mapping and observation The mapping from the actor network to the value function. Both the actor network and the critic network are prefixed with an LSTM layer, which is commonly used to learn the temporal dependencies of sequential observations and predict future changes in time series. By including historical reputation data as part of the observation input, the agent can understand the dynamic changes of the reputation mechanism and make better offloading decisions.
[0081] Step S4-2: Construct a rolling buffer for training the neural network model based on the interaction between the edge node and the environment. In DPOMDP, the complete interaction process between the edge node and the environment can be divided into five stages:
[0082] 1) Task offloading decision
[0083] When the task Reaching the edge node After that, the edge node According to current observations and Actor Network Select Action .
[0084] 2) Smart Contract Generation
[0085] If the edge node Offload tasks to target edge nodes Processing, then the edge node A smart contract will be published in the blockchain network; otherwise, the task will be directly handled by the edge node Local processing.
[0086] 3) Token Staking and Task Offloading
[0087] Target edge node A certain amount of tokens must be pledged based on the task requirements and one’s own reputation value in order to process the task. The smart contract needs to verify the completion of the pledge and lock the tokens. Encryption technology will be used to offload the actual task to the target edge node If the target edge node If a node refuses to provide services or pledge, its reputation will be punished in the future, and the task will be handled by the edge node. Local processing.
[0088] 4) Task result verification and rewards.
[0089] After the task processing is completed, the target edge node Submit the result in the form of hash commitment to the smart contract for verification. If the correct result is submitted within the deadline, the staked tokens will be returned, and additional reward tokens will be obtained as incentives, and the reputation will also increase. Errors or delays in processing tasks will be considered malicious behavior, and the staked tokens will be transferred to edge nodes. To compensate for the losses caused by mission failure, and their reputation will be punished later.
[0090] 5) Credit value update.
[0091] Smart contracts are based on target edge nodes The current performance of the target edge node is dynamically updated using the three-factor reputation evaluation method in step S2. After the transaction ends, the smart contract is destroyed.
[0092] Every action Execution and time slot After the end, it will enter the next time slot and get a new observation . And, in the next few time slots, the edge nodes According to the task Whether it is completed successfully, discarded due to timeout, or failed, and gets a delay reward At the same time, the node reputation will be updated according to the three-factor reputation evaluation method based on beta distribution. Store to rolling buffer , used for the subsequent training of actor network and critic network.
[0093] Step S4-3: Train the neural network and fit the optimal task offloading strategy:
[0094] In order to train the neural network, the Proximal Policy Optimization (PPO) algorithm is introduced. After the size exceeds the training threshold, the buffer-based Mini-batch updates for the actor network and critics network conduct Subgradient update. To avoid large updates of the offloading policy, PPO uses a truncated proxy objective function ,as follows:
[0095]
[0096] in, E is the expected value; is the probability ratio of the new strategy to the old strategy; is the advantage function; is the truncation function.
[0097] The probability ratio of the new strategy to the old strategy The expression is:
[0098]
[0099] in, and is the strategy function, which means that Select Action probability.
[0100] Truncation function The expression is:
[0101]
[0102] in, is an adjustable cutoff factor.
[0103] To balance the bias of value estimation and the variance of returns to stabilize training, the generalized advantage estimate is used as the advantage function , defined as:
[0104]
[0105] in, is the trade-off between bias and variance; For the moment The value function of .
[0106] In order to optimize the value function, the following mean square error loss is introduced :
[0107]
[0108] in, Represents cumulative returns.
[0109] In order to increase the exploration ability of the strategy, the entropy regularization term is introduced :
[0110]
[0111] in, For strategy Entropy; Represents a given observation Uninstall strategy under .
[0112] In summary, the gradient update formulas for the actor network and the critic network are as follows:
[0113]
[0114]
[0115] in, and are the learning rates of the actor network and the critic network, respectively; is the weight coefficient of entropy regularization.
[0116] Step S4-4: After completing After the subgradient update, the rolling buffer It will be cleared to prevent old data from affecting the training stability. Repeat steps S4-2 and S4-3 to and critics network Perform training and use the trained actor network Obtain the optimal task offloading strategy based on the observations of edge nodes.
[0117] In summary, this paper proposes a method for collaborative task offloading in edge computing based on dynamic reputation. The goal is to ensure safe and effective collaborative task offloading in dynamic and potentially adversarial edge computing environments. This method not only minimizes the possibility of offloading tasks to malicious nodes, but also efficiently optimizes resource competition and utilization, maximizing task success rates.
[0118] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for offloading collaborative tasks in edge computing based on dynamic reputation value, characterized by: The following steps are involved: A task model for collaborative task offloading is established; the task model includes a task generation model, a task transmission model, a task calculation model, and a task delay model, and a task offloading decision optimization problem is formulated with the goal of maximizing the task success rate. Based on the constructed task model, the task offloading decision optimization problem is transformed into a decentralized partially observable Markov decision process. In this process, an observation space is constructed based on the observations of the edge node in each time slot; an action space is constructed based on all possible actions of the edge node in each time slot; and a reputation value is introduced to set the reward for the current edge node to offload the task to the target edge node. Add task difficulty factor to edge node reputation value , reputation value The method to obtain is as follows: ; in, and are the number of positive feedback and negative feedback of the jth edge node after the uth reputation value update; For the forgetting factor; is the historical task completion rate factor; is the action feedback. If the edge node processes the task correctly within the deadline, If the edge node performs malicious behavior or fails to process the task within the deadline, ; ; N is the number of edge nodes; If a task is discarded or fails to be processed, a negative reward is imposed on the current edge node; if the task is processed correctly, a positive reward is imposed on the current edge node, which includes the reputation value of the target edge node. The reputation value of the edge node is stored using blockchain technology. A neural network model for generating task offloading strategies is constructed, and experience fragments consisting of observations, actions, and rewards are constructed based on the interaction between edge nodes and the environment in a decentralized partially observable Markov decision process. The experience fragments at different times are stored in a rolling buffer, and the rolling buffer is used to train the neural network model. Based on the observations of edge nodes, the trained neural network model is used to obtain the optimal task offloading strategy.
2. The method for offloading edge computing collaborative tasks based on dynamic reputation value according to claim 1, characterized in that: The task difficulty factor The method to obtain is as follows: ; in, Indicates the task size; Indicates task processing density; Indicates the maximum task size; Indicates the maximum value of task processing density.
3. The method for offloading edge computing collaborative tasks based on dynamic reputation value according to claim 1, characterized in that: The malicious behavior mentioned includes denial of service, refusal to pledge, delayed processing, and incorrect processing.
4. The method for offloading edge computing collaborative tasks based on dynamic reputation value according to claim 1, characterized in that: The observations of the edge nodes include task size, task processing density, task arrival probability, computing power, transmission rate between edge nodes and other edge nodes, and historical reputation data of all edge nodes.
5. The method for offloading edge computing collaborative tasks based on dynamic reputation value according to claim 1, characterized in that: The interaction process between the edge node and the environment is as follows: When the task arrives at the current edge node, the current edge node selects an action based on the current observation and actor network; if the current edge node offloads the task to the target edge node for processing, the current edge node publishes a smart contract, and the target edge node pledges tokens; after the pledge is completed, the current edge node offloads the task to the target edge node for processing; if the target edge node refuses service or refuses to pledge, the task is processed locally by the current edge node; after the task processing is completed or the deadline is reached, the reputation value of the target edge node is dynamically updated.
6. The method for offloading edge computing collaborative tasks based on dynamic reputation value according to claim 1, characterized in that: In the described task transmission model, tasks of each edge node are transmitted in a first-in-first-out order. Tasks that have completed transmission are automatically placed in the processing queue of the target computing node and are no longer forwarded to other edge nodes. The transmission time of a transmission task between different edge nodes is expressed as the task size divided by the transmission rate between different edge nodes.
7. The method for offloading edge computing collaborative tasks based on dynamic reputation value according to claim 1, characterized in that: The task delay model constructs the total delay for task processing completion based on transmission delay, processing delay and waiting delay. , depending on the task processing method, the total delay The method to obtain is as follows: ; in, Indicates that the task is on the edge node Processing wait time in; Indicates that the task is from the edge node To the edge node Transmission waiting time; Indicates that the task is processed at the current edge node. Indicates that the task is offloaded to other edge nodes for processing; If the total delay of a task is greater than the deadline, the task will be discarded from the network and the deadline will be used as the total delay of the task.
8. The method for offloading edge computing collaborative tasks based on dynamic reputation value according to claim 1, characterized in that: The neural network model adopts a proximal policy optimization algorithm; during training, the parameters in the actor network are updated by a truncated proxy objective function and an entropy regularization term, and the parameters in the critic network are updated by a mean squared error loss of the value function.
9. A dynamic reputation-based edge computing collaborative task offloading system, comprising an IoT device, an edge node, and a smart contract; the IoT device generates tasks and transmits them to the edge node for processing; the edge node processes tasks from the IoT device and tasks offloaded from other edge nodes; characterized by: Used to execute the edge computing collaborative task offloading method based on dynamic reputation value as described in claim 1; the smart contract is responsible for the execution and supervision of the task offloading process and stores relevant information; the relevant information includes task description, task requirements, token staking rules, task result verification rules and reputation update rules.
Citation Information
Patent Citations
Sensing edge cloud blockchain network trusted offload cooperation node selection system and method
CN112202928A
Mobile edge computing task unloading method based on block chain
CN116669111A