Edge server computing power resource information diffusion method supporting multi-hop task unloading
By constructing a tree-topology IDM model and a deep reinforcement learning framework, a reinforcement learning agent is designed to diffuse information about edge server computing resources. This solves the problems of high information acquisition cost and latency in multi-hop task offloading, and optimizes system performance and stability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-21
- Publication Date
- 2026-04-10
AI Technical Summary
In edge computing, existing task offloading solutions struggle to efficiently offload multi-hop tasks in dynamic, large-scale IoT environments lacking a central controller, and the high cost of information acquisition leads to task execution delays and suboptimal resource allocation.
By adopting a limited computing power information diffusion mechanism, edge server computing power resource information is distributed through the construction of a tree-topology IDM model. Combined with a deep reinforcement learning framework and a progressive adjustment strategy, a reinforcement learning agent is designed to make task offloading decisions, thereby optimizing the overhead of computing power information diffusion and the latency of task offloading.
It achieves coordinated optimization of system computing power diffusion overhead and task unloading latency, improves the accuracy and efficiency of task unloading decisions, ensures the scalability and stability of large-scale networks, and reduces information diffusion costs and task unloading latency.
Smart Images

Figure CN121833069A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of edge server technology, and more particularly, to a method for disseminating information about limited computing resources that supports multi-hop task offloading for edge computing servers. Background Technology
[0002] Edge devices provide users with a channel to access the network and communicate with other server devices. The main purpose of edge servers is to store tasks as close as possible to the requesting client device, thereby reducing latency and shortening page load time. For the system architecture and real-time data processing architecture and applications of edge computing, refer to "Research on Edge Computing Resource Optimization Configuration Technology," Xi'an: Northwestern Polytechnical University Press, November 2020, author Shao Yanling, pp. 3-9. Edge computing is a distributed open platform that provides computing, storage, and networking services at the network edge, close to the data source. The computing power information of an edge server is far more than a simple performance number; it is a multi-dimensional comprehensive description of its processing capabilities, resource status, and applicable scenarios.
[0003] Edge computing is a new computing model that moves computing power and data storage from remote cloud data centers to locations closer to the data source, significantly reducing data transmission latency, greatly saving network costs, and protecting data security and privacy. Edge servers are key infrastructure components in edge computing systems, serving as the core computing power carrier located at the logical edge of the network. They are typically deployed physically near terminal devices or data sources, ranging in form from micro edge gateways to rack-mounted servers. Their core function is to provide near-field computing support for terminal devices within the local area. Internet of Things (IoT) devices are physical objects embedded with electronic components, software, sensors, actuators, and network connectivity. Their core function is to achieve data interaction and remote monitoring with external systems, devices, or cloud platforms through network connections. In edge computing, IoT devices have evolved from simple data collection terminals into edge nodes with certain local computing and intelligent decision-making capabilities, enabling them to collaborate with other edge devices or local edge servers to perform real-time data processing directly at the network edge, close to the data source.
[0004] Edge computing offloading is a key optimization technology in edge computing. It refers to the technical strategy of selectively migrating all or part of the computationally intensive tasks and related data of IoT devices to nearby edge servers for execution. By reasonably offloading tasks, the limitations of terminal devices in terms of computing power, storage and energy consumption can be made up for, thereby optimizing the overall system performance.
[0005] A key issue in edge computing offloading is the selection of target edge servers for IoT devices. The core objective is to achieve Pareto optimization of overall system performance by optimizing the execution location and resource allocation of tasks among IoT devices, edge servers, and the cloud, while ensuring application service quality. Specific objectives can be summarized as multi-dimensional trade-offs: first, minimizing task execution latency to ensure the responsiveness of real-time applications; second, optimizing energy efficiency to extend the battery life of restricted terminal devices; and third, balancing computational load to avoid overloading edge nodes and improve overall system throughput and reliability. This decision-making process typically requires dynamic consideration of multiple constraints, including task characteristics, network status, node computing power, and energy consumption costs, making it a typical multi-objective optimization problem.
[0006] Edge computing offloading target decision-making can be divided into two schemes: single-hop and multi-hop. Single-hop and multi-hop are two basic modes describing the topology of task migration paths. Single-hop offloading refers to IoT devices directly offloading computing tasks to edge servers. Compared with single-hop offloading, multi-hop task offloading allows IoT devices to access a wider range of edge server resources through inter-device collaborative relay, which may further reduce task execution costs and improve load balancing, and has been widely studied in recent years. However, most existing task offloading schemes are based on complete information design, that is, there is a central controller that collects real-time status information of all devices. In real-world application scenarios (such as outdoor wireless sensor networks), a central controller may be lacking. In addition, in dynamic, large-scale IoT environments, the overhead of acquiring complete information is huge, which can lead to performance degradation of solutions based on complete information. Therefore, an efficient distributed offloading strategy is needed, enabling user devices to actively collect computing power information and make autonomous offloading decisions. Summary of the Invention
[0007] To design an efficient distributed offloading strategy in edge servers, enabling user devices to proactively collect computing power information and autonomously make offloading decisions, while achieving efficient multi-hop task offloading with lower computing power information propagation costs and task offloading latency, this invention proposes a constrained computing power information diffusion mechanism, namely, an edge server computing power resource information diffusion method that supports multi-hop task offloading. The method of this invention: (S1) Constructing an IDM model based on a tree-like topology for server computing power information diffusion paths, and applying the IDM model to distribute edge server computing power resource information (i.e., ...) among IoT devices. (S2) Based on the IDM model, the computing power information of the servers along the diffusion path (i.e. The freshness of the information is assessed to obtain the age. (S3) Based on , spread jump number (i.e. (S4) Based on the IDM and ODM models, a task unloading decision-making model for task unloading and task processing is established. A JOM model is constructed to optimize the computational power information diffusion cost and task unloading latency. The JOM model is then applied to optimize the diffusion of computational power resource information on edge servers supporting multi-hop task unloading in scenarios with limited computational power information. This invention employs a deep reinforcement learning framework to determine the optimal computational power information diffusion distance for each edge server. It combines a progressive adjustment strategy and a near-end strategy optimization algorithm to design and train the agent, achieving the minimum computational power information diffusion cost and average task unloading latency.
[0008] In this invention, the goal of the JOM (Joint Operational Model) model (also known as the computing power information diffusion overhead and task unloading latency model) is to minimize the overall computing power information diffusion cost and task unloading latency of the network system by optimizing the diffusion distance of computing power information for each edge server. The JOM model is a model for IoT devices based on edge servers. Choose an edge server that minimizes average task unloading latency to unload tasks.
[0009] This invention trains an edge server computing power resource information diffusion method that supports multi-hop task offloading into a reinforcement learning agent. This reinforcement learning agent is embedded in the edge application rule and request management module. The agent is designed as an efficient distributed offloading strategy. On the one hand, it enables user devices to actively collect computing power information and make autonomous offloading decisions. On the other hand, it achieves efficient multi-hop task offloading with lower computing power information propagation costs and task offloading latency. Furthermore, it can obtain the optimal server computing power information diffusion distance. The agent is designed and trained by combining a progressive adjustment strategy and a near-end strategy optimization algorithm to achieve the minimum computing power information diffusion cost and average task offloading latency.
[0010] Technical effects of the present invention:
[0011] (1) The method of the present invention achieves the coordinated optimization of system computing power diffusion overhead and task unloading latency. By constructing computing power information diffusion overhead and task unloading latency as a joint optimization problem, and using this as the objective to determine the optimal information diffusion distance for each server, the present invention breaks through the bottleneck of traditional methods where it is difficult to balance the two, and achieves the optimal balance of overall performance at the system level, effectively reducing the overall overhead of the system.
[0012] (2) The method of the present invention improves the accuracy and efficiency of task unloading decision-making. By establishing a precise quantification function of server computing power information freshness, diffusion hops, and update frequency, IoT devices can make decisions based on more timely and accurate server computing power information, which effectively avoids the problem of suboptimal server selection caused by outdated information.
[0013] (3) The method of this invention ensures scalability and stability in large-scale networks. The constructed tree-based diffusion path structure is clear, avoiding broadcast storms and effectively controlling the diffusion cost of computing power information. Combined with an innovative deep reinforcement learning agent (embedded in the edge application rules and request management module) with a progressive adjustment strategy for diffusion distance, it can efficiently and stably solve the above-mentioned complex joint optimization problem. Even in large-scale IoT networks, it can adaptively find a near-optimal diffusion strategy, ensuring the robustness and practicality of the method of this invention. Attached Figure Description
[0014] Figure 1 This is a structural block diagram of a traditional edge reasoning system.
[0015] Figure 2 This is a flowchart of the edge server computing resource information diffusion method that supports multi-hop task offloading according to the present invention.
[0016] Figure 3 This is a block diagram showing the combination of the reinforcement learning agent designed in this invention with the edge application rules and request management module.
[0017] Figure 4 This is a flowchart of the training process for the reinforcement learning agent of this invention.
[0018] Figure 5 This is a schematic diagram of the edge computing offloading scenario of the present invention.
[0019] Figure 6 This is a line graph showing the change in computing power information diffusion overhead of various methods with the average processing rate of the server.
[0020] Figure 7 This is a line graph showing the change in task unloading latency for each method as a function of the server's average processing speed.
[0021] Figure 8 This is a line graph showing the change in total system overhead for each method as a function of the average server processing speed. Detailed Implementation
[0023] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. The examples of the parameters listed are merely preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
[0024] In this invention, any edge computing server is denoted as... A set of multiple edge computing servers is called a server set. ,and subscript The subscript indicates the identifier of the edge computing server. This indicates the total number of edge computing servers.
[0025] In this invention, any IoT device is denoted as... A collection of multiple IoT devices is called a device set. ,and subscript The identifier of the IoT device, subscript This indicates the total number of IoT devices. (IoT devices) With IoT devices Not the same device.
[0026] In this invention, any task is denoted as... A set of multiple tasks is called a task set. ,and subscript The identifier number of the task, subscript This indicates the total number of tasks.
[0027] In this invention, any task is defined. On the edge server and IoT devices The transmission delay between them is denoted as Define any task The transmission latency between IoT devices is denoted as . For ease of explanation, the task transmission latency will be standardized to [value missing]. subscript Represents a specific edge server or IoT device, and , It represents a specific Internet of Things (IoT) device.
[0028] In this invention, the system runtime is divided into K time slices of equal length, and any one of these time slices is denoted as... The last time slice is recorded as .
[0029] In this invention, the edge server computing power information diffusion path based on tree topology is denoted as... .
[0030] In this invention, the computing resource information on the edge server is denoted as... .
[0031] In this invention, the The diffusion hops from the edge server to the IoT device are denoted as .
[0032] In this invention, the freshness of edge server computing power information is denoted as... The aforementioned It also refers to the information age of edge server computing power information.
[0033] In this invention, the The diffusion rate is denoted as Each IoT device receives the aforementioned... The receiving rate is denoted as The rate at which tasks are offloaded from IoT devices to edge servers is denoted as... (Abbreviated as Task First Unloading Rate). The unloading rate of a task to the edge server is denoted as... (Referred to as the second unloading rate of the task). The unloading rate of the task from the IoT device is denoted as... (Referred to as the third unloading rate of the task). The unloading rate of the task from the IoT device to the cloud is denoted as... (Abbreviated as Task 4 Unloading Rate).
[0034] In this invention, the single-hop transmission delay of a task is denoted as... The multi-hop transmission delay of the task is denoted as... The time a task resides on the edge server is denoted as... The unloading latency of the task in the cloud is denoted as... .
[0035] In this invention, the average multi-hop transmission delay of the task is denoted as... The average time a task resides on the edge server is denoted as... The average unloading latency of the task in the cloud is denoted as... The average unloading latency of a task within a single time slice is denoted as... The average unloading latency of the task over all time slices is denoted as... .
[0036] In this invention, the The diffusion overhead on a single device is denoted as... The aforementioned The total diffusion overhead across all devices is denoted as .
[0037] like Figure 1The edge computing system architecture shown depicts an edge computing system where each edge computing node (i.e., edge host, edge server, an intermediate computing element located between the cloud and smart terminal devices) is inherently heterogeneous and deployed in various environments. Edge service providers offer edge computing infrastructure services, edge computing platforms, and edge computing software through heterogeneous edge micro data centers deployed near user devices. Edge computing infrastructure services (IaaS) utilize a shared cluster of collaborative edge computing nodes to provide authorized customers with edge resources such as network, storage, and computing, enabling localized intelligent business processing and thus improving the quality of service (QoS). Edge computing platforms (PaaS) provide edge service customers with programming languages, libraries, and development tools to create or execute applications on the cluster of collaborative edge computing nodes. Edge computing software (SaaS) is a service model that provides software services to customers via the Internet. Edge providers uniformly deploy application software on a collaborative edge computing cluster, and users access the applications deployed on the edge computing nodes through client interfaces or application programming interfaces. The edge computing platform manager includes modules for managing edge platform elements, edge application rules and requests, and edge application lifecycle management.
[0038] like Figure 1 , Figure 2 , Figure 3 As shown, this invention trains an edge server computing power resource information diffusion method that supports multi-hop task offloading into a reinforcement learning agent, which is embedded in the edge application rule and request management module. The agent enables user devices to proactively collect computing power information and autonomously make offloading decisions using an efficient distributed offloading strategy, achieving efficient multi-hop task offloading with lower computing power information propagation costs and task offloading latency. This invention addresses the problem of IoT devices selecting suitable target edge servers for task offloading using a limited computing power information diffusion mechanism. When the edge server's computing power information diffusion overhead is low but the task offloading latency is high, the server's computing power information diffusion distance is increased to achieve lower task offloading latency. Conversely, when the edge server's computing power information diffusion overhead is high but the task offloading latency is low, the edge server's computing power information diffusion distance is reduced to lower the edge server's computing power information diffusion overhead and obtain an acceptable task offloading latency. By jointly considering the edge server's computing power information diffusion overhead and task offloading latency, the low energy efficiency and low service quality problems of computing power information diffusion in edge scenarios are solved.
[0039] like Figure 2 As shown, the edge server computing resource information diffusion method supporting multi-hop task offloading of the present invention includes the following steps:
[0040] Step S1: Construct an IDM model based on a tree-structured topology for the diffusion of server computing power information, and apply the IDM model to distribute edge server computing power resource information among IoT devices. ;
[0041] The steps for constructing the IDM model are as follows:
[0042] Step 11: Construct the computing power information diffusion topology of the edge server;
[0043] Each edge computing server The computing power information is diffused outwards, and in edge computing offloading scenarios, all computing power information diffusion topology forms a tree centered on edge servers. Computing power information diffusion tree with root node This computing power information diffusion tree The shape is given based on a breadth-first search.
[0044] Step 12, Representation of the computing power information diffusion behavior of each node in the computing power information diffusion tree;
[0045] Computing power information diffusion tree Each parent node in the process propagates the edge server's computing resource information to its child nodes. .
[0046] edge server computing resource information With diffusion rate The process of disseminating computing power information from a parent node to its child nodes will be carried out. Randomly announce to a child node.
[0047] Each IoT device Receive from edge server The computing power resource information is denoted as ,and .
[0048] Step S2: Evaluate the freshness of server computing power information along the diffusion path based on the IDM model to obtain the information age. ;
[0049] The maximum timeliness limit for edge server computing resource information is denoted as: .
[0050] Computing IoT devices Own edge servers The expected value of the information age of computing power resource information is denoted as... ,and .when Greater than At that time, IoT devices Edge servers will not be used. The computing power information is used to prevent tasks from being offloaded to the edge server. The above refers to the information age of the computing power resources of edge servers. Exceeded The computing power resource information belonging to this edge server will no longer be trusted, and IoT devices will no longer use this computing power resource information.
[0051] By computing IoT devices The expected value of the information age of the computing power resources of all edge servers is obtained to acquire information belonging to IoT devices. The set of available edge servers, denoted as ,and .
[0052] Step S3, based on , diffusion jump number Based on the unloading rate, an ODM model for task unloading decisions in the task processing process is established.
[0053] The steps for constructing the ODM model in this invention are as follows:
[0054] Step 21: Calculate the data transmission latency of the task;
[0055] The multi-hop transmission delay of the computation task is denoted as . .
[0056] Computing tasks from IoT devices Offload to edge server The multi-hop transmission delay is denoted as ,and , For any node on the diffusion path, For the diffusion path located at The previous node.
[0057] Step 22, task unloading and distribution;
[0058] In time slice Internally, from IoT devices To edge server The task unloading process follows the task's first unloading rate. The Poisson process.
[0059] Calculation in time slice Offload to edge server The task arrival rate, i.e., the second task unloading rate. ,and .
[0060] Step 23, Edge server task processing;
[0061] Edge server The processing procedure can be modeled as The time a task resides on the edge server during the queuing process is denoted as . .
[0062] In time slice Internally, the task is on the edge server. The length of stay in China is ,and ,in For edge servers Task processing speed, This is the second unloading rate for the task.
[0063] Step 24, uninstall from the cloud;
[0064] The load distribution follows the following method ,in Indicates Internet of Things (IoT) devices The fourth uninstallation rate for tasks uninstalled to the cloud. Indicates the task originates from the IoT device. The third unloading rate of the task.
[0065] The total latency for a single task to be offloaded to the cloud is Obviously, when the edge server... The length of stay exceeds In such cases, offloading tasks to the cloud becomes a better option. Therefore, the load on each edge server must meet the following constraints. This condition is equivalent to .
[0066] Step 25, average task unloading latency;
[0067] In time slice Within, the average dwell time of the mission in accordance with calculate.
[0068] In time slice Within, the average transmission latency of the task in accordance with calculate.
[0069] The average offload latency of a task is the sum of the average dwell latency and the average transmission latency. (In the time slice) Internally, cloud latency is recorded as The time slice is obtained by weighted summing the average unloading latency of tasks offloaded to the edge server and the average unloading latency of tasks offloaded to the cloud. The average unloading latency for all tasks within the system is ,and Therefore, the overall average latency of the task unloading process is... ,and .
[0070] Step S4: Based on the IDM model and the ODM model, construct a JOM model to optimize the overhead of computing power information diffusion and the latency of task unloading, and apply the JOM model to optimize the diffusion of edge server computing power resource information that supports multi-hop task unloading in scenarios with limited computing power information.
[0071] See Figure 3 As shown, the reinforcement learning agent designed in this invention uses a deep reinforcement learning framework to determine the maximum computing power information diffusion distance for each edge server, and combines a progressive adjustment strategy and a proximal strategy to optimize the diffusion of edge server computing power resource information that supports multi-hop task offloading in scenarios with limited computing power information. The Agent of this invention is an artificial entity capable of perceiving its environment, making decisions, and taking actions; it is a tool with the following characteristics.
[0072] Autonomy: An agent can perform tasks independently without human intervention or input.
[0073] Perception: Agents can perceive and interpret their environment through various sensors.
[0074] Response: An agent can assess the environment and respond accordingly to achieve its goals.
[0075] Reasoning and Decision-Making: Agents are intelligent tools that analyze data and make decisions to achieve goals. They use reasoning techniques and algorithms to process information and take appropriate actions.
[0076] Learning: Agents learn and improve their performance through machine, deep and reinforcement learning elements and techniques.
[0077] Communication: Agents can communicate with other intelligent agents or humans in different ways, such as understanding and responding to natural language, recognizing speech, and exchanging messages via text.
[0078] Goal-oriented: The agent aims to achieve specific goals, which can be predefined or learned through interaction with the environment.
[0079] In this invention, the reinforcement learning agent includes: an edge server computing power information diffusion path model (IDM model), an IoT device offloading optimization problem model (ODM model), and a computing power information diffusion overhead and task offloading latency model (JOM model).
[0080] (i) Constructing an edge server computing power information diffusion path model (IDM model);
[0081] In this invention, the constructed edge server computing power information diffusion path model is referred to as the Information Dissemination Model (IDM). The IDM model has a tree-like topology and is used to diffuse the computing power resource information of edge servers. Steps 11 and 12 of this invention for constructing the IDM model, and step 13 for evaluating the freshness of server computing power information on the diffusion path based on the IDM model, to obtain the information age.
[0082] Step 11: Construct the computing power information diffusion topology of the edge server;
[0083] In this invention, each edge computing server disseminates its own computing power information outward, in edge computing offloading scenarios (such as...). Figure 4 The topology of all computing power information diffusion in the diagram forms a computing power information diffusion tree with the edge server as the root node. The shape of this computing power information diffusion tree is given by breadth-first search.
[0084] edge servers The root node is built as a class The computing power information diffusion tree is ;
[0085] Initialize computing power information diffusion tree ;
[0086] Use Breadth-first search is used to construct the root of the root. diffusion tree The diffusion tree For diffusion Breadth-first search is used primarily to construct a diffusion tree. It is the information passed between nodes in the tree.
[0087] Record Chinese belongs to Internet of Things devices The set of sibling nodes, denoted as ;
[0088] Record Chinese belongs to Internet of Things devices The parent node is denoted as ;
[0089] Record Mid-range IoT devices Edge servers The information diffusion hop count, i.e., the information diffusion distance of computing power, is denoted as... .
[0090] Step 12, Representation of the computing power information diffusion behavior of each node in the computing power information diffusion tree;
[0091] In this invention, computing power information diffusion tree Each parent node in the process propagates the edge server's computing resource information to its child nodes. .
[0092] In this invention, the computing resource information of the edge server With diffusion rate The process of disseminating computing power information from a parent node to its child nodes will be carried out. The notification is randomly sent to a child node. Therefore, in In this system, each child node receives computing power information from its parent node. For example, a computing power information diffusion tree. The diffusion rate of computing power resource information is .
[0093] In this invention, each IoT device receives data at a certain rate. Receive computing resource information For example, IoT devices Receive from edge server The computing power resource information is denoted as ,and .
[0094] Step 13, Timeliness Analysis of Edge Server Computing Power Information;
[0095] Information age ( The definition of information freshness is the time elapsed since the latest received information was generated at the source; it is widely used to quantify information freshness. In this invention, it is primarily used... Quantitative analysis of the freshness of edge server computing power information.
[0096] The maximum timeliness limit for edge server computing resource information is denoted as: .
[0097] In this invention, when the information age of the edge server... Exceed At that time, the computing power resource information belonging to that edge server will no longer be trusted, and IoT devices will no longer use that computing power resource information.
[0098] In this invention, a diffusion tree is recorded. Mid-edge server To IoT devices The set of computing power information dissemination paths, denoted as ,and , For any node on the diffusion path, the subscript The node identifier on the diffusion path, This refers to the distance at which computing power information spreads. In the aforementioned... The starting point is the edge server. The endpoint is IoT devices. The preceding node on the path is the parent node of the following node.
[0099] In this invention, computing IoT devices Own edge servers The expected value of the information age of computing power resource information is denoted as... ,and .when Greater than At that time, IoT devices Edge servers will not be used. The computing power information is used to prevent tasks from being offloaded to the edge server. superior.
[0100] In this invention, IoT devices are computed. Owned with all edge servers ( The expected value of the information age of computing power resources is obtained to acquire information belonging to Internet of Things devices. The set of available edge servers, denoted as ,and .
[0101] (ii) Constructing an optimization model for IoT device offloading (ODM model);
[0102] In this invention, the constructed IoT device offloading optimization problem model is referred to as OffloadingDecision Model (ODM).
[0103] IoT devices need to select appropriate target edge servers for offloading. This invention models the task offloading process of IoT devices (i.e., an ODM model), which analyzes the impact of edge server computing power information diffusion distance on the average task offloading latency. The steps for constructing the ODM model in this invention are as follows:
[0104] Step 21: Calculate the data transmission latency of the task;
[0105] The multi-hop transmission delay of the computation task is denoted as . .
[0106] Computing tasks from IoT devices Offload to edge server The multi-hop transmission delay is denoted as ,and , For any node on the diffusion path, For the diffusion path located at The previous node.
[0107] Step 22, task unloading and distribution;
[0108] IoT devices will split the Poisson arrival flow of a task into multiple sub-Poisson task flows that are offloaded to different edge servers.
[0109] In time slice Internally, from IoT devices To edge server The task unloading process follows a rate of The Poisson process of (task first unloading rate).
[0110] In this invention, calculation is performed in time slices. Offload to edge server The task arrival rate, i.e., the second task unloading rate. ,and .
[0111] Step 23, Edge server task processing;
[0112] Assume that the processing time of each task follows an exponential distribution. Edge server. The processing procedure can be modeled as The queuing process, where the task resides on the edge server (including queue waiting time and processing time), is denoted as... .
[0113] In this invention, during the time slice Internally, the task is on the edge server. The length of stay in China is ,and ,in For edge servers Task processing speed, This is the second unloading rate for the task.
[0114] Step 24, uninstall from the cloud;
[0115] When the edge server is overloaded, task traffic can be further offloaded to the cloud. In this invention, load balancing follows the following method: ,in Indicates Internet of Things (IoT) devices The rate at which tasks are unloaded to the cloud (i.e., the fourth unload rate of tasks). Indicates the task originates from the IoT device. The unloading rate on the task (i.e., the third unloading rate of the task).
[0116] In this invention, the total latency for a single task to be offloaded to the cloud is: Obviously, when the edge server... The length of stay exceeds In such cases, offloading tasks to the cloud becomes a better option. Therefore, the load on each edge server must meet the following constraints. This condition is equivalent to .
[0117] Step 25, average task unloading latency;
[0118] In time slice Within, the average dwell time of the mission in accordance with calculate.
[0119] In time slice Within, the average transmission latency of the task in accordance with calculate.
[0120] In this invention, the average offloading latency of a task is the sum of the average dwell latency and the average transmission latency. Furthermore, given the relatively stable operating state of the cloud, without loss of generality, within the time slice... Internally, cloud latency is recorded as The time slice is obtained by weighted summing the average unloading latency of tasks offloaded to the edge server and the average unloading latency of tasks offloaded to the cloud. The average unloading latency for all tasks within the system is ,and Therefore, the overall average latency of the task unloading process is... ,and .
[0121] In this invention, the completion process for each task offloaded to the edge server mainly includes three stages: (1) input data transmission stage; (2) queuing and waiting for processing stage; and (3) task calculation stage. In the input data transmission stage, the required task data is relayed from the source IoT device to the target edge server through a multi-hop transmission path. Newly arrived tasks must wait for all previously arrived tasks to be completed before they can begin processing.
[0122] (III) Optimize the construction of the computing power information diffusion overhead and task unloading latency model (JOM model).
[0123] In this invention, the constructed optimized computing power information diffusion overhead and task unloading latency model is abbreviated as the Joint Optimization Model (JOM). The JOM model considers the diffusion range of edge server computing power information as a joint optimization of computing power information diffusion overhead and task unloading latency. The steps for constructing the JOM model in this invention are as follows:
[0124] Step 31: Calculate the overhead of computing power information diffusion;
[0125] Edge servers need to promptly disseminate computing power information to maintain a certain level of information freshness (i.e., information age) on the IoT device side. However, this will incur additional overhead on the corresponding computing power information forwarding equipment, such as energy consumption or spectrum occupation. This represents the cost of single-hop computing power information diffusion. (In the computing power information diffusion tree) In the process, the diffusion rate of computing power information of each non-leaf node follows the parameter: The Poisson distribution is followed. Therefore, the expected computational power diffusion cost rate of each non-leaf node can be calculated as follows: .
[0126] In this invention, computing power information diffusion tree The set of all non-leaf nodes is In edge computing offloading scenarios (such as...) Figure 4 The total expected computing power information diffusion cost rate (as shown) is the sum of the expected information diffusion cost rates of all non-leaf nodes in each computing power information diffusion tree, denoted as . ,and .
[0127] Step 32: Jointly optimize the overhead of computing power information diffusion and the latency of task unloading;
[0128] The present invention aims to determine the optimal computing power information diffusion distance for each edge server and the optimal task offloading target server for each IoT device, so as to minimize the combined overhead of computing power information diffusion and task offloading latency of the system.
[0129] In this invention, the sequence formed by the furthest diffusion distance of all computing power information diffusion trees is: ,and ,in Information diffusion tree for computing power The furthest diffusion distance, i.e. the depth of the computing power information diffusion tree.
[0130] In this invention, the sequence of task flow offloading decisions for all IoT devices is as follows: ,and ,in Indicates time slice Internally, from IoT devices To edge server The task unloading process follows a certain rate.
[0131] In this invention, the joint optimization decision of computing power information diffusion overhead and task unloading latency is given by formula (1).
[0132] .
[0133] These are constraints.
[0134] It is to minimize.
[0135] It is a coefficient for the overhead of computing power information diffusion, and its value range is... .
[0136] It is a coefficient for task unloading latency, and its value range is... .
[0137] Step 33: Set the decision for the target edge server to be unloaded by the task;
[0138] For those with fixed parameters Given the diffusion distance of computing power information, the The task unloading decision subproblem is a convex optimization problem, which can be effectively solved using the interior-point method. This represents the optimal solution to the subproblem, i.e., given the parameters. The minimum average task unloading latency is given. Therefore, the problem in formula (1) can be transformed into the problem in formula (2).
[0139] .
[0140] In the problem described in Equation (2), the state space is highly complex, consisting of multiple dimensions such as network topology, edge server computing power, and task arrival rate. Furthermore, the solution space grows exponentially with network size, leading to combinatorial explosion. These characteristics make traditional heuristic search algorithms converge slowly.
[0141] (iv) Construction and training of reinforcement learning agents
[0142] See Figure 4As shown, in order to implement the use of agents to make decisions on the optimal distribution of computing power information for edge servers in the edge application rules and request management module, the constructed agents need to be trained. During the training process, the training environment is provided by the IDM model and the ODM model. Then, the agent can perceive the environmental state and make edge server computing power information distribution decisions based on the JOM model.
[0143] The Agent employs deep reinforcement learning to solve formula (2), and its core idea lies in using the Agent to continuously adjust the diffusion distance of computing power information from the edge server. Specifically, it is divided into interactive environment design, agent construction, agent training, and optimal computing power information diffusion distance decision acquisition.
[0144] Step 41, Interactive Environment Design;
[0145] Let the interaction iteration step size be denoted as The number of interaction rounds is denoted as Let the state space be denoted as The action space is recorded as .
[0146] The interactive environment consists of three parts: state, action, and reward. The state reflects the current characteristics of the environment. State space. It consists of states containing multidimensional network environment information, represented as , This refers to the state label in the state space.
[0147] The agent's action involves determining the diffusion distance parameters of computing power information for each edge server. However, in real-world applications, when the network topology is large, the number of edge servers is numerous, and the combinations of the diffusion tree depth parameters for computing power information for different edge servers are diverse, leading to an excessively large action space and hindering effective learning by the agent. To reduce the size of the action space, this invention restricts the agent to adjusting the diffusion distance of computing power information for only a single edge server at each step, rather than all edge servers. Therefore, the action space is designed as follows: .
[0148] any action This represents an operation on the depth of computing power diffusion within a specific computing power information diffusion tree. For example, Indicates the step size Adjusting the computing power information diffusion tree The parameters, and Updated to Through the action space The action space size of each interaction iteration step can be limited to 1. Furthermore, by applying action masks, invalid actions are eliminated in each interaction iteration step, thereby further reducing the action space and improving sampling efficiency. Representing state The total system overhead value is calculated below. The objective of this invention is to gradually reduce the system overhead through continuous interaction between the agent and the environment. The numerical value is obtained to obtain the optimal solution for formula (2). To incentivize the agent to minimize... The goal will be rewarded instantly. The design is the difference in total system overhead before and after the action, and .
[0149] This represents a weighted topology graph of the network, where IoT device nodes have a task arrival rate attribute, and edge server nodes have a task processing rate, computing power information diffusion rate, and computing power information diffusion tree overhead attribute.
[0150] This represents the agent's current decision-making scheme regarding the diffusion distance of edge server computing power information.
[0151] It represents global state information, mainly including control plane and data plane overhead, cloud offloading latency, and effective information age threshold.
[0152] Information diffusion tree for computing power The furthest diffusion distance, i.e. the depth of the computing power information diffusion tree.
[0153] Indicates the step size of the interaction iteration Any environmental state.
[0154] This represents a positive weighting factor. Used to adjust learning rate and stability.
[0155] Representing state The total system overhead value.
[0156] Representing state The total system overhead value.
[0157] Step 42, Agent Construction;
[0158] An agent is constructed and trained using the Proximal Policy Optimization (PPO) algorithm. This algorithm is based on the Actor-Critic framework, where the Actor network interacts with the environment to learn the policy, and the Critic network evaluates the state value to guide the Actor network's updates. The Proximal Policy Optimization algorithm introduces a pruning mechanism to limit the policy update magnitude, thereby avoiding performance degradation and maintaining high sample efficiency, thus improving training stability.
[0159] In this invention, the Actor-Critic framework is referenced from Konda, V., & Tsitsiklis, J. (1999). Actor-critic algorithms. Advances in neural information processing systems, 12.
[0160] Step 43, Agent Training;
[0161] Reinforcement learning generates historical experience through continuous interaction between the agent and the environment, and uses this historical experience to train the agent.
[0162] (A) Initialize the agent and training parameters. After initialization, start the interaction process between the agent and the environment, and update the parameters of the agent during the interaction process.
[0163] (B) The entire interaction process can be described using two loops, with the outer loop iterating through the interaction rounds between the agent and the environment. The inner loop iterates through each interaction round. Interaction iteration step size between intelligent agents and the environment Each interaction round Step (C) needs to be performed.
[0164] Initialize the interaction round variable to 0, and in each interaction round... After the interaction round ends, the interaction round number variable is incremented by 1. The training process of the agent is completed when the interaction round number variable reaches the preset maximum number of interaction rounds.
[0165] (C) In each interaction round First, initialize the interactive iteration step size. , The step count variable is 0, each After the interaction ends, increment the interaction iteration step count variable by 1. The current interaction round is completed when the interaction iteration step count variable reaches the preset maximum interaction iteration step count. Step size for each interaction iteration Execution step (D);
[0166] (D) In the interactive iteration step size In the process, the intelligent system first obtains the current interaction iteration step size based on environmental state sampling. Actions to be performed And perform the action. After execution, the environment will calculate based on both the IDM and ODM models. and And obtain according to the JOM model After obtaining the total system overhead Afterwards, the environment will be based on Calculate the single-step reward value of the action performed by the intelligent agent. And store single-step experience To the experience buffer, where For the Actor in state The probability distribution of each action. For Critic's state The estimated value function value. After each step of experience is stored in the experience buffer, it is determined whether the experience buffer is full. If the experience buffer is not full, the process returns to step (C) to continue iterating; if the experience buffer is full, step (E) is executed.
[0167] (E) The agent training algorithm updates the agent's parameters based on the PPO algorithm using historical experience data in the experience buffer. After the agent parameters are updated, the experience buffer is cleared, and then the process returns to step (C) to continue iterating.
[0168] Step 44: Obtain the optimal computing power information diffusion distance;
[0169] After the agent is trained, during its interaction with the environment, it selects the action with the highest probability in each interaction iteration step to gradually obtain the optimal solution of formula (2). On the other hand, when the maximum iteration step size is reached, it selects the last iteration step size. As the optimal solution of formula (2). Example
[0170] This invention utilizes computer hardware, and the software is developed and implemented using the Python language, and simulated in the Python environment (version 3.12).
[0171] In the example, it was based on the edge computing scenario. Figure 5This invention presents a method for disseminating server computing power information and offloading computing tasks. In the diagram, Cloud Servers represent cloud servers, MEC Servers represent edge servers, IoT devices represent Internet of Things (IoT) devices, and Offloading probability represents the offloading probability. The colored arrows in the lower left corner indicate the computing power information dissemination path for each edge server, and the black dashed arrows represent the task offloading path for IoT devices. Each edge server disseminates its own computing power information outward according to a tree structure. After receiving the server's computing power information, the IoT device determines whether the information has timed out and considers offloading the task to an edge server or cloud server where the information has not timed out. The computing power information dissemination distance of the edge servers is set using the reinforcement learning algorithm proposed in this invention. The offloading target decision for the IoT device is solved by a convex optimization problem that minimizes task offloading latency, as proposed in this invention.
[0172] Based on Figure 5 In the scenario shown, the number of edge server nodes in the edge network is set to 6, and the number of IoT devices is 60. During performance simulation, the average processing rate of the edge servers is set to different values, and the changes in computing power information diffusion overhead and task unloading latency are observed. A comparison of the effects of the method of this invention (Ours) with existing methods (SHID method, FAID method, Greedy method) is shown below. Figure 6 , Figure 7 and Figure 8 As shown.
[0173] Three comparison schemes were set up during the simulation. The first was the single-hop information diffusion scheme, namely the SHID method, such as the method proposed in the paper (T. Gan, S. Zhang, W. Xu, X. Li, Z. Wang, and H. Luo, “Backoff RevealsPriority: A Distributed Status-Aware Response Method for Device-Assisted TaskOffloading,” in 2025 IEEE International Conference on Web Services (ICWS), Helsinki, Finland, Jul. 2025, pp. 1–7.). In this method, the edge server only diffuses computing power information to a single hop range.
[0174] The FAID method is the furthest information diffusion scheme. In this method, each edge server diffuses its computing power information to the farthest possible range, thereby enabling IoT devices to obtain the maximum amount of edge server computing power information.
[0175] The Greedy method is a greedy information diffusion scheme. In this method, each IoT device only offloads its tasks to the edge server closest to it, and the information diffusion range of each edge server precisely covers the IoT device that it is the offloading target.
[0176] Figure 6 The graph shows the variation in average processing rate of computing power information diffusion overhead for each method. The horizontal axis, AverageServer Processing Rate, represents the average task processing rate of all edge servers, and the vertical axis, Dissemination Cost, represents the total computing power information diffusion overhead. The results show that the method of the present invention (Ours) achieves task offloading latency comparable to the full information diffusion scheme, while reducing computing power information diffusion overhead by 66.88% to 95.14% compared to the full information diffusion scheme.
[0177] Figure 7 Line graphs showing the change in task offloading latency as a function of the average server processing rate for each method are presented. The horizontal axis, AverageServer Processing Rate, represents the average task processing rate of all edge servers, and the vertical axis, TaskOffloading Delay, represents the average offloading latency of all tasks. The results show that the method of this invention (Ours) achieves a task offloading latency comparable to the complete information diffusion scheme. Furthermore, compared to the single-hop information diffusion scheme and the greedy information diffusion scheme, this scheme can reduce task offloading latency by up to 97.14% to 86.77%, respectively.
[0178] Figure 8 The graph shows the total system cost as a function of the average server processing rate. The horizontal axis, AverageServer Processing Rate, represents the average task processing rate of all edge servers, and the vertical axis, Overall Cost, represents the total system cost. Compared with other existing technologies, the method of this invention can reduce the total system cost by up to 66.17% to 82.37%, indicating that the proposed solution achieves the best balance between computing power information diffusion overhead and task offloading latency.
[0179] This invention discloses a constrained computing power information diffusion scheme for multi-hop task unloading. The scheme aims to reduce the overhead of computing power information diffusion and the latency of task unloading. Specifically, this invention first studies the impact of a tree-based topology-based computing power information diffusion mechanism on the freshness of user-side computing power information; then, it analyzes the mechanism by which the computing power information diffusion distance affects the latency of task unloading using queuing theory. Based on this, this invention constructs the optimization problem of computing power information diffusion distance into an optimization model with the objective of jointly minimizing the overhead of computing power information diffusion and the latency of task unloading. To solve this model, a deep reinforcement learning algorithm based on near-end policy optimization is proposed. This algorithm integrates an innovative progressive adjustment strategy for diffusion distance to dynamically optimize the computing power diffusion distance.
Claims
1. A method for disseminating edge server computing resource information supporting multi-hop task offloading, wherein the method for disseminating edge server computing resource information supporting multi-hop task offloading is trained into a reinforcement learning agent, which is embedded in the edge application rule and request management module; characterized in that: The method for disseminating edge server computing resource information that supports multi-hop task offloading includes the following steps; Step S1: Construct an IDM model based on a tree-structured topology for the diffusion of server computing power information, and apply the IDM model to distribute edge server computing power resource information among IoT devices. ; The steps for constructing the IDM model are as follows: Step 11: Construct the computing power information diffusion topology of the edge server; Each edge computing server The computing power information is diffused outwards, and in edge computing offloading scenarios, all computing power information diffusion topology forms a tree centered on edge servers. Computing power information diffusion tree with root node ; This computing power information diffusion tree The shape is given based on a breadth-first search. Step 12, Representation of the computing power information diffusion behavior of each node in the computing power information diffusion tree; Computing power information diffusion tree Each parent node in the process propagates the edge server's computing resource information to its child nodes. ; edge server computing resource information With diffusion rate The information is diffused; whenever a parent node diffuses computing power information to a child node, it will... Randomly notify a child node; Each IoT device Receive from edge server The computing power resource information is denoted as ,and ; Step S2: Evaluate the freshness of server computing power information along the diffusion path based on the IDM model to obtain the information age. ; The maximum timeliness limit for edge server computing resource information is denoted as: ; Computing IoT devices Own edge servers The expected value of the information age of computing power resource information is denoted as... ,and ;when Greater than At that time, IoT devices Edge servers will not be used. The computing power information is used to prevent tasks from being offloaded to the edge server. The above refers to the information age of the computing power resources of edge servers. Exceeded The computing power resource information belonging to this edge server will no longer be trusted, and IoT devices will no longer use this computing power resource information; By computing IoT devices The expected value of the information age of the computing power resources of all edge servers is obtained to acquire information belonging to IoT devices. The set of available edge servers, denoted as ,and ; Step S3, based on , diffusion jump number Based on the unloading rate, an ODM model for task unloading and task processing decision-making is established; the steps for constructing the ODM model are: Step 21: Calculate the data transmission latency of the task; The multi-hop transmission delay of the computation task is denoted as . ; Computing tasks from IoT devices Offload to edge server The multi-hop transmission delay is denoted as ,and , For any node on the diffusion path, For the diffusion path located at A previous node; Step 22, task unloading and distribution; In time slice Internally, from IoT devices To edge server The task unloading process follows the task's first unloading rate. The Poisson process; Calculation in time slice Offload to edge server The task arrival rate, i.e., the second task unloading rate. ,and ; Step 23, Edge server task processing; Edge server The processing procedure can be modeled as The time a task resides on the edge server during the queuing process is denoted as . ; In time slice Internally, the task is on the edge server. The length of stay in China is ,and ,in For edge servers Task processing speed, The second unloading rate for the task; Step 24, uninstall from the cloud; The load distribution follows the following method ,in Indicates Internet of Things (IoT) devices The fourth uninstallation rate for tasks uninstalled to the cloud. Indicates the task originates from the IoT device. The third unloading rate of the task on the platform; The total latency for a single task to be offloaded to the cloud is Obviously, when the edge server The length of stay exceeds In such cases, offloading tasks to the cloud becomes the preferred option; therefore, the load on each edge server must meet the following constraints. This condition is equivalent to ; Step 25, average task unloading latency; In time slice Within, the average dwell time of the mission in accordance with calculate; In time slice Within, the average transmission latency of the task in accordance with calculate; The average offload latency of a task is the sum of the average dwell latency and the average transmission latency; in the time slice Internally, cloud latency is recorded as The time slice is obtained by weighted summing the average offload latency of tasks offloaded to the edge server and the average offload latency of tasks offloaded to the cloud. The average unloading latency for all tasks within the system is ,and ; Therefore, the overall average latency of the task unloading process is ,and ; Step S4: Based on the IDM model and the ODM model, construct a JOM model to optimize the overhead of computing power information diffusion and the latency of task unloading, and apply the JOM model to optimize the diffusion of edge server computing power resource information that supports multi-hop task unloading in scenarios with limited computing power information.
2. The edge server computing resource information diffusion method supporting multi-hop task offloading according to claim 1, characterized in that: The JOM model considers the spread range of edge server computing power information as a joint optimization of computing power information spread overhead and task unloading latency; Step 31: Calculate the overhead of computing power information diffusion; Computing power information diffusion tree The set of all non-leaf nodes is denoted as ; In edge computing offloading scenarios, the total expected computing power information diffusion overhead rate is the sum of the expected information diffusion overhead rates of all non-leaf nodes in each computing power information diffusion tree, denoted as . ,and ; Step 32: Jointly optimize the overhead of computing power information diffusion and the latency of task unloading; The sequence formed by the furthest diffusion distance of all computing power information diffusion trees is as follows: ,and ,in Information diffusion tree for computing power The furthest diffusion distance, i.e. the depth of the computing power information diffusion tree; The sequence of task offloading decisions for all IoT devices is as follows: ,and ,in Indicates time slice Internally, from IoT devices To edge server The rate at which the task unloading process follows; The joint optimization decision of computing power information diffusion overhead and task unloading latency is given by formula (1); ; These are constraints; It is to minimize; It is a coefficient for the overhead of computing power information diffusion, and its value range is... ; It is a coefficient for task unloading latency, and its value range is... ; Step 33: Set the decision for the target edge server to be unloaded by the task; The The task unloading decision subproblem is a convex optimization problem, which can be effectively solved using the interior-point method; using This represents the optimal solution to the subproblem, i.e., given the parameters. The minimum average task unloading latency under the given conditions; then the problem in formula (1) can be transformed into the problem in formula (2); 。 3. The edge server computing resource information diffusion method supporting multi-hop task offloading according to claim 1 or 2, characterized in that: During the training process of the agent, the training environment is provided by the IDM model and the ODM model. Then the agent perceives the environmental state and makes the edge server computing power information diffusion decision of the JOM model. Step 41, Interactive Environment Design; Let the interaction iteration step size be denoted as The number of interaction rounds is denoted as Let the state space be denoted as The action space is recorded as ; The interactive environment consists of three parts: state, action, and reward; the state reflects the current characteristics of the environment; the state space... It consists of states containing multidimensional network environment information, represented as , State labels in the state space; The agent's actions are used to determine the propagation distance parameters of computing power information from each edge server; therefore, the action space is designed as follows: ; any action This represents an operation on the computing power diffusion depth of a certain computing power information diffusion tree; Indicates the step size Adjusting the computing power information diffusion tree The parameters, and Updated to Through action space The action space size of each interaction iteration step can be limited to 1. ;use Representing state The total system overhead value is gradually reduced. The numerical value is obtained to obtain the optimal solution for formula (2). To incentivize the agent to minimize The goal will be rewarded instantly. The design is the difference in total system overhead before and after the action, and ; A weighted topology diagram representing a network; Represents global status information; Information diffusion tree for computing power The furthest diffusion distance; Indicates the step size of the interaction iteration Any environmental state; Indicates a positive weighting factor; Representing state The total system overhead value below; Representing state The total system overhead value below; Step 42, Agent Construction; An agent is constructed and trained using a proximal policy optimization algorithm; Step 43, Agent Training; Reinforcement learning generates historical experience through continuous interaction between the agent and the environment, and uses this historical experience to train the agent. (A) Initialize the agent and training parameters. After initialization, start the interaction process between the agent and the environment, and update the parameters of the agent during the interaction process. (B) The entire interaction process can be described using two loops, with the outer loop iterating through the interaction rounds between the agent and the environment. The inner loop iterates through each interaction round. Interaction iteration step size between intelligent agents and the environment Each interaction round Step (C) needs to be performed. Initialize the interaction round variable to 0, and in each interaction round... After the interaction round ends, the interaction round number variable is incremented by 1. The training process of the agent is completed when the interaction round number variable reaches the preset maximum number of interaction rounds. (C) In each interaction round First, initialize the interactive iteration step size. , The step count variable is 0, each After the interaction ends, increment the interaction iteration step count variable by 1. The current interaction round is completed when the interaction iteration step count variable reaches the preset maximum interaction iteration step count. Step size for each interaction iteration Execution step (D); (D) In the interactive iteration step size In the process, the intelligent system first obtains the current interaction iteration step size based on environmental state sampling. Actions to be performed And perform the action; the action After execution, the environment will calculate based on both the IDM and ODM models. and And obtain according to the JOM model After obtaining the total system overhead Afterwards, the environment will be based on Calculate the single-step reward value of the action performed by the intelligent agent. And store single-step experience To the experience buffer, where For the Actor in state The probability distribution of each action. For Critic's state The estimated value function value; after each single-step experience is stored in the experience buffer, it is determined whether the experience buffer is full. If the experience buffer is not full, the process returns to step (C) to continue iterating; if the experience buffer is full, step (E) is executed. (E) The agent training algorithm updates the agent's parameters based on the PPO algorithm using historical experience data in the experience buffer; after the agent parameters are updated, the experience buffer will be cleared, and then the process will return to step (C) to continue iterating. Step 44: Obtain the optimal computing power information diffusion distance; After the agent is trained, during its interaction with the environment, the agent selects the action with the highest probability in each interaction iteration step to gradually obtain the optimal solution of formula (2); on the other hand, when the maximum iteration step size is reached, it selects the last iteration step size. As the optimal solution of formula (2).