Dependent task unloading method based on deep reinforcement learning under end-side cloud architecture
By using the advantageous action comment model of directed acyclic graph and deep reinforcement learning under the end-edge cloud architecture, the task offload decision is optimized, and the task offloading efficiency problem under complex dependencies is solved, achieving more efficient resource utilization and system performance improvement.
Patent Information
- Application Number
- CN202510275713.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-07-04
AI Technical Summary
The existing technology lacks a general description and analysis of task offloading with complex dependencies under the peer-end edge cloud architecture. Traditional optimization methods cannot effectively utilize system resources, resulting in inexpensive offloading efficiency and poor adaptability, especially in dynamic environments.
Directed acyclic graph (DAG) modeling is used to determine the priority of subtasks, combine the advantageous action comment model of deep reinforcement learning, and build a state space for offloading, and use the A2C model for strategy optimization and feedback adjustment.
It significantly improves the stability and adaptability of task offloading, reduces latency and energy consumption, improves system performance, especially under complex dependencies.
Smart Images

Figure CN120256094A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of task offloading in edge computing technology, and particularly relates to a method, device, electronic device, computer-readable storage medium, and computer program product for dependent task offloading based on deep reinforcement learning under an end-edge-cloud architecture. Background Art
[0002] With the continuous progress of information technology, the global informatization level has been continuously improved, and electronic intelligent devices have become increasingly popular. According to the latest statistical data, the total number of terminal connections in the Chinese mobile network has reached billions, and the intelligent Internet of Things system is being steadily built. This trend has not only promoted the process of social informatization but also brought a huge demand for computing resources. With the rapid increase in the number of terminals, the demand for computing power from users has also expanded. A large number of compute-intensive and latency-critical applications, such as autonomous vehicles, online games, and augmented reality, have put forward higher requirements for the computing power of terminal devices. As a solution, cloud computing provides powerful computing resources through remote data centers, making up for the deficiencies of user devices. However, the high latency characteristics of cloud computing have limited the user experience to a certain extent, especially in application scenarios that require quick responses.
[0003] To solve this problem, edge computing technology has emerged. A task can be processed locally, offloaded to an edge server, or offloaded to the cloud. Edge computing technology significantly reduces latency and improves data processing speed by deploying edge servers near users. The introduction of edge computing enables data processing to be closer to the data source, reducing data transmission latency and bandwidth consumption. On this basis, the end-edge-cloud architecture has been proposed, aiming to combine the rich computing resources of cloud servers and the low-latency advantages of edge servers to achieve more efficient utilization of computing resources. In this architecture, the cloud server serves as the core module, providing powerful computing and storage capabilities; while the edge server takes advantage of its geographical location to provide fast-response computing services for users.
[0004] At the same time, in practical applications, the management of dependencies between tasks is particularly crucial. These dependencies may manifest as data dependencies, control dependencies, or time dependencies. Among them, data dependency means that the output of one task becomes the input of another task; control dependency involves the execution order of tasks; and time dependency focuses on the real-time requirements of task execution. To effectively manage these dependencies, researchers have proposed various strategies to analyze and optimize task dependencies. For example, complex tasks are decomposed into multiple subtasks, and models such as directed acyclic graphs are used to clarify the dependencies between tasks.
[0005] In the development of the edge-cloud architecture, the allocation of computing resources and the coordination of user tasks are key issues. With the advancement of technology, computing offloading methods are also constantly updated. Initially, offloading solutions were mainly based on heuristic algorithms, which mainly focused on minimizing latency, minimizing energy consumption, or weighing latency and energy consumption in terms of optimization objectives. However, these traditional methods may encounter limitations in efficiency and adaptability when faced with large-scale mobile edge computing (MEC) systems or complex optimization problems. With the development of artificial intelligence technology, deep learning and reinforcement learning methods have been introduced into offloading decisions, further improving the performance of the system. For example, an autonomous management framework based on deep Q-learning technology can model problems through Markov decision processes and solve them through deep reinforcement learning to minimize the latency of service computing. The emergence of deep reinforcement learning (DRL) further combines the perception ability of deep learning and the decision-making ability of reinforcement learning. For example, some studies have proposed an online computing offloading scheme based on DRL, which considers both blockchain data mining tasks and data processing tasks, and introduces an adaptive genetic algorithm into the exploration of deep reinforcement learning, effectively accelerating the convergence speed and improving the robustness.
[0006] In existing technical solutions, most focus on task offloading in specific scenarios, such as vehicle-to-everything (V2X) and mobile edge computing. These technical solutions usually lack generality. Therefore, a more general offloading technical solution is needed to improve the applicability and flexibility of task offloading in different environments, better adapt to diverse computing requirements, enhance the scalability and reliability of the system; in addition, many existing technical solutions do not fully consider the dependencies between tasks and usually regard tasks as independent individuals for offloading. However, in practical applications, tasks usually have complex dependencies, and offloading without considering dependencies will lead to low offloading efficiency. Facing task offloading with complex dependencies, traditional optimization algorithms used in traditional technical solutions, such as heuristic algorithms like genetic algorithms and ant colony algorithms, although they can find feasible solutions to problems in some cases, usually require a large number of iterations, have poor adaptability to dynamically changing environments, and cannot handle continuous action spaces and high-dimensional state spaces.
[0007] All in all, existing technical solutions lack a general description and analysis of task offloading with complex dependencies under the edge-cloud architecture. At the same time, deep reinforcement learning models are rarely applied in this scenario. Traditional optimization methods cannot stably output efficient offloading solutions and cannot effectively utilize the respective advantages and computing power resources of cloud servers, edge servers, and terminal devices in the system. Summary of the Invention
[0008] The object of the present invention is to overcome the above-mentioned existing technical defects, such asFigure 4 As shown in the figure, a dependency task offloading method based on deep reinforcement learning under an edge-cloud architecture is provided, including:
[0009] Initial step: Obtain the application to be executed and the edge-cloud architecture Internet of Things for executing the application; the application includes multiple subtasks, and the edge-cloud architecture Internet of Things includes a terminal device, an edge server, and a cloud server; extract the dependency relationships between the subtasks to obtain a directed acyclic graph, and determine the priority of each subtask by analyzing the ready time and dependency conditions of the subtasks, and generate a queue containing all ready tasks;
[0010] Offloading step: During the time gap, collect the computing resources, earliest executable time, and network distance between devices of the terminal device, the edge server, and the cloud server to form system information; obtain the data size, computing requirements, and priority of the currently to-be-executed subtask to form task information; fuse the system information and the task information to construct a state space; based on the state space, use an advantage actor-critic model to evaluate the performance metrics of each offloading strategy to generate a task offloading decision;
[0011] Execution step: According to the task offloading decision, offload the currently to-be-executed subtask to the specified device. After execution, obtain the actual completion time, resource consumption, and delay of the subtask, and store them in the buffer pool; optimize the offloading strategy based on the data in the buffer pool;
[0012] Loop step: Repeat the offloading step and the execution step until all subtasks on all devices in the edge-cloud architecture Internet of Things are completed, obtain the execution result of the application, and end the offloading process.
[0013] The dependency task offloading method based on deep reinforcement learning under the edge-cloud architecture, where the initial step includes:
[0014] At the beginning of each time gap, initialize or update the states of all devices, including the current states of the terminal device, the edge server, and the cloud server, as well as the available computing resources and their respective task queues;
[0015] For the directed acyclic graph G=(V, E), V represents the set of subtask nodes of each terminal device, E represents the dependency relationships between subtasks, and e i,j represents that subtask J u,i has a dependency relationship with subtask J u,j ;
[0016] Traverse the task sets of each terminal device one by one, and generate the ready time of the current task according to the completion status of the predecessor tasks; collect all the ready but unoffloaded subtasks, sort them according to the priority, and generate a ready task queue to provide input for subsequent offloading decisions.
[0017] The method for offloading dependent tasks based on deep reinforcement learning under the edge-cloud architecture, wherein the offloading step includes: inputting the state space into a pre-trained advantage actor-critic model; the policy network in the advantage actor-critic model generates a probability distribution of possible actions, and the evaluation network evaluates the performance of these actions, calculates their latency and energy consumption, samples according to the probability distribution generated by the policy network, and determines whether the subtask is executed locally or offloaded to the edge server or the cloud server in combination with task dependencies, system resource status and historical experience.
[0018] The method for offloading dependent tasks based on deep reinforcement learning under the edge-cloud architecture, wherein the execution step includes:
[0019] If the resources of the specified device are insufficient, the subtask will queue in the execution task sequence of the specified device; the data in the buffer pool is used to evaluate and feedback the effectiveness of the offloading policy, and the advantage actor-critic model optimizes the parameters of its policy network according to the effectiveness.
[0020] As Figure 5 shown, the present invention also proposes a device for offloading dependent tasks based on deep reinforcement learning under the edge-cloud architecture, which includes:
[0021] An initial module that obtains the application program to be executed and the edge-cloud architecture Internet of Things for executing the application program; the application program includes multiple subtasks, and the edge-cloud architecture Internet of Things includes terminal devices, edge servers and cloud servers; extract the dependency relationships between the subtasks to obtain a directed acyclic graph, and determine the priority of each subtask by analyzing the ready time and dependency conditions of the subtasks, and generate a queue containing all the ready tasks;
[0022] An offloading module that, during the time gap, collects the computing resources, earliest executable time, and network distance between devices of the terminal device, the edge server, and the cloud server to form system information; obtains the data size, computing requirements, and priority of the current subtask to be executed to form task information; fuses the system information and the task information to construct a state space; based on the state space, uses the advantage actor-critic model to evaluate the performance metrics of each offloading policy to generate a task offloading decision;
[0023] The execution module unloads the currently to-be-executed subtask to the specified device according to the task offloading decision. After completion, it obtains the actual completion time, resource consumption, and latency of the subtask and stores them in the buffer pool. Based on the data in the buffer pool, it optimizes the offloading strategy.
[0024] The loop module repeatedly executes the offloading module and the execution module until all subtasks on all devices in the edge-cloud architecture IoT are completed, obtains the execution result of the application program, and ends the offloading process.
[0025] The described device for dependent task offloading based on deep reinforcement learning under the edge-cloud architecture, where the initial module includes:
[0026] At the beginning of each time interval, initialize or update the status of all devices, including the current status of terminal devices, edge servers, and cloud servers, as well as the available computing resources and their respective task queues.
[0027] The directed acyclic graph G=(V, E), where V represents the set of subtask nodes of each terminal device, E represents the dependency relationship between subtasks, and e i,j represents subtask J u,i and subtask J u,j have a dependency relationship;
[0028] Traverse the task sets of each terminal device one by one, generate the ready time of the current task according to the completion status of the predecessor tasks. Collect all the ready but unoffloaded subtasks and sort them according to the priority to generate a ready task queue, providing input for subsequent offloading decisions.
[0029] The offloading module includes: inputting the state space into the pre-trained advantage actor-critic model; the policy network in the advantage actor-critic model generates the probability distribution of possible actions, the evaluation network evaluates the performance of these actions, calculates their latency and energy consumption, samples according to the probability distribution generated by the policy network, and determines whether the subtask is executed locally or offloaded to the edge server or cloud server in combination with task dependencies, system resource status, and historical experience.
[0030] The described device for dependent task offloading based on deep reinforcement learning under the edge-cloud architecture, where the execution module includes:
[0031] If the resources of the specified device are insufficient, the subtask will queue in the execution task sequence of the specified device. The data in the buffer pool is used to evaluate and feedback the effectiveness of the offloading strategy, and the advantage actor-critic model optimizes the parameters of its policy network according to this effectiveness.
[0032] The present invention also provides an electronic device, which includes the above-described device for dependent task offloading based on deep reinforcement learning under the edge-cloud architecture. The electronic device is either connected to an information display device, which is configured to display the execution result of the application program with display parameters, attributes set by the user, or through an artificial intelligence model.
[0033] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method for dependent task offloading based on deep reinforcement learning under the edge-cloud architecture are implemented.
[0034] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the method for dependent task offloading based on deep reinforcement learning under the edge-cloud architecture are implemented.
[0035] As can be seen from the above solutions, the advantages of the present invention are as follows:
[0036] In summary, for the multiple subtasks and their dependencies in each application program, the technical solution of the present invention uses a directed acyclic graph (DAG) to model each subtask set, determines the ready state of each subtask, and thus decides whether it enters offloading; constructs a task offloading model for the edge-cloud architecture and abstracts the problem into a mixed integer programming problem of NP-hardness; finally, combining the above model and problem, with the goal of minimizing the weighted average of the latency and energy consumption of all terminal devices, a dependent task offloading strategy based on deep reinforcement learning under the edge-cloud architecture is proposed. The technical solution of the present invention has better convergence and stability compared with other DRL algorithms in different scenarios. Compared with three baseline algorithms, the costs are reduced by 13.81%, 67.33%, and 81.04% respectively. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 It is an application scenario diagram of the method for dependent task offloading under the edge-cloud architecture of the present invention;
[0038] Figure 2 It is a structural diagram of the deep reinforcement learning model of the method for dependent task offloading under the edge-cloud architecture of the present invention;
[0039] Figure 3 It is a schematic flowchart of the method for dependent task offloading under the edge-cloud architecture of the present invention;
[0040] Figure 4 It is a flowchart of the method of the present invention;
[0041] Figure 5 It is a module diagram of the device of the present invention;
[0042] Figure 6 Structural schematic diagram of the first electronic device of the present invention;
[0043] Figure 7 Structural schematic diagram of the application environment of the first electronic device of the present invention;
[0044] Figure 8 Structural schematic diagram of the second electronic device of the present invention.
[0045] Reference numerals:
[0046] A - First electronic device;
[0047] B - Dependency task offloading device based on deep reinforcement learning under the edge cloud architecture;
[0048] C - Data acquisition device;
[0049] D - Information display device;
[0050] 1000 - Second electronic device;
[0051] Ⅰ - Computing unit;
[0052] Ⅱ - ROM;
[0053] Ⅲ - RAM;
[0054] Ⅳ - Bus;
[0055] Ⅴ - Interface;
[0056] Ⅵ - Input unit;
[0057] Ⅶ - Output unit;
[0058] Ⅷ - Storage medium;
[0059] Ⅸ - Communication unit. Detailed implementation manners
[0060] It should be noted that in this application, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non - exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device.
[0061] Without further limitations, an element qualified by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes the said element.
[0062] The processor described in the present invention is the control center of an electronic device, which can be a single processor or a collective term for multiple processing elements. For example, it can be one or more central processing units (CPUs), or a specific application integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention. For instance: one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs).
[0063] Optionally, the processor can execute various functions of the electronic device by running or executing software programs stored in the memory and by invoking data stored in the memory.
[0064] In a specific implementation, as an embodiment, the processor can include one or more CPUs. Each of these processors can be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). Here, the processor can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions). The electronic device can include: servers, desktop computers, laptop computers, smartphones, tablet computers, embedded computers, etc., where the embedded computer includes vehicles and robots, etc.
[0065] The memory is used to store the software program for implementing the solution of the present invention and is controlled by the processor for execution. The specific implementation manner can refer to the above method embodiments and will not be elaborated here.
[0066] It should be noted that the structure of the electronic device shown in the drawings of the present invention does not constitute a limitation thereto. The actual knowledge structure recognition device may include more or fewer components than those shown in the drawings, or combine certain components, or have different component arrangements.
[0067] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.
[0068] It should also be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood with reference to the context.
[0069] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.
[0070] It should also be understood that in various embodiments of the present invention, the sequence numbers of the above processes do not imply the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0071] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in electrical, mechanical, or other forms.
[0072] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0073] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0074] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0075] The present invention proposes a method for dynamically offloading dependent tasks implemented through deep reinforcement learning under an edge-cloud collaborative architecture, which is used to optimize the latency and energy consumption of task offloading and adapt to complex and changing system environments and task requirements. For example, when the task requirement of an intelligent program is intelligent driving, the subtask set can be driving-related tasks, including perception tasks (image recognition, data processing, etc.), decision-making tasks (path planning, traffic flow analysis), control tasks (vehicle acceleration, steering, braking), auxiliary tasks, etc. The terminal device is the in-vehicle system of an intelligent vehicle; the edge server: a network node close to the vehicle, such as an edge server; the cloud server: a remote data center, a server cluster.
[0076] This method includes the following main steps:
[0077] Step 1: System state initialization and task modeling. During the task offloading process, both the system state and the task sequence are in dynamic changes. To ensure the accuracy of the offloading decision, the system needs to update the current state in real time and construct a task dependency model. State initialization: The agent initializes or updates in real time key information such as the task queue, computing resources, and network conditions in the system to ensure that all data is accurate and up-to-date; Dependency modeling: Use a directed acyclic graph (DAG) to model the subtask set of each application program, representing the dependency relationship between the subtasks of the application program. By analyzing the ready time of the task (the time when all the prerequisite tasks required by the task have been completed.) and the dependency conditions, determine the priority of each subtask, and generate a queue containing all the ready tasks as the input for subsequent decisions. Here, the agent refers to a system or entity that can autonomously perceive the environment, make decisions, and execute actions. For example, the agent in intelligent driving should only be the entire intelligent driving system.
[0078] Step 2: State space generation and offloading decision. The task offloading decision depends on the current state and task information of the system. This step aims to construct the state space and make an offloading decision based on the optimized strategy. State space generation: During each time interval, the agent collects system state information such as the computing resources, earliest executable time, and network distance between devices of the terminal device, edge server, and cloud server. At the same time, obtain the relevant attributes of the current task, including task information such as data size, computing requirements, and priority. By fusing the system state information with the task information, construct a complete state space; Decision generation: Based on the input of the state space, the agent uses the Advantage Actor-Critic (A2C) model to evaluate the potential impact of different offloading strategies, including the latency, energy consumption of task execution, and the impact on the overall system performance, and selects the optimal task offloading decision.
[0079] Among them, the state space consists of the global system state and the current offloading task state. The global state includes the terminal device where the current task is located, as well as the computing resources F and the earliest executable time T of each edge server and cloud server. The offloading task state includes the basic attribute J u,k and the distance matrix L between the terminal device and each edge server u,k .
[0080] Step 3: Task execution and dynamic feedback optimization. Task execution: According to the generated optimal decision, the agent offloads the task to the specified device (terminal, edge or cloud) for execution and monitors the completion of the task in real time; Feedback collection: After the task execution is completed, the agent collects data such as the actual completion time, resource consumption, delay and reward of the task. This information will be used to evaluate the effect of the current offloading strategy and stored as feedback data in the buffer pool of the learning model; Strategy optimization: Based on the collected execution data, the agent dynamically optimizes the offloading strategy to gradually improve the adaptability and efficiency of the decision-making. Through multiple iterative learning, the agent can effectively cope with the dynamic changes of the system state and task requirements and optimize the offloading effect.
[0081] Step 4: Global task completion and process termination. When all subtasks on all terminal devices are completed, the entire offloading process ends. Through the full-process optimization of task allocation and execution, this method realizes the efficient offloading of dependent tasks in the edge-cloud-terminal collaborative environment, significantly reduces the delay and energy consumption, and improves the system performance.
[0082] In an instance, Step 1 includes:
[0083] Step 11: System state initialization and update. At the beginning of each time slot, the agent first initializes or updates the states of all devices in the system, including the current states of terminal devices, edge servers and cloud servers, as well as their available computing resources and their respective task queues. By obtaining this information in real time, the agent can comprehensively master the current system operation and ensure the accuracy and timeliness of the state data. This provides a reliable basis for the subsequent formulation of task offloading decisions and enables the system to adapt to the dynamically changing environment.
[0084] Step 12: Task dependency modeling and generation of the ready subtask queue. The agent models the task dependencies of each terminal device through a directed acyclic graph (DAG) to describe the dependencies between tasks. Specifically, G=(V, E), where V represents the set of subtask nodes of each terminal device, E represents the dependency relationship between subtasks, and e i,j represents subtask J u,i and subtask J u,jThere is a dependency relationship. The agent traverses the task sets of each terminal device one by one, calculates the ready time of the current task according to the completion status of the predecessor tasks. Then, the agent collects all the ready but unoffloaded subtasks, sorts them according to the priority, generates a ready task queue, and provides the input for the subsequent offloading decision.
[0085] In one instance, step 2 includes:
[0086] Step 21: Generation of the state space. The agent integrates the collected system information and task information to generate the state space of the current time slot. The state space contains key information that has an important impact on the offloading decision, including the computing resources of the terminal device, edge server, and cloud server, the earliest executable time, the target distance matrix between devices, and the specific attributes of the current task (such as data size, computing requirements, and priority). By flattening and concatenating this information, a complete state space is formed to provide input for the decision-making model.
[0087] Step 22: Generation of the offloading decision. The agent inputs the generated state space into the pre-trained A2C model. The policy network in the model generates the probability distribution of possible actions based on the current state, and the evaluation network evaluates the performance of these actions, calculates their potential time delay, energy consumption, and impact on system performance. The agent samples according to the probability distribution generated by the policy network, combines task dependencies, system resource status, and historical experience to determine whether the task is executed locally, or offloaded to the edge server or cloud server to ensure that the selected decision can optimize the overall system performance.
[0088] In one instance, step 3 includes:
[0089] Step 31: Action execution. According to the decision provided by the A2C decision model, the agent guides the execution of the task. For each subtask, there are M + 2 offloading methods: execute locally, or offload to the edge server or cloud server. The agent selects the most suitable executor according to the current resource status and task characteristics. If the resources of the executor are insufficient, the task will queue in its execution task sequence. For each offloading decision, we have conducted a complete and comprehensive computational modeling analysis to obtain the formula for its execution time delay and energy consumption. M is a positive integer greater than or equal to 1.
[0090] Step 32: Feedback collection and model optimization. After unloading the current task, the agent checks whether there are still tasks to be unloaded. If so, it will restart the unloading process until all tasks are unloaded. Meanwhile, the execution results of each task, including metrics such as completion time, reward, latency, and energy consumption, are used to evaluate and feedback the effectiveness of the offloading strategy. These execution results are stored as feedback information in the buffer pool of the A2C model (i.e., the experience replay buffer) to provide data support for subsequent decision-making offloading learning. Through this iterative learning mechanism, the A2C model continuously optimizes the parameters of the policy network, enabling the agent to adapt to the dynamic changes of the system state and the evolution of task requirements.
[0091] Step 4: Global task completion and process termination. As time goes by, the decision-making and optimization processes in Steps 1 to 3 are iteratively carried out until all subtasks of all terminal devices are effectively processed, completing the entire task offloading process. In this way, the efficient offloading of dependent tasks in the edge-cloud-end collaborative environment is achieved, significantly reducing latency and energy consumption and improving system performance.
[0092] To make the above features and effects of the present invention more clearly understandable, specific embodiments are hereinafter given and described in detail in conjunction with the accompanying drawings of the specification. This specification discloses one or more embodiments incorporating the features of the present invention. The disclosed embodiments are merely for illustrative purposes. The scope of protection of the present invention is not limited to the disclosed embodiments, and the present invention is defined by the appended claims.
[0093] Figure 1 The application scenario of the present invention is shown, mainly applied to the task offloading in the edge-cloud-end collaborative architecture. During the task offloading process, the agent unloads each subtask to different executors respectively. Since the cloud server has rich computing power resources, it can execute multiple subtasks simultaneously, but there is an additional transmission latency due to the long distance; the edge server has relatively less computing power and can only execute tasks equal to the number of processors simultaneously, but it is closer and has a lower transmission latency; the local processor has the least computing power and can only process one subtask at a time, but there is no transmission latency. As shown in the system model, a terminal device runs an application program with multiple subtasks, and there are dependencies between the subtasks, which are represented by arrows. Therefore, an application program is a directed acyclic graph.
[0094] Figure 2Shows the most core model architecture of the present invention. The A2C model is a complex decision-making model used to optimize the task offloading process under the edge-cloud architecture. The system consists of two core components: a policy network and a value network. The policy network is the main body of decision-making. It receives detailed state information including global state, subtask attributes, and distance matrix as input. This information covers the network topology and resource status from each terminal device to the edge server and the cloud data center, providing the policy network with a comprehensive environmental awareness. Based on these inputs, the policy network generates specific offloading decisions, determining whether each subtask is processed locally or offloaded to a specific edge server or cloud server. At the same time, the value network plays a role in value evaluation. It evaluates the expected value of the decisions proposed by the policy network by calculating the temporal difference error (TD Error). This evaluation process involves quantifying the difference between the actual reward and the expected reward, where the reward signal reflects the results of task offloading, including key performance indicators such as task completion time, energy consumption, and latency. The evaluation results of the value network are used to guide the learning and optimization process of the policy network. Through the calculation of policy gradients and parameter updates, the policy network gradually improves its decision-making strategy to reduce the temporal difference error and increase the obtained reward.
[0095] During the entire interaction process between the agent and the environment, the A2C model continuously learns and adapts, optimizing the task offloading decision through an iterative approach. This adaptive mechanism enables the agent to dynamically make decisions that result in higher rewards when facing real-time changing environmental states and subtask requirements. Over time, through this continuous learning and optimization, the A2C model significantly improves the efficiency of task offloading and the performance of the entire system, enabling it to make more low-latency and low-energy consumption offloading decisions for tasks with complex dependencies under the edge-cloud architecture.
[0096] Figure 3 Shows a flowchart of an example of offloading dependent tasks under the edge-cloud architecture of the present invention. The offloading method includes the following steps:
[0097] Step S310: Initialize or update the system state. For each terminal device, each edge server, and the cloud server, initialize or update information such as its computing resources and task queue according to the initial system requirements or the results returned in the previous time slot.
[0098] S311: Update the task state. In this step, the agent traverses the subtask sets of all terminal devices and then updates the task state according to the current time state and task dependencies.
[0099] Each subtask has three state flags which are the ready state flag, the offloading state flag, and the completion state flag respectively. Among them Indicates that all tasks before the subtask have been completed, Indicates that the subtask has been unloaded by the agent, Indicates that the subtask has been completed. Therefore, subtask J u,k has four states: not ready, ready but not unloaded, unloaded but not completed, and completed.
[0100]
[0101] When all the predecessor nodes P(J u,k ) of a certain subtask have been completed, the subtask enters the ready state. The ready time RT u,k of each subtask can be calculated by the following formula:
[0102]
[0103] S312: Access all terminal devices, collect all subtasks that are ready but not unloaded, and then generate a ready subtask queue according to the priority of the subtasks.
[0104] Step S320: Read the first subtask in the ready subtask queue as the unloading task, then generate a state space based on the current unloading task and system information, and input it into the A2C model based on deep reinforcement learning for unloading decision-making.
[0105] In this step, the collected information includes the global system state and the current unloading task state. The global state includes the terminal device where the current task is located and the computing resources F and the earliest executable time T of each edge server and cloud server. The unloading task state includes the basic attributes J u,k and the distance matrix L u,k .
[0106]
[0107] J u,k ={d u,k ,w u,k ,q u,k}
[0108] S321: The agent first uses the policy network to generate the probability distribution π(a t |s t ) of all possible actions in the current state space s t ), randomly selects an action a t according to this probability distribution and passes it to the evaluation network to update the network parameters and optimize the task unloading strategy.
[0109] Step S330: According to the output action a t, we offload the task to the corresponding location. There are M + 2 offloading strategies:
[0110]
[0111] In this step, a t = 0 indicates local execution, and the task is offloaded to the local sub-task queue for execution; a t = i, i ∈ {1, …, M} indicates edge computing, and the task is offloaded to edge server i and enters the sub-task queue for execution on this edge server; a t = M + 1 indicates cloud computing, and the task is offloaded to the sub-task queue for cloud execution.
[0112] Based on the output action of the network as the offloading strategy, specific task execution is carried out, and its execution delay and energy consumption are calculated. Assuming there is 1 cloud server and M edge servers with N cores, in the time slot t, for offloading task J u,k is processed. Each sub-task includes three inherent attributes J u,k = {d u,k , w u,k , q u,k}: task data size, total number of loops required for calculation, and execution priority. In this method, the specific offloading strategies are local execution, edge computing, and cloud computing:
[0113] a) Local execution: Since the local terminal device has less computing resources, it can only execute one task at a time, and there is a queuing delay in the local sub-task queue L(u). The start time of its task can be calculated according to the following formula:
[0114]
[0115] The local execution time is: where f u represents the computing power of the local terminal device in each cycle. From this, we can obtain the final completion time of this task as:
[0116]
[0117] So the delay for completing this task is:
[0118]
[0119] At the same time, by subtracting the idle power consumption during the execution time from the computing power consumption, the additional power consumption of local computing for this offloading task can be obtained, where ζ u = 10 -27 (f u ) 2 is the energy consumption generated in each CPU cycle:
[0120]
[0121] b) Edge computing: When offloading to edge server m, the task completion process goes through two stages:
[0122] Task uploading and task computing. At the same time, edge servers usually have limited computing resources. This method assumes that each edge server has N cores, representing that N subtasks can be executed simultaneously. Therefore, for offloading task J u,k The start time on edge server m is:
[0123]
[0124] where R(m) represents the execution subtask queue of edge server m, and the function f N represents finding the Nth largest number in the set. During the task uploading stage, the uplink transmission rate v between the terminal device u and the edge server m u,m can be obtained according to the Shannon formula:
[0125]
[0126] The uploading time is: The computing time is: From this, the final completion time of the task executed on the edge server can be obtained is:
[0127]
[0128] Therefore, the latency of the task during edge computing and the energy consumed by the terminal device are:
[0129]
[0130] c) Cloud computing: Cloud servers have abundant computing resources. Since tasks offloaded to cloud servers do not need to wait and are executed immediately.
[0131] During the cloud computing process, first, the task is uploaded from the local to the nearest edge server m through a wireless connection. The uploading time is: Then, it is transmitted through the wired link between the edge server and the cloud service. The transmission time is: Finally, the task is computed on the cloud server: From this, the cloud computing latency of the task can be calculated is:
[0132]
[0133] During the cloud computing process, the additional energy consumption of the terminal device is the same as that of edge computing.
[0134] S321: After the agent takes an action, the state transitions from s t to s t+1 , and a reward in this state is generated. The goal of deep reinforcement learning is to maximize the accumulated reward. In this method, our goal is to minimize the total delay and energy consumption, so the reward function is designed as the negative of the weighted sum of the current delay and energy consumption difference.
[0135] The offloading policy for each subtask J u,k can be expressed as: a u,k = 1 indicates local execution, indicates offloading to edge server i, and Z u,k = 1 indicates offloading to the cloud server. Therefore, the total delay and energy consumption are calculated as:
[0136]
[0137] In this method, our goal is to minimize the total delay and energy consumption, so the reward function is designed as the negative of the weighted sum of the current delay and energy consumption difference.
[0138]
[0139] S332: After the current task offloading process ends, the agent will check to confirm whether all tasks have been successfully offloaded. If it is found that there are still tasks that have not been completed offloading, the agent will loop back to the starting point of the offloading process and repeat the above steps. This process will continue until all tasks are properly offloaded, ensuring the integrity and continuity of task offloading.
[0140] S333: At the same time, in the task offloading scenario, the agent follows a series of carefully designed steps to perform the offloading task and stores the result of each offloading in the experience replay buffer of the A2C model. These data are then used as the key parameters for model optimization. The agent also sets a fixed update frequency to perform model updates at an appropriate time.
[0141] When the update period arrives, the agent will first use the evaluation network to evaluate the value V(s t ) of the current state. Then, it will calculate the current temporal difference error and use this error to approximately calculate the advantage function A π (s t ,a t ), and its formula is as follows:
[0142] A π(s t ,a t )=δ t =r t+1 +γV(s t+1 )-V(s t )
[0143] where δ t represents the time difference error, r t+1 is the immediate reward obtained after executing the action, γ is the discount factor for future rewards, and V(s t+1 ) and V(s t ) are the estimated values of the new state and the current state respectively.
[0144] Subsequently, the agent will apply the policy gradient formula in the reinforcement learning algorithm to calculate the gradient of the policy network. The gradient calculation formula is as follows:
[0145]
[0146] Through this formula, the agent can evaluate the gradient of the policy network parameters θ and update them accordingly to optimize its decision-making policy. The agent uses the time difference error to update the parameters of the evaluation network, further improving the accuracy of the state value estimation. In this way, the policy and evaluation networks are gradually improved in continuous iterations to better adapt to the requirements of task offloading.
[0147] S4: Global task completion and process termination. As time goes by, the decision-making and optimization processes of S1 to S3 are continuously iterated until all subtasks of all terminal devices are effectively processed, completing the entire task offloading process.
[0148] The following is a system embodiment corresponding to the above method embodiment. This embodiment can be implemented in cooperation with the above embodiment. The relevant technical details mentioned in the above embodiment are still valid in this embodiment. To avoid repetition, they will not be elaborated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied in the above embodiment.
[0149] As Figure 5 shown, the present invention also proposes a dependency task offloading device based on deep reinforcement learning under an edge-cloud architecture, which includes:
[0150] An initial module that obtains the application program to be executed and the edge-cloud architecture Internet of Things for executing the application program; the application program includes multiple subtasks, and the edge-cloud architecture Internet of Things includes terminal devices, edge servers, and cloud servers; extracts the dependency relationships between the subtasks to obtain a directed acyclic graph, and determines the priority of each subtask by analyzing the ready time and dependency conditions of the subtasks, and generates a queue containing all ready tasks;
[0151] The unloading module collects the computing resources, earliest executable time, and network distance between devices of the terminal device, the edge server, and the cloud server during a time interval to form system information; obtains the data size, computing requirements, and priority of the currently pending subtask to form task information; fuses the system information and the task information to construct a state space; based on the state space, uses an advantage actor-critic model to evaluate the performance metrics of each unloading strategy to generate a task unloading decision.
[0152] The execution module unloads the currently pending subtask to the specified device according to the task unloading decision. After execution, it obtains the actual completion time, resource consumption, and delay of the subtask and stores them in the buffer pool; optimizes the unloading strategy based on the data in the buffer pool.
[0153] The loop module repeatedly executes the unloading module and the execution module until all subtasks on all devices in the terminal-edge-cloud architecture Internet of Things are completed, obtains the execution result of the application program, and ends the unloading process.
[0154] The described device for dependent task unloading based on deep reinforcement learning under the terminal-edge-cloud architecture, where the initial module includes:
[0155] At the beginning of each time interval, initialize or update the states of all devices, including the current states of the terminal device, the edge server, and the cloud server, as well as the available computing resources and their respective task queues.
[0156] The directed acyclic graph G=(V, E), where V represents the set of subtask nodes of each terminal device, E represents the dependency relationship between subtasks, and e i,j represents that subtask J u,i has a dependency relationship with subtask J u,j ;
[0157] Traverse the task sets of each terminal device one by one, generate the ready time of the current task according to the completion status of the predecessor tasks; collect all the ready but unloaded subtasks and sort them according to the priority to generate a ready task queue, providing input for subsequent unloading decisions.
[0158] The unloading module includes: inputting the state space into a pre-trained advantage actor-critic model; the policy network in the advantage actor-critic model generates the probability distribution of possible actions, the evaluation network evaluates the performance of these actions, calculates their delay and energy consumption, samples according to the probability distribution generated by the policy network, and determines whether the subtask is executed locally or unloaded to the edge server or the cloud server in combination with task dependencies, system resource status, and historical experience.
[0159] The described device for dependent task offloading based on deep reinforcement learning under the edge-cloud architecture, wherein the execution module includes:
[0160] If the resources of the specified device are insufficient, the subtasks will queue up in the execution task sequence of the specified device; the data in the buffer pool is used to evaluate and feedback the effectiveness of the offloading strategy, and the advantage-actor critic model optimizes the parameters of its policy network according to this effectiveness.
[0161] As Figure 6 shown, in another embodiment of the present invention, a first electronic device A is also proposed, including the described device for dependent task offloading based on deep reinforcement learning under the edge-cloud architecture.
[0162] As Figure 7 shown, the first electronic device A can also be connected to the data acquisition device C and the information display device D through a wired or wireless information transmission scheme. The data acquisition device C is used to acquire the application programs to be executed, such as the autonomous driving task described in the embodiments of the present invention, and the information display device D is used to display the execution results of the application programs analyzed by the present invention.
[0163] Among them, the information display device D can process and organize the data output by the first electronic device A based on the information display mechanism to improve the readability of the data output by the first electronic device A. The information display mechanism can be preset manually. For example, the data output by the first electronic device A is visually displayed, and it can display according to the display parameters and / or attributes set by the user. The display parameters can be, for example, the display data range, and the display attributes can be, for example, the display font, color, whether to scroll and play, etc. Present the key information specified by the user to the user. For example, vehicle control instructions, road surface information, driving information, etc., so that the user can understand this information more timely without having to access a secondary page or scroll the page, saving the user's operations. Or the information display mechanism can be an artificial intelligence AI display model, which can learn the key information that the user focuses on according to the user's previous usage habits, such as viewing duration, click times, editing times, etc., and then automatically present rich and necessary key information to the user.
[0164] The present invention also provides a computer program product, the computer program product includes a computer program, the computer program can be stored on a readable storage medium, and when the computer program is executed by a processor, the computer can execute the method for dependent task offloading based on deep reinforcement learning under the edge-cloud architecture provided by the above-mentioned various methods.
[0165] In another embodiment, the present invention further provides a storage medium VIII for storing a computer program for executing the method for offloading dependent tasks based on deep reinforcement learning in the edge-cloud architecture. It should be understood that the storage medium in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0166] Figure 8 FIG. shows a schematic block diagram of a second electronic device 1000 that may be used to implement the embodiments of the present invention. The second electronic device 1000 is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The second electronic device 1000 may also represent various forms of mobile devices, such as, for example, personal digital assistants, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present invention described and / or claimed herein. The second electronic device 1000 may be the same as or different from the first electronic device A.
[0167] The second electronic device 1000 includes a computing unit I, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory II (ROM) or a computer program loaded from a storage medium VIII into a random access memory (RAM) III. In the RAM III, various programs and data required for the operation of the device 1000 can also be stored. The computing unit I, the ROM II, and the RAM III are connected to each other via a bus IV. An input / output (I / O) interface V is also connected to the bus IV.
[0168] Multiple components in the second electronic device 1000 are connected to the I / O interface V, including: an input unit VI, such as a keyboard, a mouse, etc.; an output unit VII, such as various types of displays, speakers, etc.; a storage medium VIII, such as a magnetic disk, an optical disc, etc.; and a communication unit IX, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit IX allows the second electronic device 1000 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0169] The computing unit I can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit I include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit I executes the various methods and processes described above, such as method steps S1 - S4. For example, in some embodiments, the method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage medium VIII. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 1000 via the ROM II and / or the communication unit IX. When the computer program is loaded into the RAM III and executed by the computing unit I, one or more steps of the method described above can be executed. Alternatively, in other embodiments, the computing unit I can be configured to execute the method in any other appropriate manner (e.g., by means of firmware).
[0170] Although the embodiments of the present invention have been disclosed as above, they are not limited only to the applications listed in the specification and the embodiments. It can be fully applied to various fields suitable for the present invention. For those familiar with the art, additional modifications can be easily made. Therefore, without departing from the general concept defined by the claims and the equivalent scope, the present invention is not limited to the specific details and the illustrated and described examples here.
Claims
1. A dependence task offloading method based on deep reinforcement learning under an edge-cloud architecture, characterized in that, Including: An initial step of obtaining an application to be executed and an edge-cloud architecture Internet of Things for executing the application; The application includes multiple subtasks, and the edge-cloud architecture Internet of Things includes a terminal device, an edge server, and a cloud server; Extracting the dependencies between the subtasks to obtain a directed acyclic graph, determining the priority of each subtask by analyzing the ready time and dependency conditions of the subtasks, and generating a queue containing all the ready tasks; An offloading step of, during a time gap, collecting the computing resources, earliest executable time, and network distance between devices of the terminal device, the edge server, and the cloud server to form system information; obtaining the data size, computing requirements, and priority of the currently to-be-executed subtask to form task information; fusing the system information and the task information to construct a state space; Based on the state space, using an advantage actor-critic model to evaluate the performance metrics of each offloading strategy to generate a task offloading decision; An execution step of, according to the task offloading decision, offloading the currently to-be-executed subtask to a specified device, and after execution is completed, obtaining the actual completion time, resource consumption, and latency of the subtask and storing them in a buffer pool; Optimizing the offloading strategy based on the data in the buffer pool; A loop step of repeatedly executing the offloading step and the execution step until all subtasks on all devices in the edge-cloud architecture Internet of Things are completed, obtaining the execution result of the application, and ending the offloading process.
2. The method for offloading dependent tasks based on deep reinforcement learning under the edge-cloud architecture according to claim 1, characterized in that, The initial step includes: At the start of each time gap, initializing or updating the states of all devices, including the current states of the terminal device, the edge server, and the cloud server, as well as the available computing resources and their respective task queues; The directed acyclic graph G=(V, E), where V represents the set of sub-task nodes of each terminal device, E represents the dependency relationship between sub-tasks, and e i,j represents sub-task J u,i and sub-task J u,j have a dependency relationship; Traversing the task sets of each terminal device one by one, generating the ready time of the current task according to the completion status of the predecessor tasks; collecting all the ready but unoffloaded subtasks and sorting them according to the priority to generate a ready task queue, providing input for subsequent offloading decisions.
3. The method for offloading dependent tasks based on deep reinforcement learning under the edge-cloud architecture according to claim 1, characterized in that The offloading step includes: inputting the state space into a pre-trained advantage actor-critic model; the policy network in the advantage actor-critic model generates a probability distribution of possible actions, the evaluation network evaluates the performance of these actions, calculates their latency and energy consumption, samples according to the probability distribution generated by the policy network, and determines whether the subtask is executed locally, or offloaded to the edge server or the cloud server in combination with task dependencies, system resource status, and historical experience.
4. The method for offloading dependent tasks based on deep reinforcement learning under the edge-cloud architecture according to claim 1 or 2 or 3, characterized in that, The execution step includes: If the resources of the specified device are insufficient, the subtask will queue in the execution task sequence of the specified device; the data in the buffer pool is used to evaluate and feedback the effectiveness of the offloading strategy, and the advantage actor-critic model optimizes the parameters of its policy network according to the effectiveness.
5. A dependency task offloading device based on deep reinforcement learning under an edge-cloud architecture, characterized in that, Including: An initial module that obtains an application to be executed and an edge-cloud architecture Internet of Things for executing the application; The application includes multiple subtasks, and the edge-cloud architecture Internet of Things includes a terminal device, an edge server, and a cloud server; Extract the dependencies between the subtasks to obtain a directed acyclic graph, and determine the priority of each subtask by analyzing the ready time and dependency conditions of the subtasks, and generate a queue containing all the ready tasks; Unloading module, during the time gap, collect the computing resources, earliest executable time, and network distance between devices of the terminal device, the edge server, and the cloud server to form system information; obtain the data size, computing requirements, and priority of the currently pending subtask to form task information; fuse the system information and the task information to construct a state space; Based on the state space, use the advantage actor-critic model to evaluate the performance metrics of each offloading strategy to generate a task offloading decision; Execution module, according to the task offloading decision, offload the currently pending subtask to the specified device. After execution, obtain the actual completion time, resource consumption, and latency of the subtask and store them in the buffer pool; Optimize the offloading strategy based on the data in the buffer pool; Loop module, repeatedly execute and call the offloading module and the execution module until all subtasks on all devices in the terminal-edge-cloud architecture Internet of Things are completed, obtain the execution result of the application program, and end the offloading process.
6. The dependency task offloading device based on deep reinforcement learning under the edge-cloud architecture according to claim 1, characterized in that, The initial module includes: At the beginning of each time gap, initialize or update the status of all devices, including the current status of the terminal device, the edge server, and the cloud server, as well as the available computing resources and their respective task queues; The directed acyclic graph G=(V, E), where V represents the set of sub-task nodes of each terminal device, E represents the dependency relationship between sub-tasks, and e i,j represents sub-task J u,i and sub-task J u,j have a dependency relationship; Traverse the task set of each terminal device one by one, and generate the ready time of the current task according to the completion status of the predecessor task; collect all the ready but unoffloaded subtasks and sort them according to the priority to generate a ready task queue, providing input for subsequent offloading decisions; The offloading module includes: inputting the state space into a pre-trained advantage actor-critic model; the policy network in the advantage actor-critic model generates the probability distribution of possible actions, the evaluation network evaluates the performance of these actions, calculates their latency and energy consumption, samples according to the probability distribution generated by the policy network, and determines whether the subtask is executed locally, or offloaded to the edge server or the cloud server in combination with task dependencies, system resource status, and historical experience.
7. The apparatus for dependent task offloading based on deep reinforcement learning under the edge-cloud architecture according to claim 5 or 6, characterized in that, The execution module includes: If the resources of the specified device are insufficient, the subtask will queue up in the execution task sequence of the specified device; the data in the buffer pool is used to evaluate and feedback the effectiveness of the offloading strategy, and the advantage actor-critic model optimizes the parameters of its policy network according to the effectiveness.
8. An electronic device, characterized in that, An apparatus for dependent task offloading based on deep reinforcement learning under a terminal-edge-cloud architecture according to claims 5-7, the electronic device is connected to an information display device, and the information display device is used to display the execution result of the application program with display parameters, attributes set by the user, or through an artificial intelligence model.
9. A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method for dependent task offloading based on deep reinforcement learning under a terminal-edge-cloud architecture according to any one of claims 1-4 are implemented.
10. A computer program product comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method for offloading dependent tasks based on deep reinforcement learning under the edge-cloud architecture described in any one of claims 1-4.
Citation Information
Cited By
Edge computing task unloading method based on deep reinforcement learning
CN122219999A
An edge computing task offloading method based on deep reinforcement learning
CN122219999B