Task unloading instant decision-making method for micro-service architecture

By using deep Q network to build a task offload decision model under the microservice architecture, the problem of insufficient adaptability of task offload is solved, efficient and real-time task offload decisions are achieved, and the stability and resource utilization of the system are improved.

CN120281812APending Publication Date: 2025-07-08CHONGQING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510643059.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing technology has insufficient adaptability to task offloading under the microservice architecture, simple scheduling rules lead to poor accuracy, and traditional heuristic algorithms have high computing overhead, making it difficult to meet real-time decision-making needs.

Method used

The task offload decision model built on the deep Q network is adopted, and the task offload decision model is built through the deep Q network, and the environment changes are sensed in real time and decision strategies are adjusted to adapt to the task offload environment under the microservice architecture, and to improve the accuracy and real-timeness of task offloading.

Benefits of technology

Through intelligent decision-making of deep Q networks, the accuracy and efficiency of task offloading are improved, and a large number of state-action pairs can be quickly processed, decision-making time is shortened, and system stability and reliability are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120281812A_ABST
    Figure CN120281812A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of internet big data, in particular to a task unloading instant decision-making method for a micro-service architecture, which comprises the following steps: S1, acquiring an SFC request to form a task group; one SFC request is formed by combining a plurality of micro-services in a chained mode; s2, sorting all the micro-services requested by the SFC in the task group according to the micro-service execution dependency relationship to form a task queue; s3, constructing a task unloading decision model based on the deep Q network, and training the task unloading decision model; s4, generating a task unloading decision for the micro-service in the task queue through the trained task unloading decision model; the task unloading decision comprises an edge server and a computing unit which are distributed for the micro-service; and S5, executing processing of the corresponding micro-service based on the task unloading decision. The method can adapt to a task unloading environment under a micro-service architecture, and meanwhile, the accuracy and the real-time performance of task unloading can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of Internet big data, and particularly to an instant decision-making method for task offloading in a microservice architecture. Background Art

[0002] With the exponential expansion of the scale of Internet of Things terminals, the traditional monolithic service architecture is prone to problems such as redundant data storage, redundant deployment of service functions, and low utilization rate of computing resources, and it is difficult to form a collaborative mechanism between heterogeneous systems.

[0003] The microservice architecture divides traditional monolithic tasks into several independent microservices according to the implemented functions. Each microservice module focuses on implementing a specific function. When deploying a new application, only the required function modules need to be combined and then deployed, effectively solving the problems of high coupling between functions in the traditional monolithic service architecture, redundant resource deployment, and difficult development and maintenance.

[0004] Although there are many solutions for task offloading under the traditional monolithic architecture, its adaptability to task offloading in the microservice architecture is insufficient. At the same time, the scheduling rules of traditional rule-based scheduling methods (such as the polling method based on load balancing) are simple, and the optimization of the scheduling time arrangement for multi-tasks is insufficient, resulting in poor accuracy of task offloading and relatively high overall latency. In addition, heuristic algorithms based on genetic algorithms or particle swarm algorithms have the problem of large computational overhead, are difficult to meet the real-time decision-making requirements in task offloading, and cannot adapt to the characteristics of dynamic arrival of SFC requests.

[0005] Therefore, how to design a method that can effectively adapt to task offloading in the microservice architecture and ensure the accuracy and real-time performance of task offloading is a technical problem that needs to be solved urgently. Summary of the Invention

[0006] Aiming at the deficiencies of the above-mentioned prior art, the technical problem to be solved by the present invention is: how to provide an instant decision-making method for task offloading in a microservice architecture, which realizes the task offloading decision in the microservice architecture based on a task offloading decision model constructed by a deep Q-network, can sense environmental changes in real time and adjust the decision-making strategy through learning to adapt to the task offloading environment in the microservice architecture, and at the same time can improve the accuracy and real-time performance of task offloading.

[0007] To solve the above technical problem, the present invention adopts the following technical solutions:

[0008] An instant decision-making method for task offloading in a microservice architecture, comprising:

[0009] S1: Obtain a task group composed of SFC requests; an SFC request is composed of multiple microservices combined in a chain manner;

[0010] S2: Sort the microservices of all SFC requests within the task group according to the microservice execution dependencies to form a task queue;

[0011] S3: Build a task offloading decision model based on the deep Q-network and train the task offloading decision model;

[0012] S4: Generate task offloading decisions for the microservices in the task queue through the trained task offloading decision model; the task offloading decisions include the edge servers and computing units assigned to the microservices;

[0013] S5: Execute the processing of the corresponding microservices based on the task offloading decisions.

[0014] Preferably, in step S1, classify the SFC requests arriving within a continuous number of time slots into one task group.

[0015] Preferably, in step S2, sort according to the earliest start time of all unscheduled microservices in the task group to form a task queue.

[0016] Preferably, in step S3, the task offloading decision model built based on the deep Q-network includes a state space, an action space, and a reward function;

[0017] 1) State space

[0018] The formula representation of the state space is:

[0019]

[0020] In the formula: o t represents the action at time slot t; represents the microservices to be scheduled and the task information included in the microservices to be scheduled in the task queue at time slot t; τ t represents the information on the execution of microservice instances in each computing unit of the edge server at time slot t;

[0021] 2) Action space

[0022] The action is the task offloading decision for the microservice σ m during the m-th task scheduling in the task queue at time slot t. The task offloading decision represents the operation of allocating the microservice to be scheduled to a computing unit in the edge server; the set of computing units of the edge server is L, and there are |L| computing units in total, so there are |L| kinds of task offloading decisions, that is

[0023] sfc i The execution time of the k-th microservice on the server is expressed as:

[0024]

[0025] In the formula: represents the computing intensity of the edge server computing unit ; c i,k is the number of CPU instructions executed per second;

[0026] 3) Reward function

[0027] The formula of the reward function is expressed as:

[0028]

[0029] In the formula: r * represents the reward; τ m-1 represents the maximum value of time in the completion progress of all computing units of the edge server before the m-th round of task scheduling; represents the microservice σ at the m-th task scheduling m whether it is offloaded to the same computing unit as σ, represents σ m is offloaded to the same computing unit as σ, represents σ m is offloaded to a computing unit different from σ; σ represents all the microservices scheduled in the computing unit that requires the longest time to the next computing unit idle among all computing units before the m-th task scheduling; represents the time elapsed from the (m - 1)-th round of scheduling to the m-th round of scheduling.

[0030] Preferably, in step S3, in the state space, τ t includes and They are jointly used to indicate the time series information of microservice instance scheduling;

[0031] τ m-1 represents the maximum value of time in the completion progress of all computing units before the m-th round of task scheduling;

[0032] τ m represents the maximum value of time from the actual start time of the scheduled microservice to the completion progress time of all computing units after the m-th round of task scheduling;

[0033] represents the earliest start time of the scheduled microservice in the m-th round of scheduling;

[0034] is the computing execution time of the scheduled microservice in the m-th round of scheduling.

[0035] Preferably, in step S3, the processing steps for training the task offloading decision model include:

[0036] S301: Initialize the parameters θ of the online learning network, the parameters of the target policy network,

[0037] and the experience replay buffer; t S302: Determine the current state o t ; Select an action a through the ε-greedy policy

[0038] S303: Execute the selected action a in the environment t , observe the new state o t ' and the reward r * ;

[0039] S304: Store the experience sample (o t , a t , r * , o t ') into the experience replay buffer;

[0040] S305: Randomly sample from the experience replay buffer: For each sample, update the parameters of the online learning network by minimizing the loss function L(θ) using gradient descent;

[0041] S306: Every certain number of time steps or training epochs, copy the parameters θ of the online learning network to update the parameters of the target network ;

[0042] S307: Repeat steps S302 to S306 to iteratively train the task offloading decision model until the model converges or reaches the maximum number of training epochs.

[0043] Preferably, in step S305, the calculation formula of the loss function is expressed as:

[0044]

[0045] Where: γ represents the discount factor of the reward; r is the reward value designed according to the experience; represents the Q-value prediction of the target policy network for the next state o t ' and possible actions a'; Q(o t , a; θ) represents the Q-value prediction of the online learning network for the state o t and action a.

[0046] Preferably, in step S305, the Q-value prediction y of the online learning network is calculated by the following formula i :

[0047]

[0048] Where: ri is the reward value designed based on experience; γ represents the discount factor of the reward; is the Q-value prediction of the target policy network for state and action a t ′; the termination state is defined as that in the current time slot t, there is no microservice to be scheduled in .

[0049] Preferably, in step S4, the processing steps for generating task offloading decisions for the microservices in the task queue include:

[0050] S401: Move to a new time slot, check whether the current time slot belongs to the current task group. If not, execute step S405; if so, check whether there is an SFC request arriving: if there is no SFC request arriving, execute step S402; if there is an SFC request arriving, determine the earliest start time of the first microservice of each newly arriving SFC request, add the microservices in the newly arriving SFC request to the task queue of the current task group, and execute step S402;

[0051] S402: Sort the microservices with determined earliest start times in descending order of the earliest start times to form a task queue; the microservice with the smallest earliest start time in the task queue is the next microservice to be scheduled;

[0052] S403: Check whether the current time slot has reached or exceeded the earliest start time of the next microservice to be scheduled in step S402: if not, execute step S401; if so, execute step S404;

[0053] S404: Use the trained task offloading decision model to schedule the microservice. After scheduling, the microservice is removed from the task queue, and calculate the earliest start time of the successor microservice of the microservice; check whether there are unscheduled microservices in the current task queue. If not, execute step S401; if so, execute step S402;

[0054] S405: Check whether there is an SFC request arriving in the current time slot. If there is no SFC request arriving, execute step S406; if there is an SFC request arriving, classify the SFC requests arriving in this time slot into the next task group, and then execute step S406;

[0055] S406: Check whether the task queue of the current task group is empty. If the task queue is not empty, execute step S402; if the task queue is empty, the immediate task offloading decision of the current task group ends.

[0056] Preferably, in step S404, the task offloading decision model schedules microservices, which means that the target policy network of the task offloading decision model generates the best action for the microservices, i.e., the task offloading decision, based on the current state.

[0057] Compared with the prior art, the task offloading instant decision method for microservice architecture in the present invention has the following beneficial effects:

[0058] By classifying SFC requests arriving within a continuous number of time slots into a task group, the present invention can manage and allocate system resources more effectively. And by centrally processing SFC requests with similar or related functional requirements, it can more reasonably plan computing resources, storage resources, and network bandwidth, avoiding fragmented use of resources, thereby improving resource utilization. At the same time, the present invention sorts the microservices within the task group according to the execution dependency relationship to form a task queue, which helps to reduce the overhead of task scheduling. And the construction of the task queue can clearly understand the execution order and dependency relationship of each microservice, enabling more rapid and accurate decision-making during task scheduling, thereby improving the response speed and processing efficiency. In addition, through the management of task groups and task queues, the present invention can better cope with sudden traffic and load changes. When the load increases, the load can be balanced by adjusting the size of the task group and the priority of the task queue to ensure that critical tasks can be processed in a timely manner, enhancing the stability and reliability of the system.

[0059] The task offloading decision model based on the deep Q-network in the present invention can learn and adapt to the complex task offloading environment of the microservice architecture to achieve intelligent decision-making. First, by continuously interacting with the environment and learning the optimal strategy, the deep Q-network can generate the optimal task offloading decision for the microservices according to the current task state, task requirements, and resource status, including the allocated edge server and computing unit, thereby improving the accuracy and efficiency of task offloading. Second, compared with traditional task offloading decision methods, the task offloading decision model based on the deep Q-network has higher decision-making efficiency, can quickly process a large number of state-action pairs, and give the optimal decision in a short time, thereby shortening the decision-making time of task offloading and improving the real-time performance of task offloading. Finally, in the microservice architecture, the task offloading environment is often dynamically changing, including network conditions, the load of edge servers, fluctuations in task requirements, etc. The task offloading decision model based on the deep Q-network can perceive these changes in real time and adapt to the complex task offloading environment of the microservice architecture by learning and adjusting the decision-making strategy, thereby improving the stability and reliability of task offloading. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] In order to make the objectives, technical solutions, and advantages of the invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings, where:

[0061] Figure 1 It is a logic block diagram of an instant decision-making method for task offloading in a microservices architecture.

[0062] Figure 2 They are possible task scheduling situations.

[0063] Figure 3 It is a schematic diagram of a single task offloading at an edge node.

[0064] Figure 4 It is a logic block diagram of a deep Q network.

[0065] Figure 5 It is a flowchart of task offloading decision-making.

[0066] Figure 6 They are task offloading results.

[0067] Figure 7 They are computing unit utilization rates. Detailed implementation manners

[0068] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Components of the embodiments of the present invention usually described and illustrated in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0069] It should be noted that similar reference numerals and letters denote similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the figures, or the orientation or positional relationship in which the product of the invention is usually placed during use. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention. In addition, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance. In addition, terms such as "horizontal" and "vertical" do not mean that the components are required to be absolutely horizontal or hanging, but can be slightly inclined. For example, "horizontal" only means that its direction is more horizontal relative to "vertical", and does not mean that the structure must be completely horizontal, but can be slightly inclined. In the description of the present invention, it should also be noted that unless otherwise clearly specified and limited, the terms "set", "install", "connect", and "couple" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0070] The following is a more detailed description through specific embodiments:

[0071] Embodiment:

[0072] This embodiment discloses a task offloading instant decision method for a microservices architecture.

[0073] As Figure 1 shown, a task offloading instant decision method for a microservices architecture includes:

[0074] S1: Obtain SFC requests to form a task group; an SFC request is composed of multiple microservices in a chained manner, and different combination methods correspond to different functions and are related to the functions actually required to be implemented;

[0075] In this embodiment, SFC requests arriving within a continuous number of time slots are classified into a task group.

[0076] S2: Sort the microservices of all SFC requests in the task group according to the microservice execution dependency relationship to form a task queue;

[0077] In this embodiment, a task queue is formed by sorting according to the earliest start times of all unscheduled microservices in the task group.

[0078] S3: Construct a task offloading decision model based on the deep Q-network and train the task offloading decision model;

[0079] S4: Generate task offloading decisions for the microservices in the task queue through the trained task offloading decision model; the task offloading decisions include the edge servers and computing units assigned to the microservices;

[0080] S5: Execute the processing of the corresponding microservices based on the task offloading decisions.

[0081] By classifying SFC requests arriving within a continuous number of time slots into one task group, the present invention can manage and allocate system resources more effectively. And by centrally processing SFC requests with similar or related functional requirements, it can more reasonably plan computing resources, storage resources, and network bandwidth, avoiding fragmented use of resources, thereby improving resource utilization. At the same time, the present invention sorts the microservices within the task group according to the execution dependency relationship to form a task queue, which helps reduce the overhead of task scheduling. And the construction of the task queue can clearly understand the execution order and dependency relationship of each microservice, enabling faster and more accurate decision-making during task scheduling, thereby improving the response speed and processing efficiency. In addition, through the management of task groups and task queues, the present invention can better cope with sudden traffic and load changes. When the load increases, the load can be balanced by adjusting the size of the task group and the priority of the task queue to ensure that critical tasks can be processed in a timely manner, enhancing the stability and reliability of the system.

[0082] The task offloading decision model based on the deep Q-network of the present invention can learn and adapt to complex microservice architecture task offloading environments to achieve intelligent decision-making. First, by continuously interacting with the environment and learning the optimal strategy, the deep Q-network can generate optimal task offloading decisions for microservices according to the current task status, task requirements, and resource status, including the assigned edge servers and computing units, thereby improving the accuracy and efficiency of task offloading. Second, compared with traditional task offloading decision methods, the task offloading decision model based on the deep Q-network has higher decision-making efficiency, can quickly process a large number of state-action pairs, and give the optimal decision in a short time, thereby shortening the decision-making time of task offloading and improving the real-time performance of task offloading. Finally, in a microservice architecture, the task offloading environment is often dynamically changing, including network conditions, the load of edge servers, fluctuations in task requirements, etc. The task offloading decision model based on the deep Q-network can perceive these changes in real time and adapt to complex microservice architecture task offloading environments by learning and adjusting decision-making strategies, thereby improving the stability and reliability of task offloading.

[0083] To better introduce the technical solution of the present invention, this embodiment will be described through the following several parts.

[0084] I. Scenario Design

[0085] In this embodiment, the targeted application scenario is as Figure 1 shown. Multiple terminals can send SFC requests to the edge server at any time. After receiving the SFC requested by the terminal, the edge server provides the corresponding computing service for the terminal and finally returns the computing result to the corresponding terminal. Each SFC is composed of multiple microservices combined in a chain-like manner, and different combination methods correspond to different functions and are related to the functions that actually need to be implemented.

[0086] This embodiment is implemented based on a discrete-time system. The basic unit of discrete time is a time slot, which is a relatively short period of time, and the task offloading algorithm is guided by the time slot. SFCs may arrive in different time slots. Define the SFCs that arrive within a continuous number of time slots (the number is a parameter that can be set) as a task group. For example, all the SFCs that arrive in time slots 1, 2, and 3 form task group 1, and all the SFCs that arrive in time slots 4, 5, and 6 form task group 2, and so on. Task offloading is carried out in units of a task group, and the scheduling of tasks in the next task group is only carried out after the tasks in the previous task group have been scheduled. Among them, task scheduling refers to the operation of allocating computing units for the unscheduled microservices within the task group on the edge server. Task scheduling is completed when the operation of arranging computing units for all the microservices that make up the SFC within the task group is completed, rather than when the microservices have been executed.

[0087] The task queue is defined as an arrangement of microservices formed after sorting the unscheduled microservices in the task group. The task queue contains the unscheduled microservices in all the SFCs that have arrived in the current and previous time slots of the current task group. The sorted microservices use a deep Q-network (DQN)-based reinforcement learning method to schedule the microservices in the task queue.

[0088] The set of discrete time slots is represented as T = {1,..., t,...}. The terminal is represented by u i where i ∈ I = {1, 2,..., |I|}, and |I| is the number of terminals. In this embodiment, it is assumed that each terminal has only one SFC request. For the case where a terminal requests multiple SFCs, it is equivalent to the requests of multiple terminals. In this way, there is a one-to-one correspondence between the terminal and the SFC request. Subsequently, the differences among the three terms of terminal, SFC request, and request will no longer be distinguished.

[0089] The SFC requested by terminal u i can be expressed as where |K i | represents the composition of sfci The number of microservices. sfc i Each microservice of is represented as where c i,k , respectively represent the size of the input data required for microservice computing, the size of the output data of microservice computing, the number of CPU cycles required to complete the task, the earliest start time of the microservice, and the indices i and k represent the k-th microservice of the i-th request.

[0090] Since task sorting is involved, the original index values i and k of the microservices will be arranged in a new order after sorting. Therefore, each microservice in the sorted task queue is called a record, and σ m represents the microservice at the m-th task scheduling in the task queue. To facilitate the representation of the conversion between the two types of indices, the function tr(i,k) is used to represent the conversion from the microservice index to the corresponding record index after sorting. For example, if the 3rd microservice in sfc2 is at the 4th record in the task queue after task sorting, then tr(2,3) = 4. Note that since the microservice at the head of the task queue will be removed from the task queue after a task scheduling is completed, σ m represents at the m-th scheduling.

[0091] II. Task Sorting

[0092] In this embodiment, task sorting refers to the sorting of all unscheduled microservices within a task group, and the sorting basis is the earliest start time of the microservices The earliest start time of the first microservice in the i-th SFC is the arrival time of the SFC request The earliest start time of the subsequent microservices is the execution end time of their predecessor microservices, that is

[0093] III. Task Offloading Decision Model

[0094] In this embodiment, task offloading is defined as the operation of determining the actual execution computing unit of the microservice to be scheduled in the task queue, which is completed by the Deep Q Network (DQN) algorithm. DQN is a type of deep reinforcement learning algorithm. The basic framework of reinforcement learning mainly consists of four key elements: agent, state, action, and reward. The agent is the proxy and executor of the entire algorithm, and only the remaining three elements need to be designed.

[0095] The task offloading decision model based on the deep Q network includes a state space, an action space, and a reward function;

[0096] 1) State space

[0097] The formula of the state space is expressed as:

[0098]

[0099] In the formula: o t represents the action at time slot t; represents the microservices to be scheduled and the task information included in the microservices to be scheduled in the task queue at time slot t; τ t represents the information of the microservice instances executed in each computing unit of the edge server at time slot t;

[0100] τ t includes and They are jointly used to indicate the time series information of microservice instance scheduling; τ m-1 represents the maximum value of time in the completion progress of all computing units before the m-th round of task scheduling, and the corresponding time slice occupation is represented by the letter σ in Figure 2 ; τ m represents the maximum value of time from the actual start time of the scheduled microservice to the completion progress time of all computing units after the m-th round of task scheduling; represents the earliest start time of the microservice to be scheduled in the m-th round of scheduling, and it is equal to the earliest start time of the microservice σ m at the m-th task scheduling where tr(i,k) = m; is the computing execution time of the microservice to be scheduled in the m-th round of scheduling, and it is equal to the execution time of the microservice σ m at the m-th task scheduling where tr(i,k) = m.

[0101] represents the elapsed time from the (m - 1)-th round of scheduling to the m-th round of scheduling;

[0102] τ elp The calculation formula of is:

[0103]

[0104] In the formula: is the earliest start time of the microservice σ m-1 at the (m - 1)-th task scheduling.

[0105] 2) Action space

[0106] The action of the agent is defined as the microservice σ at the m-th task scheduling in the task queue at time slot t mThe task offloading decision represents the operation of allocating the microservice to be scheduled to a computing unit in the edge server. A schematic diagram of a single task offloading is as shown in Figure 3 as follows.

[0107] Assume that the set of computing units in the edge server is \(L\), and there are \(|L|\) computing units in total. Then there are \(|L|\) actions for the task offloading decision, which are denoted by the symbol Another way of writing is used to emphasize that the \(m\)-th record in the task queue within time slot \(t\) is the \(k\)-th microservice in the sorted \(sfc\) i obtained after sorting;

[0108] According to the above definition of actions, the execution time of the \(k\)-th microservice in \(sfc\) i on the server is expressed as:

[0109]

[0110] In the formula: represents the computing intensity of the edge server computing unit ; \(c\) i,k is the number of CPU instructions executed per second.

[0111] 3) Reward function

[0112] The reward function is related to the scheduling process. Possible situations of scheduling are as shown in Figure 2 as follows. represents the computing execution time of the \(m\)-th round of scheduling tasks, is equal to the corresponding execution time , and it is a period of time. Correspondingly, the earliest start time is the time when the previous microservice in the SFC where \(\sigma\) m is located finishes execution. It is a moment value and can be represented on the number line, as shown in the number line annotation in Figure 2 as follows. That is, the moment when \(\sigma\) m starts task scheduling. It should be noted that the moment when task scheduling starts may not be equal to the moment when the microservice actually starts execution. If the microservices arranged in the current computing unit are backlogged, the actual start time of the microservice will be greater than the earliest start time. Therefore, the actual start time of the scheduled microservice is defined as the time when the computing unit allocated to this microservice first becomes idle.

[0113] The above variables are elements of \(\tau\) t in the state space and are used to measure the quality of the agent's decision. If the agent can make the most use of \(\tau\) m-1If the time allows for parallel execution of multiple tasks, the scheduling is relatively optimal.

[0114] According to the above design concept, the formula of the reward function is expressed as:

[0115]

[0116] In the formula: r * represents the reward; τ m-1 represents the maximum value of time among the completion progress of all computing units in the edge server before the m-th round of task scheduling; represents the microservice σ at the m-th task scheduling m whether it is offloaded to the same computing unit as σ, represents σ m being offloaded to the same computing unit as σ, represents σ m being offloaded to a computing unit different from σ; σ represents all the microservices scheduled in the computing unit that requires the longest time until the next computing unit becomes idle among all computing units before the m-th task scheduling.

[0117] Figure 2 shows possible task scheduling situations. The first row and the second row of the formula for calculating the reward function correspond to Figure 2 the (a) and (b) diagrams in Figure 2 respectively, the third row corresponds to Figure 2 the (c) and (d) diagrams in Figure 2 (a) and Figure 2 (b), the reward value is the overlapping part of the time occupied by the microservice to be scheduled with τ m-1 respectively being and σ and σ m The situation where the time slices do not overlap is as shown in Figure 2 (c) and (d). At this time, regardless of the value of , it is equivalent to a task scheduling starting from the instance idle state, and the reward is set to 0. Figure 2 (e) and (f) diagrams are for the situation. The two situations shown in the diagrams occur when there is no idle computing unit, and the reward is set to 0.

[0118] IV. Model Training

[0119] In this embodiment, a deep Q-network is used to solve the problem of dynamic task offloading of microservices under edge collaboration. The goal of this method is to maximize the long-term cumulative reward, and the cumulative reward refers to the accumulation of the above reward function. As shown in Figure 4 , the processing steps during the training of the task offloading decision model include:

[0120] S301: Initialize the parameters θ of the online learning network, the parameters of the target policy network and the experience replay buffer;

[0121] In this embodiment, the parameters θ of the online learning network are randomly initialized; initialize the parameters of the target policy network to be consistent with the parameters of the online learning network, and the structure of the target policy network is the same as that of the online learning network. Initialize the experience replay buffer to store experience samples.

[0122] S302: Determine the current state o t (i.e., the information of the microservice to be scheduled currently); select an action a through the ε-greedy policy t : with probability 1 - ε, select the optimal action evaluated by the online learning network Q(o t , a t ; θ) in the current state, or randomly select an action with probability ε to ensure exploration;

[0123] S303: Execute the selected action a in the environment t , observe the new state o t ' and the reward r * ;

[0124] S304: Store the experience sample (o t , a t , r * , o t ') into the experience replay buffer;

[0125] S305: Randomly sample from the experience replay buffer: i ∈ {1, 2, …, minibatch}; for each sample, update the parameters of the online learning network by minimizing the loss function L(θ) using gradient descent;

[0126] The calculation formula of the loss function is expressed as:

[0127]

[0128] In the formula: γ represents the discount factor of the reward, taking a value between 0 and 1. The larger γ is, the more past Q values will be referred to in the calculation; represents the expected value of the long-term behavior distribution, used to calculate the average loss over the entire data distribution; r is the reward value designed according to experience, which is a positive value and needs to be determined by trial and can be set to +1; represents the Q-value prediction of the target policy network for the next state o t ' and possible actions a'; is the parameter of the target policy network, which is periodically copied from the parameters of the online learning network; Q(o t , a; θ) is the Q-value prediction of the online learning network for the state o t and the action a;

[0129] Among them, the Q-value prediction y i , that is, the target value of the on-policy network (or the Q-value of the on-policy network), y i uses the Q-value of the target policy network during calculation. Since the parameters of the target policy network are periodically copied from the parameters of the on-policy network (the on-policy network and the target policy network have exactly the same structure, only the parameter update methods and frequencies are different), and the task offloading decision is obtained from the target policy network, the on-policy network target value y i is actually implicitly used.

[0130]

[0131] In the formula: r i is the reward value designed according to experience, which is a positive value and needs to be determined by trial and is set to +1; γ represents the discount factor of the reward, taking a value between 0 and 1, generally taking a value from 0.95 to 0.99. The larger γ is, the more past Q-values will be referred to during calculation; is the Q-value prediction of the target policy network for the state and the action a t '; represents the next state at time step t, a t ' represents the action that may be taken in the next state ; is the parameter of the target policy network, indicating that the Q-value is obtained from the target policy network; the termination state is defined as that in the current time slot t, of there are no microservices to be scheduled;

[0132] S306: Every certain number of time steps or training times, copy the parameters θ of the online learning network to the target network to update the parameters to maintain the stability of the target network. The update frequency of the target network parameters is t tgt ;

[0133] S307: Repeat steps S302 to S306 to iteratively train the task offloading decision model until the model converges or reaches the maximum number of training times.

[0134] V. Task Offloading Decision

[0135] In this embodiment, a single task scheduling only schedules one microservice, and the task scheduling is completed by the reinforcement learning agent. For exampleFigure 5 As shown in Figure 5 , the processing steps for generating task offloading decisions for microservices in the task queue include:

[0136] S401: Move to a new time slot and check if the current time slot belongs to the current task group. If it does not, execute step S405; if it does, check if there is an SFC request arriving: if no SFC request arrives, execute step S402; if an SFC request arrives, determine the earliest start time of the first microservice in each newly arrived SFC request, add the microservices in the newly arrived SFC request to the task queue of the current task group, and execute step S402;

[0137] S402: Sort the microservices with determined earliest start times in descending order of the earliest start time to form a task queue; the microservice with the smallest earliest start time in the task queue is the next microservice to be scheduled.

[0138] S403: Check if the current time slot has reached or exceeded (the exceeded situation occurs when waiting for the previous task group for all microservices in the arrived SFC to complete scheduling) the earliest start time of the next microservice to be scheduled in step S402: if it has not reached or exceeded, execute step S401; if it has reached or exceeded, execute step S404;

[0139] S404: Use the trained task offloading decision model to schedule the microservice. After scheduling, the microservice is removed from the task queue, and calculate the earliest start time of the successor microservice of this microservice (if this microservice has a successor); check if there are unscheduled microservices in the current task queue. If not, execute step S401; if there are, execute step S402;

[0140] The task offloading decision model scheduling the microservice means that the target policy network of the task offloading decision model generates the best action for the microservice, that is, the task offloading decision, based on the current state.

[0141] S405: Check if there is an SFC request arriving in the current time slot. If no SFC request arrives, execute step S406; if an SFC request arrives, classify the SFC requests arriving in this time slot into the next task group, and then execute step S406;

[0142] S406: Check if the task queue of the current task group is empty. If the task queue is not empty, execute step S402; if the task queue is empty, the immediate task offloading decision for the current task group ends.

[0143] VI. Experiment Description

[0144] To better illustrate the advantages of the technical solution of the present invention, the following experiment is disclosed in this embodiment.

[0145] This experiment was conducted when the task group was one time slot, with the time slot size being 0.25 s. In this time slot, 15 SFCs with random arrival times were generated, and each SFC consisted of 4 microservices.

[0146] 1. System model parameters

[0147]

[0148]

[0149] SFC request arrival time Unit: second (s). 15 SFC requests were designed. The values in the matrix represent the arrival times of the 15 SFCs, that is, the arrival times of the first microservice of each SFC. The arrival time matrix is expressed as:

[0150] [0.0007, 0.0009, 0.0978, 0.0133, 0.0000, 0.0004, 0.0469, 0.0717, 0.0987, 0.0200, 0.0018, 0.0016, 0.0013, 0.0805, 0.0004]

[0151] The number of computing units |L| = 8, and the computing processing intensity vectors of computing units 1 to 8 are [2, 1, 2, 2, 2, 1, 1, 3], unit: GHz.

[0152] The number of CPU computing cycles required for unit input data volume is 214 cycles / MB. The calculation formula for the number of CPU cycles required for the corresponding microservice computing is

[0153] The parameters of the reinforcement learning model are set as follows: discount factor γ = 0.995, exploration factor ∈ = 0.3, the maximum capacity of the experience replay pool is 1×104, the experience replay sampling batch size minibatch = 256, and the update period t tgt = 100.

[0154] 2. Task offloading results

[0155] The simulation was implemented on a computer equipped with an Intel Core i5-10400F, 2.9 GHz processor and 16 GB RAM, in the MATLAB 2023b environment. The offloading result matrix is shown as follows, where the rows represent 15 SFCs, the columns correspond to the four microservices in the SFC, and the corresponding values represent the computing units to which the corresponding microservices are offloaded.

[0156] [6, 8, 5, 4

[0157] 3,4,3,7

[0158] 4,4,4,8

[0159] 1,8,1,8

[0160] 8,4,3,8

[0161] 1,3,7,6

[0162] 2,8,8,8

[0163] 5,5,1,3

[0164] 2,6,8,8

[0165] 6,3,1,1

[0166] 2,7,8,8

[0167] 7,6,3,3

[0168] 4,8,5,5

[0169] 3,6,1,3

[0170] 5,5,8,8]

[0171] The visualization result of task offloading is as Figure 6 shown. Among them, the rectangle represents the computing execution time occupied by the microservice, and "x - y" represents the y-th microservice of the x-th SFC request. The computing unit utilization rate is as Figure 7 shown. Among them, since the computing unit 8 has a high allocated computing rate, this computing unit is utilized more fully; at the same time, this method also makes full use of other computing units to optimize the overall task scheduling delay. In summary, the task offloading method proposed by the present invention can efficiently achieve task offloading under the microservice architecture.

[0172] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit them. Those of ordinary skill in the art should understand that any modifications or equivalent replacements made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions shall be covered by the scope of the claims of the present invention.

Claims

1. An immediate decision-making method for task offloading in a microservices architecture, characterized in that Including: S1: Obtain SFC requests to form a task group; an SFC request is composed of multiple microservices combined in a chained manner; S2: Sort the microservices of all SFC requests within the task group according to the microservice execution dependency relationship to form a task queue; S3: Build a task offloading decision model based on the deep Q-network and train the task offloading decision model; S4: Generate task offloading decisions for the microservices in the task queue through the trained task offloading decision model; The task offloading decision includes the edge server and computing unit assigned to the microservice; S5: Execute the processing of the corresponding microservice based on the task offloading decision.

2. The instant decision-making method for task offloading in a microservices architecture according to claim 1, wherein: In step S1, classify the SFC requests arriving within a continuous number of time slots into one task group.

3. The task offloading instant decision-making method for microservice architecture according to claim 1, wherein: In step S2, sort according to the earliest start time of all unscheduled microservices in the task group to form a task queue.

4. The task offloading instant decision-making method for a microservices architecture according to claim 3, wherein: In step S3, the task offloading decision model built based on the deep Q-network includes a state space, an action space, and a reward function; 1) State space The formula representation of the state space is: Where: o t represents the action at time slot t; represents the microservices to be scheduled and the task information contained in the microservices to be scheduled in the task queue at time slot t; τ t represents the information of the microservice instances executed in each computing unit of the edge server at time slot t; 2) Action space Action For the microservice σ during the m-th task scheduling in the task queue at time slot t m The task offloading decision, which represents the operation of allocating the microservice to be scheduled to a computing unit in the edge server; The set of edge server computing units is L, and there are |L| computing units in total. Then there are |L| kinds of actions for the task offloading decision, that is sfc i The execution time of the k-th microservice on the server is expressed as: In the formula: represents the computing intensity of the edge server computing unit ; c i,k is the number of CPU instructions executed per second; 3) Reward function The formula representation of the reward function is: Where: r * represents the reward; τ m-1 represents the maximum value of time among the completion progress of all computing units of the edge server before the m-th round of task scheduling; represents the microservice σ at the m-th task scheduling m whether it is unloaded to the same computing unit as σ, represents σ m unloaded to the same computing unit as σ, represents σ m unloaded to a computing unit different from σ; σ represents all the microservices scheduled in the computing unit that requires the longest time to the next free computing unit among all computing units before the m-th task scheduling; represents the time elapsed from the (m - 1)-th round of scheduling to the m-th round of scheduling.

5. The task offloading instant decision method for a microservices architecture according to claim 4, wherein: In step S3, in the state space, τ t includes and They are jointly used to indicate the time series information of microservice instance scheduling; τ m-1 represents the maximum value of time among the progress of all computing units before the m-th round of task scheduling; τ m represents the maximum value of the time from the actual start time of the scheduled microservice to the completion progress time of all computing units after the m-th round of task scheduling; Indicates the earliest start time of the microservice to be scheduled in the m-th round of scheduling; is the computing execution time of the microservice to be scheduled in the m-th round of scheduling.

6. The instant decision-making method for task offloading in a microservices architecture according to claim 5, wherein: In step S3, the processing steps during the training of the task offloading decision model include: S301: Initialize the parameters θ of the online learning network, the parameters of the target policy network and the experience replay buffer; S302: Determine the current state o t ; Select an action a through the ∈-greedy policy t ; S303: Execute the selected action a in the environment t , observe the new state o t ' and the reward r * ; S304: Store the experience sample (o t , a t , r * , o t ′) into the experience replay buffer; S305: Randomly sample from the experience replay buffer: For each sample, update the parameters of the online learning network by minimizing the loss function L(θ) using gradient descent; S306: At regular time steps or training iterations, copy the parameters θ of the online learning network to update the parameters of the target network S307: Repeatedly execute steps S302 to S306 to iteratively train the task offloading decision model until the model converges or reaches the maximum number of training times.

7. The task offloading instant decision method for microservice architecture according to claim 6, wherein: In step S305, the calculation formula of the loss function is: Where: γ represents the discount factor of the reward; r is the reward value designed according to experience; represents the Q-value prediction of the target policy network for the next state o t ′ and possible action a′; Q(o t , a; θ) represents the Q-value prediction of the online learning network for state o t and action a.

8. The task offloading instant decision method for microservice architecture according to claim 7, characterized in that: In step S305, the Q-value prediction y of the online learning network is calculated by the following formula i :[[]]END]] where: r i is the reward value designed based on experience; γ represents the discount factor of the reward; is the Q-value prediction of the target policy network for the state and the action a t '; the termination state is defined as that in the current time slot t, of there is no microservice to be scheduled.

9. The instant decision-making method for task offloading in a microservices architecture according to claim 1, characterized in that: In step S4, the processing steps for generating task offloading decisions for the microservices in the task queue include: S401: Move to a new time slot, check whether the current time slot belongs to the current task group. If not, execute step S405; if so, check whether there is an SFC request arriving: if there is no SFC request arriving, execute step S402; if there is an SFC request arriving, determine the earliest start time of the first microservice of each newly arrived SFC request, add the microservices in the newly arrived SFC request to the task queue of the current task group, and execute step S402; S402: Sort the microservices with the determined earliest start time in descending order of the earliest start time to form a task queue; the microservice with the smallest earliest start time in the task queue is the next microservice to be scheduled; S403: Check whether the current time slot has reached or exceeded the earliest start time of the next microservice to be scheduled in step S402: if not, execute step S401; if so, execute step S404; S404: Use the trained task offloading decision model to schedule the microservice. After scheduling, the microservice is removed from the task queue, and calculate the earliest start time of the successor microservice of the microservice; check whether there are unscheduled microservices in the current task queue. If not, execute step S401; if so, execute step S402; S405: Check whether there is an SFC request arriving in the current time slot. If there is no SFC request arriving, execute step S406; if there is an SFC request arriving, classify the SFC requests arriving in this time slot into the next task group, and then execute step S406; S406: Check whether the task queue of the current task group is empty. If the task queue is not empty, execute step S402; if the task queue is empty, the task offloading immediate decision of the current task group ends.

10. The task offloading instant decision method for microservice architecture according to claim 9, characterized in that: In step S404, the task offloading decision model schedules microservices, which means that the target policy network of the task offloading decision model generates the best action for the microservices based on the current state, that is, the task offloading decision.