Multi-user task unloading scheduling and resource allocation method and system

Through the collaborative optimization of DQN and TD3 networks, the problem of insufficient consideration of inter-task dependencies in multi-user edge computing task processing is solved, more efficient task offload decisions and resource allocation are achieved, and industrial production efficiency and resource utilization efficiency are improved.

CN120011014APending Publication Date: 2025-05-16ANQING NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510081933.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing research on edge computing task processing has failed to fully consider the dependencies between tasks in multi-user scenarios, and there are shortcomings in the coordinated optimization of task offload decisions and resource allocation.

Method used

Through the collaborative optimization of DQN and TD3 networks, multi-user task offloading scheduling and resource allocation are realized. The specific steps include obtaining the calculation tasks of industrial equipment, decomposing them into multiple subtasks, and using Markov decision-making process model and reinforcement learning algorithm to perform binary offloading decisions and bandwidth computing resources on DAG node tasks.

Benefits of technology

It significantly improves the success rate of task offloading, meets the strict requirements for real-time performance of industrial task processing, improves industrial production efficiency, avoids resource waste, improves resource utilization efficiency, and is suitable for dynamically changing industrial environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011014A_ABST
    Figure CN120011014A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-user task unloading scheduling and resource allocation method and system, and relates to the technical field of industrial internet computing unloading and resource allocation, and the method comprises the steps: obtaining a computing task of industrial equipment, decomposing the computing task of the industrial equipment into a plurality of subtasks, and representing the relation between the subtasks through a directed acyclic graph; inputting the plurality of sub-tasks into a pre-established Markov decision process model, and carrying out collaborative optimization on a DAG node task binary unloading decision and bandwidth computing resources based on a DQN of a reinforcement learning discrete action space and a TD3 network of a continuous action space; according to the task unloading scheduling and resource allocation method, the task processing efficiency and success rate in edge computing can be effectively improved, and the real-time requirement in an industrial scene is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of industrial Internet computing offloading and resource allocation, and specifically to a multi-user task offloading scheduling and resource allocation method and system. Background Art

[0002] In recent years, the vigorous development of information technologies such as 5G networks, the Internet of Things, big data, and artificial intelligence has profoundly changed the way of life and production in society, and is gradually moving towards the era of the Internet of Everything. Against this background, the Industrial Internet has also ushered in a new development situation. Industrial equipment often faces insufficient computing power when processing computationally intensive and delay-sensitive applications; although traditional cloud computing architecture can break through hardware limitations and provide complex, high-computing services to end users, with the access of massive networked devices, backbone network bandwidth, cloud center resources, and low-latency computing requirements have brought serious challenges to cloud computing.

[0003] In order to resolve the many defects faced by local computing and cloud computing, Mobile Edge Computing (MEC) was born, and its research mainly includes offloading decisions and resource allocation. Offloading decisions are to determine the location of task execution, and offload to the local, edge or cloud; resource allocation aims to reasonably allocate bandwidth and computing resources according to task needs to maximize benefits. MEC greatly reduces the delay caused by data transmission by "decentralizing" servers with computing resources and storage functions to the vicinity of end users. Especially in 5G networks, there is almost no difference in delay between local computing and edge computing. At the same time, it can reduce the battery power consumption caused by local computing, so that users can get high bandwidth and ultra-low latency experience quality.

[0004] However, existing research on edge computing task processing has certain limitations. For example, most DAG (directed acyclic graph) task modeling only considers single-user scenarios, and the dependencies between tasks are not fully considered in multi-user task offloading scheduling. There is also room for further improvement in the coordinated optimization of task offloading decisions and resource allocation. Summary of the invention

[0005] In order to solve the deficiencies mentioned in the above background technology, the purpose of the present invention is to provide a multi-user task offloading scheduling and resource allocation method and system, which can realize accurate task offloading decision and reasonable resource allocation through the collaborative optimization of DQN and TD3 network, and improve the real-time performance, success rate and resource utilization efficiency of industrial task processing.

[0006] In a first aspect, the purpose of the present invention can be achieved by the following technical solution: a multi-user task offloading scheduling and resource allocation method, the method comprising the following steps:

[0007] Obtaining a computing task of the industrial equipment, and decomposing the computing task of the industrial equipment into a plurality of subtasks, wherein the relationship between the subtasks is represented by a directed acyclic graph;

[0008] Multiple subtasks are input into the pre-established Markov decision process model, and the DAG node task binary offloading decision and bandwidth computing resources are collaboratively optimized based on the DQN network in the discrete action space of reinforcement learning and the TD3 network in the continuous action space.

[0009] In combination with the first aspect, in certain implementations of the first aspect, the method further includes: the industrial equipment i={1,2,3,...,n} provides computing services through an edge server, discretizing continuous time into time slots t∈(0,1,2...,T) of equal length, where the length is represented by ▽t.

[0010] In combination with the first aspect, in some implementations of the first aspect, the method further includes: the process of decomposing the computing task of the industrial equipment into a plurality of subtasks:

[0011] Defined as a triple:

[0012] V ij (t) = {d ij (t),z ij (t),λ ij (t)}, i=1,2,....,n,j=0,1,2,....,m,t=0,1,2....,T

[0013] where d ij (t) indicates that time slot t device i generates task V i The data size of the jth subtask in (t); z ij (t) indicates that time slot t device i generates task V i The computing resources required for the jth subtask in (t); ij (t)∈{0,1} represents the task V ij (t) is offloaded to edge nodes for processing or local processing;

[0014] d ij (t) and z ij (t) is in linear proportion, which can be expressed as:

[0015] z ij (t) = cd ij (t)

[0016] c represents the computing resources required per unit data volume.

[0017] In combination with the first aspect, in some implementations of the first aspect, the method further includes: when λ ijWhen (t) = 1, it means task V ij (t) local processing;

[0018] The processing delay of the computing task of the local device is:

[0019]

[0020] where f i local Indicates the processing capability of the local device CPU;

[0021] The amount of data released during local processing is:

[0022]

[0023] When the task V ij (t) After After time slots, the remaining task size Indicates that the task has been completed and the subsequent task becomes ready;

[0024] Mission V ij (t) The end time calculated locally is:

[0025]

[0026] in For Task V ij (t) Calculate the queue waiting time locally.

[0027] In combination with the first aspect, in some implementations of the first aspect, the method further includes: when λ ij (t) = 0, indicating that task V ij (t) is offloaded to edge nodes for processing;

[0028] According to Shannon's theorem, the data transmission rate between device i and edge node in time slot t′ is:

[0029]

[0030] Where B is the total channel bandwidth resource, p i represents the transmission power of device i, σ 2 represents the Gaussian white noise power, g is the channel gain between the device and the edge server, l is the straight-line distance between the device and the edge server, α is the loss factor, and w i (t′) is the proportion of bandwidth resources allocated to device i in time slot t′.

[0031] In combination with the first aspect, in some implementations of the first aspect, the method further includes: task V ij (t) When transmitted to the processing queue of the edge node via the wireless link:

[0032] Processing tasks V in the queue ij (t) The amount of data released in time slot t′ is;

[0033]

[0034] where f edge Indicates the processing power of the edge server CPU. The proportion of computing resources allocated to device i for time slot t′;

[0035] Mission V ij (t) The end time at the edge node is:

[0036]

[0037] Among them is the edge node queuing time, It is the time for edge nodes to process tasks.

[0038] In combination with the first aspect, in some implementations of the first aspect, the method further includes: the process of inputting the multiple subtasks into a pre-established Markov decision process model:

[0039] Define MDP as a triple, where the triple includes state, action, and reward;

[0040] Analyze the current real-time task size and system queue information by interacting with the device and the current environment;

[0041] Execute different actions to achieve state transfer and obtain instantaneous rewards, maximize long-term cumulative rewards and obtain the optimal strategy.

[0042] In combination with the first aspect, in some implementations of the first aspect, the method further includes: a process in which the DQN based on the discrete action space of reinforcement learning and the TD3 network in the continuous action space collaboratively optimize the binary offloading decision of the DAG node task and the bandwidth computing resources:

[0043] Use DQN network to optimize the binary offloading decision of each node task in DAG;

[0044] Adopt TD3 network to optimize bandwidth and computing resource allocation ratio;

[0045] By regularly synchronizing network parameters and memory data to train the DQN and TD3 networks, we can ultimately obtain the optimal task offloading and resource allocation strategy.

[0046] In a second aspect, in order to achieve the above-mentioned purpose, the present invention discloses a multi-user task offloading scheduling and resource allocation system, comprising:

[0047] A task processing module, used for obtaining a computing task of the industrial equipment, and decomposing the computing task of the industrial equipment into a plurality of subtasks, wherein the relationship between the subtasks is represented by a directed acyclic graph;

[0048] The collaborative optimization module is used to input multiple subtasks into the pre-established Markov decision process model, and collaboratively optimize the DAG node task binary offloading decision and bandwidth computing resources based on the DQN network in the discrete action space of reinforcement learning and the TD3 network in the continuous action space.

[0049] In another aspect of the present invention, in order to achieve the above-mentioned purpose, a terminal device is disclosed, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor, wherein the memory stores a computer program capable of running on the processor, and when the processor loads and executes the computer program, a multi-user task unloading scheduling and resource allocation method as described above is adopted.

[0050] Beneficial effects of the present invention:

[0051] The present invention significantly improves the success rate of task offloading through the collaborative optimization of DQN and TD3 networks, can better meet the strict real-time requirements of industrial task processing, and effectively improve industrial production efficiency; it performs fine modeling of multi-user DAG tasks, deeply analyzes the internal dependencies of tasks, and achieves more accurate task offloading decisions, avoids resource waste, and improves resource utilization efficiency; compared with traditional methods, the present invention can quickly generate offloading strategies with short reasoning time, is suitable for resource allocation in dynamically changing industrial environments, enhances the adaptability and stability of industrial systems, and promotes the development of industrial intelligence. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0053] Figure 1 It is a schematic flow chart of the method of the present invention;

[0054] Figure 2 It is a schematic diagram of the structure of the multi-user task offloading scheduling and resource allocation system model of the present invention;

[0055] Figure 3 It is a logical framework diagram of unloading scheduling and resource allocation of the present invention;

[0056] Figure 4 It is a comparison of the convergence performance under different learning rates of the present invention;

[0057] Figure 5 It is a comparison of the effects of the present invention and other solutions;

[0058] Figure 6 is the relationship between the arrival rate of different tasks and the unloading success rate of the present invention;

[0059] Figure 7 is the inference time of each time slot of the present invention;

[0060] Figure 8 It is a schematic diagram of the system structure of the present invention. DETAILED DESCRIPTION

[0061] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0062] Embodiment 1:

[0063] like Figure 1 As shown, a multi-user task offloading scheduling and resource allocation method comprises the following steps:

[0064] S101: Obtain a computing task of an industrial device, and decompose the computing task of the industrial device into a plurality of subtasks, wherein the relationship between the subtasks is represented by a directed acyclic graph;

[0065] Specifically, a system architecture consisting of a single edge server, a single industrial base station, and multiple industrial devices is constructed, including:

[0066] like Figure 2 As shown, the edge server can provide computing services for industrial devices i = {1, 2, 3, ..., n};

[0067] Discretize continuous time into time slots of equal length t∈(0,1,2...,T), the length of which is represented by ▽t;

[0068] Assuming that each industrial device has a new task arriving with a certain probability every 10▽t, its computing task can be processed locally or offloaded to the edge node for processing.

[0069] The industrial equipment i={1,2,3,...,n} provides computing services through the edge server, discretizing the continuous time into time slots t∈(0,1,2...,T) of equal length, with the length represented by ▽t.

[0070] like Figure 3 As shown, a mobile application task V i (t) can be decomposed into multiple subtasks, defined as triples:

[0071] V ij (t) = {d ij (t),z ij (t),λ ij (t)}, i=1,2,....,n,j=0,1,2,....,m,t=0,1,2....,T

[0072] where d ij (t) indicates that time slot t device i generates task V i The data size of the jth subtask in (t); z ij (t) indicates that time slot t device i generates task V i The computing resources required for the jth subtask in (t). ij (t)∈{0,1} represents the task V ij (t) is offloaded to edge nodes for processing or local processing;

[0073] d ij (t) and z ij (t) is in linear proportion, which can be expressed as:

[0074] z ij (t) = cd ij (t)

[0075] c represents the computing resources required per unit data volume;

[0076] When ij (t) = 1, indicating that task V ij (t) Local processing, each device has a queue to be processed and process queues

[0077] First V ij (t) is placed in the local waiting queue

[0078] The task scheduler processes the task based on the current local queue. Whether the remaining task size is met

[0079] From the local pending queue In the order in which tasks are submitted, tasks in the ready state (i.e. tasks whose predecessor tasks have been completed) are searched and placed in the local processing queue.

[0080] The processing delay of the computing task of the local device is:

[0081]

[0082] where f i local Indicates the processing capability of the local device CPU;

[0083] Local processing time one time slot local processing queue The amount of released data is:

[0084]

[0085] When the task V ij (t) After After time slots, the remaining task size Indicates that the task has been completed and the subsequent task becomes ready;

[0086] Mission V ij (t) The end time calculated locally is:

[0087]

[0088] in For Task V ij (t) Calculate the queue waiting time locally.

[0089] When ij (t) = 0, indicating that task V ij (t) is offloaded to edge nodes for processing;

[0090] First, the data is transferred to the transmission queue through the wireless network link.

[0091] Task Scheduler detects the transfer processing queue Whether the remaining task size R is satisfied i tran ≤0;

[0092] From the transfer pending queue Search for tasks in the ready state in the order in which they are submitted and put them into the transmission processing queue

[0093] Consider a wireless communication model where the device and the base station use OFDMA for communication, that is, the subchannels are orthogonal to each other and there is no interference. According to Shannon's theorem, the data transmission rate between device i and the edge node in time slot t′ is:

[0094]

[0095] Where B is the total channel bandwidth resource, p irepresents the transmission power of device i, σ 2 represents the Gaussian white noise power, g is the channel gain between the device and the edge server, l is the straight-line distance between the device and the edge server, α is the loss factor, and w i (t′) is the proportion of bandwidth resources allocated to device i in time slot t′.

[0096] When the task V ij (t) The queue to be processed is transmitted to the edge node via the wireless link hour;

[0097] The task scheduler will first detect the processing queue Whether the remaining task size is met

[0098] The pending queue The tasks are placed in the processing queue in FIFO mode. implement;

[0099] Each time slot edge server can process up to n tasks simultaneously;

[0100] The computing resources allocated to device i in time slot t′ are Processing queue Medium Mission V ij (t) The amount of data released in time slot t′ is:

[0101]

[0102] where f edge Indicates the processing power of the edge server CPU. The proportion of computing resources allocated to device i for time slot t′;

[0103] After y time slots Medium Mission V ij (t) Remaining task size Represents task V ij (t) Execution completed;

[0104] When processing queue Medium Mission V ij (t) After the calculation is completed, it means that the entire task V i (t) Completed or task V ij (t)'s subsequent driving task is in the ready state;

[0105] Mission V ij (t) The end time at the edge node is:

[0106]

[0107] Among them is the edge node queuing time, It is the time for edge nodes to process tasks.

[0108] S102: Input multiple subtasks into a pre-established Markov decision process model, and collaboratively optimize the DAG node task binary offloading decision and bandwidth computing resources based on the DQN network of the discrete action space of reinforcement learning and the TD3 network of the continuous action space.

[0109] Define MDP as a triple (state, action, reward);

[0110] By interacting with the current environment through the device, the current real-time task size, system queue information and other states are analyzed, different actions are performed to achieve state transfer and obtain instant rewards, and the long-term cumulative rewards are maximized to obtain the optimal strategy;

[0111] The specific expression of the status is as follows:

[0112] The input of the DQN network state is the queue of all local devices to be processed in time slot t Number of tasks queued

[0113] The size d of the task generated in time slot t ij (t);

[0114] The TD3 network state input for allocating bandwidth resources is the time slot t transmission queue Number of tasks queued

[0115] Time slot t transmission processing queue The size of the remaining task data in

[0116] The TD3 network input state for allocating computing resources is the current edge node waiting queue Number of tasks queued

[0117] Time slot t edge node processing queue The size of the remaining task data in

[0118] The specific expression of the action is as follows:

[0119] Task V i The four subtask offloading decisions of (t) are replaced by four-bit binary numbers, which are converted into decimal numbers as the output actions of the DQN network;

[0120] A mission V i (t)Total 24 Unloading decision, the output action set of the DQN network is:

[0121] Z∈{0,1,...,14,15}

[0122] The output layer of the TD3 network that allocates bandwidth resources outputs the bandwidth resource ratio after the activation function:

[0123] [w1(t),w2(t),....,w n-1 (t),w n (t)]

[0124] The TD3 network output computing resource ratio for allocating computing resources is:

[0125]

[0126] The specific expression of the reward is as follows:

[0127] Since the common goal of all networks is to maximize the number of successfully unloaded tasks within T time slots, the reward function of all networks is defined as the same, that is, the instantaneous reward for each time slot is the calculation of the completed task V i3 (t) if the task V is not completed in this time slot. i3 (t), the reward is 0.

[0128] The Eval_net estimation network in the DQN algorithm is used to output the estimated value of the current state s-action a pair:

[0129] Eval_Q=Sum(a*μ θ (s′)

[0130] The Target_net target network in the DQN algorithm is used to output the next state s' value;

[0131] The actual value of the current state s is:

[0132]

[0133] Use the error between the actual value and the estimated value to correct the estimation network model, and define the loss function as: loss = MES (Eval_Q, Target_Q)

[0134] Continuously optimize the parameters in the estimation network Eval_net through the back-propagation mechanism;

[0135] Regularly synchronize the Eval_net network parameters to the Target_net network.

[0136] The TD3 algorithm uses the Actor_net strategy network to output bandwidth and computing resource ratio values, and the Critic_net network to output Q values.

[0137] The specific training process of DQN and TD3 algorithms is as follows:

[0138] Input the current state s into the policy network Actor_net to get the output μ θ (s);

[0139] Combine the current state s and the output μ θ (s) is used as the network input of Critic_net0 and Critic_net1 to get the output of two Eval_Q values:

[0140]

[0141]

[0142] Similarly, the next state s′ and the output μ of Target Actor_net θ (s′) is jointly input into the TargetCritic_net0 and TargetCritic_net1 networks, and two Q values ​​are obtained, and the smallest Q value Q is selected min ;

[0143] Calculate Target_Q = r + γ * Q min , respectively construct two loss functions as:

[0144]

[0145]

[0146] Optimize the network parameters in Critic_net0 and Critic_net1 through the back-propagation mechanism;

[0147] Use the Actor_net strategy network to maximize the Q value, and the loss function is:

[0148]

[0149] The network parameters of Actor_net and Critic_net are synchronized to Target_Actor_net and Target_Critic_net by delayed soft update to complete the offloading decision and resource allocation.

[0150] In order to verify the convergence performance of the DQN+TD3 algorithm under different learning rates, several comparative tests were conducted. The learning rate of the DQN network was fixed to 0.0001, and the influence of the Actor_net learning rate act_lr and the Critic_net learning rate cri_lr in the TD3 network on the algorithm convergence was tested. Figure 4 It can be seen that when act_lr=1e-5 and cri_lr=1e-5, the curve converges better. Usually the setting value of cri_lr is larger than act_lr, so the learning rate is set to act_lr=1e-5 and cri_lr=1e-4.

[0151] In order to verify the superiority of the DQN+TD3 algorithm, we compared it with the other two algorithms such as Figure 5 As shown in the figure. Only_DQN algorithm: Use the DQN network to output the offloading decision and evenly distribute the bandwidth resources and computing resources; DQN+DDPG algorithm: Use the DQN network to output the offloading decision and use the DDPG network to allocate bandwidth resources and computing resources. In order to ensure the fairness of the experiment, all experimental tasks are generated under the same batch of random seeds. It can be seen from the figure that the DQN+TD3 algorithm is close to convergence at 600 Episodes, and can obtain more rewards, that is, a higher number of successful offloading, compared with the other two algorithms.

[0152] In order to test the relationship between different task arrival rates (0.5-0.9) and task processing success rates and real-time performance, in addition to the DQN+TD3 algorithm used in this invention, three benchmark methods are given for comparison:

[0153] Benchmark 1: Randomly assign offloading decisions and evenly distribute bandwidth resources and computing resources;

[0154] Benchmark 2: All Tasks V ij (t) All local calculations, i.e., λ ij (t) = 1;

[0155] Benchmark 3: All Tasks V ij (t) All edge nodes are calculated, that is, λ ij (t)=0, and the bandwidth resources and computing resources are evenly divided.

[0156] like Figure 6 As shown in the figure, compared with the all-local and all-edge offloading strategies, the DQN+TD3 algorithm greatly increases the offloading success rate; in terms of resource allocation, the DQN+TD3 algorithm can also improve the offloading success rate compared with evenly dividing bandwidth and computing resources.

[0157] like Figure 7As shown in the figure, the time consumption of the strategy generated by DQN+TD3 network reasoning in each time slot is maintained at about 0.4ms, which can quickly generate the offloading strategy, and the time consumption is very short to meet the real-time requirements of the task.

[0158] Embodiment 2: In the second aspect, as Figure 8 As shown, in order to achieve the above-mentioned purpose, the present invention discloses a multi-user task offloading scheduling and resource allocation system, comprising:

[0159] The task processing module 11 is used to obtain the computing task of the industrial equipment and decompose the computing task of the industrial equipment into a plurality of subtasks, wherein the relationship between the subtasks is represented by a directed acyclic graph;

[0160] The collaborative optimization module 12 is used to input multiple subtasks into a pre-established Markov decision process model, and collaboratively optimize the DAG node task binary offloading decision and bandwidth computing resources based on the DQN network of the discrete action space of reinforcement learning and the TD3 network of the continuous action space.

[0161] Based on the same inventive concept, the present invention also provides a computer device, which includes: one or more processors, and a memory for storing one or more computer programs; the program includes program instructions, and the processor is used to execute the program instructions stored in the memory. The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, which is used to implement one or more instructions, specifically for loading and executing one or more instructions in a computer storage medium to implement the above method.

[0162] It needs to be further explained that, based on the same inventive concept, the present invention also provides a computer storage medium, on which a computer program is stored, and the computer program is executed by a processor to execute the above method. The storage medium can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electrical, magnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples (non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, a computer-readable storage medium can be any tangible medium containing or storing a program, which can be used by an instruction execution system, device or device or used in combination with it.

[0163] In the description of this specification, the description with reference to the terms "one embodiment", "example", "specific example", etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0164] The above shows and describes the basic principles, main features and advantages of the present disclosure. Those skilled in the art should understand that the present disclosure is not limited by the above embodiments, and the above embodiments and descriptions are only for explaining the principles of the present disclosure. Without departing from the spirit and scope of the present disclosure, the present disclosure may have various changes and improvements, and these changes and improvements fall within the scope of the present disclosure to be protected.

Claims

1. A multi-user task offloading scheduling and resource allocation method, characterized in that: The method comprises the following steps: Obtaining a computing task of the industrial equipment, and decomposing the computing task of the industrial equipment into a plurality of subtasks, wherein the relationship between the subtasks is represented by a directed acyclic graph; Multiple subtasks are input into the pre-established Markov decision process model, and the DAG node task binary offloading decision and bandwidth computing resources are collaboratively optimized based on the DQN network in the discrete action space of reinforcement learning and the TD3 network in the continuous action space.

2. A multi-user task offloading scheduling and resource allocation method according to claim 1, characterized in that: The industrial equipment i = {1, 2, 3, ..., n} provides computing services through edge servers, discretizing continuous time into time slots t∈(0, 1, 2 ..., T) of equal length, with lengths expressed as express.

3. A multi-user task offloading scheduling and resource allocation method according to claim 1, characterized in that: The process of decomposing the computing task of industrial equipment into multiple subtasks: Defined as a triple: V ij (t)={d ij (t),z ij (t),λ ij (t)},i=1,2,....,n,j=0,1,2,....,m,t=0,1,2...,T where d ij (t) indicates that time slot t device i generates task V i The data size of the jth subtask in (t); z ij (t) indicates that time slot t device i generates task V i The computing resources required for the jth subtask in (t); ij (t)∈{0,1} represents the task V ij (t) is offloaded to edge nodes for processing or local processing; d ij (t) and z ij (t) is in linear proportion, which can be expressed as: With ij (t)=continued ij (t) c represents the computing resources required per unit data volume.

4. A multi-user task offloading scheduling and resource allocation method according to claim 3, characterized in that: When ij When (t) = 1, it means task V ij (t) local processing; The processing delay of the computing task of the local device is: where f i local Indicates the processing capability of the local device CPU; The amount of data released during local processing is: When the task V ij (t) After After time slots, the remaining task size Indicates that the task has been completed and the subsequent task becomes ready; Mission V ij (t) The end time calculated locally is: in For Task V ij (t) Calculate the queue waiting time locally.

5. A multi-user task offloading scheduling and resource allocation method according to claim 4, characterized in that: When ij (t) = 0, indicating that task V ij (t) is offloaded to edge nodes for processing; According to Shannon's theorem, the data transmission rate between device i and edge node in time slot t′ is: Where B is the total channel bandwidth resource, p i represents the transmission power of device i, σ 2 represents the Gaussian white noise power, g is the channel gain between the device and the edge server, l is the straight-line distance between the device and the edge server, α is the loss factor, and w i (t′) is the proportion of bandwidth resources allocated to device i in time slot t′.

6. A multi-user task offloading scheduling and resource allocation method according to claim 5, characterized in that: The task V ij (t) When transmitted to the processing queue of the edge node via the wireless link: Processing tasks V in the queue ij (t) The amount of data released in time slot t′ is; where f edge Indicates the processing power of the edge server CPU. Calculate the proportion of resources allocated to device i for time slot t′; Mission V ij (t) The end time at the edge node is: Among them is the edge node queuing time, It is the time for edge nodes to process tasks.

7. A multi-user task offloading scheduling and resource allocation method according to claim 1, characterized in that: The process of inputting multiple subtasks into a pre-established Markov decision process model: Define MDP as a triple, where the triple includes state, action, and reward; Analyze the current real-time task size and system queue information by interacting with the device and the current environment; Execute different actions to achieve state transfer and obtain instantaneous rewards, maximize long-term cumulative rewards and obtain the optimal strategy.

8. A multi-user task offloading scheduling and resource allocation method according to claim 1, characterized in that: The process of collaboratively optimizing the DAG node task binary offloading decision and bandwidth computing resources based on the DQN in discrete action space of reinforcement learning and the TD3 network in continuous action space: Use DQN network to optimize the binary offloading decision of each node task in DAG; Adopt TD3 network to optimize bandwidth and computing resource allocation ratio; By regularly synchronizing network parameters and memory data to train the DQN and TD3 networks, we can ultimately obtain the optimal task offloading and resource allocation strategy.

9. A multi-user task offloading scheduling and resource allocation system, characterized in that: include: A task processing module, used for obtaining a computing task of the industrial equipment, and decomposing the computing task of the industrial equipment into a plurality of subtasks, wherein the relationship between the subtasks is represented by a directed acyclic graph; The collaborative optimization module is used to input multiple subtasks into the pre-established Markov decision process model, and collaboratively optimize the DAG node task binary offloading decision and bandwidth computing resources based on the DQN network in the discrete action space of reinforcement learning and the TD3 network in the continuous action space.

10. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that: The memory stores a computer program that can be run on the processor. When the processor loads and executes the computer program, a multi-user task offloading scheduling and resource allocation method according to any one of claims 1 to 8 is adopted.