Task scheduling method and device based on distributed large language model, terminal equipment and storage medium
By adopting a task scheduling method based on a distributed large language model in a multi-layer network and using an integer linear programming algorithm for model segmentation and task scheduling, the problems of low efficiency and large latency in multi-layer networks and large language models in the existing technology are solved, and more efficient resource management and performance improvement are achieved.
Patent Information
- Application Number
- CN202510256770.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-03-05
AI Technical Summary
The prior art fails to effectively optimize task scheduling in multi-layer networks and large language models (LLM), resulting in low system efficiency and large latency.
The task scheduling method based on the distributed large language model is adopted, and the communication data and real-time monitoring data of each device node of the multi-layer network are obtained, and the model segmentation and task scheduling are used using an integer linear planning algorithm. The goal is to minimize the total delay of the multi-layer network.
Resource management of multi-layer networks within large language models is realized, network performance is improved, and latency is reduced.
Smart Images

Figure CN120179359A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of model task scheduling, and in particular to a task scheduling method, device, terminal device and storage medium based on a distributed large language model. Background Art
[0002] In the rapidly developing field of multi-tier cloud-edge computing, how to perform efficient task scheduling is crucial for optimizing system throughput and resource utilization. The multi-tier network architecture aims to meet the requirements of modern artificial intelligence applications, especially large language models (LLMs), by distributing computing tasks across the cloud, intermediate nodes, and the edge layer.
[0003] In the field of edge computing, graph neural networks (GNNs) have been used to optimize task offloading strategies and improve resource utilization. However, existing GNN methods do not optimize task scheduling for the multi-tier network and the segmented optimization results of LLMs, resulting in low efficiency and large latency when the system processes tasks. Summary of the Invention
[0004] The present invention provides a task scheduling method, device, terminal device and storage medium based on a distributed large language model, which can achieve resource management of the multi-tier network inside the large language model, improve the performance of the network, and reduce latency.
[0005] An embodiment of the present invention provides a task scheduling method based on a distributed large language model, including:
[0006] Obtaining communication data, tasks to be scheduled, task characteristics of tasks to be scheduled, and real-time monitoring data of each device node in the multi-tier network of the distributed large language model; wherein, the above task characteristics include: task computing power, network conditions, task size, and computing power requirements; the above real-time monitoring data includes: computing power, memory, bandwidth, and latency;
[0007] According to the above communication data, tasks to be scheduled, task characteristics of tasks to be scheduled, and real-time monitoring data of each device node, the multi-tier network is segmented and task-scheduled with the goal of minimizing the total latency of the multi-tier network, obtaining target device nodes corresponding to each task to be scheduled, and allocating each task to be scheduled to the corresponding target device node for scheduling;
[0008] Among them, the multi-tier network is segmented according to the above communication data and the branch-and-bound integer linear programming algorithm to obtain the allocation scheme of each layer; the above allocation scheme includes: model size and computing requirements;
[0009] All tasks to be scheduled are assigned to each layer according to the above allocation scheme, and the target tasks to be scheduled corresponding to each layer are obtained;
[0010] According to the above target tasks to be scheduled, the target task characteristics corresponding to the above target tasks to be scheduled, the above real-time monitoring data, and a preset task scheduling model, task scheduling is performed with the goal of minimizing the sum of the scheduling cost and the violation of the computing power constraint, and the target device nodes corresponding to each task to be scheduled are obtained; among them, the sum of the scheduling cost and the violation of the computing power constraint is calculated according to the above real-time monitoring data.
[0011] Further, when performing model segmentation and task scheduling on the above multi-layer network with the goal of minimizing the total delay of the multi-layer network, the corresponding first objective function is:
[0012]
[0013] R u =B k log2(1 + SINR us )
[0014]
[0015] In the formula, T total represents the total delay, L represents the total number of layers, l represents the l-th layer, represents the transmission delay of the l-th layer, represents the computing delay of the l-th layer, u l represents the device node of the l-th layer, U l represents the total number of device nodes of the l-th layer, u l′ represents the device node of the l ′ -th layer, U l′ represents the total number of device nodes of the l ′ -th layer, represents whether the scheduling task is scheduled from the l-th layer to u l′ -th layer, β l represents the computing demand allocation coefficient of the l-th layer, M c represents the computing demand, C l represents the computing power of the device node in the l-th layer, γ l represents the model size allocation coefficient of the l-th layer, M s represents the model size, R u represents the achievable transmission rate of node u, B k represents the bandwidth of sub-channel k, SINR us represents the transmission signal interference noise ratio, represents the interference plus noise power, represents other users g in the same sub-channel set to the l′ Hierarchical node u l ′ Interference Indicates the transmission signal-to-interference-plus-noise ratio between layer l’ and layer l Indicates from the l-th layer node u l To the l’-th layer node u l′ Channel gain Indicates the transmission power of the l-th layer node u l To node u l′ Transmission power
[0016] Furthermore, based on the above communication data and the branch-and-bound integer linear programming algorithm, the allocation schemes for each layer are obtained, including:
[0017] Initialize the priority queue; among them, the above priority queue includes several nodes used to represent the above allocation schemes; the initialized priority queue includes: the root node without model segmentation;
[0018] Repeat the model segmentation operation until the priority queue is empty, and take the optimal allocation scheme corresponding to the empty priority queue as the above allocation scheme;
[0019] Among them, the above model segmentation operation includes:
[0020] Obtain the current priority queue, and arbitrarily select a node from the current priority queue as the target node; among them, the initial node is the above root node;
[0021] If the current target node is a leaf node, calculate the total delay corresponding to the current target node according to the communication data corresponding to the current target node; among them, the above leaf node represents the allocation scheme corresponding to the above multi-layer network;
[0022] When the current total delay is less than the minimum total delay corresponding to the current optimal solution, update the minimum total delay according to the current total delay, take the allocation scheme corresponding to the current target node as the current optimal allocation scheme, and delete the current target node from the current priority queue; otherwise, delete the current target node from the current priority queue; among them, the initial optimal solution is infinity, and the initial optimal allocation scheme is empty;
[0023] If the current target node is not a leaf node, select an undetermined variable in the current target node for branching to obtain several child nodes; among them, the above variables include: model size and computing requirements;
[0024] Calculate the total delay of each child node according to the communication data of the current child node. If the current maximum total delay of the child nodes is less than the minimum total delay corresponding to the current optimal solution, or the current minimum total delay of the child nodes is greater than the maximum total delay corresponding to the current optimal solution, then delete the current target node and the child nodes corresponding to the current target node from the current priority queue; otherwise, add the current child node as a node of the priority queue to the current priority queue, and at the same time delete the current target node to update the current priority queue.
[0025] Further, according to the above target tasks to be scheduled, the target task characteristics corresponding to the above target tasks to be scheduled, the above real-time monitoring data, and a preset task scheduling model, with the goal of minimizing the sum of the scheduling cost and the violation of the computing power constraint, obtain the target device nodes corresponding to each task to be scheduled, including:
[0026] Initialize the list of tasks to be scheduled; among them, the initialized list of tasks to be scheduled includes all the above target tasks to be scheduled;
[0027] Repeat the task scheduling operation until the above list of tasks to be scheduled is empty, and determine the target device nodes corresponding to each task to be scheduled;
[0028] Among them, the above task scheduling operation includes:
[0029] Select a device node as the current node to be scheduled, and obtain the current task to be scheduled from the current list of tasks to be scheduled;
[0030] Construct the current structure diagram according to the current task to be scheduled and the corresponding current target task characteristics;
[0031] Obtain the current real-time monitoring data and the above preset task scheduling model, and input the current structure diagram, the current task characteristics, and the current real-time monitoring data into the above preset task scheduling model, so that the above preset task scheduling model calculates the current degree matrix and adjacency matrix according to the current structure diagram;
[0032] Obtain the current node feature matrix of the last layer in the above preset task scheduling model according to the current degree matrix, the above adjacency matrix, and the current target task characteristics;
[0033] According to the current node feature matrix and the activation function, with the goal of minimizing the sum of the scheduling cost and the violation of the computing power constraint, obtain the scheduling probability distribution between the current node to be scheduled and each current task to be scheduled;
[0034] Take the current node to be scheduled as the target device node of the task to be scheduled with the highest scheduling probability, and delete this node to be scheduled from the current list of tasks to be scheduled to update the list of tasks to be scheduled.
[0035] Further, constructing the current structure diagram according to the current task to be scheduled and the corresponding current target task characteristics includes:
[0036] Using the current task to be scheduled as the graph nodes of the current structure diagram, and connecting each graph node according to the current target task characteristics to obtain the current structure diagram.
[0037] Further, when performing task scheduling with the goal of minimizing the sum of the scheduling cost and the violation of the computing power constraint, the corresponding second objective function is:
[0038]
[0039] In the formula, H represents the second objective function, represents the task size, represents whether the task to be scheduled is scheduled from the device node l , represents the penalty parameter, K l represents the computing power constraint of the l-th layer.
[0040] Based on the above method item embodiments, the present invention correspondingly provides device item embodiments;
[0041] The present invention provides a task scheduling device based on a distributed large language model, including:
[0042] A data acquisition module and a target device node determination module;
[0043] The above data acquisition module is used to acquire the communication data, tasks to be scheduled, task characteristics of the tasks to be scheduled, and real-time monitoring data of each device node in the multi-layer network of the distributed large language model; wherein, the above task characteristics include: task computing power, network conditions, task size, and computing power requirements; the above real-time monitoring data includes: computing power, memory, bandwidth, and latency;
[0044] The above target device node determination module is used to perform model segmentation and task scheduling on the above multi-layer network with the goal of minimizing the total latency of the multi-layer network according to the above communication data, tasks to be scheduled, task characteristics of the tasks to be scheduled, and real-time monitoring data of each device node, obtain the target device nodes corresponding to each task to be scheduled, and allocate each task to be scheduled to the corresponding target device node for scheduling;
[0045] Among them, model segmentation is performed according to the above communication data and the branch and bound integer linear programming algorithm to obtain the allocation scheme of each layer; the above allocation scheme includes: model size and computing requirements;
[0046] According to the above allocation scheme, allocate each task to be scheduled to each layer to obtain the target tasks to be scheduled corresponding to each layer;
[0047] According to the above target tasks to be scheduled, the target task characteristics corresponding to the above target tasks to be scheduled, the above real-time monitoring data, and a preset task scheduling model, perform task scheduling with the goal of minimizing the sum of the scheduling cost and the violation of the computing power constraint to obtain the target device nodes corresponding to each task to be scheduled; wherein, the sum of the above scheduling cost and the violation of the computing power constraint is calculated according to the above real-time monitoring data.
[0048] Furthermore, the above target device node determination module includes:
[0049] A priority queue initialization unit and a model splitting unit;
[0050] The above priority queue initialization unit is used to initialize the priority queue; wherein, the above priority queue includes several nodes for representing the above allocation scheme; the initialized priority queue includes: a root node without model splitting;
[0051] The above model splitting unit is used to repeatedly perform model splitting operations until the priority queue is empty, and use the optimal allocation scheme corresponding to when the priority queue is empty as the above allocation scheme;
[0052] Among them, the above model splitting operation includes:
[0053] Obtain the current priority queue, and arbitrarily select a node from the current priority queue as the target node; wherein, the initial node is the above root node;
[0054] If the current target node is a leaf node, calculate the total delay corresponding to the current target node according to the communication data corresponding to the current target node; wherein, the above leaf node represents the allocation scheme corresponding to the above multi-layer network;
[0055] When the current total delay is less than the minimum total delay corresponding to the current optimal solution, update the minimum total delay according to the current total delay, use the allocation scheme corresponding to the current target node as the current optimal allocation scheme, and delete the current target node from the current priority queue; otherwise, delete the current target node from the current priority queue; wherein, the initial optimal solution is infinity, and the initial optimal allocation scheme is empty;
[0056] If the current target node is not a leaf node, select an undetermined variable in the current target node for branching to obtain several child nodes; wherein, the above variables include: model size and computing requirements;
[0057] Calculate the total delay of each child node according to the communication data of the current child node. If the current maximum total delay of the child nodes is less than the minimum total delay corresponding to the current optimal solution, or the current minimum total delay of the child nodes is greater than the maximum total delay corresponding to the current optimal solution, then delete the current target node and the child nodes corresponding to the current target node from the current priority queue; otherwise, add the current child node as a node of the priority queue to the current priority queue, and at the same time delete the current target node to update the current priority queue.
[0058] Based on the above method item embodiments, the present invention correspondingly provides a terminal device item embodiment;
[0059] The present invention provides a terminal device, including a processor, a memory, and a computer program stored in the above memory and configured to be executed by the above processor. When the above processor executes the above computer program, it implements a task scheduling method based on a distributed large language model in any one of the embodiments of the present invention.
[0060] Based on the above method item embodiments, the present invention correspondingly provides a storage medium item embodiment;
[0061] The present invention provides a storage medium, including a processor, a memory, and a computer program stored in the above memory and configured to be executed by the above processor. When the above processor executes the above computer program, it implements a task scheduling method based on a distributed large language model in any one of the embodiments of the present invention.
[0062] The embodiments of the present invention have the following beneficial effects:
[0063] The present invention provides a task scheduling method, apparatus, terminal device, and storage medium based on a distributed large language model. The method includes: obtaining communication data, tasks to be scheduled, task characteristics of the tasks to be scheduled, and real-time monitoring data of each device node in the multi-layer network of the distributed large language model; wherein, the task characteristics include: task computing power, network conditions, task size, and computing power requirements; the real-time monitoring data includes: computing power, memory, bandwidth, and latency; subsequently, according to the communication data, tasks to be scheduled, task characteristics of the tasks to be scheduled, and real-time monitoring data of each device node, the multi-layer network is subjected to model segmentation and task scheduling with the goal of minimizing the total latency of the multi-layer network, obtaining the target device nodes corresponding to each task to be scheduled, and allocating each task to be scheduled to the corresponding target device node for scheduling; wherein, model segmentation is performed according to the communication data and the integer linear programming algorithm of branch and bound, obtaining the allocation scheme of each layer; the allocation scheme includes: model size and computing requirements; according to the allocation scheme, each task to be scheduled is allocated to each layer, obtaining the target tasks to be scheduled corresponding to each layer; according to the target tasks to be scheduled, the target task characteristics corresponding to the target tasks to be scheduled, the real-time monitoring data, and a preset task scheduling model, task scheduling is performed with the goal of minimizing the sum of the scheduling cost and the violation of the computing power constraint, obtaining the target device nodes corresponding to each task to be scheduled; wherein, the sum of the scheduling cost and the violation of the computing power constraint is calculated according to the real-time monitoring data. Therefore, when the present invention performs task scheduling, the entire process is divided into two parts: model segmentation and task scheduling. With the goal of minimizing the total latency of the multi-layer network, first, the model is segmented according to the integer linear programming algorithm of branch and bound and the communication data, and then, based on the model segmentation result, and based on the target tasks to be scheduled, the target task characteristics corresponding to the target tasks to be scheduled, the real-time monitoring data, and a preset task scheduling model, task scheduling is performed with the goal of minimizing the sum of the scheduling cost and the violation of the computing power constraint, determining the target device nodes corresponding to each task to be scheduled to complete the entire task scheduling. Thereby, resource management of the multi-layer network is realized, the performance of the network is improved, and the latency is reduced. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 FIG. is a flowchart of a task scheduling method based on a distributed large language model provided by an embodiment of the present invention.
[0065] Figure 2 FIG. is a schematic diagram of a two-stage multi-level multi-node scheduling of an LLM based on a collaborative AI computing method provided by an embodiment of the present invention.
[0066] Figure 3It is a schematic diagram of the automatic decoupling deployment of the inter-layer LLM based on the Transformer architecture provided by an embodiment of the present invention.
[0067] Figure 4 It is a schematic diagram of the comparison result of the inference task performance based on the OPT model in the three-layer framework provided by an embodiment of the present invention.
[0068] Figure 5 It is a schematic diagram of the comparison result of the inference task performance based on the Qwen2 model in the three-layer framework provided by an embodiment of the present invention.
[0069] Figure 6 It is a schematic diagram of the comparison result of the inference task performance based on the OPT model in the four-layer framework provided by an embodiment of the present invention.
[0070] Figure 7 It is a schematic diagram of the comparison result of the inference task performance based on the Qwen2 model in the four-layer framework provided by an embodiment of the present invention.
[0071] Figure 8 It is a schematic diagram of the structure of a task scheduling device based on a distributed large language model provided by an embodiment of the present invention. Detailed implementation manners
[0072] Next, the technical solutions in the present invention will be clearly and completely described in conjunction with the accompanying drawings in the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.
[0073] As Figure 1 shown, a task scheduling method based on a distributed large language model provided by an embodiment of the present invention includes:
[0074] Step S101: Obtain the communication data, tasks to be scheduled, task characteristics of the tasks to be scheduled, and real-time monitoring data of each device node in the multi-layer network of the distributed large language model; wherein, the above task characteristics include: task computing power, network conditions, task size, and computing power requirements; the above real-time monitoring data includes: computing power, memory, bandwidth, and latency;
[0075] Specifically, this paper considers a multi-layer cloud-edge network architecture, which contains L levels, denoted as l ∈ {1, 2,..., L}. Each level contains multiple computing nodes (mNodes), namely the above-mentioned "device nodes". These nodes can be edge devices, intermediate servers, or cloud data centers. The inference tasks of large language models (LLMs) need to be allocated and scheduled among these levels to achieve optimal performance and resource utilization. Use U l denote the set of device nodes at the l-th level. Each device node has different computing capabilities and network connection conditions.
[0076] Step S102: Based on the above communication data, tasks to be scheduled, task characteristics of the tasks to be scheduled, and real-time monitoring data of each device node, perform model segmentation and task scheduling on the above multi-layer network with the goal of minimizing the total latency of the multi-layer network, obtain the target device nodes corresponding to each task to be scheduled, and allocate each task to be scheduled to the corresponding target device node for scheduling;
[0077] Among them, model segmentation is performed according to the above communication data and the branch-and-bound integer linear programming algorithm to obtain the allocation scheme for each level; the above allocation scheme includes: model size and computing requirements;
[0078] Allocate each task to be scheduled to each level according to the above allocation scheme to obtain the target tasks to be scheduled corresponding to each level;
[0079] Based on the above target tasks to be scheduled, the target task characteristics corresponding to the above target tasks to be scheduled, the above real-time monitoring data, and a preset task scheduling model, perform task scheduling with the goal of minimizing the sum of the scheduling cost and the violation of the computing capacity constraint to obtain the target device nodes corresponding to each task to be scheduled; among them, the sum of the above scheduling cost and the violation of the computing capacity constraint is calculated according to the above real-time monitoring data.
[0080] In a preferred embodiment, when performing model segmentation and task scheduling on the above multi-layer network with the goal of minimizing the total latency of the multi-layer network, the corresponding first objective function is:
[0081]
[0082] R u =B k log2(1 + SINR us )
[0083]
[0084]
[0085] In the formula, T totalDenote the total latency as \(T\), the total number of layers as \(L\), and the \(l\)-th layer as \(l\). Denote the transmission latency of the \(l\)-th layer as \(T_{l}^{t}\). Denote the computing latency of the \(l\)-th layer as \(T_{l}^{c}\). l Denote the device node of the \(l\)-th layer as \(U_{l}\). l Denote the total number of device nodes in the \(l\)-th layer as \(u_{l}\). l′ Denote the \(l\)-th ′ layer's device node as \(U_{l}\). l′ Denote the \(l\)-th ′ layer's total number of device nodes. Denote whether the scheduling task is scheduled from layer \(l\) to layer \(u\) as \(\beta_{l,u}\). l′ layer as \(\beta_{l,u}\). l Denote the computing demand allocation coefficient of the \(l\)-th layer as \(M_{l}\). c Denote the computing demand as \(C\). l Denote the computing capacity of the device node in the \(l\)-th layer as \(\gamma_{l}\). l Denote the model size allocation coefficient of the \(l\)-th layer as \(M_{l}\). s Denote the model size as \(R\). u Denote the achievable transmission rate of node \(u\) as \(B_{u}\). k Denote the bandwidth of sub-channel \(k\) as \(SINR_{k}\). us Denote the transmission signal interference noise ratio as \(SINR\). Denote the interference plus noise power as \(I + N\). Denote the interference from other users \(g\) in the same sub-channel set to node \(u\) in the \(l\)-th ′ layer as \(I_{l,u,g}\). l ′ of node \(u\). Denote the transmission signal interference noise ratio between layer \(l'\) and layer \(l\) as \(SINR_{l',l}\). Denote from node \(u\) in the \(l\)-th layer l to node \(u\) in the \(l'\)-th layer l′ as the channel gain \(h_{l,u,l',u}\). Denote the transmission power of node \(u\) in the \(l\)-th layer l to node \(u\) l′ as \(P_{l,u}\).
[0086] Specifically, consider the entire LLM model as a divisible task, and its model size and computing demand are denoted as \(M_s\) and \(M_c\) respectively. According to the model splitting decision \(\pi\), the model size and computing demand are allocated to different layers. The model size and computing demand allocated to the \(l\)-th layer are \(\gamma_{l}\) l \(M_s\) and \(\beta_{l}\) l \(M_c\), where \(\gamma_{l}\) l and \(\beta_{l}\) l are the corresponding allocation coefficients. The LLM model can be decomposed into multiple independent parts, and these independent parts can be executed in parallel on different mNodes. This decomposition can be represented by the following formula:
[0087]
[0088] In the formula, M i represents the i-th segment of the LLM model that is segmented, and n represents the total number of segments.
[0089] Specifically, the transmission delay in a multi-layer network mainly depends on the size of the model segments, bandwidth, channel gain, and channel noise. First, it is necessary to calculate the transmission rate affected by noise. Assume u l represents the node at the l-th layer, and u' l represents the node at the l'-th layer. Then, the signal-to-interference-plus-noise ratio (SINR) between layers l and l' can be expressed as:
[0090]
[0091] In the formula, the interference plus noise power, which is defined as follows:
[0092]
[0093] Specifically, the achievable transmission rate R u of node l can be expressed as:
[0094] R u = B k log2(1 + SINR us )
[0095] Specifically, the transmission delay of the l-th layer can be expressed as:
[0096]
[0097] Specifically, the calculation delay depends on the calculation requirements of the model segments, the computing power of the nodes, and the scheduling decision. Therefore, the calculation delay of the l-th layer can be expressed as:
[0098]
[0099] Specifically, according to the above various formulas, the total delay in the multi-layer network can be obtained:
[0100]
[0101] Therefore, according to the above various formulas, when performing model segmentation and task scheduling on the above multi-layer network with the goal of minimizing the total delay of the multi-layer network, the corresponding first objective function is obtained. The entire optimization problem can be transformed into minimizing the total delay, and its goal is to find the optimal model segmentation decision π and scheduling decision A.
[0102] Schematically, the schematic diagram of the two-stage multi-level multi-node scheduling of the LLM based on the collaborative AI computing method is as Figure 2 shown, Figure 2 This paper elaborates on the multi-level multi-node LLM scheduling (MMSL) method, which is the task scheduling method based on the distributed large language model shown in the present invention, aiming to optimize the inference efficiency of the large language model (LLM) in the multi-level network. This method decomposes the complex optimization problem (P0) into two sub-problems: P1 and P2, thus simplifying the solution process. P0 represents the overall MMSL problem, whose goal is to calculate and optimize the LLM inference in the multi-level network. The core challenge lies in how to efficiently allocate computing resources among different levels to minimize the total latency.
[0103] Specifically, to address the complexity of the original problem P0, the MMSL method decomposes it into two sub-problems. P1, namely cross-layer LLM automatic decoupling and slicing (CT-ADS), focuses on determining the optimal model segmentation strategy π. This involves deciding the model size and computing requirements that each level should handle, and is achieved through integer linear programming techniques, especially the branch and bound algorithm-based (ILP-BB). ILP-BB searches the solution space systematically to find the optimal resource allocation scheme, including monitoring the resource situation at each level, solving the optimization problem using the ILP-BB algorithm, and transforming the resource allocation problem into a mathematical model for solution. P2, namely intra-layer LLM adaptive scheduling and distribution (IT-ASD), is responsible for finding the optimal task scheduling scheme A within each level. This requires considering node resource utilization and network conditions to achieve load balancing and minimum latency, and is achieved through the graph neural network (GNN)-based task scheduling algorithm (TS-GNN). GNN optimizes task scheduling decisions by propagating information and updating node features, including monitoring the device resource utilization of each layer and using the GNN algorithm to optimize the intra-layer task allocation.
[0104] Preferably, by sequentially solving the two sub-problems P1 and P2, the MMSL method can effectively minimize the total latency of large model processing in the entire multi-level computing network. This decomposition greatly reduces the computational complexity and also provides a structured framework for addressing the challenges in large model deployment and resource management, ultimately achieving efficient scheduling of LLM inference in the multi-level network, highlighting the key roles of problem decoupling, mathematical formulation and optimization, and graph neural networks in solving complex optimization problems.
[0105] In this preferred embodiment, by calculating the total latency within the multi-level network, a first objective function is constructed for model segmentation and task scheduling of the above multi-level network with the goal of minimizing the total latency of the multi-level network.
[0106] In another preferred embodiment, the above-mentioned allocation schemes for each level are obtained according to the above-mentioned communication data and the branch-and-bound integer linear programming algorithm, including:
[0107] Initialize the priority queue; wherein, the above-mentioned priority queue includes several nodes for representing the above-mentioned allocation schemes; the initialized priority queue includes: the root node without model segmentation;
[0108] Repeat the model segmentation operation until the priority queue is empty, and use the optimal allocation scheme corresponding to the empty priority queue as the above-mentioned allocation scheme;
[0109] Among them, the above-mentioned model segmentation operation includes:
[0110] Obtain the current priority queue, and arbitrarily select a node from the current priority queue as the target node; wherein, the initial node is the above-mentioned root node;
[0111] If the current target node is a leaf node, calculate the total delay corresponding to the current target node according to the communication data corresponding to the current target node; wherein, the above-mentioned leaf node represents the allocation scheme corresponding to the above-mentioned multi-layer network;
[0112] When the current total delay is less than the minimum total delay corresponding to the current optimal solution, update the minimum total delay according to the current total delay, use the allocation scheme corresponding to the current target node as the current optimal allocation scheme, and delete the current target node from the current priority queue; otherwise, delete the current target node from the current priority queue; wherein, the initial optimal solution is infinity, and the initial optimal allocation scheme is empty;
[0113] If the current target node is not a leaf node, select an undetermined variable in the current target node for branching to obtain several child nodes; wherein, the above-mentioned variables include: model size and computing requirements;
[0114] Calculate the total delay of each child node according to the communication data of the current child node. If the current maximum total delay of the child nodes is less than the minimum total delay corresponding to the current optimal solution, or the current minimum total delay of the child nodes is greater than the maximum total delay corresponding to the current optimal solution, then delete the current target node and the child nodes corresponding to the current target node from the current priority queue; otherwise, add the current child node as a node of the priority queue to the current priority queue, and at the same time delete the current target node to update the current priority queue.
[0115] Schematically, the schematic diagram of the automatic decoupling deployment of the inter-layer LLM based on the Transformer architecture is as Figure 3 shown Figure 3Depicts the deployment architecture of a Transformer-based large language model (LLM) in a multi-layer computing environment, showing how the LLM model can be efficiently segmented and processed in a network composed of cloud, intermediate layer, and edge devices to optimize performance and reduce latency. Figure 3 Shows a multi-layer network structure, including a user layer and the first, second, and third layers, each containing mNodes in either an idle or busy state. These layers can be regarded as representatives of the cloud, intermediate nodes, and edge nodes, each responsible for different types of computing tasks. The cloud is mainly responsible for model training and management, intermediate nodes handle preprocessing and preliminary inference, and edge nodes focus on local inference and personalized fine-tuning.
[0116] Specifically, Figure 3 Also details the components in a Transformer-based LLM model, including an Encoder, a Decoder, a Multi-head Attention mechanism, a Feed Forward Layer, and Add&Normalization steps. The input to the Encoder is an Embed. The Decoder structure is similar to the Encoder but adds a Masked Multi-head Attention and an Encoder-Decoder Attention layer to complete specific processing tasks. The output of the Decoder passes through a softmax layer to generate the probabilities of target words. Different computing tasks or data streams are assigned to different mNodes in the multi-layer network for parallel processing, making full use of idle computing resources and thus reducing inference latency.
[0117] Specifically, the flexibility of the Transformer architecture allows the LLM model to be segmented into multiple smaller fragments (Mi), which can be independently executed on different device nodes in the multi-layer edge network. This flexible segmentation strategy enables the model to be elastically deployed on different mNodes according to computing resources and task requirements.
[0118] Specifically, the integer linear programming algorithm based on branch and bound, namely the ILP-BB algorithm, whose core goal is to achieve cross-layer automatic decoupling and segmentation for the deployment of a large language model (LLM) in a multi-layer network environment. This decoupling and segmentation dynamically optimize the allocation of model size and computing requirements according to the computing capabilities and resource conditions of different layers in the network, i.e., the corresponding coefficients γ l and β lThe size. Through this refined resource allocation, the latency of the entire system is minimized, and the inference efficiency is improved. Due to the complexity of cross-layer resource allocation, directly solving this problem is computationally infeasible. Therefore, the algorithm adopts the branch and bound method. This method cleverly decomposes a complex optimization problem into a series of more manageable and solvable sub-problems, and then gradually approaches the optimal solution.
[0119] Specifically, the execution process of the algorithm can be understood as a systematic search process. First, initialize a priority queue (PQ) and add the root node representing the initial state of the model size and computing allocation to the queue. This root node represents the initial state of the entire optimization problem, where the model has not been partitioned and resources have not been allocated. At the same time, initialize the objective function value z of the optimal solution to infinity, and the optimal allocation scheme π to be empty. Here, z represents the total latency. Subsequently, as long as the priority queue is not empty, continuously extract nodes from it for processing. If the currently processed node is a leaf node, it means that all parts of the model have been allocated. At this time, the algorithm calculates the objective function value z corresponding to this node, that is, the total latency under the current allocation scheme. If the current z value is less than the objective function value z* of the current optimal solution, it means that a better allocation scheme has been found, and the algorithm will update z* and the optimal allocation scheme π*. Among them, the objective function value z is calculated through the following formula:
[0120]
[0121] In the formula, C l represents the computing power of the l-th layer.
[0122] If the current node is not a leaf node, it means that there are still parts of the model unallocated. The algorithm will select an unallocated variable, that is, γ l or β l to branch. Subsequently, two sub-problems are created based on the selected variable, that is, the above-mentioned "child nodes", which represent the lower bound and upper bound of this variable respectively. For example, if γ l is selected for branching, the algorithm will try to set the value of γ l to the floor and ceiling of its current value respectively. Then, for each sub-problem, the algorithm calculates its corresponding objective function value and boundary. Here, the algorithm evaluates the sub-problem by calculating z and updates the lower bound and upper bound. The lower bound takes the minimum value of the objective function values of all sub-problems, and the upper bound takes the maximum value of the objective function values of all sub-problems. These boundary values are used for pruning operations to avoid unnecessary searches. If the lower bound is greater than the upper bound of the current optimal solution, or the upper bound is less than the lower bound of the current optimal solution, it means that this branch cannot contain the optimal solution, and the algorithm will perform pruning and no longer continue to explore this branch.
[0123] Preferably, by continuously selecting the node with the highest priority for branching and pruning, the algorithm gradually approaches the optimal solution. This process can be regarded as searching in the solution space, reducing the search scope through pruning operations, and the whole process continues until the optimal solution is found or it is confirmed that there is no solution that meets the conditions. All in all, the ILP-BB algorithm uses the branch and bound method to efficiently search the solution space of the integer linear programming problem, providing an effective cross-layer model segmentation and resource allocation scheme for the LLM deployment in the multi-layer network. This helps to maximize the performance and efficiency of the LLM in resource-constrained environments.
[0124] In this preferred embodiment, the integer linear programming algorithm based on branch and bound is used to perform model segmentation on the LLM, and an allocation scheme with the minimum total delay of the multi-layer network is obtained.
[0125] In another preferred embodiment, according to the above-mentioned target tasks to be scheduled, the target task characteristics corresponding to the above-mentioned target tasks to be scheduled, the above-mentioned real-time monitoring data, and a preset task scheduling model, with the goal of minimizing the sum of the scheduling cost and the violation of the computing power constraint, the target device nodes corresponding to each task to be scheduled are obtained, including:
[0126] Initialize the list of tasks to be scheduled; among them, the initialized list of tasks to be scheduled includes all the above-mentioned target tasks to be scheduled;
[0127] Repeat the task scheduling operation until the list of tasks to be scheduled is empty, and determine the target device nodes corresponding to each task to be scheduled;
[0128] Among them, the above-mentioned task scheduling operation includes:
[0129] Select a device node as the current node to be scheduled, and obtain the current task to be scheduled from the current list of tasks to be scheduled;
[0130] Construct the current structure diagram according to the current task to be scheduled and the corresponding current target task characteristics;
[0131] Obtain the current real-time monitoring data and the above-mentioned preset task scheduling model, and input the current structure diagram, the current task characteristics, and the current real-time monitoring data into the above-mentioned preset task scheduling model, so that the above-mentioned preset task scheduling model calculates the current degree matrix and adjacency matrix according to the current structure diagram;
[0132] Obtain the current node feature matrix of the last layer in the above-mentioned preset task scheduling model according to the current degree matrix, the above-mentioned adjacency matrix, and the current target task characteristics;
[0133] According to the current node feature matrix and activation function, with the goal of minimizing the sum of scheduling cost and violation of computing power constraints, obtain the scheduling probability distribution between the current nodes to be scheduled and each current task to be scheduled;
[0134] Take the current node to be scheduled as the target device node of the task to be scheduled with the highest scheduling probability, and delete this node to be scheduled from the current list of nodes to be scheduled to update the list of nodes to be scheduled.
[0135] Specifically, according to the allocation scheme obtained previously, the size of the segmented models deployed in each layer of the LLM and the corresponding computing requirements can be determined. Subsequently, based on the model size and computing requirements, the tasks that can be scheduled at this layer can be determined.
[0136] Specifically, the core goal of the task scheduling algorithm based on graph neural network is to achieve adaptive task scheduling within each layer for the inference tasks of large language models (LLMs) in a multi-layer network environment. This is not simply about allocating tasks to each device node, but rather making optimal task offloading decisions dynamically based on the real-time resource status of each node (such as computing power, memory usage, etc.) and network conditions (such as bandwidth, latency, etc.). This algorithm aims to maximize the resource utilization rate of the entire system, reduce the task execution time, and thus improve the overall performance through this intelligent task scheduling. Different from the branch-and-bound based integer linear programming algorithm that focuses on cross-layer model partitioning and resource allocation, the task scheduling algorithm based on graph neural network focuses on fine-grained task scheduling within the same layer.
[0137] Specifically, the core idea of the algorithm is to use graph neural network (GNN) to process network topology structure data and perform propagation and update of node features to optimize task scheduling decisions. This enables the algorithm to understand the relationships between various nodes in the network and the resource status of each node, thus making more informed scheduling decisions. Specifically, the algorithm first initializes the parameters of the task scheduling model, including weight matrices, bias terms, etc. Then, the algorithm continuously monitors network resources, including but not limited to: the size of tasks Penalty parameters used to constrain task scheduling decisions and the computing capacity limit (K) of each layer of nodes l)。The task size here can be understood as the amount of computation required for the LLM inference task; the penalty parameter is used to penalize behaviors that exceed the computing capacity of the device node during the optimization process; while the computing capacity limit represents the maximum processing capacity of the device node. After monitoring the network resources, the algorithm performs separate task scheduling for each layer in the network. For each layer, the algorithm collects the tasks to be scheduled within the current layer and places these tasks in a waiting list (i.e., the aforementioned list of tasks to be scheduled). These tasks to be scheduled may be part of the model or data that needs further processing. Subsequently, the algorithm schedules these tasks one by one until the waiting list is empty.
[0138] The algorithm calculates the degree matrix (D) and the adjacency matrix (Λ). Among them, the degree matrix describes the number of other nodes connected to each node in the structure diagram; while the adjacency matrix describes the connection relationship between nodes in the structure diagram. Both are basic representations of graph-structured data. Subsequently, the algorithm performs feature propagation through multiple GNN layers. In each GNN layer, a node aggregates and updates its own feature information and the feature information from adjacent nodes, thereby learning more comprehensive network state information. This process is carried out through the following formula:
[0139] H (l+1) =σ(D -1 / 2 ΛD -1 / 2 H (l) W (l) )
[0140] In the formula, H (l) is the node feature matrix of the l-th GNN layer, Λ is the adjacency matrix, D is the degree matrix, W (l) is the weight matrix of the l-th GNN layer, and σ is the activation function (such as ReLU). Feature propagation enables each node in the network to "perceive" the states of other nodes, thereby providing a basis for subsequent scheduling decisions. After being processed by multiple GNN layers, on the output of the last GNN layer, the Softmax activation function is used to determine the optimal target device node for the task. The Softmax function converts the feature vector of the node into a probability distribution, indicating the likelihood of each node corresponding to the task scheduling target as the target device node. The task to be scheduled corresponding to the node with the highest probability is scheduled to that target device node.
[0141] Specifically, when training the task scheduling model, to optimize the parameters of the task scheduling model, the algorithm also defines a loss function (L) based on the objective function of the Quadratic Unconstrained Binary Optimization (QUBO) model. The QUBO model is a mathematical framework for solving combinatorial optimization problems, which allows the algorithm to transform complex scheduling problems into an optimizable mathematical form. The loss function L is designed to reflect the cost of task scheduling decisions and the penalty for violating the computing capacity limit. By using the backpropagation algorithm, the algorithm updates the model parameters according to the value of the loss function L, thereby minimizing the loss function and achieving the optimization of the task scheduling strategy. In this way, the algorithm transforms a discrete optimization problem (task scheduling) into a continuous optimization problem (minimizing the loss function). This transformation enables the GNN to be effectively applied to large-scale combinatorial optimization problems in various practical scenarios. The final output of the algorithm is the scheduling decision variable It is a binary value used to indicate whether a task is transferred from one node l to another node u
[0142] In this preferred embodiment, based on the target task to be scheduled, the target task features corresponding to the target task to be scheduled, the real-time monitoring data, and the preset task scheduling model, with the goal of minimizing the sum of the scheduling cost and the violation of the computing capacity constraint, the target device node corresponding to each task to be scheduled is obtained
[0143] In another preferred embodiment, the construction of the current structure diagram according to the current task to be scheduled and the corresponding current target task features includes
[0144] Using the current task to be scheduled as the graph node of the current structure diagram and connecting each graph node according to the current target task features to obtain the current structure diagram
[0145] Specifically, to capture the spatial and temporal correlations between tasks to be scheduled, the tasks to be scheduled within each level are represented as nodes in the structure diagram. The edges represent the dependencies between tasks, and these relationships are affected by characteristics such as task computing power, network conditions, task size, and computing power requirements. For example, when there are time conflicts between task executions, edge connections are generated between the corresponding nodes in the graph
[0146] In this preferred embodiment, a structure diagram is constructed based on each task to be scheduled and the dependencies between the tasks to be scheduled.
[0147] In another preferred embodiment, when performing task scheduling with the goal of minimizing the sum of the scheduling cost and the violation of the computing power constraint, the corresponding second objective function is:
[0148]
[0149] In the formula, H represents the second objective function, represents the task size, represents whether the task to be scheduled is scheduled from the device node to the device node u l , represents the penalty parameter, K l represents the computing power constraint of the l-th layer.
[0150] Specifically, the above uses quadratic unconstrained binary optimization (QUBO) to model the task scheduling problem. The scheduling decision between the nodes of a specific layer and the node u of layer l l can be represented by the binary variable . If then it means that the task is scheduled from the node to the node u l , otherwise
[0151] Specifically, represents the scheduling cost, represents the penalty for violating the computing power constraint.
[0152] In another preferred embodiment, to comprehensively evaluate a task scheduling method based on a distributed large language model proposed by the present invention, three computing levels are designed, and multiple heterogeneous device nodes are deployed at each level. These nodes have different computing and communication capabilities, aiming to simulate a complex heterogeneous computing environment in the real world. The specific configuration is as follows: The first level: Six Jetson AGX Orin nodes are deployed, each node having a powerful computing capability of up to 275 TFLOPS, suitable for processing tasks with high computing loads; The second level: Six Jetson Orin NX nodes are deployed, each node having a computing capability of 70 TFLOPS, serving as an intermediate layer between the cloud and the edge, responsible for processing some computationally intensive tasks; The third level: Six Jetson AGX Xavier nodes are deployed, each node having a computing capability of 32 TFLOPS, located at the network edge, directly serving end users, and processing local inference tasks. To simulate a real LLM application scenario, Qwen2 and OPT models are selected as test objects. These models are split into multiple parts, and then these parts are deployed to different mNode nodes for inference tasks. The number of input tokens for the inference tasks is set to 4096 or 2048, simulating large-scale and medium / small-scale input scenarios respectively.
[0153] The network communication is based on the Orthogonal Frequency Division Multiple Access (OFDMA) protocol, which is widely used in wireless communication networks. The bandwidth is set to 5 MHz, and the channel noise is -117 dBm, simulating the actual communication environment. In addition, to simulate the diversity of user devices, the device communication power randomly varies between 0.3W, 0.5W, and 0.7W. To more realistically reflect the characteristics of the multi-layer network, the delays of each level are introduced, which are set to 10 ms for the first level, 20 ms for the second level, and 30 ms for the third level, simulating the delays generated when data is transmitted and processed at different levels. To ensure the reliability of the system, redundant nodes are added. When a certain node fails, the redundant node can take over its work to ensure the continuous operation of the system. Comparisons are made based on five algorithms respectively:
[0154] The first algorithm: Uniform Model Partitioning and Random Assignment Algorithm (UMPRA): The LLM model is evenly partitioned into multiple tasks of the same size, and then randomly assigned to the nodes in the multi-layer network. This method is simple and direct, but may not fully utilize the computing capabilities of different nodes.
[0155] The second algorithm: Random Model Partitioning and Random Scheduling Algorithm (RMPSS): The LLM model is randomly partitioned into tasks of different sizes, and then randomly scheduled to the nodes in the multi-layer network. This method is more flexible, but may lead to unbalanced loads.
[0156] The third algorithm: Model Partitioning and Arbitrary Task Distribution Algorithm (MPATD): It uses Integer Linear Programming (ILP) to partition the LLM model and then arbitrarily assigns tasks to the mNode nodes. This is a baseline considering model partitioning, but the task assignment is not optimized.
[0157] The fourth algorithm: Model Partitioning and Heuristic Task Scheduling Algorithm (MPHTS): It uses Integer Linear Programming (ILP) to partition the LLM model and uses heuristic rules for task scheduling. This is a relatively optimized baseline that performs task assignment through heuristic rules.
[0158] The algorithm proposed in this application: Multi - layer Multi - node LLM Collaborative Computing Scheduling Algorithm (MMSL): It adopts dynamic model partitioning and graph - based analysis to optimize task scheduling. This is the algorithm proposed in this application, which realizes more efficient resource utilization and task scheduling.
[0159] Schematically, the performance comparison results of the inference tasks based on the OPT model in the three - layer framework are as Figure 4 shown, and the performance comparison results of the inference tasks based on the Qwen2 model in the three - layer framework are as Figure 5 shown. Figure 4 and Figure 5 show the performance of different algorithms when executing LLM inference tasks under the three - layer network architecture. The horizontal axis represents the number of tasks to be processed (from 5 to 45), and the vertical axis represents the total inference time required to complete all tasks. Figure 4 shows the performance when using the OPT model with an input of 4096 tokens. The experimental results clearly show that the completion times of UMPRA and RMPSS are significantly higher than that of MMSL. Specifically, when the number of tasks is 40, UMPRA and RMPSS respectively need 1219 seconds and 1042 seconds to complete all tasks, while MMSL only needs 708 seconds. This indicates that the efficiency of the MMSL algorithm is significantly better than that of UMPRA and RMPSS, with improvements of 72.2% and 47.2% respectively. Figure 5 shows the results when using the Qwen2 model. As the number of tasks increases, the completion times of UMPRA and RMPSS increase rapidly, while the time growth rate of MMSL is significantly slower, indicating that MMSL has better scalability when dealing with more tasks.
[0160] Schematically, the performance comparison results of the inference tasks based on the OPT model in the four - layer framework are as Figure 6 shown, and the performance comparison results of the inference tasks based on the Qwen2 model in the four - layer framework are as Figure 7 shown. Figure 6 and Figure 7Shows the performance under a four - layer network architecture. This network adds an edge computing layer on the basis of the original three - layer structure and also uses Jetson AGX Xavier nodes. The experimental results are similar to those of the three - layer network. The MMSL algorithm always maintains the lowest task execution time, further proving that MMSL has efficient model segmentation performance under different network structures. Figure 6 uses the OPT model, Figure 7 is the result of using the Qwen2 model. The results show that through an efficient model segmentation mechanism, MMSL can still obtain the shortest task execution time in a four - layer network.
[0161] Based on the above - mentioned method - item embodiments, the present invention correspondingly provides device - item embodiments.
[0162] As Figure 8 shown, an embodiment of the present invention provides a task scheduling device based on a distributed large - language model, including:
[0163] a data acquisition module and a target device node determination module;
[0164] The above - mentioned data acquisition module is used to acquire the communication data, tasks to be scheduled, task characteristics of the tasks to be scheduled, and real - time monitoring data of each device node in the multi - layer network of the distributed large - language model; wherein, the above - mentioned task characteristics include: task computing power, network conditions, task size, and computing power requirements; the above - mentioned real - time monitoring data includes: computing power, memory, bandwidth, and latency;
[0165] The above - mentioned target device node determination module is used to perform model segmentation and task scheduling on the above - mentioned multi - layer network with the goal of minimizing the total latency of the multi - layer network according to the above - mentioned communication data, tasks to be scheduled, task characteristics of the tasks to be scheduled, and real - time monitoring data of each device node, obtain the target device nodes corresponding to each task to be scheduled, and allocate each task to be scheduled to the corresponding target device node for scheduling;
[0166] Among them, model segmentation is performed according to the above - mentioned communication data and the branch - and - bound integer linear programming algorithm to obtain the allocation scheme for each level; the above - mentioned allocation scheme includes: model size and computing requirements;
[0167] Each task to be scheduled is allocated to each level according to the above - mentioned allocation scheme to obtain the target tasks to be scheduled corresponding to each level;
[0168] According to the above target tasks to be scheduled, the target task characteristics corresponding to the above target tasks to be scheduled, the above real-time monitoring data, and a preset task scheduling model, task scheduling is performed with the goal of minimizing the sum of the scheduling cost and the violation of the computing power constraint, and the target device nodes corresponding to each task to be scheduled are obtained; wherein, the sum of the above scheduling cost and the violation of the computing power constraint is calculated based on the above real-time monitoring data.
[0169] In a preferred embodiment, the above target device node determination module includes:
[0170] A priority queue initialization unit and a model splitting unit;
[0171] The above priority queue initialization unit is used to initialize the priority queue; wherein, the above priority queue includes several nodes for representing the above allocation scheme; the initialized priority queue includes: a root node without model splitting.
[0172] The above model splitting unit is used to repeatedly perform model splitting operations until the priority queue is empty, and use the optimal allocation scheme corresponding to when the priority queue is empty as the above allocation scheme.
[0173] Wherein, the above model splitting operation includes:
[0174] Obtain the current priority queue, and arbitrarily select a node from the current priority queue as the target node; wherein, the initial node is the above root node.
[0175] If the current target node is a leaf node, calculate the total delay corresponding to the current target node according to the communication data corresponding to the current target node; wherein, the above leaf node represents the allocation scheme corresponding to the above multi-layer network.
[0176] When the current total delay is less than the minimum total delay corresponding to the current optimal solution, update the minimum total delay according to the current total delay, use the allocation scheme corresponding to the current target node as the current optimal allocation scheme, and delete the current target node from the current priority queue; otherwise, delete the current target node from the current priority queue; wherein, the initial optimal solution is infinity, and the initial optimal allocation scheme is empty.
[0177] If the current target node is not a leaf node, select an undetermined variable in the current target node for branching to obtain several child nodes; wherein, the above variables include: model size and computing requirements.
[0178] Calculate the total delay of each child node according to the communication data of the current child node. If the current maximum total delay of the child nodes is less than the minimum total delay corresponding to the current optimal solution, or the current minimum total delay of the child nodes is greater than the maximum total delay corresponding to the current optimal solution, then delete the current target node and the child nodes corresponding to the current target node from the current priority queue; otherwise, add the current child node as a node of the priority queue to the current priority queue, and at the same time delete the current target node to update the current priority queue.
[0179] It should be noted that the device embodiments described above are only illustrative. The modules described as separate components above may or may not be physically separated. The components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided by the present invention, the connection relationship between the modules indicates that they have a communication connection, which can be specifically implemented as one or more communication buses or signal lines. Those of ordinary skill in the art can understand and implement without creative efforts. The above schematic diagram is only an example of a task scheduling device based on a distributed large language model, and does not constitute a limitation on a task scheduling device based on a distributed large language model. It may include more or fewer components than shown in the figure, or combine some components, or different components.
[0180] Based on the above method item embodiments, the present invention correspondingly provides terminal device item embodiments.
[0181] Another embodiment of the present invention provides a terminal device, including a processor, a memory, and a computer program stored in the above memory and configured to be executed by the above processor. When the above processor executes the above computer program, it implements the method for task scheduling based on a distributed large language model in any one of the above embodiments of the present invention.
[0182] Exemplarily, in this embodiment, the above computer program can be divided into one or more modules. The above one or more modules are stored in the above memory and executed by the above processor to complete the present invention. The above one or more module elements can be a series of computer program instruction segments capable of completing specific functions, and this instruction segment is used to describe the execution process of the above computer program in the above device;
[0183] The above terminal device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The above device may include, but is not limited to, a processor and a memory;
[0184] The so-called processor may be a Central Processing Unit (CPU), or it may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The above-mentioned processor is the control center of the above-mentioned device, and connects various parts of the entire device through various interfaces and lines;
[0185] The above-mentioned memory can be used to store the above-mentioned computer programs and / or modules. The above-mentioned processor realizes various functions of the above-mentioned device by running or executing the computer programs and / or modules stored in the above-mentioned memory, and by calling the data stored in the memory. The above-mentioned memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function, etc. In addition, the memory can include high-speed random access memory, and can also include non-volatile memory, such as a hard disk, memory, plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices.
[0186] Based on the above method item embodiments, the present invention correspondingly provides storage medium item embodiments.
[0187] Another embodiment of the present invention provides a storage medium. The above storage medium includes a stored computer program, wherein when the above computer program runs, it controls the device where the above storage medium is located to execute the task scheduling method based on a distributed large language model in any one of the above embodiments of the present invention.
[0188] In this embodiment, the above storage medium is a computer-readable storage medium, and the above computer program includes computer program code, which can be in the form of source code, object code, executable file or some intermediate form, etc. The above computer-readable medium may include: any entity or device capable of carrying the above computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0189] Compared with the prior art, by implementing the above various embodiments of the present invention, resource management of multiple internal layers of the large language model can be achieved, the performance of the network is improved, and the latency is reduced.
[0190] The above are the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art of this technology, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present invention.
Claims
1. A task scheduling method based on a distributed large language model, characterized in that: include: Obtain communication data of each device node in the multi-layer network of the distributed large language model, tasks to be scheduled, task characteristics of the tasks to be scheduled, and real-time monitoring data of each device node; wherein the task characteristics include: task computing power, network conditions, task size, and computing power requirements; the real-time monitoring data includes: computing power, memory, bandwidth, and latency; According to the communication data, the tasks to be scheduled, the task characteristics of the tasks to be scheduled, and the real-time monitoring data of each device node, the multi-layer network is model segmented and task scheduled with the goal of minimizing the total delay of the multi-layer network, the target device node corresponding to each task to be scheduled is obtained, and each task to be scheduled is assigned to the corresponding target device node for scheduling; The model is segmented according to the communication data and the branch-and-bound integer linear programming algorithm to obtain allocation schemes for each level; the allocation schemes include: model size and computing requirements; Allocate each task to be scheduled to each level according to the allocation scheme to obtain the target task to be scheduled corresponding to each level; According to the target tasks to be scheduled, the target task characteristics corresponding to the target tasks to be scheduled, the real-time monitoring data, and the preset task scheduling model, task scheduling is performed with the goal of minimizing the sum of scheduling cost and violation of computing power constraints, and the target device nodes corresponding to each task to be scheduled are obtained; wherein the sum of scheduling cost and violation of computing power constraints is calculated based on the real-time monitoring data.
2. The task scheduling method based on a distributed large language model according to claim 1, characterized in that: When the multi-layer network is model segmented and task scheduled with the goal of minimizing the total delay of the multi-layer network, the corresponding first objective function is: R u =B k log2(1+SINR us ) Where, T total represents the total delay, L represents the total number of levels, l represents the lth level, represents the transmission delay at level l, represents the computation delay at level l, u l Indicates the device node at level l, U l Indicates the total number of device nodes at level l, u l′ Indicates the first ′ Hierarchical device nodes, U l′ Indicates the first ′ The total number of device nodes in the hierarchy, Indicates whether the scheduled task is scheduled from level l to level u l′ Level, β l represents the computing demand allocation coefficient of the lth level, M c represents the computational demand, C l represents the computing power of the device node in the lth level, γ l represents the size distribution coefficient of the l-th level model, M s Indicates the model size, R u represents the achievable transmission rate of node u, B k represents the bandwidth of subchannel k, SINR us represents the transmission signal-to-interference-noise ratio, represents the interference plus noise power at level l, Indicates the other users g in the same subchannel set have ′ Level node u l ′ interference, represents the transmission signal interference noise ratio between level l' and level l, Indicates that from the l-th level node u l To the l'th level node u l′ The channel gain, Represents the node u at layer l l To node u l′ transmission power.
3. The task scheduling method based on a distributed large language model according to claim 2, characterized in that: The allocation schemes of each level are obtained according to the communication data and the branch-and-bound integer linear programming algorithm, including: Initializing a priority queue; wherein the priority queue includes a plurality of nodes for representing the allocation scheme; the initialized priority queue includes: a root node that has not been model-segmented; Repeat the model splitting operation until the priority queue is empty, and use the optimal allocation scheme corresponding to the empty priority queue as the allocation scheme; The model segmentation operation includes: Obtain the current priority queue, and select any node from the current priority queue as the target node; wherein the initial node is the root node; If the current target node is a leaf node, the total delay corresponding to the current target node is calculated according to the communication data corresponding to the current target node; wherein the leaf node represents the allocation scheme corresponding to the multi-layer network; When the current total delay is less than the minimum total delay corresponding to the current optimal solution, the minimum total delay is updated according to the current total delay, the allocation scheme corresponding to the current target node is used as the current optimal allocation scheme, and the current target node is deleted from the current priority queue; otherwise, the current target node is deleted from the current priority queue; wherein, the initial optimal solution is infinite, and the initial optimal allocation scheme is empty; If the current target node is not a leaf node, select an undetermined variable in the current target node for branching to obtain several child nodes; wherein the variables include: model size and computing requirements; According to the communication data of the current child nodes, the total delay of the child nodes corresponding to each child node is calculated. If the current maximum total delay of the child nodes is less than the minimum total delay corresponding to the current optimal solution, or the current minimum total delay of the child nodes is greater than the maximum total delay corresponding to the current optimal solution, the current target node and the child nodes corresponding to the current target node are deleted from the current priority queue; otherwise, the current child node is added to the current priority queue as a node of the priority queue, and the current target node is deleted at the same time to update the current priority queue.
4. The task scheduling method based on a distributed large language model according to claim 3 is characterized in that: The target device node corresponding to each task to be scheduled is obtained based on the target task to be scheduled, the target task characteristics corresponding to the target task to be scheduled, the real-time monitoring data, and the preset task scheduling model, with the scheduling cost and the sum of the violation of the computing power constraint as the minimum, including: Initialize a list of tasks to be scheduled; wherein the initialized list of tasks to be scheduled includes all the target tasks to be scheduled; Repeat the task scheduling operation until the list of tasks to be scheduled is empty, and determine the target device node corresponding to each task to be scheduled; The task scheduling operation includes: Select a device node as the current node to be scheduled, and obtain the current task to be scheduled from the current list of tasks to be scheduled; Construct the current structure diagram according to the current tasks to be scheduled and the corresponding current target task characteristics; Acquire current real-time monitoring data and the preset task scheduling model, input the current structure diagram, current task characteristics and current real-time monitoring data into the preset task scheduling model, so that the preset task scheduling model calculates the current degree matrix and adjacency matrix according to the current structure diagram; According to the current degree matrix, the adjacency matrix and the current target task characteristics, the current node feature matrix of the last layer in the preset task scheduling model is obtained; According to the current node feature matrix and activation function, the scheduling probability distribution between the current node to be scheduled and each current task to be scheduled is obtained with the goal of minimizing the sum of the scheduling cost and the violation of the computing power constraint; The current node to be scheduled is used as the target device node of the task to be scheduled with the highest scheduling probability, and the node to be scheduled is deleted from the current list to be scheduled, so as to update the list to be scheduled.
5. The task scheduling method based on a distributed large language model according to claim 4 is characterized in that: The current structure diagram is constructed according to the current tasks to be scheduled and the corresponding current target task characteristics, including: The current task to be scheduled is used as the graph node of the current structure graph, and each graph node is connected according to the characteristics of the current target task to obtain the current structure graph.
6. The task scheduling method based on a distributed large language model according to claim 5, characterized in that: When scheduling tasks with the goal of minimizing the sum of scheduling cost and violation of computing capacity constraints, the corresponding second objective function is: Where H represents the second objective function, Indicates the size of the task, Indicates whether the task to be scheduled is from the device node Schedule to device node u l , represents the penalty parameter, K l Represents the computing capacity constraint of the lth level.
7. A task scheduling device based on a distributed large language model, characterized in that: include: Data acquisition module and target device node determination module; The data acquisition module is used to acquire the communication data of each device node in the multi-layer network of the distributed large language model, the tasks to be scheduled, the task characteristics of the tasks to be scheduled, and the real-time monitoring data of each device node; wherein the task characteristics include: task computing power, network conditions, task size and computing power requirements; the real-time monitoring data includes: computing power, memory, bandwidth and latency; The target device node determination module is used to perform model segmentation and task scheduling on the multi-layer network with the goal of minimizing the total delay of the multi-layer network according to the communication data, the tasks to be scheduled, the task characteristics of the tasks to be scheduled, and the real-time monitoring data of each device node, obtain the target device node corresponding to each task to be scheduled, and assign each task to be scheduled to the corresponding target device node for scheduling; The model is segmented according to the communication data and the branch-and-bound integer linear programming algorithm to obtain allocation schemes for each level; the allocation schemes include: model size and computing requirements; Allocate each task to be scheduled to each level according to the allocation scheme to obtain the target task to be scheduled corresponding to each level; According to the target tasks to be scheduled, the target task characteristics corresponding to the target tasks to be scheduled, the real-time monitoring data, and the preset task scheduling model, task scheduling is performed with the goal of minimizing the sum of scheduling cost and violation of computing power constraints, and the target device nodes corresponding to each task to be scheduled are obtained; wherein the sum of scheduling cost and violation of computing power constraints is calculated based on the real-time monitoring data.
8. The task scheduling device based on a distributed large language model according to claim 7, characterized in that: The target device node determination module includes: Priority queue initialization unit and model segmentation unit; The priority queue initialization unit is used to initialize the priority queue; wherein the priority queue includes a number of nodes for representing the allocation scheme; the initialized priority queue includes: a root node that has not been model-segmented; The model segmentation unit is used to repeatedly perform the model segmentation operation until the priority queue is empty, and use the optimal allocation scheme corresponding to the empty priority queue as the allocation scheme; The model segmentation operation includes: Obtain the current priority queue, and select any node from the current priority queue as the target node; wherein the initial node is the root node; If the current target node is a leaf node, the total delay corresponding to the current target node is calculated according to the communication data corresponding to the current target node; wherein the leaf node represents the allocation scheme corresponding to the multi-layer network; When the current total delay is less than the minimum total delay corresponding to the current optimal solution, the minimum total delay is updated according to the current total delay, the allocation scheme corresponding to the current target node is used as the current optimal allocation scheme, and the current target node is deleted from the current priority queue; otherwise, the current target node is deleted from the current priority queue; wherein, the initial optimal solution is infinite, and the initial optimal allocation scheme is empty; If the current target node is not a leaf node, select an undetermined variable in the current target node for branching to obtain several child nodes; wherein the variables include: model size and computing requirements; According to the communication data of the current child nodes, the total delay of the child nodes corresponding to each child node is calculated. If the current maximum total delay of the child nodes is less than the minimum total delay corresponding to the current optimal solution, or the current minimum total delay of the child nodes is greater than the maximum total delay corresponding to the current optimal solution, the current target node and the child nodes corresponding to the current target node are deleted from the current priority queue; otherwise, the current child node is added to the current priority queue as a node of the priority queue, and the current target node is deleted at the same time to update the current priority queue.
9. A terminal device, characterized in that: It comprises a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, a task scheduling method based on a distributed large language model as described in any one of claims 1 to 6 is implemented.
10. A storage medium, characterized in that: The storage medium includes a stored computer program, wherein when the computer program is running, the device where the storage medium is located is controlled to execute the task scheduling method based on a distributed large language model as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Scheduling processing method, device and equipment for deep learning training or reasoning tasks
CN118445045A
Large model reasoning scheduling method based on off-network computing power server
CN119537032A
Internet of things system
US20250071040A1
Intelligent-computing-oriented adaptive adjustment system and method for pipeline-parallel training
WO2024060788A1