Distributed reasoning task allocation method for edge computing large model
By using precise quadratic reconstruction and mathematical model optimization in distributed large-scale model deployment, the problems of limited edge server resources and uncertain resource utilization of inference task are solved, and efficient inference task allocation and maximizing service provider benefits are achieved.
Patent Information
- Application Number
- CN202510320529.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-13
AI Technical Summary
In distributed large-scale deployment, edge server resources are limited, inference task resource utilization is uncertain, and service providers find it difficult to maximize service benefits, especially in scenarios where low latency requirements and heterogeneous user latency requirements are required.
Accurate quadratic reconstruction (ECR) is used to solve the uncertainty of resource occupation of inference tasks, and a mathematical model is constructed to maximize the benefits of service providers, provide theoretically guaranteed approximate solutions through dual methods, and optimize the allocation decisions of inference tasks.
It improves the utilization rate of edge network resources, meets the latency needs of user inference services, and maximizes the service benefits of service providers.
Smart Images

Figure CN120144310A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of mobile edge computing, and particularly to a distributed inference task allocation method for edge computing large models. Background Art
[0002] Due to the long transmission distance and frequent congestion of the backbone network during peak periods, deploying large language models (LLMs) in a central Ized Cloud Server is insufficient to meet users' low-latency requirements. Therefore, service providers have started to deploy LLMs on edge servers, bringing the inference tasks closer to users. Due to the large size of LLMs, their deployment usually spans multiple edge servers. However, after the distributed LLM deployment, the limited resources of edge servers and the nature of resource requirements for LLM inference tasks still pose severe challenges to service providers in allocating users' tasks to appropriate edge servers and obtaining corresponding service benefits. Although there have been some related works on such inference allocation problems, three main key issues have not been well studied.
[0003] First, the distributed deployment of LLMs for inference generates a communication and computing workflow different from traditional distributed task allocation. Second, the uncertainty in the amount of output tokens generated by LLMs is ignored. Finally, only focusing on minimizing latency while ignoring the heterogeneity of users' latency requirements will prevent service providers from obtaining the maximum service benefits. Summary of the Invention
[0004] The present invention optimizes the service provider's inference task allocation decision for the scenarios of distributed large model deployment, uncertain resource occupancy of generative tasks, heterogeneous service benefits, and edge network scenarios with limited resources, thereby maximizing the service provider's benefits.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] Step S1: Considering the distributed deployment scenario of large language models, construct a workflow in which multiple edge servers cooperate to complete inference tasks.
[0007] Step S2: Considering the generative characteristics of inference tasks, use Exact Conic Reformulation (ECR) to address the uncertainty in resource occupancy of inference tasks.
[0008] Step S3: Construct a mathematical model expression for the edge inference task allocation problem in this scenario, with the goal of maximizing the service provider's benefits.
[0009] Step S4: Reconstruct the constructed mathematical model into a combinatorial selection problem, and use the dual method to give an approximately guaranteed solution with theoretical support.
[0010] Preferably, in step S1, a large model is stacked by multiple transformer layers with the same structure. A transformer block includes a multi-head attention (MHA) part and a multi-layer perceptron (MLP) part. In a single transformer layer, let d E represent the dimension of the large model embedding layer, d A represent the multi-head attention (MHA) dimension, and d P represent the MLP layer attention dimension. Then the weight parameter matrix of this layer is
[0011] During the inference process, it needs to go through two stages: the initialization stage and the auto-regressive stage. One operation passing through all transformer layers is called one iteration.
[0012] In the initialization stage, the first output token will be generated through one iteration. Let n in represent the number of input tokens. At this time, the calculation operation of a single-layer transformer is
[0013] X Q = Xw Q , X K = Xw K , X V = Xw V , (1a)
[0014]
[0015] x OUT = f gelu (x MHA w F1 w F2 + X MHA , (1c)
[0016] where, (1a) and (1b) represent the calculations of the MHA part, and (1c) represents the calculations of the MLP part. Let c IA and c IP represent the calculation amounts of the MHA and MLP parts respectively, which can be measured by floating-point operations as:
[0017] c IA = 6n in d A dE +4(n in ) 2 d A +2n in d A d E , (2a)
[0018] c IP =4n in d E d P , (2b)
[0019] Use c I =c IA +c IP to represent the total computational overhead in the initial stage. To accelerate the inference in the autoregressive stage, the KV pairs (X K , X V ) generated in the initial stage need to be saved, and the storage space m I occupied by them is expressed as:
[0020] m I =4n in d A . (3)
[0021] The autoregressive stage runs iteratively. In each iteration, the token generated in the previous iteration is used as the input, and a token is generated as the output. In this case, the computational operation of a single-layer transformer is:
[0022]
[0023] g OUT =f gelu (g MHA w F1F2 +g MHA . (4c)
[0024] Let n out represent the number of output tokens. Then the computational costs c RA and c RP of the corresponding MHA and MLP parts are:
[0025]
[0026] c RP =4(n out -1)d E d P , (5b)
[0027] The total computational cost in the autoregressive stage is c R =c RA +c RP.
[0028] In the autoregressive stage, the generated KV pairs are updated each time, and the final storage space occupied is:
[0029] m R = 4(n out - 1)d A . (6)
[0030] Finally, assuming that the large language model has a total of L layers, the total storage space occupied and computational overhead are m = L(m I + m R ) and c = L(c I + c R ).
[0031] Due to limited resources on edge servers, it is usually impossible to host the deployment of large models with a huge number of parameters. Therefore, it is usually necessary to split the model's weight matrix and deploy it distributively on multiple edge servers. Based on the deployed parameter matrix, each edge server performs MHA and MLP computational operations. However, due to the lack of a complete parameter matrix, the MHA (MLP) calculation results on each edge server are incomplete. Further inference based on incomplete results may lead to incorrect answers. Therefore, a communication pattern called Ring AllReduce is applied after each MHA (MLP) operation to restore the complete results of each edge server.
[0032] Given d as the data size and K as the number of edge servers participating in Ring AllReduce, each Edge server will transmit 2(K - 1) rounds of d / K data size. Since each transmission round is only completed when the slowest server finishes transmitting, the latency of Ring AllReduce is:
[0033]
[0034] where j is the edge server in the transmission ring, j ′ is the next edge server of j in the transmission ring, and l jj’ is the transmission rate between j and j'.
[0035] In step S2, Exact Conic Reformulation (ECR) is adopted to solve the uncertainty of resource occupancy in the inference task.
[0036] Since the inference answers of large models are generative, it is impossible to determine the amount of tokens it generates before the task is completed. And the size of the token amount determines the size of resource occupancy, which will thus affect the allocation decision. Let the task set be The edge server set is Define the probability forms of resource and latency constraints as:
[0037]
[0038] where, ∈ j is the risk level of edge server j, representing the probability of tolerable storage space overload; ∈ i is the risk level of task i, representing the probability of tolerable latency violation for user i.
[0039] Since the distribution of output tokens during large language model inference is user-dependent and usually unknown, this probability constraint is difficult to handle. Considering that the service provider can access the service history data of users, we assume that it can obtain the mean and variance of the number of output tokens of task i. In this case, by adopting exact quadratic reconstruction (ECR), the first constraint can be transformed into the following form:
[0040]
[0041] where:
[0042]
[0043] Similarly, the second constraint is transformed into the following form:
[0044]
[0045] where,
[0046]
[0047] Preferably: In step S3, construct a mathematical model expression for the edge inference task allocation problem in this scenario, with the goal of maximizing the service provider's revenue.
[0048] Suppose there are multiple users in the edge system and edge servers Each server has limited computing resources C j and storage resources D j .
[0049] Each user has an inference task. To simplify the concept, we reuse i to represent the inference task of user . Task i can be specified as where b i represents the service benefit, and represent the number of input tokens and output tokens respectively. Ti This is the latency requirement.
[0050] The entire LLM model is divided into K chunks according to tensor parallelism, denoted as Let the model ratio of the k-th part be At this time, the hidden layers of MHA and MLP in the k-th part can be calculated as and Assume that each edge server j only accommodates one part of the model, denoted as k j . Based on this, we group the edge servers that have deployed the same part, and divide the original edge server set into K subsets
[0051] Let x i indicate whether the inference task i is assigned (x i = 1) or not assigned (x i = 0), and x ij indicate whether the task i is executed by the edge server j (x ij = 1 yes) (x ij = 0 no). If the task i is assigned, the service provider should select one edge server from each subset to recover the complete LLM model for inference. Denote the set of edge servers selected for task i as . The next edge server of j in Ring AllReduce is denoted as R i (j). Then the latency in the initial stage of inference is:
[0052]
[0053] where f j is the computing power provided by edge server j for each task assigned to it. represents the computational overhead of the MHA and MLP parts when executing task i on edge server j in the initial stage. It can be obtained by substituting and in i into formula (2).
[0054] According to formula (3), the resource occupancy of task i on edge server j in the initial stage is:
[0055]
[0056] Similarly, by substituting and into equation (4), the computational overhead of the MHA and MLP parts when executing task i on edge server j in the autoregressive stage can be obtained and The latency of the autoregressive stage is:
[0057]
[0058] According to formula (6), during the autoregressive stage, the storage occupancy of task i on edge server j is:
[0059]
[0060] Before the inference starts, the user needs to transfer the input tokens to the edge server, and after the inference ends, the server needs to send the output tokens back to the user. Therefore, the total user latency is:
[0061]
[0062] where t ij is the latency of transferring one token between user i and server j.
[0063] The total storage occupancy of task i on edge server j is expressed as:
[0064]
[0065] The optimization problem is finally formulated as:
[0066]
[0067] where:
[0068] The goal of optimization is to maximize the service revenue of the service provider. The first constraint indicates that task i must select K servers to perform the inference because the LLM model is split into K parts. In addition, the second constraint shows that the selected edge servers have different model parts, so that the complete LLM model can be restored. The third and fourth constraints ensure that the computing and storage resources are not overused. The fifth constraint ensures that the latency requirement of inference task i is met. The last constraint shows the domain of the decision variables.
[0069] In step S4, the constructed mathematical model is reconstructed into a combinatorial selection problem, and the dual method is used to give an approximately optimal solution with theoretical guarantee.
[0070] From the first and second ones, it is found that x i and x ij The final decision represented is to select K edge servers from all subsets of task i. In addition, only when K edge servers are successfully selected for task i can the revenue of b i be generated in the objective. This means that the optimal number of edge servers for each task should be K or 0.
[0071] In this case, the decision variables (x i , x ij ) can be packed. Let g be a combination of K edge servers, where each edge server comes from a different subset All different combinations form the set G, |G| = QK, k = 1|Jk|. Then, the decision for task i is reduced to whether to select a Denoted by z ig to indicate whether it is selected for task i, the problem can be reformulated as:
[0072]
[0073] where the first constraint ensures that K or 0 edge servers are selected for each task. The third and fourth constraints correspond to the resource limits of each edge server. Denote The value when R i = g. This constraint ensures that g can only be selected when it meets the latency requirement. In other words, selecting g that violates the latency requirement Ti i of task i cannot obtain its benefit bi i . Inspired by this insight, introduce the indicator variable h ig = 1 to indicate that is satisfied, h ig = 0 to indicate that it is not satisfied. Subsequently, we incorporate h i g as a penalty term into the objective function bi i z ig to ensure that selecting any g that exceeds the latency requirement of task i will not generate benefits. Finally, the reformulated problem is:
[0074]
[0075] To solve this multi-constraint problem, the primal-dual theory is adopted. First, relax the domain of z ig to z ig ≥ 0 and introduce the dual variables α i , β j and γ j . The dual problem is constructed as:
[0076]
[0077] According to the dual relaxation theorem, z ig is non-zero only when the first constraint is tight. Therefore, introduce:
[0078]
[0079] The dual relaxation definition can be reformulated as: if z ig = 1, then α i = RH ig . Combining with the constraint that α i must be non - negative, we can define α i as:
[0080]
[0081] where Based on this definition, we design an allocation rule, that is, the inference task i will be allocated according to g* only when α i > 0. After each round of allocation, β j , γ j are updated as follows:
[0082]
[0083] where e j and q j represent the amounts of computing and storage resources occupied on the edge server j respectively. is the lower (upper) limit of β j . and play the same role for γ j .
[0084] The specific process of the algorithm is as follows:
[0085] Step S4 - 1: For task Calculate h according to the exact quadratic reconstruction ig Step S4 - 2: Generate a candidate set by selecting
[0086] that satisfy the latency and resource constraints ig Step S4 - 3: Select the
[0087] with the maximum RH Step S4 - 4: If
[0088] generates a positive net profit, the task i will be allocated according to g*, and the resources occupied on each edge server j ∈ g* will be updated accordingly; otherwise, discard the task;
[0089] Step S4 - 5: Repeat the above process until all tasks have been processed.
[0090] The present invention designs a working mode for collaborative inference between edge servers based on tensor parallelism. In addition, the present invention can allocate appropriate edge servers for processing when the resource occupancy of large model inference tasks is unknown, which helps to improve the resource utilization rate of the edge network and meet the latency requirements of user inference services. Description of the Drawings
[0091] Figure 1 It is the specific flowchart in the embodiment of the present invention;
[0092] Figure 2 It is the schematic diagram of the algorithm flow in the embodiment of the present invention. Detailed Description of the Invention
[0093] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings of the specification.
[0094] Figure 1 It is the flowchart of the reliability-based edge computing offloading and resource allocation method according to the embodiment of the present invention, as Figure 1 The method includes:
[0095] Step S1: Considering the distributed deployment scenario of large language models, construct a workflow for multiple edge servers to collaborate to complete inference tasks.
[0096] A large model is composed of multiple stacked transformer layers with the same structure. A transformer block contains a multi-head attention (MHA) part and a multi-layer perceptron (MLP) part. In a single transformer layer, let d E represent the dimension of the embedding layer of the large model, d A represent the dimension of the multi-head attention (MHA), and d P represent the attention dimension of the multi-layer perceptron (MLP) layer. Then the weight parameter matrix of this layer is
[0097] Let c IA and c IP represent the computational amounts of the MHA and MLP parts respectively, which can be measured by floating-point operations as:
[0098] c IA = 6n in d A d E + 4(n in ) 2 d A + 2n in d A d E ,
[0099] c IP = 4n in d E d P ,
[0100] Use c I = c IA + c IP to represent the total computational overhead in the initial stage. To accelerate the inference in the autoregressive stage, the KV pairs (X K , X V ) generated in the initial stage need to be saved, and the storage space m I occupied by them is expressed as:
[0101] m I = 4n in d A .
[0102] Let n out represent the number of output tokens. Then, the computational amounts c RA and c RP of the MHA and MLP parts corresponding to the autoregressive stage are:
[0103]
[0104] c RP = 4(n out - 1)d E d P , (5b)
[0105] The total computational amount in the autoregressive stage is c R = c RA + c RP .
[0106] In the autoregressive stage, the generated KV pairs are updated each time, and the final occupied storage space is
[0107] m R = 4(n out - 1)d A . (6)
[0108] Finally, assuming that the large language model has a total of L layers, the total storage space occupation and computational overhead are m = L(m I + m R ) and c = L(c I + c R ).
[0109] Due to the limited resources of edge servers, it is usually impossible to host the deployment of large models with a huge number of parameters. Therefore, it is usually necessary to split the model's weight matrices and deploy them distributively on multiple edge servers.
[0110] Therefore, after each MHA (MLP) operation, a communication pattern called Ring AllReduce is applied to recover the complete results of each edge server.
[0111] Given d as the data size and K as the number of edge servers participating in Ring AllReduce, each Edge server will transmit 2(K - 1) rounds of d / K data size. Since each transmission round is only completed when the slowest server finishes transmitting, the latency of Ring AllReduce is
[0112]
[0113] where j is the edge server in the transmission ring, j′ is the next edge server of j in the transmission ring, and l jj’ is the transmission rate between j and j'.
[0114] Step S2 uses Exact Conic Reformulation (ECR) to address the uncertainty of resource occupancy in the inference task. Let the task set be the edge server set be Define the probabilistic forms of resource and latency constraints as
[0115]
[0116] where ∈ j is the risk level of edge server j, representing the probability of storage space overload that can be tolerated; ∈ i is the risk level of task i, representing the probability of latency violation that user i can tolerate.
[0117] Since the distribution of output tokens during large language model inference is user-dependent and usually unknown, this probabilistic constraint is difficult to handle. Considering that the service provider can access the user's service history data, we assume that it can obtain the mean and variance of the number of output tokens of task i. In this case, by using Exact Conic Reformulation (ECR), the first constraint can be transformed into the following form:
[0118]
[0119] where,
[0120] Similarly, the second constraint is transformed into the following form:
[0121]
[0122] Among them,
[0123]
[0124] Step S3 constructs a mathematical model expression for the edge inference task allocation problem in this scenario, with the goal of maximizing the service provider's revenue.
[0125] Suppose there are multiple users in the edge system and edge servers Each server has limited computing resources C j and storage resources D j .
[0126] Each user has an inference task. To simplify the concept, we reuse i to represent the inference task of the user of. Task i can be specified as where b i represents the service benefit, and represent the number of input tokens and output tokens respectively. T i is the latency requirement.
[0127] The entire LLM model is divided into K blocks according to tensor parallelism, denoted as Let the model ratio of the k-th part be At this time, the hidden layers of MHA and MLP in the k-th part can be calculated as and Assume that each edge server j only accommodates one part of the model, denoted as k j . Based on this, we group the edge servers that have deployed the same part into one group, and divide the original edge server set into K subsets
[0128] Let x i indicate whether inference task i is assigned (x i = 1) or not assigned (x i = 0), and x ij indicate whether task i is executed by edge server j (x ij = 1 yes) (x ij = 0 no). If task i is assigned, the service provider should select one edge server from each subset to recover the complete LLM model for inference. Denote as the set of edge servers selected for task i. The next edge server of j in Ring AllReduce is denoted as R i(j). The delay in the initial stage of inference is:
[0129]
[0130] where f j is the computing power provided by edge server j for each task assigned to it. represents the computational overhead of the MHA and MLP parts when executing task i on edge server j in the initial stage. It can be obtained by substituting and in i into formula (2).
[0131] According to formula (3), the resource occupancy of task i on edge server j in the initial stage is:
[0132]
[0133] Similarly, by substituting and into equation (4), the computational overhead of the MHA and MLP parts when executing task i on edge server j in the autoregressive stage can be obtained and Then the delay in the autoregressive stage is:
[0134]
[0135] According to formula (6), in the autoregressive stage, the storage occupancy of task i on edge server j is
[0136] The total user delay is:
[0137]
[0138] where t ij is the delay in transmitting one token between user i and server j. The total storage occupancy of task i on edge server j is expressed as The optimization problem is finally constructed as:
[0139]
[0140] where,
[0141] Step S4 reconstructs the constructed mathematical model into a combinatorial selection problem and uses the dual method to give an approximately optimal solution with theoretical guarantee. Let g be a combination of K edge servers, where each edge server comes from a different subset All different combinations form a set G. |G| = QK k = 1|Jk|. Then, the decision for task i is simplified to whether to select a Denoted by z ig to indicate whether it is selected for task i, the problem can be reformulated as:
[0142]
[0143] Introduce an indicator variable h ig = 1 indicates that is satisfied, h ig = 0 indicates not satisfied. Subsequently, we incorporate h i g as a penalty term into the objective function b i z ig so as to ensure that selecting any g that exceeds the latency requirement of task i will not generate benefits. Finally, the reformulated problem is
[0144]
[0145] To solve this multi-constraint problem, the primal-dual theory is adopted. First, relax the domain of z ig to z ig ≥ 0 and introduce dual variables α i , β j and γ j . The dual problem is constructed as
[0146]
[0147] According to the dual relaxation theorem, z ig is non-zero only when the first constraint is tight. Therefore, introduce:
[0148]
[0149] The dual relaxation definition can be reformulated as: if z ig = 1, then α i = RH ig . Combining the constraint that α i must be non-negative, we can define α i as
[0150]
[0151] where Based on this definition, we design an allocation rule, that is, the inference task i will be allocated according to g* only when α i > 0. After each round of allocation, β j , γ j are updated as follows:
[0152]
[0153] where e j and q j respectively represent the amounts of computing and storage resources occupied on edge server j. is the lower (upper) limit of β j and play the same role for γ j
[0154] Next, execute according to the algorithm shown in Figure 2 The algorithm includes the following processes:
[0155] Step S4-1: For task Calculate h according to exact quadratic reconstruction ig Step S4-2: Generate a candidate set by selecting
[0156] that satisfy the latency and resource constraints ig Step S4-3: Select the one with the maximum RH
[0157] Step S4-4: If generates a positive net benefit, task i will be allocated according to g*, and the resources occupied on each edge server j ∈ g* will be updated accordingly; otherwise, discard the task.
[0158] Step S4-5: Repeat the above process until all tasks have been processed.
[0159] It should be noted that the above content only illustrates the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. For those of ordinary skill in the art of this technology, without departing from the principle of the present invention, several improvements and refinements can still be made, and these improvements and refinements all fall within the protection scope of the claims of the present invention.
Claims
1. A distributed reasoning task allocation method for edge computing large models, characterized by: The method comprises the following steps: Step S1: Considering the distributed deployment scenario of a large language model, a workflow is constructed in which multiple edge servers collaborate to complete the inference task; Step S2: Considering the generative characteristics of the reasoning task, accurate secondary reconstruction is used to solve the uncertainty of the resource occupation of the reasoning task; Step S3: construct a mathematical model to express the edge reasoning task allocation problem in this scenario, with the goal of maximizing the service provider's revenue; Step S4: Reconstruct the constructed mathematical model into a combinatorial selection problem, and use the dual method to give an approximate solution with theoretical guarantee.
2. The distributed reasoning task allocation method for edge computing large models according to claim 1 is characterized in that: In step S1, the workflow of multi-edge server collaborative reasoning is defined: a large model is composed of multiple transformer layers with the same structure. A transformer block contains a multi-head attention part and a multi-layer perceptron part. In a single transformer layer, let d E represents the embedding layer dimension of the large model, d A represents the multi-head attention dimension, d P represents the attention dimension of the multi-layer perceptron layer, then the weight parameter matrix of this layer is The inference process needs to go through two stages: the initialization stage and the autoregression stage. One operation through all transformer layers is called an iteration. The initialization phase will generate the first output token through a round of iteration. Let n in Indicates the number of input tokens. At this time, the calculation operation of a single-layer transformer is: X Q =Xw Q ,X K =Xw K ,X V =Xw V , (la) Among them, formula (1a) and formula (1b) represent the calculation of the MHA part, and formula (1c) represents the calculation of the MLP part. Let c IA and c IP They represent the computational effort of the MHA and MLP parts, respectively, and are measured using floating-point operations: c IA =6n in d A d E +4(n in ) 2 d A +2n in d A d E ,(2a) c IP =4n in d E d P ,(2b) Use c I =c IA +c IP represents the total computational overhead of the initial stage. In order to speed up the reasoning of the autoregressive stage, the KV pairs (X K , X V ) needs to be saved, and the storage space it occupies is m I It is expressed as: m I =4n in d A (3) The autoregression phase runs in an iterative manner, where each iteration takes the token produced by the previous iteration as input and generates a token as output. In this case, the computational operation of a single-layer transformer is: g OUT =f gelu (g MHA w F1 )w F2 +g MHA (4c) Assume n out represents the number of output tokens, then the corresponding MHA and MLP calculation amount c RA and c RP for: c RP =4(n out -1)d E d P ,(5b) The total amount of computation in the autoregressive stage is c R =c RA +c RP ; In the autoregressive phase, each generated KV pair will be updated, and the final storage space occupied is: m R =4(n out -1)d A (6) Finally, assuming that the large language model has a total of L layers, the total storage space occupied and computational overhead is: m=L(m I +m R ) and c=L(c I +c R ); Given d as the data size and K as the number of edge servers participating in Ring AllReduce, each Edge server will transmit 2(K-1) rounds of d / K data size. Since each transmission round is completed only when the slowest server completes the transmission, the latency of Ring AllReduce is: Where j is the edge server in the transmission ring, j ′ is the next edge server of j in the transmission ring, l jj’ is the transmission rate between j and j'.
3. The distributed reasoning task allocation method for edge computing large models according to claim 1 is characterized in that: In step S2, accurate quadratic reconstruction is used to resolve the uncertainty of the resource usage of the reasoning task: Since the reasoning answer of the large model is generative, the amount of tokens it generates cannot be determined before the task is completed. The size of the token determines the size of the resource usage. Suppose the task set is The edge server set is The probabilistic form of defining resource and delay constraints is: Among them, ∈ j is the risk level of edge server j, indicating the tolerable probability of storage space overload; ∈ i is the risk level of task i, indicating the delay violation probability that user i can tolerate, m ij is the storage resource usage of task i on server j, D j is the total storage resource of server j, V i is the total delay of task i, T i is the delay requirement of task i; Since the distribution of output tokens during inference of a large language model is user-dependent and usually unknown, this probability constraint is difficult to handle. Considering that the service provider has access to the user's service history data, it is assumed that it can obtain the mean number of output tokens for task i and variance In this case, using exact quadratic reconstruction, the first constraint can be transformed into the following form: in: Similarly, the second constraint is transformed into the following form: in, They represent the server with the strongest computing power and the server closest to the user. The distribution represents the server with the largest transmission delay in Ring AllReduce and its next server. represents the proportion of large models deployed on server j, f j represents the computing power of server j, t ij represents the unit transmission delay from user i to server j, l jj’ represents the unit transmission delay between servers j and j', R i is the set of edge servers serving task i, Indicates that in the Ring AllReduce ring, The next server, and Server j * The amount of data in the upper MHA layer and MLP layer.
4. The distributed reasoning task allocation method for edge computing large models according to claim 1 is characterized in that: In step S3, a mathematical model is constructed to express the edge reasoning task allocation problem in this scenario, with the goal of maximizing the service provider's revenue: Assume there are multiple users in the edge system and edge servers Each server With limited computing resources C j and storage resources D j , Each user has a reasoning task. To simplify the concept, we reuse i to represent the user. The reasoning task of task i is concretely defined as where b i Expressing service interests, and Represent the number of input tokens and output tokens, T i It is a delay requirement; The entire LLM model is divided into K blocks according to tensor parallelism, expressed as Assume that the model ratio of the kth part is At this point, the hidden layers of the MHA and MLP of the kth part are calculated as and Assume that each edge server j only hosts a part of the model, denoted as k j Based on this, the edge servers that deploy the same parts are grouped together, and the original edge server set Divide into K subsets Let x i Indicates whether the reasoning task i is assigned (x i =1) or not assigned (x i =0), x ij Indicates whether task i is executed by edge server j (x ij =1Yes)(x ij = 0 No), if task i is assigned, the service provider should select Select an edge server in to restore the complete LLM model for reasoning, and use represents the set of edge servers selected for task i, and the next edge server of j in Ring AllReduce is represented by R i (j), then the delay in the initial stage of reasoning is: Among them, f j is the computing power provided by edge server j for each task assigned to it, represents the computational overhead of the MHA and MLP parts when executing task i on edge server j in the initial stage. and Substituting into formula (2) we get; According to formula (3), in the initial stage, the resource usage of task i on edge server j is: Similarly, by and Substituting into equation (4), we can obtain the computational overhead of the MHA and MLP parts when executing task i on edge server j in the autoregressive phase: and Then the delay in the autoregressive phase is: According to formula (6), in the autoregressive stage, the storage usage of task i on edge server j is: Before the inference starts, the user needs to transmit the input token to the edge server. After the inference is completed, the server needs to transmit the output token back to the user. Therefore, the total user delay is: Among them, t ij is the delay of transmitting a token between user i and server j; The total storage usage of task i on edge server j is expressed as The optimization problem is finally constructed as: in, The optimization goal is to maximize the service provider's service revenue. The first constraint indicates that task i must select K servers to perform reasoning because the LLM model is split into K parts. In addition, the second constraint indicates that the selected edge servers have different model parts. The third and fourth constraints ensure that computing and storage resources are not overused. The fifth constraint ensures that the latency requirements of reasoning task i are met. The last constraint indicates the domain of the decision variable.
5. The distributed reasoning task allocation method for edge computing large models according to claim 1 is characterized in that: In step S4, the constructed mathematical model is reconstructed into a combinatorial selection problem, and a theoretically guaranteed approximate solution is given by using the dual method: From the first and second items, we find that x i and x ij The final decision is that in all subsets of task i In addition, only if K edge servers are successfully selected for task i, can b be generated in the target. i This means that the optimal number of edge servers for each task should be K or 0; In this case, the decision variable (x i , x ij ) package, let g be a combination of K edge servers, where each edge server comes from a different subset All the different combinations make up the set Then, the decision of task i is simplified to whether to choose a Use z ig Indicates whether to select for task i, the problem can be restated as: The first constraint ensures that K or 0 edge servers are selected for each task, and the third and fourth constraints correspond to the resource limits of each edge server. express In R i = g, this constraint ensures that g can be selected only when it satisfies the delay requirement, and the delay requirement T of task i is violated. i g cannot obtain its benefits b i , introduce indicator variables h ig =1 means Satisfied, h ig =0 means not satisfied, h i g is included in the objective function b as a penalty term i z ig , thus ensuring that any choice of g that exceeds the delay requirement of task i will not generate any benefit. Finally, the restated problem is: In order to solve this multi-constraint problem, the master-dual theory is adopted. First, z ig The domain is relaxed to z ig ≥0 and introduce the dual variable α i , β j and γ j , the dual problem is constructed as: According to the dual relaxation theorem, z can be ig is non-zero, so we introduce: The dual relaxation definition can be restated as follows: If z ig =1, then α i =RH ig , combined with α i must be non-negative, we can replace α i Defined as: in Based on this definition, a distribution rule is designed, that is, only when α i >0, the reasoning task i will be allocated according to g*. After each round of allocation, β j , γ j The updates are as follows: where e j and q j They represent the amount of computing and storage resources occupied by edge server j, Is β j The lower limit of Is β j The upper limit of and γ j Plays the same role; The specific process of the algorithm is as follows: Step S4-1: For the task According to the exact quadratic reconstruction h ig Step S4-2: Select a To generate candidate sets Step S4-3: Select the one with the maximum RH ig of Step S4-4: If If a positive net benefit is generated, task i will be allocated according to g*, and the resources occupied by each edge server j∈g* will be updated accordingly. Otherwise, the task will be abandoned. Step S4-5: Repeat the above process until all tasks have been processed.
Citation Information
Cited By
Task processing model training and task processing method and device
CN121543768A
Training of task processing model and task processing method and device
CN121543768B
Large model verifiable combinatorial reasoning method and system based on combined certificate lattice and program product
CN122433923A
A method, system and program product for verifiable combinatorial reasoning of large models based on combinatorial certificate lattices
CN122433923B