Large language model inference service optimization method and system, device, and medium

WO2026199616A1PCT designated stage Publication Date: 2026-10-01SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/086421
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-24
Filing Date
2025-03-31
Publication Date
2026-10-01

Smart Images

  • Figure CN2025086421_01102026_PF_FP_ABST
    Figure CN2025086421_01102026_PF_FP_ABST
Patent Text Reader

Abstract

A large language model inference service optimization method and system, a device, and a medium. The method comprises: acquiring a current inference request queue; dynamically sorting inference requests in the current inference request queue; using an SLO-BS algorithm to perform dynamic batch processing on the inference requests in the inference request queue to obtain a current batch of inference requests; and inputting the current batch of inference requests into a large language model. The method, system, device, and medium can solve the problem of video memory fragmentation occurring during concurrent inference of multiple tasks, while also reducing the SLO violation rate.
Need to check novelty before this filing date? Find Prior Art

Description

A method, system, device, and medium for optimizing reasoning services for large language models. Technical Field

[0001] This invention belongs to the field of big data technology and relates to a method, system, device and medium for optimizing large language model reasoning services. Background Technology

[0002] Generative large language models (LLMs), such as ChatGPT, Llama, and ChatGLM, leverage the autoregressive nature of the Transformer architecture. They generate token sequences by predicting them sequentially. During inference, the model dynamically captures historical context information through self-attention and utilizes a key-value cache (KV cache) to store the key-value pairs of generated tokens to accelerate computation. However, as the length *n* of the output sequence increases, the memory usage of the KV cache grows linearly, leading to a significant increase in peak memory consumption for a single inference task. For example, when the Llama-7B model generates 512 tokens, the KV cache consumes up to 4.2GB of GPU memory, accounting for 37% of the total GPU memory consumption. This characteristic makes LLMs face a severe memory bottleneck in long text generation scenarios, especially during multi-task concurrent inference, where memory fragmentation further exacerbates resource waste.

[0003] Existing generative large language model inference optimization techniques have significant drawbacks: traditional deployment schemes employ static GPU (Graphics Processing Unit) memory allocation strategies, which cannot adapt to the dynamic growth characteristics of KV cache in LLM inference, leading to memory fragmentation and increased latency in cross-device data transmission; mainstream schedulers use output sequence length as a single constraint and do not incorporate SLO (Service Level Objective) metrics; at the same time, existing configuration search frameworks rely on full-scale stress testing, resulting in excessively high latency in single decisions, failing to support second-level elastic scaling, and resource contention in hybrid deployment scenarios leading to increased SLO violation rates. Technical issues

[0004] The purpose of this invention is to overcome the shortcomings of the prior art and provide a method, system, device and medium for optimizing large language model inference services. This method, system, device and medium can solve the problem of memory fragmentation that occurs during multi-task concurrent inference and reduce the SLO violation rate. Technical solutions

[0005] To achieve the above objectives, this invention discloses a method for optimizing large language model inference services, comprising:

[0006] Get the current inference request queue;

[0007] The inference requests in the current inference request queue are dynamically sorted.

[0008] The SLO-BS algorithm (Service Level Objective-Output Driven) is used to dynamically batch each inference request in the inference request queue to obtain the current batch of inference requests.

[0009] The inference requests for the current batch are input into the large language model.

[0010] A further improvement of the large language model inference service optimization method described in this invention lies in:

[0011] Furthermore, the process of dynamically sorting the inference requests in the current inference request queue is as follows:

[0012] Extract the length s of the input sequence and the preset SLO delay from each inference request. ;

[0013] The length s of the input sequence is input into the ChatGLM3-6B model to predict the output length corresponding to each inference request. ;

[0014] Based on the output length corresponding to each inference request and preset SLO delay Determine the priority weight W of each inference request;

[0015] The inference requests in the inference request queue are sorted according to their priority weight W.

[0016] Furthermore, the process of using the SLO-BS algorithm to dynamically batch process each inference request in the inference request queue to obtain the current batch of inference requests is as follows:

[0017] In order, traverse the reasoning request queue, and for the j-th reasoning request... Calculate the delay increment and length increment ;

[0018] when If the j-th inference request is added to the current batch of inference requests, then the j-th inference request will be added to the current batch of inference requests.

[0019] Furthermore, the delay increment and length increment They are respectively:

[0020]

[0021] in, For the j-th reasoning request SLO latency; Maximum SLO for the current batch; For the j-th reasoning request The corresponding predicted output length; This indicates the maximum output length for the current batch.

[0022] Furthermore, the process of inputting the current batch of inference requests into the large language model also includes:

[0023] Optimize the deployment of large language models.

[0024] Furthermore, the process of optimizing the deployment of the large language model is as follows:

[0025] Calculate the video memory used by the key-value cache. ;

[0026] Build a GPU cluster topology ;

[0027] Based on the total number of layers in the large language model and the GPU cluster topology. Several deployment schemes for deploying large language models to GPU clusters were identified;

[0028] Based on the video memory occupied by the Key-Value cache Construct the state transition equation;

[0029] Calculate the total delay for each deployment scheme based on the state transition equation;

[0030] The deployment scheme corresponding to the minimum total latency is taken as the optimal deployment scheme, and the large language model is deployed on the GPU cluster according to the optimal deployment scheme.

[0031] Furthermore, the state transition equation is:

[0032]

[0033]

[0034] in, The status is marked as Time allocation to nodes The smallest delay, For binary tags, Represents a node To the node Communication delay, Indicates allocation to a node The number of model layers, For nodes Available video memory, The number of layers in a large language model. For the first The average latency of deploying a layer in a large language model on a node. This represents the average memory usage per layer of a large language model. The status is marked as Time allocation to nodes The minimum delay, It represents 0 or 1.

[0035] This invention discloses a large language model inference service optimization system, comprising:

[0036] The acquisition module is used to acquire the current inference request queue;

[0037] The sorting module is used to dynamically sort the inference requests in the current inference request queue;

[0038] The selection module is used to dynamically batch process each inference request in the inference request queue using the SLO-BS algorithm to obtain the current batch of inference requests.

[0039] The input module is used to input the inference requests of the current batch into the large language model.

[0040] The present invention discloses a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the large language model inference service optimization method.

[0041] This invention discloses a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the large language model inference service optimization method. Beneficial effects

[0042] Compared with the prior art, the present invention has the following beneficial technical effects:

[0043] In practical operation, the large language model inference service optimization method, system, device and medium described in this invention use the SLO-BS algorithm to dynamically batch process each inference request in the inference request queue to obtain the inference request of the current batch. By considering the SLO requirements for scheduling, the SLO violation rate is reduced, and the batch size can be dynamically adjusted to reduce the redundancy of key-value pair cache, thereby improving processing efficiency and solving the problem of memory fragmentation that occurs during multi-task concurrent inference.

[0044] Furthermore, this invention optimizes the deployment of large language models to improve device mapping and enhance resource utilization. Attached Figure Description

[0045] The accompanying drawings, which form part of this application, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an undue limitation of the invention. In the drawings:

[0046] Figure 1 is a flowchart of the method of the present invention;

[0047] Figure 2 is a system structure diagram of the present invention. Embodiments of the present invention

[0048] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0049] In the description of this invention, it should be understood that the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.

[0050] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.

[0051] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. Additionally, the character " / " in this invention generally indicates that the preceding and following objects have an "or" relationship.

[0052] It should be understood that although terms such as first, second, third, etc., may be used in the embodiments of the present invention to describe the preset range, these preset ranges should not be limited to these terms. These terms are only used to distinguish the preset ranges from one another. For example, without departing from the scope of the embodiments of the present invention, the first preset range may also be referred to as the second preset range, and similarly, the second preset range may also be referred to as the first preset range.

[0053] Depending on the context, the word "if" as used here can be interpreted as "when," "when," "in response to determination," or "in response to detection." Similarly, depending on the context, the phrase "if determination" or "if detection (of the stated condition or event)" can be interpreted as "when determination," "in response to determination," "when detection (of the stated condition or event)," or "in response to detection (of the stated condition or event)."

[0054] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0055] The accompanying drawings illustrate various structural schematic diagrams according to embodiments disclosed in this invention. These drawings are not to scale, and some details have been enlarged for clarity, and some details may have been omitted. The shapes of the various regions and layers shown in the drawings, as well as their relative sizes and positional relationships, are merely exemplary and may deviate from reality due to manufacturing tolerances or technical limitations. Furthermore, those skilled in the art can design regions / layers with different shapes, sizes, and relative positions as needed.

[0056] Example 1

[0057] Referring to Figure 1, the large language model inference service optimization method of the present invention includes the following steps:

[0058] 1) Obtain the current inference request queue;

[0059] 2) Dynamically sort the inference requests in the inference request queue;

[0060] Step 2) is as follows:

[0061] 21) Extract request features from each inference request and obtain the SLO;

[0062] The specific operation of step 21) is as follows:

[0063] 211) Extract the length s of the input sequence and the preset SLO delay from the inference request. ;

[0064] 212) Predict the output length corresponding to the extraction request ;

[0065] The length s of the input sequence is input into the ChatGLM3-6B model to predict the output length corresponding to the extraction request. for:

[0066]

[0067] in, Indicates the length of the input sequence. Indicates the first The probability value of each output token. Indicates the previous generation Each token column This represents taking the value of n that maximizes the product.

[0068] 213) Constructing tuples .

[0069] 22) Dynamically sort the inference requests in the inference request queue;

[0070] The specific operation of step 22) is as follows:

[0071] Calculate the priority weight W of each inference request in the inference request queue, where W is:

[0072]

[0073] The inference requests in the sequence are sorted according to their priority weights to ensure that high-priority inference requests are listed first.

[0074] 3) Use the SLO-BS algorithm to dynamically batch process each inference request in the inference request queue;

[0075] Step 3) is as follows:

[0076] 31) Batch initialization;

[0077] Create an empty batch container Initialize parameters, specifically, set the maximum SLO for the current batch: ; Set the maximum output length for the current batch: Set the cumulative delay for each batch: ;

[0078] 32) Calculate the increment;

[0079] Specifically, for the j-th inference request Calculate the delay increment and length increment for:

[0080]

[0081] 33) Merging determination;

[0082] When satisfied When that happens, the j-th inference request will be... Add to the current batch of inference requests and update:

[0083]

[0084] in, For the j-th reasoning request SLO latency, The maximum SLO for the current batch. For the j-th reasoning request The predicted output length, This indicates the maximum output length for the current batch.

[0085] 34) Input the inference requests of the current batch into the large language model.

[0086] In addition, this invention also includes: optimizing the deployment of large language models using the HLR algorithm (High-efficiency Low-Latency Resource), specifically including the following steps:

[0087] 4) Calculate the video memory used by the Key-Value cache. ;

[0088] Assuming the Transformer architecture used in the large language model has 28 layers (l=28) and the hidden layer dimension (h=4096), then the GPU memory occupied by the Key-Value cache... for:

[0089]

[0090] in, The maximum input length for the task. This represents the maximum output length of the task. Indicates the batch size of the task.

[0091] 5) Construct the GPU cluster topology. ;

[0092] GPU cluster topology diagram , express , express N represents the number of nodes in the GPU cluster topology graph. Indicates the first Attribute triples for each GPU node Represents a node arrive Communication delay, , Represents a node The video memory capacity, Represents a node Computational performance Represents a node power consumption, Represents a node arrive The specific delay value.

[0093] 6) Dynamically plan and deploy large language models;

[0094] The specific process of step 6) is as follows:

[0095] 61) Perform state initialization;

[0096] State The status is marked as Time allocation to nodes The smallest delay; The large language model to be deployed is ; The number of layers of the large language model to be deployed is ; Obtain the GPU memory usage of the large language model to be deployed. Average memory usage per layer of a large language model ;No. The average latency for deploying a layer in a large language model on a node is ;

[0097] Based on the total number of layers in the large language model and the GPU cluster topology. Several deployment schemes were identified that allow large language models to be deployed to GPU clusters.

[0098] 62) The state transition equation is constructed as follows:

[0099]

[0100] in, The status is marked as Time allocation to nodes The smallest delay, For binary tags, Represents a node To the node Communication delay, The number of layers in a large language model. For the first The average latency of deploying a layer in a large language model on a node. Indicates allocation to a node The number of model layers, For nodes Available video memory, This represents the average memory usage per layer of a large language model. The status is marked as Time allocation to nodes The minimum delay, It represents 0 or 1.

[0101] 7) Select the optimal path;

[0102] The total latency of each deployment scheme is calculated based on the state transition equation. The deployment scheme with the minimum total latency is taken as the optimal deployment scheme. The large language model is then deployed to the GPU cluster according to the optimal deployment scheme.

[0103] It should be noted that this invention performs resource usage feature analysis on each inference request and collects the SLO requirements and output length information of the inference requests. It performs batch processing scheduling of inference requests based on the SLO-ODBS algorithm, optimizes batch combinations according to the predicted output length, and considers SLO requirements during scheduling to reduce the violation rate. It can also dynamically adjust batch size to improve processing efficiency. Furthermore, this invention employs the HLR algorithm for large language model deployment optimization, allocates computing resources based on hardware topology, performs hierarchical deployment considering the characteristics of large language models, and optimizes device mapping to improve resource utilization. Experiments show that this invention can reduce inference latency by 72.3%-90.3%, improve GPU utilization by 1.2-4.1 times, and increase throughput by 1.92-4.98 times. It achieves a resource utilization improvement of over 30%, a SLO default rate reduction of over 50%, and a system throughput increase of over 40%, fully demonstrating the practical value of this invention in large-scale AI inference service platforms, cloud-based inference service systems, and other scenarios.

[0104] Example 2

[0105] Referring to Figure 2, the large language model inference service optimization system of the present invention includes:

[0106] The acquisition module is used to acquire the current inference request queue;

[0107] The sorting module is used to dynamically sort the inference requests in the current inference request queue;

[0108] The selection module is used to dynamically batch process each inference request in the inference request queue using the SLO-BS algorithm to obtain the current batch of inference requests.

[0109] The input module is used to input the inference requests of the current batch into the large language model.

[0110] In this embodiment, the process of dynamically sorting the inference requests in the current inference request queue is as follows:

[0111] Extract the length s of the input sequence and the preset SLO delay from each inference request. ;

[0112] The length s of the input sequence is input into the ChatGLM3-6B model to predict the output length corresponding to each inference request. ;

[0113] Based on the output length corresponding to each inference request and preset SLO delay Determine the priority weight W of each inference request;

[0114] The inference requests in the inference request queue are sorted according to their priority weight W.

[0115] In this embodiment, the process of using the SLO-BS algorithm to dynamically batch process each inference request in the inference request queue to obtain the current batch of inference requests is as follows:

[0116] In order, traverse the reasoning request queue, and for the j-th reasoning request... Calculate the delay increment and length increment ;

[0117] when If the j-th inference request is added to the current batch of inference requests, then the j-th inference request will be added to the current batch of inference requests.

[0118] In this embodiment, the delay increment and length increment They are respectively:

[0119]

[0120] in, For the j-th reasoning request SLO latency; Maximum SLO for the current batch; For the j-th reasoning request The corresponding predicted output length; This indicates the maximum output length for the current batch.

[0121] In this embodiment, the step of inputting the current batch of inference requests into the large language model further includes:

[0122] Optimize the deployment of large language models.

[0123] In this embodiment, the process of optimizing the deployment of the large language model is as follows:

[0124] Calculate the video memory used by the key-value cache. ;

[0125] Build a GPU cluster topology ;

[0126] Based on the total number of layers in the large language model and the GPU cluster topology. Several deployment schemes for deploying large language models to GPU clusters were identified;

[0127] Based on the video memory occupied by the Key-Value cache Construct the state transition equation;

[0128] Calculate the total delay for each deployment scheme based on the state transition equation;

[0129] The deployment scheme corresponding to the minimum total latency is taken as the optimal deployment scheme, and the large language model is deployed on the GPU cluster according to the optimal deployment scheme.

[0130] In this embodiment, the state transition equation is:

[0131]

[0132]

[0133] in, The status is marked as Time allocation to nodes The smallest delay, For binary tags, Represents a node To the node Communication delay, Indicates allocation to a node The number of model layers, For nodes Available video memory, The number of layers in a large language model. For the first The average latency of deploying a layer in a large language model on a node. This represents the average memory usage per layer of a large language model. The status is marked as Time allocation to nodes The minimum delay, It represents 0 or 1.

[0134] The module division in this embodiment of the invention is illustrative and represents only one logical functional division. In actual implementation, other division methods may be used. Furthermore, the functional modules in the various embodiments of the invention can be integrated into a single processor, exist as separate physical entities, or be integrated into a single module. The integrated modules described above can be implemented in hardware or as software functional modules.

[0135] Example 3

[0136] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the large language model inference service optimization method, including, for example,: obtaining a current inference request queue; dynamically sorting the inference requests in the current inference request queue; using the SLO-BS algorithm to dynamically batch process the inference requests in the inference request queue to obtain the current batch of inference requests; and inputting the current batch of inference requests into a large language model. The memory may include main memory, such as high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device. The processor, network interface, and memory are interconnected via an internal bus, which may be an industry-standard architecture bus, a peripheral component interconnection standard bus, an extended industry-standard architecture bus, etc., and the bus may be divided into an address bus, a data bus, a control bus, etc. The memory is used to store the program; specifically, the program may include program code, which includes computer operation instructions. The memory may include main memory and non-volatile memory, and provides instructions and data to the processor.

[0137] Example 4

[0138] A computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the large language model inference service optimization method. For example, the method includes: obtaining a current inference request queue; dynamically sorting the inference requests in the current inference request queue; using the SLO-BS algorithm to dynamically batch process the inference requests in the inference request queue to obtain a current batch of inference requests; and inputting the current batch of inference requests into a large language model. Specifically, the computer-readable storage medium includes, but is not limited to, volatile memory and / or non-volatile memory. The volatile memory may include random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include read-only memory (ROM), hard disk, flash memory, optical disk, magnetic disk, etc.

[0139] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0140] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in one or more blocks of the flowchart illustrations and / or one or more blocks of the block diagrams.

[0141] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means that implement the functions specified in one or more flowcharts and / or one or more block diagrams.

[0142] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable apparatus, provide steps for implementing the functions specified in one or more flowcharts and / or one or more block diagrams.

[0143] Other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and disclosure of the invention. This invention is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of the invention are indicated by the following claims.

[0144] It should be understood that the present invention is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.

[0145] The above description is merely a preferred embodiment of the present invention and does not constitute any limitation on the present invention. Any simple modifications, alterations, or equivalent structural changes made to the above embodiments based on the technical essence of the present invention shall still fall within the protection scope of the present invention.

Claims

1. A method for optimizing reasoning services using large language models, characterized in that, include: Get the current inference request queue; The inference requests in the current inference request queue are dynamically sorted. The SLO-BS algorithm is used to dynamically batch each inference request in the inference request queue to obtain the current batch of inference requests. The inference requests for the current batch are input into the large language model.

2. The method for optimizing large language model inference services according to claim 1, characterized in that, The process of dynamically sorting the inference requests in the current inference request queue is as follows: Extract the length s of the input sequence and the preset SLO delay from each inference request. ; The length s of the input sequence is input into the ChatGLM3-6B model to predict the output length corresponding to each inference request. ; Based on the output length corresponding to each inference request and preset SLO delay Determine the priority weight W of each inference request; The inference requests in the inference request queue are sorted according to their priority weight W.

3. The method for optimizing large language model inference services according to claim 1, characterized in that, The process of using the SLO-BS algorithm to dynamically batch-process each inference request in the inference request queue to obtain the current batch of inference requests is as follows: In order, traverse the reasoning request queue, and for the j-th reasoning request... Calculate the delay increment and length increment ; when If the j-th inference request is added to the current batch of inference requests, then the j-th inference request will be added to the current batch of inference requests.

4. The method for optimizing large language model inference services according to claim 3, characterized in that, The delay increment and length increment They are respectively: in, For the j-th reasoning request SLO latency; Maximum SLO for the current batch; For the j-th reasoning request The corresponding predicted output length; This indicates the maximum output length for the current batch.

5. The method for optimizing large language model inference services according to claim 1, characterized in that, The process of inputting the current batch of inference requests into the large language model also includes: Optimize the deployment of large language models.

6. The method for optimizing large language model inference services according to claim 5, characterized in that, The process of optimizing the deployment of large language models is as follows: Calculate the video memory used by the key-value cache. ; Build a GPU cluster topology ; Based on the total number of layers in the large language model and the GPU cluster topology. Several deployment schemes for deploying large language models to GPU clusters were identified; Based on the video memory occupied by the Key-Value cache Construct the state transition equation; Calculate the total delay for each deployment scheme based on the state transition equation; The deployment scheme corresponding to the minimum total latency is taken as the optimal deployment scheme, and the large language model is deployed on the GPU cluster according to the optimal deployment scheme.

7. The method for optimizing large language model inference services according to claim 6, characterized in that, The state transition equation is: in, The status is marked as Time allocation to nodes The smallest delay, For binary tags, Represents a node To the node Communication delay, Indicates allocation to a node The number of model layers, For nodes Available video memory, The number of layers in a large language model. For the first The average latency of deploying a layer in a large language model on a node. This represents the average memory usage per layer of a large language model. The status is marked as Time allocation to nodes The minimum delay, It represents 0 or 1.

8. A large language model inference service optimization system, characterized in that, include: The acquisition module is used to acquire the current inference request queue; The sorting module is used to dynamically sort the inference requests in the current inference request queue; The selection module is used to dynamically batch process each inference request in the inference request queue using the SLO-BS algorithm to obtain the current batch of inference requests. The input module is used to input the inference requests of the current batch into the large language model.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the large language model inference service optimization method as described in any one of claims 1-7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the large language model inference service optimization method as described in any one of claims 1-7.