Large language model service request scheduling method and system

By building instruction fine-tuning data sets and establishing response time models, the problem that the service request scheduling method of large language model in the prior art is difficult to accurately predict response time, and efficient resource utilization and user experience improvement are achieved.

CN119960939AActive Publication Date: 2025-05-09SHANDONG UNIV OF FINANCE & ECONOMICS

Patent Information

Application Number
CN202510029675.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-09
Estimated Expiration
2045-01-08

AI Technical Summary

Technical Problem

The existing large language model service request scheduling method is difficult to accurately predict response time, ignoring the complexity and dynamics of the model internally, resulting in inefficient resource allocation and affecting user experience and system performance.

Method used

By constructing a instruction fine-tuning data set, the large language model is fine-tuned by efficient parameter fine-tuning method, predict the response text length, and establish a relationship model between the response time and the response text length, thereby realizing accurate prediction and batch scheduling of the service response time of the large language model.

Benefits of technology

It realizes accurate prediction of service response time of large language model, improves resource utilization, reduces the average request response time, and improves user experience and system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119960939A_ABST
    Figure CN119960939A_ABST
Patent Text Reader

Abstract

The invention provides a large language model service request scheduling method and system, and belongs to the technical field of large language models and service computing. Fine-tuning the large language model by constructing an instruction fine-tuning data set and predicting the length of a response text, and constructing a relation model between the response time of the fine-tuned large language model and the length of the response text to obtain an approximate solution of the response time; performing batch processing on the service request of the large language model and setting a batch scheduling strategy; an error processing mechanism is introduced, the predicted response text length is compared with the response text length output after the large language model receives the service request, and the service response state of the large language model is determined according to the comparison result. According to the method, the task characteristics in the service response process can be fully utilized on the basis of ensuring accurate prediction of the service response time, so that effective and ordered scheduling of the large language model service response is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of large language models and service computing technology, and in particular, relates to a large language model service request scheduling method and system. Background Art

[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.

[0003] Large Language Model (LLM) refers to a deep learning model trained with a large amount of text data, which enables the model to generate natural language text or understand the meaning of language text. Large language models and their derived intelligent services have been widely used in many fields, such as dialogue generation, machine translation, text classification, content creation, etc. In order to ensure service quality and reduce resource costs, such large language model services are usually deployed in cloud environments. However, since the service quality of large language model services is closely related to hardware conditions such as computing power resources, the response time of large language model service requests under limited computing power conditions is prone to fluctuations, which seriously affects the user experience and application performance. At the same time, the existing static batch scheduling system for large language model services often finds it difficult to accurately predict the scale of tasks. It usually adopts a first-come, first-served (FCFS) strategy to process requests in a fixed batch size, which easily leads to inefficient resource allocation, and even problems such as task queue blocking and service interruption, which seriously restricts the large-scale application of large language model services.

[0004] In order to solve these problems, some research works are devoted to improving the efficiency of large language models themselves, such as through model compression, quantization and other technologies to reduce the time and resources required for model inference. However, these methods often require retraining or adjustment of the model, which is costly and difficult to solve problems such as queue blocking caused by concurrent requests. In addition, some works focus on optimizing scheduling algorithms, such as adopting more complex queue management strategies and resource pre-allocation strategies; however, these methods are often difficult to adapt to the complexity and diversity of large language model service tasks, and their prediction accuracy and resource utilization still have a lot of room for improvement. In addition, existing technologies usually lack an accurate evaluation mechanism for request complexity and resource requirements, resulting in inaccurate resource allocation strategies, affecting the efficiency and stability of the service. Therefore, there is an urgent need for a more effective, more accurate and more reliable large model service response time prediction and task scheduling system to improve system resource utilization in cloud environments, reduce the average request response time, and improve user experience.

[0005] It can be seen that the existing large language model service request scheduling method in the cloud environment still has significant defects, and the reason for these significant defects is that there are some technical problems in the method itself, such as:

[0006] (1) Existing large language model service request scheduling methods in cloud environments only consider simple queue models or resource utilization factors, but ignore the internal complexity and dynamics of large language models; however, for different types of request tasks, the attention mechanism and parameter update processes within the large language model will affect the execution time of the task. These factors are generally ignored in existing methods, which leads to insufficient response time prediction accuracy for large language model service response scheduling, which in turn affects the quality of request scheduling.

[0007] (2) Some existing large language model service request scheduling methods in cloud environments take the request input length into consideration, but they usually use simple linear relationships for modeling, which cannot effectively capture the nonlinear relationship between task complexity and response time. However, for some complex natural language inference tasks, the response time may have a nonlinear relationship with the input length. Therefore, the simple linear model constructed by the existing large language model service request scheduling methods in cloud environments still has difficulty in accurately predicting the response time.

[0008] (3) Even if the existing large language model service request scheduling methods or systems in the cloud environment can accurately predict the task response time, they still need an effective scheduling strategy to optimize the execution order of the request tasks to minimize the average response time while improving the system throughput. However, the commonly used static batch scheduling algorithms are often inefficient when processing large language model service request tasks. They cannot effectively cope with the response time fluctuations caused by requests of different sizes, and even cause long processing waits, resulting in low resource utilization and difficulty in dynamic optimization according to the size of different requests, which ultimately affects the overall system performance and user experience. Summary of the invention

[0009] In order to overcome the shortcomings of the above-mentioned prior art, the present invention provides a large language model service request scheduling method and system, which can fully utilize the task characteristics in the service response process on the basis of ensuring accurate prediction of the service response time, so as to achieve effective and orderly scheduling of large language model service responses.

[0010] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions:

[0011] A first aspect of the present invention provides a large language model service request scheduling method.

[0012] A large language model service request scheduling method, comprising:

[0013] Build instruction fine-tuning dataset;

[0014] Fine-tune the data set according to the instructions, and fine-tune the parameters of the large language model using an efficient parameter fine-tuning method; predict the length of the response text based on the fine-tuned large language model;

[0015] Constructing a relationship model between the fine-tuned large language model response time and the length of the obtained response text, and obtaining an approximate solution for the response time of the large language model based on the relationship model;

[0016] Based on the known approximate solution of the response time of the large language model, the service requests of the large language model are processed in batches and a batch scheduling strategy is set; and the service requests are sent to the large language model according to the set batch scheduling strategy;

[0017] An error handling mechanism is introduced to compare the predicted response text length with the actual response text length output by the large language model after receiving the service request and the length of the response text currently generated in each batch, and the service response status of the large language model is determined based on the comparison results.

[0018] Furthermore, the instruction fine-tuning dataset includes a series of data pairs, and the data pairs include instruction requirements, input texts, and corresponding response text length labels.

[0019] Furthermore, an efficient parameter fine-tuning method is used to fine-tune the parameters of the large language model. Specifically, a low-rank adaptive method is used to optimize the parameters in the large language model.

[0020] Furthermore, the response time of the large language model is the total time required for the user to generate a complete response; wherein the total time required for the user to generate a complete response includes the sum of the time from the user submitting the request to the generation of the complete text.

[0021] Furthermore, the batch scheduling strategy is set based on triple conditions. Specifically, the triple conditions include the request length, response length and response time of the service request sent by the user.

[0022] Furthermore, the batch scheduling strategy is deployed in the service scheduler of the large language model, and the service scheduler is provided with a scheduling strategy option; the request expectation is set, and the set request expectation is used as the scheduling strategy option in the service scheduler.

[0023] Furthermore, the service response status of the large language model is determined based on the comparison results, including: setting an identifier upper limit and a text length threshold. When, in the batch service requests, the request output identifier in any batch is greater than the set identifier upper limit, and the difference between the current maximum generated text length in the batch and the minimum generated text length corresponding to the generated request is greater than the set text length threshold, the inference service of the batch is terminated.

[0024] A second aspect of the present invention provides a large language model service request scheduling system.

[0025] A large language model service request scheduling system, comprising:

[0026] The instruction fine-tuning dataset generation module is configured to: construct an instruction fine-tuning dataset;

[0027] The response text length prediction module is configured to: fine-tune the data set according to the instruction, and fine-tune the parameters of the large language model using an efficient parameter fine-tuning method; and predict the response text length based on the fine-tuned large language model;

[0028] A relational model building module is configured to: build a relational model between the fine-tuned large language model response time and the length of the obtained response text, and obtain an approximate solution for the response time of the large language model based on the relational model;

[0029] The batch processing module is configured to: process the service requests of the large language model in batches based on the known approximate solution of the response time of the large language model, and set a batch scheduling strategy; send the service request to the large language model according to the set batch scheduling strategy;

[0030] The service request scheduling module is configured to: introduce an error handling mechanism to compare the predicted response text length with the actual response text length output by the large language model after receiving the service request and the length of the response text currently generated in each batch, and determine the service response status of the large language model based on the comparison results.

[0031] A third aspect of the present invention provides a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the steps in a large language model service request scheduling method as described in the first aspect of the present invention.

[0032] The fourth aspect of the present invention provides an electronic device, including a memory, a processor, and a program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps in the large language model service request scheduling method as described in the first aspect of the present invention are implemented.

[0033] One or more of the above technical solutions have the following beneficial effects:

[0034] (1) The present invention constructs an instruction fine-tuning dataset, and according to the constructed instruction fine-tuning dataset, uses an efficient parameter fine-tuning method to fine-tune the parameters of the large language model; and then predicts the length of the response text based on the fine-tuned large language model. The present invention utilizes the powerful semantic understanding and modeling capabilities of the large language model itself, and can accurately predict the length of the large model output text at a relatively low cost, providing a basis for response time prediction. Therefore, the present invention can better pay attention to the internal complexity and dynamics of the model based on the fine-tuned large language model, thereby achieving accurate prediction of the large language model service response time.

[0035] (2) Based on the preliminary prediction of the response text length of the large language model, the present invention constructs a relationship model between the fine-tuned large language model response time and the obtained response text length, and obtains an approximate solution to the response time of the large language model based on the obtained relationship model. This relationship model can effectively capture the nonlinear relationship between task complexity and response time, thereby providing a basis for service scheduling. Therefore, the relationship model constructed by the present invention can provide further guarantee for predicting the response time.

[0036] (3) Based on the known approximate solution of the response time of the large language model, the present invention processes the service requests of the large language model in batches and sets a batch scheduling strategy; according to the set batch scheduling strategy, the service request is sent to the large language model. This service request scheduling mode can realize efficient batch processing operations and effectively deal with response time fluctuations caused by requests of different sizes. Therefore, the present invention can dynamically optimize according to the scale of different requests, make full use of the task characteristics in the service response process, so as to realize the effective and orderly scheduling of the service response of the large language model; thereby improving the performance of the overall system and user experience.

[0037] Advantages of additional aspects of the present invention will be given in part in the following description, and in part will become obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.

[0039] Figure 1 The present invention is a flowchart of a large language model service request scheduling method in the first embodiment of the present invention.

[0040] Figure 2 This is a schematic diagram of the arrival of a large language model service request and task scheduling in the first embodiment of the present invention.

[0041] Figure 3This is a flowchart of constructing an instruction fine-tuning dataset and fine-tuning a large language model in Embodiment 1 of the present invention.

[0042] Figure 4 This is an example diagram of data samples in Embodiment 1 of the present invention.

[0043] Figure 5 This is a flowchart of large language model service response time prediction in Embodiment 1 of the present invention.

[0044] Figure 6 This is a flowchart of the iterative algorithm in Embodiment 1 of the present invention.

[0045] Figure 7 This is an overall architecture diagram of a large language model service request scheduling system in a cloud environment in Embodiment 2 of the present invention. DETAILED DESCRIPTION

[0046] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.

[0047] It should be noted that the terms used herein are for describing specific embodiments only and are not intended to be limiting of exemplary embodiments according to the present invention.

[0048] In the absence of conflict, the embodiments of the present invention and the features of the embodiments may be combined with each other.

[0049] The overall idea proposed by the present invention: The present invention provides a method for scheduling service requests of a large language model. First, the large language model is fine-tuned by means of instruction fine-tuning data sets, and the length of the output text of the large language model is accurately predicted at a relatively low cost by utilizing the powerful semantic understanding and modeling capabilities of the large language model itself; then, an approximate solution of the service response time of the large language model is predicted by establishing a relationship model between the response time and the output length. On this basis, requests with similar request sizes are batched, and request expectations are introduced as a request scheduling strategy option for the service scheduler; the request scheduling strategy can effectively reduce the average response time and prevent batches with longer service times from waiting in the queue for a long time. At the same time, the present invention designs an error handling mechanism for timely handling prediction deviations that occur in the system, thereby effectively ensuring the service quality of the large language model. The method provided by the present invention does not conflict with other existing inference acceleration technologies, and can effectively supplement the existing large language model inference acceleration tool library to further improve the overall performance of large language model services.

[0050] Embodiment 1

[0051] This embodiment discloses a large language model service request scheduling method.

[0052] like Figure 1 As shown, a large language model service request scheduling method includes:

[0053] Step S1, constructing an instruction fine-tuning data set;

[0054] Step S2, fine-tuning the data set according to the instruction, and fine-tuning the parameters of the large language model using an efficient parameter fine-tuning method; predicting the length of the response text based on the fine-tuned large language model;

[0055] Step S3, constructing a relationship model between the fine-tuned large language model response time and the length of the obtained response text, and obtaining an approximate solution for the response time of the large language model based on the relationship model;

[0056] Step S4: based on the known approximate solution of the response time of the large language model, the service requests of the large language model are processed in batches, and a batch scheduling strategy is set; according to the set batch scheduling strategy, the service request is sent to the large language model;

[0057] Step S5: Introduce an error handling mechanism to compare the predicted response text length with the actual response text length output by the large language model after receiving the service request and the length of the response text currently generated in each batch, and determine the service response status of the large language model based on the comparison results.

[0058] Based on the above steps, the present invention can fully utilize the task characteristics in the service response process on the basis of ensuring accurate prediction of the service response time, so as to achieve effective and orderly scheduling of the service response of the large language model. To facilitate the understanding of the technical solution of the present invention, the specific implementation steps of the technical solution of the present invention are further explained and illustrated below.

[0059] It should be noted that the solution provided in this embodiment is implemented on a cloud server cluster as the background platform. The cloud platform supports the deployment of multiple large language model services such as ChatGLM3-6B, Qwen1.5-7B, and LLaMA2-7B. At the same time, the cloud platform is an experimental environment composed of 6 cloud servers, each of which is equipped with an NVIDIA GeForce RTX 3090 GPU and 24GB of video memory, and supports on-demand elastic expansion. Figure 2 As shown, the large language model in the cloud environment (i.e. Figure 2The large language model service described in the specification intelligently batches the concurrent requests submitted by users and optimizes the order of task execution through an efficient batch scheduling strategy; and dynamically adjusts the batch size to adapt to real-time load changes, ultimately improving the service quality of the large language model, reducing the average response time, and ensuring the stability of the large language model service. This embodiment performs the following operations in the above cloud environment.

[0060] Step S1: construct an instruction fine-tuning dataset.

[0061] According to the goal of the large language model response text length, a fine-tuning dataset is constructed; wherein the instruction fine-tuning dataset contains a series of data pairs, which include instruction requirements, input text, and corresponding response text length labels.

[0062] Furthermore, if Figure 3 As shown, the instruction fine-tuning dataset used for instruction fine-tuning the large language model in this embodiment includes two parts; one part is based on 15,000 samples after the public dataset Alpaca is modified; the other part is 5,000 samples constructed based on seed samples using the generation ability of the large language model. Specifically, the seed samples are simulated samples defined based on real data samples that can reflect the characteristics and distribution of real data, and are basic and representative.

[0063] Furthermore, the transformation of the Alpaca dataset includes cleaning and normalizing the instructions and performing quality control on the "input-output" to improve the quality and consistency of the instruction fine-tuning dataset. The selection of seed samples is based on considerations of instruction diversity and coverage, and strives to cover different types of instructions and tasks. Specifically: by deforming, expanding, and combining seed samples, the generation capabilities of the large language model are used to construct and generate them to ensure the diversity and scale of the dataset; at the same time, for the generated samples, in order to effectively evaluate the various possibilities of the large language model's response text length, this embodiment chooses to sample each instruction 5 times, and uses the maximum length of the 5 samples as the target length of each instruction, and uses the target length as the length of the text output by the large language model, that is:

[0064] L = max{L1,L2,L3,L4,L5};

[0065] Where L represents the length of the output text of the large language model. Each sample data contains the instruction requirement, input text, and the corresponding response text length label. The dataset covers texts of different lengths, different styles, and different fields, thus ensuring the generalization ability of the model.

[0066] like Figure 4As shown in Figure 1, the dataset contains sample data in the format of "instruction-input-output"; the instruction is the task to be performed by the large language model, the input is the user request, and the output is the length of the response text predicted by the large language model. During instruction supervision fine-tuning, the content corresponding to the instruction will be concatenated with the content corresponding to the input as the human instruction, and the content corresponding to the output is the response answer of the large language model.

[0067] Step S2: fine-tune the parameters of the large language model according to the established instruction fine-tuning data set by using an efficient parameter fine-tuning method; and predict the length of the response text based on the fine-tuned large language model.

[0068] Based on the constructed fine-tuning dataset, the large language model is fine-tuned, and the parameters of the large language model are optimized by using the Low-Rank Adaptation (LoRA) method to improve the accuracy of the large language model in predicting the length of the response text. Figure 3 As shown, this can be achieved through the following steps:

[0069] Step S2-1, load the pre-trained large language model. In order to reduce the amount of training parameters and speed up the training speed, this embodiment uses the LoRA fine-tuning method based on LLaMA Factory; wherein, LLaMA Factory is a large language model factory, which is an open source fine-tuning framework that provides developers with a simple and efficient tool to quickly adapt to specific task requirements and improve model performance based on the pre-trained model. This fine-tuning method is to add a small number of trainable parameters on the basis of the pre-trained model, rather than directly fine-tuning all the parameters of the large model. This can significantly reduce computing resource consumption and memory usage, so that large language models can be efficiently fine-tuned in resource-constrained environments and quickly adapted to specific tasks. Specifically, the calculation formula of the LoRA fine-tuning method is as follows, namely:

[0070] h=W0x+ΔWx;

[0071] Among them, h represents the output of the model, and x represents the input of the model. W0 represents the pre-trained weight matrix, and ΔW represents the additional matrix added in the bypass, which can be further refined as ΔW=BA. During initialization, matrix A is initialized with a random Gaussian, and matrix B is initialized to 0; this ensures that only the left backbone is effective in the initial stage. During training, all pre-trained weight matrices are frozen, and only the ΔW part is iteratively optimized.

[0072] Step S2-2: Before fine-tuning the instruction, set the necessary parameters required for fine-tuning; wherein the necessary parameters required for fine-tuning include learning rate, number of training rounds and calculation type, etc. Specifically, the setting of the necessary parameters is shown in Table 1:

[0073] Table 1 Necessary parameters for fine-tuning

[0074] Parameter Type Parameter Value Learning Rate 5e-5 Number of training rounds 3 Calculation Type bf16 Batch size 2 Validation set ratio 0.1 Learning rate regulator Cosine

[0075] Among them, the learning rate regulator Cosine, namely the cosine annealing learning rate regulator, is a learning rate adjustment strategy based on the cosine function; the calculation type bf16 is an efficient data format with short calculation time and less storage requirements. After the parameter setting is completed, the instruction fine-tuning dataset constructed in step S1 is input into the large language model for instruction fine-tuning.

[0076] Step S2-3: During instruction fine-tuning, the large language model predicts the length of the response text based on the input text.

[0077] Use the back propagation algorithm to update the trainable parameters introduced by the LoRA method; at the same time, use the AdamW optimizer to adjust the learning rate and other hyperparameters to continuously optimize the performance of the large language model. Among them, back propagation starts from the loss value of the output layer, and uses the chain rule to back propagate the gradient of the loss to the parameters of each layer, so as to update the model parameters according to the gradient and reduce the loss. The core is to guide the adjustment of model parameters by calculating the gradient; AdamW optimizer is a commonly used parameter optimization algorithm. The specific adjustment and calculation process has been encapsulated as a method for direct call. In principle, it adjusts the learning rate and other hyperparameters by combining the adaptive learning rate adjustment strategy and setting the hyperparameters such as weight decay according to certain rules. Evaluate the large language model regularly, and adjust the hyperparameters and data preprocessing strategies according to the evaluation results until the performance of the large language model reaches the expected effect; after the model training is completed, merge the trained LoRA-related weight files with the basic model weights to export and generate a new model; save the fine-tuned large language model. The saved large language model can be deployed, and after deployment, it can be used to predict the response text length of the large language model.

[0078] This embodiment fine-tunes the large language model and uses the large language model's powerful semantic understanding and modeling capabilities to accurately predict the length of the large model output text at a relatively low cost, providing a basis for response time prediction. By establishing a relationship model between response time and response text length, the response time of the large language model service can be obtained, thereby providing a basis for service scheduling.

[0079] Step S3: construct a relationship model between the fine-tuned large language model response time and the length of the obtained response text, and obtain an approximate solution for the response time of the large language model based on the relationship model.

[0080] The response time of the large language model is the total time required for the user to generate a complete response; the total time required for the user to generate a complete response includes the sum of the time from the user submitting the request to the generation of the complete text. By establishing the relationship between the response time of the large language model and the length of the resulting response text, an approximate solution for the response time of the large language model can be obtained. Figure 5 As shown, this can be achieved through the following steps:

[0081] Step S3-1: Collect sample data of large language model request services. Each sample contains the length of user input text, model response text length, CPU / GPU utilization, memory utilization and other server operating resource conditions; the sample label is the response time of the large model. The response time is the time from the user submitting the request to the model returning the complete text.

[0082] Step S3-2: Clean and preprocess the collected data, such as processing outliers and missing values, and perform standardization.

[0083] Step S3-3, respectively construct decision tree regression, support vector regression, XGBoost regression, LightGBM regression and KNN regression models to train the preprocessed data to obtain multiple independent prediction models. In this process, the 5-fold cross-validation method is used to avoid overfitting and improve the generalization ability of the model; the model parameters are tuned by the grid search method until the best hyperparameter combination is obtained to maximize the prediction accuracy of the large language model.

[0084] Step S3-4: Evaluate the large language model. Use an independent test data set to evaluate the prediction accuracy of the large language model trained above. Specifically, through the analysis of experimental results, select three models with higher performance as the prediction basic models; at the same time, the average prediction value based on the above three basic models is used as the final response time prediction result to improve the stability and accuracy of the prediction.

[0085] The performance evaluation results of the five cross-validations are averaged to obtain the final evaluation result. Specifically, this embodiment selects the root mean square error RMSE as the model performance evaluation indicator. The evaluation results of the model are shown in Table 2:

[0086] Table 2 Evaluation results of large language model

[0087] Model Root mean square error RMSE Decision Tree Regression 3.32 Support Vector Regression 3.28 KNN Regression 2.37 XGBoost Regression 1.66 LightGBM Regression 1.31

[0088] Based on the above results, this embodiment selects KNN regression, XGBoost regression and LightGBM regression with better performance as the prediction basic models; at the same time, the average prediction value based on the above three basic models is used as the final response time prediction result to further improve the accuracy and stability of the large language model prediction. The integrated large language model will be used as the final prediction model to predict new input data.

[0089] Step S4: based on the known approximate solution of the response time of the large language model, the service requests of the large language model are processed in batches, and a batch scheduling strategy is set; and the service requests are sent to the large language model according to the set batch scheduling strategy.

[0090] After obtaining an approximate solution to the response time, the large language model service requests are batched, and the batch scheduling strategy is set based on three conditions, where the three conditions include the request length, response length, and response time of the service request sent by the user. The batch scheduling strategy is deployed in the service scheduler of the large language model, and the service scheduler is provided with a scheduling strategy option. Then, the request expectation is set, and the set request expectation is used as the scheduling strategy option in the service scheduler. The batch scheduling strategy can be implemented by the following steps:

[0091] Step S4-1: Approximate the response time T according to the prediction obtained in the above steps r , batch the requests entering the service queue. In this embodiment, the specific batching strategy adopts the K-Means clustering method, specifically: comprehensively consider the request length L r , expected response length L output And the predicted response time approximate solution T r , construct a three-dimensional feature vector (L r ,L output ,T r ); Use the K-Means algorithm to divide the requests into K groups. The batch size K can be dynamically adjusted based on historical data analysis and system load conditions; the number of requests in a group B k It can be determined dynamically based on the clustering results. In order to avoid excessive number of requests in some groups and affect scheduling efficiency, the maximum number of requests N in the group can be set to control the maximum size of the scheduling batch. The maximum number of requests N is set by the maximum throughput of the system.

[0092] Step S4-2: Introduce request expectation E i As a scheduling policy option. Request expectation E i The predicted response time T of the requests in the group r and the request waiting time W r The weighted calculation is used to determine:

[0093]

[0094] Among them, α and β represent weight coefficients, satisfying 0≤α≤1, 0≤β≤1, and α+β=1. The weight can be adjusted dynamically according to the actual situation; the higher the request expectation, the higher the scheduling priority, that is, the system gives priority to batches with higher request expectations.

[0095] Step S4-3: Initialize the batch size B of the large language model o ; At the same time, the requests in each group are packaged into batches and processed in batches. i To set the scheduling priority; according to the scheduling priority of each batch, the service scheduler selects batches for processing in priority order. When the LLM (Large Language Model) instance completes the batch service and becomes idle, the service scheduler continues to select a target batch from the waiting queue according to the scheduling priority and then schedules it to the LLM instance.

[0096] Step S4-4, system evaluation. The performance of the scheduling system is evaluated through indicators such as average response time and average waiting time. Specifically, by monitoring and analyzing these indicators, the service quality of large language model service request scheduling can be effectively evaluated, and then the batch strategy and scheduling algorithm can be continuously optimized to achieve the best system performance.

[0097] Through the above process, this embodiment constructs an efficient batch scheduling method. By batching requests with similar request sizes and introducing request expectations as a scheduling strategy option for the service scheduler, it can comprehensively consider the request processing time and waiting time. On the one hand, it prevents batches with longer service times from waiting in the queue for a long time, and on the other hand, it effectively reduces the average response time of requests and improves the overall service quality of the large language model.

[0098] Step S5: Introduce an error handling mechanism to compare the predicted response text length with the actual response text length output by the large language model after receiving the service request and the length of the response text currently generated in each batch, and determine the service response status of the large language model based on the comparison results.

[0099] Although after fine-tuning the large language model through instructions and establishing a relationship model between response time and response text length, the accuracy of the system's prediction of response text length and response time can be greatly improved; however, prediction deviations still occur. In this regard, this embodiment introduces an error handling mechanism to capture abnormal situations and optimize them in a timely manner, thereby improving the robustness of the system. By continuously monitoring the batch response status of the large language model service; once an abnormal deviation is detected, the system will automatically trigger the error handling mechanism, collect the request processing anomalies that occur in the system in real time, and initiate corresponding optimization measures. Through the error handling mechanism, it is possible to capture abnormal situations and optimize them in a timely manner, thereby improving the robustness of the system.

[0100] In the actual implementation process, the system monitors the predicted response text length L in real time. output , Actual response text length The length of the text currently generated in the batch is L MAX The difference between them. The service response status of the large language model is determined according to the comparison results presented by the difference, that is: the identifier upper limit and the text length threshold are set. When, in the service requests under batches, the request output identifier in any batch is greater than the set identifier upper limit, and the difference between the current maximum generated text length in the batch and the minimum generated text length corresponding to the generated request is greater than the set text length threshold, the inference service of the batch is terminated. Preferably, the identifier upper limit and the text length threshold are set to δ and D respectively; when there are more than δ requests in the batch output the [end] identifier, and the current maximum generated text length L MAX The minimum length L of generated text for the δ requests that have been generated in the batch δ When the difference is greater than the text length threshold D, the system determines it as a prediction error, terminates the reasoning service of the batch, and marks the request set R that has not completed the reasoning F .

[0101] Furthermore, for abnormal requests, they are collected in real time and packaged into batches; through iterative processing, a new abnormal request batch R is formed. F(r) When the proportion of completed requests x in the same batch exceeds P, the reasoning service of the batch is terminated and the set of requests R that have not completed reasoning is marked F(r-x) Where P represents the threshold that can be set in the exception handling mechanism.

[0102] In order to continuously improve the service quality of the large language model, this embodiment continuously optimizes the parameters and models in the above steps by formulating an iterative optimization mechanism. According to the server load and actual response time, the batch size, batch strategy, weight coefficient and other parameters are dynamically adjusted to achieve the best scheduling effect. The above process is iterated until all requests are completed. Moreover, after each iteration, the system will generate a request set R F(x) Feedback to users; specifically, the implementation process of the iterative algorithm is as follows Figure 6 As shown. Due to the inherent response length limit of the large language model service, when the maximum response length is exceeded, the request will be automatically completed to avoid infinite iterations. Through the above error handling mechanism, the present invention can effectively avoid the scheduling cost caused by the response text length prediction error, and even in the face of complex requests and abnormal response text length, it can ensure the stable operation of the system and efficient service quality. In addition, the present invention does not conflict with other existing inference acceleration technologies, and can effectively supplement the existing large language model inference acceleration tool library, and can effectively enhance the functions and performance of the existing large language model inference acceleration toolkit, and further improve the overall performance of the large model service; this complementarity can provide new directions and possibilities for future research and development in related fields.

[0103] Furthermore, in order to verify the comprehensive effect of the large language model service request scheduling method provided by the present invention, the present embodiment conducted the following two verification experiments:

[0104] Experiment 1: Verify the performance of the fine-tuned large language model in predicting the length of the response text. Specifically, the comparison models selected in this embodiment include the unfine-tuned large language model and other prediction models; for the unfine-tuned large language model: directly verify the effectiveness of fine-tuning by comparing the performance difference between the fine-tuned model and the unfine-tuned same benchmark large language model in predicting the length of the response text. For other prediction models: comparing the performance of the fine-tuned model with other models for predicting text length helps analyze the relative advantages and disadvantages of different prediction models in solving the same task; the selected models include decision trees, Bi-LSTM, Bi-LSTM+attention, etc.

[0105] This embodiment constructs and enhances features based on user prompts, and realizes text length prediction by mining the relationship between user input and response text length. In the comparative experiment, Acc±25 and Acc±50 are selected as evaluation indicators, which represent the accuracy of the prediction results within the range of 25 and 50 above and below the label value, respectively. The experimental results are shown in Table 3; the experiment uses ChatGLM3-6B as the benchmark large language model.

[0106] Table 3 Comparison results of experiment 1

[0107] Model Acc±25 Acc±50 Untuned ChatGLM3-6B 31% 52% Decision Tree 11% 19% Bi-LSTM 16% 34% Bi-LSTM+attention 21% 38% ChatGLM3-6B after fine-tuning 43% 66%

[0108] The experimental results show that the performance of the fine-tuned ChatGLM3-6B is significantly better than that of the unfine-tuned ChatGLM3-6B; the Acc±25 and Acc±50 are improved by 12% and 14% respectively, indicating that fine-tuning effectively improves the prediction accuracy of the model. In addition, compared with the decision tree, Bi-LSTM, and Bi-LSTM+attention, the fine-tuned ChatGLM3-6B and the unfine-tuned ChatGLM3-6B, the large language model performs better in terms of Acc±25 and Acc±50 indicators, indicating that its accuracy in predicting the length of the response text is significantly higher than these traditional models.

[0109] Experiment 2 verifies the performance of the large model service scheduling method and error handling mechanism proposed in the present invention. Before verification, the following experimental settings are performed in this embodiment:

[0110] (1) An experimental environment was built based on six cloud servers in the cloud platform, two of which were used to deploy a fine-tuned large model and a service request scheduler to ensure request processing capabilities.

[0111] (2) Design a request generator to simulate the arrival of concurrent requests from different services in the experiment. The request generator can be dynamically adjusted according to the set parameters (such as the number of requests, request interval, request type, etc.) to simulate the load conditions in real scenarios; the request types cover a variety of scenarios such as text conversations, comparative analysis, and translation.

[0112] (3) Set experimental evaluation indicators, mainly including: Average Response Time (ART), which is the average time from submitting a request to returning the result; Average Wait Time (AWT), which is the average time from submitting a request to the start of request processing.

[0113] (4) Selection of comparison methods. This embodiment selects commonly used static batch scheduling algorithms as benchmarks, including a simple first-come-first-served scheduling method, a shortest job first method, and a priority scheduling method, so as to better evaluate the relative performance advantages of the scheduling methods proposed in this embodiment. The first-come-first-served scheduling method focuses on the arrival order of requests; the shortest job first method focuses on the size of requests; and the priority scheduling method focuses on the priority of requests, and the priority is determined by the request expectation of a single request.

[0114] Furthermore, this embodiment selected 3,000 sample data sets generated based on the public data set Alpaca data set transformation and seed samples as the validation set, covering requests of different lengths and types. The experimental results are shown in Table 4:

[0115] Table 4 Comparison results of Experiment 2

[0116]

[0117]

[0118] From the analysis of the above experimental results, it can be seen that the scheduling method and error handling mechanism of the present invention significantly reduce the average response time and the average waiting time, indicating that by effectively constructing and scheduling request batches, the overall efficiency and service quality of large model services can be improved, which provides strong support for further promoting the feasibility of this method in practical applications.

[0119] Embodiment 2

[0120] This embodiment discloses a large language model service request scheduling system.

[0121] A large language model service request scheduling system, comprising:

[0122] The instruction fine-tuning dataset generation module is configured to: construct an instruction fine-tuning dataset;

[0123] The response text length prediction module is configured to: fine-tune the data set according to the instruction, and fine-tune the parameters of the large language model using an efficient parameter fine-tuning method; and predict the response text length based on the fine-tuned large language model;

[0124] A relational model building module is configured to: build a relational model between the fine-tuned large language model response time and the length of the obtained response text, and obtain an approximate solution for the response time of the large language model based on the relational model;

[0125] The batch processing module is configured to: process the service requests of the large language model in batches based on the known approximate solution of the response time of the large language model, and set a batch scheduling strategy; send the service request to the large language model according to the set batch scheduling strategy;

[0126] The service request scheduling module is configured to: introduce an error handling mechanism to compare the predicted response text length with the actual response text length output by the large language model after receiving the service request and the length of the response text currently generated in each batch, and determine the service response status of the large language model based on the comparison results.

[0127] like Figure 7As shown, a large language model service request scheduling system provided by the present invention is placed in the control chain layer of the entire cloud environment to control the large language model service scheduling strategy; the control chain layer interacts with the application layer and the application access layer in turn upward; and interacts with the cloud service cluster layer downward. Among them, the user-visible layer is the top layer, and the user-visible layer is the application access layer.

[0128] Embodiment 3

[0129] The purpose of this embodiment is to provide a computer-readable storage medium.

[0130] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps in a large language model service request scheduling method as described in Embodiment 1 of the present disclosure.

[0131] Embodiment 4

[0132] The purpose of this embodiment is to provide an electronic device.

[0133] An electronic device includes a memory, a processor, and a program stored in the memory and executable on the processor. When the processor executes the program, the steps in a large language model service request scheduling method as described in Embodiment 1 of the present disclosure are implemented.

[0134] The steps involved in the apparatuses of the above embodiments 2, 3 and 4 correspond to the method embodiment 1, and the specific implementation methods can refer to the relevant description part of embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood to include any medium that can store, encode or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.

[0135] Those skilled in the art should understand that the modules or steps of the present invention described above can be implemented by a general-purpose computer device, or alternatively, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.

[0136] Although the above describes the specific implementation mode of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without creative work are still within the scope of protection of the present invention.

Claims

1. A large language model service request scheduling method, characterized in that: include: Build instruction fine-tuning dataset; Fine-tune the data set according to the instructions, and use an efficient parameter fine-tuning method to fine-tune the parameters of the large language model; Predict the length of the response text based on the fine-tuned large language model; Constructing a relationship model between the fine-tuned large language model response time and the length of the obtained response text, and obtaining an approximate solution for the response time of the large language model based on the relationship model; Based on the known approximate solution of the response time of the large language model, the service requests of the large language model are processed in batches and a batch scheduling strategy is set; and the service requests are sent to the large language model according to the set batch scheduling strategy; An error handling mechanism is introduced to compare the predicted response text length with the actual response text length output by the large language model after receiving the service request and the length of the response text currently generated in each batch, and the service response status of the large language model is determined based on the comparison results.

2. A large language model service request scheduling method as claimed in claim 1, characterized in that: The instruction fine-tuning dataset includes a series of data pairs, wherein the data pairs include instruction requirements, input texts, and corresponding response text length labels.

3. A large language model service request scheduling method as claimed in claim 1, characterized in that: An efficient parameter fine-tuning method is used to fine-tune the parameters of the large language model. Specifically, a low-rank adaptive method is used to optimize the parameters in the large language model.

4. A large language model service request scheduling method as claimed in claim 1, characterized in that: The response time of the large language model is the total time required for the user to generate a complete response; the total time required for the user to generate a complete response includes the sum of the time from the user submitting the request to the generation of the complete text.

5. A large language model service request scheduling method as claimed in claim 1, characterized in that: The batch scheduling strategy is set based on three conditions. Specifically, the three conditions include the request length, response length and response time of the service request sent by the user.

6. A large language model service request scheduling method as claimed in claim 5, characterized in that: The batch scheduling strategy is deployed in the service scheduler of the large language model, and the service scheduler is provided with a scheduling strategy option; the request expectation is set, and the set request expectation is used as the scheduling strategy option in the service scheduler.

7. A large language model service request scheduling method as claimed in claim 1, characterized in that: The service response status of the large language model is determined based on the comparison results, including: setting an identifier upper limit and a text length threshold. When, in the batch service requests, the request output identifier in any batch is greater than the set identifier upper limit, and the difference between the current maximum generated text length in the batch and the minimum generated text length corresponding to the generated request is greater than the set text length threshold, the inference service of the batch is terminated.

8. A large language model service request scheduling system, characterized in that: include: The instruction fine-tuning dataset generation module is configured to: construct an instruction fine-tuning dataset; The response text length prediction module is configured to: fine-tune the data set according to the instruction, and fine-tune the parameters of the large language model using an efficient parameter fine-tuning method; and predict the response text length based on the fine-tuned large language model; A relational model building module is configured to: build a relational model between the fine-tuned large language model response time and the length of the obtained response text, and obtain an approximate solution for the response time of the large language model based on the relational model; The batch processing module is configured to: process the service requests of the large language model in batches based on the known approximate solution of the response time of the large language model, and set a batch scheduling strategy; send the service request to the large language model according to the set batch scheduling strategy; The service request scheduling module is configured to: introduce an error handling mechanism to compare the predicted response text length with the actual response text length output by the large language model after receiving the service request and the length of the response text currently generated in each batch, and determine the service response status of the large language model based on the comparison results.

9. A computer-readable storage medium having a program stored thereon, characterized in that: When the program is executed by a processor, the steps in the large language model service request scheduling method as described in any one of claims 1-7 are implemented.

10. An electronic device comprising a memory, a processor, and a program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps in the large language model service request scheduling method as described in any one of claims 1-7 are implemented.

Citation Information

Patent Citations

  • Application scheduling deployment method based on big data cluster and storage medium

    CN116680062A

  • Large language model reasoning system, method and equipment without perception of server

    CN116702907A

  • Large language model fine tuning method and device, equipment and storage medium

    CN117592535A

  • Request scheduling method and device for large model

    CN118819771A

Cited By

  • Large model service-oriented task parallel processing intelligent scheduling method and system

    CN120256068A

  • A task parallel processing intelligent scheduling method and system for large model services

    CN120256068B