Multi-model time division multiplexing and parallel loading reasoning service method and system
Through the inference service method of multi-model time-sharing multiplexing and parallel loading, the problems of low resource utilization and inference delay in multi-LLM systems are solved, efficient resource utilization and low-latency inference service are achieved, and the user's service-level goals are met.
Patent Information
- Application Number
- CN202510173996.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-06-27
AI Technical Summary
In multi-LLM systems, static deployment strategies lead to low resource utilization and high deployment costs, and the scarecrow solution has inference delay problems, which cannot meet strict service level goals (SLO).
The inference service method of multi-model time-sharing multiplexing and parallel loading is adopted. By receiving user model inference requests, the model parameter layer is deployed to the GPU, parallel loading and calculation are performed, the parameter layer is unloaded after completion of inference, and the time-sharing multiplexing and parallel loading technology is used to collaborate on resource configuration.
It realizes reducing inference latency while reducing deployment costs, meeting users' SLO requirements, improving resource utilization, and optimizing overall resource configuration through adaptive model deployment technology.
Smart Images

Figure CN120218228A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of machine learning technology, and in particular, to an inference service method and system for multi-model time-sharing multiplexing and parallel loading. Background Art
[0002] Large language models are now widely used to build many online applications. However, for a multi-LLM system that serves multiple model inferences simultaneously, it is a challenge to achieve high resource utilization while meeting strict service level objectives (SLO). Static deployment strategies have problems of low resource utilization and high deployment costs. To reduce deployment costs, an intuitive strawman solution is to only deploy some parameter layers of the model in the GPU during the deployment process, while storing other parameter layers in the CPU. When a model request is received, the model parameter layers stored in the GPU will perform calculations, and at the same time, the remaining parameter layers will be loaded from the CPU to the GPU. Once these initial parameter layers complete the calculation, subsequent layers will also be loaded to enable continuous calculation. This strategy allows the released GPU memory to be used to deploy other models, serving more models with the same number of devices, thereby improving resource utilization and reducing deployment costs.
[0003] Although the strawman solution can reduce deployment costs, since there is no collaboration between multiple models, each model can only use its respective GPU cluster for inference. The calculation speed of the parameter layers in the GPU is usually much higher than the speed of loading these parameter layers from the CPU to the GPU. When the calculation is completed, if the next parameter layer has not been loaded into the GPU, the calculation cannot continue due to the dependency relationship, and the inference process will stall. Therefore, there are a large number of pipeline stalls, which will cause extremely high inference latency, which is unacceptable for many online applications built using large language models with strict latency requirements. Summary of the Invention
[0004] In view of this, this application proposes an inference service method and system for multi-model time-sharing multiplexing and parallel loading, which reduces the inference latency existing in the strawman solution while reducing deployment costs, meeting the SLO requirements of users.
[0005] According to one aspect of this application, there is provided an inference service method for multi-model time-sharing multiplexing and parallel loading, including:
[0006] Receiving a user model inference request, and finding the corresponding model according to the user model inference request;
[0007] Deploying the model parameter layers required for inference to the GPU; inputting the user model inference request text into the parameter layers deployed in the GPU for calculation;
[0008] While calculating in the partial parameter layer, load the remaining parameter layers of the model into the corresponding specified GPU;
[0009] After completing the inference calculation, unload the parameter layer back to the CPU and return the inference result.
[0010] In a possible implementation, receive a user model inference request, and find the corresponding model according to the model information provided in the user model inference request; wherein, the model information includes at least one of a model name and a model type.
[0011] In a possible implementation, when deploying partial parameter layers of the model required for inference to the GPU, it includes:
[0012] Based on the deployment algorithm and the model placement algorithm, calculate the number of parameter layers that the model should be deployed, and the device locations where each parameter layer should be placed during the deployment and loading process; deploy partial parameter layers of the model to the corresponding specified GPU devices; determine the model loading strategy and the model inference strategy.
[0013] In a possible implementation, when calculating the text of the user model inference request by inputting it into the parameter layer deployed on the GPU, first encode the text of the user model inference request, including: splitting the text of the user model inference request through a model tokenizer; then mapping the split text to a digital sequence through a token table; input the digital sequence mapped from the text of the user model inference request into the parameter layer deployed on the GPU for calculation.
[0014] In a possible implementation, the loading of the remaining parameter layers of the model into the corresponding specified GPU includes: loading the remaining parameter layers in the CPU to the corresponding specified GPU device according to the model loading strategy and the parallel loading technology.
[0015] In a possible implementation, before calculating the number of parameter layers that the model should be deployed and the device locations where each parameter layer should be placed during the deployment and loading process, it further includes: if a model list is provided, load the corresponding model, tokenizer, and configuration file from the storage layer to the CPU; for models not saved in the CPU, download the models from the model download platform according to the model list; if no model list is provided, scan the model-related information folder in the storage layer and load the valid models into the CPU.
[0016] In a possible implementation, loading the remaining parameter layers in the CPU to the corresponding specified GPU device according to the model loading strategy and the parallel loading technology further includes: during the calculation, sequentially transfer the current GPU calculation result to the next parameter layer's GPU until the calculation is completed.
[0017] According to another aspect of the present application, an inference service system for multi-model time-sharing multiplexing and parallel loading is provided, including: a service access layer, an execution layer, a storage layer, and a scheduling layer;
[0018] The service access layer includes a model submission service module and an inference service module, which are used to receive user model inference requests and return inference results;
[0019] The execution layer includes a model submission execution module, an initialization deployment module, a model inference execution module, and an uninstallation module, which are used to find the corresponding model according to the user model inference request, encode the user model inference request, load the encoded user model inference request into the GPU where the required model parameter layer is located, and perform model inference calculations in the GPU;
[0020] The storage layer includes a model storage module and a related information storage module, which are used to store the model and model-related information;
[0021] The scheduling layer includes a parameter layer deployment scheduling module and a model placement scheduling module, which are used to schedule the specific location of the GPU where the model parameter layer is located.
[0022] According to another aspect of the present application, a multi-model inference service device is provided, including: a processor; a memory for storing processor-executable instructions; wherein, the processor is configured to execute the above method.
[0023] According to another aspect of the present application, a non-volatile computer-readable storage medium is provided, on which computer program instructions are stored, wherein the computer program instructions implement the above method when executed by a processor.
[0024] Advantages of the present application: By taking advantage of the characteristic that the peak requests of multiple models are staggered, when an inference request for a certain model arrives, time-sharing multiplexing technology is used, and the idle GPU device cluster reserved for other models is utilized to perform inference services for the model currently in the peak request period, enabling the model to use more GPU device clusters for inference, thereby realizing cooperation among the device clusters of multiple models. As the number of GPU device clusters used for inference increases, loading the model parameter layers into these GPU device clusters simultaneously can better mask the pipeline stalls that occur during the inference process and reduce the inference latency of each model request. With fewer parameter layers deployed in the GPU device and less system video memory reserved and occupied, the same or even better inference speed can be achieved, thereby reducing the inference latency while reducing the deployment cost and meeting the service level objective (SLO) requirements of users. Runtime adaptive model deployment technology. This technology automates the model deployment and scheduling process during system operation. Using the minimum video memory deployment algorithm, it minimizes the model deployment cost while meeting the SLO requirements of each model request; in addition, a model placement algorithm is designed to determine the optimal placement position of the model during deployment and loading, optimizing the overall resource configuration of the system.
[0025] Other features and aspects of the present application will become clear from the following detailed description of the exemplary embodiments with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The accompanying drawings, which are included in and constitute a part of this specification, illustrate exemplary embodiments, features, and aspects of the present application together with the specification and are used to explain the principles of the present application.
[0027] Figure 1 Flowchart showing the inference service method of multi-model time-sharing multiplexing and parallel loading according to an embodiment of the present application;
[0028] Figure 2 Block diagram showing the structure of the inference service system of multi-model time-sharing multiplexing and parallel loading according to an embodiment of the present application;
[0029] Figure 3 Flowchart showing the model submission process of the inference service system of multi-model time-sharing multiplexing and parallel loading according to an embodiment of the present application;
[0030] Figure 4 Flowchart showing the model inference process of the inference service system of multi-model time-sharing multiplexing and parallel loading according to an embodiment of the present application;
[0031] Figure 5 Schematic diagram showing the inference of time-sharing multiplexing and parallel loading according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0032] Various exemplary embodiments, features, and aspects of the present application will be described in detail below with reference to the accompanying drawings. The same reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings do not have to be drawn to scale unless otherwise specified.
[0033] Among them, it should be understood that the terms "center", "longitudinal", "transverse", "length", "width", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", "axial", "radial", "circumferential", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings. They are only for the convenience of describing the present application or simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation to the present application.
[0034] The special term "exemplary" here means "serving as an example, embodiment, or illustrative". Any embodiment described as "exemplary" here does not have to be construed as superior to or better than other embodiments.
[0035] In addition, for a better description of the present application, numerous specific details are given in the following specific implementation manners. Those skilled in the art should understand that the present application can also be implemented without certain specific details. In some instances, methods, means, elements, and circuits well-known to those skilled in the art are not described in detail in order to highlight the gist of the present application.
[0036] The present application is applicable to multi-model inference services. Based on time-sharing multiplexing and parallel loading technologies, it realizes the cooperation between device clusters of multiple models. Loading the model parameter layer into these GPU device clusters simultaneously can better mask the pipeline stalls during the inference process, thereby reducing the inference latency of each model request. With less system video memory reserved and occupied, the same or even better inference speed can be achieved, meeting the user's SLO requirements and reducing the overall pre-deployment cost of multiple models.
[0037] Embodiment 1
[0038] Figure 1 A flowchart showing an inference service method for multi-model time-sharing multiplexing and parallel loading according to an embodiment of the present application is shown. As Figure 1 shown, it includes:
[0039] Step S100, receiving a user model inference request, and searching for a corresponding model according to the user model inference request. The user model inference request includes the name or type of the model, and optional model SLO requirements. Taking advantage of the staggered peak values of model requests, the GPU cluster devices of other idle models can be used for inference calculations based on time-sharing multiplexing technology, thereby achieving collaboration between device clusters of multiple models. Time-sharing multiplexing technology is a model request scheduling strategy. When the peak request period of a certain model arrives, the GPU clusters of other models that are currently in an idle period are used to assist it in providing inference services, thereby improving the overall performance of the system.
[0040] Step S200, deploy the parameter layer of the model part required for reasoning into the GPU; input the user model reasoning request text into the parameter layer deployed by the GPU for calculation. The model parameter layer refers to the model preprocessing stage, which divides the model into multiple parameter layers according to the model type, performs reasoning tests on each parameter layer of the model and collects the relevant model information during this test, and finally stores all the model parameter layers, the model's word segmenter and configuration file, and the model's related information in the storage layer. The user model reasoning request text input to the GPU needs to be encoded. Since the model cannot process the text directly, the request text is split by the word segmenter and then mapped to a string of ids through the token table. The process of converting text into a digital sequence is encoding.
[0041] Step S300, while calculating some parameter layers, the remaining parameter layers of the model are loaded into the corresponding designated GPU. Parallel loading technology is used. Parallel loading technology is a model reasoning scheduling strategy. The model parameter layers in the CPU memory are loaded into multiple GPU device clusters in parallel to participate in the calculation. On the basis of time-sharing multiplexing and parallel loading, the reasoning workflow loading strategy is obtained according to the model deployment algorithm, so that the calculation and loading of the reasoning process are better overlapped, thereby better covering up the pipeline pauses that occur in the reasoning process, reducing deployment costs while reducing reasoning latency, and meeting the user's SLO requirements. Pipeline pause refers to when a parameter layer of the model is calculated on the GPU cluster device, if the next parameter layer has not been loaded into the GPU for calculation, the calculation will not be able to continue due to the dependency relationship. Reasoning latency refers to the time required from the initiation of a model request to the request being responded to by the system and returning the result to the user.
[0042] Step S400, after completing the inference calculation, the parameter layer is unloaded and returned to the CPU, and the inference result is returned.
[0043] The remaining parameter layers in the CPU except the corresponding model parameter layers are loaded into the GPU.
[0044] In a possible implementation, a user model inference request is received and encoded based on a model tokenizer for the user model inference request;
[0045] First, the model is initialized and deployed. If a model list is provided in the user model inference request, the corresponding model is loaded from the storage layer to the CPU;
[0046] Among them, the models not saved in the CPU are downloaded from the model download platform according to the model list;
[0047] If a model list is not provided in the user model inference request, the model-related information folder is scanned and the valid models therein are loaded to the CPU;
[0048] After the corresponding model is loaded to the CPU, calculate the number of parameter layers required for model deployment and the location of each parameter layer placed on the GPU;
[0049] Each parameter layer required by the model is correspondingly deployed in its corresponding GPU, and the remaining parameter layers of the model in the CPU are also loaded onto the corresponding GPU device.
[0050] Find the corresponding model parameter layer according to the user model inference request;
[0051] Load the encoded user model inference request to the GPU where the model parameter layer is located;
[0052] Perform model inference calculation in the GPU, unload the parameter layers loaded to the GPU during the inference process back to the CPU, and return the inference result.
[0053] As Figure 2 shown in the multi-model inference service system, the system includes: a storage layer, a service access layer, an execution layer, and a scheduling layer;
[0054] The storage layer is the part of the system that needs to store data locally and is responsible for storing the models submitted by users. Generally, a model is split into multiple model parameter layers for storage. In addition, the storage layer also stores the model configuration, tokenizer, and model-related information, generally using local disks for storage;
[0055] The storage layer includes: a model storage module, a related information storage module;
[0056] The model storage module is used to store all parameter layers of the model submitted by the user after being divided according to the model structure, as well as the tokenizer and model configuration files required during the model inference process;
[0057] The relevant information storage module is used to store the model-related information files collected by the system during the process of the user submitting the model, including the name, type, storage path, inference latency at full deployment, and single-layer inference-related events of each parameter layer, etc.
[0058] The service access layer is the interface between the client and the system backend, processes requests from the client, forwards them to the corresponding backend services for execution, and finally returns the processing execution results of the requests to the client. The service access layer acts as a relay between the client and the system execution layer.
[0059] The service access layer includes a model submission service module and an inference service module;
[0060] The model submission service module is used to process the requests for the user to submit models. According to the model type or name submitted by the user and the SLO requirements for the model inference latency, it sends the model submission requests to specific submission execution functions for execution, and returns corresponding results to the user according to whether the model submission method is successfully executed;
[0061] The inference service module is used to process the inference requests of the user for a certain model. According to the model type or name and the input text submitted by the user, it sends the model inference requests to specific inference execution functions for execution, and returns the inference results generated after the execution of the model inference method as a response to the user;
[0062] The execution layer is the part where the system functions are specifically executed, including the division of the model and the acquisition of model-related information after the user submits the model, the encoding process of the input of the request and the execution of model inference, etc. The main functional methods of the system are all in this layer;
[0063] The execution layer includes a model submission execution module, an initialization deployment module, an inference execution module, and a parameter layer offloading module;
[0064] The model submission execution module is used to process the model submitted by the user. After the user submits the model, this module will divide the model into multiple parameter layers, conduct tests and collect model-related information, and finally store all the model parameter layers and model-related information of the model in the storage layer;
[0065] Specifically, the model division directly divides the parameter layers according to the model structure using the pytorch and transformers libraries. For example, in the embodiments of the present application, methods in the pytorch and transformers libraries are used to obtain a bert-base-uncased model, and its variable name is model, which contains 1 embeddings layer and 12 encoderlayer layers.
[0066] Then, the method model.bert.embeddings.state_dict() can be directly used to obtain the parameters of the embeddings layer of this model, and model.bert.encoder.layer[i].state_dict() can be directly used to obtain the parameters of each encoder layer. The torch.save method is used to store these parameters on the hard disk, so as to achieve the division and storage of the model according to the model structure.
[0067] The initialization and deployment module is used to load all the models stored in the storage layer into the host when the system starts, and deploy the parameter layers of the models and determine the inference strategy according to the deployment algorithm and the model placement algorithm.
[0068] Specifically, the greedy search strategy is adopted in the deployment algorithm to iteratively deploy different numbers of model parameter layers. For each deployment, calculate the total inference time after adopting this deployment, the number of GPU device clusters used, and also consider the possible number of pipeline stalls. In the greedy search process, first select the number of parameter layers for the initial deployment, starting from half of the total number of model parameters. Among them, the number of initial deployment parameter layers can be adjusted accordingly according to the actual situation. Based on this initial deployment, according to the inference workflow, the total inference time and the number of GPU device clusters required can be calculated. If the obtained inference time meets the SLO constraint, the algorithm will experiment with fewer initial deployment parameter layers. If the inference time exceeds the SLO, the algorithm will increase the number of initial deployment parameter layers and repeat the process. This iterative adjustment will continue until the optimal number of initial deployment parameter layers and GPU device clusters is obtained under the condition of meeting the SLO.
[0069] It should be noted that if the number of required GPU device clusters exceeds the number of available GPU device clusters, the calculation and loading processes cannot be fully overlapped, resulting in pipeline stalls. In this case, the algorithm gradually increases the number of parameter layers according to the previously determined optimal value. In the case of pipeline stalls, re-derive the total inference time and the number of GPU device clusters to be used until the number of GPU device clusters reaches the limit number of available devices, so as to obtain the final optimal deployment memory and the number of parameter layers results.
[0070] Finally, the algorithm determines the minimum GPU memory and parameter layers required for each model under the condition of meeting the SLO constraint. At the same time, the algorithm can also obtain the model loading strategy during the derivation process, providing guidance for the inference workflow when model requests arrive. The deployment algorithm minimizes the model deployment cost under the condition of meeting the SLO requirements of each model request;
[0071] In the model placement algorithm, all model parameter layers that perform inference calculations on the same GPU device cluster are grouped together, and the total number of groups corresponds to the number of GPU device clusters required for the model inference process. The algorithm first obtains the memory size occupied by each group based on the memory size occupied by each parameter layer. Next, the algorithm uses the system API to collect the available memory size of each GPU device cluster and sorts the clusters in descending order by available memory. Then, the algorithm sorts the groups from large to small by the memory size occupied, and places them by assigning the group with the largest memory requirement to the GPU device cluster with the most available memory. Finally, this placement process is performed in descending order of the memory size occupied by the group, so that each group is assigned to the device cluster according to its memory requirement, so that as many models as possible can be accommodated on the GPU device cluster, while ensuring that each model can perform inference normally. The model placement algorithm determines the optimal placement location for the model when deploying and loading, and optimizes the overall resource configuration of the system.
[0072] Based on adaptive model deployment technology, the deployment and scheduling scheme of the model is automatically generated when the system is running, thereby optimizing resource utilization. Based on the greedy strategy, the load balancing deployment algorithm generates deployment and scheduling schemes according to indicators such as model information, SLO, and the number of GPUs. It does not preempt the resources being executed for calculation, and gives priority to GPUs with low occupancy to deploy more model layers.
[0073] In addition, the system needs to be preheated to eliminate the impact of system cold start on model inference latency as much as possible; since system preheating is a conventional technical means in this field, it will not be described in detail in this application.
[0074] The inference execution module is used to execute the inference task of the model. After finding the specific location of the model in the system according to the model name, the model word segmenter is used to encode the user input text, and based on the predetermined model parameter layer inference strategy, multi-model time-sharing multiplexing and parallel loading technology are used for inference to obtain the corresponding inference results; time-sharing multiplexing utilizes the characteristics of staggered request peaks of multiple models. When the inference request of a certain model arrives, the idle GPU device cluster reserved for other models is used to perform inference for the model that is currently in the peak request period. The time-sharing multiplexing technology enables the model to use more GPU device clusters for inference, and the parallel loading technology loads the model parameter layer into these GPU device clusters at the same time, which can better cover up the pipeline pauses that occur during the inference process, thereby reducing the deployment cost while reducing the inference latency to meet the user's SLO requirements.
[0075] The parameter layer unloading module is used to unload the part of the parameter layer that was loaded into the GPU device during the model inference process after the model inference process ends and it is determined that the model keep-alive time has expired. This restores the model deployment to its original state to save deployment costs according to the pre-determined model deployment strategy.
[0076] The scheduling layer is the part that schedules the entire system, including scheduling the number of model parameter layers to be deployed and the loading inference strategy at system startup, and dynamically adjusting the model deployment strategy and loading strategy during system operation to adapt to system changes and meet the user's requirements for inference latency. It includes a deployment algorithm that minimizes the model deployment video memory under the premise of meeting the user's SLO requirements, and a scheduling algorithm for the placement location and loading strategy of the model on the device cluster during deployment and system operation.
[0077] The scheduling layer includes a parameter layer deployment scheduling module and a model placement scheduling module.
[0078] The parameter layer deployment scheduling module is used to schedule the number of parameter layers of a model that are pre-deployed on the GPU device during system initialization deployment or system operation. It includes a deployment video memory algorithm that is modeled based on the inference workflow and obtains the optimal number of model parameter layers to be deployed under the premise of meeting the user's SLO requirements according to parameters such as the number of available GPU devices, SLO, and model-related information to reduce deployment costs.
[0079] The model placement scheduling module is used to schedule the specific location where the parameter layer of a model should be placed on the system GPU device during system initialization deployment or system operation. It obtains the GPU device location and loading strategy where the model parameter layer should be placed based on parameters such as the current GPU device usage of the system and model-related information, thereby optimizing the system's resource configuration.
[0080] In a possible implementation, the multi-model inference service system includes two main functional operations: model submission and model inference.
[0081] The model submission process is as Figure 3 shown:
[0082] The client user submits a model submission request (1), which includes the name or type of the model, and optionally the model SLO requirement.
[0083] After the service access layer receives the user's model submission request, the model submission service module hands it over to the model submission execution module in the execution layer for further processing (2).
[0084] After the model submission and execution module in the execution layer receives a task, it downloads the corresponding model, tokenizer, and configuration file (3) from the model download platform according to the model name or type submitted by the user.
[0085] After the model is downloaded, the execution module divides the model into multiple parameter layers according to the model type, conducts inference tests on each parameter layer of the model, and collects relevant model information during this test process. Finally, all the model parameter layers, the tokenizer and configuration file of the model, and the collected relevant model information are stored in the storage layer (4). The storage layer usually uses the local disk as the storage medium for storage;
[0086] If the model has been successfully stored in the storage layer, the service access layer will return the information that the model submission is successful as the request response result to the client (5); on the contrary, if the model submission fails due to reasons such as download failure or the corresponding model name does not exist on the website, the service access layer will return the response information that the model submission fails to the client.
[0087] The model inference process is as Figure 4 shown:
[0088] Before model inference, the system needs to start and complete the initialization and deployment of each model. The steps are as follows:
[0089] During the system startup process, if the user provides a specific list of models to be loaded, then the initialization and deployment module in the execution layer will load the corresponding model, tokenizer, and configuration file from the storage layer into the CPU (1). The models not saved in the CPU will be downloaded from the model download platform; if no specific model list is provided at startup, then the initialization and deployment module in the execution layer will scan the model-related information folder in the storage layer and load the models and the corresponding tokenizer and configuration file in the valid model-related information files into the CPU.
[0090] After all the parameter layers of all models are loaded into the CPU, the system will calculate the number of parameter layers that each model should be deployed according to the deployment algorithm and model placement algorithm, and the device location where each parameter layer should be placed during the deployment and loading process. Then, some parameter layers of each model will be deployed to the specified GPU device (2), and the loading and inference strategies for each model will be determined.
[0091] After the system initialization and deployment is completed, when a model inference request arrives:
[0092] After the service access layer receives the user's model inference request, it is further processed by the model inference execution module in the execution layer by the inference service module (3).
[0093] Based on time-division multiplexing technology, the model during the peak request period can utilize the reserved GPU devices that are currently in an idle state. Inference is performed on multiple groups of GPU cluster devices through collaborative computing and parallel loading.
[0094] The inference execution module in the execution layer first locates the position of the corresponding model parameter layer in the system according to the model name or type in the inference request. Then, it uses the tokenizer of the model to encode the request input text submitted by the user, loads the encoded text into the GPU device where the parameter layer of the model is deployed, and then performs inference calculations. At the same time, the inference execution module loads the remaining parameter layer part of the model stored in the CPU onto the corresponding GPU device according to the pre-set loading strategy (4). Finally, after completing the inference calculation and obtaining the result, the model inference execution module organizes and returns the calculation result to the service access layer.
[0095] After the inference is completed, the unloading module unloads the parameter layer temporarily loaded into the GPU device during the inference process back to the CPU after the model has passed the keep-alive time, restoring the initial deployment state, that is, the parameter layer does not occupy resources in the GPU. The service access layer returns the inference result to the user of the client (5), ensuring the efficiency of inference and resource utilization.
[0096] Figure 5 The figure shows a schematic diagram of the model inference process. Based on the inference workflow loading strategy, the optimal deployment memory and the number of parameter layers are calculated for the model. For a model with a total of 14 parameter layers, the calculation of layers 1 - 7 is first performed, and at the same time, layers 8 and 9 are loaded into the GPU. When the calculation of layer 7 is completed, layer 8 has been loaded and can directly proceed to the next layer of calculation. At the same time, the remaining parameter layers are loaded using other idle GPU devices, and the calculation results are transmitted to the next loaded parameter layer, and so on until the inference calculation is completed. Compared with the strawman solution in the background technology, the calculation and loading in the inference process are better overlapped, reducing pipeline stalls and lowering the inference latency.
[0097] To illustrate the multi-model inference service system proposed by the present invention in more detail, the following gives a specific embodiment 2:
[0098] In a scenario where a model is submitted to the system, the client user wants to submit a model named bert-base-uncased to the system. At this time, the user only needs to set the model name to bert-base-uncased in the model submission request and then send the request. After the service access layer receives the user's model submission request, it is handed over to the model submission execution module in the execution layer by the model submission service module. The model submission execution module in the system execution layer downloads the corresponding model, tokenizer, and configuration file from the model download platform website according to the model name or type submitted by the user, and will automatically download the bert-base-uncased model from the huggingface website of the model download platform and store it in the system for subsequent use.
[0099] When classifying the model, first divide the model into multiple parameter layers according to the model type, perform inference tests on each parameter layer of the model, and collect relevant model information during this test process. Finally, store all the model parameter layers, the tokenizer and configuration file of the model, and the relevant information of the model in the storage layer (4), and the storage layer uses the local disk for storage.
[0100] If the model has been successfully stored in the storage layer, the service access layer will return the information of successful model submission as the request response result to the client (5); if the model submission fails due to reasons such as download failure or non-existence of the corresponding model name on the website, the service access layer will return the information of failed model submission to the client.
[0101] During the initialization deployment process when the system starts, this model named bert-base-uncased will be deployed in the system. When the user wants to use this model named bert-base-uncased for inference operations, only need to set the model name to bert-base-uncased and set the input text when initiating the model inference request, and then send the request, and the system will return the result obtained by this model for inferring the input text.
[0102] Furthermore, according to another aspect of the present application, an inference service system for multi-model time-sharing multiplexing and parallel loading is also provided, including: a service access layer, an execution layer, a storage layer, and a scheduling layer.
[0103] The service access layer includes a model submission service module and an inference service module, which are used to receive user model inference requests and return inference results.
[0104] The execution layer includes a model submission and execution module, an initialization and deployment module, a model inference execution module, and an uninstallation module, which are used to find the corresponding model according to the user's model inference request; encode the user's model inference request; load the encoded user's model inference request into the GPU where the model parameter layer is located; and perform model inference calculations in the GPU.
[0105] The storage layer includes a model storage module and a relevant information storage module, which are used to store the model parameter layer and model-related information.
[0106] The scheduling layer includes a parameter layer deployment scheduling module and a model placement scheduling module, which are used to schedule the specific location of the GPU where the model parameter layer is located.
[0107] Furthermore, according to another aspect of the present application, an inference service device for multi-model time-sharing multiplexing and parallel loading is also provided, including a processor and a memory for storing processor-executable instructions. The processor is configured to implement the multi-model inference service method described in any one of the foregoing when executing the executable instructions. Here, it should be noted that the number of processors can be one or more.
[0108] As a computer-readable storage medium, the memory can be used to store software programs, computer-executable programs, and various modules, such as the programs or modules corresponding to the multi-model inference service method of the embodiments of the present application. The processor executes various functional applications and data processing of the inference service device for multi-model by running the software programs or modules stored in the memory.
[0109] According to another aspect of the present application, a non-volatile computer-readable storage medium is also provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor 210, the multi-model inference service method described in any one of the foregoing is implemented.
[0110] The embodiments of the present application have been described above. The above description is exemplary and not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art in the technical field without departing from the scope and spirit of the described embodiments. The selection of the terms used herein is intended to best explain the principles of the embodiments, practical applications, or technical improvements to the technologies in the market, or to enable other ordinary skill in the art in the technical field to understand the embodiments disclosed herein.
Claims
1. A multi-model time-sharing multiplexing and parallel loading reasoning service method, characterized in that: include: Receiving a user model inference request, and searching for a corresponding model according to the user model inference request; Deploy some parameter layers of the model required for inference to the GPU; Input the user model inference request text into the parameter layer deployed by GPU for calculation; While the partial parameter layers are being calculated, the remaining parameter layers of the model are loaded into the corresponding designated GPU; After the inference calculation is completed, the parameter layer is unloaded back to the CPU and the inference result is returned.
2. The method according to claim 1, characterized in that Receive a user model inference request, and find a corresponding model according to the model information provided in the user model inference request; The model information includes at least one of a model name and a model type.
3. The method according to claim 1, characterized in that When deploying some parameter layers of the model required for inference to the GPU, it includes: Based on the deployment algorithm and model placement algorithm, calculate the number of parameter layers that the model should deploy and the device location where each parameter layer should be placed during the deployment and loading process; Deploy some parameter layers of the model to the corresponding specified GPU device; Determine the model loading strategy and model inference strategy.
4. The method according to claim 1, characterized in that: When the user model inference request text is input into the parameter layer of GPU deployment for calculation, the user model inference request text is first encoded, including: Segment the user model inference request text by a model word segmenter; Then use the token table to map the segmented text into a digital sequence; The digital sequence mapped from the user model reasoning request text is input into the parameter layer deployed in the GPU for calculation.
5. The method according to claim 1, characterized in that The step of loading the remaining parameter layers of the model into the corresponding designated GPU includes: According to the model loading strategy and parallel loading technology, the remaining parameter layers in the CPU are loaded onto the corresponding designated GPU device.
6. The method according to claim 3, characterized in that The calculation model should deploy the number of parameter layers and the device location where each parameter layer should be placed during deployment and loading, and also includes: If a model list is provided, the corresponding model, tokenizer, and configuration file are loaded from the storage layer to the CPU; The models not stored in the CPU are downloaded from the model download platform according to the model list; If no model list is provided, the model-related information folder of the storage layer is scanned and the valid models are loaded into the CPU.
7. The method according to claim 5, characterized in that According to the model loading strategy and parallel loading technology, the remaining parameter layers in the CPU are loaded onto the corresponding designated GPU device, and further comprising: When performing calculations, the current GPU calculation results are transmitted to the GPU where the next parameter layer is located in sequence until the calculation is completed.
8. A multi-model time-sharing multiplexing and parallel loading inference service system, characterized in that: include: Service access layer, execution layer, storage layer, and scheduling layer; The service access layer includes a model submission service module and an inference service module, which are used to receive user model inference requests; Return the inference result; The execution layer includes a model submission execution module, an initialization deployment module, a model reasoning execution module, and an unloading module, which are used to search for a corresponding model according to the user model reasoning request; encode the user model reasoning request; load the encoded user model reasoning request to the GPU where the model parameter layer is located; and perform model reasoning calculation in the GPU; The storage layer includes a model storage module and a related information storage module, which are used to store the model and model related information; The scheduling layer includes a parameter layer deployment scheduling module and a model placement scheduling module, which are used to schedule the specific location of the GPU where the model parameter layer is located.
9. A multi-model reasoning service device, characterized in that: include: processor; a memory for storing processor-executable instructions; Wherein, the processor is configured to implement the method described in any one of claims 1 to 7 when executing the executable instructions.
10. A non-volatile computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
GPU computing power scheduling method, device and equipment, medium and program product
CN120994411A
GPU computing power scheduling method and device, equipment, medium and program product
CN120994411B
Large model reasoning platform scheduling method of parallel architecture, platform and medium
CN121255404A
Data processing method and system
CN121458523A