Large model cluster deployment method and device, equipment and storage medium

By performing parallel quantization and metric evaluation of the large model, combined with dynamic programming strategies and heterogeneous cluster resource information, the problems of GPU performance squeezing and CPU idleness in heterogeneous clusters are solved, and a larger model deployment with lower cost and higher stability are achieved.

CN120045196APending Publication Date: 2025-05-27SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510159553.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

During the deployment of large-scale models, GPU performance swelling and idle CPU computing resources are prone to occur in heterogeneous clusters, resulting in high deployment costs and poor stability.

Method used

By quantifying the large model in parallel based on preset quantization tools, and evaluating the quantized model after metrics, we will determine the list of models to be deployed. Combining the computing resource information and operating status of heterogeneous clusters, a dynamic planning strategy is used to determine the target deployment plan, and the model is reasonably allocated to GPU and CPU resources.

Benefits of technology

It effectively avoids the situation of GPU performance squeezing in heterogeneous clusters, reduces deployment costs, and improves the operational stability of large models in heterogeneous clusters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045196A_ABST
    Figure CN120045196A_ABST
Patent Text Reader

Abstract

The invention discloses a large model cluster deployment method and device, equipment and a storage medium, and relates to the field of artificial intelligence, and the method comprises the steps: carrying out the parallel quantification of a large model corresponding to a deployment task based on a preset quantification tool, and carrying out the index evaluation of the quantized model, so as to obtain a corresponding model evaluation result; determining a to-be-deployed model list corresponding to the deployment task based on the model evaluation result; obtaining computing resource information and a cluster running state of a heterogeneous cluster corresponding to the deployment task at present; and determining a target deployment scheme by using the computing resource information, the cluster running state, the to-be-deployed model list and a dynamic planning strategy, and completing a large model deployment operation corresponding to the deployment task based on the target deployment scheme. According to the method and the device, the condition of GPU performance occupation in the heterogeneous cluster can be effectively avoided, so that the deployment cost is reduced, and the operation stability of the large model in the heterogeneous cluster is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and particularly to a method, device, equipment and storage medium for deploying a large model cluster. Background Art

[0002] With the wide application of large models, the demand for using large models to process tasks such as consultations is increasing day by day.

[0003] However, the implementation of large models in these tasks often requires multiple models to cooperate, and these models have different sizes. When deploying large models using existing implementation means, more manual operations are required to deploy large models on GPUs (Graphics Processing Units). Moreover, if all are deployed on GPU computing cards, not only will there be a situation of GPU performance occupation in the cluster, but also there will be an idle problem of CPU (Central Processing Unit) computing resources. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide a method, device, equipment and storage medium for deploying a large model cluster, which can effectively avoid the situation of GPU performance occupation in a heterogeneous cluster, thereby reducing the deployment cost and improving the stability of the large model running in the heterogeneous cluster. The specific solutions are as follows:

[0005] In a first aspect, the present application provides a method for deploying a large model cluster, including:

[0006] Performing parallel quantization on a large model corresponding to a deployment task based on a preset quantization tool, and performing index evaluation on the quantized model to obtain a corresponding model evaluation result;

[0007] Determining a list of models to be deployed corresponding to the deployment task based on the model evaluation result;

[0008] Obtaining the computing resource information and cluster running status of the current heterogeneous cluster corresponding to the deployment task;

[0009] Using the computing resource information, the cluster running status, the list of models to be deployed, and a dynamic programming strategy to determine a target deployment plan, and completing the deployment operation of the large model corresponding to the deployment task based on the target deployment plan.

[0010] Optionally, the performing parallel quantization on a large model corresponding to a deployment task based on a preset quantization tool includes:

[0011] Triggering a first model quantization operation on a large model corresponding to a deployment task based on an automatic adaptive weight quantization framework and a weight quantization algorithm based on activation values to obtain a first quantization result;

[0012] Trigger the second model quantization operation of the large model corresponding to the deployment task based on the offline quantization algorithm to obtain the second quantization result.

[0013] Optionally, the evaluating the metrics of the quantized model to obtain the corresponding model evaluation result includes:

[0014] Use the WikiText dataset and the first quantization result to evaluate the model confusion metric, the first-character latency metric of model inference, and the character generation speed metric of the quantized model to obtain the corresponding first model evaluation result;

[0015] Use the WikiText dataset and the second quantization result to evaluate the model confusion metric, the first-character latency metric of model inference, and the character generation speed metric of the quantized model to obtain the corresponding second model evaluation result.

[0016] Optionally, the determining the list of models to be deployed corresponding to the deployment task based on the model evaluation result includes:

[0017] Visualize the model evaluation result on the front-end interface to determine the list of models to be deployed corresponding to the deployment task.

[0018] Optionally, the obtaining the computing resource information and the cluster running status of the heterogeneous cluster corresponding to the deployment task currently includes:

[0019] Obtain the computing resource information and the cluster running status of the heterogeneous cluster corresponding to the deployment task currently by integrating the Grafana monitoring component and the Prometheus monitoring component in the container cluster management system.

[0020] Optionally, the determining the target deployment plan using the computing resource information, the cluster running status, the list of models to be deployed, and the dynamic programming strategy includes:

[0021] Based on the computing resource information, the cluster running status, and the dynamic programming strategy, determine the resource requirement information and the target running platform corresponding to each large model to be deployed in the list of models to be deployed to obtain the target deployment plan.

[0022] Optionally, the completing the large model deployment operation corresponding to the deployment task based on the target deployment plan includes:

[0023] Deploy each large model to be deployed to the graphics processor computing resources in the heterogeneous cluster corresponding to the deployment task based on the resource requirement information and the target running platform corresponding to each large model to be deployed;

[0024] When deploying the collaborative model corresponding to each of the to-be-deployed large models to the heterogeneous cluster, deploy the collaborative model to the computing resources of the central processing unit.

[0025] In a second aspect, the present application provides a large model cluster deployment device, including:

[0026] A parallel quantization module, configured to perform parallel quantization on a large model corresponding to a deployment task based on a preset quantization tool, and evaluate the metrics of the quantized model to obtain a corresponding model evaluation result;

[0027] A list determination module, configured to determine a list of to-be-deployed models corresponding to the deployment task based on the model evaluation result;

[0028] An information acquisition module, configured to acquire the computing resource information and the cluster running status of the heterogeneous cluster currently corresponding to the deployment task;

[0029] A large model deployment module, configured to determine a target deployment plan by using the computing resource information, the cluster running status, the list of to-be-deployed models, and a dynamic programming strategy, and complete the large model deployment operation corresponding to the deployment task based on the target deployment plan.

[0030] In a third aspect, the present application provides an electronic device, including:

[0031] A memory, configured to store a computer program;

[0032] A processor, configured to execute the computer program to implement the steps of the foregoing large model cluster deployment method.

[0033] In a fourth aspect, the present application provides a computer-readable storage medium, configured to store a computer program, and when the computer program is executed by a processor, implement the steps of the foregoing large model cluster deployment method.

[0034] It can be seen that in this application, the large model corresponding to the deployment task is parallelly quantized based on a preset quantization tool, and the quantized model is evaluated for metrics to obtain the corresponding model evaluation result; a list of models to be deployed corresponding to the deployment task is determined based on the model evaluation result; the computing resource information and the cluster running status of the heterogeneous cluster corresponding to the deployment task are obtained currently; a target deployment plan is determined by using the computing resource information, the cluster running status, the list of models to be deployed, and a dynamic programming strategy, and the large model deployment operation corresponding to the deployment task is completed based on the target deployment plan. That is to say, when dealing with the deployment task of a heterogeneous cluster in this application, first, a preset quantization tool is used for quantization, and the quantized model is evaluated for metrics. Then, a list of models to be deployed is determined according to the model evaluation result. After that, a target deployment plan corresponding to the deployment task is determined according to the dynamic programming strategy, the list of models to be deployed, and the computing resource information and the cluster running status of the heterogeneous cluster, and the large model deployment corresponding to the deployment task is completed. In this way, the situation of GPU performance occupation in the heterogeneous cluster can be effectively avoided, thereby reducing the deployment cost and improving the stability of the large model running in the heterogeneous cluster. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0036] Figure 1 It is a flowchart of a method for deploying a large model cluster provided by this application;

[0037] Figure 2 It is a flow architecture diagram of a large model cluster deployment provided by this application;

[0038] Figure 3 It is a schematic structural diagram of a device for deploying a large model cluster provided by this application;

[0039] Figure 4 It is a structural diagram of an electronic device provided by this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0041] When existing implementation means deploy large models, more manual operations are required to deploy large models on GPUs. Moreover, if all are deployed on GPU computing cards, not only will there be a situation of GPU performance occupation in the cluster, but also the problem of idle CPU computing resources will occur. For this reason, the present application provides a large model cluster deployment solution, which can effectively avoid the situation of GPU performance occupation in a heterogeneous cluster, thereby reducing the deployment cost and improving the stability of the large model running in the heterogeneous cluster.

[0042] See Figure 1 As shown, an embodiment of the present invention discloses a large model cluster deployment method, including:

[0043] Step S11, perform parallel quantization on the large model corresponding to the deployment task based on a preset quantization tool, and evaluate the metrics of the quantized model to obtain the corresponding model evaluation result.

[0044] In this embodiment, combined with Figure 2 As shown in the architecture diagram, it should be understood that in a computing power cluster, there are multiple nodes, and each node is configured with multiple GPU cards and single or multiple high-performance CPUs. In a heterogeneous cluster, a heterogeneous high-performance structure is adopted, and there are various combinations of computing resources. When the service starts, models with different amounts of computation are distributed to different computing resources in the cluster nodes for inference. For example Figure 2 As shown in the heterogeneous cluster, in this embodiment, an automatic quantization tool (i.e., the preset quantization tool) in the automatic quantization system is used to generate low-bit quantization models with multiple quantization schemes in parallel for the models in the service model list in the cluster, and at the same time evaluate the model metrics and feedback them to the front-end interface to compress the model to reduce the communication cost during the inference service deployment process of the model, and at the same time effectively utilize the high-performance tensor computing core of the GPU card. The processed models in the service model list and the computing resource information of the cluster are received through the scheduling management system, and the models to be deployed are deployed to the corresponding computing resources through the dynamic programming algorithm to effectively improve the utilization efficiency of the service for the cluster. The running state of the current cluster and the information related to the computing resources are also obtained through the monitoring component and fed back to the scheduling management system.

[0045] Specifically, some quantization schemes supported by the preset quantization tool are as follows:

[0046] (1)Implement a 4-bit model integer quantization algorithm for only the quantized model based on the automatic adaptive weight quantization framework (i.e., the AutoAWQ framework, the Automatic Adaptive Weight Quantization framework), that is, trigger the first model quantization operation of the large model corresponding to the deployment task based on the activation value-based weight quantization algorithm (i.e., the AWQ algorithm) to obtain the first quantization result. This method can effectively reduce the video memory occupancy cost during the inference process of the model, reducing the video memory occupancy of the model by 70%. This solution can also effectively utilize the tensor cores of the GPU card to achieve high-performance matrix calculations of W4A16 (weight 4 bits + activation 16 bits), effectively improving the throughput performance of the model inference service.

[0047] (2)Trigger the second model quantization operation of the large model corresponding to the deployment task based on the offline quantization algorithm (i.e., the smoothquant algorithm, the full name is the Smooth Quantization algorithm) to obtain the second quantization result. By implementing an 8-bit quantization algorithm for model parameters and activation parameters, this method can effectively utilize the 8-bit tensor core computing performance of the GPU card while effectively ensuring the stability of the content generated by the model. Among them, the computing power of the 8-bit tensor computing core is twice that of the original 16-bit floating-point tensor computing core. This quantization method can effectively improve the computing performance of the inference service and effectively reduce the service latency.

[0048] Further, in this embodiment, it is also necessary to evaluate the metrics of the quantized model to obtain the corresponding model evaluation results, including: using the WikiText dataset (i.e., the WikiText dataset) and the first quantization result to evaluate the model confusion metric, the first-word latency metric of model inference, and the character generation speed metric of the quantized model to obtain the corresponding first model evaluation results; using the WikiText dataset and the second quantization result to evaluate the model confusion metric, the first-word latency metric of model inference, and the character generation speed metric of the quantized model to obtain the corresponding second model evaluation results.

[0049] Step S12: Determine the list of models to be deployed corresponding to the deployment task based on the model evaluation results.

[0050] In this embodiment, after obtaining the model evaluation results, the list of models to be deployed corresponding to the deployment task will be determined based on the model evaluation results, including: visualizing the model evaluation results on the front-end interface to determine the list of models to be deployed corresponding to the deployment task. In this way, by visualizing the metrics, users can quickly screen the most suitable quantization method for the current task, and then determine the list of models to be deployed.

[0051] Step S13: Obtain the computing resource information and cluster running status of the heterogeneous cluster corresponding to the current deployment task.

[0052] In this embodiment, the cluster is monitored by a monitoring component, and when the model is deployed, the data monitored by the component is obtained through the scheduling and management system to participate in the deployment decision-making. That is, the obtaining of the computing resource information and cluster running status of the heterogeneous cluster corresponding to the current deployment task includes: obtaining the computing resource information and cluster running status of the heterogeneous cluster corresponding to the current deployment task through the Grafana monitoring component and the Prometheus monitoring component integrated in the container cluster management system (i.e., the kubernetes platform). In this way, the scheduling and management system can obtain the running status information of each node in the cluster in real time, complete the unified management of the cluster computing resources, and achieve the efficient utilization of the cluster computing resources.

[0053] Step S14: Determine the target deployment plan by using the computing resource information, the cluster running status, the list of models to be deployed, and the dynamic programming strategy, and complete the large model deployment operation corresponding to the deployment task based on the target deployment plan.

[0054] Specifically, in this embodiment, after obtaining the computing resource information, the cluster running status, and the list of models to be deployed in real time, the large model and the collaborative model are dynamically matched and deployed in the CPU computing resources and GPU computing resources of the cluster to determine the target deployment plan. Then, the single instruction multiple data computing instruction set of the CPU in the heterogeneous cluster is effectively called through the software system corresponding to the cluster to effectively utilize the original idle CPU computing resources and complete the model deployment.

[0055] It should be further understood that the determining of the target deployment plan by using the computing resource information, the cluster running status, the list of models to be deployed, and the dynamic programming strategy includes: determining the resource requirement information and the target running platform corresponding to each large model to be deployed in the list of models to be deployed based on the computing resource information, the cluster running status, and the dynamic programming strategy to obtain the target deployment plan. The completing of the large model deployment operation corresponding to the deployment task based on the target deployment plan includes: deploying each large model to be deployed to the graphics processor computing resources in the heterogeneous cluster corresponding to the deployment task based on the resource requirement information and the target running platform corresponding to each large model to be deployed; when deploying the collaborative model corresponding to each large model to be deployed to the heterogeneous cluster, deploying the collaborative model to the central processing unit computing resources.

[0056] In a specific embodiment, when using a GPU card to deploy a large model, the large model inference deployment framework VLLM (Virtual Large Language Model) is used. Based on the high-concurrency inference advantage of the VLLM framework, a high-performance large model inference service implementation based on the GPU is constructed. The specific operations are as follows:

[0057] (1) Use the Page Attention of the VLLM framework (a memory management technology for optimizing the large model inference process) and the Scheduler implementation to complete the paging management of the GPU card video memory within the server and nodes, effectively alleviating the waste of GPU video memory caused by multiple concurrent applications during the inference process, and effectively increasing the throughput performance of the GPU card in the inference service.

[0058] (2) Use the Flash Attention kernel function implementation of the VLLM framework (a key component for accelerating Transformer models in natural language processing and other applications that require attention mechanisms) to effectively handle the Attention calculation process during the large model inference process. By optimizing the problem of excessive memory access frequency of softmax (an activation function), the processing latency of the large model in handling long sequences is effectively reduced.

[0059] (3) Based on the tensor parallelism feature implemented by the VLLM framework, solve the problem that the large model has too many parameters to be inferred on a single GPU card. Tensor parallelism can also efficiently utilize the strong computing power of multiple cards to alleviate the computing pressure when inferring large models.

[0060] (4) Use the quantization and function implementation of the VLLM framework. By quantizing the large model parameters, the parameters are compressed from the half-precision floating-point data type to 4-bit integer data variables. This reduces the inference transmission cost during the inference process and can effectively utilize the high-performance tensor cores of the GPU card to further enhance the computing performance of the inference service.

[0061] Based on the above technologies, by deploying the quantized large model on the server with an invocation interface adapted to OpenAI (Open Artificial Intelligence, an artificial intelligence), users can realize the function of conversing with the large model by sending requests to this service. The specific parameters can be referred to as follows:

[0062] 1) "model": Specify the large model access path.

[0063] 2) "host", "port": Specify the server IP (Internet Protocol Address) and port for the large model service request.

[0064] 3) "quantization": Specifies the quantization method and provides the parameter quantization method for the large model.

[0065] 4) "tensor-parallel-size": Specifies the tensor parallel size, and the model will be distributed on the specified number of GPUs according to the specified size to complete parallelism.

[0066] 5) "gpu-memory-utilization": Specifies the proportion of video memory occupied by the large model service on the GPU card.

[0067] In another specific implementation, when deploying the large model by calling the GPU card, the large model inference deployment framework Llama.CPP (Llama Language Model Inference Framework, an inference framework for loading and running the LLaMA language model) is used, and a CPU large model service deployment module is built based on the high-performance CPU implementation and functions of the Llama.CPP inference framework and various model quantization methods. In the usage scenario of large model tasks, multiple models often need to cooperate rather than a single large model for text generation. In actual deployment services, multiple models perform pre-processing and post-processing on the input and output data of the large model service. Among them, in order to reduce the occupancy of GPU card resources by non-generation models, the method of the patent deploys non-generation models in the CPU computing resources. The specific operations are as follows:

[0068] (1) Use the high-performance matrix multiplication operator of llama.cpp to support the high-performance computing requirements of large model inference.

[0069] (2) Use various model quantization methods in llama.cpp to quantize non-generation models in the deployment service, effectively reducing the memory access pressure during model inference and at the same time effectively utilizing the computing instruction set supported by the processor architecture.

[0070] (3) By calibrating the resource requirements and annotating the deployment platform for the models to be deployed in the service, at the beginning of the large model service deployment, the model deployment scheduling system implemented based on the dynamic programming algorithm calculates and plans the operation platform of each model and starts relevant parameter information in the current computing resources. Different models are deployed on different computing platforms through the above information to obtain the maximum utilization rate of computing resources.

[0071] (4)Build CUDA (Compute Unified Device Architecture, a parallel computing platform and programming model) into the base image, enabling the container to have the ability to schedule the host graphics card driver.

[0072] (5)Build the deep learning framework into the container image and run it in the form of a container during inference.

[0073] (6)Use the container orchestration framework to automatically schedule tasks, and isolate each model resource of the inference service. Deploy the model in the heterogeneous computing cluster through the task planning system based on the resource dynamic programming algorithm to prevent task failures caused by idle waste or occupation of the cluster computing resources.

[0074] In addition, when deploying the model in this embodiment, the model running storage occupancy mark will also be used to obtain the storage size required by the large model and the collaborative model during the service deployment process and the corresponding running platform to complete the data support work for the service deployment model.

[0075] In summary, in the model deployment solution described in this embodiment, the large model is deployed in the current computing cluster, and a variety of inference frameworks and computing resources are used to support heterogeneous computing to form a heterogeneous inference service cluster; the large model is divided into multiple model scales and allocated to different computing cores for inference; multiple automatic quantization schemes are provided for the large model to reduce the resource occupancy for realizing the large model inference service; a dynamic resource scheduling and management system is adopted to coordinate the matching of the computing resource requirements of each model in the large model deployment service and the heterogeneous computing resources in different clusters, and improve the utilization efficiency of the cluster computing resources. Its beneficial effects include:

[0076] 1. Improve the utilization rate of cluster computing resources: While deploying the large model on the GPU, deploy the common collaborative models of the service on the CPU computing resources to reduce the possibility of multiple models in the service squeezing the GPU resources.

[0077] 2. Reduce costs: By integrating multiple quantization methods and multiple inference frameworks, while ensuring the quality of the inference service, call the original idle computing resources through the software system to effectively reduce the cost requirements for service deployment.

[0078] 3. System universality: It can be effectively applied to the computing clusters of various heterogeneous computing architectures.

[0079] 4. Improve service stability: Through the automatic quantization scheme to achieve and avoid GPU performance squeezing, it is possible to avoid a large number of problems such as high service deployment costs and runtime problems caused by GPU computing resource squeezing.

[0080] It can be seen that in the embodiments of the present application, the large model corresponding to the deployment task is parallelly quantized based on a preset quantization tool, and the quantized model is evaluated for metrics to obtain the corresponding model evaluation result; the list of models to be deployed corresponding to the deployment task is determined based on the model evaluation result; the computing resource information and the cluster running status of the heterogeneous cluster corresponding to the deployment task are obtained; the target deployment plan is determined by using the computing resource information, the cluster running status, the list of models to be deployed, and the dynamic programming strategy, and the large model deployment operation corresponding to the deployment task is completed based on the target deployment plan. That is, when processing the deployment task of the heterogeneous cluster in the present application, first, a preset quantization tool is used for quantization, and the quantized model is evaluated for metrics, and then the list of models to be deployed is determined according to the model evaluation result. Then, according to the dynamic programming strategy, the list of models to be deployed, and the computing resource information and the cluster running status of the heterogeneous cluster, the target deployment plan corresponding to the deployment task is determined, and the large model deployment corresponding to the deployment task is completed. In this way, the situation of GPU performance occupation in the heterogeneous cluster can be effectively avoided, thereby reducing the deployment cost and improving the stability of the large model running in the heterogeneous cluster.

[0081] See Figure 3 As shown, the embodiments of the present application also correspondingly disclose a large model cluster deployment device, including:

[0082] The parallel quantization module 11 is used to parallelly quantize the large model corresponding to the deployment task based on a preset quantization tool, and evaluate the metrics of the quantized model to obtain the corresponding model evaluation result;

[0083] The list determination module 12 is used to determine the list of models to be deployed corresponding to the deployment task based on the model evaluation result;

[0084] The information acquisition module 13 is used to acquire the computing resource information and the cluster running status of the heterogeneous cluster corresponding to the deployment task;

[0085] The large model deployment module 14 is used to determine the target deployment plan by using the computing resource information, the cluster running status, the list of models to be deployed, and the dynamic programming strategy, and complete the large model deployment operation corresponding to the deployment task based on the target deployment plan.

[0086] Among them, for the more specific working processes of the above-mentioned various modules, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be elaborated herein.

[0087] It can be seen that when this application processes the deployment task of a heterogeneous cluster, it first uses a preset quantization tool for quantization, evaluates the metrics of the quantized model, and then determines the list of models to be deployed according to the model evaluation results. After that, according to the dynamic programming strategy, the list of models to be deployed, the computing resource information of the heterogeneous cluster, and the cluster running status, it determines the target deployment plan corresponding to the deployment task, and completes the deployment of the large model corresponding to the deployment task. In this way, it can effectively avoid the situation of GPU performance occupation in the heterogeneous cluster, thereby reducing the deployment cost and improving the stability of the large model running in the heterogeneous cluster.

[0088] In some specific embodiments, the parallel quantization module 11 can specifically be used to trigger the first model quantization operation of the large model corresponding to the deployment task based on the automatic adaptive weight quantization framework and the weight quantization algorithm based on the activation value to obtain the first quantization result; trigger the second model quantization operation of the large model corresponding to the deployment task based on the offline quantization algorithm to obtain the second quantization result.

[0089] In some specific embodiments, the parallel quantization module 11 can specifically be used to use the WikiText dataset and the first quantization result to evaluate the model confusion degree index, the first word latency index of model inference, and the character generation speed index of the quantized model to obtain the corresponding first model evaluation result; use the WikiText dataset and the second quantization result to evaluate the model confusion degree index, the first word latency index of model inference, and the character generation speed index of the quantized model to obtain the corresponding second model evaluation result.

[0090] In some specific embodiments, the list determination module 12 can specifically be used to visualize the model evaluation results on the front-end interface to determine the list of models to be deployed corresponding to the deployment task.

[0091] In some specific embodiments, the information acquisition module 13 can specifically be used to obtain the computing resource information and the cluster running status of the current heterogeneous cluster corresponding to the deployment task by integrating the Grafana monitoring component and the Prometheus monitoring component in the container cluster management system.

[0092] In some specific embodiments, the large model deployment module 14 can specifically be used to determine the resource requirement information and the target running platform corresponding to each large model to be deployed in the list of models to be deployed based on the computing resource information, the cluster running status, and the dynamic programming strategy to obtain the target deployment plan.

[0093] In some specific embodiments, the large model deployment module 14 may specifically be configured to deploy each of the large models to be deployed to the graphics processor computing resources in the heterogeneous cluster corresponding to the deployment task based on the resource requirement information corresponding to each of the large models to be deployed and the target operating platform; when deploying the collaboration models corresponding to each of the large models to be deployed to the heterogeneous cluster, deploy the collaboration models to the central processor computing resources.

[0094] Furthermore, an embodiment of the present application also discloses an electronic device. Figure 4 It is a structural diagram of an electronic device 20 shown according to an exemplary embodiment, and the content in the figure should not be considered as any limitation to the scope of use of the present application.

[0095] Figure 4 It is a schematic structural diagram of an electronic device 20 provided by an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the large model cluster deployment method disclosed in any of the foregoing embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0096] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows is any communication protocol applicable to the technical solution of the present application, and no specific limitation is imposed on it here; the input / output interface 25 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application requirements, and no specific limitation is made here.

[0097] In addition, as a carrier for resource storage, the memory 22 may be a read-only memory, a random access memory, a magnetic disk, or an optical disc, etc., and the resources stored thereon may include an operating system 221, a computer program 222, etc., and the storage method may be temporary storage or permanent storage.

[0098] Among them, the operating system 221 is used to manage and control each hardware device and the computer program 222 on the electronic device 20, and it may be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program that can be used to complete the large model cluster deployment method executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 222 may further include computer programs that can be used to complete other specific tasks.

[0099] Further, the present application also discloses a computer-readable storage medium for storing a computer program. When the computer program is executed by a processor, the large model cluster deployment method disclosed above is implemented. For the specific steps of this method, reference may be made to the corresponding content disclosed in the foregoing embodiments, and details will not be elaborated herein.

[0100] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference may be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference may be made to the description in the method section.

[0101] Those skilled in the art can further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0102] The steps of the methods or algorithms described in combination with the embodiments disclosed herein can be directly implemented by hardware, software modules executed by a processor, or a combination of the two. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.

[0103] Finally, it should also be noted that in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.

[0104] The above has introduced the technical solution provided by the present application in detail. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A large model cluster deployment method, characterized in that: include: Based on the preset quantization tool, the large model corresponding to the deployment task is quantified in parallel, and the quantified model is evaluated by indicators to obtain the corresponding model evaluation results; Determine a list of models to be deployed corresponding to the deployment task based on the model evaluation results; Obtain computing resource information and cluster operation status of the heterogeneous cluster currently corresponding to the deployment task; The target deployment plan is determined by using the computing resource information, the cluster operation status, the list of models to be deployed, and the dynamic programming strategy, and the large model deployment operation corresponding to the deployment task is completed based on the target deployment plan.

2. The large model cluster deployment method according to claim 1, characterized in that: The parallel quantization of the large model corresponding to the deployment task based on the preset quantization tool includes: Triggering a first model quantization operation of a large model corresponding to the deployment task based on an automatic adaptive weight quantization framework and an activation value-based weight quantization algorithm to obtain a first quantization result; A second model quantization operation of the large model corresponding to the deployment task is triggered based on an offline quantization algorithm to obtain a second quantization result.

3. The large model cluster deployment method according to claim 2, characterized in that: The indicator evaluation of the quantized model to obtain the corresponding model evaluation results includes: Using the Wikipedia text dataset and the first quantization result, the quantized model is evaluated on a model confusion index, a first word delay index of model reasoning, and a character generation speed index to obtain a corresponding first model evaluation result; The Wikipedia text dataset and the second quantization result are used to evaluate the model confusion index, the first word delay index of model reasoning, and the character generation speed index of the quantized model to obtain the corresponding second model evaluation result.

4. The large model cluster deployment method according to claim 1, characterized in that: The determining, based on the model evaluation results, a list of models to be deployed corresponding to the deployment task comprises: The model evaluation results are visualized on the front-end interface to determine the list of models to be deployed corresponding to the deployment task.

5. The large model cluster deployment method according to claim 1, characterized in that: The obtaining of computing resource information and cluster operation status of the heterogeneous cluster currently corresponding to the deployment task includes: The computing resource information and cluster operation status of the heterogeneous cluster currently corresponding to the deployment task are obtained by integrating the Grafana monitoring component and the Prometheus monitoring component in the container cluster management system.

6. The large model cluster deployment method according to any one of claims 1 to 5, characterized in that: The determining of the target deployment scheme by using the computing resource information, the cluster operation status, the list of models to be deployed, and the dynamic programming strategy includes: The resource requirement information and the target operation platform corresponding to each large model to be deployed in the list of models to be deployed are determined based on the computing resource information, the cluster operation status and the dynamic programming strategy to obtain a target deployment plan.

7. The large model cluster deployment method according to claim 6, characterized in that: The completing the large model deployment operation corresponding to the deployment task based on the target deployment scheme includes: Deploy each of the large models to be deployed to the graphics processor computing resources in the heterogeneous cluster corresponding to the deployment task based on the resource demand information and the target operating platform respectively corresponding to each of the large models to be deployed; When the collaborative models corresponding to the large models to be deployed are deployed to the heterogeneous cluster, the collaborative models are deployed to the central processing unit computing resources.

8. A large model cluster deployment device, characterized in that: include: The parallel quantization module is used to perform parallel quantization on the large model corresponding to the deployment task based on the preset quantization tool, and to perform indicator evaluation on the quantized model to obtain the corresponding model evaluation results; A list determination module, used to determine a list of models to be deployed corresponding to the deployment task based on the model evaluation results; An information acquisition module, used to obtain computing resource information and cluster operation status of the heterogeneous cluster currently corresponding to the deployment task; The large model deployment module is used to determine the target deployment plan by using the computing resource information, the cluster operation status, the list of models to be deployed and the dynamic programming strategy, and complete the large model deployment operation corresponding to the deployment task based on the target deployment plan.

9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the large model cluster deployment method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: Used to store a computer program, which, when executed by a processor, implements the large model cluster deployment method as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Model deployment method, electronic equipment, storage medium and program product

    CN121560346A

  • Resource planning method and device, electronic equipment, storage medium and program product

    CN122044892A