Server-agnostic computing-based large model deployment method, device and product

CN117271057BActive Publication Date: 2026-09-22TSINGHUA UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311249504.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-26
Publication Date
2026-09-22
Estimated Expiration
2043-09-26

AI Technical Summary

Technical Problem

[0003]然而,部署大模型需要选择合适的硬件和操作系统的特定类型虚拟机实例,即云虚拟环境方案,而公有云所提供的云虚拟环境方案数量庞大,很难人为地从中选择出合适的云虚拟环境方案,容易导致模型部署成本过高资金浪费,或者配置过低降低模型的推理速度

Benefits of technology

[0073]本申请实施例结合了贝叶斯优化算法和深度强化学习模型,自适应地发现最优的云虚拟环境方案和模型并行化方案。具体的,本申请实施例通过多次迭代,不断地对贝叶斯优化算法的参数进行调整,从而使得贝叶斯优化算法能够学习到目标大模型的推理性能,进而从多个云虚拟环境方案中确定出合适的云虚拟环境方案;并且利用深度强化学习模型,生成最优模型并行化方案,以平衡计算并行性和设备间通信开销。由此,本申请实施例能够同时找到经济高效的云虚拟环境方案和模型并行化方案,实现云虚拟环境和模型并行化的双重联合优化,进一步提高大模型云部署的成本效益。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117271057B_ABST
    Figure CN117271057B_ABST
Patent Text Reader

Abstract

The application provides a large model deployment method, device and product based on server non-aware computing, relates to the technical field of deep learning, and comprises the following steps: obtaining a deployment request of a target large model; sampling from a plurality of cloud virtual environment schemes by using a Bayesian optimization algorithm to obtain a target cloud virtual environment scheme; optimizing a computation graph of the target large model to obtain an optimized computation graph; generating an optimal model parallelization scheme by using a deep reinforcement learning model; deploying the target large model according to the target cloud virtual environment scheme and the optimal model parallelization scheme, executing inference service by using the target large model to obtain an inference performance result; adjusting parameters of the Bayesian optimization algorithm according to the inference performance result; performing multiple iterations; and deploying the target large model according to the target cloud virtual environment scheme and the optimal model parallelization scheme generated in the last iteration process, and executing inference service by using the target large model by adopting a server non-aware computing method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of deep learning technology, and in particular to a method, apparatus and product for deploying large models based on server-insensitive computing. Background Technology

[0002] With the development of deep learning technology, large-scale deep neural network (DNN) models have been widely applied in various fields such as computer vision and speech recognition, becoming an important backend support for many real-time online services. To improve service efficiency, many real-time online services choose to deploy their pre-trained large-scale models in public clouds (such as Alibaba Cloud and Tencent Cloud) and provide corresponding inference services to users.

[0003] However, deploying large models requires selecting specific types of virtual machine instances with suitable hardware and operating systems, i.e., cloud virtual environment solutions. Public clouds offer a vast number of cloud virtual environment solutions, making it difficult to manually select the appropriate one. This can easily lead to excessively high model deployment costs and wasted funds, or under-configuration that reduces the model's inference speed.

[0004] Therefore, it is necessary to develop a method, apparatus, and product for deploying large models based on server-insensitive computing to improve the cost-effectiveness of large model cloud deployment. Summary of the Invention

[0005] In view of the above problems, embodiments of this application provide a method, apparatus and product for deploying large models based on server-insensitive computing, so as to overcome the above problems or at least partially solve the above problems.

[0006] A first aspect of this application provides a method for deploying a large model based on server-insensitive computing, the method comprising:

[0007] Obtain the deployment request of the target large model; the target large model refers to the basic model whose computation graph contains more than 100 million parameters;

[0008] Based on the deployment request, a target cloud virtual environment scheme is obtained by sampling from multiple cloud virtual environment schemes using a Bayesian optimization algorithm.

[0009] The computational graph of the target large model is optimized to obtain the optimized computational graph;

[0010] Using a deep reinforcement learning model, an optimal model parallelization scheme is generated based on the optimized computation graph; the optimal model parallelization scheme represents the parallelization scheme of multiple computation nodes of the target large model on multiple computing devices.

[0011] According to the target cloud virtual environment scheme and the optimal model parallelization scheme, the target large model is deployed, and the inference service is executed using the target large model to obtain the inference performance results of the target large model;

[0012] Based on the inference performance results, the parameters of the Bayesian optimization algorithm are adjusted.

[0013] Repeat the above steps multiple times until the iteration stops.

[0014] Based on the target cloud virtual environment scheme and the optimal model parallelization scheme generated in the last iteration, the target large model is deployed, and the server-insensitive computing method is used to execute inference services using the target large model.

[0015] In one optional implementation, the step of sampling from multiple cloud virtual environment schemes using a Bayesian optimization algorithm to obtain a target cloud virtual environment scheme includes:

[0016] Using the Bayesian optimization algorithm, the inference cost of each cloud virtual environment scheme is predicted; each cloud virtual environment scheme includes at least the following information corresponding to that environment: CPU core count, CPU clock speed, CUDA core count, GPU clock speed, GPU count, and cloud virtual environment memory; the inference cost represents the product of the inference time and the price of the cloud virtual environment scheme.

[0017] The cloud virtual environment scheme with the lowest inference cost among the multiple cloud virtual environment schemes is determined as the target cloud virtual environment scheme.

[0018] In one optional implementation, the Bayesian optimization algorithm includes constraints that indicate that the Bayesian optimization algorithm needs to determine the target cloud virtual environment scheme within a preset search time.

[0019] In one optional implementation, the step of generating an optimal model parallelization scheme based on the optimized computation graph using a deep reinforcement learning model includes:

[0020] The optimized computation graph is encoded into a deep learning operator sequence and input into the deep reinforcement learning model. The deep reinforcement learning model outputs a device sequence, and the devices in the device sequence correspond one-to-one with the deep learning operators in the deep learning operator sequence. The deep learning operator sequence and the device sequence constitute a model parallelization scheme.

[0021] According to the target cloud virtual environment scheme and the model parallelization scheme, the target large model is deployed, and the inference service is executed using the target large model to obtain the inference performance results of the target large model;

[0022] Based on the inference performance results, adjust the model parameters of the deep reinforcement learning model;

[0023] Repeat the above steps multiple times to complete the iterative training of the deep reinforcement learning model and obtain the trained deep reinforcement learning model.

[0024] The deep learning operator sequence is input into the trained deep reinforcement learning model to obtain the target device sequence. The target device sequence and the deep learning operator sequence constitute the optimal model parallelization scheme.

[0025] In one optional implementation, the method of employing server-aware computing to perform inference services using the target large model includes:

[0026] The inference service request input from the user is obtained through a server-side interface that is not aware of the user's input.

[0027] The inference service request is abstracted into a feature vector, and the feature vector is sent to the target large model;

[0028] The target large model performs inference services based on the feature vectors to obtain the inference result;

[0029] The reasoning result is returned to the user terminal.

[0030] In one alternative implementation, the step of performing inference service using the target large model includes:

[0031] In the case of the target large model being launched for the first time, a probe is used to identify and launch the target code block to complete the launch of the target large model; the target code block includes at least: a hardware detection code block, a computation graph construction code block, and a CUDA initialization code block;

[0032] Preserve the startup state of the target code block;

[0033] In the case that the target large model is being launched for the Kth time, the launch status of the target code block is directly obtained to complete the launch of the target large model; where K represents any constant greater than 1.

[0034] After the target large model is started, the inference service is executed.

[0035] In one optional implementation, optimizing the computational graph of the target large model to obtain an optimized computational graph includes:

[0036] Using a tensor algebra super optimizer, one or more source subgraphs in the computation graph of the target large model are replaced with target subgraphs, which are functionally equivalent to the source subgraphs and have higher computational performance than the source subgraphs.

[0037] A second aspect of this application also provides a large-scale model deployment apparatus based on server-insensitive computing, the apparatus comprising:

[0038] The deployment request acquisition module is used to acquire the deployment request of the target large model; the target large model refers to the basic model whose computation graph contains more than 100 million parameters.

[0039] The cloud virtual environment scheme sampling module is used to sample from multiple cloud virtual environment schemes based on the deployment request using a Bayesian optimization algorithm to obtain the target cloud virtual environment scheme.

[0040] The computation graph optimization module is used to optimize the computation graph of the target large model to obtain an optimized computation graph.

[0041] The model parallelization scheme generation module is used to generate an optimal model parallelization scheme based on the optimized computation graph using a deep reinforcement learning model; the optimal model parallelization scheme represents the parallelization scheme of multiple computation nodes of the target large model on multiple computing devices.

[0042] The inference module is used to deploy the target large model according to the target cloud virtual environment scheme and the optimal model parallelization scheme, execute inference services using the target large model, and obtain the inference performance results of the target large model.

[0043] The parameter adjustment module is used to adjust the parameters of the Bayesian optimization algorithm based on the inference performance results.

[0044] The iteration module is used to perform multiple iterations according to the above steps until the iteration stops.

[0045] The deployment module is used to deploy the target large model according to the target cloud virtual environment scheme and the optimal model parallelization scheme generated in the last iteration, and to perform inference services using the target large model by adopting a server-insensitive computing method.

[0046] In one optional implementation, the cloud virtual environment scheme sampling module includes:

[0047] The inference cost prediction submodule is used to predict the inference cost of each cloud virtual environment scheme using the Bayesian optimization algorithm. Each cloud virtual environment scheme includes at least the following information corresponding to the environment: CPU core count, CPU clock speed, CUDA core count, GPU clock speed, GPU count, and cloud virtual environment memory. The inference cost represents the product of the inference time and the price of the cloud virtual environment scheme.

[0048] The determination submodule is used to determine the cloud virtual environment scheme with the lowest inference cost among the multiple cloud virtual environment schemes as the target cloud virtual environment scheme.

[0049] In one optional implementation, the Bayesian optimization algorithm includes constraints that indicate that the Bayesian optimization algorithm needs to determine the target cloud virtual environment scheme within a preset search time.

[0050] In one optional implementation, the model parallelization scheme generation module includes:

[0051] The device sequence output submodule is used to encode the optimized computation graph into a deep learning operator sequence, input it into the deep reinforcement learning model, and the deep reinforcement learning model outputs a device sequence. The devices in the device sequence correspond one-to-one with the deep learning operators in the deep learning operator sequence. The deep learning operator sequence and the device sequence form a model parallelization scheme.

[0052] The inference submodule is used to deploy the target large model according to the target cloud virtual environment scheme and the model parallelization scheme, use the target large model to perform inference services, and obtain the inference performance results of the target large model;

[0053] The parameter tuning submodule is used to adjust the model parameters of the deep reinforcement learning model based on the inference performance results.

[0054] The iterative submodule is used to repeat the above steps multiple times to complete the iterative training of the deep reinforcement learning model and obtain the trained deep reinforcement learning model.

[0055] The optimal model parallelization scheme generation submodule is used to input the deep learning operator sequence into the trained deep reinforcement learning model to obtain the target device sequence, and the target device sequence and the deep learning operator sequence constitute the optimal model parallelization scheme.

[0056] In one optional implementation, the deployment module includes:

[0057] The Acquisition submodule is used to acquire inference service requests input by the user through a server-invisible interface;

[0058] The feature vector generation submodule is used to abstract the inference service request into a feature vector and send the feature vector to the target large model;

[0059] The inference service submodule is used to perform inference services based on the feature vectors of the target large model to obtain inference results;

[0060] The return submodule is used to return the inference result to the user terminal.

[0061] In one optional implementation, the inference module includes:

[0062] The first startup submodule is used to identify and start a target code block using a probe to complete the startup of the target large model when the target large model is being started for the first time; the target code block includes at least: a hardware detection code block, a computation graph construction code block, and a CUDA initialization code block;

[0063] A reserved submodule is used to retain the startup state of the target code block;

[0064] The second startup submodule is used to directly obtain the startup status of the target code block and complete the startup of the target large model when the target large model is being started for the Kth time; where K represents any constant greater than 1.

[0065] The service submodule is used to execute inference services after the target large model has been started.

[0066] In one optional implementation, the computational graph optimization module includes:

[0067] The subgraph replacement submodule is used to replace one or more source subgraphs in the computation graph of the target large model with target subgraphs using the quantitative algebra super optimizer. The target subgraphs are functionally equivalent to the source subgraphs, and the computational performance of the target subgraphs is higher than that of the source subgraphs.

[0068] A third aspect of this application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps in the large-scale model deployment method based on server-insensitive computing described in the first aspect of this application.

[0069] The fourth aspect of this application also provides a computer-readable storage medium having a computer program / instructions stored thereon, which, when executed by a processor, implements the steps in the large-scale model deployment method based on server-insensitive computing described in the first aspect of this application.

[0070] The fifth aspect of this application also provides a computer program product that, when run on an electronic device, causes a processor to execute the steps in the large-scale model deployment method based on server-insensitive computing as described in the first aspect of this application.

[0071] This application provides a method, apparatus, and product for deploying a large model based on server-insensitive computing. The method includes: obtaining a deployment request for a target large model; the target large model represents a basic model whose computation graph contains more than 100 million parameters; based on the deployment request, sampling from multiple cloud virtual environment schemes using a Bayesian optimization algorithm to obtain a target cloud virtual environment scheme; optimizing the computation graph of the target large model to obtain an optimized computation graph; and using a deep reinforcement learning model, generating an optimal model parallelization scheme based on the optimized computation graph; the optimal model parallelization scheme represents multiple computation graphs of the target large model. The parallelization scheme of computing nodes on multiple computing devices; according to the target cloud virtual environment scheme and the optimal model parallelization scheme, the target large model is deployed, and the inference service is executed using the target large model to obtain the inference performance result of the target large model; based on the inference performance result, the parameters of the Bayesian optimization algorithm are adjusted; the above steps are performed multiple times until the iteration stopping condition is reached; according to the target cloud virtual environment scheme and the optimal model parallelization scheme generated in the last iteration, the target large model is deployed, and the server-agnostic computing method is used to execute the inference service using the target large model.

[0072] The specific beneficial effects are as follows:

[0073] This application combines a Bayesian optimization algorithm and a deep reinforcement learning model to adaptively discover the optimal cloud virtual environment scheme and model parallelization scheme. Specifically, this application iterates multiple times to continuously adjust the parameters of the Bayesian optimization algorithm, enabling it to learn the inference performance of the target large model and thus determine a suitable cloud virtual environment scheme from multiple options. Furthermore, it utilizes a deep reinforcement learning model to generate an optimal model parallelization scheme to balance computational parallelism and inter-device communication overhead. Therefore, this application can simultaneously find cost-effective cloud virtual environment and model parallelization schemes, achieving dual joint optimization of cloud virtual environment and model parallelization, further improving the cost-effectiveness of large model cloud deployment. Attached Figure Description

[0074] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0075] Figure 1 This is a flowchart illustrating the steps of a large-scale model deployment method based on server-insensitive computing, as provided in an embodiment of this application.

[0076] Figure 2 This is a flowchart illustrating a large-scale model deployment method provided in an embodiment of this application;

[0077] Figure 3 This is a schematic diagram of the structure of a large-scale model deployment device provided in an embodiment of this application;

[0078] Figure 4 This is a schematic diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0079] Exemplary embodiments of this application will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of this application are shown in the drawings, it should be understood that this application may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of this application and to fully convey the scope of this application to those skilled in the art.

[0080] Deep learning is currently a standard technology used in many fields such as computer vision, speech recognition, and natural language processing. In recent years, large deep neural network models have increasingly become an important backend support for many real-time online services such as Siri and Instagram. These models are pre-trained general-purpose models; specifically, they are general-purpose base models upon which many task-specific models are built through lightweight adaptation (without needing to train from scratch). Simultaneously, due to the massive scale of operations and parameters, they exhibit high prediction accuracy (for example, GPT-3 is a base model with 175 billion parameters, which can be adapted through natural language prompts to build models for language translation, question answering, formula generation, etc.). Since prediction accuracy is guaranteed by large models, the performance of real-time online services mainly depends on the response time for processing user requests, including network transmission time, task scheduling time, and inference time (i.e., the execution time of the DNN inference service). Inference time usually accounts for the majority of the response time, especially for large models. Therefore, inference time is generally considered the main limiting factor for Quality of Service (QoS) in DNN-driven real-time online services.

[0081] To improve service efficiency, many real-time online services choose to deploy their pre-trained large models in public clouds (such as Alibaba Cloud and Tencent Cloud) and provide corresponding inference services to users. However, deploying large models requires selecting specific types of virtual machine instances with suitable hardware and operating systems, i.e., cloud virtual environment solutions. Public clouds offer a vast number of cloud virtual environment solutions (often exceeding 100), making it difficult to manually select a suitable one. This can easily lead to excessively high model deployment costs and wasted funds, or under-configured systems that reduce model inference speed and extend inference service time.

[0082] In view of the above problems, this application proposes a method, apparatus, and product for deploying large models based on server-invisible computing to solve the problem of the inapplicability of the above-mentioned cloud virtual environment solutions, thereby improving the cost-effectiveness of large model cloud deployment. The following, in conjunction with the accompanying drawings, provides a detailed description of the large model deployment method based on server-invisible computing provided by this application through some embodiments and application scenarios.

[0083] The first aspect of this application provides a method for deploying large-scale models based on server-insensitive computing, referring to... Figure 1 , Figure 1 A flowchart illustrating the steps of a large-scale model deployment method based on server-insensitive computing, as provided in this application embodiment, is shown below. Figure 1 As shown, the method includes:

[0084] Step S101: Obtain the deployment request of the target large model, where the target large model refers to the basic model whose computation graph contains more than 100 million parameters.

[0085] In this embodiment, the target large model is a large model, namely a general base model (the computation graph of this base model often contains a huge number of operators and parameters, with the number of parameters known to be over 100 million). After receiving a deployment request from the user, a model for a specific task is built on top of this base model through lightweight adaptation (without needing to train from scratch), such as a language translation, question answering, or formula generation model. Based on the deployment request, the computation graph of the target large model and various cloud virtual environment solutions available for the target large model can be obtained.

[0086] Step S102: Based on the deployment request, a Bayesian optimization algorithm is used to sample from multiple cloud virtual environment schemes to obtain the target cloud virtual environment scheme.

[0087] This application proposes a Bayesian optimization algorithm to adaptively select a suitable target cloud virtual environment scheme from multiple cloud virtual environment schemes. A cloud virtual environment scheme represents a combination of computing resources, i.e., a combination of specific types of cloud servers with different hardware and operating systems. The cloud virtual environment scheme plays a crucial role in the inference performance and operating cost of large models. The Bayesian optimization algorithm is a sequential design strategy for globally optimizing a black-box function. It does not assume any functional form; the goal of the black-box function is to optimize an objective function f(x). This algorithm can query the value of the objective function f(x) at x without knowing any other information (e.g., gradient information) or the specific formula of the objective function f(x).

[0088] In related technologies, the desired cloud virtual environment solution is often selected manually from multiple options. However, when the number of cloud virtual environment solutions is large (generally exceeding 100), manual selection is insufficient to find the optimal solution. Over-configuration leads to wasted resources, while under-configuration reduces inference speed and prolongs inference time. This embodiment utilizes a Bayesian optimization algorithm to adaptively select the optimal target cloud virtual environment solution from multiple options.

[0089] In an optional implementation, step S102, which uses a Bayesian optimization algorithm to sample from multiple cloud virtual environment schemes to obtain a target cloud virtual environment scheme, includes:

[0090] Step S1021: Using the Bayesian optimization algorithm, predict the inference cost of each cloud virtual environment scheme; each cloud virtual environment scheme includes at least the following information corresponding to the environment: CPU core count, CPU clock speed, CUDA core count, GPU clock speed, GPU count, and cloud virtual environment memory; the inference cost represents the product of the inference time and the price of the cloud virtual environment scheme.

[0091] Inference cost is the product of the time required for inference (i.e., the time required for the target large model to execute inference services according to the cloud virtual environment solution) and the price of the cloud virtual environment solution (i.e., the monetary cost of deploying according to the cloud virtual environment solution).

[0092] Step S1022: Among the multiple cloud virtual environment solutions, the cloud virtual environment solution with the lowest inference cost is determined as the target cloud virtual environment solution.

[0093] In this embodiment, the expected gain (i.e. inference cost) of each cloud virtual environment scheme is predicted using a Bayesian optimization algorithm. Then, the cloud virtual environment scheme with the highest expected gain (i.e. the scheme with the lowest inference cost) is selected as the target cloud virtual environment scheme, thereby completing the sampling of cloud virtual environment schemes.

[0094] In one optional implementation, the Bayesian optimization algorithm includes constraints that indicate that the Bayesian optimization algorithm needs to determine the target cloud virtual environment scheme within a preset search time.

[0095] Due to the large number of cloud virtual environment (VM) solutions, simply searching directly for the VM solution with the lowest inference cost would result in excessively high search costs and long search times. To solve the problem within a limited search time, this embodiment proposes adding constraints to the Bayesian optimization algorithm. These constraints limit the search time, requiring the determination of a suitable target VM solution within that timeframe. The search time can be adjusted based on the specific application; however, this embodiment does not impose any restrictions on it.

[0096] Step S103: Optimize the computation graph of the target large model to obtain the optimized computation graph.

[0097] Machine learning frameworks typically abstract the computational operations of large models into computational graphs. A computational graph consists of multiple computational operations, and its structure represents the execution order of these operations. Specifically, a computational graph can be represented as G, consisting of n operations (denoted as O = {O1, O2, ... O...}). n In G, there is a set of directed edges, each edge connecting two operations, representing the dependency between these two operations. If a directed edge connects O...i and O j Then operation O j Only when operating O i Only after completion can we begin. To achieve real-time inference for large models, the configuration cost of cloud virtual environment solutions obtained through Bayesian optimization is generally high, leading to high inference costs after the target large model is deployed. Therefore, this embodiment proposes optimizing the computation graph of the target large model to improve its computational performance and increase the efficiency of the target large model's inference service, thereby reducing inference costs.

[0098] Step S103 involves optimizing the computational graph of the target large model to obtain an optimized computational graph, including:

[0099] Using a tensor algebra super optimizer, one or more source subgraphs in the computation graph of the target large model are replaced with target subgraphs, which are functionally equivalent to the source subgraphs and have higher computational performance than the source subgraphs.

[0100] In this embodiment, some inefficient subgraphs exist in the computational graph of the target large model, which significantly reduce the speed of model inference. This embodiment employs the adaptive graph replacement method of the Tensor Algebra Super Optimizer (TASO). This method optimizes the computational graph by adaptively replacing inefficient subgraphs (source subgraphs) with computationally more efficient subgraphs (target subgraphs) to reduce inference costs.

[0101] Step S104: Using a deep reinforcement learning model, an optimal model parallelization scheme is generated based on the optimized computation graph; the optimal model parallelization scheme represents the parallelization scheme of multiple computation nodes of the target large model on multiple computing devices.

[0102] When deploying a large target model, in addition to a suitable cloud virtual environment solution, the parallelism of computation also needs to be considered. Specifically, for a large model to perform inference services, each computational operation needs to be placed on a corresponding computing device (such as a CPU core or GPU), and the computing device executes the corresponding computational operation. In related technologies, the model parallelization scheme is usually designed manually by the service provider. For large models with very large computational graphs, there are multiple model parallelization schemes to choose from, and it is difficult to determine the optimal or near-optimal model parallelization scheme manually. Deep reinforcement learning refers to an artificial intelligence method that combines the perception ability of deep learning with the decision-making ability of reinforcement learning, and directly controls based on the input. This embodiment uses a deep reinforcement learning model to determine the optimal model parallelization scheme, which can be represented as P = (p1, p2, ... p... nFor any computation operation O in computation graph G, i There exists a corresponding p in P. i O i The devices are placed accordingly. Therefore, this embodiment, through an optimal model parallelization scheme, explicitly places the computational operations of the target large model computation graph on multiple computing devices such as GPUs and CPUs to accelerate inference and balance computational parallelism with inter-device communication overhead.

[0103] Step S104, utilizing a deep reinforcement learning model and based on the optimized computation graph, generates an optimal model parallelization scheme, including:

[0104] Step S1041: Encode the optimized computation graph into a deep learning operator sequence, input it into the deep reinforcement learning model, and the deep reinforcement learning model outputs a device sequence. The devices in the device sequence correspond one-to-one with the deep learning operators in the deep learning operator sequence. The deep learning operator sequence and the device sequence constitute a model parallelization scheme.

[0105] In this embodiment, the information of all computational operations in the optimized computation graph is encoded as a data sequence (deep learning operator sequence), for example, O = {O1, O2, ... O}. n This allows the input to a deep reinforcement learning model to construct the output of the model as a sequence of devices corresponding to the input deep learning operator sequence, for example, P = (p1, p2, ... p...). n The elements in the two sequences correspond one-to-one. For any deep learning operator O in the input deep learning operator sequence... i There exists a corresponding element p in the output device sequence P. i O i The equipment that is placed there.

[0106] Step S1042: Deploy the target large model according to the target cloud virtual environment scheme and the model parallelization scheme, and use the target large model to execute inference services to obtain the inference performance results of the target large model.

[0107] Following the target cloud virtual environment scheme, the target large model is deployed on corresponding CPUs and GPUs. Furthermore, according to the correspondence between elements in the deep learning operator sequence and device sequence in the model parallelization scheme, the computational operations of the target large model are deployed on the corresponding hardware devices. After deployment, the target large model can be used to execute corresponding inference services. Based on the results of the inference service, the inference performance of the target large model can be tested. These performance results represent the inference performance of the target large model, such as the time required to execute the inference service, inference speed, and accuracy of the inference results.

[0108] Step S1043: Adjust the model parameters of the deep reinforcement learning model based on the inference performance results.

[0109] Step S1044: Repeat the above steps multiple times to complete the iterative training of the deep reinforcement learning model and obtain the trained deep reinforcement learning model. Specifically, training can be stopped when a preset number of training iterations or model convergence conditions are reached, resulting in the trained deep reinforcement learning model.

[0110] Step S1045: Input the deep learning operator sequence into the trained deep reinforcement learning model to obtain the target device sequence. The target device sequence and the deep learning operator sequence constitute the optimal model parallelization scheme.

[0111] This embodiment utilizes a Bayesian optimization algorithm to determine the next cloud virtual environment scheme to be sampled (the target cloud virtual environment scheme) to reduce inference costs. After sampling the target cloud virtual environment scheme, the computation graph of the target large model is determined based on this scheme. Based on this computation graph, a deep reinforcement learning model is iteratively trained to determine the optimal model parallelization scheme. In this embodiment, the model parameters of the deep reinforcement learning model are adjusted based on the obtained inference performance results. After multiple iterations, the training of the deep reinforcement learning model is completed, enabling it to learn the inference performance of the target large model. This allows it to determine the optimal model parallelization scheme from multiple model parallelization schemes based on the input deep learning operator sequence, thereby improving the inference performance of the target large model after deployment.

[0112] Step S105: Deploy the target large model according to the target cloud virtual environment scheme and the optimal model parallelization scheme, and use the target large model to execute inference services to obtain the inference performance results of the target large model.

[0113] According to the target cloud virtual environment scheme, corresponding CPU and GPU resources are configured for the target large model. Following the correspondence between the deep learning operator sequence and device sequence elements in the optimal model parallelization scheme, the computational operations of the target large model are deployed on the corresponding hardware devices. After deployment, the target large model can be used to execute corresponding inference services. Based on the results of the inference services, the inference performance of the target large model can be tested. This performance result represents the inference performance of the target large model after deployment according to the target cloud virtual environment scheme and the optimal model parallelization scheme, such as the time required to execute the inference service, the inference speed, and the accuracy of the inference results.

[0114] Step S106: Adjust the parameters of the Bayesian optimization algorithm based on the inference performance results.

[0115] Step S107: Perform multiple iterations following the steps described above until the iteration stopping condition is met. Specifically, the iteration stopping condition can be a preset number of iterations or an iteration time.

[0116] Reference Figure 2 , Figure 2 A flowchart illustrating a large-scale model deployment method is shown, such as... Figure 2 As shown, for the deployment request of the target large model, several cloud virtual environment schemes (corresponding to) are selected from the cloud server pool. Figure 2 In process node 1), the Bayesian optimization algorithm samples from multiple cloud virtual environment schemes to determine the target cloud virtual environment scheme (corresponding to...). Figure 2 In process node 2), the computational graph of the target large model is optimized using the TASO tensor algebra super optimizer. The optimized computational graph is then input into the deep reinforcement learning model. Through multiple iterations of training, the optimal model parallelization scheme is selected from multiple model parallelization schemes (corresponding to...). Figure 2 In process node 3), the target large model is then deployed according to the obtained target cloud virtual environment scheme and optimal model parallelization scheme. Real-time online inference services are executed in an execution environment based on server-insensitive computing. Inference performance results are generated based on the inference results, and the parameters of the Bayesian optimization algorithm are adjusted and optimized using these results (corresponding to...). Figure 2 In process node 4), when performing the next sampling to determine the target cloud virtual environment scheme, a more suitable cloud virtual environment scheme can be selected. For example... Figure 2 As shown, completing one process node 2, 3, and 4 is considered as completing one training cycle. Through multiple iterations, the parameters of the Bayesian optimization algorithm are fine-tuned multiple times to enable it to learn the performance of the target large model. Thus, within a limited search time, the cloud virtual environment scheme with the lowest inference cost is selected as the target cloud virtual environment scheme from multiple cloud virtual environment schemes.

[0117] Step S108: Deploy the target large model according to the target cloud virtual environment scheme and the optimal model parallelization scheme generated in the last iteration, and use the server-insensitive computing method to execute inference services using the target large model.

[0118] This application combines Bayesian optimization algorithms and deep reinforcement learning models to adaptively discover the optimal cloud virtual environment scheme and model parallelization scheme. Specifically, this application iterates multiple times to continuously adjust the parameters of the Bayesian optimization algorithm, enabling it to learn the inference performance of the target large model and thus determine a suitable cloud virtual environment scheme from multiple options. Furthermore, it utilizes a deep reinforcement learning model to generate the optimal model parallelization scheme, balancing computational parallelism and inter-device communication overhead. Therefore, this application, by jointly utilizing Bayesian optimization and deep reinforcement learning techniques, can adaptively find cost-effective cloud virtual environment and model parallelization schemes within a limited search time, achieving dual joint optimization of cloud virtual environment and model parallelization. This minimizes inference costs while meeting search time constraints, further improving the cost-effectiveness of large model cloud deployment.

[0119] In one optional implementation, the method of employing server-aware computing to perform inference services using the target large model includes:

[0120] The inference service request input from the user is obtained through a server-side interface that is not aware of the user's input.

[0121] The inference service request is abstracted into a feature vector, and the feature vector is sent to the target large model.

[0122] The target large model performs inference services based on the feature vectors to obtain the inference results.

[0123] The reasoning result is returned to the user terminal.

[0124] During the execution of inference services, the inference service request input by the user is often directly transmitted to the server, which then uses the deployed target model to perform all the calculations to complete the inference service. For example, in a face recognition service, the user sends the collected face image and the face recognition request to the server, which analyzes and calculates the face image to generate the face recognition result. However, this method does not consider the protection of user data privacy and is prone to leaking user information to the server.

[0125] This embodiment proposes utilizing server-insensitive computing technology to obtain inference service requests input from the user client through a server-insensitive interface. These requests include user data required to execute the inference service. The user-input data is stored on the client-side of the large-scale model inference service via the server-insensitive interface. The inference service request is abstracted into a feature vector that cannot be reverse-decoded, and this feature vector is then sent to the target large-scale model deployed on the server. Finally, the target large-scale model performs inference based on the input feature vector and returns the inference result to the user client. For example, in a face recognition service, the user client sends the collected face image and face recognition request to the server-insensitive interface. The server-insensitive interface abstracts the face image into a feature vector that cannot be reverse-decoded, and this feature vector is input into the target large-scale model on the server for face recognition analysis, generating a face recognition result. This avoids the server obtaining the user-input data, protecting user privacy and improving user information security.

[0126] In one alternative implementation, the step of performing inference service using the target large model includes:

[0127] When the target large model is being launched for the first time, a probe is used to identify and launch the target code block to complete the launch of the target large model; the target code block includes at least: a hardware detection code block, a computation graph construction code block, and a CUDA initialization code block.

[0128] Preserve the startup state of the target code block.

[0129] When the target large model is launched for the Kth time, the launch status of the target code block is directly obtained to complete the launch of the target large model; where K represents any constant greater than 1.

[0130] After the target large model is started, the inference service is executed.

[0131] This application's embodiments combine Bayesian optimization algorithms and deep reinforcement learning models, requiring multiple iterations and repeated runs of time-consuming inference service experiments, resulting in low deployment efficiency for the target large model. Dynamic tracing of the basic code blocks within the target large model reveals that the root cause of the runtime issue lies in the need to initiate heavy startup code blocks each time the target large model is started. These include hardware detection code blocks, computation graph construction code blocks, and Compute Unified Device Architecture (CUDA) initialization code blocks. These code blocks consume a significant amount of startup time; while unrelated to the inference service input, they account for over 98% of the execution time for the inference service.

[0132] To improve inference service efficiency, this embodiment improves the original target large model startup mechanism. Specifically, during the first startup of the target large model, a probe is used to identify target code blocks within the code block structure. These code blocks are then pre-executed to complete the initial startup of the target large model. Furthermore, the startup state of the target code blocks is retained. Therefore, during subsequent startups of the target large model, the startup state of these target code blocks can be directly utilized to complete the startup, eliminating the need to restart these heavy code blocks and saving time spent repeatedly starting them. This allows for reuse of the startup foundation during subsequent inference service executions, thereby reducing search overhead.

[0133] For example, in multiple iterations, after each execution of step S105, following the deployment of the target large model according to the target cloud virtual environment scheme and the optimal model parallelization scheme, the target large model needs to be started to perform inference services and obtain the inference performance results of the target large model. During the first execution of step S105, a probe is used to identify and start the target code block to complete the startup of the target large model. During the k-th execution of step S105, the startup status of the target code block retained from the first execution of step S105 is directly obtained to complete the startup of the target large model, thereby reducing the time for the target large model to perform inference services during multiple iterations.

[0134] In one optional implementation, during step S104, the deep reinforcement learning model needs to be trained iteratively multiple times. During each training iteration, step S1042 needs to be executed to deploy and start the target large model according to the target cloud virtual environment scheme and model parallelization scheme. This allows the target large model to perform inference services and obtain its inference performance results. Therefore, during step S104, the target large model needs to be started multiple times. Specifically, during the first execution of step S1042, a probe is used to identify and start the target code block to complete the startup of the target large model. During the k-th execution of step S1042, the startup state of the target code block retained from the first execution of step S1042 is directly obtained, and the startup basis is reused to complete the startup of the target large model. This reduces the time required for the target large model to perform inference services during multiple iterations of training the deep reinforcement learning model, thus improving inference efficiency.

[0135] A second aspect of this application also provides a large-scale model deployment apparatus based on server-insensitive computing, referring to... Figure 3 , Figure 3 A schematic diagram of a large-scale deployment device is shown, such as... Figure 3 As shown, the device includes:

[0136] The deployment request acquisition module is used to acquire the deployment request of the target large model; the target large model refers to the basic model whose computation graph contains more than 100 million parameters.

[0137] The cloud virtual environment scheme sampling module is used to sample from multiple cloud virtual environment schemes based on the deployment request using a Bayesian optimization algorithm to obtain the target cloud virtual environment scheme.

[0138] The computation graph optimization module is used to optimize the computation graph of the target large model to obtain an optimized computation graph.

[0139] The model parallelization scheme generation module is used to generate an optimal model parallelization scheme based on the optimized computation graph using a deep reinforcement learning model; the optimal model parallelization scheme represents the parallelization scheme of multiple computation nodes of the target large model on multiple computing devices.

[0140] The inference module is used to deploy the target large model according to the target cloud virtual environment scheme and the optimal model parallelization scheme, execute inference services using the target large model, and obtain the inference performance results of the target large model.

[0141] The parameter adjustment module is used to adjust the parameters of the Bayesian optimization algorithm based on the inference performance results.

[0142] The iteration module is used to perform multiple iterations according to the above steps until the iteration stops.

[0143] The deployment module is used to deploy the target large model according to the target cloud virtual environment scheme and the optimal model parallelization scheme generated in the last iteration, and to perform inference services using the target large model by adopting a server-insensitive computing method.

[0144] In one optional implementation, the cloud virtual environment scheme sampling module includes:

[0145] The inference cost prediction submodule is used to predict the inference cost of each cloud virtual environment scheme using the Bayesian optimization algorithm. Each cloud virtual environment scheme includes at least the following information corresponding to the environment: CPU core count, CPU clock speed, CUDA core count, GPU clock speed, GPU count, and cloud virtual environment memory. The inference cost represents the product of the inference time and the price of the cloud virtual environment scheme.

[0146] The determination submodule is used to determine the cloud virtual environment scheme with the lowest inference cost among the multiple cloud virtual environment schemes as the target cloud virtual environment scheme.

[0147] In one optional implementation, the Bayesian optimization algorithm includes constraints that indicate that the Bayesian optimization algorithm needs to determine the target cloud virtual environment scheme within a preset search time.

[0148] In one optional implementation, the model parallelization scheme generation module includes:

[0149] The device sequence output submodule is used to encode the optimized computation graph into a deep learning operator sequence, input it into the deep reinforcement learning model, and the deep reinforcement learning model outputs a device sequence. The devices in the device sequence correspond one-to-one with the deep learning operators in the deep learning operator sequence. The deep learning operator sequence and the device sequence form a model parallelization scheme.

[0150] The inference submodule is used to deploy the target large model according to the target cloud virtual environment scheme and the model parallelization scheme, use the target large model to perform inference services, and obtain the inference performance results of the target large model;

[0151] The parameter tuning submodule is used to adjust the model parameters of the deep reinforcement learning model based on the inference performance results.

[0152] The iterative submodule is used to repeat the above steps multiple times to complete the iterative training of the deep reinforcement learning model and obtain the trained deep reinforcement learning model.

[0153] The optimal model parallelization scheme generation submodule is used to input the deep learning operator sequence into the trained deep reinforcement learning model to obtain the target device sequence, and the target device sequence and the deep learning operator sequence constitute the optimal model parallelization scheme.

[0154] In one optional implementation, the deployment module includes:

[0155] The Acquisition submodule is used to acquire inference service requests input by the user through a server-invisible interface;

[0156] The feature vector generation submodule is used to abstract the inference service request into a feature vector and send the feature vector to the target large model;

[0157] The inference service submodule is used to perform inference services based on the feature vectors of the target large model to obtain inference results;

[0158] The return submodule is used to return the inference result to the user terminal.

[0159] In one optional implementation, the inference module includes:

[0160] The first startup submodule is used to identify and start a target code block using a probe to complete the startup of the target large model when the target large model is being started for the first time; the target code block includes at least: a hardware detection code block, a computation graph construction code block, and a CUDA initialization code block;

[0161] A reserved submodule is used to retain the startup state of the target code block;

[0162] The second startup submodule is used to directly obtain the startup status of the target code block and complete the startup of the target large model when the target large model is being started for the Kth time; where K represents any constant greater than 1.

[0163] The service submodule is used to execute inference services after the target large model has been started.

[0164] In one optional implementation, the computational graph optimization module includes:

[0165] The subgraph replacement submodule is used to replace one or more source subgraphs in the computation graph of the target large model with target subgraphs using a tensor algebra super optimizer. The target subgraphs are functionally equivalent to the source subgraphs, and the computational performance of the target subgraphs is higher than that of the source subgraphs.

[0166] For example, a server-agnostic computation-based large model deployment system as described in this embodiment is implemented based on TensorFlow, and a prototype system is built on a server instance rented from Microsoft Azure. To comprehensively evaluate the effectiveness and adaptability of the system, extensive experiments were conducted using commonly used lightweight DNN models and large models, including natural language processing models (Recurrent Neural Network Language Model RNNLM and Bidirectional Encoder Representation Large Model BERT) and online image classification models (Third-generation Inception Deep Convolutional Neural Network Model Inception-V3 and 19-layer Deep Convolutional Neural Network Model VGG19). Experimental results show that, compared with related technologies, the large model deployment system proposed in this embodiment improves the inference speed of lightweight DNN models by 60% and reduces optimization overhead by 98%; it improves the inference speed of large models by 21% and reduces optimization overhead by 84%. Furthermore, compared with heuristic baselines such as greedy search, it reduces inference cost by up to 51% and search time by 57% for lightweight DNN models, and saves 74% inference cost and 38% search time for the base model.

[0167] This application also provides an electronic device, see embodiments thereof. Figure 4 , Figure 4 This is a schematic diagram of the electronic device proposed in an embodiment of this application. Figure 4 As shown, the electronic device 100 includes a memory 110 and a processor 120. The memory 110 and the processor 120 are connected via a bus. The memory 110 stores a computer program that can run on the processor 120 to implement the steps in the server-insensitive computing-based large model deployment method disclosed in the embodiments of this application.

[0168] This application also provides a computer-readable storage medium storing a computer program / instructions thereon, which, when executed by a processor, implements the steps in the large-scale model deployment method based on server-insensitive computing disclosed in this application.

[0169] This application also provides a computer program product that, when run on an electronic device, enables the processor to execute the steps in the large-scale model deployment method based on server-insensitive computing disclosed in this application.

[0170] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0171] This application describes embodiments with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, electronic devices, and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0172] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0173] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal equipment, causing a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0174] Although preferred embodiments of the present application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present application.

[0175] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.

[0176] The above provides a detailed description of a large-scale model deployment method, apparatus, and product based on server-insensitive computing provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for deploying large-scale models based on server-insensitive computing, characterized in that, The method includes: Step S101: Obtain the deployment request of the target large model, where the target large model refers to the basic model whose computation graph contains more than 100 million parameters. Step S102: Based on the deployment request, the target cloud virtual environment scheme is obtained by sampling from multiple cloud virtual environment schemes using the Bayesian optimization algorithm. Step S103: Optimize the computation graph of the target large model to obtain the optimized computation graph; Step S104: Using a deep reinforcement learning model, an optimal model parallelization scheme is generated based on the optimized computation graph; the optimal model parallelization scheme represents the parallelization scheme of multiple computation nodes of the target large model on multiple computing devices. Step S105: Deploy the target large model according to the target cloud virtual environment scheme and the optimal model parallelization scheme, and use the target large model to execute inference services to obtain the inference performance results of the target large model; Step S106: Adjust the parameters of the Bayesian optimization algorithm based on the inference performance results; Step S107: Perform multiple iterations according to steps S102 to S106 above until the iteration stop condition is met. Step S108: According to the target cloud virtual environment scheme and the optimal model parallelization scheme generated in the last iteration, deploy the target large model, adopt the server-insensitive computing method, and use the target large model to perform inference services; The process of sampling from multiple cloud virtual environment schemes using a Bayesian optimization algorithm to obtain the target cloud virtual environment scheme includes: Using the Bayesian optimization algorithm, the inference cost of each cloud virtual environment scheme is predicted; each cloud virtual environment scheme includes at least the following information corresponding to that environment: CPU core count, CPU clock speed, CUDA core count, GPU clock speed, GPU count, and cloud virtual environment memory; the inference cost represents the product of the inference time and the price of the cloud virtual environment scheme. The cloud virtual environment scheme with the lowest inference cost among the multiple cloud virtual environment schemes is determined as the target cloud virtual environment scheme.

2. The large-scale model deployment method based on server-insensitive computing according to claim 1, characterized in that, The Bayesian optimization algorithm includes constraints, which indicate that the Bayesian optimization algorithm needs to determine the target cloud virtual environment scheme within a preset search time.

3. The method for deploying large-scale models based on server-insensitive computing according to claim 1, characterized in that, The step of generating an optimal model parallelization scheme based on the optimized computation graph using a deep reinforcement learning model includes: Step S1041: Encode the optimized computation graph into a deep learning operator sequence, input it into the deep reinforcement learning model, and the deep reinforcement learning model outputs a device sequence. The devices in the device sequence correspond one-to-one with the deep learning operators in the deep learning operator sequence. The deep learning operator sequence and the device sequence form a model parallelization scheme. Step S1042: Deploy the target large model according to the target cloud virtual environment scheme and the model parallelization scheme, and use the target large model to execute inference services to obtain the inference performance results of the target large model; Step S1043: Adjust the model parameters of the deep reinforcement learning model based on the inference performance results; Step S1044: Repeat steps S1041 to S1043 multiple times to complete the iterative training of the deep reinforcement learning model and obtain the trained deep reinforcement learning model. Step S1045: Input the deep learning operator sequence into the trained deep reinforcement learning model to obtain the target device sequence. The target device sequence and the deep learning operator sequence constitute the optimal model parallelization scheme.

4. The method for deploying large-scale models based on server-insensitive computing according to claim 1, characterized in that, The method of server-insensitive computing, which utilizes the target large model to perform inference services, includes: The inference service request input from the user is obtained through a server-side interface that is not aware of the user's input. The inference service request is abstracted into a feature vector, and the feature vector is sent to the target large model; The target large model performs inference services based on the feature vectors to obtain the inference result; The reasoning result is returned to the user terminal.

5. The method for deploying large models based on server-insensitive computing according to claim 1, characterized in that, The execution of inference services using the target large model includes: In the case of the target large model being launched for the first time, a probe is used to identify and launch the target code block to complete the launch of the target large model; the target code block includes at least: a hardware detection code block, a computation graph construction code block, and a CUDA initialization code block; Preserve the startup state of the target code block; In the case that the target large model is being launched for the Kth time, the launch status of the target code block is directly obtained to complete the launch of the target large model; where K represents any constant greater than 1. After the target large model is started, the inference service is executed.

6. The method for deploying large models based on server-insensitive computing according to claim 1, characterized in that, The optimization of the computational graph of the target large model to obtain the optimized computational graph includes: Using a tensor algebra super optimizer, one or more source subgraphs in the computation graph of the target large model are replaced with target subgraphs, which are functionally equivalent to the source subgraphs and have higher computational performance than the source subgraphs.

7. A large-scale model deployment device based on server-insensitive computing, characterized in that, The device includes: The deployment request acquisition module is used to acquire the deployment request of the target large model; the target large model refers to the basic model whose computation graph contains more than 100 million parameters. The cloud virtual environment scheme sampling module is used to sample from multiple cloud virtual environment schemes based on the deployment request using a Bayesian optimization algorithm to obtain the target cloud virtual environment scheme. The computation graph optimization module is used to optimize the computation graph of the target large model to obtain an optimized computation graph. The model parallelization scheme generation module is used to generate an optimal model parallelization scheme based on the optimized computation graph using a deep reinforcement learning model; the optimal model parallelization scheme represents the parallelization scheme of multiple computation nodes of the target large model on multiple computing devices. The inference module is used to deploy the target large model according to the target cloud virtual environment scheme and the optimal model parallelization scheme, execute inference services using the target large model, and obtain the inference performance results of the target large model. The parameter adjustment module is used to adjust the parameters of the Bayesian optimization algorithm based on the inference performance results. The iteration module is used to repeatedly execute the cloud virtual environment scheme sampling module, the computation graph optimization module, the model parallelization scheme generation module, the inference module, and the parameter adjustment module until the iteration stop condition is met; The deployment module is used to deploy the target large model according to the target cloud virtual environment scheme and the optimal model parallelization scheme generated in the last iteration, and to use the target large model to perform inference services by employing a server-insensitive computing method. The cloud virtual environment solution sampling module includes: The inference cost prediction submodule is used to predict the inference cost of each cloud virtual environment scheme using the Bayesian optimization algorithm. Each cloud virtual environment scheme includes at least the following information corresponding to the environment: CPU core count, CPU clock speed, CUDA core count, GPU clock speed, GPU count, and cloud virtual environment memory. The inference cost represents the product of the inference time and the price of the cloud virtual environment scheme. The determination submodule is used to determine the cloud virtual environment scheme with the lowest inference cost among the multiple cloud virtual environment schemes as the target cloud virtual environment scheme.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps in the large-scale model deployment method based on server-insensitive computing as described in any one of claims 1-6.

9. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps in the large-scale model deployment method based on server-insensitive computing as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Marketing activity prediction model structure and prediction method based on knowledge distillation

    CN112967088A

  • Method and system for developing a machine learning model

    US20210304073A1