Model deployment method and device, computer equipment and storage medium
By automatically determining the target model from the model square, and using dynamic mapping tables and computing resource indicator data for resource allocation, the resource shortage caused by user estimation deviation is solved, and efficient resource utilization and fault self-healing capabilities are achieved.
Patent Information
- Application Number
- CN202510419833.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-25
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing model deployment system relies on users to estimate computing resources, resulting in inaccurate resource allocation, shortage or excess of resources, and reduces the overall resource utilization rate.
By responding to model deployment requests, the target model is automatically determined from the model square, and the dynamic mapping table is used to obtain model support information and computing power resource index data of the computing nodes. The dynamic mapping table and scheduling algorithm are used to accurately allocate resources, and dynamically adjust resource allocation to deal with failures.
It realizes flexible allocation of computing power resources, avoids insufficient resources, improves overall resource utilization, and realizes efficient self-healing of faults when computing node failures, ensuring business continuity.
Smart Images

Figure CN120371470A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technologies, and in particular, to a model deployment method, apparatus, computer device, and storage medium. Background Art
[0002] With the rapid development of large models, the demand for model training and deployment has shown an exponential growth trend. Model training and deployment require a large amount of computing resources (such as GPUs, CPUs, memory, etc.). The scheduling system is responsible for managing and allocating these computing resources, and it needs to reasonably allocate resources according to the requirements of model training tasks to ensure that each task can obtain sufficient resources to operate efficiently.
[0003] The current scheduling system mainly relies on users to estimate the computing power resources required for deploying the model and manually select specific computing power brands and inference frameworks for resource allocation. However, due to possible deviations in users' understanding of the required computing power resources, the resource estimation is inaccurate. And when the computing power resources of the specified brand are insufficient or fail, the idle computing power resources cannot be effectively utilized, thus reducing the overall resource utilization rate. Summary of the Invention
[0004] This application proposes a model deployment method, apparatus, computer device, and storage medium to improve the overall resource utilization rate.
[0005] In a first aspect, a model deployment method is provided, including:
[0006] Respond to a model deployment request and determine a target model from a model plaza;
[0007] Obtain model support information of the target model based on a dynamic mapping table;
[0008] Obtain the computing power resource index data of each computing node corresponding to the model support information;
[0009] Deploy the target model based on the computing power resource index data of each computing node and the dynamic mapping table.
[0010] Optionally, the step of responding to a model deployment request and determining a target model from a model plaza includes:
[0011] Respond to a model deployment request and obtain the classification label in the model deployment request;
[0012] Determine the target model from the model plaza based on the classification label.
[0013] Optionally, the step of responding to a model deployment request and determining a target model from a model plaza includes:
[0014] In response to a model deployment request, obtain user preference information and / or historical record information corresponding to the model deployment request;
[0015] Determine the target model from the model square based on the user preference information and / or the historical record information.
[0016] Optionally, the deploying the target model based on the computing power resource metric data of each computing node and the dynamic mapping table includes:
[0017] Determine a target computing node based on the computing power resource metric data of each computing node and the dynamic mapping table, and determine a target inference framework in the node information of the target computing node;
[0018] Deploy the target model based on the node information of the target computing node and the target inference framework.
[0019] Optionally, the determining a target computing node based on the computing power resource metric data of each computing node and the dynamic mapping table includes:
[0020] Determine a target scheduling algorithm based on the computing power resource metric data of each computing node and the dynamic mapping table;
[0021] Determine the target computing node according to the target scheduling algorithm.
[0022] Optionally, the model deployment method further includes:
[0023] Monitor the target computing node;
[0024] If the target computing node fails, update the dynamic mapping table;
[0025] Update the target computing node of the target model according to the updated dynamic mapping table;
[0026] Migrate the workload of the failed target computing node to the updated target computing node for execution.
[0027] Optionally, the model deployment method further includes:
[0028] In response to a model viewing request, obtain model detailed information corresponding to the model viewing request;
[0029] Display the model detailed information.
[0030] In a second aspect, there is provided a model deployment apparatus, including:
[0031] A determination module, configured to determine a target model from a model square in response to a model deployment request;
[0032] A first acquisition module, configured to acquire model support information of the target model based on a dynamic mapping table;
[0033] A second acquisition module, configured to acquire computing power resource index data of each computing node corresponding to the model support information;
[0034] A deployment module, configured to deploy the target model based on the computing power resource index data of each computing node and the dynamic mapping table.
[0035] In a third aspect, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above model deployment method are implemented.
[0036] In a fourth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above model deployment method are implemented.
[0037] This application provides a model deployment method, device, computer device, and storage medium. By responding to a model deployment request, a target model is determined from a model square; model support information of the target model is acquired based on a dynamic mapping table; computing power resource index data of each computing node corresponding to the model support information is acquired; and the target model is deployed based on the computing power resource index data of each computing node and the dynamic mapping table. In the model deployment solution provided in this application, when a model deployment request is received, first, the model deployment request can be responded to and the target model can be automatically determined from the model square. Then, the model support information of the target model can be automatically determined through the dynamic mapping table, without the user manually selecting a computing power brand and an inference framework, avoiding insufficient or excessive computing power resources caused by the user's estimation deviation of computing power resources. Next, the computing power resource index data of each computing node corresponding to the model support information can be automatically acquired, and the computing power resources can be accurately and flexibly allocated to the target model and deployed in combination with the dynamic mapping table, which can avoid the occurrence of insufficient computing power resources. And when a computing node fails, the computing power resource allocation of the target model can be dynamically adjusted according to the computing power resource index data of each computing node, thereby improving the overall resource utilization rate. Description of the Drawings
[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0039] Figure 1 It is an application environment diagram of the model deployment method provided by an embodiment of this application;
[0040] Figure 2 It is a flowchart of the model deployment method provided by an embodiment of this application;
[0041] Figure 3 It is a flowchart of the model deployment method provided by another embodiment of this application;
[0042] Figure 4 It is a structural block diagram of the model deployment system provided by an embodiment of this application;
[0043] Figure 5 It is a structural block diagram of the model deployment device provided by an embodiment of this application;
[0044] Figure 6 It is a structural block diagram of the computer device provided by an embodiment of this application. Detailed implementation manners
[0045] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0046] In addition, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, many specific details are provided to give a thorough understanding of the embodiments of the present application. However, those skilled in the art will realize that the technical solutions of the present application may be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be used. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of the present application.
[0047] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0048] The flowcharts shown in the drawings are only illustrative descriptions and do not necessarily include all contents and operations / steps, nor do they necessarily need to be executed in the described order. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined. Therefore, the actual execution order may change according to the actual situation.
[0049] The model deployment method provided by the embodiments of the present invention can be applied in an application environment such as Figure 1 . Among them, the computer device 110 communicates with the server 120 through the network 130. The computer device 110 can respond to the model deployment request, determine the target model from the model square; obtain the model support information of the target model based on the dynamic mapping table; obtain the computing power resource index data of each computing node corresponding to the model support information; deploy the target model based on the computing power resource index data of each computing node and the dynamic mapping table, and display it through the computer device 110. In the present invention, when receiving the model deployment request, first, the target model can be automatically determined from the model square in response to the model deployment request. Then, the model support information of the target model can be automatically determined through the dynamic mapping table, without the user manually selecting the computing power brand and inference framework, avoiding insufficient or excessive computing power resources caused by the user's estimation deviation of the computing power resources. Then, the computing power resource index data of each computing node corresponding to the model support information can be automatically obtained, and the computing power resources can be accurately and flexibly allocated to the target model and deployed in combination with the dynamic mapping table, which can avoid the occurrence of insufficient computing power resources. And when a computing node fails, the computing power resource allocation of the target model can be dynamically adjusted according to the computing power resource index data of each computing node, so as to improve the overall resource utilization rate. Among them, the computer device 110 can be but is not limited to various smart phones 110-1, tablet computers 110-2, and laptop computers 110-3. The present invention will be described in detail through specific embodiments below.
[0050] Please refer to Figure 2 shown in Figure 2 is a schematic flow chart of the model deployment method provided by the embodiments of the present invention. This method can be applied to both terminals and servers. In this embodiment, it is exemplified by being applied to the server side. The model deployment method includes the following steps:
[0051] S101: Respond to the model deployment request and determine the target model from the model square.
[0052] Among them, the model deployment request represents a request to deploy the trained target model to the production environment to provide actual application services. The model square represents a platform or resource library that centrally displays, manages, and provides various pre-registered models to be trained.
[0053] Exemplarily, assume that the model deployment request is to deploy a deep learning-based image classification model for identifying animal categories (such as cats, dogs, birds, etc.) in images. If the model square contains deep learning models, then the model deployment request can be responded to, and the deep learning model can be selected from the model square as the target model.
[0054] In one embodiment, the steps of responding to a model deployment request and determining a target model from a model plaza include:
[0055] Respond to the model deployment request and obtain the classification label in the model deployment request;
[0056] Based on the classification label, determine the target model from the model plaza.
[0057] Among them, when registering a model to be trained in the model plaza, a classification label can be set for the model to be trained in advance. The classification label can include a manufacturer label (such as Qianwen), a workload label, a model type label (such as a large language model, an LM model, a vector model, a deep learning model, etc.), etc. Thus, when receiving a model deployment request, the classification label in the model deployment request can be matched with the preset classification label in the model plaza, and the model to be trained corresponding to the successfully matched classification label is determined as the target model.
[0058] Exemplarily, assume that the model deployment request is to deploy a deep learning-based image classification model for identifying animal categories (such as cats, dogs, birds, etc.) in images. By parsing this model deployment request, the model classification label of the model deployment request can be obtained as a deep learning model. Then, the corresponding deep learning model can be selected from the model plaza as the target model according to this model classification label.
[0059] In this embodiment, the user can quickly locate the target model through classification label retrieval.
[0060] In one embodiment, the steps of responding to a model deployment request and determining a target model from a model plaza include:
[0061] Respond to the model deployment request and obtain the user preference information and / or historical record information corresponding to the model deployment request;
[0062] Based on the user preference information and / or the historical record information, determine the target model from the model plaza.
[0063] Among them, the user preference information refers to the preference information preset by the initiating account of the model deployment request. Exemplarily, the preference information may include model type preference (such as deep learning model), task type preference (such as image classification), model vendor preference (such as Qianwen), performance metric preference (such as high accuracy of the model), etc. The historical record information refers to the user behavior records and model deployment records corresponding to the initiating account of the model deployment request. Among them, the user behavior records include model usage records, that is, the records of the user's past deployment and use of the model, such as the name, version, usage time, usage frequency, etc. of the model; task type records, that is, the task types completed by the user in the past, such as image classification, natural language processing, speech recognition, etc.; performance feedback records, that is, the user's feedback on the model performance, such as evaluations in terms of accuracy, inference speed, resource consumption, etc. The model deployment records may include deployment environment records, that is, the hardware environment (such as CPU, GPU, memory) and software environment (such as operating system, Python version, dependency library version) used by the user when deploying the model in the past; deployment method records, that is, the deployment methods selected by the user in the past, such as Web API, containerization (Docker), local deployment, etc.; deployment result records, that is, whether the deployment is successful, and the problems encountered and solutions during the deployment process, etc. Based on this, the target model can be determined from the model square based on the user preference information and / or historical record information.
[0064] S102: Obtain the model support information of the target model based on the dynamic mapping table.
[0065] Among them, the dynamic mapping table records the compatibility relationships (i.e., mapping relationships) between each model to be trained in the model square and various computing power brands (such as computing power cards, graphics cards) and training and inference frameworks. The training and inference framework is a software tool or platform used to build, train, and deploy models. In this application, it is mainly used to train the target model determined from the model square and deploy the model after training is completed. The dynamic mapping table can be maintained and provided by the XFrameworkMatching service of the training and inference framework, that is, synchronize the computing power resource index data of each computing relay collected to the dynamic mapping table to update the resource usage of each computing node in the dynamic mapping table.
[0066] The model support information can be information that supports the training and deployment of the target model, and this information may include computing power brand, training and inference framework, application scenario (such as animal image classification), etc.
[0067] Please refer to Figure 3 , after the user determines the target model, the model support information such as the computing power brand, resource specification, and inference framework supported by the target model can be matched.
[0068] Exemplarily, assume that the model name of the model to be trained registered in the model square is "Qwen2.5-32B-Instruct". The dynamic mapping table records the model name "Qwen2.5-32B-Instruct" of this model, the model parameter "32B", the computing power brands supported, such as the computing power brand named "nvidia", the device model "A100-SXM4-80GB" of this brand, the resource specification "16vcpu-64G-1*80G" of this brand, the computing power brand named "ascend", the device model "Ascend-910B" of this brand, and the resource specification "16vcpu-64G-2*64G" of this brand; the list of hardware brands compatible with the inference frameworks supported by the model, such as, for example, the framework name "vllm", the framework brand "nvidia", the framework description "Requires CUDA11.7 or higher", and the framework name "vllm"; the framework name "mindie", the framework brand "ascend", the framework description "Only supports Ascend 910B devices", etc. Based on this, the support information of the target model that can be obtained includes the model name, model parameters, the computing power name, device model, resource specification of the supported computing power brands, the list of hardware brands compatible with the inference frameworks supported by the model, etc. That is, according to the target model selected from the model square, the computing power brands, specifications, and corresponding inference frameworks supported by this target model can be matched. There is no need for users to be familiar with the deployment specifications of multiple frameworks, avoiding increasing the learning cost and technical operation and maintenance burden and reducing the deployment complexity.
[0069] S103: Obtain the computing power resource index data of each computing node corresponding to the model support information.
[0070] Among them, a computing node refers to a physical or virtual device used to execute the model deployment task. A computing node usually refers to a server or a virtual machine, which has independent computing resources (such as CPU, GPU, memory, storage, etc.) and can execute the model deployment task corresponding to the model deployment request.
[0071] The computing power resource index data refers to the resource index of the corresponding computing power brand on the computing node. This resource index includes computing power brand manufacturers, device models, the number of allocable devices, the allocated video memory, the allocable video memory, device status and other computing power information, as well as the health status.
[0072] Optionally, a corresponding Deviceplugin service plugin is set for each computing power brand. The Deviceplugin service plugin is used to manage non-standard hardware resources (such as GPUs, FPGAs, NICs, etc.) in the cluster. Through DevicePlugin, hardware resources on nodes can be dynamically discovered. Specifically, please refer to Figure 3 , in this application, the Deviceplugin service plugin is used to obtain in real time the computing power resource metric data of the corresponding computing power brand on each computing node, so as to provide real-time data support for the scheduler scheduler.
[0073] S104: Deploy the target model based on the computing power resource metric data of each computing node and the dynamic mapping table.
[0074] The scheduler can be based on the computing power resource metric data and the dynamic mapping table provided in real time by the DevicePlugin service plugin. Please refer to Figure 3 , the scheduler evaluates and pre-selects at least one computing node that meets the conditions (that is, meets the model support information of the target model and the computing power resources required for deploying the target model), and specifies the inference framework in the node information of the at least one computing node, and deploys the target model based on the selected computing node and the inference framework. Exemplarily, if the target model deployment task requires 3 / 4 of the remaining computing power resources (such as the number of CPU cores, the number of GPUs, the memory size, or other dedicated hardware resources) on a computing node, then the remaining 3 / 4 of the computing power resources of the computing node will be scheduled to process the target model deployment task; if the target model deployment task requires the GPU computing power of a computing power card, then an idle computing power card on the computing node will be scheduled; if the target model deployment task requires 5 loads, and only a part of the loads is satisfied on the computing power card of one computing node, while the computing power card of another computing node just meets the other part, then the two computing power cards on the two computing nodes will be called to execute the target model deployment task; if the target model deployment task requires 500 loads, if the computing power resources of a computing power card on a computing node meet this requirement, then the computing power card of the computing node can be called to execute the model deployment task.
[0075] In one embodiment, the deploying the target model based on the computing power resource metric data of each computing node and the dynamic mapping table includes:
[0076] Determine the target computing node based on the computing power resource metric data of each computing node and the dynamic mapping table, and determine the target inference framework in the node information of the target computing node;
[0077] Deploy the target model based on the node information of the target computing node and the target inference framework.
[0078] Among them, the inference framework has the functions of training and inference, and can support the entire process of the model to be trained in the model square from training to inference. In this application, the target inference framework supports the entire process of the target model from training to inference.
[0079] Exemplarily, the scheduler is enabled to determine that the node name of the target computing node of the target model based on the computing power resource index data of each computing node and the dynamic mapping table is "node-c01", the computing power brand is "nvidia", the device model is "A100-SXM4-80GB", the resource specification is "16vcpu-64G-1*80G", the target inference framework is "vllm", and the node information of the target computing node includes target inference framework information, and the target inference framework information includes the framework version "0.7.2" and the framework description "Requires CUDA 11.7 or higher"; and the node name of another target computing node is "node-c05", the computing power brand is "ascend", the device model is "Ascend-910B", the resource specification is "6vcpu-64G-2*64G", the target inference framework is "mindie", and the node information of the other target computing node includes target inference framework information, and the target inference framework information includes the framework version "1.0" and the framework description "Only supports Ascend 910B devices (i.e., only supports Ascend 910B devices)". It should be noted here that the computing nodes "node-c01" and "node-c05" determined by the scheduler based on the computing power resource index data of each computing node and the dynamic mapping table are compatible in terms of the inference framework and the resources are adapted when deploying the target model. Therefore, the model deployment can be completed through the computing nodes "node-c01" and "node-c05" and the corresponding target inference framework. Optionally, a configuration file can be generated based on the node information of the determined target computing node, the computing power card information, the model details of the target model, etc., and the configuration file can be used to deploy the target model through an automated deployment tool or platform (such as Kubernetes, Docker Swarm, Terraform, etc.).
[0080] Since the computing power resources (such as CPU, GPU, memory) of each computing node are different, in order to reasonably utilize the computing power resources, a suitable scheduling algorithm is selected according to the requirements of the target model deployment (such as model support information and required resource specifications) and the computing power resource index data of the computing node, and the target model deployment task is scheduled to the target computing node. Optionally, in one embodiment, determining the target computing node based on the computing power resource index data of each computing node and the dynamic mapping table includes:
[0081] Determine a target scheduling algorithm based on the computing power resource index data of each computing node and the dynamic mapping table;
[0082] Determine a target computing node according to the target scheduling algorithm.
[0083] Multiple scheduling algorithms are pre-set in the scheduler, such as bin-packing scheduling algorithm, preemption scheduling algorithm, priority scheduling algorithm, load balancing scheduling algorithm, batch processing scheduling algorithm, etc.
[0084] Exemplarily, assume there are the following three target computing nodes: Node A: 32 CPU cores, 4 NVIDIA Tesla V100 GPUs, 128 GB of memory, current load: 50% CPU, 30% GPU, 60% memory; Node B: 16 CPU cores, 2 NVIDIA Tesla V100 GPUs, memory: 64 GB, current load: 80% CPU, 70% GPU, 80% memory; Node C: 64 CPU cores, 8 NVIDIA Tesla V100 GPUs, 256 GB of memory, current load: 30% CPU, 20% GPU, 40% memory. Model support information of the target model: Supported hardware is NVIDIA GPU, supported inference framework is vLLM; Required resources: 16 CPU cores, 4 GPUs, 128 GB of memory. According to the resource requirements of the task (16 CPU cores, 4 GPUs, 128 GB of memory) and the computing power resource index data of the target computing node, select a node that can meet the requirements and has a lower load. For example, since the batch processing scheduling algorithm is applicable to the situation where nodes need to be selected according to the specific resource requirements of the task to ensure the smooth operation of the task, by selecting the batch processing scheduling algorithm, it is evaluated that Node A meets the resource requirements (32 CPU cores, 4 GPUs, 128 GB of memory) and the current load is moderate; while Node C meets the resource requirements (64 CPU cores, 8 GPUs, 256 GB of memory) and the current load is the lowest, thus determining Node C as the target computing node. When multiple computing nodes or computing power brands on multiple computing nodes are required to cooperate to execute a task, the priority scheduling algorithm can be used, such as giving priority to the same computing power brand and the next priority to different but compatible computing power brands.
[0085] After the model is deployed, during operation, once a server or computing power card fails, since the computing power brand and training and inference framework are fixed during the initial deployment and cannot be dynamically switched to other available computing power brands and frameworks, resulting in recovery delay, a fault detection and status saving and recovery mechanism is embedded in the scheduler, which can respond in a timely manner when a node or computing power card has an abnormality, reschedule the affected workload to other healthy nodes, avoid recovery delay, and ensure high availability. That is, in one embodiment, the model deployment method further includes:
[0086] Monitor the target computing node;
[0087] If the target computing node fails, update the dynamic mapping table;
[0088] Update the target computing node of the target model according to the updated dynamic mapping table;
[0089] Migrate the workload of the failed target computing node to the updated target computing node for execution.
[0090] Among them, the target computing node can be monitored through a preset fault detection and status saving and recovery mechanism. Optionally, the preset fault detection and status saving and recovery mechanism can monitor the target computing node through monitoring tools (such as Prometheus, Nagios, Zabbix, Shell scripts, etc.).
[0091] As Figure 3 shown, after the target model is deployed, if it is monitored that the deployed target model fails during operation, the fault detection and status saving and recovery mechanism in the scheduler is used to detect the fault (computing node fault (such as sudden power failure) or computing power card, graphics card fault (computing power card, graphics card is broken)), and trigger the recovery mechanism, that is, request the compatibility relationship between the target model in the dynamic mapping table and various computing power brands (such as computing power cards, graphics cards) and training and inference frameworks, as well as the mapping relationship of the real-time computing power indicators of each computing node, so that the scheduler can re-plan the computing power brand and inference framework of the target model, pre-select new healthy nodes (that is, re-determine the target computing node of the target model), migrate the workload of the failed target computing node to the new healthy node for execution, and continuously monitor the node status.
[0092] In this embodiment, the node status can be monitored in real time during the model operation. When the server or computing power card fails, the scheduler can automatically re-plan the mapping relationship between the computing power brand and the inference framework, and reschedule the affected workload to other healthy nodes, so as to achieve an efficient fault self-healing function and ensure business continuity.
[0093] In one embodiment, the model deployment method further includes:
[0094] In response to a model viewing request, obtain the model details corresponding to the model viewing request;
[0095] Display the model details.
[0096] A model viewing request refers to a request for viewing the details of a target model. The details of the model can include the model manufacturer, the number of parameters, the supported computing power brands, the training and inference frameworks, the application scenarios, etc.
[0097] When a user initiates a model viewing request, display the details of the target model, such as the model manufacturer, the number of parameters, the supported computing power brands, the training and inference frameworks, the application scenarios, etc. to the user.
[0098] As Figure 4 shown, a deployment system is provided. The model system includes a model square, an intelligent scheduling system, and a device layer (i.e., device). Among them, when responding to a model deployment request, determine a target model from the model square. The intelligent scheduling system obtains the model support information of the target model based on the dynamic mapping table; obtain the computing power resource index data of each computing node corresponding to the model support information on the device layer; deploy the target model based on the computing power resource index data of each computing node and the dynamic mapping table, innovatively integrating the DevicePlugin heterogeneous computing power resource perception, Scheduler intelligent scheduling, and XFrameworkMatching training and inference framework matching technologies. Through this deployment method, users only need to select a model to achieve efficient deployment, and at the same time have a strong self-healing ability for faults, greatly improving the flexibility and diversity of the deployment method during the model training and inference processes. This method not only simplifies the operation process, but also significantly enhances the robustness and resource utilization rate of the system, providing strong technical support for the rapid development of the AI field.
[0099] In this application, when the user selects a target model, there is no need to care about the specific computing power brand or inference framework. It can intelligently match the most suitable combination of computing power brand, resource specifications, and inference framework according to the requirements of the target model and the currently available resources, realizing intelligent scheduling. This method not only significantly reduces the operation complexity of the user, but also greatly improves the flexibility of the system and the overall resource utilization rate. In addition, the scheduler has a built-in fault detection and status preservation and recovery mechanism, which can monitor the node status in real time during the model operation. When a server or computing power card fails, the scheduler can automatically re-plan the mapping relationship between the target model, computing power brand, and inference framework, and re-schedule the affected workloads to other healthy nodes, so as to achieve an efficient fault self-healing function, ensure business continuity, and realize fault self-healing. In summary, the model deployment method of this application simplifies the model deployment process, realizes flexible allocation of computing power resources, improves the utilization rate of computing power resources, and enhances the fault tolerance ability.
[0100] The above is the model deployment process of this application.
[0101] As described above, this application provides a model deployment method, device, computer device, and storage medium. By responding to a model deployment request, a target model is determined from a model square; model support information of the target model is obtained based on a dynamic mapping table; computing power resource index data of each computing node corresponding to the model support information is obtained; and the target model is deployed based on the computing power resource index data of each computing node and the dynamic mapping table. In the model deployment solution provided by this application, when a model deployment request is received, first, the target model can be automatically determined from the model square in response to the model deployment request. Then, the model support information of the target model can be automatically determined through the dynamic mapping table, without the user manually selecting the computing power brand and inference framework, avoiding insufficient or excessive computing power resources caused by the user's estimation deviation of computing power resources. Next, the computing power resource index data of each computing node corresponding to the model support information can be automatically obtained, and the computing power resources can be accurately and flexibly allocated to the target model and deployed in combination with the dynamic mapping table, which can avoid the occurrence of insufficient computing power resources. And when a computing node fails, the computing power resource allocation of the target model can be dynamically adjusted according to the computing power resource index data of each computing node, thereby improving the overall resource utilization rate.
[0102] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0103] In one embodiment, a model deployment device is provided, and this model deployment device corresponds one-to-one to the model deployment method in the above embodiment. Please refer to Figure 5 As shown, this model deployment device includes:
[0104] A determination module 201, configured to determine a target model from a model plaza in response to a model deployment request;
[0105] A first acquisition module 202, configured to acquire model support information of the target model based on a dynamic mapping table;
[0106] A second acquisition module 203, configured to acquire computing power resource index data of each computing node corresponding to the model support information;
[0107] A deployment module 204, configured to deploy the target model based on the computing power resource index data of each computing node and the dynamic mapping table.
[0108] In the model deployment solution provided in this application, when receiving a model deployment request, first, the target model can be automatically determined from the model plaza in response to the model deployment request. Then, the model support information of the target model can be automatically determined through the dynamic mapping table, without the need for the user to manually select a computing power brand and an inference framework, avoiding insufficient or excessive computing power resources caused by the user's estimation deviation of the computing power resources. Next, the computing power resource index data of each computing node corresponding to the model support information can be automatically acquired, and the computing power resources can be accurately and flexibly allocated to the target model and deployed in combination with the dynamic mapping table, which can avoid the occurrence of insufficient computing power resources. Moreover, when a computing node fails, the computing power resource allocation of the target model can be dynamically adjusted according to the computing power resource index data of each computing node, thereby improving the overall resource utilization rate.
[0109] Optionally, the determination module includes a first determination sub-module, and the first determination sub-module is specifically configured to:
[0110] Respond to the model deployment request and acquire a classification label in the model deployment request;
[0111] Determine the target model from the model plaza based on the classification label.
[0112] Optionally, the determination module includes a second determination sub-module, and the second determination sub-module is specifically configured to:
[0113] Respond to the model deployment request and acquire user preference information and / or historical record information corresponding to the model deployment request;
[0114] Determine the target model from the model plaza based on the user preference information and / or the historical record information.
[0115] Optionally, the deployment module includes:
[0116] A third determination sub-module, configured to determine a target computing node based on the computing power resource index data of each computing node and the dynamic mapping table, and determine a target inference framework in the node information of the target computing node;
[0117] A deployment sub-module, configured to deploy the target model based on the node information of the target computing node and the target inference framework.
[0118] Optionally, the third determination sub-module includes:
[0119] A first determination unit, configured to determine a target scheduling algorithm based on the computing power resource index data of each computing node and the dynamic mapping table;
[0120] A second determination unit, configured to determine a target computing node according to the target scheduling algorithm.
[0121] Optionally, the model deployment device further includes a monitoring module, and the monitoring module is specifically configured to:
[0122] Monitor the target computing node;
[0123] If the target computing node fails, update the dynamic mapping table;
[0124] Update the target computing node of the target model according to the updated dynamic mapping table;
[0125] Migrate the workload of the failed target computing node to the updated target computing node for execution.
[0126] Optionally, the model deployment device further includes a display module, and the display module is specifically configured to:
[0127] In response to a model viewing request, obtain the detailed model information corresponding to the model viewing request;
[0128] Display the detailed model information.
[0129] In one embodiment, a computer device is provided. The internal structure diagram of the computer device can be as Figure 6 shown. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server through a network connection. When the computer program is executed by the processor, it realizes the functions or steps of a model deployment method.
[0130] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented:
[0131] In response to a model deployment request, determine a target model from a model square; obtain model support information of the target model based on a dynamic mapping table; obtain computing power resource index data of each computing node corresponding to the model support information; and deploy the target model based on the computing power resource index data of each computing node and the dynamic mapping table.
[0132] In this embodiment, when a model deployment request is received, first, the target model can be automatically determined from the model square in response to the model deployment request. Then, the model support information of the target model can be automatically determined through the dynamic mapping table, without the user manually selecting a computing power brand and an inference framework, avoiding insufficient or excessive computing power resources caused by the user's estimation deviation of computing power resources. Next, the computing power resource index data of each computing node corresponding to the model support information can be automatically obtained, and the computing power resources can be accurately and flexibly allocated to the target model and deployed in combination with the dynamic mapping table, which can avoid the occurrence of insufficient computing power resources. Moreover, when a computing node fails, the computing power resource allocation of the target model can be dynamically adjusted according to the computing power source index data of each computing node, thereby improving the overall resource utilization rate.
[0133] In one embodiment, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the following steps are implemented:
[0134] In response to a model deployment request, determine a target model from a model square; obtain model support information of the target model based on a dynamic mapping table; obtain computing power resource index data of each computing node corresponding to the model support information; and deploy the target model based on the computing power resource index data of each computing node and the dynamic mapping table.
[0135] In this embodiment, when a model deployment request is received, first, the target model can be automatically determined from the model square in response to the model deployment request. Then, the model support information of the target model can be automatically determined through the dynamic mapping table, without the user manually selecting a computing power brand and an inference framework, avoiding insufficient or excessive computing power resources caused by the user's estimation deviation of computing power resources. Next, the computing power resource index data of each computing node corresponding to the model support information can be automatically obtained, and the computing power resources can be accurately and flexibly allocated to the target model and deployed in combination with the dynamic mapping table, which can avoid the occurrence of insufficient computing power resources. Moreover, when a computing node fails, the computing power resource allocation of the target model can be dynamically adjusted according to the computing power source index data of each computing node, thereby improving the overall resource utilization rate.
[0136] It should be noted that for the functions or steps that can be achieved by the above computer-readable storage medium or computer device, reference can be made to the relevant descriptions on the server side and the client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0137] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above various methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0138] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.
[0139] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention and should all be included in the protection scope of the present invention.
Claims
1. A model deployment method, characterized in that, Including: In response to a model deployment request, determine a target model from a model plaza; Obtain model support information of the target model based on a dynamic mapping table; Obtain computing power resource metric data of each computing node corresponding to the model support information; Deploy the target model based on the computing power resource metric data of each computing node and the dynamic mapping table.
2. The model deployment method according to claim 1, wherein The step of determining a target model from a model plaza in response to a model deployment request includes: In response to a model deployment request, obtain a classification label in the model deployment request; Determine the target model from the model plaza based on the classification label.
3. The model deployment method according to claim 1, wherein The step of determining a target model from a model plaza in response to a model deployment request includes: In response to a model deployment request, obtain user preference information and / or historical record information corresponding to the model deployment request; Determine the target model from the model plaza based on the user preference information and / or the historical record information.
4. The model deployment method according to claim 1, wherein The step of deploying the target model based on the computing power resource metric data of each computing node and the dynamic mapping table includes: Determine a target computing node based on the computing power resource metric data of each computing node and the dynamic mapping table, and determine a target inference framework in the node information of the target computing node; Deploy the target model based on the node information of the target computing node and the target inference framework.
5. The model deployment method according to claim 4, wherein The step of determining a target computing node based on the computing power resource metric data of each computing node and the dynamic mapping table includes: Determine a target scheduling algorithm based on the computing power resource metric data of each computing node and the dynamic mapping table; Determine the target computing node according to the target scheduling algorithm.
6. The model deployment method according to claim 4, wherein The model deployment method further includes: Monitor the target computing node; If the target computing node fails, update the dynamic mapping table; Update the target computing node of the target model according to the updated dynamic mapping table; Migrate the workload of the failed target computing node to the updated target computing node for execution.
7. The model deployment method according to any one of claims 1 to 6, characterized in that, The model deployment method further includes: In response to a model viewing request, obtain model detailed information corresponding to the model viewing request; Display the model detailed information.
8. A model deployment device, characterized in that, Including: A determination module, configured to determine a target model from a model plaza in response to a model deployment request; A first acquisition module, configured to obtain model support information of the target model based on a dynamic mapping table; A second acquisition module, configured to obtain computing power resource metric data of each computing node corresponding to the model support information; A deployment module, configured to deploy the target model based on the computing power resource metric data of each computing node and the dynamic mapping table.
9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the model deployment method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, the steps of the model deployment method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Learned model provision method, and learned model provision device
CN109074607A
Customized service deployment method and device, equipment and storage medium
CN118803045A
Large model scheduling system and method, server, medium and product
CN119473534A
Calculation power scheduling method and device for model training, storage medium and electronic equipment
CN119537031A
Computing power resource scheduling method and related device
CN119576503A
Cited By
Model deployment method
CN121300870A
Model deployment method, electronic equipment, storage medium and program product
CN122240133A
Model deployment methods, electronic devices, storage media, and application products
CN122240133B