Model inference service deployment method and apparatus, device, and storage medium

WO2026174669A1PCT designated stage Publication Date: 2026-08-27X STAR TECHNOLOGY PTE LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/094486
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-20
Filing Date
2025-05-13
Publication Date
2026-08-27

Smart Images

  • Figure CN2025094486_27082026_PF_FP_ABST
    Figure CN2025094486_27082026_PF_FP_ABST
Patent Text Reader

Abstract

A model inference service deployment method and apparatus, a device, and a storage medium. The method comprises: in a model inference process, acquiring inference metric parameters of all model inference services in a computing resource cluster (S101); for each model inference service, on the basis of the inference metric parameters corresponding to the model inference service and pre-constructed mode determination thresholds, determining a target scheduling mode corresponding to the model inference service, wherein the target scheduling mode comprises a quality of service priority mode and a resource utilization priority mode (S102); and on the basis of a service deployment policy in the target scheduling mode, determining a service deployment mode corresponding to the model inference service, and on the basis of the service deployment mode, performing service deployment on the model inference service (S103).
Need to check novelty before this filing date? Find Prior Art

Description

Model inference service deployment methods, devices, equipment and storage media

[0001] This application claims priority to Chinese Patent Application No. 202510188154.6, filed on February 20, 2025, the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of computing resource allocation technology, such as a method, apparatus, device, and storage medium for deploying model inference services. Background Technology

[0003] With the development of machine learning and artificial intelligence technologies, various businesses in the automotive finance sector are increasingly leveraging these technologies to develop and train AI models for their specific business verticals, deploying these models as inference services for their regular operations. However, due to the strong isolation between these vertical businesses and the significant differences in models used across them, there is virtually no model reuse between businesses. This has led to a surge in demand for Graphics Processing Units (GPUs) for model inference as more AI models are integrated into these verticals, putting increasing pressure on GPU resources. Given the scarcity of GPU resources and the high cost of expanding them, there is an urgent need to find a way to improve the utilization rate of existing GPU resources. Summary of the Invention

[0004] This application provides a method, apparatus, device, and storage medium for deploying a model inference service, which dynamically adjusts the scheduling strategy as needed to ensure that the model inference service can meet the predetermined performance requirements, while improving the utilization rate of existing GPU resources.

[0005] This application provides a method for deploying a model inference service. The method includes:

[0006] During the model inference process, obtain the inference metric parameters of all model inference services in the computing resource cluster;

[0007] For each of the model inference services, a threshold is determined based on the inference index parameters corresponding to the model inference service and multiple pre-built patterns to determine the target scheduling mode corresponding to the model inference service. The target scheduling mode includes a service quality priority mode and a resource utilization priority mode.

[0008] Based on the service deployment strategy in the target scheduling mode, the service deployment method corresponding to the model inference service is determined, and the model inference service is deployed based on the service deployment method.

[0009] This application provides a model inference service deployment apparatus. The apparatus includes:

[0010] The inference metric parameter acquisition module is configured to acquire inference metric parameters of all model inference services in the computing resource cluster during the model inference process.

[0011] The target scheduling mode determination module is configured to determine the target scheduling mode corresponding to each model inference service based on the inference index parameters corresponding to the model inference service and multiple pre-built modes. The target scheduling mode includes a service quality priority mode and a resource utilization priority mode.

[0012] The service deployment method determination module is configured to determine the service deployment method corresponding to the model inference service based on the service deployment strategy in the target scheduling mode, and deploy the model inference service based on the service deployment method.

[0013] This application provides an electronic device, the electronic device comprising:

[0014] At least one processor; and

[0015] A memory communicatively connected to the at least one processor; wherein,

[0016] The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the model inference service deployment method described in any embodiment of this application.

[0017] This application provides a computer-readable storage medium storing computer instructions that, when executed by a processor, implement the model inference service deployment method described in any embodiment of this application. Attached Figure Description

[0018] Figure 1 is a flowchart of the model inference service deployment method provided according to Embodiment 1 of this application;

[0019] Figure 2 is a flowchart of the model inference service deployment method according to Embodiment 2 of this application;

[0020] Figure 3 is a structural diagram of a model inference service deployment device provided according to Embodiment 3 of this application;

[0021] Figure 4 is a schematic diagram of the structure of an electronic device that implements the model inference service deployment method of the present application. Detailed Implementation

[0022] The technical solutions of the embodiments of this application will now be described with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort should fall within the scope of protection of this application.

[0023] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0024] Example 1

[0025] Figure 1 is a flowchart of a model inference service deployment method provided in Embodiment 1 of this application. This embodiment is applicable to situations where the computing resources required for a model inference service are optimally allocated. This method can be executed by a model inference service deployment device, which can be implemented in hardware and / or software and can be configured in an electronic device. As shown in Figure 1, the method includes:

[0026] S101. During the model inference process, obtain the inference metric parameters of all model inference services in the computing resource cluster.

[0027] In the fields of machine learning and artificial intelligence, model inference services refer to the process of using pre-trained AI models to predict or analyze new input data. These services are typically deployed on computing resources, such as GPUs, to provide efficient inference capabilities. In sectors like automotive finance, model inference services are widely used across various vertical business areas to support intelligent decision-making and operations.

[0028] A computing resource cluster can refer to a group of interconnected computing resources (such as servers, GPUs, etc.) that work together to provide model inference services. These computing resources are organized together to efficiently handle a large number of model inference requests.

[0029] Inference metrics can refer to the inference parameters of the model inference service during the model inference process, such as response time, throughput, resource utilization (e.g., GPU memory usage) and error rate.

[0030] By selecting the corresponding monitoring metrics and time periods through the inference service scheduling policy controller, you can filter out the monitoring items related to the model inference service and export or view the required inference metric parameters.

[0031] S102. For each of the model inference services, a threshold is determined based on the inference index parameters corresponding to the model inference service and multiple pre-built patterns to determine the target scheduling mode corresponding to the model inference service.

[0032] The target scheduling modes include a quality-of-service (QoS) priority mode and a resource utilization priority mode. The mode determination threshold can be set according to actual conditions; this application does not limit the value of the mode determination threshold.

[0033] The attributes used to determine the threshold in the mode include at least: the average number of requests per time period, the average response time per percentile, the longest response time per time period, whether the default scheduling method is always used within a specified time period, and whether the default scheduling method is never used within a specified time period. For example, the average number of requests per time period may include requests per second, requests per minute, and requests per hour. The average response time per percentile may include P99 response time, P95 response time, and P90 response time. The longest response time per time period may include the longest response time per minute and the longest response time per hour.

[0034] For each model inference service, the relationship between thresholds can be determined based on inference metric parameters and multiple pre-built patterns, thereby determining whether the target scheduling mode for the model inference service is a service quality priority mode or a resource utilization priority mode.

[0035] For example, determining the target scheduling mode corresponding to the model inference service based on the inference metric parameters corresponding to the model inference service and multiple pre-built patterns includes:

[0036] The inference metric parameters and the mode determination threshold are compared and processed. Based on the matching result of the inference metric parameters and the mode determination threshold, the target scheduling mode corresponding to the model inference service is determined.

[0037] In other words, the inference metric parameters are compared and matched with the threshold determined by the mode. Specifically, the inference metric parameters are determined from the dimensions of average number of requests with a time period as the window, average response time with a percentile as the window, longest response time with a time period as the window, always using the default scheduling method within a specified time, and whether the default scheduling method is never used within a specified time. This determines whether the target scheduling mode for the model inference service is a service quality priority mode or a resource utilization priority mode.

[0038] This application does not specify the value of the threshold for multiple modes; it can be set according to actual usage or historical experience.

[0039] S103. Based on the service deployment strategy in the target scheduling mode, determine the service deployment method corresponding to the model inference service, and deploy the model inference service based on the service deployment method.

[0040] In the technical solution of this application, when the target scheduling mode is a quality of service priority mode, the service deployment method includes an independent deployment mode and a joint deployment mode; when the target scheduling mode is a resource utilization priority mode, the service deployment method includes a joint deployment mode and a card sharing deployment mode.

[0041] The independent deployment method involves a single model inference service occupying a single computing resource in the computing resource cluster. The joint deployment method involves multiple model inference services of the same inference type using the same inference service framework, sharing a context to occupy the same computing resource in the computing resource cluster. The card-sharing deployment method involves multiple model inference services taking turns occupying the same computing resource in the computing resource cluster using a time-slice rotation approach.

[0042] Independent deployment means the model inference service exclusively uses the entire GPU card resources. The model inference service's use of the GPU is exclusive, and its GPU usage continues until the model inference service's lifecycle ends (when the model inference service is officially taken offline). Joint deployment involves multiple model inference services using the same inference service framework (e.g., Nvidia Triton framework) in a shared context (Nvidia MPS) manner, jointly using the same GPU on the same compute node. This deployment method avoids the latency caused by context switching when multiple worker processes use the GPU. Shared GPU deployment involves multiple model inference services directly configured to use the same GPU on the same compute node. This deployment method involves explicit GPU context switching when different processes access the GPU; multiple processes acquire GPU usage rights through a time-slice round-robin approach.

[0043] In this application, different target scheduling modes employ different service deployment strategies. For each target scheduling mode, a corresponding service deployment strategy is adopted to determine the service deployment method corresponding to the model inference service. After determining the service deployment method, the model deployment service is notified to redeploy the model inference service according to the new service deployment method.

[0044] The technical solution of this application embodiment obtains inference metric parameters of all model inference services in the computing resource cluster during the model inference process. For each model inference service, a threshold is determined based on the corresponding inference metric parameters and multiple pre-built modes to determine the target scheduling mode corresponding to the model inference service. The target scheduling mode includes a service quality priority mode and a resource utilization priority mode. Based on the service deployment strategy in the target scheduling mode, the service deployment method corresponding to the model inference service is determined, and the model inference service is deployed based on the service deployment method. This solves the problem of insufficient or unbalanced service resource utilization, enables dynamic adjustment of the scheduling strategy as needed, ensures that the model inference service can meet predetermined performance requirements, optimizes resource allocation, improves overall resource utilization, and reduces operating costs.

[0045] Example 2

[0046] Figure 2 is a flowchart of a model inference service deployment method provided in Embodiment 2 of this application. Based on the above embodiments, this embodiment describes how to determine the service deployment method corresponding to the model inference service according to the service deployment strategy in the target scheduling mode. As shown in Figure 2, the method includes:

[0047] S201. During the model inference process, obtain the inference metric parameters of all model inference services in the computing resource cluster.

[0048] S202. For each of the model inference services, a threshold is determined based on the inference index parameters corresponding to the model inference service and multiple pre-built modes to determine the target scheduling mode corresponding to the model inference service, wherein the target scheduling mode includes a service quality priority mode and a resource utilization priority mode.

[0049] S203. When the target scheduling mode is the quality of service priority mode, the service deployment mode corresponding to the model inference service is determined based on the deployment mode threshold of the quality of service priority mode and the inference index parameters.

[0050] The deployment method threshold can refer to a judgment threshold for determining the deployment method. It is used to determine whether the service deployment method is an independent deployment or a collaborative deployment when the target scheduling mode is a quality-of-service (QoS) priority mode. For example, the deployment method threshold can be set to a single model inference using more than 75% of the GPU memory of a single card. It is worth noting that in QoS priority mode, card-sharing deployment is not used.

[0051] For example, determining the service deployment method corresponding to the model inference service based on the deployment method threshold of the service quality priority mode and the inference indicator parameter includes: if the inference indicator parameter is greater than the deployment method threshold, determining the service deployment method corresponding to the model inference service as an independent deployment method; if the inference indicator parameter is less than or equal to the deployment method threshold, determining the service deployment method corresponding to the model inference service as a joint deployment method.

[0052] In other words, under the quality-of-service (QoS) priority mode, the QoS scheduling strategy for model inference services should be satisfied as much as possible. This strategy will schedule model inference services with high resource consumption (i.e., inference metric parameters greater than the deployment method threshold) as standalone deployments, and schedule model inference services with medium and low resource consumption (i.e., inference metric parameters less than or equal to the deployment method threshold) as joint deployments.

[0053] S204. When the target scheduling mode is the resource utilization priority mode, the service deployment method corresponding to the model inference service is determined based on the scheduling strategy that maximizes resource utilization.

[0054] The resource utilization priority mode is a mode that maximizes resource utilization. It requires a scheduling strategy that maximizes resource utilization to determine the service deployment method corresponding to the model inference service.

[0055] For example, based on a greedy algorithm and referring to a scheduling strategy that maximizes resource utilization, the deployment ratios of the joint deployment method and the card-sharing deployment method in the service deployment methods are determined respectively. Alternatively, the first resource utilization rate corresponding to the joint deployment method and the second resource utilization rate corresponding to the card-sharing deployment method are calculated respectively. The joint deployment method is adopted when the first resource utilization rate is greater than the second resource utilization rate, and the card-sharing deployment method is adopted when the first resource utilization rate is less than or equal to the second resource utilization rate.

[0056] In other words, when the target scheduling mode is the resource utilization priority mode, the service deployment mode can be determined in multiple ways, including at least determining the deployment ratio of the joint deployment mode and the card sharing deployment mode in the service deployment mode based on a greedy algorithm; or determining the service deployment mode corresponding to the model inference service based on the resource utilization of the joint deployment mode and the card sharing deployment mode.

[0057] When the target scheduling mode is resource utilization priority mode, a scheduling strategy that maximizes resource utilization should be implemented. This strategy, by default, uses a greedy algorithm to deploy model inference services with different resource requirements onto the same computing resource using both joint deployment and card-sharing deployment methods. On the other hand, the joint deployment method is used when the first resource utilization is greater than the second resource utilization, and the card-sharing deployment method is used when the first resource utilization is less than or equal to the second resource utilization.

[0058] For example, when the target scheduling mode is a resource utilization priority mode, the detailed scheduling process based on a greedy algorithm and referring to a scheduling strategy that maximizes resource utilization includes:

[0059] Submit a query request to the resource management service to obtain a list of all resource nodes in the current cluster that are marked as being used for resource utilization priority strategies;

[0060] Iterate through the resource node list and obtain the resource node with the largest remaining available video memory and its corresponding GPU card number.

[0061] Query the deployment affinity of the inference service being deployed to obtain whether the inference service is deployed using federated deployment or shared card deployment. If the inference service does not specify a deployment method, the deployment method with the most remaining resources will be used by default.

[0062] Bind the inference service with the corresponding resource node and GPU card number, generate a scheduling plan, and notify the model deployment service to deploy the model inference service.

[0063] If the resource management service does not return a list of nodes being used for the resource utilization priority strategy, or if none of the nodes in the returned resource node list meet the GPU resource requirements of the inference service:

[0064] Submit a query request to the resource management service to obtain a list of idle resource nodes in the current cluster;

[0065] Iterate through the resource node list and obtain the resource node with the largest total video memory and its corresponding GPU card number.

[0066] Submit a resource allocation request to the resource management service and mark the specified node as a node that is being used for the resource utilization priority scheduling strategy.

[0067] Query the deployment affinity of the inference service being deployed to obtain whether the inference service is deployed using federated deployment or shared card deployment. If the inference service does not specify a deployment method, federated deployment will be used by default.

[0068] Bind the inference service with the corresponding resource node and GPU card number, generate a scheduling plan, and notify the model deployment service to deploy the model inference service.

[0069] S205. Based on the service deployment method, the model inference service is deployed.

[0070] The technical solution of this application, when the target scheduling mode is a quality-of-service (QoS) priority mode, determines the service deployment method corresponding to the model inference service based on the QoS priority mode deployment method threshold and the inference index parameters. When the target scheduling mode is a resource utilization priority mode, the service deployment method corresponding to the model inference service is determined based on a scheduling strategy that maximizes resource utilization. By adopting different service deployment methods according to different target scheduling modes, the allocation of tasks on hardware resources can be rationally planned, avoiding performance degradation or system crashes caused by multiple tasks simultaneously contending for the same resource, thus promoting the efficient, stable, and reliable operation of the model inference service.

[0071] Example 3

[0072] Figure 3 is a schematic diagram of a model inference service deployment device provided in Embodiment 3 of this application. As shown in Figure 3, the device includes:

[0073] The inference metric parameter acquisition module 301 is configured to acquire the inference metric parameters of all model inference services in the computing resource cluster during the model inference process.

[0074] The target scheduling mode determination module 302 is configured to determine the target scheduling mode corresponding to each model inference service based on the inference index parameters corresponding to the model inference service and multiple pre-built modes. The target scheduling mode includes a service quality priority mode and a resource utilization priority mode.

[0075] The service deployment method determination module 303 is configured to determine the service deployment method corresponding to the model inference service based on the service deployment strategy in the target scheduling mode, and deploy the model inference service based on the service deployment method.

[0076] The technical solution of this application embodiment obtains inference metric parameters of all model inference services in the computing resource cluster during the model inference process. For each model inference service, a threshold is determined based on the corresponding inference metric parameters and multiple pre-built modes to determine the target scheduling mode corresponding to the model inference service. The target scheduling mode includes a service quality priority mode and a resource utilization priority mode. Based on the service deployment strategy in the target scheduling mode, the service deployment method corresponding to the model inference service is determined, and the model inference service is deployed based on the service deployment method. This solves the problem of insufficient or unbalanced service resource utilization, enables dynamic adjustment of the scheduling strategy as needed, ensures that the model inference service can meet predetermined performance requirements, optimizes resource allocation, improves overall resource utilization, and reduces operating costs.

[0077] Optionally, the attributes for determining the threshold in the mode include at least: the average number of requests in a time period window, the average response time in a percentile window, the longest response time in a time period window, always using the default scheduling method within a specified time period, and never using the default scheduling method within a specified time period.

[0078] Optionally, the target scheduling mode determination module 302 is configured as follows:

[0079] The inference metric parameters and the mode determination threshold are compared and processed. Based on the matching result of the inference metric parameters and the mode determination threshold, the target scheduling mode corresponding to the model inference service is determined.

[0080] Optionally, when the target scheduling mode is a quality of service priority mode, the service deployment method includes an independent deployment method and a joint deployment method. In the independent deployment method, a single model inference service occupies a computing resource in the computing resource cluster on its own. In the joint deployment method, multiple model inference services of the same inference type using the same inference service framework occupy the same computing resource in the computing resource cluster in a shared context manner.

[0081] When the target scheduling mode is the resource utilization priority mode, the service deployment method includes a joint deployment mode and a card-sharing deployment mode. In the card-sharing deployment mode, multiple model inference services take turns occupying the same computing resource in the computing resource cluster in a time-slice rotation manner.

[0082] Optionally, the service deployment method determination module 303 includes:

[0083] The first deployment method determination unit is configured to determine the service deployment method corresponding to the model inference service based on the deployment method threshold of the service quality priority mode and the inference index parameters when the target scheduling mode is the service quality priority mode.

[0084] The second deployment method determination unit is configured to determine the service deployment method corresponding to the model inference service based on a scheduling strategy that maximizes resource utilization when the target scheduling mode is a resource utilization priority mode.

[0085] Optionally, the first deployment method for determining the unit is set as follows:

[0086] If the inference metric parameter is greater than the deployment method threshold, the service deployment method corresponding to the model inference service is determined to be an independent deployment method;

[0087] If the inference metric parameter is less than or equal to the deployment method threshold, the service deployment method corresponding to the model inference service is determined to be a joint deployment method.

[0088] Optionally, the second deployment method determining unit is configured as follows:

[0089] Based on a greedy algorithm and referring to a scheduling strategy that maximizes resource utilization, the deployment proportions of the joint deployment method and the card-sharing deployment method in the service deployment methods are determined respectively.

[0090] or,

[0091] Calculate the first resource utilization rate when using the joint deployment method and the second resource utilization rate when using the card sharing deployment method. If the first resource utilization rate is greater than the second resource utilization rate, the joint deployment method is used. If the first resource utilization rate is less than or equal to the second resource utilization rate, the card sharing deployment method is used.

[0092] The model inference service deployment apparatus provided in this application embodiment can execute the model inference service deployment method provided in any embodiment of this application, and has the corresponding functional modules for executing the method.

[0093] Example 4

[0094] Figure 4 illustrates a schematic diagram of an electronic device 10 that can be used to implement embodiments of this application. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the application described and / or claimed herein.

[0095] As shown in Figure 4, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a read-only memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0096] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0097] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 11 performs several of the methods and processes described above, such as model inference service deployment methods.

[0098] In some embodiments, the model inference service deployment method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the model inference service deployment method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the model inference service deployment method by any other suitable means (e.g., by means of firmware).

[0099] The various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard parts (ASSPs), systems-on-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0100] Computer programs used to implement the methods of this application may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0101] In the context of this application, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. Examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fiber, compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0102] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a cathode ray tube (CRT) or liquid crystal display (LCD) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0103] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0104] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system. It addresses the shortcomings of traditional physical hosts and Virtual Private Server (VPS) services, such as high management difficulty and weak business scalability.

[0105] It should be understood that the various processes shown above can be used to rearrange, add, or delete steps. For example, the multiple steps described in this application can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this application can be achieved, and this is not limited herein.

Claims

1. A method for deploying a model inference service, comprising: During the model inference process, obtain the inference metric parameters of all model inference services in the computing resource cluster; For each of the model inference services, a threshold is determined based on the inference index parameters corresponding to the model inference service and multiple pre-built patterns to determine the target scheduling mode corresponding to the model inference service. The target scheduling mode includes a service quality priority mode and a resource utilization priority mode. Based on the service deployment strategy in the target scheduling mode, the service deployment method corresponding to the model inference service is determined, and the model inference service is deployed based on the service deployment method.

2. The method according to claim 1, wherein, The attributes used to determine the threshold in the mode include at least: the average number of requests in a time period window, the average response time in a percentile window, the longest response time in a time period window, always using the default scheduling method within a specified time period, and never using the default scheduling method within a specified time period.

3. The method according to claim 1, wherein, The step of determining the target scheduling mode corresponding to the model inference service by determining a threshold based on the inference index parameters corresponding to the model inference service and multiple pre-built modes includes: The inference metric parameters and the mode determination threshold are compared and processed. Based on the matching result of the inference metric parameters and the mode determination threshold, the target scheduling mode corresponding to the model inference service is determined.

4. The method according to claim 1, further comprising: When the target scheduling mode is the quality of service priority mode, the service deployment method includes independent deployment and joint deployment. In the independent deployment mode, a single model inference service occupies a computing resource in the computing resource cluster on its own. In the joint deployment mode, multiple model inference services of the same inference type using the same inference service framework occupy the same computing resource in the computing resource cluster in a shared context manner. When the target scheduling mode is the resource utilization priority mode, the service deployment method includes a joint deployment mode and a card-sharing deployment mode. In the card-sharing deployment mode, multiple model inference services take turns occupying the same computing resource in the computing resource cluster in a time-slice rotation manner.

5. The method according to claim 4, wherein, The step of determining the service deployment method corresponding to the model inference service based on the service deployment strategy in the target scheduling mode includes: When the target scheduling mode is the quality of service priority mode, the service deployment mode corresponding to the model inference service is determined based on the deployment mode threshold of the quality of service priority mode and the inference index parameters. When the target scheduling mode is the resource utilization priority mode, the service deployment method corresponding to the model inference service is determined based on the scheduling strategy that maximizes resource utilization.

6. The method according to claim 5, wherein, The process of determining the service deployment method corresponding to the model inference service based on the deployment method threshold of the service quality priority mode and the inference indicator parameters includes: If the inference metric parameter is greater than the deployment method threshold, the service deployment method corresponding to the model inference service is determined to be an independent deployment method; If the inference metric parameter is less than or equal to the deployment method threshold, the service deployment method corresponding to the model inference service is determined to be a joint deployment method.

7. The method according to claim 5, wherein, The scheduling strategy based on maximizing resource utilization determines the service deployment method corresponding to the model inference service, including: Based on a greedy algorithm and referring to a scheduling strategy that maximizes resource utilization, the deployment proportions of the joint deployment method and the card-sharing deployment method in the service deployment methods are determined respectively. or, Calculate the first resource utilization rate when using the joint deployment method and the second resource utilization rate when using the card sharing deployment method. If the first resource utilization rate is greater than the second resource utilization rate, the joint deployment method is used. If the first resource utilization rate is less than or equal to the second resource utilization rate, the card sharing deployment method is used.

8. A model inference service deployment device, comprising: The inference metric parameter acquisition module is configured to acquire inference metric parameters of all model inference services in the computing resource cluster during the model inference process. The target scheduling mode determination module is configured to determine the target scheduling mode corresponding to each model inference service based on the inference index parameters corresponding to the model inference service and multiple pre-built modes. The target scheduling mode includes a service quality priority mode and a resource utilization priority mode. The service deployment method determination module is configured to determine the service deployment method corresponding to the model inference service based on the service deployment strategy in the target scheduling mode, and deploy the model inference service based on the service deployment method.

9. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the model inference service deployment method according to any one of claims 1-7.

10. A computer-readable storage medium storing computer instructions, said computer instructions being configured to cause a processor to execute the model inference service deployment method of any one of claims 1-7.