Instance scheduling method and related equipment
By introducing the maximum concurrency optimization node selection strategy in the serverless computing platform, the execution delay prediction model is used to estimate the maximum concurrency on the node, and efficient and high-density instance scheduling is achieved, which solves the problem of business performance degradation caused by resource competition and improves resource utilization and user experience.
Patent Information
- Application Number
- CN202410175722.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-22
- Filing Date
- 2024-02-07
- Publication Date
- 2025-05-23
AI Technical Summary
In a serverless computing platform, how to achieve efficient and high-density instance scheduling to ensure resource utilization and user experience, while solving the problem of business performance degradation caused by resource competition.
The maximum concurrency optimization node selection strategy is introduced, and the maximum concurrency on a single node is estimated through the execution delay prediction model of the target service, and the nodes used to schedule instances are directly decided, so that a single inference can be used to perform batch scheduling.
On the premise of ensuring execution delay, efficient and high-density instance scheduling is achieved, the frequency of inference during the scheduling process is reduced, the overall scheduling performance and resource utilization rate is improved, the user experience is optimized, and the business performance decline caused by resource competition is solved.
Smart Images

Figure CN120029726A_ABST
Abstract
Description
[0001] This application claims the priority of the Chinese patent application filed with the State Intellectual Property Office on November 22, 2023, with application number 202311566651.2 and invention name “An instance scheduling method and related equipment”, all contents of which are incorporated by reference in this application. Technical Field
[0002] The present application relates to the field of cloud computing technology, and in particular to an instance scheduling method, a scheduling system, a scheduler, a computing device cluster, a computer-readable storage medium, and a computer program product. Background Art
[0003] With the continuous development of cloud computing, various new computing paradigms are constantly emerging, and serverless computing (abbreviated as serverless) is a relatively popular computing paradigm. Serverless computing is based on Platform-as-a-Service (PaaS) and provides a micro-architecture. End customers do not need to deploy, configure or manage servers. The server services required for code operation, such as elastic scaling of backend resources, can be provided by the serverless platform in the cloud, which frees developers from managing servers.
[0004] Serverless platforms usually rely on dedicated schedulers to create new instances on nodes based on various factors, such as resource and performance requirements. Usually, when a function call request arrives, if there are no idle instances, the scheduler can create a new instance in a virtual machine based on the scheduling policy. After the instance is successfully created, it can be used to execute the function call request.
[0005] The scheduler of the serverless platform plays an important role in maintaining resource efficiency, function density (usually represented by the number of instances running concurrently in a node), and user-oriented experience (usually represented by service quality or execution latency). Each platform needs to adopt a high-density scheduling method to improve the overall resource utilization of the platform. How to achieve efficient and reliable high-density instance scheduling has gradually become a challenge. Summary of the invention
[0006] The present application provides an instance scheduling method, which introduces a maximum concurrency optimization node selection strategy, estimates the maximum concurrency of a target service on a single node based on the execution delay prediction of the target service, and directly decides on the node used to schedule the instance based on the maximum concurrency, so that batch scheduling can be performed for a single reasoning, and can schedule instances efficiently and densely while ensuring the execution delay. The present application also provides a scheduling system, a scheduler, a computing device cluster, a computer-readable storage medium, and a computer program product corresponding to the above method.
[0007] In the first aspect, the present application provides an instance scheduling method. The method can be performed by a scheduling system. The scheduling system can be a software system for implementing instance scheduling, which can be an independent software system or a software system integrated with other software, such as a software system integrated with other software in the form of plug-ins, functional modules, applets, etc. The software system can be provided to customers in the form of a software package, and customers can deploy the software package on a computing device or computing device cluster such as a local data center or a private cloud. Alternatively, the software system can be provided to users in the form of a cloud service. The above software system can be deployed in a computing device cluster, and the computing device cluster executes the program code of the software system, thereby executing the instance scheduling method of the present application. In some examples, the scheduling system can also be a hardware system, such as a computing device cluster with instance scheduling capabilities, and when the computing device cluster is running, the instance scheduling method of the present application is executed.
[0008] Specifically, the scheduling system receives an instance creation request for the target service, and queries the node on which the instance of the target service is deployed from the node cluster. When there is a first node on which the instance of the target service is deployed, the scheduling system obtains the maximum concurrency and current concurrency of the target service in the first node. Among them, the current concurrency of the target service in the first node is characterized by the number of instances of the target service currently deployed on the first node, and the maximum concurrency of the target service in the first node is determined according to the delay prediction model of the first node. The delay prediction model of the first node takes the resource utilization of the instance in the first node and the concurrency of the service in the first node as input, and takes the execution delay of the service in the first node as output. When there is a target node in the first node whose current concurrency of the target service is less than the maximum concurrency, the scheduling system deploys the instance of the target service at the target node.
[0009] This method introduces a maximum concurrency optimization node selection strategy, estimates the maximum concurrency of the target service (such as a function) on a single node (such as a virtual machine) based on the execution delay prediction of the target service, and directly decides on the node used to schedule the instance based on the maximum concurrency, so that batch scheduling can be performed for a single inference, and instances can be scheduled efficiently and densely while ensuring the execution delay. This method significantly reduces the frequency of inferences during the scheduling process, effectively improves the overall scheduling performance and resource utilization, optimizes the user experience, solves the problem of user business performance degradation caused by resource competition, and optimizes the long-tail delay phenomenon of the execution delay of services such as functions.
[0010] In some possible implementations, the scheduling system may traverse the first node to determine whether the current concurrency of the target service in the current node is less than the maximum concurrency. If so, the scheduling system deploys an instance of the target service in the current node, and if not, the scheduling system determines whether the current concurrency of the next node is less than the maximum concurrency.
[0011] In this method, the scheduling system can determine the node whose current concurrency of the target service is less than the maximum concurrency in the first node by traversing the first node, so as to realize the deployment of the instance in the node whose current concurrency of the target service is less than the maximum concurrency. It is not necessary to perform an inference every time scheduling, but it can realize one inference and multiple scheduling, thereby reducing the frequency of inference in the scheduling process and improving scheduling efficiency. Moreover, the scheduling system can stop traversing when determining the node whose current concurrency of the target service is less than the maximum concurrency, so as to avoid waste of resources.
[0012] In some possible implementations, when there is no node in the first node whose current concurrency of the target service is less than the maximum concurrency, or there is no first node in the node cluster where an instance of the target service is deployed, the scheduling system can determine the target node from the second node and deploy the instance of the target service at the target node. The second node is a node where no instance of the target service is deployed.
[0013] Taking into account the situation that there may be nodes in the node cluster where no instance of the target service is deployed, or the nodes where the instance of the target service is deployed do not meet the concurrency conditions (for example, the node where the instance of the target service is deployed does not meet the situation that the current concurrency of the target service is less than the maximum concurrency), the scheduling system can determine the target node from the second node where the instance of the target service is not deployed to deploy the instance of the target service, so that high-density instance scheduling can be achieved.
[0014] In some possible implementations, the scheduling system may also construct a delay prediction model for the target node based on the resource utilization of each instance in the target node and the concurrency and execution delay of the service in the target node. Then, the scheduling system predicts at least one execution delay of the target service based on at least one candidate concurrency of the target service in the target node through the delay prediction model of the target node. At least one execution delay corresponds to at least one candidate concurrency. Then, the scheduling system estimates the maximum concurrency of the target service in the target node based on at least one execution delay of the target service.
[0015] This method introduces service concurrency, reduces the dimension of the delay prediction model from the instance level to the service level, and integrates the historical resource utilization data of each service, such as CPU utilization, memory utilization, and bandwidth utilization, to train the delay prediction model, thereby predicting the execution delay of the service, optimizing the inference performance, and greatly saving the consumption of inference resources, which can guarantee performance during large-scale instance scheduling. Moreover, this method can provide assistance for subsequent instance scheduling by estimating the maximum concurrency of the target service in the target node. When the instance of the target service needs to be deployed later, it can be decided whether to deploy the instance of the target service in the node based on the maximum concurrency of the target service in the node, so as to achieve efficient instance scheduling.
[0016] In some possible implementations, when the target service is deployed, the scheduling system can update the maximum concurrency of the service in the target node. This method optimizes the scheduling link by updating the maximum concurrency (or concurrency capacity) of other services in the target node after the instance scheduling of the target service is completed, without affecting the original service scheduling. This method reduces the scheduling link link, reduces the scheduling delay, and optimizes the user experience.
[0017] In some possible implementations, the scheduling system determines the target node from the second node, including:
[0018] The scheduling system determines the target node according to the resource utilization rate or the number of deployment instances of the second node, and the resource utilization rate or the number of deployment instances of the target node meets the set conditions.
[0019] In some possible implementations, the target service is an application, a microservice, or a function. The application can be software as a service, the microservice can be a small service in an application of a microservice architecture, and the function can be a function as a service. In this way, high-density instance scheduling of applications, microservices, or functions can be achieved to meet the needs of different businesses.
[0020] In a second aspect, the present application provides a scheduling system. The scheduling system includes a scheduler and an estimator;
[0021] The scheduler is used to receive a request to create an instance of a target service, and query a node on which an instance of the target service is deployed from a node cluster;
[0022] The scheduler is further configured to, when there is a first node on which an instance of the target service is deployed, obtain a maximum concurrency and a current concurrency of the target service in the first node, wherein the current concurrency of the target service in the first node is characterized by the number of instances of the target service currently deployed by the first node, and the maximum concurrency of the target service in the first node is determined by the estimator according to a delay prediction model of the first node, wherein the delay prediction model of the first node takes the resource utilization of the instance in the first node and the concurrency of the service in the first node as input, and takes the execution delay of the service in the first node as output;
[0023] The scheduler is further configured to deploy an instance of the target service at a target node whose current concurrency of the target service is less than a maximum concurrency when there is a target node in the first node.
[0024] In some possible implementations, the scheduler is specifically used to:
[0025] Traversing the first node, determining whether the current concurrency of the target service in the current node is less than the maximum concurrency;
[0026] If so, deploy an instance of the target service at the current node; if not, determine whether the current concurrency of the next node is less than the maximum concurrency.
[0027] In some possible implementations, the scheduler is further configured to:
[0028] When there is no node in the first node whose current concurrency of the target service is less than the maximum concurrency, or there is no first node in the node cluster on which an instance of the target service is deployed, the target node is determined from a second node, and the instance of the target service is deployed on the target node, and the second node is a node on which an instance of the target service is not deployed.
[0029] In some possible implementations, the scheduling system further includes:
[0030] A trainer, configured to construct a delay prediction model of the target node according to the resource utilization of each instance in the target node and the concurrency and execution delay of the service in the target node;
[0031] The estimator is also used to:
[0032] According to at least one candidate concurrency of the target service in the target node, predicting at least one execution delay of the target service by using a delay prediction model of the target node, wherein the at least one execution delay corresponds to the at least one candidate concurrency;
[0033] A maximum concurrency of the target service in the target node is estimated according to at least one execution delay of the target service.
[0034] In some possible implementations, the estimator is further configured to:
[0035] When the target service is deployed, the maximum concurrency of the service in the target node is updated.
[0036] In some possible implementations, the scheduler is specifically used to:
[0037] The target node is determined according to the resource utilization rate or the number of deployment instances of the second node, and the resource utilization rate or the number of deployment instances of the target node meets the set conditions.
[0038] In some possible implementations, the target service is an application, a microservice, or a function.
[0039] In a third aspect, the present application provides a scheduler, wherein the scheduler includes a processor and a memory, wherein the processor is configured to execute instructions in the memory, so that the scheduler executes the example scheduling method described in the first aspect or any implementation of the first aspect.
[0040] In a fourth aspect, the present application provides a computing device cluster. The computing device cluster includes at least one computing device, and the at least one computing device includes at least one processor and at least one memory. The at least one processor and the at least one memory communicate with each other. The at least one processor is used to execute instructions stored in the at least one memory, so that the computing device or the computing device cluster executes the instance scheduling method as described in the first aspect or any implementation of the first aspect.
[0041] In a fifth aspect, the present application provides a computer-readable storage medium, in which instructions are stored, and the instructions instruct a computing device or a computing device cluster to execute the instance scheduling method described in the above-mentioned first aspect or any implementation of the first aspect.
[0042] In a sixth aspect, the present application provides a computer program product comprising instructions, which, when executed on a computing device or a computing device cluster, enables the computing device or the computing device cluster to execute the instance scheduling method described in the first aspect or any one of the implementations of the first aspect.
[0043] Based on the implementations provided in the above aspects, this application can also be further combined to provide more implementations. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the technical method of the embodiments of the present application, the drawings required for use in the embodiments are briefly introduced below.
[0045] Figure 1 A schematic diagram of a flow chart of instance scheduling of a function provided in this application;
[0046] Figure 2 A schematic diagram of a process flow of instance scheduling based on resource evaluation provided in this application;
[0047] Figure 3 A schematic diagram of the architecture of a scheduling system provided for this application;
[0048] Figure 4 A flowchart of an example scheduling method provided in this application;
[0049] Figure 5 A schematic diagram of an application scenario of an example scheduling method provided in this application;
[0050] Figure 6 A schematic diagram of the training process of a concurrency and execution delay model provided in this application;
[0051] Figure 7 A schematic diagram of the structure of a computing device provided for this application;
[0052] Figure 8 A schematic diagram of the structure of a computing device cluster provided for this application;
[0053] Fig. 9 A schematic diagram of the structure of another computing device cluster provided for this application;
[0054] Fig.10 A schematic diagram of the structure of another computing device cluster provided in this application. DETAILED DESCRIPTION
[0055] The terms "first" and "second" in the embodiments of the present application are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features.
[0056] First, some technical terms involved in the embodiments of the present application are introduced.
[0057] Serverless computing (serverless) is a new paradigm for the next generation of cloud computing. Based on Platform-as-a-Service (PaaS), serverless computing provides a micro-architecture where end customers do not need to deploy, configure or manage servers. The server services required for code execution are all provided by the cloud platform, and users can achieve "pay-per-usage".
[0058] Function as a service (FaaS), also known as "Function as a Service", is a representative service form of Serverless. FaaS allows developers to build, calculate, run and manage their own application packages in the form of functions without maintaining the backend infrastructure. Among them, function refers to the code logic that can be run and can usually be managed by the user. FaaS can be an event-driven execution model running in a stateless container. Functions use the services of FaaS providers to manage server-side logic and state. It should be noted that FaaS is only a service form of serverless. In other possible implementations, serverless can also include other service forms, such as software as a service (SaaS).
[0059] A virtual machine is a node where a service is executed, such as a node where FaaS is executed, or a node where SaaS is executed. Specifically, a virtual machine can be an emulator of a computer system, which can provide the functions of a physical computer by simulating a complete computer system with complete hardware system functions and running in a completely isolated environment through software. In the present application, a virtual machine may include two states, one is a deployed state, which specifically refers to an instance of the service (such as function as a service FaaS or simply referred to as function) that has been deployed on the current virtual machine, and the other is an undeployed state, which specifically refers to an instance of the service (such as function) that does not exist on the current virtual machine. Among them, an instance can be an executable environment allocated by the platform for a service such as a function, which can be used to execute a service such as a function, and the instance can usually exist in the form of a container. It should be noted that the above-mentioned executable environment can also be a virtual machine, and the instance can also exist in the form of a virtual machine.
[0060] Serverless platforms often include dedicated schedulers, such as schedulers. Figure 1As shown in the figure, when a function call request arrives at the serverless platform, the function call request can be forwarded by the gateway to the scheduler. If there is no idle instance at present, the scheduler can create a new instance in a virtual machine according to the scheduling policy. After the instance is successfully created, it can be used to execute the function request.
[0061] In order to improve resource utilization, all platforms have tried to adopt high-density instance scheduling methods. At present, the high-density instance scheduling method widely used in the industry is the instance scheduling method based on resource evaluation. Figure 2 As shown in the figure, when receiving an instance creation request, the scheduler can evaluate the virtual machines based on the idle resources on each virtual machine, including the central processing unit (CPU), memory, and bandwidth, and sort them based on the evaluation results (such as scores). The scheduler can select virtual machines in descending order of scores, and determine whether the resources of the virtual machine can accommodate the current instance, that is, whether the resource requirements for deploying the current instance are met. If so, the instance is created in the target virtual machine.
[0062] The resource evaluation-based scheduling method in the above scheme usually over-allocates resources within the virtual machine, which can lead to resource preemption between different instances within the virtual machine. In CPU-intensive application scenarios (such as computing, gaming, etc.), execution latency may deteriorate, making it impossible to guarantee the user's business performance requirements.
[0063] In view of this, the present application provides an instance scheduling method. The method can be performed by a scheduling system. The scheduling system can be a software system for implementing instance scheduling, and the software system can be an independent software system, or a software system integrated with other software, such as a software system integrated with other software in the form of plug-ins, functional modules, applets, etc. The software system can be provided to customers in the form of a software package, and customers can deploy the software package by themselves on computing devices or computing device clusters such as local data centers or private clouds. Alternatively, the software system can be provided to users in the form of cloud services, for example, the software system can be a function workflow (Function Graph), and the function workflow can be an event-driven function managed computing service. The above software system can be deployed in a computing device cluster, and the computing device cluster executes the program code of the software system, thereby executing the instance scheduling method of the present application. In some examples, the scheduling system can also be a hardware system, such as a computing device cluster with instance scheduling capabilities, and the computing device cluster executes the instance scheduling method of the present application when it is running.
[0064] Specifically, the scheduling system receives an instance creation request for the target service, and then queries the node on which the instance of the target service is deployed from the node cluster. When there is a first node on which an instance of the target service is deployed, the scheduling system obtains the maximum concurrency (or concurrency capacity) and current concurrency of the target service in the first node. Among them, the current concurrency of the target service in the first node is characterized by the number of instances of the target service currently deployed by the first node, and the maximum concurrency of the target service in the first node is determined according to the delay prediction model of the first node. The delay prediction model of the first node takes the resource utilization of the instance in the first node and the concurrency of the service in the first node as input, and uses the execution delay of the service in the first node as the output concurrency and execution delay model to determine. When there is a target node in the first node whose current concurrency of the target service is less than the maximum concurrency, the scheduling system deploys an instance of the target service at the target node.
[0065] This method introduces a maximum concurrency optimization node selection strategy, estimates the maximum concurrency of the target service (such as a function) on a single node (such as a virtual machine) based on the execution delay prediction of the target service, and directly decides on the node used to schedule the instance based on the maximum concurrency, so that batch scheduling can be performed for a single inference, and instances can be scheduled efficiently and densely while ensuring the execution delay. This method significantly reduces the frequency of inferences during the scheduling process, effectively improves the overall scheduling performance and resource utilization, optimizes the user experience, solves the problem of user business performance degradation caused by resource competition, and optimizes the long-tail delay phenomenon of the execution delay of services such as functions.
[0066] This application can be applied to the serverless platform, specifically to the scheduling system in the serverless platform that needs to schedule instances in the resource pool. Figure 3 As shown, the scheduling system is used to receive an instance creation request for a target service, for example, an instance creation request forwarded by a receiving gateway, and then querying the node on which the instance of the target service is deployed from the node cluster. The scheduling system is also used to obtain the maximum concurrency and current concurrency of the target service in the first node when there is a first node on which an instance of the target service is deployed. The current concurrency of the target service in the first node is characterized by the number of instances of the target service currently deployed by the first node, and the maximum concurrency of the target service in the first node is determined according to the delay prediction model of the first node. The delay prediction model of the first node takes the resource utilization of the instance in the first node and the concurrency of the service in the first node as input, and takes the execution delay of the service in the first node as output. When there is a target node in the first node whose current concurrency of the target service is less than the maximum concurrency, an instance of the target service is deployed at the target node.
[0067] It should be noted that the latency prediction model of the present application can be a service-level latency prediction model. Among them, the service can be a function (such as FaaS), an application (such as SaaS) or a microservice. When the service is a function, the latency prediction model is a function-level latency prediction model. The present application reduces the instance-level latency prediction model to a service-level latency prediction model (such as a function-level latency prediction model), which can optimize the single-time inference performance, greatly save the inference resource consumption, and ensure performance during large-scale instance scheduling.
[0068] The following is an introduction to the structure of the scheduling system. Figure 3 As shown, the scheduling system 300 may include a scheduler 302. Furthermore, the scheduling system 300 may also include an estimator 304, a data collector 306, and a trainer 308. Next, the above core units of the scheduling system are introduced.
[0069] When an instance creation request arrives, the scheduler 302 is used to determine the target node for deploying the instance of the target service according to the node status in the node cluster, such as the virtual machine status, and based on the maximum concurrency estimated by the estimator 304. Specifically, the scheduler 302 is used to receive the instance creation request of the target service, query the node where the instance of the target service is deployed from the node cluster, and when there is a first node where the instance of the target service is deployed, obtain the maximum concurrency and current concurrency of the target service in the first node, and when there is a target node in the first node where the current concurrency of the target service is less than the maximum concurrency, deploy the instance of the target service at the target node.
[0070] The estimator 304 is used to estimate the maximum concurrency of services deployed by nodes in the node cluster, for example, the maximum concurrency of the target service deployed in the first node, to provide a reference for the decision of the scheduler 302. Specifically, the estimator 304 is used to determine the maximum concurrency of the target service in the first node according to the delay prediction model of the first node. Among them, the delay prediction model of the first node takes the resource utilization of the instance in the first node and the concurrency of the service in the first node as input, and takes the execution delay of the service in the first node as output. The estimator 304 can use the delay prediction model of the target node to predict at least one execution delay of the target service for at least one candidate concurrency of the target service in the first node, and estimate the maximum concurrency of the target service in the first node according to the at least one execution delay of the target service.
[0071] The data collector 306 is used to collect the resource utilization of each instance in the node and the concurrency and execution delay of the service. The trainer 308 is used to build a delay prediction model of the first node according to the resource utilization of each instance in the node (for example, the first node) and the concurrency and execution delay of the service. The trainer 308 can use a regression algorithm to train the model. The regression algorithm can include but is not limited to a random forest (RF) and a support vector machine algorithm.
[0072] Based on the above scheduling system, the present application also provides an instance scheduling method. The instance scheduling method of the present application is introduced from the perspective of the scheduling system in conjunction with the accompanying drawings.
[0073] See also Figure 4 The flowchart of an example scheduling method shown in FIG. 1 includes the following steps:
[0074] S402: The scheduling system 300 receives a request to create an instance of a target service.
[0075] The target service refers to the service whose instance is to be created. The target service can be a service in a distributed system, including but not limited to an application, a microservice, or a function. Among them, the application can be Software as a Service (SaaS), the microservice can be a microservice in an application based on a microservice architecture, and the function can be Function as a Service (FaaS). The instance creation request can carry the identifier of the target service to request the creation of an instance of the target service. Among them, the identifier of the target service can include the name or identifier (ID) of the target service. For example, when the target service is a function, the instance creation request can include the function name.
[0076] Specifically, the scheduling system 300 may receive an instance creation request of a target service forwarded by a gateway. In some examples, the scheduling system 300 may also receive a service execution request forwarded by a gateway, such as a function call request, and may generate an instance creation request to create an instance if there is no idle instance of the target service.
[0077] S404, the scheduling system 300 searches the node cluster for nodes where the instance of the target service is deployed. If there is a first node where the instance of the target service is deployed, S406 is executed. If there is no first node where the instance of the target service is deployed in the node cluster, S410 is executed.
[0078] Specifically, the scheduling system 300 can traverse the node cluster to query the nodes in the node cluster where the instances of the target service are deployed, wherein the nodes where the instances of the target service are deployed are also referred to as deployed nodes. The scheduling system 300 can first traverse the node cluster to query all nodes in the node cluster where the instances of the target service are deployed, and then execute S406. Alternatively, the scheduling system 300 can also execute S406 when it queries the nodes where the instances of the target service are deployed in the node cluster. When the scheduling system 300 queries the nodes where the instances of the target service are deployed, it can reduce the number of traversed nodes and save the overhead of traversing nodes by executing S406. If the first node where the instance of the target service is deployed does not exist in the node cluster, S410 can be executed.
[0079] S406, the scheduling system 300 obtains the maximum concurrency and current concurrency of the target service in the first node. When there is a target node in the first node whose current concurrency of the target service is less than the maximum concurrency, S408 is executed. When there is no target node in the first node whose current concurrency of the target service is less than the maximum concurrency, S410 is executed.
[0080] The current concurrency (also denoted as concurrency) of the target service in the first node is represented by the number of instances of the target service currently deployed by the first node, and the maximum concurrency (denoted as max-con) of the target service in the first node represents the maximum number of instances of the target service that can be deployed by the first node, which can be determined according to the latency prediction model of the first node. Taking the target service as a function as an example, the current concurrency of the function in the first node can be the number of instances of the function currently running in the first node, and the maximum concurrency of the function in the first node can be the maximum number of instances that the function can run in the first node.
[0081] In specific implementation, the scheduling system 300 can obtain the current concurrency of the target service in the first node by detecting the number of instances of the target service currently running in the first node. The scheduling system 300 can also estimate the maximum concurrency of the service in the first node and store the maximum concurrency, so that when scheduling the instance of the target service, the maximum concurrency of the target service in the first node can be directly obtained to achieve one-time reasoning and batch scheduling.
[0082] For any node in the node cluster, the scheduling system 300 can construct a delay prediction model for the node, and then determine the maximum concurrency of the service in the node based on the delay prediction model of the node. Taking the first node as an example, the scheduling system 300 can construct a delay prediction model for the first node based on the resource utilization of each instance in the first node and the concurrency and execution delay of the service in the first node. The scheduling system 300 can predict at least one execution delay based on at least one candidate concurrency of the target service in the first node through the delay prediction model of the first node. Then, the scheduling system 300 can estimate the maximum concurrency of the target service in the first node based on at least one execution delay. Among them, the scheduling system 300 can determine the concurrency of the preset proportion of the delay not exceeding the historical tail delay (or tail delay) based on at least one predicted delay, and determine the maximum value of the above concurrency as the maximum concurrency. Among them, the preset proportion can be set according to demand, and in some examples, the preset proportion can be 95%.
[0083] The scheduling system 300 can determine whether the current concurrency of the target service in the first node meets the concurrency condition, thereby deciding the node for scheduling the instance of the target service. Among them, the concurrency condition can be that the current concurrency is less than the maximum concurrency, or the current concurrency + 1 ≤ maximum concurrency. For ease of description, this application uses the example of the concurrency condition being that the current concurrency is less than the maximum concurrency. If the current concurrency is less than the maximum concurrency, it means that the node can still deploy at least one instance of the target service. If the current concurrency is greater than or equal to the maximum concurrency, it means that the remaining resources of the node are insufficient to deploy at least one instance of the target service, and you can choose to deploy it on other nodes.
[0084] Among them, the scheduling system 300 can traverse the first node to determine whether the current concurrency of the target service in the current node is less than the maximum concurrency. If so, the scheduling system 300 deploys an instance of the target service at the current node. If not, the scheduling system 300 determines whether the current concurrency of the target service in the next node is less than the maximum concurrency. In some possible implementations, the scheduling system 300 can stop traversing when it is determined that the current concurrency of the target service is less than the target node with the maximum concurrency. If the scheduling system 300 traverses the first node and does not query the first node whose current concurrency of the target service is less than the maximum concurrency, that is, when the first node is unavailable, it can trigger execution S410 to deploy an instance of the target service at the second node.
[0085] S408. The scheduling system 300 deploys an instance of the target service on the target node.
[0086] The scheduling system 300 can deploy the code of the target service in the target node, thereby implementing the deployment of the instance of the target service in the target node. Specifically, the scheduling system 300 can initialize a container in the target node, such as a virtual machine node (or simply referred to as a virtual machine), and deploy the code or image of the target service in the container, thereby implementing the deployment of the instance of the target service in the target node.
[0087] S410. The scheduling system 300 determines a target node from the second nodes, and deploys an instance of a target service at the target node.
[0088] The second node is a node where an instance of the target service is not deployed, and the node where the instance of the target service is not deployed is also called an undeployed node. The scheduling system 300 can select a target node from the second node, estimate the maximum concurrency of the target service in the target node, and deploy an instance of the target service in the target node. Among them, similar to estimating the maximum concurrency of the target service in the first node, the scheduling system 300 can build a delay prediction model for the target node based on the resource utilization of each instance in the target node and the concurrency and execution delay of the service in the target node. It should be noted that when building the delay prediction model of the target node, the resource utilization, concurrency and execution delay of the instance of the target service can be initialized to zero. Accordingly, the scheduling system 300 can determine the maximum concurrency of the target service in the target node according to the delay prediction model of the target node. For example, the scheduling system 300 can input at least one candidate concurrency of the target service in the target node into the delay prediction model of the target node, obtain at least one execution delay, and estimate the concurrency of the preset proportion of the delay not exceeding the tail delay based on at least one delay, thereby obtaining the maximum concurrency of the target service in the target node.
[0089] In order to achieve high-density deployment, the scheduling system 300 can determine the target node according to the resource utilization rate or the number of deployment instances of the second node. The resource utilization rate or the number of deployment instances of the target node meets the set conditions. The set conditions can be set based on experience. In some examples, the set conditions can be that the resource utilization rate is the highest, the resource utilization rate is greater than a set ratio, or the number of deployment instances reaches a set value.
[0090] When the instance of the target service is deployed, the resources that can be allocated to other services in the target node become fewer, and the maximum number of instances that can be deployed will also change accordingly. For this reason, the scheduling system 300 can also update the maximum concurrency of the service in the target node. Similar to estimating the maximum concurrency of the target service, the scheduling system 300 can update the maximum concurrency of the service in the target node through the delay prediction model of the target node. Taking a service other than the target service in the target node as an example, the scheduling system 300 can input at least one candidate concurrency of the service into the delay prediction model of the target node, obtain at least one execution delay, and then determine the concurrency with a delay that does not exceed a preset proportion of the historical tail delay based on at least one predicted delay, and determine the maximum value of the above concurrency as the maximum concurrency, thereby achieving the update of the maximum concurrency of the service.
[0091] It should be noted that the above S410 is an optional step of the embodiment of the present application, and the instance scheduling method of the embodiment of the present application may not perform the above step. For example, when there is a node in the node cluster that deploys an instance of the target service, the scheduling system 300 may not perform the above S410.
[0092] Based on the above description, the instance scheduling method provided by this application introduces the maximum concurrency of the target service, optimizes the node selection strategy based on the maximum concurrency, and can realize batch creation of instances according to the maximum concurrency, for example, giving priority to nodes where instances of the target service have been deployed, realizing multiple scheduling for a single reasoning, effectively reducing the number of reasoning times, and solving the problems of low scheduling efficiency and large instance scheduling delay overhead caused by computational reasoning and data updating during instance scheduling, optimizing the scheduling link, improving scheduling performance, and improving the overall scheduling efficiency of the platform. Moreover, the method can predict the execution delay of the target service in the node, estimate the maximum concurrency that the target service can be deployed on the node while ensuring that the execution delay does not deteriorate, optimize the long-tail delay phenomenon of the target service execution delay, and ensure that the execution delay does not deteriorate while improving the platform resource utilization, thereby ensuring business performance.
[0093] Moreover, after completing the instance scheduling, the maximum concurrency of other services on the node can also be updated without affecting the original service scheduling. This can reduce the scheduling link links, optimize the scheduling link, reduce the scheduling delay, and optimize the user experience.
[0094] In order to make the technical solution of the present application clearer and easier to understand, the example scheduling method of the present application is introduced below in conjunction with a specific application scenario.
[0095] See also Figure 5Schematic diagram of an application scenario of an example scheduling method shown. This method is applied to a scheduling system 300, and the scheduling system 300 includes a scheduler 302, an estimator 304, a data collector 306, and a trainer 308. The scheduling system 300 can perform the following steps to implement instance scheduling:
[0096] ① When a request to create an instance of function f1 arrives, if there is a virtual machine with the status of deployed, the scheduler 302 traverses the currently deployed virtual machines and selects a virtual machine VM-1 for condition judgment.
[0097] Specifically, the scheduler 302 can first traverse the virtual machines in the node cluster to query whether the instances deployed on the virtual machines include the instances of function f1, so as to determine whether there is a virtual machine with the status of deployed. If there is a virtual machine with the status of deployed, the scheduler 302 can select a virtual machine such as VM-1 and determine whether the current concurrency of the virtual machine VM-1 is less than the maximum concurrency.
[0098] Among them, the scheduler 302 can randomly select a virtual machine from the virtual machines with the status of deployed, or select a virtual machine in combination with the resource utilization rate and the number of deployed instances of the virtual machine. For example, it can select a virtual machine with a high resource utilization rate to improve the deployment density as much as possible. In some examples, the scheduler 302 can also select the virtual machine for condition judgment when traversing to a virtual machine with the status of deployed.
[0099] ② The maximum concurrency of function f1 on virtual machine VM-1 is 4, and the current concurrency is 1 (there is 1 instance of function f1). The current concurrency meets the concurrency condition (concurrency < max_con), and the scheduler 302 directly deploys the instance on VM-1.
[0100] In this case, after the instance scheduling of function f1 is completed, the scheduling link does not need to be inferred and updated, and the scheduling efficiency is relatively high.
[0101] ③ When a request to create an instance of function f2 arrives, if there is a virtual machine with the status of deployed, the scheduler 302 traverses the currently deployed virtual machines and selects virtual machine VM-1 for condition judgment.
[0102] For function f2, the virtual machine with the status of deployed refers to the virtual machine on which the instances of function f2 are deployed. The specific implementation of the scheduler 302 traversing the deployed virtual machines and selecting virtual machine VM-1 for condition judgment can refer to the specific implementation of virtual machine traversal and condition judgment when the request to create an instance of function f1 arrives, which will not be elaborated here.
[0103] ④ The maximum concurrency of function f2 on virtual machine VM-1 is 2, the current concurrency is 2, VM-1 has reached the upper limit of f2 concurrency, and scheduler 302 cannot directly deploy the instance on VM-1.
[0104] ⑤ There is currently no available deployed virtual machine, and the scheduler 302 selects a virtual machine VM-2 in the undeployed state to obtain the current concurrency of each function on the virtual machine.
[0105] Specifically, the scheduler 302 can select a virtual machine from the second node of the instance of the undeployed function f2. Similar to selecting a virtual machine for conditional judgment, the scheduler 302 can randomly select a virtual machine from the virtual machines in the undeployed state, or select a virtual machine in combination with resource utilization and the number of deployed instances. Taking into account the deployment density and resource utilization, the scheduler 302 can select a virtual machine with high resource utilization or a virtual machine with a large number of deployed instances based on the resource utilization of the virtual machine or the number of deployed instances of the virtual machine. For example, the scheduler 302 can select a virtual machine with the highest resource utilization, a virtual machine with the largest number of deployed instances, or a virtual machine with a resource utilization higher than a set ratio and a virtual machine with a number of deployed instances greater than a set value from the virtual machines in the undeployed state. The scheduler 302 counts the number of instances of the function deployed on the virtual machine to obtain the current concurrency of each function.
[0106] ⑥ The data collector 306 collects the CPU utilization, memory utilization, bandwidth utilization, concurrency and execution delay of each function, and then trains the model through the random forest algorithm based on the above data to obtain a function-level delay prediction model.
[0107] Specifically, see Figure 6 Schematic diagram of a function-level delay prediction model shown in FIG. 1 , the data collector 306 obtains the resource utilization of each instance on the node (virtual machine), including but not limited to CPU utilization, memory (memory, denoted as mem) utilization, and bandwidth utilization, denoted as X k =(x k1 , x k2 , …x k13 ). Among them, X k Represents the input of a function, x k1 to x k13 To characterize the resource utilization of the instance of the function, the trainer 308 may add the concurrency of the function from the function dimension, such as the current concurrency, specifically represented by X k =(x k1 , x k2 , …x k13 , c k ). Among them, ck The trainer 308 can combine the inputs of each function as the input of the model, denoted as X=(X k , X 1 , …X k-1 ). The trainer 308 may use the execution delay of each function as the output of the model.
[0108] The trainer 308 can be constructed with input X=(X k , X 1 , …X k-1 ), the output is the model of T, as shown below:
[0109]
[0110] Among them, X 1 To X k Characterize the input of functions f1 to fk, T 1 To T k Characterizes the execution delay of functions f1 to fk.
[0111] The trainer 308 can train the above model through a regression algorithm. For example, the trainer 308 can perform model training through a random forest algorithm to obtain a function-level delay prediction model. Random forest is a supervised algorithm that uses an integrated learning method composed of many decision trees, and the output of this method is a consensus on the best answer to the problem. Random forest can be used for classification or regression. Specifically in this embodiment, random forest can be used for the execution delay of the regression function. The model trained by the random forest algorithm is suitable for smaller data sets. Each tree in the random forest can randomly sample a subset of the training data during the self-service aggregation (bagging) process and summarize the prediction results. Through sampling with replacement, multiple instances of the same data can be reused, and multiple decision trees are not only trained based on different data sets, but also use different features to make decisions. Among them, randomness ensures that the correlation between decision trees is low, thereby reducing the risk of bias. The presence of a large number of decision trees also reduces the overfitting problem and improves accuracy.
[0112] ⑦ The estimator 304 uses the concurrency of the current function on VM-2 to predict the execution delay of function f2. Based on the predicted execution delay of function f2, the estimator 304 estimates that the maximum concurrency of function f2 is 4 when the function execution delay does not exceed the user's historical 95% tail delay.
[0113] When estimating the maximum concurrency of function f2, the estimator 304 may directly determine the concurrency with the delay closest to 95% but not exceeding 95 as the maximum concurrency. Alternatively, the estimator 304 may also fit the execution delay and concurrency of function f2, and determine the concurrency corresponding to the delay of 95% of the tail delay as the maximum concurrency.
[0114] ⑧ Scheduler 302 deploys the instance of function f2 on virtual machine VM-2 and updates the maximum concurrency on VM-2 to 4. At this point, the instance scheduling of f2 is completed, and subsequent instance creation requests of f2 can be scheduled according to steps ①-② until the instance concurrency of f2 function on VM-2 reaches the upper limit.
[0115] ⑨ The scheduler 302 delays updating the maximum concurrency of the function on the virtual machine.
[0116] After the instance scheduling of the function f2 is completed, the estimator 304 may also update the maximum concurrency of other functions on the virtual machine VM-2 to optimize the overall scheduling link.
[0117] Based on the above description, the instance scheduling method of this application introduces the maximum concurrency of the function on the basis of ensuring the function execution latency performance, which not only optimizes the single reasoning performance, but also batch schedules the function instances based on the maximum concurrency of the function, greatly reducing the frequency of reasoning during the scheduling process, optimizing the overall scheduling performance, and improving the scheduling efficiency. Moreover, by predicting the function execution latency and combining it with the concurrency estimation, it can ensure that the function latency performance is effectively guaranteed on the basis of improving the deployment density and improving the user experience.
[0118] In addition, this method introduces the concurrency of functions, reduces the dimension of the latency prediction model from the instance level to the function level, and integrates the CPU, memory, bandwidth and other resource utilization data of each function to perform model training, so as to predict the execution latency at the function level, optimize the single inference performance, greatly save the inference resource consumption, and ensure performance during large-scale instance scheduling.
[0119] Based on the above-mentioned example scheduling method, the present application also provides a scheduling system. Figure 3 As shown, scheduling system 300 includes a scheduler 302 and an estimator 304 .
[0120] Scheduler 302, used to receive a request to create an instance of a target service, and query a node in a node cluster where an instance of the target service is deployed;
[0121] The scheduler 302 is further used to obtain the maximum concurrency and current concurrency of the target service in the first node when there is a first node on which an instance of the target service is deployed, the current concurrency of the target service in the first node is characterized by the number of instances of the target service currently deployed by the first node, and the maximum concurrency of the target service in the first node is determined by the estimator 304 according to the delay prediction model of the first node, and the delay prediction model of the first node takes the resource utilization of the instance in the first node and the concurrency of the service in the first node as input, and takes the execution delay of the service in the first node as output;
[0122] The scheduler 302 is further configured to deploy an instance of the target service on the target node when there is a target node in the first node whose current concurrency of the target service is less than the maximum concurrency.
[0123] The above scheduler 302 and estimator 304 may be implemented by software or hardware. The implementation by software and the implementation by hardware are described below respectively.
[0124] When implemented by software, the scheduler 302 and the estimator 304 may be applications running on a computer device, such as a computing engine. The application may be provided to users through virtualization services. Virtualization services may include virtual machine (VM) services, bare metal server (BMS) services, and container services. Among them, the VM service may be a service that virtualizes a virtual machine (VM) resource pool on multiple physical hosts through virtualization technology to provide users with VMs on demand for use. The BMS service is a service that virtualizes a BMS resource pool on multiple physical hosts to provide users with BMS on demand for use. The container service is a service that virtualizes a container resource pool on multiple physical hosts to provide users with containers on demand for use. VM is a simulated virtual computer, that is, a logical computer. BMS is a high-performance computing service that can be elastically scalable, and its computing performance is no different from that of a traditional physical machine, and it has the characteristics of secure physical isolation. Containers are a kernel virtualization technology that can provide lightweight virtualization to achieve the purpose of isolating user space, processes, and resources. It should be understood that the VM service, BMS service and container service in the above-mentioned virtualization services are only specific examples. In actual applications, virtualization services can also be other lightweight or heavyweight virtualization services, which are not specifically limited here.
[0125] When implemented by hardware, the scheduler 302 and the estimator 304 may include at least one computing device, such as a server, etc. Alternatively, the scheduler 302 and the estimator 304 may also be implemented by using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0126] In some possible implementations, the scheduler 302 is specifically configured to:
[0127] Traversing the first node, determining whether the current concurrency of the target service in the current node is less than the maximum concurrency;
[0128] If so, deploy an instance of the target service at the current node; if not, determine whether the current concurrency of the next node is less than the maximum concurrency.
[0129] In some possible implementations, the scheduler 302 is further configured to:
[0130] When there is no node in the first node whose current concurrency of the target service is less than the maximum concurrency, or there is no first node in the node cluster on which an instance of the target service is deployed, the target node is determined from a second node, and the instance of the target service is deployed on the target node, and the second node is a node on which an instance of the target service is not deployed.
[0131] In some possible implementations, the scheduling system 300 further includes:
[0132] A trainer 308, configured to construct a delay prediction model of the target node according to the resource utilization of each instance in the target node and the concurrency and execution delay of the service in the target node;
[0133] The estimator 304 is also used to:
[0134] According to at least one candidate concurrency of the target service in the target node, predicting at least one execution delay of the target service by using a delay prediction model of the target node, wherein the at least one execution delay corresponds to the at least one candidate concurrency;
[0135] A maximum concurrency of the target service in the target node is estimated according to at least one execution delay of the target service.
[0136] The resource utilization rate of each instance in the target node, the concurrency of services in the target node, the execution delay, etc. may be collected by the data collector 306 .
[0137] The data collector 306 and the trainer 308 may be implemented through software or hardware.
[0138] When implemented by software, the data collector 306 and the trainer 308 may be an application running on a computer device, such as a computing engine. The application may be provided to the user through a virtualization service. The virtualization service may include a VM service, a BMS service, or a container service. When implemented by hardware, the data collector 306 and the trainer 308 may include at least one computing device, such as a server. Alternatively, the data collector 306 and the trainer 308 may also be a device implemented using an application-specific integrated circuit ASIC or a programmable logic device PLD.
[0139] In some possible implementations, the estimator is further configured to:
[0140] When the target service is deployed, the maximum concurrency of the service in the target node is updated.
[0141] In some possible implementations, the scheduler is specifically used to:
[0142] The target node is determined according to the resource utilization rate or the number of deployment instances of the second node, and the resource utilization rate or the number of deployment instances of the target node meets the set conditions.
[0143] In some possible implementations, the target service is an application, a microservice, or a function.
[0144] The present application also provides a computing device 700. Figure 7 As shown, computing device 700 includes: bus 702, processor 704, memory 706 and communication interface 708. Processor 704, memory 706 and communication interface 708 communicate through bus 702. Computing device 700 can be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in computing device 700.
[0145] The bus 702 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 7 The bus 702 may include a path for transmitting information between various components of the computing device 700 (eg, the memory 706, the processor 704, and the communication interface 708).
[0146] The processor 704 may include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0147] The memory 706 may include a volatile memory, such as a random access memory (RAM). The memory 706 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD) or a solid state drive (SSD). The memory 706 stores executable program code, and the processor 704 executes the executable program code to implement the aforementioned example scheduling method. Specifically, the memory 706 stores instructions for the scheduling system 300 to execute the example scheduling method.
[0148] The communication interface 708 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 700 and other devices or communication networks.
[0149] The embodiment of the present application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.
[0150] like Figure 8As shown, the computing device cluster includes at least one computing device 700. The memory 706 in one or more computing devices 700 in the computing device cluster may store the same instructions of the scheduling system 300 for executing the example scheduling method.
[0151] In some possible implementations, one or more computing devices 700 in the computing device cluster may also be used to execute some instructions of the scheduling system 300 for executing the example scheduling method. In other words, a combination of one or more computing devices 700 may jointly execute the instructions of the scheduling system 300 for executing the example scheduling method.
[0152] It should be noted that the memory 706 in different computing devices 700 in the computing device cluster may store different instructions for executing partial functions of the scheduling system 300 .
[0153] Fig. 9 A possible implementation is shown. Fig. 9 As shown, two computing devices 700A and 700B are connected via a communication interface 708. The memory in computing device 700A stores instructions for executing the functions of scheduler 302. The memory in computing device 700B stores instructions for executing the functions of estimator 304. In other words, the memory 706 of computing devices 700A and 700B jointly stores instructions for scheduling system 300 to execute the example scheduling method. Further, the memory 706 of computing device 700B can also store instructions for executing the functions of data collector 306 and trainer 308.
[0154] Fig. 9 The connection mode between the computing device clusters shown may be considered to be that the example scheduling method provided in this application requires a lot of computing power for model training and reasoning. Therefore, it is considered to hand over the functions implemented by the estimator 304 to the computing device 700B for execution.
[0155] It should be understood that Fig. 9 The functions of the computing device 700A shown in FIG. 7 may also be completed by multiple computing devices 700. Similarly, the functions of the computing device 700B may also be completed by multiple computing devices 700.
[0156] In some possible implementations, one or more computing devices in the computing device cluster may be connected via a network, which may be a wide area network or a local area network. Fig.10 A possible implementation is shown. Fig.10As shown, two computing devices 700C and 700D are connected via a network. Specifically, the network is connected via a communication interface in each computing device. In this type of possible implementation, the memory 706 in the computing device 700C stores instructions for executing the functions of the scheduler 302. At the same time, the memory 706 in the computing device 700D stores instructions for executing the functions of the estimator 304. Further, the memory 706 of the computing device 700D may also store instructions for executing the functions of the data collector 306 and the trainer 308.
[0157] Fig.10 The connection method between the computing device clusters shown may be that considering that the example scheduling method provided in the present application requires a large amount of computing power for model training and reasoning, it is considered that the functions implemented by the estimator 304 are handed over to the computing device 700D for execution.
[0158] It should be understood that Fig.10 The functions of the computing device 700C shown in FIG. 700A may also be completed by multiple computing devices 700. Similarly, the functions of the computing device 700D may also be completed by multiple computing devices 700.
[0159] In addition, an embodiment of the present application further provides a scheduler, which includes a processor and a memory. The memory stores instructions, and the processor is used to execute the instructions in the memory so that the scheduler executes the aforementioned example scheduling method.
[0160] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk). The computer-readable storage medium includes instructions that instruct the computing device to execute the above-mentioned application to the scheduling system 300 for executing the example scheduling method.
[0161] The embodiment of the present application also provides a computer program product including instructions. The computer program product may be software or a program product including instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device executes the above-mentioned example scheduling method.
[0162] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.
Claims
1. An instance scheduling method, characterized in that: The method comprises: The scheduling system receives a request to create an instance of the target service; The scheduling system queries the node where the instance of the target service is deployed from the node cluster; When there is a first node on which an instance of the target service is deployed, the scheduling system obtains a maximum concurrency and a current concurrency of the target service in the first node, the current concurrency of the target service in the first node being represented by the number of instances of the target service currently deployed by the first node, and the maximum concurrency of the target service in the first node being determined according to a delay prediction model of the first node, the delay prediction model of the first node taking resource utilization of the instances in the first node and the concurrency of the service in the first node as input, and taking execution delay of the service in the first node as output; When there is a target node whose current concurrency of the target service is less than the maximum concurrency among the first nodes, the scheduling system deploys an instance of the target service at the target node.
2. The method according to claim 1, characterized in that The scheduling system deploys an instance of the target service at the target node, including: The scheduling system traverses the first node to determine whether the current concurrency of the target service in the current node is less than the maximum concurrency; If so, the scheduling system deploys an instance of the target service at the current node; if not, the scheduling system determines whether the current concurrency of the next node is less than the maximum concurrency.
3. The method according to claim 1 or 2, characterized in that: The method further comprises: When there is no node in the first node whose current concurrency of the target service is less than the maximum concurrency, or there is no first node in the node cluster on which an instance of the target service is deployed, the scheduling system determines the target node from the second node and deploys the instance of the target service on the target node, and the second node is a node on which an instance of the target service is not deployed.
4. The method according to claim 3, characterized in that The method further comprises: The scheduling system constructs a delay prediction model for the target node according to the resource utilization rate of each instance in the target node and the concurrency and execution delay of the service in the target node; The scheduling system predicts at least one execution delay of the target service in the target node according to at least one candidate concurrency of the target service in the target node through a delay prediction model of the target node, wherein the at least one execution delay corresponds to the at least one candidate concurrency; The scheduling system estimates a maximum concurrency of the target service in the target node according to at least one execution delay of the target service.
5. The method according to claim 4, characterized in that The method further comprises: When the target service is deployed, the scheduling system updates the maximum concurrency of the service in the target node.
6. The method according to any one of claims 3 to 5, characterized in that: The scheduling system determines the target node from the second node, including: The scheduling system determines the target node according to the resource utilization rate or the number of deployment instances of the second node, and the resource utilization rate or the number of deployment instances of the target node meets the set conditions.
7. The method according to any one of claims 1 to 6, characterized in that: The target service is an application, a microservice, or a function.
8. A scheduling system, characterized in that: The scheduling system includes a scheduler and an estimator; The scheduler is used to receive a request to create an instance of a target service, and query a node on which an instance of the target service is deployed from a node cluster; The scheduler is further configured to, when there is a first node on which an instance of the target service is deployed, obtain a maximum concurrency and a current concurrency of the target service in the first node, wherein the current concurrency of the target service in the first node is characterized by the number of instances of the target service currently deployed by the first node, and the maximum concurrency of the target service in the first node is determined by the estimator according to a delay prediction model of the first node, wherein the delay prediction model of the first node takes the resource utilization of the instance in the first node and the concurrency of the service in the first node as input, and takes the execution delay of the service in the first node as output; The scheduler is further configured to deploy an instance of the target service at a target node whose current concurrency of the target service is less than a maximum concurrency when there is a target node in the first node.
9. The dispatching system according to claim 8, characterized in that: The scheduler is specifically used for: Traversing the first node, determining whether the current concurrency of the target service in the current node is less than the maximum concurrency; If so, deploy an instance of the target service at the current node; if not, determine whether the current concurrency of the next node is less than the maximum concurrency.
10. The dispatching system according to claim 8 or 9, characterized in that: The scheduler is also used to: When there is no node in the first node whose current concurrency of the target service is less than the maximum concurrency, or there is no first node in the node cluster on which an instance of the target service is deployed, the target node is determined from a second node, and the instance of the target service is deployed on the target node, and the second node is a node on which an instance of the target service is not deployed.
11. The dispatching system according to claim 10, characterized in that: The scheduling system also includes: A trainer, configured to construct a delay prediction model of the target node according to the resource utilization of each instance in the target node and the concurrency and execution delay of the service in the target node; The estimator is also used to: According to at least one candidate concurrency of the target service in the target node, predicting at least one execution delay of the target service by using a delay prediction model of the target node, wherein the at least one execution delay corresponds to the at least one candidate concurrency; A maximum concurrency of the target service in the target node is estimated according to at least one execution delay of the target service.
12. The dispatching system according to claim 11, characterized in that: The estimator is also used to: When the target service is deployed, the maximum concurrency of the service in the target node is updated.
13. The dispatching system according to any one of claims 10 to 12, characterized in that: The scheduler is specifically used for: The target node is determined according to the resource utilization rate or the number of deployment instances of the second node, and the resource utilization rate or the number of deployment instances of the target node meets the set conditions.
14. The dispatching system according to any one of claims 8 to 13, characterized in that: The target service is an application, a microservice, or a function.
15. A scheduler, characterized in that: The scheduler includes a processor and a memory, wherein the memory stores computer-readable instructions; the processor executes the computer-readable instructions to implement the functions of the scheduler in the scheduling system according to any one of claims 8 to 14.
16. A computing device cluster, characterized in that: The computing device cluster includes at least one computing device, and the at least one computing device includes at least one processor and at least one memory, wherein the at least one memory stores computer-readable instructions; the at least one processor executes the computer-readable instructions so that the computing device cluster executes the instance scheduling method as described in any one of claims 1 to 7.
17. A computer-readable storage medium, characterized in that: The method comprises computer-readable instructions; the computer-readable instructions are used to implement the example scheduling method according to any one of claims 1 to 7.
18. A computer program product, characterized in that The method comprises computer-readable instructions; the computer-readable instructions are used to implement the example scheduling method according to any one of claims 1 to 7.
Citation Information
Cited By
Instance scheduling method and device, hybrid scheduling component, storage medium and product
CN121996360A