Task scheduling method and related equipment

By perceiving the state and performance models of nodes in the inference cluster, precise scheduling of tasks is achieved, and the problems of low scheduling accuracy and low resource utilization in the AI ​​computing cluster are solved, which improves response speed and reduces operation and maintenance costs.

CN120045293APending Publication Date: 2025-05-27HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410231514.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-24
Filing Date
2024-02-29
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

It is difficult for AI computing clusters to select appropriate instance nodes to process requests, resulting in low scheduling accuracy, long response time, low computing resource utilization, and increased operation and maintenance costs.

Method used

By perceiving the state and performance models of nodes such as proxy servers in the inference cluster, precise scheduling of tasks is realized, so that the computing resources of the inference server can be fully utilized.

Benefits of technology

It improves resource utilization, shortens request response time, meets business needs, and reduces operation and maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045293A_ABST
    Figure CN120045293A_ABST
Patent Text Reader

Abstract

The invention provides a task scheduling method, which is applied to a scheduling system, the scheduling system comprises a scheduling cluster and a reasoning cluster, the scheduling cluster comprises a proxy server, and the reasoning cluster comprises a plurality of reasoning servers. And the reasoning server and the proxy server establish full duplex connection based on the TCP. The scheduling cluster obtains the state of at least one inference server, receives a task request based on a bidirectional communication protocol, determines a target inference server matched with the task request from the inference cluster according to the task request, the state of the inference server in the inference cluster and a performance model of the inference server, and sends the determined target inference server to the scheduling cluster; and then scheduling the task request to a target reasoning server through full duplex connection with the target reasoning server. According to the method, the state of the reasoning server in the reasoning cluster can be sensed, and the proper reasoning server can be matched for the task request based on the state, so that the response speed of the request can be improved, and the response time delay can be shortened as much as possible.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims the priority of a Chinese patent application with the application number 202311602182.5 and the invention title "A Task Scheduling Method and Related Devices", which was filed with the National Intellectual Property Administration on November 24, 2023, and the entire content of which is incorporated herein by reference. Technical Field

[0002] This application relates to the field of computer technologies, and in particular, to a task scheduling method, a scheduling system, a computing device cluster, a computer-readable storage medium, and a computer program product. Background Art

[0003] Artificial intelligence (AI) technology is playing an increasingly important role in the software development process and has become an important tool for developers to improve software development efficiency. With the continuous growth of usage scenarios, model types, and access volumes, there will be problems such as how to improve the utilization rate of computing resources and shorten the request response time to provide a better user experience.

[0004] In applications developed based on AI technology, the utilization rate of computing resources and the request response time are usually related to the scheduling ability of the AI computing cluster. If the AI computing cluster can schedule requests to appropriate nodes, it can improve the utilization rate of computing resources and shorten the response time. However, the network addresses of the instance nodes in the AI computing cluster are usually Internet Protocol (IP) addresses in the local area network, such as the host IP in a Virtual Private Cloud (VPC) network. This makes it difficult to send requests to the specified instance nodes, and instead, they are scheduled through a load balancing component. It is difficult for the AI computing cluster to select appropriate instance nodes to process requests, resulting in a relatively low scheduling accuracy, usually a long request response time, which is difficult to meet business requirements, and a low utilization rate of computing resources, increasing the operation and maintenance costs. Summary of the Invention

[0005] This application provides a task scheduling method. This method can sense the status and performance models of nodes such as proxy servers in the inference cluster. Based on the status and performance models of the inference servers, accurate scheduling of tasks can be achieved, enabling the computing resources of the inference servers to be fully utilized, improving the resource utilization rate, shortening the request response time, meeting business requirements, and reducing operation and maintenance costs. Moreover, this application also provides a scheduling system, a computing device cluster, a computer-readable storage medium, and a computer program product corresponding to the above task scheduling method.

[0006] In a first aspect, the present application provides a task scheduling method. This method is applied to a scheduling system. The scheduling system includes a scheduling cluster and an inference cluster. Among them, the scheduling cluster includes at least one proxy server, and the inference cluster includes multiple inference servers. The inference servers establish full-duplex connections based on the Transmission Control Protocol (TCP) with at least one proxy server.

[0007] Specifically, the scheduling cluster can obtain the status of at least one inference server in the inference cluster, for example, continuously obtain the status of at least one inference server in the inference cluster. The scheduling cluster can receive task requests based on a two-way communication protocol, and determine a target inference server in the inference cluster that matches the task request according to the task request, the status of at least one inference server in the inference cluster, and the performance model of at least one inference server. The scheduling cluster schedules the task request to the target inference server through the full-duplex connection with the target inference server.

[0008] In this method, the scheduling cluster can perceive the status of the inference servers in the inference cluster, and based on this status, it can match a suitable inference server for the task request and schedule the task request to the suitable inference server for processing. In this way, the response speed of the request can be improved, and the response latency can be shortened as much as possible. Moreover, different from the scheduling scheme based on the HyperText Transfer Protocol (HTTP), this method can return the result to the user through a full-duplex connection without using the polling method to query the result, reducing the high latency and high network overhead caused by using the polling method to query the result in the HTTP communication mode.

[0009] In some possible implementation manners, the performance model includes at least one of an inference ability model or a robustness model. The scheduling cluster can also collect the inference results or online time of the inference servers in the inference cluster in a historical time period, model based on the inference results of the inference servers in the inference cluster in the historical time period to obtain an inference ability model, and / or the scheduling cluster models based on the online time of the inference servers to obtain a robustness model.

[0010] By collecting the inference results or online time of the inference servers in the historical time period, this method can model the inference ability or robustness of the inference servers. Based on the modeling results, a suitable inference server can be matched for subsequent inference tasks, avoiding the situation where the inference ability or robustness of the inference server is difficult to meet the task requirements.

[0011] In some possible implementation manners, the target inference server includes a first inference server. Correspondingly, the scheduling cluster may also receive feedback from the first inference server, where the feedback is used to indicate that the first inference server is in a busy state. The scheduling cluster may also determine a second inference server that matches the task request from the inference cluster according to the task request, the states of the at least one inference server in the inference cluster, and the performance models of the at least one inference server. The scheduling cluster may schedule the task request to the second inference server through a full-duplex connection with the second inference server.

[0012] This method supports allowing the inference server to return a busy state (or a busy status) in the case of high concurrency and resource preemption, so that the scheduling cluster can rematch the inference server to ensure high availability.

[0013] In some possible implementation manners, if the number of times the scheduling cluster determines a target inference server that matches the task request reaches the maximum number of matching times, and each time the target inference server that is matched returns feedback indicating a busy state, a quick failure event is returned to the user.

[0014] This method supports that when the system is overloaded, the request is thrown back N times (customizable by the user), quickly returning the system busy state, such as returning a quick failure event, to avoid the user's waiting time from timing out or the system burden being aggravated by continuous retries during waiting.

[0015] In some possible implementation manners, the scheduling cluster may also retrieve similar corpora of the input corpus in the task request from the vector database. The scheduling cluster may obtain prompt information according to the input corpus and the similar corpora, for example, by splicing the similar corpora and the input corpus to obtain the prompt information, and the prompt information is used for inference by the model in the inference server. In this way, the quality of the prompt information can be improved, and further the accuracy of the inference result of the model in the inference server can be improved.

[0016] In some possible implementation manners, the full-duplex connection is a long connection. When the user enables word-by-word inference, the scheduling cluster may also receive multiple inference results of the target inference server through the long connection with the target inference server. Among them, the multiple inference results are the inference results of word-by-word inference, and then the scheduling cluster may return the multiple inference results in sequence.

[0017] This method uses a long connection for the transmission of inference results, solving the problems of high latency and high network overhead caused by using polling to query results in the traditional communication method, which makes it difficult to support word-by-word generation of tokens. It can smoothly support the interactive method of word-by-word generation of tokens and improve the user experience.

[0018] In some possible implementation manners, the scheduling cluster may also register the address of the scheduling cluster with the registration center. Correspondingly, the inference cluster may obtain the address of the scheduling cluster from the registration center and send a connection establishment request to the scheduling cluster. The connection establishment request is used to establish a full-duplex connection.

[0019] In this method, the full-duplex connection is initiated by the inference server in the inference cluster, which can accurately connect to the machines in the scheduling cluster. In this way, only a pair of virtual private cloud terminal node services are required to connect the two clusters, reducing the network connection cost between the clusters and eliminating the security risks brought by peer-to-peer connections. Moreover, an inference server (inference service) may include multiple connection channels, and a reconnection mechanism can be provided based on the multiple connection channels to improve fault tolerance and reliability.

[0020] In some possible implementation manners, the full-duplex connection is a connection based on a network socket, and the connection establishment request is a request based on a network socket. The connection establishment request includes the address of the proxy server in the scheduling cluster. Correspondingly, the inference cluster may route the connection establishment request to the proxy server identified by the address in the connection establishment request through the proxy component.

[0021] The proxy component can proxy multiple inference servers and establish connections with the proxy servers in the scheduling machines. The proxy component uses header files or paths for screening and can accurately connect to the proxy servers in the scheduling cluster, reducing the network connection cost between the clusters and eliminating the security risks brought by peer-to-peer connections, and realizing automatic discovery, interconnection, and connection maintenance of cluster nodes.

[0022] In some possible implementation manners, the inference cluster may also send the status of at least one inference server to the registration center. Correspondingly, the scheduling cluster obtains the status of at least one inference server from the registration center.

[0023] This method supports the scheduling cluster to automatically discover inference servers through the registration discovery mechanism and obtain the status of the inference servers from the registration center, providing reference information for subsequent task scheduling decisions.

[0024] In some possible implementation manners, the proxy server in the scheduling cluster may also create a client connection pool and update the client connection pool according to the connection information with the client. Alternatively, the proxy server in the scheduling cluster may also create an inference service connection pool and update the inference service connection pool according to the connection information with the inference server.

[0025] This method manages the connections with the client and the inference server by creating connection pools, providing reference information for subsequent task scheduling to achieve accurate scheduling.

[0026] Second aspect, the present application provides a scheduling system. The scheduling system includes a scheduling cluster and an inference cluster. The scheduling cluster includes at least one proxy server, and the inference cluster includes multiple inference servers. The inference servers establish full-duplex connections with the at least one proxy server based on the Transmission Control Protocol (TCP).

[0027] The scheduling cluster is configured to obtain the status of at least one inference server in the inference cluster, receive task requests based on a bidirectional communication protocol, and determine a target inference server that matches the task request from the inference cluster according to the task request, the status of the at least one inference server in the inference cluster, and the performance model of the at least one inference server.

[0028] The scheduling cluster is further configured to schedule the task request to the target inference server through the full-duplex connection with the target inference server.

[0029] In some possible implementation manners, the performance model includes at least one of an inference ability model or a robustness model. The scheduling cluster is further configured to:[[]]

[0030] Collect the inference results or the online time of the inference servers in the inference cluster during a historical time period.

[0031] Build a model based on the inference results of the inference servers in the inference cluster during the historical time period to obtain the inference ability model, and / or build a model based on the online time of the inference servers to obtain the robustness model.

[0032] In some possible implementation manners, the target inference server includes a first inference server. The scheduling cluster is further configured to:[[]]

[0033] Receive feedback from the first inference server, where the feedback is used to indicate that the first inference server is in a busy state.

[0034] Determine a second inference server that matches the task request from the inference cluster according to the task request, the status of the at least one inference server in the inference cluster, and the performance model of the at least one inference server.

[0035] Schedule the task request to the second inference server through the full-duplex connection with the second inference server.

[0036] In some possible implementation manners, the scheduling cluster is further configured to:[[]]

[0037] If the number of times to determine the target inference server that matches the task request reaches the maximum number of matches, and the target inference server for each match returns feedback indicating a busy state, a quick failure event is returned to the user.

[0038] In some possible implementation manners, the scheduling cluster is further configured to:

[0039] Retrieve similar corpora of the input corpus in the task request from the vector database;

[0040] Obtain prompt information according to the input corpus and the similar corpora, where the prompt information is used for inference by a model in the inference server.

[0041] In some possible implementation manners, the full-duplex connection is a long connection, and the scheduling cluster is further configured to:

[0042] When the user enables word-by-word inference, receive multiple inference results of the target inference server through the long connection with the target inference server, where the multiple inference results are inference results of word-by-word inference, and return the multiple inference results in sequence.

[0043] In some possible implementation manners, the scheduling cluster is further configured to:

[0044] Register the address of the scheduling cluster with the registration center;

[0045] The inference cluster is configured to:

[0046] Obtain the address of the scheduling cluster from the registration center, and send a connection establishment request to the scheduling cluster, where the connection establishment request is used to establish a full-duplex connection.

[0047] In some possible implementation manners, the full-duplex connection is a connection based on a network socket, the connection establishment request is a request based on a network socket, the connection establishment request includes the address of a proxy server in the scheduling cluster, and the inference cluster is specifically configured to:

[0048] Route the connection establishment request to the proxy server identified by the address in the connection establishment request through a proxy component.

[0049] In some possible implementation manners, the inference cluster is further configured to:

[0050] Send the status of the at least one inference server to the registration center;

[0051] The scheduling cluster is specifically configured to:

[0052] Obtain the status of the at least one inference server from the registration center.

[0053] In some possible implementations, the scheduling cluster is further configured to:

[0054] Create a client connection pool and update the client connection pool according to the connection information with the client; or,

[0055] Create an inference service connection pool and update the inference service connection pool according to the connection information with the inference server.

[0056] In a third aspect, the present application provides a computing device cluster. The computing device cluster includes at least one computing device, and the at least one computing device includes at least one processor and at least one memory. The at least one processor and the at least one memory communicate with each other. The at least one processor is configured to execute instructions stored in the at least one memory, so that the computing device or the computing device cluster executes the task scheduling method as described in the first aspect or any implementation manner of the first aspect.

[0057] In a fourth aspect, the present application provides a computer-readable storage medium, in which instructions are stored, and the instructions instruct a computing device or a computing device cluster to execute the task scheduling method as described in the first aspect or any implementation manner of the first aspect.

[0058] In a fifth aspect, the present application provides a computer program product including instructions, which, when running on a computing device or a computing device cluster, causes the computing device or the computing device cluster to execute the task scheduling method as described in the first aspect or any implementation manner of the first aspect.

[0059] Based on the implementation manners provided in the above aspects, the present application can be further combined to provide more implementation manners. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] To more clearly illustrate the technical solutions in the embodiments of the present application, the accompanying drawings required for use in the embodiments will be briefly introduced below.

[0061] Figure 1 FIG. is a schematic flowchart of task scheduling based on HTTP provided by the present application;

[0062] Figure 2 FIG. is a schematic architecture diagram of a scheduling system provided by the present application;

[0063] Figure 3 FIG. is a schematic diagram of a scheduling cluster establishing a connection with an inference cluster provided by the present application;

[0064] Figure 4 FIG. is a schematic flowchart of connection establishment and inference provided by the present application;

[0065] Figure 5 Flowchart of a task scheduling method provided by this application

[0066] Figure 6 Schematic diagram of the process for matching an inference server for a task request provided by this application

[0067] Figure 7 Schematic diagram of the process of a task scheduling method provided by this application in different scenarios

[0068] Figure 8 Schematic diagram of the application scenario of a task scheduling method provided by this application

[0069] Figure 9 Schematic diagram of the structure of a scheduling cluster provided by this application

[0070] Figure 10 Schematic diagram of the structure of a computing device provided by this application

[0071] Figure 11 Schematic diagram of the structure of a computing device cluster provided by this application

[0072] Figure 12 Another schematic diagram of the structure of a computing device cluster provided by this application

[0073] Figure 13 Yet another schematic diagram of the structure of a computing device cluster provided by this application Detailed implementation manners

[0074] In the embodiments of this application, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features

[0075] First, some technical terms involved in the embodiments of this application are introduced

[0076] Artificial intelligence (AI) refers to machines or computers that mimic human cognitive functions related to human thinking. Depending on the technical routes adopted, AI can be classified into different types. Currently, widely used AI in the industry includes Generative AI, and the content generated by Generative AI is called Artificial Intelligence Generative Content (AIGC). AIGC can include different types such as text, images, or videos. For example, users can input a piece of text, and the AI can generate an image or video corresponding to that text.

[0077] AI can be applied to software development to improve software development efficiency. Specifically, based on AI, an AI model for performing specific tasks can be trained, and an application with corresponding capabilities or functions can be built through the AI model. Taking a text-image conversion application (such as text-to-picture) as an example, an AI model for generating images based on text can be trained based on AI, and a text-image conversion application can be built based on the above AI model.

[0078] An AI computing cluster refers to a cluster used to provide AI computing capabilities such as model training and model inference capabilities. An AI computing cluster can include a scheduler cluster and a worker cluster. Among them, the scheduler cluster is also called a collector cluster or a scheduling cluster. The worker cluster includes multiple workers, and each worker deploys an instance of the application to execute tasks, such as AIGC tasks. Therefore, a worker can also be called an instance node. The scheduler cluster can schedule task requests from the clients of a large number of users to the workers in the worker cluster.

[0079] The network address of an instance node (such as the above worker) in an AI computing cluster is usually an IP address in the local area network, such as the host IP in the VPC network. It is difficult to send requests to a specified instance node, but it is scheduled through a load balancing component. Currently, the AI computing cluster provides a scheduling solution based on the Hyper Text Transfer Protocol (HTTP).

[0080] The scheduling solution based on HTTP is as Figure 1 shown. The scheduler cluster receives the tasks sent by the client and actively allocates the tasks to the workers in the worker cluster. The worker processes the tasks scheduled to this worker and sends the results to the scheduler. The client can query the results from the scheduler.

[0081] However, it is difficult for the scheduler cluster to perceive the status of nodes in the worker cluster, making it difficult to achieve precise scheduling of requests. Requests can be scheduled to busy workers, resulting in untimely responses. At the same time, the computing resources of idle workers are not fully utilized, increasing the operation and maintenance costs.

[0082] In view of this, the present application provides a task scheduling method. This method can be executed by a scheduling system. The scheduling system can be a software system deployed in a computing device cluster. The computing device cluster executes the program code of the software system to execute the task scheduling method of the present application. The scheduling system can also be a hardware system. When the hardware system runs, it executes the task scheduling method of the present application. Among them, the tasks scheduled by the scheduling system can include but are not limited to AIGC tasks. Considering that AIGC tasks require a large amount of computing power to execute, the scheduling system can rely on cloud services, or the scheduling system can be provided to users in the form of cloud services. Among them, the scheduling system can include a scheduling cluster and an inference cluster. The scheduling cluster can be implemented through a cloud server or a computing cluster. For example, the scheduling cluster can be implemented through the Elastic Cloud Server (ECS) of a cloud provider, such as through the elastic cloud server of a public cloud or a hybrid cloud. The inference cluster can be implemented through an AI development platform service.

[0083] Specifically, the scheduling cluster includes at least one agent server, and the inference cluster includes multiple inference servers. Among them, the inference server is also called an inference server. The inference server establishes a full-duplex connection based on the Transmission Control Protocol (TCP) with at least one agent server. The scheduling cluster can receive task requests based on a two-way communication protocol. Then, the scheduling cluster determines a target inference server that matches the task request from the inference cluster according to the task request, the status of at least one inference server in the inference cluster, and the performance model of at least one inference server. Next, the scheduling cluster schedules the task request to the target inference server through the full-duplex connection with the target inference server.

[0084] In this method, the scheduling cluster can perceive the status of the inference servers in the inference cluster. Based on this status, it can match a suitable inference server for the task request and schedule the task request to the appropriate inference server for processing. In this way, the response speed of the request can be improved, and the response delay can be shortened as much as possible. Moreover, different from the scheduling scheme based on HTTP, this method can return the result to the user through a full-duplex connection without using the polling method to query the result, reducing the high latency and high network overhead caused by using the polling method to query the result in the HTTP communication mode.

[0085] To make the technical solution of this application clearer and easier to understand, the system architecture of this application will be introduced first below.

[0086] See Figure 2 As shown in the schematic diagram of the architecture of a scheduling system, the scheduling system 10 includes a scheduling cluster 100 and an inference cluster 200. Among them, the scheduling cluster 100 and the inference cluster 200 can be distributed in different VPC networks. The scheduling cluster 100 includes at least one agent server, and the agent server implements the function of the scheduler. The inference cluster 200 includes multiple inference servers. The inference server establishes a full-duplex connection with at least one agent server.

[0087] The full-duplex connection can be a connection based on full-duplex communication. Full-duplex communication means that in communication transmission, it is allowed to transmit data in both directions simultaneously. Further, the full-duplex connection can be a long connection, such as a full-duplex long connection. Among them, a long connection means that after the connection is established, the connection is maintained even if there is no data transmission in the channel until the disconnection is not actively triggered.

[0088] In some possible implementation manners, the full-duplex connection can be a connection based on WebSocket (WebSocket). WebSocket is a TCP-based connection, which is in the application layer of the Open Systems Interconnection (OSI) and is a network transmission protocol that realizes full-duplex communication.

[0089] What the scheduling system 10 presents externally is the entrance of the scheduling cluster 100. Specifically, the scheduling cluster 100 is used to obtain the status of at least one inference server in the inference cluster, for example, continuously obtain the status of at least one inference server, receive task requests sent by the client (requesting party) based on a two-way communication protocol, for example, a task request based on WebSocket, and determine a target inference server that matches the task request from the inference cluster according to the task request, the status of at least one inference server in the inference cluster 200, and the performance model of at least one inference server, and schedule the task request to the target inference server through the full-duplex connection with the target inference server. It should be noted that the WebSocket communication protocol can be used throughout the entire link of the scheduling system 10, and polling is not involved throughout the process. Based on this, the agent server can also be called a WebSocket server, and the client can also be called a WebSocket client.

[0090] For the scheduling cluster 100 and the inference cluster 200 deployed in different VPCs respectively, a full-duplex connection can be established at low cost through the VPC Endpoints (VPC-EP) service. The VPC endpoint can privately connect the VPC to the endpoint service (cloud service, user private service), enabling the cloud resources in the VPC to access the endpoint service without an elastic public IP, improving the access efficiency and providing a more flexible and secure networking method. As Figure 3 shown, the scheduling cluster 100 and the inference cluster 200 may also include a proxy component, such as nginx. The connection establishment request sent by the inference cluster 200 can be transmitted to the VPC-EP of the scheduling cluster 100 through the VPC-EP of the inference cluster 200. The proxy component can extract the address of the proxy server in the connection establishment request, such as the small network IP of the target node, and route (or forward) the connection establishment request to the corresponding proxy server based on this address to accurately establish a connection.

[0091] The following will describe the connection establishment process in detail with reference to the accompanying drawings.

[0092] Refer to Figure 4 a schematic diagram of a connection establishment and inference process as shown. The scheduling cluster 100 can register the address of the scheduling cluster 100 with the registration center. Among them, the address registered by the scheduling cluster 100 can be the small network IP of the proxy server in the scheduling cluster 100, such as Figure 4 shown in ① in the figure. Then, the inference cluster 200 can obtain the address of the scheduling cluster 100 from the registration center, such as Figure 4 shown in ② in the figure. Next, as Figure 4 shown in ③ in the figure, the inference cluster 200 sends a connection establishment request to the scheduling cluster 100. This connection establishment request is used to establish a full-duplex connection.

[0093] The full-duplex connection is initiated by the inference service in the inference cluster 200, and nginx's header / path is used for screening to accurately connect to the machines in the scheduling cluster 100. In this way, only a pair of VPC-EP services are required to connect the two clusters, reducing the network connection cost between the clusters and eliminating the security risks brought by peer connections.

[0094] When initiating the connection establishment, the inference cluster 200 can construct a WebSocket full-connection two-way communication structure based on the WebSocket protocol. Moreover, an inference service can include multiple connection channels, each connecting to a different proxy server, which can increase the reconnection mechanism and improve fault tolerance and reliability.

[0095] In some possible implementation manners, the inference cluster 200 may further send the status of the inference server to the registration center, such as the status of the computing cards in the inference server. The computing cards may include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), and a neural network processing unit (NPU). As Figure 4 shown in ④, the inference cluster 200 may further register the inference server with the registration center, so that when subsequent task scheduling is performed, the target inference server can be matched according to the inference server registered in the registration center.

[0096] Specifically, the scheduling cluster 100 may obtain the status of the inference server from the registration center, and determine the idle inference server based on the status of the inference server. As Figure 4 shown in ⑤, the scheduling cluster 100 may determine the target inference server from the idle inference servers, and send a task request to the target inference server, such as a task request for an inference task. The target inference server executes the inference task according to the task request.

[0097] Based on the registration center, the address registration of a proxy server (such as a WebSocket Server) can be implemented to support dynamic adaptive expansion of the connection channel after the cluster is expanded. Moreover, based on the registration center, the status reporting and discovery of the inference server (such as the GPU in the inference server) can be performed, and the availability of the inference server connection channel can be updated in a timely manner.

[0098] In some possible implementation manners, a connection manager may also be introduced on the proxy server side to complete the connection pool management of the client and the inference service. Specifically, the proxy server in the scheduling cluster 100 creates a client connection pool and updates the client connection pool according to the connection information with the client. Alternatively, the proxy server in the scheduling cluster 100 creates an inference service connection pool and updates the inference service connection pool according to the connection information with the inference server. When the scheduling cluster 100 and the inference cluster 200 are expanded, Rebalance does not occur frequently, avoiding the overall unavailability of the service caused by rebalance.

[0099] The scheduling system 10 realizes efficient probing of the connection channel based on the full-duplex, long-connection multi-cluster node automatic discovery, interconnection, and connection maintenance mechanism, ensures the high reliability of the full connection, and guarantees the real-time update of the connection status through the registration center. Moreover, full-duplex message passing can greatly reduce the polling idle time.

[0100] Based on the aforementioned scheduling system 10, the present application also provides a task scheduling method. The following will introduce the specific implementation of the task scheduling method of the present application in conjunction with the accompanying drawings.

[0101] See Figure 5 The flowchart of a task scheduling method shown. This method can be applied to the scheduling system 10, and the scheduling system 10 includes a scheduling cluster 100 and an inference cluster 200. The scheduling cluster 100 includes at least one proxy server, and the inference cluster 200 includes multiple inference servers. The inference servers establish full-duplex connections based on TCP with at least one proxy server. This method includes the following steps:

[0102] S501. The scheduling cluster 100 obtains the status of at least one inference server in the inference cluster 200.

[0103] The status of the inference server is used to characterize the load condition of the inference server. For example, the status of the inference server may include idle, overloaded, or moderately loaded. Specifically, the inference servers in the inference cluster 200, such as GPUs, can be registered in the registration center. Correspondingly, the registration center can store the status of the inference servers. The scheduling cluster 100 can discover at least one inference server from the registration center, such as an inference server that has been successfully registered in the registration center, and obtain the status of at least one inference server.

[0104] Considering that the status of the inference server can be dynamically changed, the scheduling cluster 100 can continuously obtain the status of at least one inference server in the inference cluster 200. For example, the scheduling cluster 100 can regularly obtain the status of at least one inference server in the inference cluster 200 from the registration center according to a set period. This can provide a reference for subsequent task scheduling.

[0105] It should be noted that obtaining the status of the inference server from the registration center is only a specific implementation of the present application. In actual applications, the scheduling cluster 100 can also obtain the status of the inference server through other means, and the present application does not limit this.

[0106] S502. The scheduling cluster 100 receives a task request based on a bidirectional communication protocol.

[0107] The bidirectional communication protocol refers to a protocol that enables bidirectional communication. For example, it can be the WebSocket protocol. The client can generate a task request based on the WebSocket protocol, and the scheduling cluster 100 can receive the task request based on the WebSocket protocol sent by at least one client. Among them, the task in the task request can be an inference task executed by the inference server. According to different application scenarios, the inference task can include tasks with different functions. For example, the inference task can be an AIGC task.

[0108] S504. According to the task request, the status of at least one inference server in the inference cluster 200, and the performance model of at least one inference server, the scheduling cluster 100 determines a target inference server that matches the task request from the inference cluster 200.

[0109] The status of the inference server is used to characterize the load condition of the inference server. For example, the status of the inference server may include idle, overloaded, or moderately loaded. The performance model of the inference server may include at least one of an inference ability model or a robustness model. Among them, the inference ability model is used to characterize the inference ability of the inference server, such as inference accuracy and inference time. The robustness model is used to characterize the robustness of the inference server.

[0110] Specifically, the scheduling cluster 100 can collect the inference results or online time of the inference servers in the inference cluster 200 in a historical time period. Among them, the online time can be characterized by the online time and offline time of the inference server. The scheduling cluster 100 can model based on the inference results of the inference servers in the inference cluster 200 in a historical time period to obtain an inference ability model. The scheduling cluster 100 can also model based on the online time of the inference server to obtain a robustness model. Among them, the scheduling cluster 100 can use statistical analysis to model the above data to obtain an inference ability model and a robustness model.

[0111] The scheduling cluster 100 can sort the inference servers in the resource pool according to the performance model constructed from the historical data of the inference server, such as an inference ability model or a robustness model, in combination with the status of the inference server (such as the status indicating whether the inference server is currently available), and return a target inference server that matches the task request. Among them, the scheduling cluster 100 can return the optimal N target inference servers.

[0112] Among them, the scheduling cluster 100 can first determine a candidate server set according to the status of the inference server. The candidate server set includes idle inference servers, or the candidate server set includes inference servers with loads sorted from small to the top n or with loads less than a set value. Then the scheduling cluster 100 can select inference servers with performance meeting the conditions from the candidate server set and determine them as target inference servers that match the task request. Specifically, the scheduling cluster 100 can evaluate or sort the inference servers according to the inference ability model and / or robustness model of the inference server, and determine a target inference server that matches the task request according to the evaluation or sorting results.

[0113] In some possible implementation manners, the scheduling cluster 100 can also support adaptive traffic control (flow control) based on the load of the scheduling system 10. Such as Figure 6As shown, the scheduling cluster 100 can model based on the historical traffic data of the scheduling system 10 and the resource utilization rate of the resource pool in the inference cluster 200 (reflecting the node load status), such as GPU utilization rate, to obtain a traffic load model. When the scheduling cluster 100 receives a task request, it can determine whether the task request can be processed based on the traffic load model. The scheduling cluster 100 can perform flow control on task requests that exceed its processing capacity. For example, the scheduling cluster 100 can allow the first proportion of task requests to pass through for task requests that exceed its processing capacity, and reject the remaining task requests. Among them, the first proportion can be set according to experience. In some examples, the first proportion can be set to 50%.

[0114] The scheduling cluster 100 also supports preprocessing and optimization of task requests. Specifically, the scheduling cluster 100 can obtain the trigger time of the task request. Similar requests that are continuously triggered within a short period can be regarded as invalid requests, and the scheduling cluster 100 can eliminate the invalid requests according to the trigger time. Among them, the scheduling cluster 100 can retain one of the task requests. For example, it retains the task request that is triggered first and eliminates other similar requests that are continuously triggered within a short period. In this way, it can achieve filtering of invalid requests at the user dimension. In addition to eliminating invalid requests, the scheduling cluster 100 can also perform intelligent duplicate checking on task requests. Specifically, the scheduling cluster 100 can query the cache results for task requests that are similar to historical requests.

[0115] In some possible implementation manners, the scheduling system 10 also maintains a vector database. The vector database is used to convert text into vectors and then store them. The vector database is mainly used to index relevant vectors before the user sends a prompt, so as to improve the quality of the prompt. Among them, the prompt is the request text input to the model (such as a large model with a large number of parameters), which guides the model to give a result that meets expectations.

[0116] The scheduling cluster 100 can also retrieve similar corpora of the input corpus in the task request from the vector database, and then obtain a prompt based on the input corpus and the similar corpora. This prompt is used to input into the model in the inference server, such as a generative model, for inference. Among them, the scheduling cluster 100 can splice the input corpus and the similar corpora, and the similar corpora serve as a supplement to the input corpus. In this way, the accuracy of the prompt can be improved, and by matching the target inference server with this prompt, the quality of the inference result can be improved. It should be noted that the above capabilities can be selected by the user whether to enable or disable. For example, if the user selects to turn on the switch (control) of "enable intelligent corpus", the scheduling cluster 100 can retrieve similar corpora of the input corpus in the task request from the vector database and enhance the prompt according to the input corpus and the similar corpora.

[0117] Among them, based on historical request data modeling, intercepting abnormal traffic can ensure the stable operation of the scheduling system 10, pre-processing and optimizing requests, such as eliminating similar requests and checking cached results, so as to avoid the waste of resources caused by repeated calculation of invalid requests and similar requests. Vectorizing the input corpus in the task request and querying the vector database to obtain similar corpus to enhance the prompt can improve the accuracy of the request.

[0118] S506 . The scheduling cluster 100 schedules the task request to the target inference server through a full-duplex connection with the target inference server.

[0119] Specifically, the scheduling cluster 100 may determine a full-duplex connection with a target inference server through an inference service resource pool, and schedule the task request to the target inference server through the full-duplex connection with the target inference server.

[0120] Based on the above description, it can be seen that the task scheduling method of the present application can perceive the status of the inference server in the inference cluster 200, and based on the status, it can match the appropriate inference server for the task request, and schedule the task request to the appropriate inference server for processing, thereby improving the response speed of the request and shortening the response delay as much as possible.

[0121] In some possible implementations, the present application also provides a request response mode that prioritizes response speed, such as a request response mode that prioritizes response speed based on a fail-fast mechanism. Fail-fast refers to a design method that fails quickly when an abnormal situation occurs during system operation, rather than trying to retry or wait.

[0122] Specifically, the target inference server (such as a GPU) that the scheduling system 10 matches for the task request based on the intelligent task scheduling strategy can usually be the optimal processing node with the shortest expected processing time. Taking into account the situations of high concurrency and resource preemption, this application allows the inference server (such as a GPU) to return to a busy state and rematch. Furthermore, when the scheduling system 10 is overloaded, the task request is thrown back N times (user-defined), and the scheduling cluster 100 can quickly return to the system busy state to avoid the user waiting time timeout, or continuous retries while waiting to increase the system burden.

[0123] The above-mentioned request response mode with priority on response speed is described in detail below with reference to the accompanying drawings.

[0124] See also Figure 7The flowchart of a task scheduling method is shown. In a normal scheduling scenario, after the client establishes a connection with the scheduling cluster 100, the scheduling cluster 100 can process the task request and match the target inference server for the task request according to the status of the inference server, the inference ability model of the inference server, and the robustness model. For the sake of description, in this application, the target inference server includes the first inference server (such as Figure 7 the inference server 1 in

[0125] is taken as an example). Then, the scheduling cluster 100 schedules the task request to the first inference server through a full-duplex connection with the first inference server to invoke the inference service of the first inference server for inference. The inference service can return the inference result to the scheduling cluster 100, and the scheduling cluster 100 can return the inference result to the client, thereby realizing the return of the inference result to the user.

[0126] Specifically, when the user enables token-by-token inference, the scheduling cluster 100 can receive multiple inference results of the target inference server through the long connection with the target inference server. The multiple inference results are the inference results of token-by-token inference, and the scheduling cluster 100 can return the multiple inference results in sequence.

[0127] In an extremely busy scenario, the first inference server matched by the scheduling cluster 100 for the task request may be unable to process the task request. Based on this, the scheduling cluster 100 can also accept the feedback from the first inference server, and this feedback is used to indicate that the first inference server is in a busy state. Correspondingly, the scheduling cluster 100 can also determine the second inference server (such as the inference server 2) that matches the task request from the inference cluster according to the task request, the status of the inference servers in the inference cluster, and the performance model of the inference servers.

[0128] When the maximum number of matches is configured in the scheduling system 10, for example, it is N times, the scheduling cluster 100 can repeat the above process N - 1 times. When the number of times the scheduling cluster 100 determines the target inference server that matches the task request reaches the maximum number of matches, and the target inference server returned each time returns feedback indicating a busy state, the scheduling cluster 100 can return a quick failure event to the user.

[0129] To make the technical solution of this application clearer and easier to understand, the task scheduling method of this application will be introduced below in combination with a specific application scenario.

[0130] See Figure 8 The schematic diagram of the application scenario of a task scheduling method shown in the figure. In this scenario, three Agent Servers can form a scheduling cluster 100, and three Inference Servers (including GPUs) can form an inference cluster 200. Among them, the scheduling cluster 100 can register the address of the scheduling cluster 100 with the registration center. For example, it registers the small network IP of the Agent Server in the scheduling cluster 100. Then, the inference cluster 200 can obtain the address of the scheduling cluster 100 from the registration center. Next, the inference cluster 200 sends a connection establishment request to the scheduling cluster 100. This connection establishment request is used to establish a full-duplex connection, such as a WebSocket connection.

[0131] The entire scheduling system 10 presents the entrance of the scheduling cluster 100 to the outside. As Figure 8 shown, the client can send a task request to the scheduling cluster 100 in the way of the WebSocket protocol. Since the WebSocket protocol is adopted, this task request is also called a WebSocket request.

[0132] The Agent Server in the scheduling cluster 100 can preprocess and optimize the task request. Specifically, the AgentServer can query the request identifier request_id of the task request from the cache server cluster. If the request identifier exists in the cache service cluster, the Agent Server can directly obtain the inference result from the cache service cluster and return the inference result to the client. If the request identifier does not exist in the cache service cluster, the Agent Server can store the task request in the database, for example, record the task request in MySql. Among them, the Agent Server can also enhance the prompt for the task request.

[0133] The Agent Server can match an available Inference service for the task request for inference. As Figure 8As shown, the agent can obtain available links from the GPU connection channel for message forwarding, so as to forward the task request to the available Inference service. The Inference service performs inference based on the task request and obtains the inference result. The inference result is then returned to the Agent Server via the WebSocket protocol, and the Agent Server actively pushes it to the user. The communication protocol used throughout the entire link of this scheduling system 10 is the WebSocket communication protocol, without involving polling for inference results throughout the process, and actively pushing them to the user after the inference results are generated. In this way, it can solve the problems of high latency and high network overhead caused by using the polling method to query inference results in the HTTP communication method.

[0134] Furthermore, after the user enables the switch for word-by-word inference, the inference results from the Inference Server can also be pushed to the client multiple times during the inference process until the inference is completed. Among them, the inference results can be sent via the WebSocket protocol without repeatedly sending the header files, solving the problems of network overhead (repeatedly sending header files) and slow response in the HTTP communication method.

[0135] Based on the foregoing task scheduling method, this application also provides a scheduling system. As Figure 2 shown, this scheduling system 10 includes a scheduling cluster 100 and an inference cluster 200. The scheduling cluster 100 includes at least one proxy server, and the inference cluster 200 includes multiple inference servers. The inference servers establish full-duplex connections based on the Transmission Control Protocol (TCP) with at least one proxy server;

[0136] The scheduling cluster 100 is used to obtain the status of at least one inference server in the inference cluster, receive task requests based on the two-way communication protocol, and determine a target inference server that matches the task request from the inference cluster according to the task request, the status of the at least one inference server in the inference cluster, and the performance model of the at least one inference server;

[0137] The scheduling cluster 100 is also used to schedule the task request to the target inference server through the full-duplex connection with the target inference server.

[0138] As Figure 9As shown in the figure, the scheduling cluster 100 may include or deploy the following modules: an interaction module 902, a matching module 904, and a scheduling module 906. The interaction module 902 is used to obtain the status of at least one inference server in the inference cluster and receive task requests based on a bidirectional communication protocol. The matching module 904 is used to determine a target inference server that matches the task request from the inference cluster according to the task request, the status of the at least one inference server in the inference cluster, and the performance model of the at least one inference server. The scheduling module 906 is used to schedule the task request to the target inference server through a full-duplex connection with the target inference server.

[0139] The above interaction module 902, matching module 904, and scheduling module 906 may be implemented by hardware or may be implemented by software.

[0140] When implemented by software, the interaction module 902, matching module 904, and scheduling module 906 may be application programs running on a computing device, such as a computing engine, etc. Among them, the application program may also be virtualized through a virtualization service for users to use. The virtualization service may include virtual machine (VM) service, bare metal server (BMS) service, and container service. Among them, the VM service may be a service that virtualizes a virtual machine (VM) resource pool on multiple physical hosts to provide VMs for users on demand. The BMS service is a service that virtualizes a BMS resource pool on multiple physical hosts to provide BMS for users on demand. The container service is a service that virtualizes a container resource pool on multiple physical hosts to provide containers for users on demand. A VM is a simulated virtual computer, that is, a computer logically. A BMS is an elastic and scalable high-performance computing service, and its computing performance is no different from that of a traditional physical machine, with the characteristics of secure physical isolation. A container is a kernel virtualization technology that can provide lightweight virtualization to achieve the purpose of isolating user space, processes, and resources. It should be understood that the VM service, BMS service, and container service in the above virtualization service are only used as specific examples. In actual applications, the virtualization service may also be other lightweight or heavyweight virtualization services, which are not specifically limited here.

[0141] When implemented by hardware, the interaction module 902, the matching module 904, and the scheduling module 906 may include at least one computing device, such as a server. Alternatively, the interaction module 902, the matching module 904, and the scheduling module 906 may also be devices implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). Among them, the above PLD may be implemented by a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0142] In some possible implementation manners, the performance model includes at least one of an inference capability model or a robustness model, and the scheduling cluster 100 is further configured to:

[0143] Collect the inference results or the online time of the inference servers in the inference cluster in a historical time period;

[0144] Model according to the inference results of the inference servers in the inference cluster in a historical time period to obtain the inference capability model, and / or model according to the online time of the inference servers to obtain the robustness model.

[0145] The scheduling cluster 100 may further include or deploy the following module: a modeling module 908. The modeling module 908 is configured to collect the inference results or the online time of the inference servers in the inference cluster in a historical time period, model according to the inference results of the inference servers in the inference cluster in a historical time period to obtain the inference capability model, or model according to the online time of the inference servers to obtain the robustness model.

[0146] Similar to the interaction module 902, the matching module 904, and the scheduling module 906, the modeling module 908 may be implemented by software or by hardware. When implemented by software, the modeling module 908 may be an application program running on a computing device. When implemented by hardware, the modeling module 908 may include at least one computing device, such as a server. Alternatively, the modeling module 908 may also be a device implemented using an ASIC or a PLD.

[0147] In some possible implementation manners, the target inference server includes a first inference server, and the scheduling cluster 100 is further configured to:

[0148] Receive the feedback from the first inference server, where the feedback is used to indicate that the first inference server is in a busy state;

[0149] Determine a second inference server that matches the task request from the inference cluster according to the task request, the status of the at least one inference server in the inference cluster, and the performance model of the at least one inference server;

[0150] Schedule the task request to the second inference server through a full-duplex connection with the second inference server.

[0151] In some possible implementation manners, the scheduling cluster 100 is further configured to:

[0152] If the number of times of determining the target inference server that matches the task request reaches the maximum number of matches, and each time the target inference server that is matched returns feedback indicating a busy state, then return a quick failure event to the user.

[0153] In some possible implementation manners, the scheduling cluster 100 is further configured to:

[0154] Retrieve similar corpora of the input corpus in the task request from the vector database;

[0155] Obtain prompt information according to the input corpus and the similar corpora, where the prompt information is used for inference by a model in the inference server.

[0156] The scheduling cluster 100 may further include or deploy the following module: a preprocessing module 909. The preprocessing module 909 is configured to retrieve similar corpora of the input corpus in the task request from the vector database, and obtain prompt information according to the input corpus and the similar corpora. The preprocessing module 909 may be implemented by software or by hardware. When implemented by software, the preprocessing module 909 may be an application program running on a computing device. When implemented by hardware, the preprocessing module 909 may include at least one computing device, such as a server, etc. Alternatively, the preprocessing module 909 may also be a device implemented using ASIC or PLD, etc.

[0157] In some possible implementation manners, the full-duplex connection is a long connection, and the scheduling cluster 100 is further configured to:

[0158] When the user enables word-by-word inference, receive multiple inference results of the target inference server through the long connection with the target inference server, where the multiple inference results are inference results of word-by-word inference, and return the multiple inference results in sequence.

[0159] In some possible implementation manners, the scheduling cluster 100 is further configured to:

[0160] Register the address of the scheduling cluster with the registration center;

[0161] The inference cluster 200 is used for:

[0162] Obtain the address of the scheduling cluster from the registration center, and send a connection establishment request to the scheduling cluster, where the connection establishment request is used to establish a full-duplex connection.

[0163] The scheduling cluster 100 may further include or deploy the following module: a registration module 907. The registration module 907 is used to register the address of the scheduling cluster with the registration center. The registration module 907 may be implemented by software or by hardware. When implemented by software, the registration module 907 may be an application running on a computing device. When implemented by hardware, the registration module 907 may include at least one computing device, such as a server, etc. Alternatively, the registration module 907 may also be a device implemented using ASIC or PLD, etc.

[0164] In some possible implementation manners, the full-duplex connection is a connection based on a network socket, the connection establishment request is a request based on a network socket, the connection establishment request includes the address of a proxy server in the scheduling cluster, and the inference cluster is specifically used for:

[0165] Route the connection establishment request to the proxy server identified by the address in the connection establishment request through a proxy component.

[0166] In some possible implementation manners, the inference cluster 200 is further used for:

[0167] Send the status of the inference server to the registration center;

[0168] The scheduling cluster 100 is further used for:

[0169] Obtain the status of the inference server from the registration center.

[0170] In some possible implementation manners, the scheduling cluster 100 is further used for:

[0171] Create a client connection pool and update the client connection pool according to the connection information with the client; or,

[0172] Create an inference service connection pool and update the inference service connection pool according to the connection information with the inference server.

[0173] The scheduling cluster 100 may further include or deploy the following modules: a creation module 905. The creation module 905 is used to create a client connection pool and update the client connection pool according to the connection information with the client, or create an inference service connection pool and update the inference service connection pool according to the connection information with the inference server. The creation module 905 may be implemented by software or by hardware. When implemented by software, the creation module 905 may be an application running on a computing device. When implemented by hardware, the creation module 905 may include at least one computing device, such as a server, etc. Alternatively, the creation module 905 may also be a device implemented using ASIC or PLD, etc.

[0174] This application also provides a computing device 1000. As Figure 10 shown, the computing device 1000 includes: a bus 1002, a processor 1004, a memory 1006, and a communication interface 1008. The processor 1004, the memory 1006, and the communication interface 1008 communicate with each other through the bus 1002. The computing device 1000 may be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 1000.

[0175] The bus 1002 may be a peripheral component interconnect (PCI) bus, an extended industry standard architecture (EISA) bus, or the like. The bus may be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, Figure 10 only one line is shown here, but it does not mean that there is only one bus or one type of bus. The bus 1002 may include a path for transmitting information between various components of the computing device 1000 (for example, the memory 1006, the processor 1004, the communication interface 1008).

[0176] The processor 1004 may include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP), etc.

[0177] The memory 1006 may include a volatile memory, such as a random access memory (RAM). The memory 1006 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD). Executable program codes are stored in the memory 1006, and the processor 1004 executes the executable program codes to implement the foregoing task scheduling method. Specifically, instructions for the scheduling system 10 to execute the task scheduling method are stored on the memory 1006. Among them, the memory 1006 may store instructions for implementing the functions of the interaction module 902, the matching module 904, and the scheduling module 906. Further, the memory 1006 may also store instructions for implementing the functions of the modeling module 908, the preprocessing module 909, the registration module 907, and the creation module 905. The communication interface 1008 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement the communication between the computing device 1000 and other devices or communication networks.

[0178] The embodiment of the present application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device may be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device may also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.

[0179] As Figure 11 shown, the computing device cluster includes at least one computing device 1000. Instructions for the same scheduling system 10 to execute the task scheduling method may be stored in the memory 1006 of one or more of the computing devices 1000 in the computing device cluster.

[0180] In some possible implementation manners, one or more of the computing devices 1000 in the computing device cluster may also be used to execute some of the instructions of the scheduling system 10 for executing the task scheduling method. In other words, a combination of one or more computing devices 1000 may jointly execute the instructions of the scheduling system 10 for executing the task scheduling method.

[0181] It should be noted that the memories 1006 in different computing devices 1000 in the computing device cluster may store different instructions for executing some functions of the scheduling system 10.

[0182] Figure 12 shows a possible implementation manner. AsFigure 12 As shown, two computing devices 1000A and 1000B are connected through a communication interface 1008. Instructions for executing the functions of the interaction module 902 and the matching module 904 in the scheduling cluster 100 are stored in the memory of the computing device 1000A. Instructions for executing the function of the scheduling module 906 are stored in the memory of the computing device 1000B. In other words, the memories 1006 of the computing devices 1000A and 1000B jointly store instructions for the scheduling system 10 to execute the task scheduling method, such as instructions for the scheduling cluster to execute the task scheduling method. Further, the memory 1006 of the computing device 1000A may also store instructions for the functions of the modeling module 908 and the preprocessing module 909, and the memory of the computing device 1000B may also store instructions for the functions of the registration module 907 and the creation module 905.

[0183] Figure 12 The connection method between the computing device clusters shown may be considered because the task scheduling method provided in this application requires more resources for task scheduling decisions. Therefore, it is considered to hand over the function implemented by the scheduling module 906 to the computing device 1000B for execution.

[0184] It should be understood that Figure 12 the functions of the computing device 1000A shown in can also be completed by multiple computing devices 1000. Similarly, the functions of the computing device 1000B can also be completed by multiple computing devices 1000.

[0185] In some possible implementation manners, one or more computing devices in the computing device cluster may be connected through a network. Among them, the network may be a wide area network or a local area network, etc. Figure 13 shows a possible implementation manner. As Figure 13 shown, two computing devices 1000C and 1000D are connected through a network. Specifically, they are connected to the network through the communication interfaces in each computing device. In this type of possible implementation manner, instructions for executing the functions of the interaction module 902 and the matching module 904 are stored in the memory 1006 of the computing device 1000C. At the same time, instructions for executing the function of the scheduling module 906 are stored in the memory 1006 of the computing device 1000D.

[0186] Figure 13 The connection method between the computing device clusters shown may be considered because the task scheduling method provided in this application requires more resources for task scheduling decisions. Therefore, it is considered to hand over the function implemented by the scheduling module 906 to the computing device 1000D for execution.

[0187] It should be understood that Figure 13The functions of the computing device 1000C shown can also be completed by multiple computing devices 1000. Similarly, the functions of the computing device 1000D can also be completed by multiple computing devices 1000.

[0188] An embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive), etc. The computer-readable storage medium includes instructions that direct the computing device to execute the above-mentioned task scheduling method applied to the scheduling system 10.

[0189] An embodiment of the present application also provides a computer program product containing instructions. The computer program product can be software or a program product that contains instructions and can run on a computing device or be stored in any available medium. When the computer program product runs on at least one computer device, it causes at least one computer device to execute the above-mentioned task scheduling method.

[0190] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.

Claims

1. A task scheduling method, characterized in that: Applied to a scheduling system, the scheduling system includes a scheduling cluster and an inference cluster, the scheduling cluster includes at least one proxy server, the inference cluster includes multiple inference servers, the inference server and the at least one proxy server establish a full-duplex connection based on the transmission control protocol TCP, the method includes: The scheduling cluster obtains the status of at least one inference server in the inference cluster; The scheduling cluster receives a task request based on a two-way communication protocol; The scheduling cluster determines, according to the task request, the state of at least one inference server in the inference cluster, and the performance model of the at least one inference server, a target inference server matching the task request from the inference cluster; The scheduling cluster schedules the task request to the target inference server through a full-duplex connection with the target inference server.

2. The method according to claim 1, characterized in that The performance model includes at least one of a reasoning capability model or a robustness model, and the method further includes: The scheduling cluster collects the inference results or the online time of the inference servers in the inference cluster in the historical time period; The scheduling cluster models the reasoning capability model according to the reasoning results of the reasoning servers in the reasoning cluster in the historical time period, and / or the scheduling cluster models the robustness model according to the online time of the reasoning servers.

3. The method according to claim 1 or 2, characterized in that: The target inference server includes a first inference server, and the method further includes: The scheduling cluster receives feedback from the first inference server, where the feedback is used to indicate that the first inference server is in a busy state; The scheduling cluster determines, according to the task request, a state of the at least one inference server in the inference cluster, and a performance model of the at least one inference server, a second inference server matching the task request from the inference cluster; The scheduling cluster schedules the task request to the second inference server through a full-duplex connection with the second inference server.

4. The method according to claim 3, characterized in that The method further comprises: The scheduling cluster determines that the number of target inference servers that match the task request reaches a maximum number of matches, and each matched target inference server returns feedback indicating that it is busy, then a quick failure event is returned to the user.

5. The method according to any one of claims 1 to 4, characterized in that: The method further comprises: The scheduling cluster retrieves similar corpus to the input corpus in the task request from a vector database; The scheduling cluster obtains prompt information according to the input corpus and the similar corpus, and the prompt information is used to input into a model in an inference server for inference.

6. The method according to any one of claims 1 to 5, characterized in that: The full-duplex connection is a long connection, and the method further includes: When the user turns on word-by-word reasoning, the scheduling cluster receives multiple reasoning results from the target reasoning server through a long connection with the target reasoning server, and the multiple reasoning results are the reasoning results of word-by-word reasoning, and returns the multiple reasoning results in sequence.

7. The method according to any one of claims 1 to 6, characterized in that: The method further comprises: The scheduling cluster registers the address of the scheduling cluster with the registration center; The inference cluster obtains the address of the scheduling cluster from the registration center, and sends a connection establishment request to the scheduling cluster, where the connection establishment request is used to establish a full-duplex connection.

8. The method according to claim 7, characterized in that The full-duplex connection is a connection based on a network socket, the connection establishment request is a request based on the network socket, the connection establishment request includes an address of a proxy server in the scheduling cluster, and the reasoning cluster sends the connection establishment request to the scheduling cluster, including: The inference cluster routes the connection establishment request to a proxy server identified by an address in the connection establishment request through a proxy component.

9. The method according to any one of claims 1 to 8, characterized in that: The method further comprises: The inference cluster sends the status of the at least one inference server to the registration center; The scheduling cluster obtains the state of at least one inference server in the inference cluster, including: The scheduling cluster obtains the status of the at least one inference server from the registration center.

10. The method according to any one of claims 1 to 9, characterized in that: The method further comprises: The proxy server in the scheduling cluster creates a client connection pool and updates the client connection pool according to the connection information with the client; or, The proxy server in the scheduling cluster creates an inference service connection pool, and updates the inference service connection pool according to the connection information with the inference server.

11. A scheduling system, characterized in that: The scheduling system includes a scheduling cluster and an inference cluster, the scheduling cluster includes at least one proxy server, the inference cluster includes multiple inference servers, and a full-duplex connection based on a transmission control protocol TCP is established between the inference server and the at least one proxy server; The scheduling cluster is used to obtain the state of at least one inference server in the inference cluster, receive a task request based on a two-way communication protocol, and determine a target inference server matching the task request from the inference cluster according to the task request and the state of at least one inference server in the inference cluster and a performance model of the at least one inference server; The scheduling cluster is further used to schedule the task request to the target inference server through a full-duplex connection with the target inference server.

12. The system according to claim 11, characterized in that The performance model includes at least one of a reasoning capability model or a robustness model, and the scheduling cluster is further used for: Collecting the inference results or the online time of the inference servers in the inference cluster in a historical time period; The reasoning capability model is obtained by modeling according to the reasoning results of the reasoning servers in the reasoning cluster in the historical time period, and / or the robustness model is obtained by modeling according to the online time of the reasoning servers.

13. The system according to claim 11 or 12, characterized in that The target inference server includes a first inference server, and the scheduling cluster is further used for: receiving feedback from the first inference server, wherein the feedback is used to indicate that the first inference server is in a busy state; Determine, from the inference cluster, a second inference server matching the task request according to the task request, a state of the at least one inference server in the inference cluster, and a performance model of the at least one inference server; The task request is dispatched to the second inference server through a full-duplex connection with the second inference server.

14. The system according to claim 13, characterized in that The scheduling cluster is also used for: If the number of target inference servers that match the task request reaches a maximum number of matches, and each matched target inference server returns feedback indicating that it is busy, a quick failure event is returned to the user.

15. The system according to any one of claims 11 to 14, characterized in that: The scheduling cluster is also used for: Retrieving similar corpus to the input corpus in the task request from a vector database; Prompt information is obtained according to the input corpus and the similar corpus, and the prompt information is used to input into a model in an inference server for inference.

16. The system according to any one of claims 11 to 15, characterized in that The full-duplex connection is a long connection, and the scheduling cluster is also used for: When the user turns on word-by-word reasoning, multiple reasoning results of the target reasoning server are received through a long connection with the target reasoning server, and the multiple reasoning results are the reasoning results of word-by-word reasoning, and the multiple reasoning results are returned in sequence.

17. The system according to any one of claims 11 to 16, characterized in that: The scheduling cluster is also used for: Registering the address of the scheduling cluster with a registration center; The inference cluster is used to: The address of the scheduling cluster is obtained from the registration center, and a connection establishment request is sent to the scheduling cluster, where the connection establishment request is used to establish a full-duplex connection.

18. The system according to claim 17, characterized in that The full-duplex connection is a connection based on a network socket, the connection establishment request is a request based on the network socket, the connection establishment request includes an address of a proxy server in the scheduling cluster, and the inference cluster is specifically used for: The connection establishment request is routed through a proxy component to a proxy server identified by an address in the connection establishment request.

19. The system according to any one of claims 11 to 18, characterized in that The inference cluster is also used to: sending a status of the at least one inference server to a registration center; The scheduling cluster is specifically used for: A status of the at least one inference server is obtained from the registration center.

20. The method according to any one of claims 11 to 19, characterized in that The scheduling cluster is also used for: Creating a client connection pool and updating the client connection pool according to the connection information with the client; or, An inference service connection pool is created, and the inference service connection pool is updated according to the connection information with the inference server.

21. A computing device cluster, characterized in that: The computing device cluster includes at least one computing device, and the at least one computing device includes at least one processor and at least one memory, wherein the at least one memory stores computer-readable instructions; the at least one processor executes the computer-readable instructions so that the computing device cluster executes the task scheduling method as described in any one of claims 1 to 10.

22. A computer-readable storage medium, characterized in that: Comprising computer-readable instructions; the computer-readable instructions are used to implement the task scheduling method described in any one of claims 1 to 10.

23. A computer program product, characterized in that Comprising computer-readable instructions; the computer-readable instructions are used to implement the task scheduling method described in any one of claims 1 to 10.