A task processing method, device and equipment running on a cloud computing platform

By dynamically adjusting the task scheduling strategy on the cloud computing platform, the problem of hardware resource waste in the inference process of large language model is solved, and resource saving and task processing flexibility are improved.

CN119537040BActive Publication Date: 2025-05-23SHENZHEN HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510103643.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-23
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

In the inference process of large language models, a fixed number of full and incremental instances occupy the board for a long time, resulting in wasting hardware resources.

Method used

By dynamically adjusting the task scheduling strategy on the cloud computing platform, and schedule an appropriate number of business container groups to process tasks based on the type of tasks to be inferred and the capability information of the business container group to avoid long-term use of hardware resources.

Benefits of technology

Effectively save hardware resources, while improving task processing flexibility and performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119537040B_ABST
    Figure CN119537040B_ABST
Patent Text Reader

Abstract

The present application discloses a task processing method, device and equipment running on a cloud computing platform, and relates to the field of cloud computing technology. The cloud computing platform can dynamically adjust the scheduling strategy of the task to be inferred according to the task type and the task processing performance of the business container group, which can avoid the hardware resources of the relevant board being occupied for a long time. It saves hardware resources while improving the flexibility of task processing. The method includes: obtaining task information of the task to be inferred and capability information of the business container group; determining a target scheduling strategy from the multiple scheduling strategies according to the task type of the task to be inferred and the capability information of the business container group corresponding to each scheduling strategy in the multiple scheduling strategies, each scheduling strategy is used to indicate the scheduling of the business container group corresponding to each scheduling strategy to process the inference task corresponding to each scheduling strategy; scheduling the business container group corresponding to the target scheduling strategy to process the task to be inferred.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of cloud computing technology, and in particular to a task processing method, device and equipment running on a cloud computing platform. Background Art

[0002] The reasoning process of a large language model (LLM) consists of two stages: prefill and decode. In a separate framework, the prefill and decode stages can be deployed in different boards, respectively. Specifically, the prefill stage can be executed based on a fixed number of full instances in one or more boards, and the decode stage can be executed based on a fixed number of incremental instances in other boards, so as to improve the overall reasoning performance of LLM.

[0003] However, since a fixed number of full instances and incremental instances will occupy the corresponding boards for a long time, when the hardware resources required for the LLM reasoning process are far less than the hardware resources that these boards can provide, it will cause a large waste of hardware resources. Summary of the invention

[0004] The present application provides a task processing method, device and equipment running on a cloud computing platform, which solves the technical problem in the related technology that a fixed number of full instances and incremental instances will occupy the corresponding boards for a long time, thereby causing a large waste of hardware resources.

[0005] In a first aspect, a task processing method running on a cloud computing platform is provided, the method comprising: obtaining task information of a task to be inferred and capability information of a business container group, the task information comprising a task type of the task to be inferred, the business container group being used to be assigned to execute the task to be inferred, and the capability information being used to characterize the task processing performance of the business container group; determining a target scheduling strategy from the multiple scheduling strategies according to the task type and the capability information of the business container group corresponding to each scheduling strategy in a plurality of scheduling strategies, the each scheduling strategy being used to instruct the business container group corresponding to each scheduling strategy to be scheduled to process the inference task corresponding to each scheduling strategy; and scheduling the business container group corresponding to the target scheduling strategy to process the task to be inferred.

[0006] It can be understood that the method described in the first aspect can be executed by a computing device, which can be a terminal, an apparatus including a terminal, or a chip in a terminal, or the computing device can also be a network device, an apparatus including a network device, or a chip in a network device. For ease of understanding, the following description is based on the example of the execution of a computing device.

[0007] In this application, since the target scheduling strategy is determined by the cloud computing platform from multiple scheduling strategies according to the task type of the task to be inferred and the capability information of the business container group corresponding to each scheduling strategy, the cloud computing platform can dynamically adjust the scheduling strategy of the task to be inferred according to the task type and the task processing of the business container group, which can avoid the hardware resources of the relevant board being occupied for a long time. It can save hardware resources and improve the flexibility of task processing.

[0008] In an optional implementation manner, the capability information of the service container group includes the remaining resource size of the service container group and / or the board type of the board to which the service container group belongs.

[0009] In this implementation, for a business container group, if the remaining resource size of the business container group is larger, it means that the task processing performance of the business container group is higher; on the contrary, if the remaining resource size of the business container group is smaller, it means that the task processing performance of the business container group is lower. In addition, different types of boards have different performance levels, so the performance of business container groups included (or deployed) in different types of boards also varies. In this way, the cloud computing platform can accurately and effectively determine the processing performance of the business container group based on the remaining resource size of the business container group and / or the board type of the board to which the business container group belongs, thereby improving the efficiency of determining the target scheduling strategy.

[0010] In an optional implementation, a scheduling strategy includes an instance ratio, which is the ratio between the number of full instances corresponding to the scheduling strategy and the number of incremental instances corresponding to the scheduling strategy. An instance ratio corresponds to a business container group set. Before the above-mentioned scheduling of the business container group corresponding to the target scheduling strategy to process the task to be inferred, the method also includes: determining the business container group included in the target business container group set as the business container group corresponding to the target scheduling strategy. The target business container group set is the business container group set corresponding to the target instance ratio, and the target instance ratio is the instance ratio included in the target scheduling strategy.

[0011] In this implementation, after determining the target scheduling strategy, the cloud computing platform can determine the target business container group set (i.e., the business container group set corresponding to the target instance ratio) based on the target instance ratio included in the target scheduling strategy, specifically, the ratio between the number of full instances corresponding to the target scheduling strategy and the number of incremental instances corresponding to the target scheduling strategy, thereby determining the business container group included in the target business container group set as the business container group corresponding to the target scheduling strategy, and being able to accurately and effectively determine the business container group used to process the task to be inferred, thereby improving the effectiveness of task processing.

[0012] In an optional implementation, different scheduling strategies include different instance ratios.

[0013] In this implementation, since different scheduling strategies include different instance ratios, that is, for multiple scheduling strategies, the ratio between the number of full instances corresponding to one scheduling strategy and the number of incremental instances corresponding to the scheduling strategy is different from the ratio between the number of full instances corresponding to other scheduling strategies and the number of incremental instances corresponding to the other scheduling strategies. At this time, the cloud computing platform can determine a unique business container group for processing the task to be inferred (i.e., the business container group corresponding to the target scheduling strategy) based on the unique instance ratio included in the unique scheduling strategy (i.e., the target instance ratio included in the target scheduling strategy), which can improve the accuracy of task processing.

[0014] In an optional implementation, the target business container group set includes at least one first business container group and at least one second business container group, the at least one first business container group is a full instance corresponding to the target scheduling policy, and the at least one second business container group is an incremental instance corresponding to the target scheduling policy. The above-mentioned scheduling of the business container group corresponding to the target scheduling policy to process the task to be inferred specifically includes: scheduling the at least one first business container group to process the pre-filling stage of the task to be inferred, and scheduling the at least one second business container group to process the decoding stage of the task to be inferred.

[0015] In this implementation, the cloud computing platform schedules at least one business container group to process the pre-filling stage of the task to be inferred, that is, the full instance corresponding to the target scheduling policy executes the pre-filling stage; the cloud computing platform schedules at least one second business container group to process the decoding stage of the task to be inferred, that is, the incremental instance corresponding to the target scheduling policy executes the decoding stage. By scheduling different business container groups to process different processing stages of the task to be inferred, the cloud computing platform can separate the pre-filling stage and the decoding stage for deployment and processing, specifically, the full instance focuses on executing the pre-filling stage, and the incremental instance focuses on executing the decoding stage, which can improve the processing performance of the task to be inferred.

[0016] In an optional implementation, the above-mentioned scheduling of the business container group corresponding to the target scheduling policy to process the task to be inferred specifically includes: when the current scheduling policy is different from the target scheduling policy, scheduling the business container group corresponding to the target scheduling policy to process the task to be inferred.

[0017] In this implementation, when the current scheduling strategy is different from the target scheduling strategy, it means that the current scheduling strategy is different from the scheduling strategy dynamically determined by the cloud computing platform from multiple scheduling strategies based on the task type of the task to be inferred and the capability information of the business container group corresponding to each scheduling strategy. It can also be understood that the current scheduling strategy is not the better scheduling strategy determined by the cloud computing platform (specifically, it can be a scheduling strategy with better corresponding task processing performance). At this time, the cloud computing platform can replace the current scheduling strategy with the target scheduling strategy, and execute the task to be inferred based on the target scheduling strategy, which can improve the processing performance of the task to be inferred.

[0018] In an optional implementation, the above-mentioned scheduling of the business container group corresponding to the target scheduling policy to process the task to be inferred specifically includes: when the current scheduling policy does not meet the latency requirement, scheduling the business container group corresponding to the target scheduling policy to process the task to be inferred.

[0019] In this implementation, if the current scheduling strategy does not meet the latency requirements, it means that the latency consumed by executing the inference task based on the scheduling strategy (or the business container group corresponding to the scheduling strategy) is high and cannot meet the user's low latency requirements. At this time, the cloud computing platform can replace the current scheduling strategy with the target scheduling strategy, and schedule the business container group corresponding to the target scheduling strategy to process the task to be inferred, which can reduce the processing latency of the task to be inferred, meet the relevant latency requirements, and improve the user experience.

[0020] In an optional implementation, the task information of the task to be inferred also includes the task length of the task to be inferred, and the above-mentioned determining the target scheduling strategy from the multiple scheduling strategies based on the task type and the capability information of the business container group corresponding to each scheduling strategy in the multiple scheduling strategies specifically includes: determining the task load of the task to be inferred based on the task type of the task to be inferred and the task length of the task to be inferred; determining the reasoning performance of each scheduling strategy based on the task load of the task to be inferred and the capability information of the business container group corresponding to each scheduling strategy; and determining the target scheduling strategy from the multiple scheduling strategies based on the reasoning performance of each scheduling strategy.

[0021] In this implementation, after obtaining the task information of the task to be inferred, the cloud computing platform can conveniently and quickly determine the task load of the task to be inferred based on the task type of the task to be inferred and the task length of the task to be inferred included in the task information. Since the capability information of the business container group can characterize the task processing performance of the business container group, and the task load of the task to be inferred can characterize the processing pressure of the task to be inferred, the cloud computing platform can comprehensively determine the reasoning performance of each scheduling strategy based on the processing pressure of the task to be inferred and the task processing performance of the business container group corresponding to each scheduling strategy. Then, the target scheduling strategy can be determined from multiple scheduling strategies based on the reasoning performance of each scheduling strategy. The cloud computing platform can accurately and effectively determine the target scheduling strategy through parameters such as task type, task length, and capability information of the business container group, thereby improving the effectiveness of task processing.

[0022] In an optional implementation, the above-mentioned determining the task load of the task to be reasoned according to the task type of the task to be reasoned and the task length of the task to be reasoned specifically includes: determining the task load of the task to be reasoned from an information load relationship according to the task type and the task length, and the information load relationship is used to indicate: the task type and task length corresponding to one or more task loads.

[0023] In this implementation, the cloud computing platform can query the information load relationship according to the task type and the task length, and thereby determine the task load corresponding to both the task type and the task length from the information load relationship, that is, the task load of the task to be inferred can be obtained.

[0024] In an optional implementation, the above-mentioned determining the target scheduling strategy from the multiple scheduling strategies based on the reasoning performance of each scheduling strategy specifically includes: determining a scheduling strategy among the multiple scheduling strategies whose reasoning performance is greater than or equal to a first performance threshold as the target scheduling strategy.

[0025] In this implementation, for any of the multiple scheduling strategies, if the reasoning performance of the scheduling strategy is greater than or equal to the first performance threshold, it means that the reasoning performance of the scheduling strategy is high, which can also be understood as the performance of the business container group corresponding to the scheduling strategy in processing the task to be reasoned is high / better. At this time, the cloud computing platform can determine the scheduling strategy as the target scheduling strategy. While improving the efficiency of determining the target scheduling strategy, the reasoning efficiency of the task to be reasoned can be improved.

[0026] In an optional implementation, the above-mentioned determining the target scheduling strategy from the multiple scheduling strategies based on the reasoning performance of each scheduling strategy specifically includes: determining the scheduling strategy whose corresponding number of business container groups in at least two scheduling strategies is less than or equal to a quantity threshold as the target scheduling strategy, and the at least two scheduling strategies are scheduling strategies whose reasoning performance in the multiple scheduling strategies is greater than or equal to a second performance threshold, and the second performance threshold is less than the first performance threshold.

[0027] In this implementation, at least two scheduling strategies are scheduling strategies whose reasoning performance is greater than or equal to the second performance threshold among multiple scheduling strategies, indicating that the reasoning performance of the at least two scheduling strategies is higher among multiple scheduling strategies. At this time, the cloud computing platform can determine the number of business container groups corresponding to each of the at least two scheduling strategies. For any of the at least two scheduling strategies, if the number of business container groups corresponding to the scheduling strategy is less than or equal to the number threshold, it means that the number of business container groups corresponding to the scheduling strategy is small. At this time, the cloud computing platform can determine the scheduling strategy as the target scheduling strategy. Thereby, the cloud computing platform can realize the process of scheduling fewer business container groups to process the tasks to be inferred, which can further save hardware resources.

[0028] In the second aspect, a cloud computing platform is provided. The cloud computing platform is deployed with a computing power evaluation module, a scheduler module and a hardware resource pool, and the hardware resource pool is deployed with a business container group; the computing power evaluation module is used to obtain task information of the task to be inferred and capability information of the business container group, the task information includes the task type of the task to be inferred, the business container group is used to be assigned to perform the task to be inferred, and the capability information is used to characterize the task processing performance of the business container group; the computing power evaluation module is also used to determine a target scheduling strategy from the multiple scheduling strategies according to the task type and the capability information of the business container group corresponding to each scheduling strategy in the multiple scheduling strategies, and each scheduling strategy is used to indicate the scheduling of the business container group corresponding to each scheduling strategy to process the inference task corresponding to each scheduling strategy; the scheduler module is used to schedule the business container group corresponding to the target scheduling strategy to process the task to be inferred.

[0029] In this application, since the target scheduling strategy is determined by the computing power evaluation module from multiple scheduling strategies based on the task type of the task to be inferred and the capability information of the business container group corresponding to each scheduling strategy, the scheduler module can then schedule the business container group corresponding to the target scheduling strategy to process the task to be inferred. Therefore, the computing power evaluation module can dynamically adjust the scheduling strategy of the task to be inferred according to the task type and the task processing of the business container group, which can avoid the hardware resources of the related board being occupied for a long time. It can improve the flexibility of task processing while saving hardware resources.

[0030] In an optional implementation, the task information of the task to be inferred also includes the task length of the task to be inferred, the computing power evaluation module is deployed with a load analyzer module and a separate deployment simulator module, the separate deployment simulator module is deployed with a scheduler simulator module and a model modeling module, the scheduler simulator module stores each scheduling strategy, the model modeling module is used to simulate the reasoning environment of the LLM in each scheduling strategy, and the reasoning environment of the LLM in each scheduling strategy is constructed based on the capability information of the business container group corresponding to each scheduling strategy; the load analyzer module is used to determine the task load of the task to be inferred according to the task type of the task to be inferred and the task length of the task to be inferred; the load analyzer module is also used to send a virtual task to the separate deployment simulator module, and the task load of the virtual task is the same as the task load of the task to be inferred; the model modeling module is used to obtain the reasoning performance of each scheduling strategy in response to the input virtual task.

[0031] In this implementation, since the model modeling module is used to simulate the reasoning environment of LLM in each scheduling strategy, and the reasoning environment of LLM in each scheduling strategy is constructed based on the capability information of the business container group corresponding to each scheduling strategy. That is, the model modeling module can fully and effectively simulate the real task processing level of each scheduling strategy. Therefore, when the model modeling module responds to the input virtual task with the same task load as the task to be inferred, the reasoning performance of each scheduling strategy can be accurately and effectively simulated, thereby improving the determination efficiency of the target scheduling strategy.

[0032] In an optional implementation, the computing power evaluation module is also deployed with a result searcher module; the result searcher module is used to receive the reasoning performance of each scheduling strategy sent by the separate deployment simulator module; the result searcher module is also used to determine the target scheduling strategy from the multiple scheduling strategies based on the reasoning performance of each scheduling strategy.

[0033] In this implementation, since the reasoning performance of each scheduling strategy is fully and effectively simulated by the model modeling module based on the reasoning environment of each scheduling strategy and the virtual task with the same task load as the task to be reasoned, the result searcher module determines the target scheduling strategy from multiple scheduling strategies according to the reasoning performance of each scheduling strategy, and can use a method that combines modeling with real business scenarios to determine the target scheduling strategy in real time, and then the scheduler module schedules the business container group corresponding to the target scheduling strategy to process the task to be reasoned, so as to realize separated deployment services.

[0034] In a third aspect, a task processing device running on a cloud computing platform is provided, the device comprising an acquisition module, a determination module and a scheduling module; the acquisition module is used to acquire task information of the task to be inferred and capability information of the business container group, the task information comprising the task type of the task to be inferred, the business container group being used to be assigned to execute the task to be inferred, and the capability information being used to characterize the task processing performance of the business container group; the determination module is used to determine a target scheduling strategy from the multiple scheduling strategies according to the task type and the capability information of the business container group corresponding to each scheduling strategy in the multiple scheduling strategies, the each scheduling strategy being used to instruct the business container group corresponding to each scheduling strategy to be scheduled to process the inference task corresponding to each scheduling strategy; the scheduling module is used to schedule the business container group corresponding to the target scheduling strategy to process the task to be inferred.

[0035] Optionally, the capability information of the service container group includes the remaining resource size of the service container group and / or the board type of the board to which the service container group belongs.

[0036] Optionally, a scheduling strategy includes an instance ratio, where the instance ratio is a ratio between the number of full instances corresponding to the scheduling strategy and the number of incremental instances corresponding to the scheduling strategy, and one instance ratio corresponds to a business container group set;

[0037] The determination module is further used to determine the business container group included in the target business container group set as the business container group corresponding to the target scheduling policy, the target business container group set is the business container group set corresponding to the target instance ratio, and the target instance ratio is the instance ratio included in the target scheduling policy.

[0038] Optionally, different scheduling strategies include different instance ratios.

[0039] Optionally, the above-mentioned target business container group set includes at least one first business container group and at least one second business container group, the at least one first business container group is the full instance corresponding to the target scheduling strategy, and the at least one second business container group is the incremental instance corresponding to the target scheduling strategy; the scheduling module is specifically used to schedule the at least one first business container group to process the pre-filling stage of the task to be inferred, and schedule the at least one second business container group to process the decoding stage of the task to be inferred.

[0040] Optionally, the scheduling module is specifically configured to schedule the business container group corresponding to the target scheduling strategy to process the task to be inferred when the current scheduling strategy is different from the target scheduling strategy.

[0041] Optionally, the scheduling module is specifically used to schedule the business container group corresponding to the target scheduling strategy to process the task to be inferred when the current scheduling strategy does not meet the latency requirement.

[0042] Optionally, the task information of the task to be reasoned also includes the task length of the task to be reasoned; the determination module is specifically used to determine the task load of the task to be reasoned according to the task type of the task to be reasoned and the task length of the task to be reasoned; the determination module is also specifically used to determine the reasoning performance of each scheduling strategy according to the task load of the task to be reasoned and the capability information of the business container group corresponding to each scheduling strategy; the determination module is also specifically used to determine the target scheduling strategy from the multiple scheduling strategies according to the reasoning performance of each scheduling strategy.

[0043] Optionally, the determination module is further specifically used to determine the task load of the task to be inferred from an information load relationship according to the task type and the task length, and the information load relationship is used to indicate: the task type and task length corresponding to one or more task loads.

[0044] Optionally, the determination module is further specifically configured to determine, among the multiple scheduling strategies, a scheduling strategy whose reasoning performance is greater than or equal to a first performance threshold, as the target scheduling strategy.

[0045] Optionally, the determination module is also specifically used to determine the scheduling strategy whose corresponding number of business container groups in at least two scheduling strategies is less than or equal to a quantity threshold as the target scheduling strategy, and the at least two scheduling strategies are scheduling strategies whose reasoning performance is greater than or equal to a second performance threshold among the multiple scheduling strategies, and the second performance threshold is less than the first performance threshold.

[0046] In addition, the technical effects of the task processing device described in the third aspect can refer to the technical effects of the method described in the first aspect, and will not be repeated here.

[0047] In a fourth aspect, a computing device cluster is provided, comprising at least one computing device, each computing device comprising a processor and a memory; the processor of the at least one computing device is used to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method described in the first aspect.

[0048] In a fifth aspect, a computer program product comprising instructions is provided. When the instructions are executed by a computing device cluster, the computing device cluster executes the method described in the first aspect.

[0049] In a sixth aspect, a computer-readable storage medium is provided, comprising computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster executes the method described in the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1A schematic diagram of the structure of a cloud computing platform provided in an embodiment of the present application is shown;

[0051] Figure 2 A flowchart of a task processing method provided in an embodiment of the present application is shown;

[0052] Figure 3 A flowchart of another task processing method provided in an embodiment of the present application is shown;

[0053] Figure 4 A flowchart of another task processing method provided in an embodiment of the present application is shown;

[0054] Figure 5 A flowchart of another task processing method provided in an embodiment of the present application is shown;

[0055] Figure 6 A flowchart of another task processing method provided in an embodiment of the present application is shown;

[0056] Figure 7 A schematic diagram of the structure of another cloud computing platform provided in an embodiment of the present application is shown;

[0057] Figure 8 A flowchart of another task processing method provided in an embodiment of the present application is shown;

[0058] Fig. 9 A flowchart of another task processing method provided in an embodiment of the present application is shown;

[0059] Fig.10 A schematic diagram of the structure of a task processing device provided in an embodiment of the present application is shown;

[0060] Fig.11 A schematic diagram of the structure of a computing device provided in an embodiment of the present application is shown;

[0061] Fig.12 A schematic diagram of the structure of a computing device cluster provided in an embodiment of the present application is shown;

[0062] Fig.13 A schematic diagram of a network architecture of a computing device cluster provided in an embodiment of the present application is shown. DETAILED DESCRIPTION

[0063] To facilitate understanding of the solutions provided by the embodiments of the present application, some related technologies involved in the present application are explained before introducing the solutions provided by the embodiments of the present application.

[0064] With the development of generative artificial intelligence (GAI) technology, LLM is being used more and more widely, and users' performance requirements for LLM are increasing. Therefore, how to use resources more efficiently under the existing architecture to process more reasoning requests faster and better is of great significance to users.

[0065] The reason why LLM's reasoning process limits its efficient use of resources is that the reasoning process consists of two distinct stages: the pre-filling stage and the decoding stage. The two stages have different requirements for resources. Specifically, the pre-filling stage requires high computing power, while the decoding stage requires high bandwidth. When the pre-filling stage and the decoding stage are combined for processing, no matter how the scheduling strategy is adjusted, there is a strong coupling relationship between the two indicators: time to first token (TTFT) and time per output token (TPOT). Since resources (including graphics processing units (GPUs) and neural processing units (NPUs)) are limited, when executing the pre-filling stage, the decoding stage cannot be executed at the same time. This execution process actually sacrifices TPOT to preserve TTFT, and the principle that the pre-filling stage cannot be executed at the same time when executing the decoding stage is similar.

[0066] The idea of ​​deploying the pre-filling stage and the decoding stage separately is proposed in the related technology. Specifically, in the separated framework, the pre-filling stage and the decoding stage no longer share a board / group of boards (i.e., GPU / NPU). The full instance focuses on executing the pre-filling stage, and the incremental instance focuses on executing the decoding stage. After the full instance of the pre-filling stage is calculated, the output key-value (KV) cache is transferred to the incremental instance of the decoding stage, and the incremental instance continues to complete the reasoning process. If a board only executes the pre-filling stage or only executes the decoding stage, the effective throughput (good put) is better than that of executing both the pre-filling stage and the decoding stage.

[0067] Optionally, the above separation deployment includes homogeneous separation deployment and heterogeneous separation deployment. Specifically, homogeneous separation deployment refers to the application of the same type of physical device (or machine) in the pre-filling stage and the decoding stage; heterogeneous separation deployment refers to the application of one type of physical device in the pre-filling stage and another type of physical device in the decoding stage.

[0068] However, the above-mentioned related technologies have at least the following three defects:

[0069] 1. Inference requests do not maintain a high-voltage number at all times, and not all inference requests are long-sequence inference requests. In the process of deploying the pre-filling stage and the decoding stage on different boards respectively, the number of full instances executing the pre-filling stage and the number of incremental instances executing the decoding stage are fixed, and these fixed numbers of full instances and incremental instances will occupy the corresponding boards for a long time, and flexible configuration cannot be achieved. When the hardware resources required for the LLM reasoning process are far less than the hardware resources that these boards can provide (which can also be understood as when the pressure of the reasoning request is small), it will cause a lot of waste of hardware resources.

[0070] Second, separate deployment uses only a single / fixed model. Maintaining a single / fixed model for a long time will also cause a waste of hardware resources. It is impossible to make targeted model adjustments based on the current inference performance, and the cost-effectiveness is low.

[0071] 3. The operation and maintenance / tuning process in the separated deployment requires a lot of time and manpower for testing and verification, and it is impossible to always maintain the optimal performance of the reasoning process in the real business scenarios that change at any time.

[0072] Based on this, the embodiment of the present application provides a task processing method running on a cloud computing platform. Since the target scheduling strategy is determined by the cloud computing platform from multiple scheduling strategies according to the task type of the task to be inferred and the capability information of the business container group corresponding to each scheduling strategy, the cloud computing platform can dynamically adjust the scheduling strategy of the task to be inferred according to the task type and the task processing of the business container group, which can avoid the hardware resources of the relevant board being occupied for a long time. The flexibility of task processing is improved while saving hardware resources.

[0073] The task processing method, device and equipment running on the cloud computing platform provided in the embodiments of the present application are applied to natural language processing scenarios (including but not limited to voice assistants, chatbots and text generation scenarios). After the cloud computing obtains the task information of the task to be inferred, it can determine the target scheduling strategy from multiple scheduling strategies, and schedule the business container group corresponding to the target scheduling strategy to process the task to be inferred, thereby obtaining the inference result of the task to be inferred.

[0074] In order to enable ordinary persons in the art to better understand the technical solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings.

[0075] It is understood that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the attached claims.

[0076] It should also be understood that the term “comprising” indicates the presence of described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements and / or components.

[0077] In one example, the cloud computing platform that executes the task processing method provided in the embodiment of the present application can be a terminal, which can also be called user equipment (UE), access terminal, subscriber unit (subscriber unit), user station, mobile station (MS), mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication equipment, user agent or user device. The terminal in the embodiments of the present application can be a mobile phone, a cellular phone, a smart phone, a tablet computer, a wireless data card, a personal digital assistant (PDA), a wireless modem, a handset, a laptop computer, a machine type communication (MTC) terminal, a computer with a wireless transceiver function, a virtual reality (VR) terminal, an augmented reality (AR) terminal, a smart home device (for example, a refrigerator, a television, an air conditioner, an electric meter, etc.), an intelligent robot, a robotic arm, a workshop equipment, a wireless terminal in industrial control, a wireless terminal in self driving, a wireless terminal in remote medical, a wireless terminal in a smart grid, a wireless terminal in transportation safety, a wireless terminal in a smart city, a wireless terminal in a smart home, a vehicle-mounted terminal, a road side unit with a terminal function, and a wireless terminal in a smart city. unit, RSU), etc., flying equipment (e.g., intelligent robots, hot air balloons, drones, airplanes), etc. The terminal of the present application may also be an on-board module, on-board module, on-board component, on-board chip or on-board unit built into the vehicle as one or more components or units. The terminal may also be other devices with terminal functions, for example, the terminal may also be a device that functions as a terminal in device to device (D2D) communication.

[0078] The embodiments of the present application do not limit the device form of the terminal. The device for realizing the function of the terminal can be a terminal; it can also be a device that can support the terminal to realize the function, such as a chip system. The device can be installed in the terminal, or used in combination with the terminal. In the embodiments of the present application, the chip system can be composed of chips, or it can include chips and other discrete devices.

[0079] In other examples, the cloud computing platform may also be a server, which may be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, network acceleration services (content delivery network, CDN), as well as big data and artificial intelligence platforms.

[0080] like Figure 1 As shown, in an optional implementation, a cloud computing platform that executes the task processing method provided in the embodiment of the present application may include a scheduler module 101, a computing power evaluation module 102, and a hardware resource pool 103. Among them, the hardware resource pool 103 includes five business container groups, namely business container group 1031, business container group 1032, business container group 1033, business container group 1034, and business container group 1035. Business container group 1031 and business container group 1032 are used to be assigned to perform separation deployment reasoning task A, and business container group 1033, business container group 1034, and business container group 1035 are used to be assigned to perform separation deployment reasoning task B. The computing power evaluation module 102 includes a business evaluation module 1021, a hardware resource evaluation module 1022, and a performance evaluation module 1023. Multiple scheduling strategies are stored in the computing power evaluation module 102.

[0081] The business evaluation module 1021 is used to determine the task load of the task to be inferred.

[0082] The hardware resource evaluation module 1022 is used to obtain the capability information of each of the five service container groups.

[0083] The performance evaluation module 1023 is used to determine the reasoning performance of each of the above-mentioned multiple scheduling strategies.

[0084] The scheduler module 101 is used to schedule the business container group corresponding to the target scheduling strategy to process the task to be inferred.

[0085] Optionally, the hardware resource pool 103 can feedback the resource utilization of the inference task to the computing power evaluation module 102, and the scheduler module 101 can send the service configuration of the task to be inferred to the hardware resource pool 103, that is, configure the business container group corresponding to the target scheduling strategy to process the task to be inferred.

[0086] The functions of the cloud computing platform in the embodiment of the present application may be integrated in a computer device cluster, which includes one or more computing devices (the computing device may be the computing device 1100 provided in the following embodiment). The computing device includes at least a memory and a processor.

[0087] The memory stores executable program codes, and the processor executes the executable program codes to respectively implement the functions of the computing power evaluation module 102 (including the business evaluation module 1021, the hardware resource evaluation module 1022, and the performance evaluation module 1023) and the scheduler module 101, thereby implementing the task processing method provided in the embodiment of the present application. That is, the memory stores instructions for executing the task processing method.

[0088] Combined with the description of the above embodiments, the task processing method provided in the embodiments of the present application can be executed by a cloud computing platform. Figure 2 As shown, the task processing method may include S201-S203.

[0089] S201: Obtain task information of the task to be inferred and capability information of the business container group.

[0090] The task information of the task to be inferred includes the task type of the task to be inferred, the business container group is used to be assigned to execute the task to be inferred, and the capability information of the business container group is used to characterize the task processing performance of the business container group.

[0091] It should be understood that the task to be reasoned is a task that requires LLM to perform a reasoning process and obtain a reasoning result.

[0092] Among them, LLM refers to a deep learning model trained with a large amount of text data, which can generate natural language text or understand the meaning of language text. LLM can obtain powerful general modeling and generalization capabilities by pre-training on a large data set, and can perform complex language tasks such as text summarization, machine translation, sentiment analysis, dialogue generation, and content recommendation.

[0093] Optionally, the task to be inferred may also be acquired by the cloud computing platform in the form of an inference request. In this case, the task to be inferred may be understood as a request to be inferred.

[0094] Exemplarily, the task type of the task to be reasoned may be a long sequence reasoning task, a short sequence reasoning task, or a multi-round dialogue reasoning task, etc. The embodiment of the present application does not specifically limit the task type of the task to be reasoned.

[0095] S202: Determine a target scheduling strategy from multiple scheduling strategies according to the task type of the task to be inferred and the capability information of the service container group corresponding to each scheduling strategy in the multiple scheduling strategies.

[0096] Each scheduling strategy is used to instruct the service container group corresponding to each scheduling strategy to be scheduled to process the reasoning task corresponding to each scheduling strategy.

[0097] It should be understood that the multiple scheduling strategies can be pre-stored in the cloud computing platform.

[0098] In an optional implementation manner, the capability information of the service container group (pod) includes the remaining resource size of the service container group and / or the board type of the board to which the service container group belongs.

[0099] The cloud computing platform determines the target scheduling strategy from multiple scheduling strategies based on the task type to be inferred and the capability information of the business container group corresponding to each scheduling strategy in multiple scheduling strategies. For the specific method, please refer to the following text. Figure 5 S2021-S2023 shown are not repeated here.

[0100] In this implementation, for a business container group, if the remaining resource size of the business container group is larger, it means that the task processing performance of the business container group is higher; on the contrary, if the remaining resource size of the business container group is smaller, it means that the task processing performance of the business container group is lower. In addition, different types of boards have different performance levels, so the performance of business container groups included (or deployed) in different types of boards also varies. In this way, the cloud computing platform can accurately and effectively determine the processing performance of the business container group based on the remaining resource size of the business container group and / or the board type of the board to which the business container group belongs, thereby improving the efficiency of determining the target scheduling strategy.

[0101] S203: The business container group corresponding to the scheduling target scheduling strategy processes the task to be inferred.

[0102] Specifically, the cloud computing platform can schedule the hardware resources (including GPU, NPU) in the business container group corresponding to the target scheduling strategy to complete the reasoning / processing process of the task to be reasoned in the business container group, thereby obtaining the reasoning result of the task to be reasoned.

[0103] Please refer to the following text for the specific way in which the business container group corresponding to the cloud computing platform scheduling target scheduling strategy handles the tasks to be inferred. Figure 4The S2031 shown, and the subsequent steps A and B are not described in detail here.

[0104] In the embodiment of the present application, since the target scheduling strategy is determined by the cloud computing platform from multiple scheduling strategies according to the task type of the task to be inferred and the capability information of the business container group corresponding to each scheduling strategy, the cloud computing platform can dynamically adjust the scheduling strategy of the task to be inferred according to the task type and the task processing performance of the business container group, which can avoid the hardware resources of the relevant board being occupied for a long time, thereby saving hardware resources and improving the flexibility of task processing.

[0105] In some embodiments, a scheduling strategy includes an instance ratio, which is the ratio between the number of full instances corresponding to the scheduling strategy and the number of incremental instances corresponding to the scheduling strategy. One instance ratio corresponds to a set of business container groups. Figure 2 ,like Figure 3 As shown, before the business container group corresponding to the scheduling target scheduling strategy processes the task to be inferred, the task processing method provided in the embodiment of the present application may further include S204.

[0106] S204: Determine the service container group included in the target service container group set as the service container group corresponding to the target scheduling policy.

[0107] The target business container group set is a business container group set corresponding to a target instance ratio, and the target instance ratio is an instance ratio included in the target scheduling policy.

[0108] Exemplarily, the following Table 1 is an example of a policy set relationship provided in an embodiment of the present application. Specifically, the ratio set relationship includes three instance ratios, the three instance ratios are 1:2, 1:3 and 2:1, and the business container group sets corresponding to the three instance ratios are business container group set 1, business container group set 2 and business container group set 3.

[0109] Table 1

[0110]

[0111] Assume that the instance ratio included in the target scheduling policy is 2:1, and the business container group set 3 includes business container group a, business container group b, and business container group c. Then the cloud computing platform determines that the target business container group set is business container group set 3, and the business container groups corresponding to the target scheduling policy are business container group a, business container group b, and business container group c.

[0112] In this embodiment, after determining the target scheduling strategy, the cloud computing platform can determine the target business container group set (i.e., the business container group set corresponding to the target instance ratio) based on the target instance ratio included in the target scheduling strategy, specifically, the ratio between the number of full instances corresponding to the target scheduling strategy and the number of incremental instances corresponding to the target scheduling strategy, thereby determining the business container group included in the target business container group set as the business container group corresponding to the target scheduling strategy, and being able to accurately and effectively determine the business container group used to process the task to be inferred, thereby improving the effectiveness of task processing.

[0113] In an optional implementation, for the above-mentioned multiple scheduling strategies, different scheduling strategies include different instance ratios.

[0114] In this implementation, since different scheduling strategies include different instance ratios, that is, for multiple scheduling strategies, the ratio between the number of full instances corresponding to one scheduling strategy and the number of incremental instances corresponding to the scheduling strategy is different from the ratio between the number of full instances corresponding to other scheduling strategies and the number of incremental instances corresponding to the other scheduling strategies. At this time, the cloud computing platform can determine a unique business container group for processing the task to be inferred (i.e., the business container group corresponding to the target scheduling strategy) based on the unique instance ratio included in the unique scheduling strategy (i.e., the target instance ratio included in the target scheduling strategy), which can improve the accuracy of task processing.

[0115] In some embodiments, the target business container group set includes at least one first business container group and at least one second business container group, the at least one first business container group is a full instance corresponding to the target scheduling policy, and the at least one second business container group is an incremental instance corresponding to the target scheduling policy. Figure 2 ,like Figure 4 As shown, the business container group corresponding to the above scheduling target scheduling strategy processes the task to be inferred, specifically including S2031.

[0116] S2031. Schedule at least one first service container group to process a pre-filling phase of a task to be inferred, and schedule at least one second service container group to process a decoding phase of the task to be inferred.

[0117] It should be understood that the task to be inferred in the embodiment of the present application includes at least two processing stages, namely, a pre-filling stage and a decoding stage. The pre-filling stage can be used to obtain the candidate results of the task to be inferred, and the decoding stage can be used to obtain the final result of the task to be inferred (which can also be understood as the reasoning result of the task to be inferred).

[0118] Among them, the prefill phase is a phase included in the reasoning process of LLM. The main task of this phase is to generate candidate text. Specifically, in the prefill phase, an input token sequence can be generated according to the natural language text (prompt) input by the user, and the generated input token sequence can be batch processed and the KV cache calculated to finally generate the first output token. In natural language processing, a token is the smallest unit of text data, which can be a word, a part of a word, a character, or a word.

[0119] The decoding phase is another phase included in the reasoning process of LLM. The main task of this phase is to select the final result from the candidate text. Specifically, in the decoding phase, the first output word generated in the pre-filling phase is added to the input word sequence and used as a new input to generate the next word, that is, subsequent words are generated by iteration until the end mark is generated or the maximum sequence length is reached.

[0120] A full instance refers to all the data of a database system that needs to be migrated during database system migration or data synchronization, which is usually completed using batch processing. The full instance in the embodiment of the present application is used to perform the pre-filling stage in the LLM reasoning process.

[0121] Incremental instances refer to newly generated data during database system migration or data synchronization, which are synchronized through streaming computing. The incremental instances in the embodiments of the present application are used to perform the decoding phase in the LLM reasoning process.

[0122] In the embodiment of the present application, since the full number of instances corresponding to a scheduling strategy can be used to execute the prefill phase of the task to be inferred, the incremental instances corresponding to the scheduling strategy can be used to execute the decoding phase of the task to be inferred. Therefore, the instance ratio included in the scheduling strategy can also be understood as the ratio between the number of instances used to execute the prefill phase in the scheduling strategy and the number of instances used to execute the decode phase in the scheduling strategy. That is, the instance ratio can also be called the PD ratio.

[0123] In this embodiment, the cloud computing platform schedules at least one business container group to process the pre-filling stage of the task to be inferred, that is, the full instance corresponding to the target scheduling policy executes the pre-filling stage; the cloud computing platform schedules at least one second business container group to process the decoding stage of the task to be inferred, that is, the incremental instance corresponding to the target scheduling policy executes the decoding stage. By scheduling different business container groups to process different processing stages of the task to be inferred, the cloud computing platform can separate the pre-filling stage and the decoding stage for deployment and processing, specifically, the full instance focuses on executing the pre-filling stage, and the incremental instance focuses on executing the decoding stage, which can improve the processing performance of the task to be inferred.

[0124] In an implementation manner of the embodiment of the present application, the business container group corresponding to the scheduling target scheduling strategy processes the task to be inferred, which may specifically include step A.

[0125] Step A: When the current scheduling strategy is different from the target scheduling strategy, the business container group corresponding to the target scheduling strategy is scheduled to process the task to be inferred.

[0126] It should be understood that the current scheduling strategy is the scheduling strategy currently being used by the cloud computing platform for processing reasoning tasks.

[0127] In this implementation, when the current scheduling strategy is different from the target scheduling strategy, it means that the current scheduling strategy is different from the scheduling strategy dynamically determined by the cloud computing platform from multiple scheduling strategies based on the task type of the task to be inferred and the capability information of the business container group corresponding to each scheduling strategy. It can also be understood that the current scheduling strategy is not the better scheduling strategy determined by the cloud computing platform (specifically, it can be a scheduling strategy with better corresponding task processing performance). At this time, the cloud computing platform can replace the current scheduling strategy with the target scheduling strategy, and execute the task to be inferred based on the target scheduling strategy, which can improve the processing performance of the task to be inferred.

[0128] In another implementation of the embodiment of the present application, the business container group corresponding to the scheduling target scheduling strategy processes the task to be inferred, which may specifically include step B.

[0129] Step B: When the current scheduling strategy does not meet the latency requirement, the business container group corresponding to the target scheduling strategy is scheduled to process the task to be inferred.

[0130] In this implementation, if the current scheduling strategy does not meet the latency requirements, it means that the latency consumed by executing the inference task based on the scheduling strategy (or the business container group corresponding to the scheduling strategy) is high and cannot meet the user's low latency requirements. At this time, the cloud computing platform can replace the current scheduling strategy with the target scheduling strategy, and schedule the business container group corresponding to the target scheduling strategy to process the task to be inferred, which can reduce the processing latency of the task to be inferred, meet the relevant latency requirements, and improve the user experience.

[0131] Optionally, the latency requirement in the embodiment of the present application may be a service level objective (SLO) latency requirement.

[0132] Optionally, for any scheduling strategy of the embodiments of the present application, the cloud computing platform can obtain the TTFT corresponding to the scheduling strategy and the TPOT corresponding to the scheduling strategy, and determine the delay corresponding to the scheduling strategy based on the TTFT corresponding to the scheduling strategy and the TPOT corresponding to the scheduling strategy, so as to determine whether the delay corresponding to the scheduling strategy meets the delay requirements.

[0133] Among them, TTFT refers to the time from the user input query to the LLM output of the first word, TTFT corresponds to the pre-filling stage. TTFT is an important indicator for measuring user experience and can determine the time the user waits for the first response. TPOT refers to the time for each output word, excluding the time for the first output word, TPOT corresponds to the decoding stage. TPOT is a key indicator for measuring the time consumed by the entire reasoning process and can determine the total time required to process all words.

[0134] In some embodiments, the task information of the task to be inferred also includes the task length of the task to be inferred. Figure 2 ,like Figure 5 As shown, the above-mentioned determining the target scheduling strategy from multiple scheduling strategies according to the task type of the task to be inferred and the capability information of the business container group corresponding to each scheduling strategy in the multiple scheduling strategies may specifically include S2021-S2023.

[0135] S2021. Determine the task load of the task to be reasoned according to the task type and the task length of the task to be reasoned.

[0136] In an optional implementation, the cloud computing platform can determine the task load of the task to be reasoned from the information load relationship based on the task type of the task to be reasoned and the task length of the task to be reasoned, where the information load relationship is used to indicate the task type and task length corresponding to one or more task loads.

[0137] It should be understood that the cloud computing platform can query the information load relationship according to the task type and the task length, and thereby determine the task load corresponding to the task type and the task length from the information load relationship, that is, the task load of the task to be inferred can be obtained.

[0138] S2022: Determine the reasoning performance of each scheduling strategy according to the task load of the task to be inferred and the capability information of the business container group corresponding to each scheduling strategy.

[0139] In combination with the description of the above embodiment, it should be understood that the capability information of the business container group is used to characterize the task processing performance of the business container group. When the task load of the task to be inferred is small and the task processing performance of the business container group corresponding to a certain scheduling strategy is high, it means that the processing process of the task to be inferred is relatively simple and the business container group corresponding to the scheduling strategy has a high efficiency in executing the task to be inferred. At this time, the cloud computing platform can determine that the reasoning performance of the scheduling strategy is high / superior.

[0140] Correspondingly, when the task load of the task to be inferred is large and / or the task processing performance of the business container group corresponding to a certain scheduling strategy is low, it means that the processing process of the task to be inferred is relatively complicated and / or the efficiency of the business container group corresponding to the scheduling strategy in executing the task to be inferred is low. At this time, the cloud computing platform can determine that the reasoning performance of the scheduling strategy is low / poor.

[0141] In an optional implementation, the cloud computing platform can determine the reasoning performance of each scheduling strategy from the task performance relationship based on the task load of the task to be inferred and the capability information corresponding to each scheduling strategy, where the task performance relationship is used to indicate the task load and capability information corresponding to one or more reasoning performances.

[0142] Optionally, the cloud computing platform may determine the capability information corresponding to each scheduling strategy according to the capability information of each service container group corresponding to each scheduling strategy.

[0143] Exemplarily, the following Tables 2 and 3 are examples of information load relationships and task performance relationships provided in the embodiments of the present application. Specifically, Table 2 includes three task loads, namely, load 1, load 2, and load 3. The task types corresponding to the three task loads are type 1, type 2, and type 3, and the task lengths corresponding to the three task loads are length 1, length 1, and length 2.

[0144] Table 3 includes five reasoning performances, namely performance 1, performance 2, performance 3, performance 4 and performance 5. The task loads corresponding to the five reasoning performances are load 1, load 1, load 1, load 2 and load 3, and the capability information corresponding to the five reasoning performances are information 1, information 2, information 3, information 1 and information 2.

[0145] Table 2

[0146]

[0147] Table 3

[0148]

[0149] Assuming that the task type of the task to be inferred is type 1 and the task length of the task to be inferred is length 1, the cloud computing platform determines that the task load of the task to be inferred is load 1.

[0150] Assume that the above multiple scheduling strategies include scheduling strategy a, scheduling strategy b and scheduling strategy c, and the capability information corresponding to these three scheduling strategies are information 1, information 2 and information 3. Then the cloud computing platform determines that the inference performance corresponding to scheduling strategy a, scheduling strategy b and scheduling strategy c are performance 1, performance 2 and performance 3 respectively.

[0151] S2023. Determine a target scheduling strategy from multiple scheduling strategies based on the reasoning performance of each scheduling strategy.

[0152] It should be understood that the reasoning performance of a scheduling strategy is the performance of the service container group corresponding to the scheduling strategy in processing the tasks to be reasoned.

[0153] The specific method of determining the target scheduling strategy from multiple scheduling strategies according to the inference performance of each scheduling strategy by the cloud computing platform scheduling is referred to steps C and D below and will not be described here.

[0154] In an embodiment of the present application, after the cloud computing platform obtains the task information of the task to be inferred, it can conveniently and quickly determine the task load of the task to be inferred based on the task type of the task to be inferred and the task length of the task to be inferred included in the task information. Since the capability information of the business container group can characterize the task processing performance of the business container group, and the task load of the task to be inferred can characterize the processing pressure of the task to be inferred, the cloud computing platform can comprehensively determine the reasoning performance of each scheduling strategy based on the processing pressure of the task to be inferred and the task processing performance of the business container group corresponding to each scheduling strategy. Then, the target scheduling strategy can be determined from multiple scheduling strategies based on the reasoning performance of each scheduling strategy. The cloud computing platform can accurately and effectively determine the target scheduling strategy through parameters such as task type, task length, and capability information of the business container group, thereby improving the effectiveness of task processing.

[0155] In an implementation of the embodiment of the present application, the above-mentioned determining the target scheduling strategy from multiple scheduling strategies based on the reasoning performance of each scheduling strategy may specifically include step C.

[0156] Step C: Determine a scheduling strategy whose reasoning performance is greater than or equal to a first performance threshold among multiple scheduling strategies as a target scheduling strategy.

[0157] In this implementation, for any of the multiple scheduling strategies, if the reasoning performance of the scheduling strategy is greater than or equal to the first performance threshold, it means that the reasoning performance of the scheduling strategy is high, which can also be understood as the performance of the business container group corresponding to the scheduling strategy in processing the task to be reasoned is high / better. At this time, the cloud computing platform can determine the scheduling strategy as the target scheduling strategy. While improving the efficiency of determining the target scheduling strategy, the reasoning efficiency of the task to be reasoned can be improved.

[0158] In another implementation of the embodiment of the present application, the above-mentioned determination of the target scheduling strategy from multiple scheduling strategies based on the reasoning performance of each scheduling strategy may specifically include step D.

[0159] Step D: determining the scheduling strategy for which the number of service container groups corresponding to at least two scheduling strategies is less than or equal to the number threshold as the target scheduling strategy.

[0160] Among them, the at least two scheduling strategies are scheduling strategies whose reasoning performance is greater than or equal to a second performance threshold among the above-mentioned multiple scheduling strategies, and the second performance threshold is less than the above-mentioned first performance threshold.

[0161] In this implementation, at least two scheduling strategies are scheduling strategies whose reasoning performance is greater than or equal to the second performance threshold among multiple scheduling strategies, indicating that the reasoning performance of the at least two scheduling strategies is higher among multiple scheduling strategies. At this time, the cloud computing platform can determine the number of business container groups corresponding to each of the at least two scheduling strategies. For any of the at least two scheduling strategies, if the number of business container groups corresponding to the scheduling strategy is less than or equal to the number threshold, it means that the number of business container groups corresponding to the scheduling strategy is small. At this time, the cloud computing platform can determine the scheduling strategy as the target scheduling strategy. Thereby, the cloud computing platform can realize the process of scheduling fewer business container groups to process the tasks to be inferred, which can further save hardware resources.

[0162] Optionally, the cloud computing platform may also determine a scheduling strategy with the highest reasoning performance among multiple scheduling strategies as a target scheduling strategy.

[0163] In some embodiments, a scheduling policy also includes machine model information corresponding to the scheduling policy. The machine model information corresponding to a scheduling policy includes at least one of the following information: the board type of the board to which the service container group corresponding to the scheduling policy belongs, the board number of the board to which the service container group corresponding to the scheduling policy belongs, and the number of service container groups corresponding to the scheduling policy.

[0164] The cloud computing platform can also determine the target scheduling strategy based on the inference performance of each scheduling strategy and the machine model information corresponding to each scheduling strategy.

[0165] In an optional manner, the cloud computing platform may determine as a target scheduling strategy a scheduling strategy that includes a high-performance board in at least one scheduling strategy, where the at least one scheduling strategy is a scheduling strategy whose reasoning performance is greater than or equal to a first performance threshold among the above-mentioned multiple scheduling strategies. A scheduling strategy that includes a high-performance board specifically refers to a board type of a board in a business container group corresponding to the scheduling strategy that is a high-performance board.

[0166] In another optional implementation, the cloud computing platform may also determine the scheduling strategy whose corresponding number of boards among the at least two scheduling strategies is less than or equal to the board number threshold as the target scheduling strategy. The number of boards corresponding to a scheduling strategy is specifically the number of boards of the business container group corresponding to the scheduling strategy.

[0167] Optionally, different scheduling strategies include different machine model information.

[0168] In the embodiment of the present application, the task processing method running on the cloud computing platform can also be understood as a separation deployment service. By executing the task processing method, the cloud computing platform can provide the separation deployment service to the user. The following is an explanation of the task processing method provided in the embodiment of the present application based on the separation deployment service. Figure 6 As shown, the task processing method includes S601-S607.

[0169] S601, the separation deployment service is started.

[0170] Specifically, the separation deployment service can be started on the cloud computing platform to start executing the task processing method provided in the embodiment of the present application.

[0171] S602: Request perception analysis.

[0172] Among them, step S602 can be understood as the process of the cloud computing platform obtaining the task information of the task to be inferred in the above embodiment. The task to be inferred at this time can also be understood as a request to be inferred.

[0173] S603: Determine whether the current instance ratio is appropriate.

[0174] It should be understood that the current instance ratio is the instance ratio (or PD ratio) included in the current scheduling policy, specifically the ratio between the number of full instances corresponding to the current scheduling policy and the number of incremental instances corresponding to the current scheduling policy.

[0175] Specifically, the cloud computing platform can determine whether the current instance ratio is the same as the instance ratio included in the above target scheduling strategy. If they are the same, it is determined that the current instance ratio is appropriate; if they are not the same, it is determined that the current instance ratio is inappropriate.

[0176] Optionally, the cloud computing platform may also determine whether the current instance ratio is appropriate by determining whether the current instance ratio is an optimal instance ratio.

[0177] If not, that is, when the current instance ratio is not appropriate, the cloud computing platform executes the following S604. If yes, that is, when the current instance ratio is appropriate, the cloud computing platform executes the following S605.

[0178] S604: Dynamically adjust instance ratio.

[0179] If not, that is, the current instance ratio is not appropriate, indicating that the current instance ratio is different from the instance ratio included in the target scheduling strategy. At this time, the cloud computing platform can replace the current instance ratio with the instance ratio included in the target scheduling strategy.

[0180] S605: Determine whether the current model is suitable.

[0181] It should be understood that the current model is the model information corresponding to the current scheduling strategy.

[0182] Optionally, the cloud computing platform may determine whether the current machine model is suitable by determining whether the machine model (or machine model information) corresponding to the current scheduling strategy is the most cost-effective.

[0183] If not, that is, when the current model is not suitable, the cloud computing platform executes the following S606. If yes, that is, when the current model is suitable, the cloud computing platform executes the following S607.

[0184] S606, Dynamically adjust the model.

[0185] If not, it means that the machine model information corresponding to the current scheduling strategy is different from the machine model information corresponding to the target scheduling strategy. At this time, the cloud computing platform can replace the current scheduling strategy with the target scheduling strategy.

[0186] It should be understood that after executing S606 (ie, dynamically adjusting the model), the cloud computing platform may execute the above S603, specifically, determining whether the current instance ratio is appropriate.

[0187] S607: Determine that the current scheduling strategy remains unchanged.

[0188] It should be understood that if the current instance ratio is appropriate and the current machine model is appropriate, it means that the instance ratio included in the current scheduling strategy is the same as the instance ratio included in the target scheduling strategy, and the machine model information corresponding to the current scheduling strategy is the same as the machine model information corresponding to the target scheduling strategy. It can also be understood that the current scheduling strategy is the same as the target scheduling strategy, in which case the current scheduling strategy can remain unchanged. The cloud computing platform can schedule the business container group corresponding to the current scheduling strategy to process the task to be inferred.

[0189] like Figure 7 As shown, the cloud computing platform provided by the embodiment of the present application may include a computing power evaluation module 701, a scheduler module 702 and a hardware resource pool 703, and the hardware resource pool 703 is deployed with a business container group, including a business container group 7031, a business container group 7032 and a business container group 7033.

[0190] Among them, the computing power evaluation module 701 is used to obtain task information of the task to be inferred and capability information of the business container group. The task information includes the task type of the task to be inferred. The business container group is used to be assigned to execute the task to be inferred, and the capability information is used to characterize the task processing performance of the business container group.

[0191] The computing power evaluation module 701 is also used to determine the target scheduling strategy from multiple scheduling strategies based on the task type of the task to be inferred and the capability information of the business container group corresponding to each scheduling strategy in the multiple scheduling strategies, and each scheduling strategy is used to indicate the scheduling of the business container group corresponding to each scheduling strategy to process the inference task corresponding to each scheduling strategy.

[0192] The scheduler module 702 is used to schedule the business container group corresponding to the target scheduling strategy to process the task to be inferred.

[0193] Specifically, the scheduler module 702 schedules the business container group corresponding to the target scheduling strategy to process the task to be inferred, which can also be understood as the scheduler module 702 scheduling the business container group corresponding to the target scheduling strategy to implement the separation deployment service.

[0194] In the embodiment of the present application, since the target scheduling strategy is determined by the computing power evaluation module from multiple scheduling strategies according to the task type of the task to be inferred and the capability information of the business container group corresponding to each scheduling strategy, the scheduler module can then schedule the business container group corresponding to the target scheduling strategy to process the task to be inferred. Therefore, the computing power evaluation module can dynamically adjust the scheduling strategy of the task to be inferred according to the task type and the task processing of the business container group, which can avoid the hardware resources of the related board being occupied for a long time. It can save hardware resources while improving the flexibility of task processing.

[0195] Optionally, the scheduler module in the embodiment of the present application can also be used to execute S204, S2031, step A and step B in the above embodiment. The specific process of the scheduler module executing these steps is the same or similar to the explanation of the above cloud computing platform executing these steps, which will not be repeated here.

[0196] Continue as Figure 7 As shown, the computing power assessment module 701 is deployed with a load analyzer module 7011 and a separate deployment simulator module 7012. The separate deployment simulator module 7012 is deployed with a scheduler simulator module 7012a and a model building module 7012b. The scheduler simulator module 7012a stores each of the above-mentioned multiple scheduling strategies. The model building module 7012b is used to simulate the reasoning environment of the LLM in each scheduling strategy. The reasoning environment of the LLM in each scheduling strategy is constructed based on the capability information of the business container group corresponding to each scheduling strategy.

[0197] Optionally, the task information of the task to be reasoned further includes the task length of the task to be reasoned.

[0198] The load analyzer module 7011 is used to determine the task load of the task to be reasoned according to the task type and the task length of the task to be reasoned.

[0199] The load analyzer module 7011 is also used to send a virtual task to the separate deployment simulator module 7012, and the task load of the virtual task is the same as the task load of the task to be inferred.

[0200] The model building module 7012b is used to obtain the reasoning performance of each of the above scheduling strategies in response to the input virtual tasks.

[0201] It should be understood that the scheduler module 702 is deployed with a task perception module 7021 for acquiring / perceiving the task to be inferred and the task information of the task to be inferred. The cloud computing platform is also deployed with a separate deployment task recording module 704 for saving / recording the task information of the task to be inferred. The load analyzer module 7011 can read the task information of the task to be inferred from the separate deployment task recording module 704 at regular intervals.

[0202] It can be understood that the task type of the above virtual task is the same as the task type of the task to be inferred, and the task length of the virtual task is the same as the task type of the task to be inferred.

[0203] Optionally, the separate deployment simulator module 7012 (specifically the model building module 7012b) can obtain capability information of the business container group from the hardware resource pool 703, thereby constructing the reasoning environment of the LLM in each scheduling strategy based on the capability information of the business container group corresponding to each scheduling strategy.

[0204] In the embodiment of the present application, since the model modeling module is used to simulate the reasoning environment of LLM in each scheduling strategy, and the reasoning environment of LLM in each scheduling strategy is constructed based on the capability information of the business container group corresponding to each scheduling strategy. That is, the model modeling module can fully and effectively simulate the real task processing level of each scheduling strategy. Therefore, when the model modeling module responds to the input virtual task with the same task load as the task to be inferred, the reasoning performance of each scheduling strategy can be accurately and effectively simulated, thereby improving the determination efficiency of the target scheduling strategy.

[0205] Continue as Figure 7 As shown, the computing power evaluation module 701 is also deployed with a result searcher module 7013.

[0206] Among them, the result searcher module 7013 is used to receive the reasoning performance of each of the above scheduling strategies sent by the separation deployment simulator module 7012.

[0207] The result searcher module 7013 is also used to determine a target scheduling strategy from multiple scheduling strategies based on the reasoning performance of each scheduling strategy.

[0208] In an embodiment of the present application, since the reasoning performance of each scheduling strategy is fully and effectively simulated by the model modeling module based on the reasoning environment of each scheduling strategy in LLM and the virtual task with the same task load as the task to be reasoned, the result searcher module determines the target scheduling strategy from multiple scheduling strategies according to the reasoning performance of each scheduling strategy, and can use a method that combines modeling with real business scenarios to determine the target scheduling strategy in real time, and then the scheduler module schedules the business container group corresponding to the target scheduling strategy to process the task to be reasoned, so as to realize separated deployment services.

[0209] Optionally, the result searcher module in the embodiment of the present application can also be used to perform step C and step D in the above embodiment. The specific process of the result searcher module performing these steps is the same or similar to the explanation of the above cloud computing platform performing these steps, which will not be repeated here.

[0210] In some embodiments, the above-mentioned task to be inferred may be a request to be inferred, and the computing power evaluation module in the cloud computing platform may sense the traffic change of the request to be inferred to instruct the scheduler module to schedule the service container group in the corresponding scheduling strategy. Figure 8As shown, the process may include S801-S804.

[0211] S801. The computing power evaluation module senses the traffic change of the request to be inferred.

[0212] The traffic change of the request to be inferred may be understood as request information of the request to be inferred, and the request information includes the request type of the request to be inferred and the request length of the request to be inferred.

[0213] S802: The computing power evaluation module senses the request load of the request to be inferred.

[0214] Among them, the specific process of the computing power evaluation module determining the request load of the request to be inferred according to the request type of the request to be inferred and the request length of the request to be inferred is the same or similar to the explanation of the above-mentioned cloud computing platform determining the task load of the task to be inferred according to the task type of the task to be inferred and the task length of the task to be inferred, and will not be repeated here.

[0215] S803. The computing power evaluation module sends the target scheduling strategy to the scheduler module.

[0216] The target scheduling strategy includes an instance ratio (or PD ratio).

[0217] S804: The scheduler module updates the instance ratio based on the target scheduling strategy.

[0218] Specifically, the scheduler module updates the current scheduling strategy to the target scheduling strategy, which may be to update the instance ratio included in the current scheduling strategy to the instance ratio (or PD ratio) included in the target scheduling strategy, so that the scheduler module can schedule the business container group corresponding to the target scheduling strategy to process the request to be inferred.

[0219] In other embodiments, the computing power evaluation module may also sense the change in the hardware resource pool to instruct the scheduler module to schedule the service container group in the corresponding scheduling strategy, such as Fig. 9 As shown, the process may include S901-S904.

[0220] S901. The computing power evaluation module senses changes in the hardware resource pool.

[0221] In combination with the description of the above embodiment, it should be understood that a service container group is deployed in the hardware resource pool, and the change of the hardware resource pool is specifically a change of the capability information of the service container group.

[0222] S902. The computing power evaluation module determines the reasoning performance of each scheduling strategy among multiple scheduling strategies.

[0223] Among them, different scheduling strategies include different model information.

[0224] S903. The computing power evaluation module sends the target scheduling strategy to the scheduler module.

[0225] S904. The scheduler module adjusts the model based on the target scheduling strategy.

[0226] Specifically, the scheduler module may adjust the machine model information included in the current scheduling strategy based on the machine model information included in the target scheduling strategy, so as to schedule the service container group corresponding to the target scheduling strategy to process the task to be inferred.

[0227] It is understandable that in order to implement the functions in the above embodiments, the client and the server include hardware structures and / or software modules corresponding to the execution of each function. It should be easy for those skilled in the art to realize that, in combination with the units and method steps of each example described in the embodiments disclosed in this application, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application scenario and design constraints of the technical solution.

[0228] Combined with the above Figures 1 to 9 , describes in detail the task processing method provided by this embodiment, and will now be combined with Fig.10 , describing the task processing device 100 provided according to this embodiment.

[0229] The present application also provides a task processing device 100 running on a cloud computing platform, such as Fig.10 As shown, it includes: an acquisition module 1001, a determination module 1002 and a scheduling module 1003.

[0230] Acquisition module 1001 is used to acquire task information of the task to be inferred and capability information of the business container group, wherein the task information includes the task type of the task to be inferred, the business container group is used to be assigned to execute the task to be inferred, and the capability information is used to characterize the task processing performance of the business container group.

[0231] Determination module 1002 is used to determine the target scheduling strategy from the multiple scheduling strategies based on the task type and the capability information of the business container group corresponding to each scheduling strategy in the multiple scheduling strategies, and each scheduling strategy is used to indicate the scheduling of the business container group corresponding to each scheduling strategy to process the reasoning task corresponding to each scheduling strategy.

[0232] The scheduling module 1003 is used to schedule the business container group corresponding to the target scheduling strategy to process the task to be inferred.

[0233] Optionally, the capability information of the service container group includes the remaining resource size of the service container group and / or the board type of the board to which the service container group belongs.

[0234] Optionally, a scheduling strategy includes an instance ratio, which is a ratio between the number of full instances corresponding to the scheduling strategy and the number of incremental instances corresponding to the scheduling strategy. One instance ratio corresponds to a business container group set.

[0235] The determination module 1002 is further used to determine the business container group included in the target business container group set as the business container group corresponding to the target scheduling policy, the target business container group set is the business container group set corresponding to the target instance ratio, and the target instance ratio is the instance ratio included in the target scheduling policy.

[0236] Optionally, different scheduling strategies include different instance ratios.

[0237] Optionally, the target business container group set includes at least one first business container group and at least one second business container group, the at least one first business container group is a full instance corresponding to the target scheduling policy, and the at least one second business container group is an incremental instance corresponding to the target scheduling policy.

[0238] The scheduling module 1003 is specifically configured to schedule the at least one first service container group to process the pre-filling phase of the task to be inferred, and schedule the at least one second service container group to process the decoding phase of the task to be inferred.

[0239] Optionally, the scheduling module 1003 is specifically configured to schedule the business container group corresponding to the target scheduling policy to process the task to be inferred when the current scheduling policy is different from the target scheduling policy.

[0240] Optionally, the scheduling module 1003 is specifically configured to schedule the service container group corresponding to the target scheduling strategy to process the task to be inferred when the current scheduling strategy does not meet the latency requirement.

[0241] Optionally, the task information of the task to be reasoned further includes the task length of the task to be reasoned.

[0242] The determination module 1002 is specifically configured to determine the task load of the task to be reasoned according to the task type of the task to be reasoned and the task length of the task to be reasoned.

[0243] The determination module 1002 is further specifically configured to determine the reasoning performance of each scheduling strategy according to the task load of the task to be inferred and the capability information of the service container group corresponding to each scheduling strategy.

[0244] The determination module 1002 is further specifically configured to determine the target scheduling strategy from the multiple scheduling strategies according to the reasoning performance of each scheduling strategy.

[0245] Optionally, the determination module 1002 is further specifically used to determine the task load of the task to be inferred from the information load relationship according to the task type and the task length, and the information load relationship is used to indicate: the task type and task length corresponding to one or more task loads.

[0246] Optionally, the determination module 1002 is further specifically configured to determine, among the multiple scheduling strategies, a scheduling strategy whose reasoning performance is greater than or equal to a first performance threshold, as the target scheduling strategy.

[0247] Optionally, determination module 1002 is also specifically used to determine the scheduling strategy whose corresponding number of business container groups in at least two scheduling strategies is less than or equal to a quantity threshold as the target scheduling strategy, and the at least two scheduling strategies are scheduling strategies whose reasoning performance is greater than or equal to a second performance threshold among the multiple scheduling strategies, and the second performance threshold is less than the first performance threshold.

[0248] Among them, the acquisition module 1001, the determination module 1002 and the scheduling module 1003 can all be implemented by software, or can be implemented by hardware. Exemplarily, the implementation of the determination module 1002 is described below by taking the determination module 1002 as an example. Similarly, the implementation of the acquisition module 1001 and the scheduling module 1003 can refer to the implementation of the determination module 1002.

[0249] As an example of a software functional unit, the module 1002 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Furthermore, the computing instance may be one or more. For example, the module 1002 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one data center or multiple data centers with similar geographical locations. Generally, a region may include multiple AZs.

[0250] Similarly, multiple hosts / virtual machines / containers used to run the code can be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Usually, a VPC is set up in a region. For cross-region communication between two VPCs in the same region and between VPCs in different regions, a communication gateway needs to be set up in each VPC to achieve interconnection between VPCs through the communication gateway.

[0251] As an example of a hardware functional unit, the determination module 1002 may include at least one computing device, such as a server, etc. Alternatively, the determination module 1002 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0252] The multiple computing devices included in the determination module 1002 may be distributed in the same region or in different regions. The multiple computing devices included in the determination module 1002 may be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the determination module 1002 may be distributed in the same VPC or in multiple VPCs. The multiple computing devices may be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0253] It should be noted that, in other embodiments, the determination module 1002 can be used to execute any step in the task processing method, the acquisition module 1001 can be used to execute any step in the task processing method, and the scheduling module 1003 can be used to execute any step in the task processing method. The steps that the determination module 1002, the acquisition module 1001, and the scheduling module 1003 are responsible for implementing can be specified as needed, and the entire functions of the task processing device 100 are realized by respectively implementing different steps in the task processing method through the determination module 1002, the acquisition module 1001, and the scheduling module 1003.

[0254] The present application also provides a computing device 1100. Fig.11As shown, the computing device 1100 includes: a bus 1102, a processor 1104, a memory 1106, and a communication interface 1108. The processor 1104, the memory 1106, and the communication interface 1108 communicate through the bus 1102. The computing device 1100 can be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in the computing device 1100.

[0255] The bus 1102 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Fig.11 The bus 1102 may include a path for transmitting information between various components of the computing device 1100 (eg, the memory 1106, the processor 1104, and the communication interface 1108).

[0256] The processor 1104 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0257] The memory 1106 may include a volatile memory, such as a random access memory (RAM). The processor 1104 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD).

[0258] The memory 1106 stores executable program codes, and the processor 1104 executes the executable program codes to respectively implement the functions of the aforementioned determination module 1002, acquisition module 1001, and scheduling module 1003, thereby implementing the task processing method. That is, the memory 1106 stores instructions for executing the task processing method.

[0259] The communication interface 1108 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 1100 and other devices or a communication network.

[0260] The embodiment of the present application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.

[0261] like Fig.12 As shown, the computing device cluster includes at least one computing device 1100. The memory 1106 in one or more computing devices 1100 in the computing device cluster may store the same instructions for executing the task processing method.

[0262] In some possible implementations, the memory 1106 of one or more computing devices 1100 in the computing device cluster may also store partial instructions for executing the task processing method. In other words, the combination of one or more computing devices 1100 may jointly execute instructions for executing the task processing method.

[0263] It should be noted that the memory 1106 in different computing devices 1100 in the computing device cluster can store different instructions, which are respectively used to execute part of the functions of the task processing device. That is, the instructions stored in the memory 1106 in different computing devices 1100 can implement the functions of one or more modules in the determination module 1002, the acquisition module 1001 and the scheduling module 1003.

[0264] In some possible implementations, one or more computing devices in the computing device cluster may be connected via a network, which may be a wide area network or a local area network. Fig.13 A possible implementation is shown. Fig.13 As shown, two computing devices 1100A and 1100B are connected via a network. Specifically, the network is connected via a communication interface in each computing device. In this type of possible implementation, the memory 1106 in the computing device 1100A stores instructions for executing the functions of the acquisition module 1001 and the determination module 1002. At the same time, the memory 1106 in the computing device 1100B stores instructions for executing the functions of the scheduling module 1003.

[0265] Fig.13The connection mode between the computing device clusters shown may be considered to be a process that requires a large number of scheduling service container groups in the task processing method provided in this application. Therefore, it is considered to hand over the functions implemented by the scheduling module 1003 to the computing device 1100B for execution.

[0266] It should be understood that Fig.13 The functions of the computing device 1100A shown in FIG. 1100A may also be completed by multiple computing devices 1100. Similarly, the functions of the computing device 1100B may also be completed by multiple computing devices 1100.

[0267] The present application embodiment also provides another computing device cluster. The connection relationship between the computing devices in the computing device cluster can be similar to that of Fig.12 and Fig.13 The connection mode of the computing device cluster is different in that the memory 1106 in one or more computing devices 1100 in the computing device cluster may store the same instructions for executing the task processing method.

[0268] In some possible implementations, the memory 1106 of one or more computing devices 1100 in the computing device cluster may also store partial instructions for executing the task processing method. In other words, the combination of one or more computing devices 1100 may jointly execute instructions for executing the task processing method.

[0269] The embodiment of the present application also provides a computer program product including instructions. The computer program product may be software or a program product including instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device executes the task processing method.

[0270] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center that contains one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk). The computer-readable storage medium includes instructions that instruct the computing device to execute the task processing method.

[0271] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.

Claims

1. A task processing method running on a cloud computing platform, characterized in that: The method comprises: Acquire task information of the task to be inferred and capability information of the business container group, wherein the task information includes the task type of the task to be inferred, the business container group is used to be assigned to execute the task to be inferred, and the capability information is used to characterize the task processing performance of the business container group; Determine a target scheduling strategy from the multiple scheduling strategies according to the task type and the capability information of the business container group corresponding to each scheduling strategy in the multiple scheduling strategies, wherein each scheduling strategy is used to instruct the business container group corresponding to each scheduling strategy to process the reasoning task corresponding to each scheduling strategy, and a scheduling strategy includes an instance ratio, wherein the instance ratio is the ratio between the number of full instances corresponding to the scheduling strategy and the number of incremental instances corresponding to the scheduling strategy, and an instance ratio corresponds to a business container group set; Schedule the business container group corresponding to the target scheduling strategy to process the task to be inferred, the business container group corresponding to the target scheduling strategy is the business container group included in the target business container group set, the target business container group set is the business container group set corresponding to the target instance ratio, and the target instance ratio is the instance ratio included in the target scheduling strategy.

2. The method according to claim 1, characterized in that The capability information of the service container group includes the remaining resource size of the service container group and / or the board type of the board to which the service container group belongs.

3. The method according to claim 1, characterized in that The target business container group set includes at least one first business container group and at least one second business container group, the at least one first business container group is a full instance corresponding to the target scheduling strategy, the at least one second business container group is an incremental instance corresponding to the target scheduling strategy, and the scheduling of the business container group corresponding to the target scheduling strategy to process the task to be inferred includes: The at least one first service container group is scheduled to process a pre-filling phase of the task to be inferred, and the at least one second service container group is scheduled to process a decoding phase of the task to be inferred.

4. The method according to claim 1, characterized in that: The scheduling of the business container group corresponding to the target scheduling strategy to process the task to be inferred includes: When the current scheduling strategy is different from the target scheduling strategy, the service container group corresponding to the target scheduling strategy is scheduled to process the task to be inferred.

5. The method according to claim 1, characterized in that The scheduling of the business container group corresponding to the target scheduling strategy to process the task to be inferred includes: When the current scheduling strategy does not meet the latency requirement, the service container group corresponding to the target scheduling strategy is scheduled to process the task to be inferred.

6. The method according to any one of claims 1 to 5, characterized in that The task information of the task to be inferred also includes the task length of the task to be inferred, and determining the target scheduling strategy from the multiple scheduling strategies according to the task type and the capability information of the business container group corresponding to each scheduling strategy in the multiple scheduling strategies includes: Determining a task load of the task to be reasoned according to a task type of the task to be reasoned and a task length of the task to be reasoned; Determining the reasoning performance of each scheduling strategy according to the task load of the task to be inferred and the capability information of the service container group corresponding to each scheduling strategy; The target scheduling strategy is determined from the multiple scheduling strategies according to the reasoning performance of each scheduling strategy.

7. The method according to claim 6, characterized in that The determining the task load of the task to be reasoned according to the task type of the task to be reasoned and the task length of the task to be reasoned includes: The task load of the task to be inferred is determined from an information load relationship according to the task type and the task length, wherein the information load relationship is used to indicate: a task type and a task length corresponding to one or more task loads.

8. The method according to claim 6, characterized in that The step of determining the target scheduling strategy from the plurality of scheduling strategies according to the reasoning performance of each scheduling strategy comprises: A scheduling strategy whose reasoning performance is greater than or equal to a first performance threshold among the multiple scheduling strategies is determined as the target scheduling strategy.

9. The method according to claim 6, characterized in that The step of determining the target scheduling strategy from the plurality of scheduling strategies according to the reasoning performance of each scheduling strategy comprises: The scheduling strategy whose corresponding number of business container groups in at least two scheduling strategies is less than or equal to the number threshold is determined as the target scheduling strategy, and the at least two scheduling strategies are scheduling strategies whose reasoning performance is greater than or equal to the second performance threshold among the multiple scheduling strategies, and the second performance threshold is less than the first performance threshold.

10. A cloud computing platform, characterized in that: The cloud computing platform is deployed with a computing power evaluation module, a scheduler module and a hardware resource pool, and the hardware resource pool is deployed with a business container group; The computing power evaluation module is used to obtain task information of the task to be inferred and capability information of the business container group, wherein the task information includes the task type of the task to be inferred, the business container group is used to be assigned to perform the task to be inferred, and the capability information is used to characterize the task processing performance of the business container group; The computing power evaluation module is further used to determine a target scheduling strategy from the multiple scheduling strategies according to the task type and the capability information of the business container group corresponding to each scheduling strategy in the multiple scheduling strategies, wherein each scheduling strategy is used to instruct the business container group corresponding to each scheduling strategy to process the reasoning task corresponding to each scheduling strategy, and a scheduling strategy includes an instance ratio, wherein the instance ratio is the ratio between the number of full instances corresponding to the scheduling strategy and the number of incremental instances corresponding to the scheduling strategy, and an instance ratio corresponds to a set of business container groups; The scheduler module is used to schedule the business container group corresponding to the target scheduling strategy to process the task to be inferred, the business container group corresponding to the target scheduling strategy is the business container group included in the target business container group set, the target business container group set is the business container group set corresponding to the target instance ratio, and the target instance ratio is the instance ratio included in the target scheduling strategy.

11. The cloud computing platform according to claim 10, characterized in that: The task information of the task to be inferred also includes the task length of the task to be inferred, the computing power evaluation module is deployed with a load analyzer module and a separate deployment simulator module, the separate deployment simulator module is deployed with a scheduler simulator module and a model building module, the scheduler simulator module stores each scheduling strategy, the model building module is used to simulate the reasoning environment of the large language model LLM in each scheduling strategy, and the reasoning environment of the LLM in each scheduling strategy is constructed based on the capability information of the business container group corresponding to each scheduling strategy; The load analyzer module is used to determine the task load of the task to be reasoned according to the task type of the task to be reasoned and the task length of the task to be reasoned; The load analyzer module is further used to send a virtual task to the separate deployment simulator module, wherein the task load of the virtual task is the same as the task load of the task to be inferred; The model building module is used to obtain the reasoning performance of each scheduling strategy in response to the input virtual task.

12. The cloud computing platform according to claim 11, characterized in that: The computing power evaluation module is also deployed with a result searcher module; The result searcher module is used to receive the reasoning performance of each scheduling strategy sent by the separate deployment simulator module; The result searcher module is further used to determine the target scheduling strategy from the multiple scheduling strategies according to the reasoning performance of each scheduling strategy.

13. A task processing device running on a cloud computing platform, characterized in that: The device includes an acquisition module, a determination module and a scheduling module; The acquisition module is used to acquire task information of the task to be inferred and capability information of the business container group, wherein the task information includes the task type of the task to be inferred, the business container group is used to be assigned to execute the task to be inferred, and the capability information is used to characterize the task processing performance of the business container group; The determination module is used to determine a target scheduling strategy from the multiple scheduling strategies according to the task type and the capability information of the business container group corresponding to each scheduling strategy in the multiple scheduling strategies, wherein each scheduling strategy is used to instruct the business container group corresponding to each scheduling strategy to process the reasoning task corresponding to each scheduling strategy, and a scheduling strategy includes an instance ratio, wherein the instance ratio is the ratio between the number of full instances corresponding to the scheduling strategy and the number of incremental instances corresponding to the scheduling strategy, and an instance ratio corresponds to a set of business container groups; The scheduling module is used to schedule the business container group corresponding to the target scheduling strategy to process the task to be inferred, the business container group corresponding to the target scheduling strategy is the business container group included in the target business container group set, the target business container group set is the business container group set corresponding to the target instance ratio, and the target instance ratio is the instance ratio included in the target scheduling strategy.

14. The device according to claim 13, characterized in that The capability information of the service container group includes the remaining resource size of the service container group and / or the board type of the board to which the service container group belongs.

15. The device according to claim 13, characterized in that The target business container group set includes at least one first business container group and at least one second business container group, wherein the at least one first business container group is a full instance corresponding to the target scheduling policy, and the at least one second business container group is an incremental instance corresponding to the target scheduling policy; The scheduling module is specifically used to schedule the at least one first service container group to process the pre-filling stage of the task to be inferred, and schedule the at least one second service container group to process the decoding stage of the task to be inferred.

16. The device according to claim 13, characterized in that The scheduling module is specifically used to schedule the business container group corresponding to the target scheduling strategy to process the task to be inferred when the current scheduling strategy is different from the target scheduling strategy.

17. The device according to claim 13, characterized in that The scheduling module is specifically used to schedule the business container group corresponding to the target scheduling strategy to process the task to be inferred when the current scheduling strategy does not meet the delay requirement.

18. The device according to any one of claims 13 to 17, characterized in that The task information of the task to be inferred also includes the task length of the task to be inferred; The determination module is specifically used to determine the task load of the task to be reasoned according to the task type of the task to be reasoned and the task length of the task to be reasoned; The determination module is further specifically used to determine the reasoning performance of each scheduling strategy according to the task load of the task to be inferred and the capability information of the business container group corresponding to each scheduling strategy; The determination module is further specifically configured to determine the target scheduling strategy from the multiple scheduling strategies according to the reasoning performance of each scheduling strategy.

19. The device according to claim 18, characterized in that The determination module is further specifically used to determine the task load of the task to be inferred from the information load relationship according to the task type and the task length, and the information load relationship is used to indicate: the task type and task length corresponding to one or more task loads.

20. The device according to claim 18, characterized in that The determination module is further specifically configured to determine, among the multiple scheduling strategies, a scheduling strategy whose reasoning performance is greater than or equal to a first performance threshold, as the target scheduling strategy.

21. The device according to claim 18, characterized in that The determination module is also specifically used to determine the scheduling strategy whose corresponding number of business container groups in at least two scheduling strategies is less than or equal to a quantity threshold as the target scheduling strategy, and the at least two scheduling strategies are scheduling strategies whose reasoning performance is greater than or equal to a second performance threshold among the multiple scheduling strategies, and the second performance threshold is less than the first performance threshold.

22. A computing device cluster, characterized in that: comprising at least one computing device, each computing device comprising a processor and a memory; The processor of the at least one computing device is used to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the task processing method according to any one of claims 1 to 9.

23. A computer program product comprising instructions, characterized in that When the instruction is executed by the computing device cluster, the computing device cluster executes the task processing method according to any one of claims 1 to 9.

24. A computer-readable storage medium, characterized in that: The method comprises computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster executes the task processing method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Task scheduling method and related equipment

    CN113568721A

  • Computing equipment scheduling method and device, nonvolatile storage medium and electronic equipment

    CN117311973A

  • Application scheduling method, cloud service platform and related equipment

    CN117640770A