Task scheduling method and device, central server and storage medium
By dynamically scheduling tasks to execution servers where resources are not evenly utilized, the problem of unbalanced computing resources and I/O resources on large language model servers is solved, improving system performance and user experience.
Patent Information
- Application Number
- CN202510682708.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-09-26
AI Technical Summary
In the existing technology, the computing resources and I/O resources of large language model servers are used unbalancedly, resulting in reduced computing performance and response speed and serious waste of resources.
By obtaining the difference between the number of text units to be pre-filled and the number of target processing results, combined with the remaining storage space and resource usage of the execution server, tasks are dynamically scheduled to the target execution server where resources are not evenly used, so as to balance the use of computing resources and storage and reading resources.
It achieves balanced use of server resources, improves system processing speed and response efficiency, reduces system latency and resource waste, and enhances user experience.
Smart Images

Figure CN120704859A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of large language models, and in particular to a task scheduling method, device, central server and storage medium. Background Art
[0002] In the field of large language model technology, a trained LLM (Large Language Model) can be used to respond to text input by the user. Specifically, after the text is input into the LLM, the LLM will obtain multiple tokens (discrete text units) based on the text, prefill these multiple tokens through a large amount of calculations, and determine the KV (Key Value) corresponding to the text; decode based on the KV to generate multiple tokens. The text composed of the generated tokens is the response to the text. If the tokens generated by Decoding indicate the characters "t", "o", "k", "e", and "n" in sequence, the response to the text is "token". And each time data is calculated during the Decoding process, the calculated data can be stored in KVCache (key-value cache space); when these data are needed later, the stored data can be read from KVCache.
[0003] In related technologies, after a user enters text through an online platform, the platform sends the processing task for that text to the server with the fewest concurrent tasks. After that server receives a response to the text based on its deployed LLM, it sends the response to the online platform. The online platform can then display the response on its interface.
[0004] However, a large amount of calculation is required during the Prefill process, which consumes more computing resources; data must be frequently stored / read in the KVCache during the Decode process, which consumes more I / O (Input / Output) resources. When scheduling tasks based on relevant technologies, for example, if server 1 processes two tasks in parallel, and both tasks are in the Prefill stage; server 2 processes three tasks in parallel, and both tasks are in the Decode stage, and server 1 currently processes the least number of tasks in parallel, the online platform will assign the task to server 1. After server 1 receives the text, it needs to prefill the three tasks in parallel, resulting in low utilization of server 1's I / O resources and low utilization of server 2's computing resources. It can be seen that based on relevant technologies, there will be an imbalance in the use of server computing resources and I / O resources. Summary of the Invention
[0005] The purpose of the embodiments of the present invention is to provide a task scheduling method, device, central server, and storage medium to balance the use of server computing resources and storage resources. The specific technical solution is as follows:
[0006] In a first aspect of the present invention, a task scheduling method is provided, the method comprising: obtaining the number of text units to be pre-filled as a first number; wherein the text units to be pre-filled are: text units that have not been pre-filled in the text units of the text to be processed; a text unit of a text represents the encoding result of the characters in the text; predicting the number of text units of the target processing result as a predicted number; wherein the target processing result is: the result obtained by pre-filling and decoding the text units of the text to be processed based on a large language model; calculating the difference between the predicted number and the number of generated text units to obtain a second number; wherein the generated text units are: text units that have been generated by decoding in the text units of the target processing result; determining from each execution server an execution server whose current remaining storage space satisfies the requirements for processing the text to be processed using the large language model as an alternative execution server; for each alternative Select an execution server, and based on the first number and the second number, predict how long it will take for the alternative execution server to obtain the processing result of the current specified text for the current processing stage using the large language model if the text to be processed is sent to the alternative execution server; wherein the current specified text is determined from the text received from the alternative execution server; the current processing stage of any specified text represents a pre-filling stage or a decoding stage; the time required for an execution server indicates the amount of target resources currently available to the execution server; the target resources are: the type of resources required to obtain the processing result of the text to be processed for the current stage; based on the time required for each alternative execution server, determine a target execution server from each alternative execution server, and send the text to be processed to the target execution server, so that the target execution server processes the text to be processed based on the large language model using the target resources currently available to the target execution server.
[0007] In a second aspect of the present invention, a task scheduling device is provided, comprising: an acquisition module for acquiring the number of text units to be pre-filled as a first number; wherein the text units to be pre-filled are text units that have not been pre-filled in the text units of the text to be processed; a text unit of a text represents an encoding result of a character in the text; a prediction module for predicting the number of text units of a target processing result as a predicted number; wherein the target processing result is obtained by pre-filling and decoding the text units of the text to be processed based on a large language model; a number calculation module for calculating the difference between the predicted number and the number of generated text units to obtain a second number; wherein the generated text units are text units that have been generated by decoding in the text units of the target processing result; a determination module for determining, from each execution server, an execution server whose current remaining storage space satisfies the requirements for processing the text to be processed using the large language model, as an alternative execution server ; A duration calculation module is used to predict, for each alternative execution server, based on the first number and the second number, the duration required for the alternative execution server to obtain the processing result of the current specified text for the current processing stage using the large language model if the text to be processed is sent to the alternative execution server; wherein the current specified text is determined from the text received from the alternative execution server; the current processing stage of any specified text represents a pre-filling stage or a decoding stage; the duration required for an execution server indicates the amount of target resources currently available to the execution server; the target resources are: the type of resources required to obtain the processing result of the text to be processed for the current stage; a sending module is used to determine a target execution server from each alternative execution server based on the duration required by each alternative execution server, and send the text to be processed to the target execution server, so that the target execution server processes the text to be processed based on the large language model using the target resources currently available to the target execution server.
[0008] In the third aspect of the implementation of the present invention, a central server is provided, including a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; the memory is used to store computer programs; and the processor is used to implement any task scheduling method described in the first aspect when executing the program stored in the memory.
[0009] In another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, which, when executed by a processor, implements the task scheduling method described in any one of the first aspects. In another aspect of the present invention, a computer program product comprising instructions is provided, which, when executed on a computer, causes the computer to execute the task scheduling method described in any one of the first aspects.
[0010] An embodiment of the present invention provides a task scheduling method, which obtains a first number of text units to be pre-filled; the text units to be pre-filled are text units that have not been pre-filled in the text units of the text to be processed; a text unit of a text represents the encoding result of the characters in the text; a predicted number of text units of a target processing result is predicted; the target processing result is obtained by pre-filling and decoding the text units of the text to be processed based on a large language model; a second number is obtained by calculating the difference between the predicted number and the number of generated text units; the generated text units are text units that have been generated by decoding in the text units of the target processing result; from each execution server, an alternative execution server is determined whose current remaining storage space satisfies the requirements for processing the text to be processed using the large language model; for each alternative execution server, based on the first number and the second number, number, predicting that if the text to be processed is sent to the alternative execution server, the alternative execution server will use the large language model to obtain the time required for the processing result of the current specified text for the current processing stage; the current specified text is determined from the text received from the alternative execution server; the current processing stage of any specified text represents a pre-filling stage or a decoding stage; the time required for an execution server indicates the amount of target resources currently available to the execution server; the target resources are the types of resources required to obtain the processing result of the text to be processed for the current stage; based on the time required for each alternative execution server, a target execution server is determined from each alternative execution server, and the text to be processed is sent to the target execution server, so that the target execution server uses the target resources currently available to the target execution server to process the text to be processed based on the large language model.
[0011] Based on the above processing, the text units to be pre-filled are the text units in the text to be processed that have not been pre-filled. The first number of text units to be pre-filled can indicate the amount of computing resources required to pre-fill the text units to be pre-filled. The second number is the difference between the predicted number of text units for the target processing result and the number of generated text units, that is, the predicted number of text units that still need to be generated through decoding. The second number can indicate the amount of storage and read resources required to obtain the target processing result. Since pre-filling text based on a large language model consumes more computing resources, decoding based on the pre-filled processing result consumes more storage and read resources. When processing text, the resources allocated to the text are positively correlated with the processing speed of the text. Based on the first and second numbers, the time required for each alternative execution server to obtain the processing result of the current specified text for the current processing stage using the large language model is calculated. The time required for an execution server can indicate the amount of target resources currently available to the execution server. The shorter the time required by an execution server, the more target resources it currently has available. This increases the probability that the target resources of that execution server are underutilized, meaning that its computing and storage resources are not evenly utilized. Therefore, based on the time required by each candidate execution server, the target execution server with unbalanced resource utilization can be identified from among the candidate execution servers. The target execution server then sends the pending document to be processed, which uses its currently available target resources to process it. This fully utilizes the target execution server's unbalanced resources, ensuring balanced utilization of its computing and storage resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for describing the embodiments or the prior art.
[0013] Figure 1 A first flow chart of a task scheduling method provided by an embodiment of the present invention;
[0014] Figure 2 A second flow chart of the task scheduling method provided in an embodiment of the present invention;
[0015] Figure 3 A structural diagram of a task scheduling device provided by an embodiment of the present invention;
[0016] Figure 4 A structural diagram of a central server provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0017] The technical solutions in the embodiments of the present invention will be described below with reference to the accompanying drawings in the embodiments of the present invention.
[0018] In the field of large language model technology, a trained LLM can process user-entered text and obtain processing results for that text. Specifically, existing LLMs, such as GPT-4 (Generative Pre-trained Transformer), obtain multiple tokens (i.e., text units in subsequent embodiments) based on the text when processing it. These tokens are prefilled through extensive computation to determine the key-value (KV) corresponding to the text. The KV is then decoded to generate multiple tokens. The text obtained based on each generated token is the processing result for that text. During the decoding process, each calculated data can be stored in a KV cache. When this data is needed later, the stored data can be read from the KV cache. In related technologies, after a user enters text through an online platform, the online platform sends the processing task for that text to the server with the fewest concurrent processing tasks. This server, based on its deployed LLM, obtains the processing result for the text and sends it to the online platform. The online platform can then display the processing result on a display interface. It is understandable that the KVCache occupied when prefilling a text is positively correlated with the length of the tokens obtained by the LLM based on the text; the KVCache occupied when decoding based on the KV obtained by prefill is positively correlated with the length of the tokens of the processing result. In practical scenarios, large model deployment frameworks such as VLLM (Virtual Large Language Model) can be used to process text through dynamic batch processing. However, when a server is currently processing a long text with a large number of corresponding tokens, although the server is currently processing fewer tasks in parallel, the server's KVCache is occupied. Based on relevant technologies, it is possible to continue to assign tasks to this server. If the server receives a new task and is unable to allocate sufficient KVCache for the new task due to insufficient remaining storage space, when the KVCache is full, the server will unload the relevant data calculated based on the new task stored in the KVCache and then reload the unloaded relevant data into other storage spaces, such as the server's memory. After the relevant data is unloaded, the server will not process new tasks, which means that new tasks will face the problem of being unable to process. This problem is particularly common when the server processes long or complex text.Moreover, the process of unloading and reloading data will occupy the server's bandwidth, and will also generate additional overhead and server resource loss, resulting in increased system latency, longer response time for outputting the processing results of the user-input text, and reduced throughput. In other words, the system's computing performance loss is large, and the overall performance is reduced, which in turn leads to a reduction in the processing speed of tasks and a poor user experience. In addition, since the Prefill process consumes more computing resources and the Decode process consumes more I / O resources. When scheduling tasks based on related technologies, there may be an imbalance in the use of the server's computing resources and I / O resources, that is, the server's resources are wasted, which in turn leads to a reduction in the overall processing speed of the system, that is, the system's response speed is not high, and the expansion of the system is restricted. It can be seen that for the task scheduling system, the existing task scheduling method will bring significant performance loss and resource waste to the task scheduling system.
[0019] In order to solve the above problems, the present invention provides a task scheduling method, which can be applied to a central server. The central server can interact with the execution server, and a large language model for processing text is deployed in the execution server. For example, the central server can send text to an execution server, send instructions to the execution server, obtain the processing progress of the text currently processed by the execution server, obtain the remaining storage space of the KVCache of the execution server, etc. In actual application scenarios, the large language model deployed in each execution server can be the same large language model; the hardware devices of each execution server can be different. Based on the task scheduling method provided by the present invention, the central server can determine the target execution server whose resources are not used evenly from each execution server, and the target execution server can use its currently available target resources to process the text to be processed, so as to balance the computing resources and storage resources of the target execution server.
[0020] See also Figure 1 , Figure 1 This is a first flow chart of a task scheduling method provided by an embodiment of the present invention. The method may include the following steps:
[0021] S101: Obtain the number of text units to be pre-filled as a first number.
[0022] The text unit to be pre-filled is: a text unit that has not been pre-filled among the text units of the text to be processed; a text unit of a text represents the encoding result of the characters in the text.
[0023] S102: Predicting the number of text units of the target processing result as the predicted number.
[0024] The target processing result is obtained by pre-filling and decoding the text units of the text to be processed based on the large language model.
[0025] S103: Calculate the difference between the predicted number and the number of generated text units to obtain a second number.
[0026] The generated text unit is: a text unit generated through decoding processing among the text units of the target processing result.
[0027] S104: Determine, from among the execution servers, an execution server whose current remaining storage space satisfies the requirement for processing the text to be processed using the large language model, and select the execution server as a candidate execution server.
[0028] S105: For each candidate execution server, based on the first number and the second number, predict the time required for the candidate execution server to obtain the processing result of the current specified text for the current processing stage using the large language model if the text to be processed is sent to the candidate execution server.
[0029] Among them, the current specified text is determined from the text received from the alternative execution server; the current processing stage of any specified text represents the pre-filling stage or the decoding stage; the duration required for an execution server indicates the amount of target resources currently available to the execution server; the target resources are: the type of resources required to obtain the processing results of the text to be processed for the current stage.
[0030] S106: Determine a target execution server from the candidate execution servers based on the time required by each candidate execution server, and send the to-be-processed text to the target execution server, so that the target execution server processes the to-be-processed text based on the large language model and using the target resources currently available to the target execution server.
[0031] Based on the task scheduling method provided by an embodiment of the present invention, the text units to be pre-filled are those text units in the text to be processed that have not been pre-filled. The first number of text units to be pre-filled can indicate the amount of computing resources required to pre-fill the text units. The second number is the difference between the predicted number of text units for the target processing result and the number of generated text units, i.e., the predicted number of text units that still need to be generated through decoding. The second number can indicate the amount of storage and read resources required to obtain the target processing result. Because pre-filling text based on a large language model consumes a large amount of computing resources, decoding based on the pre-filled processing result consumes a large amount of storage and read resources. Furthermore, when processing text, the resources allocated to the text are positively correlated with the processing speed of the text. Based on the first and second numbers, the time required for each candidate execution server to obtain the processing result for the current specified text for the current processing stage using the large language model is calculated. The time required by each execution server can indicate the amount of target resources currently available to that execution server. The shorter the time required by an execution server, the more target resources it currently has available. This increases the probability that the target resources of that execution server are underutilized, meaning that its computing and storage resources are not evenly utilized. Therefore, based on the time required by each candidate execution server, the target execution server with unbalanced resource utilization can be identified from among the candidate execution servers. The target execution server then sends the pending document to be processed, which uses its currently available target resources to process it. This fully utilizes the target execution server's unbalanced resources, ensuring balanced utilization of its computing and storage resources.
[0032] In step S101, in a real-world scenario, a user can enter text for which they wish to obtain a corresponding processing result through an online platform. This text is referred to as the text to be processed. The processing result for a text is also text. For example, for the text "Where is the capital of country A?", the processing result for this text may be "City B."
[0033] Alternatively, the central server can retrieve text that has not yet received a corresponding processing result from the execution server as the text to be processed. For example, if a failure occurs during the execution server 1's operation while processing text A using a large language model, resulting in a slow processing speed for text A, the execution server 1 can send text A to the central server, which will then schedule text A, thereby reducing the time required to obtain the processing result for text A and improving the efficiency of obtaining the processing result for text A.
[0034] If the execution server 1 has more computing resources but fewer storage and reading resources, in order to improve the efficiency of obtaining the processing results of text A, the KV corresponding to text A can be obtained through the execution server 1; then the central server determines the execution server 2 with more storage and reading resources, and the execution server 2 performs decoding processing based on the KV corresponding to text A.
[0035] It is understandable that migrating the text to be processed and the related data obtained by processing the text to be processed using a large language model will incur additional overhead; and multiple scheduling of the text to be processed according to the current processing stage of the text to be processed will also increase the operating pressure of the central server. Therefore, in actual application scenarios, the text to be processed can be determined in combination with the performance of the central server and the operating requirements of the system. For example, when the performance of the central server is poor and the system's additional overhead is required to be low, the central server can schedule the text to be processed, and when the execution server fails, the text currently being processed by the failed execution server can be scheduled. If the performance of the central server is good and the system is required to have a high response efficiency in obtaining the processing results of the user-entered text, the central server can schedule the text to be processed multiple times. For example, after obtaining the processing results of the text to be processed for the current processing stage, the text to be processed can be scheduled again to improve the efficiency of obtaining the processing results of the text to be processed.
[0036] A text is composed of multiple characters. The characters in the text are encoded according to the preset correspondence between characters and text units, and the encoded result is the text unit of the text. If the text to be processed is obtained from the online platform, the number of text units obtained by encoding the characters in the text to be processed is the first number of text units to be pre-filled. If the text to be processed is obtained from the execution server, the central server can determine the first number according to the processing progress of the text to be processed. For example, the text to be processed is parsed based on the prompt vector to obtain the first number of text units to be pre-filled. The processing progress can indicate the proportion of completed processing in all processing that needs to be performed in the process of obtaining the target processing result of the text to be processed. For example, when obtaining the processing result of a text, it is necessary to pre-fill 10 text units using a large language model and generate 20 text units through decoding; if 5 text units have been pre-filled, the processing progress of the text can indicate that the pre-filling stage has been completed by 50% and the decoding stage has not yet been performed.
[0037] The larger the first number, the more text units to be pre-filled. That is, when a large language model deployed on an execution server processes a pending text, the more text units that need to be pre-filled, and the greater the amount of computing resources required for the pre-filling process. Therefore, the first number can indicate the amount of computing resources required for the pre-filling process of the pending text.
[0038] For step S102 and step S103, the target processing result is the processing result obtained by pre-filling and decoding the text units of the text to be processed based on the large language model. Since the text to be processed has not yet been sent to any execution server, the target processing result has not yet been obtained. At this time, the central server can predict the predicted number of text units of the target processing result. For example, the central server can use a specified number as the predicted number. The specific way for the central server to obtain the predicted number is described in detail in the subsequent embodiments and will not be repeated here. The difference between the predicted number and the number of generated text units, that is, when the text to be processed is decoded by the large language model deployed in an execution server, the second number of text units (which can be called text units to be generated) that are predicted to be generated. The larger the second number, the more text units to be generated, and the greater the amount of storage and reading resources required for decoding. Therefore, the second number can indicate: the amount of storage and reading resources required for decoding the text to be processed.
[0039] With respect to step S104, the total storage space of an execution server is determined by the hardware of the execution server; the trained large language model deployed in the execution server will occupy the storage space of the execution server; and the data generated during the operation of the large language model will also occupy the storage space of the execution server. Therefore, the total storage space size of an execution server minus the storage space to be occupied is the current remaining storage space size. For example, if the total storage space size of an execution server is 70GB (Gigabyte), the trained large language model needs to occupy 20GB of storage space, and the storage space required to process the text currently being processed by the execution server using the large language model is 5GB, then the current remaining storage space size of the execution server is 45GB.
[0040] When determining the target execution server for processing the pending text, the central server can estimate the amount of storage space required to process the pending text using the large language model (i.e., the total storage space in subsequent embodiments) and calculate the current remaining storage space of each execution server. If the current remaining storage space of an execution server is not less than the total storage space, the execution server is considered a candidate execution server that meets the requirements for processing the pending text using the large language model.
[0041] Regarding step S105, the process of pre-filling a text unit with the large language model to obtain the KV corresponding to the text is the pre-filling process; the processing stage for text that has not yet obtained a corresponding KV is the pre-filling stage, and the processing result of the pre-filling stage is the KV corresponding to the text. The process of decoding the text unit by the large language model based on the KV is the decoding process; the processing stage for text that has obtained a corresponding KV but has not yet obtained a corresponding processing result is the decoding stage, and the processing result of the decoding stage is the processing result of the text.
[0042] It is understood that the hardware performance of an execution server is known. For example, the hardware's product parameters may include the data processing speed that the hardware can provide. The number of texts processed in parallel by an execution server is positively correlated with the total number of text units that the execution server can process. For example, the hardware performance of an execution server indicates that when the execution server processes two texts in parallel, the execution server can process 100 text units per second; when the execution server processes three texts in parallel, the execution server can process 120 text units per second. Therefore, the shorter the time required for an execution server to obtain the processing result of a text to be processed using a large language model, the more efficient it is in obtaining the target processing result, that is, the more efficient it is in responding to user input text. The shorter the total time required to obtain the processing result of all currently received texts, the more efficient the execution server's resources are. Therefore, in real-world scenarios, there may be two requirements: the need to minimize the time required to obtain the processing result of the text to be processed (requirement 1), and the need to minimize the total time required to obtain the processing results of multiple texts (requirement 2).
[0043] When the requirement is requirement 1 above, the currently designated text may be the text to be processed. The target execution server is then determined based on the time required for each candidate execution server to obtain the processing result of the text to be processed at the current processing stage using the large language model. When the requirement is requirement 2 above, the currently designated text may be all text currently received by the candidate execution server. The target execution server is then determined based on the total time required for each candidate execution server to obtain the processing result of the text currently received by the candidate execution server using the large language model.
[0044] In this way, the central server can determine the time required for each candidate execution server based on the business needs of the actual application scenario. Subsequently, the target execution server is determined based on the time required by the candidate execution server. In other words, the target execution server for processing the pending document is determined based on business needs. The determined target execution server is then used to process the pending document, further meeting business needs while balancing the computing and storage resources of each target execution server.
[0045] For each candidate execution server, the central server can predict, based on the hardware performance of the candidate execution server, the instantaneous speed at which the candidate execution server would process the text to be processed if the text was sent to the candidate execution server. The central server can then calculate the ratio of the number of text units corresponding to the current processing stage of the text to the instantaneous speed as the predicted duration required by the candidate execution server. The method by which the central server predicts the instantaneous speed corresponding to the candidate execution server can be found in the detailed description of the subsequent embodiments.
[0046] For step S106, the resources allocated to the text are positively correlated with the processing speed of the text. That is, the duration required for an alternative execution server can indicate the resource amount of the target resources currently available to the alternative execution server. The target resources are the type of resources required when obtaining the processing results of the text to be processed for the current stage. If the current processing stage of the text to be processed is the pre-filling stage, the duration required for the alternative execution server is short, which can indicate that the current resource amount of the computing resources of the alternative execution server is large, that is, the computing resources of the alternative execution server are not fully utilized. Therefore, based on the duration required for each alternative execution server, a target execution server whose resources are not used evenly can be determined from each alternative execution server. The target execution server uses its own target resources to process the text to be processed using a large language model, so as to use the computing resources and storage resources of the target execution server in a balanced manner. The specific way in which the central server determines the target execution server can be found in the detailed introduction of the subsequent embodiments, which will not be repeated here.
[0047] Understandably, online platforms are user-oriented, and the timing of user input into online platforms is uncontrollable. Therefore, in real-world scenarios, it is possible to receive multiple texts in a short period of time that require processing using a large language model.
[0048] In one approach, the central server may determine the target execution server for processing each text in sequence according to the order in which the texts are received, that is, schedule each text in sequence.
[0049] In another embodiment, before step S101, the method may further include the following steps: determining the text that meets the preset processing conditions from the texts that have not been pre-filled as the text to be processed; wherein the preset processing conditions indicate that the time taken to obtain the corresponding processing result is the shortest.
[0050] When the time required to obtain a processing result for a text is the shortest, that text needs to be processed first, that is, the target execution server for processing the text needs to be determined first. Specifically, the preset processing conditions may include at least one of the following: the number of text units included is the smallest number of text units included in each text; and the corresponding task processing priority is the highest among the task processing priorities corresponding to each text.
[0051] When the preset processing conditions include: the number of text units included is the least among the number of text units included in each text, since a text includes a small number of text units, the time consumed by using a large language model to obtain the processing result of the text may be shorter. Therefore, each time the central server schedules a task, it can use the text with the least number of text units included in each text that has not been pre-filled (hereinafter referred to as short text) as the text to be processed. That is, giving priority to short texts can reduce the probability of task congestion, improve the overall response efficiency and throughput of the task scheduling system, and improve user experience.
[0052] It is understandable that different users may have different permissions, so the task processing priorities corresponding to the texts entered by different users may also be different. When the preset processing conditions include: the corresponding task processing priority is the highest among the task processing priorities corresponding to each text, since the task processing priority corresponding to a text is higher, it means that the text is an important text with higher importance, such as the text that may be entered by a user with higher permissions. The central server can then schedule important texts first, improve the response efficiency of the task scheduling system for important texts, reduce the probability of user rights being damaged due to untimely processing of important texts, and improve user experience.
[0053] In some embodiments, the aforementioned step S102 may include the following steps: if the maximum number of text units that can be generated by the large language model has been set, and the maximum number of units is not greater than a preset number threshold, the maximum number of units is determined to be the number of text units of the target processing result, and the predicted number is obtained; if the maximum number of units is not set, and the text to be processed carries a specified number of units indicating the number of text units of the target processing result, or if the maximum number of units is greater than the preset number threshold, and the text to be processed carries the specified number of units, the specified number of units is determined to be the predicted number; if the maximum number of units is not set, and the text to be processed does not carry the specified number of units, or if the maximum number of units is greater than the preset number threshold, and the text to be processed does not carry the specified number of units, the text to be processed is input into a pre-trained output unit number prediction model to obtain a predicted number; wherein the output unit number prediction model is: trained based on sample text, and sample labels indicating the number of text units of the sample processing result of the sample text.
[0054] It's understandable that during the training process of a large language model, a maximum number of text units that the large language model can generate may be set. If the large language model has already generated the maximum number of text units, it will stop generating text units. In other words, if a maximum number of units is set, the number of predictions will inevitably be no greater than the maximum number of units.
[0055] The preset threshold can be determined based on actual business needs. For example, if an online platform typically provides users with a service for summarizing articles, the threshold can be set to a smaller value, such as 200. If an online platform typically provides users with a service for writing articles based on indicator words, the threshold can be set to a larger value. For example, if the indicator word is "fairy tale," the threshold can be set to 3000.
[0056] It can be seen that if the maximum number of units is set and is not greater than the preset number threshold, the number of text units in the target processing result may be the maximum number of units. Therefore, in this case, the maximum number of units can be determined as the predicted number.
[0057] However, if the maximum number of units is greater than the preset threshold, it cannot accurately indicate the number of text units in the target processing result. If the maximum number of units is used as the predicted number, the predicted number will be inaccurate, resulting in a low accuracy of the target execution server subsequently determined. Therefore, if the maximum number of units is not set, or if the maximum number of units is greater than the preset threshold, the large language model can continue to generate text units until the complete target processing result is generated.
[0058] It is understood that if the text to be processed itself can indicate the number of text units of the target processing result, for example, if the text to be processed is "Give me a fairy tale of 1,000 words," that is, the text to be processed indicates that the target processing result includes 1,000 characters, then processing the text to be processed using the large language model will obtain the target processing result corresponding to 1,000 characters. In this case, the specified number of text units indicating the target processing result carried by the text to be processed itself can be determined as the predicted number.
[0059] Exemplarily, the central server may input the text to be processed into a keyword extraction model, and obtain the output of the keyword extraction model: an extraction result indicating whether the text to be processed includes a specified number of units, and when the text to be processed includes the specified number of units, the specific numerical value of the specified number of units. For example, the keyword extraction model may be a CNN (Convolutional Neural Network) model. The technician may obtain a plurality of first sample texts, and a first sample label indicating whether each first sample text includes a specified number of units, and when the first sample text includes the specified number of units, the specific numerical value of the specified number of units. Each first sample text is input into the keyword extraction model of the initial structure, and an extraction result output by the keyword extraction model is obtained; a loss function value representing the difference between the extraction result and the first sample label is calculated based on a preset loss function, and the model parameters of the keyword extraction model are adjusted using the loss function value until the model converges. For example, the preset loss function may be an MSE (Mean Squared Error Loss) function, a cross entropy loss function, etc., which is not limited in the present invention.
[0060] If the maximum number of units is not set and the text to be processed does not carry the specified number of units, or if the maximum number of units is greater than the preset number threshold and the text to be processed does not carry the specified number of units, the central server can input the text to be processed into a pre-trained output unit number prediction model and use the output data of the output unit number prediction model as the predicted number. For example, the output unit number prediction model can be an LSTM (Long Short-Term Memory) or a Transformer. The technician can obtain multiple second sample texts and a second sample label indicating the number of sample units of the text unit of the processing result of each second sample text. Each second sample text is input into the output unit number prediction model of the initial structure to obtain the predicted number of text units included in the second sample text output by the output unit number prediction model; based on the preset loss function, a loss function value representing the difference between the predicted number of units and the sample number of units is calculated, and the model parameters of the output unit number prediction model are adjusted using the loss function value until the model converges. For example, the preset loss function can be an MSE function, a cross entropy loss function, etc.; the first sample text and the second sample text can be the same text or different texts, and the present invention is not limited to this.
[0061] Based on the above processing, the central server can obtain a more accurate prediction number, that is, it can improve the accuracy of the second space size required for decoding the text to be processed. The target execution server determined based on the more accurate second space size is more accurate, reducing the probability of allocating a large amount of storage space for a text resulting in other texts being unable to be processed in a timely manner.
[0062] In some embodiments, if an output unit number prediction model is used to obtain a predicted number of text units to be processed, after the target execution server obtains the target processing result using the large language model, the central server may further use the text to be processed as sample text and the actual number of text units in the target processing result as the sample number to continue training the output unit number prediction model. That is, during the use of the output unit number prediction model, the output unit number prediction model is updated based on the actual number of text units in the processing result, thereby improving the accuracy of the subsequent predictions obtained based on the output unit number prediction model, improving the accuracy of the subsequent target execution server determination, and thereby improving the overall resource utilization and response efficiency of the task scheduling system.
[0063] In some embodiments, as described above, after obtaining the processing result of the text to be processed for the current processing stage, the central server can reschedule the text to be processed to improve the efficiency of obtaining the target processing result. In the case where the current processing stage of the text to be processed is the pre-filling stage, the central server can calculate the space size required for pre-filling processing of the first number of text units through the large language model (that is, the first space size in the subsequent embodiments), and then determine from each execution server an execution server whose current remaining storage space is not less than the first space size as an alternative execution service. Subsequently, after obtaining the processing result of the pre-filling stage through the target execution server, the central server can again determine from each execution server a target execution server for decoding processing of the text to be processed.
[0064] In some embodiments, in order to reduce the system's additional overhead and reduce the operating pressure of the central server, the above-mentioned step S104 may include the following steps: calculating the first space size required for pre-filling a first number of text units through a large language model, calculating the second space size required for generating a second number of text units through the large language model, and calculating the sum of the first space size and the second space size as the total space size; from each execution server, determining the execution server whose current remaining storage space is not less than the total space size as an alternative execution server.
[0065] It's understandable that the space required by a trained large language model to process a single text unit (hereinafter referred to as the unit space size) is known. For example, if a large language model includes 70 network layers, each with 8192 model parameters, a text unit is 2 bits long, and each text unit corresponds to one key and one value, then the space required to process a single text unit using the large language model is approximately 8192 × 70 × 2 × 2, or approximately 2.5 MB (megabytes).
[0066] Therefore, by multiplying the first number by the unit space size, we can obtain the first space size required to pre-fill the first number of text units to be pre-filled using the large language model. By multiplying the second number by the unit space size, we can obtain the second space size required to generate the second number of text units to be generated using the large language model. The sum of the first and second space sizes is the total space size required to process the text to be processed using the large language model.
[0067] Furthermore, the central server can identify alternative execution servers whose current remaining storage space is no less than the total storage space. Consequently, regardless of whether the document to be processed is subsequently scheduled again, the alternative execution servers can allocate sufficient storage space for the document to be processed. This reduces the probability of the target execution server identified from the alternative execution servers needing to offload data related to the document to be processed due to insufficient storage space, thereby reducing additional overhead and the likelihood of overall system performance degradation.
[0068] In some embodiments, the above-mentioned step S105 may include the following steps: when the first number is not 0, for each alternative execution server, based on the first number, predicting how long it will take for the alternative execution server to use the large language model to obtain the processing results of the current specified text for the pre-filling stage if the text to be processed is sent to the alternative execution server; when the first number is 0, for each alternative execution server, based on the second number, predicting how long it will take for the alternative execution server to use the large language model to obtain the processing results of the current specified text for the decoding stage if the text to be processed is sent to the alternative execution server.
[0069] Because the second number is calculated based on the predicted number, while the first number is directly determined based on the known text to be processed, the first number is accurate, while the second number may have errors compared to the actual number of text units in the target processing result. Therefore, when the first number is not zero, the accuracy of the time required for each candidate execution server to obtain the processing result of the current specified text in the pre-filling stage, as predicted based on the first number, is higher than the accuracy of the time required for each candidate execution server to obtain the processing result of the current specified text in the decoding stage, as predicted based on the second number.
[0070] Therefore, when the first number is not 0, that is, the text to be processed is currently in the pre-filling stage, the pre-filling processing of the text to be processed requires more computing resources. In order to improve the accuracy of the determined target execution server, for each alternative execution server, the time required for the alternative execution server can be calculated according to the hardware performance of the alternative execution server, the processing progress of the text in the pre-filling stage currently being processed in parallel by the alternative execution server, and the first number.
[0071] For example, if the first number is 100, and the hardware performance of an alternative execution server indicates that the alternative execution server processes 120 text units per second when processing three texts in parallel, then if, after the text to be processed is sent to the alternative execution server, the alternative execution server performs pre-fill processing on the three texts in parallel, that is, the alternative execution server performs pre-fill processing on 40 text units in one text per second, then the time required for the alternative execution server to obtain the pre-fill processing result for the text to be processed using the large language model may be 2.5 seconds.
[0072] When the first number is 0, that is, the text to be processed is currently in the decoding stage, and decoding the text to be processed requires a large amount of storage and reading resources, the central server can calculate the time required for the alternative execution server based on the hardware performance of the alternative execution server, the processing progress of the text in the decoding stage currently being processed in parallel by the alternative execution server, and the second number. The method by which the central server calculates the time required for an alternative execution server to obtain the processing results of the decoding stage of the text using the large language model is similar to the method for calculating the time required for an alternative execution server to obtain the processing results of the pre-filling stage of the text using the large language model. Please refer to the relevant description of the aforementioned embodiment.
[0073] Furthermore, the central server may determine a target execution server whose resources are not used evenly from among the candidate execution servers based on the time required by each candidate execution server.
[0074] In some embodiments, the above-mentioned step S106 may include the following steps: when the first number is not 0, if the ratio of the first number to the second number is greater than the first ratio, the alternative execution server with the shortest required time is determined as the target execution server; if the ratio of the first number to the second number is not greater than the first ratio, the alternative execution server with the longest required time is determined as the target execution server; when the first number is 0, the alternative execution server with the shortest required time is determined as the target execution server.
[0075] It is understandable that in actual scenarios, the tasks indicated by the text input by the user usually include three categories: summary tasks, generation tasks, and answer tasks.
[0076] The number of text units in a summary-type task is typically much larger than the number of text units in the processed result. For example, Text 1 for a summary-type task might include an article of over a thousand words, with the prompt "Please give me a summary of this article, no more than 100 words." This means that Text 1 corresponds to over a thousand text units; however, the processed result of Text 1 might only contain a few dozen text units.
[0077] The number of text units in the text corresponding to the generation task is usually much smaller than the number of text units in the processed result of the text. For example, Text 2 for the generation task may be "Give me a fairy tale of 1,000 words." This means that Text 2 only corresponds to dozens of text units, but the processed result of Text 2 may contain thousands of text units.
[0078] The number of text units in the text corresponding to the task is usually close to the number of text units in the result of processing the text. For example, text 3 indicating the task of answering the question could be "Where is the capital of Country A?" This means that text 3 corresponds to dozens of text units, and the result of processing text 3 also only includes a dozen or so text units.
[0079] As can be seen, in actual scenarios, the number of text units to be processed during the pre-filling phase (referred to as the pre-filling number) is typically inversely proportional to the number of text units to be generated during the decoding phase (referred to as the decoding number). That is, when the pre-filling number for a text is large, the decoding number for that text is typically small. Furthermore, when the text input by the user is directly used as the text to be processed, the pre-filling number is the aforementioned first number, and the decoding number is the aforementioned second number. The ratio of the first and second numbers (hereinafter referred to as the number ratio) can indicate the type of task indicated by the text to be processed. If the number ratio is not greater than the first ratio, the task indicated by the text to be processed is likely to be a generation task; if the number ratio is greater than the first ratio but not greater than the second ratio, the task indicated by the text to be processed is likely to be a solution task; and if the number ratio is greater than the second ratio, the task indicated by the text to be processed is likely to be a summary task. The first and second ratios can be determined based on the actual application scenario. For example, the first ratio can be 0.1, and the second ratio can be 10.
[0080] Therefore, if the first number is not zero and the number ratio is greater than the first ratio, the amount of computing resources required to process the document to be processed is greater than the amount of storage and reading resources required, meaning that computing resources are primarily consumed when processing the document to be processed. To improve processing efficiency of the document to be processed, the candidate execution server with the shortest processing time can be determined as the target execution server.
[0081] In some embodiments, as described above, when the number ratio is greater than the first ratio, the task indicated by the to-be-processed text may be a solution task or a summary task. When the task indicated by the to-be-processed text is a solution task, more storage and reading resources are consumed when processing the to-be-processed text.
[0082] At this time, the central server can first determine the alternative execution server with the shorter required time from each alternative execution server according to the time required by each alternative execution server (hereinafter referred to as the execution server to be screened). For example, the alternative execution server whose required time is less than the preset time threshold can be determined as the execution server to be screened; or, the alternative execution servers can be sorted in order from small to large according to the time required by each alternative execution server, and then the first preset number of alternative execution servers in the sorting are determined as the execution servers to be screened, but not limited to this. Then, according to the second number, it is predicted that after sending the text to be processed to each execution server to be screened, the execution server to be screened will use the large language model to obtain the time required for the processing result of the current specified text for the decoding stage, and then the execution server to be screened with the shortest time required to obtain the processing result of the decoding stage is used as the target execution server.
[0083] In the case where the number ratio is not greater than the first ratio, that is, the storage and reading resources are mainly consumed when processing the text to be processed. However, as mentioned above, in the case where the first number is not 0, the accuracy of the time required for each alternative execution server predicted based on the first number is higher than the accuracy of the time required for each alternative execution server predicted based on the second number. Therefore, in order to improve the accuracy of the target execution server with uneven resource usage, the central server can use the alternative execution server with the longest time required to obtain the processing result of the current specified text for the pre-filling stage as the target execution server with the shortest time required to obtain the processing result of the current specified text for the decoding stage, that is, the target execution server whose storage and reading resources are not used evenly.
[0084] If the first number is 0, the text to be processed is currently in the decoding phase, and processing the text primarily consumes storage and reading resources. At this point, the central server can determine, from among the candidate execution servers, the candidate execution server that takes the shortest time to obtain the processing result for the decoding phase of the current specified text, and use it as the target execution server.
[0085] Based on the above process, the central server can determine the type of resources primarily consumed when processing the document to be processed based on the number ratio. Then, based on the time required by each candidate execution server, it can determine a target execution server from among the candidate execution servers whose resources of that type are not evenly utilized. After the document to be processed is sent to the target execution server, the target execution server processes the document based on its own unbalanced resources, ensuring that these resources are evenly utilized.
[0086] In some embodiments, in order to further improve the accuracy of the determined target execution server, the number of text units that can be pre-filled by the execution server per second may be different from the number of text units that can be generated per second. For each alternative execution server, the central server can calculate the time required for the alternative execution server to obtain the processing result of the pre-filling processing of the text to be processed (hereinafter referred to as the first time), and the time required for the alternative execution server to obtain the processing result of the decoding processing of the text to be processed (hereinafter referred to as the second time), according to the hardware performance of the alternative execution server, and calculate the ratio of the first time and the second time (hereinafter referred to as the time ratio). Then determine the target execution server according to the time ratio. For example, the central server can determine the target execution server based on the time ratio in a manner similar to the aforementioned determination of the target execution server based on the number ratio. If the first time is not 0, if the time ratio is greater than the first ratio, the alternative execution server with the smallest first time is determined as the target execution server; if the time ratio is not greater than the first ratio, the alternative execution server with the longest first time is determined as the target execution server. When the first duration is 0, the candidate execution server with the shortest second duration is determined as the target execution server.
[0087] In some embodiments, the above-mentioned step S105 may include the following steps: when the processing priority of the text to be processed is not less than the preset priority, for each alternative execution server, based on the first number and the second number, predict that if the text to be processed is sent to the alternative execution server, the alternative execution server will use the large language model to obtain the minimum time required for the processing result of the text to be processed for the current processing stage.
[0088] Accordingly, step S106 may include the following steps: selecting the candidate execution server with the shortest required execution time as the target execution server, and sending instruction information and the text to be processed to the target execution server, so that the target execution server, based on the instruction information, migrates the text to be migrated and the data to be migrated to a designated storage space, and then processes the text to be processed based on the large language model. The text to be migrated includes: text currently being processed by the target execution server that is at the same processing stage as the text to be processed; and the data to be migrated includes: all data obtained during the processing of the text to be migrated using the large language model.
[0089] In actual scenarios, there may be texts to be processed that need to be processed with priority, such as the texts with higher corresponding task processing priorities mentioned above. The higher the task processing priority of a text to be processed, the more important the text to be processed is, and the more priority it needs. Therefore, when the processing priority of the text to be processed is not less than the preset priority, the text to be processed can be called an important text. The preset priority can be set according to the business needs of the actual scenario, such as the highest task processing priority of each text; or, it can be a specified task processing priority. In order to improve the efficiency of obtaining the processing results of important texts, for each alternative execution server, the central server can determine the minimum time required for the alternative execution server to obtain the processing results of the important text using the large language model, that is, the time required when the alternative execution server is currently only processing important texts.
[0090] Specifically, since the alternative execution server may be currently processing other texts, the central server can determine the current stage of the other texts currently being processed by the alternative execution server, and then use other texts that are in the same processing stage as the important text as texts to be migrated; all data obtained in the process of processing the texts to be migrated using the large language model will be used as data to be migrated.
[0091] If the text to be migrated and the data to be migrated are migrated to the designated storage space, the local storage space and resources of the alternative execution server are released. At this time, if the important text is sent to the alternative execution server, the time required by the alternative execution server is the shortest time required to obtain the processing result of the important text using the large language model. The way in which the central server obtains the shortest time required by the alternative execution server is similar to the way in which the time required by each alternative execution server is obtained in the aforementioned embodiment. Please refer to the relevant introduction of the aforementioned embodiment. Furthermore, the central server can determine the target execution server from each alternative execution server based on the shortest time required by each alternative execution server, and then send the instruction information and important text to the target execution server. Accordingly, after receiving the instruction information, the target execution server can migrate the text to be migrated and the data to be migrated to the designated storage space based on the instruction information to release the storage space and resources of the target execution server. Then, the large language model is used to process the important text.
[0092] The designated storage space of an execution server is: the current remaining storage space in the designated server; or the memory of the execution server.
[0093] In one approach, the target execution server can migrate the text and data to be migrated into its own memory. Later, after processing the important text, the target execution server removes the text and data from its memory and continues processing them using the large language model. This eliminates the need for rescheduling through the central server, reducing the operational pressure on the central server.
[0094] Alternatively, the target execution server can send the documents and data to be migrated to the central server. The central server then schedules the migrated documents as pending documents. This ensures that pending documents are processed promptly, improving the overall responsiveness of the task scheduling system.
[0095] In some embodiments, after the target execution server uses the large language model to obtain the processing results for the important text in the current processing stage, it can send the important text and the obtained processing results to the central server, which then schedules the important text. This can shorten the time spent on the pre-population and decoding stages of the important text, improving the efficiency of the response for important text.
[0096] In some embodiments, the above-mentioned step S106 may include the following steps: when the text to be processed is obtained from any execution server, the text to be processed and the data to be reused are sent to the target execution server, so that the target execution server continues to process the text to be processed using the data to be reused based on the large language model; wherein the data to be reused is: all data obtained in the process of processing the text to be processed using the large language model.
[0097] At this time, the text to be processed can be the aforementioned text to be migrated, the aforementioned important text, but obviously, there is more than this in the actual scenario. In the case of obtaining the text to be processed from any execution server, that is, the text to be processed has been processed using the large language model, and some relevant data (i.e., data to be reused) has been calculated. After determining the target execution server, the central server can send the text to be processed and the data to be reused to the target execution server. Furthermore, the target execution server can use the large language model to continue processing the text to be processed using the data to be reused. There is no need to re-process the text to be processed, which improves the efficiency of obtaining the target processing results.
[0098] In some embodiments, while each execution server is currently processing a document, the central server sequentially schedules the documents to be processed, which can be called real-time scheduling. As the aforementioned task scheduling methods can all be performed in real-time scheduling, by real-time scheduling of documents, the overall response efficiency of the system can be improved.
[0099] In some embodiments, there are multiple texts to be processed; accordingly, the above step S105 may include the following steps: after each execution server has processed the texts received historically, according to the preset allocation order, determine the execution server currently used to process the texts to be processed from each execution server as the current execution server; based on the first number and the second number of each text to be processed, determine the text to be sent to the current execution server from other texts to be processed except the target text, so as to be processed by the current execution server using the large language model, as the target text corresponding to the current execution server; if all the target texts corresponding to the current execution server do not meet the exception handling conditions, return to the execution based on each text to be processed. The first number and the second number of this method are used to determine, from other texts to be processed except the target text, the step of sending the texts to be processed to the current execution server for processing by the current execution server using the large language model, until all target texts corresponding to the current execution server meet the exception handling condition, and obtain the time required for obtaining the processing results of all target texts corresponding to the current execution server using the large language model; wherein, multiple texts meet the exception handling condition, which means that if the multiple texts are sent to the current execution server, an exception occurs in the processing process of the multiple texts by the current execution server; and returning to the step of determining, from each execution server in accordance with a preset allocation order, the execution server currently used to process the texts to be processed.
[0100] In actual scenarios, in order to reduce the system's additional overhead, that is, to reduce the probability of multiple scheduling of texts, and to reduce the operating pressure of the central server, the central server can also perform package scheduling. Specifically, the text entered by the user through the online platform can be cached in the task pool first. After each execution server has processed the text received historically, that is, when the text selected during the last package scheduling has been fully processed, the central server can obtain all the texts in the current task pool as the text to be processed. Alternatively, the central server can also select multiple texts from the task pool as the text to be processed in a preset order. For example, the preset order can be the order of the caching time of the text, the order of the number of text units of the text from small to large, the order of the processing priority of the text from high to low, etc., but is not limited to this. The number of texts selected by the central server (which can be called the number of packaged texts) can be a specified number, or it can be positively correlated with the number of execution servers, and the present invention is not limited to this. Obviously, at this time, each execution server in the task scheduling system is an alternative execution server.
[0101] It is understood that in actual scenarios, the central server may determine the target text corresponding to each execution server in the task scheduling system in a predetermined allocation order. For example, the target text corresponding to each execution server may be determined in descending order of hardware performance; or the target text corresponding to each execution server may be determined in order of the execution server's serial number, etc., but the present invention is not limited thereto.
[0102] For ease of description, the central server determines the target text corresponding to the current execution server as an example for explanation. The current execution server can be any execution server in the task scheduling system. As mentioned above, the first number of a text can indicate the amount of computing resources required to pre-fill the text; the second number of the text can indicate the amount of storage and reading resources required to obtain the processing result of the text. Then, by combining the first number and the second number, the text that can be normally executed by the current execution server and the computing resources and storage and reading resources of the current execution server are fully utilized (that is, the target text corresponding to the current execution server) can be determined from other texts to be processed except the target text. Then, the target text corresponding to the current execution server is subsequently sent to the current execution server, and the resources of the current execution server can be fully utilized. Accordingly, the current execution server is the target execution server used to process the target text corresponding to the current execution server.
[0103] In one embodiment, when the target text corresponding to the current execution server does not meet the exception handling conditions, the central server can compare the amount of computing resources currently available to the current execution server (referred to as the current computing resources) with the amount of storage resources currently available to the current execution server (referred to as the current storage resources). If the current computing resources are not less than the current storage resources, the central server can use the text with the largest first number of corresponding texts to be processed other than the target text as the target text corresponding to the current execution server. Then, the difference between the current computing resources and the amount of computing resources corresponding to the first number of target texts corresponding to the current execution server (referred to as the used computing resources) is calculated to update the current computing resources; the difference between the current storage resources and the amount of storage resources corresponding to the second number of target texts corresponding to the current execution server (referred to as the used storage resources) is calculated to update the current storage resources. If the current computing resources are less than the current storage resources, the central server can use the text with the largest second number of corresponding texts to be processed other than the target text as the target text corresponding to the current execution server. Then, the difference between the current computing resource amount and the used computing resource amount is calculated to update the current computing resource amount; the difference between the current storage and read resource amount and the used storage and read resource amount is calculated to update the current storage and read resource amount.
[0104] Then, the comparison between the current computing resource amount and the current storage and reading resource amount is continued, and this cycle is repeated until the target text corresponding to the current execution server meets the exception handling condition.
[0105] If multiple texts satisfy the exception handling conditions, it means that if the multiple texts are sent to the current execution server, the current execution server experiences an exception in processing the multiple texts. For example, the exception handling conditions may include: the number of target texts corresponding to the current execution server is not less than the maximum parallel processing capacity of the current execution server; the total space occupied by the target texts corresponding to the current execution server is not less than the current remaining storage space of the current execution server; and the time required for the current execution server to obtain the processing results for its own target texts using the large language model is not less than a preset time threshold.
[0106] It is understandable that after the large language model is deployed on the execution server, there is an upper limit to the number of texts that the large language model can process in parallel. In order to reduce the probability of abnormal situations during task scheduling, the maximum number of parallel processes is usually less than this upper limit. For example, through testing, it is determined that when the large language model processes 10 texts in parallel, the large language model can operate normally; but when processing 11 texts in parallel, the operation of the large language model becomes abnormal, that is, the upper limit of the number of texts that the large language model can process in parallel is 10. At this time, the maximum number of parallel processes can be set to 8. Alternatively, in order to fully utilize the resources of each execution server in the task scheduling system, the maximum number of parallel processes can also be determined based on the number of texts to be processed and the number of execution servers. For example, the quotient of the number of texts to be processed and the number of execution servers can be calculated as the maximum number of parallel processes. This can reduce the probability of some execution servers not having corresponding target texts during the packaging scheduling process, thereby reducing the probability of the problem that the resources of these execution servers are not fully utilized.
[0107] As mentioned above, the storage space of each execution server is limited. In order to reduce the system's additional overhead, alleviate the operating pressure of the central server, and reduce the probability of multiple scheduling of texts, for each current execution server, the sum of the total space occupied by the target text corresponding to the current execution server must also be no less than the current remaining storage space size of the current execution server.
[0108] The computing resources and storage resources of each execution server are also limited. In order to improve the overall response efficiency of the system, the time required for each current execution server to use the large language model to obtain the processing results of its corresponding target text must be no less than the preset time threshold.
[0109] If the target text corresponding to the current execution server meets any of the exception processing conditions, that is, the current execution server cannot continue to process further text, then the target text corresponding to the current execution server is the text to be sent to the current execution server for processing by the current execution server using the large language model; the current execution server is the target execution server for processing its own target text.
[0110] In one embodiment, the central server may determine the target text corresponding to the current execution server based on the first number and the second number alternately.
[0111] Specifically, the central server can sort each text to be processed in the order of the size of the first number of the text to be processed, such as in the order from large to small according to the first number (hereinafter referred to as the first order). The more forward a text to be processed is in the first order, the larger the first number of the text to be processed is, that is, the larger the storage space to be occupied and the more computing resources to be consumed when obtaining the processing result of the text to be processed in the pre-filling stage. The more forward a text to be processed is in the second order, the larger the second number of the text to be processed is, that is, the larger the storage space to be occupied and the more computing resources to be consumed when obtaining the processing result of the text to be processed in the pre-filling stage (hereinafter referred to as the second order). The more forward a text to be processed is in the second order, the larger the second number of the text to be processed is, that is, the larger the storage space to be occupied and the more storage and reading resources to be consumed when decoding the text to be processed is. The central server can alternately determine the text to be processed with a larger first number (hereinafter referred to as the long pre-filled text), and the text to be processed with a larger second number (hereinafter referred to as the long decoded text).
[0112] For example, the central server may determine, from among the other to-be-processed texts excluding the target text, the to-be-processed text that is at the front of the first order in the first order as the target text corresponding to the current execution server, and then determine whether all the target texts corresponding to the current execution server meet the aforementioned exception handling conditions.
[0113] If all target texts corresponding to the current execution server do not meet the aforementioned exception handling conditions, in order to fully utilize the resources of the current execution server, the central server may, from among the other to-be-processed texts other than the target text, determine the to-be-processed text that is at the front of the second order in the second order as the target text corresponding to the current execution server. The central server then proceeds to determine whether all target texts corresponding to the current execution server meet the aforementioned exception handling conditions.
[0114] If all target texts corresponding to the current execution server do not meet the aforementioned exception handling conditions, the central server continues to determine the target text corresponding to the current execution server based on the first order. This cycle continues until the aforementioned exception handling conditions are met. The current execution server is then determined again according to the preset allocation order, and the target text corresponding to the current execution server is determined in the same manner as above.
[0115] It is understandable that the specific method of determining the target text corresponding to the current execution server based on the first number and the second number alternately can be adjusted based on the needs of the actual scenario, such as according to the aforementioned method, first based on the first order, then based on the second order, and then based on the first order, and so on, to determine the target text corresponding to the current execution server; or, it is also possible to use the second order, then based on the first order, and then based on the second order, and so on, to determine the target text corresponding to the current execution server; or, it is also possible to combine the first number and the second number to sort the texts to be processed, such as calculating the statistical values of the first number and the second number, and sorting them in order from large to small according to the statistical values corresponding to each text to be processed (referred to as the third order), and then, first based on the first order, then based on the second order, then based on the third order, and then based on the first order, and so on, alternately. Obviously, there are far more feasible ways in actual scenarios than this.
[0116] Exemplarily, the above method of determining the target text corresponding to the current execution server is as follows: first based on the first order, then based on the second order, and then based on the first order, and so on alternately.
[0117] If the first order is: Text 1, Text 2, Text 3, Text 4, Text 5; and the second order is: Text 5, Text 4, Text 3, Text 2, Text 1, then for the first execution server 1 in the preset assignment order, the central server can initially select Text 1 as the target text for execution server 1. At this point, Text 1 has been determined as the target text. If Text 1 does not meet the exception handling conditions, the central server will continue to follow the second order and select Text 5 as the target text for execution server 1 from the remaining pending texts (i.e., Text 5, Text 4, Text 3, Text 2).
[0118] At this time, both text 1 and text 5 have been determined as target texts. If text 1 and text 5 meet the exception handling conditions, execution server 1 will be the target execution server for text 1 and text 5.
[0119] According to the preset allocation order, the central server determines that execution server 2 is the current execution server. According to the first order, from the other to-be-processed texts (i.e., text 2, text 3, and text 4) except the target text, text 2 is used as the target text corresponding to execution server 2. At this time, text 1, text 2, and text 5 have all been determined as target texts. If text 2 does not meet the exception handling conditions, the central server continues to follow the second order and, from the other to-be-processed texts (i.e., text 4 and text 3) except the target text, text 4 is used as the target text corresponding to execution server 2. At this time, text 1, text 2, text 4, and text 5 have all been determined as target texts. If text 2 and text 4 do not meet the exception handling conditions, the central server continues to follow the first order and, from the other to-be-processed texts (i.e., text 3) except the target text, text 3 is used as the target text corresponding to execution server 2.
[0120] At this point, there is no text to be processed, so this round of packaging scheduling is completed, and execution server 2 is the target execution server for text 2, text 3, and text 4.
[0121] Based on the above processing, the central server can also perform unified scheduling of multiple pending documents using a packaged scheduling approach. This can reduce system overhead, namely, the need to schedule documents multiple times, and can also reduce the operating pressure on the central server, lowering the probability of operational failures due to excessive business pressure, and improving the stability of the task scheduling system.
[0122] In some embodiments, if the target texts corresponding to each execution server have been determined but there are still pending texts, the central server can reallocate the pending texts to the execution server that requires the shortest processing time. Accordingly, the execution server can store the pending texts in memory and, after completing processing of its own target texts, process the pending texts stored in memory.
[0123] In some embodiments, after obtaining the processing result of a text to be processed, the central server can also count the storage space actually required for the relevant data generated in the process of processing the text to be processed using the large language model, the performance of the execution server that pre-fills the text to be processed, the actual time consumed to obtain the pre-fill processing result of the text to be processed, the performance of the execution server that decodes the text to be processed, the actual time consumed to obtain the decoding processing result of the text to be processed, and other related information. For example, the central server can record the relevant information in a local knowledge base. Subsequently, the central server can continuously update the task scheduling system based on the information recorded in the local knowledge base, such as updating the keyword extraction model and the output unit number prediction model in the aforementioned embodiment. The accuracy of the target execution server determined subsequently can be improved, thereby improving the overall resource utilization and response efficiency of the task scheduling system.
[0124] In some embodiments, see Figure 2 , Figure 2 This is a second flow chart of the task scheduling method provided by an embodiment of the present invention. According to the execution order of the various steps included in the task scheduling method, the execution process of the task scheduling method can be divided into the following stages: input acquisition stage; input and output length parsing stage; token (text unit) length calculation stage; instance performance parsing stage; task resource requirement calculation stage; and scheduling stage.
[0125] In the input acquisition stage, the central server can acquire input, that is, the text to be processed in the aforementioned embodiment.
[0126] In the input and output length parsing stage, prompt (prompt vector) parsing is performed on the obtained input to obtain the first number of text units included in the input, that is, the input tokens length in the tokens length calculation stage is obtained.
[0127] For the output that needs to be obtained, when the max_tokens (maximum number of text units) of the large language model is set, that is, the maximum number of units in the aforementioned embodiment, and max_tokens is not greater than the preset number threshold, max_tokens is the second number of predicted output text units, that is, the output tokens length in the tokens length calculation stage is obtained. If max_tokens is not set, or, when max_tokens is greater than the preset number threshold, if keywords are extracted from the input, keywords indicating the number of text units of the target processing result carried in the input (that is, the specified number of units in the aforementioned embodiment) can be obtained, and the extracted specified number of units is the output tokens length. If max_tokens is not set, and the input does not carry the above keywords, or if max_tokens is greater than the preset number threshold, and the input does not carry the above keywords, the central server can process the input through a deep learning model (that is, the output unit number prediction model in the aforementioned embodiment) to obtain the output tokens length.
[0128] Furthermore, based on the length of the input tokens, the KVCache (key-value cache) space required to Prefill the input can be calculated, that is, the first space size in the aforementioned embodiment; based on the length of the output tokens, the KVCache space required for Decode in the process of obtaining the output can be calculated, that is, the second space size in the aforementioned embodiment.
[0129] Moreover, for each execution server, according to the hardware performance of the execution server, the central server can calculate the Prefill time of the execution server based on the Prefill performance of the execution server (that is, the number of text units that can be prefilled per second in the aforementioned embodiment) and the input tokens length; and calculate the Decode time of the execution server based on the Decode performance of the execution server (that is, the number of text units that can be generated per second in the aforementioned embodiment) and the output tokens length.
[0130] An execution server with a trained large language model can be used as an instance. The central server can obtain the real-time KVCache balance, that is, the current remaining storage space of each execution server.
[0131] Furthermore, when performing real-time scheduling, the central server can determine sufficient KVCache instances according to the real-time KVCache margin, the KVCache space required for Prefill, and the KVCache space required for Decode, that is, determine the alternative execution server whose current remaining storage space is not less than the total space size.
[0132] Furthermore, for each instance, the central server can also obtain the number of tasks and the remaining processing time for each stage of Prefill / Decode. That is, for each candidate execution server, the central server can also obtain the processing progress of each text currently being processed in parallel by the candidate execution service.
[0133] According to the Prefill time / Decode time, that is, the duration ratio in the aforementioned embodiment, the central server can determine the instance for processing the input.
[0134] When the duration ratio is much greater than 1, the input is preferentially dispatched to instances where the number of Prefill phase tasks is small or the Prefill phase tasks are about to end. That is, in the aforementioned embodiment, the candidate execution server that takes the shortest time to obtain the processing result of the to-be-processed text in the Prefill phase.
[0135] When the duration ratio is much less than 1, the input is preferentially dispatched to instances with a large number of Prefill stage tasks or Prefill stage tasks that are time-consuming. That is, in the aforementioned embodiment, when the first number is not 0, the candidate execution server with the longest processing time required to process the text to be processed in the prefill stage is obtained; when the first number is 0, the candidate execution server with the shortest processing time required to process the text to be processed in the decoding stage is obtained.
[0136] When the duration ratio is approximately 1, the appropriate instance is selected based on the duration of the Prefill phase. That is, in the aforementioned embodiment, the server that takes the shortest time to process the text in the Prefill phase is selected from the servers that take the shortest time to process the text in the Decoding phase.
[0137] When performing package scheduling, one task corresponds to one text to be processed.
[0138] The central server can sort tasks based on their Prefill and Decode times. Since the hardware performance of a single execution server is fixed, the order in which the Prefill times are taken by the execution server when prefilling tasks is consistent with the order of the first number of text units included in each task. That is, the Prefill time ranking corresponds to the first order in the aforementioned embodiment. Similarly, the Decode time ranking corresponds to the second order in the aforementioned embodiment.
[0139] Then, the central server can combine the long Prefill task with the long Decode task based on KVCache. The long Prefill task is the long prefill text in the aforementioned embodiment; the long Decode task is the long decoded text in the aforementioned embodiment. That is, the central server can determine the target text corresponding to the current execution server from other to-be-processed texts other than the target text in a first order. If all target texts corresponding to the current execution server do not meet the aforementioned exception handling conditions, the target text corresponding to the current execution server can be determined from other to-be-processed texts other than the target text in a second order; if all target texts corresponding to the current execution server do not meet the aforementioned exception handling conditions, the target text corresponding to the current execution server can continue to be determined from other to-be-processed texts other than the target text in a first order until all target texts corresponding to the current execution server meet any of the aforementioned exception handling conditions. Furthermore, the central server can integrate the remaining tasks based on the time consumed by each group of Prefill tasks and Decode tasks. That is, in the aforementioned embodiment, if the target texts corresponding to each execution server have been determined but there is still text to be processed, the central server can reallocate the text to be processed to the execution server with the shortest required time.
[0140] Based on the above processing, the text units to be pre-filled are the text units that have not been pre-filled among the text units of the text to be processed, then the first number of text units to be pre-filled can indicate: the amount of computing resources required to be consumed when pre-filling the text units to be pre-filled; the second number is the difference between the predicted number of text units of the target processing result and the number of generated text units, that is, the predicted number of text units that still need to be generated through decoding processing, then the second number can indicate: the amount of storage and reading resources that still need to be consumed in the process of obtaining the target processing result.
[0141] Because pre-populating text based on a large language model consumes significant computing resources, and decoding based on the pre-population results consumes significant storage and read resources, the resources allocated to the text during text processing are positively correlated with the text processing speed. Based on the first number and the second number, the time required for each candidate execution server to utilize the large language model to obtain the processing result for the current processing stage of the specified text is calculated. The time required by each execution server can indicate the amount of target resources currently available to that execution server. The shorter the time required by an execution server, the greater the amount of target resources currently available to that execution server, and the greater the probability that the target resources of that execution server are underutilized, i.e., the greater the probability that the computing and storage resources of that execution server are not being used evenly. Therefore, based on the time required by each candidate execution server, the target execution server with the unbalanced resource usage can be determined from among the candidate execution servers. The text to be processed is then sent to the target execution server, which then processes the text using its currently available target resources. This fully utilizes the target execution server's own unbalanced resources, thereby achieving balanced utilization of the target execution server's computing and storage resources.
[0142] That is, it realizes the intelligent task allocation scheduled among the deployment of multiple large model examples, which can optimize the use of KVCache space and the resource utilization of each execution server. That is to improve the utilization rate of KVCache space, balance the pre-filling and decoding tasks, reduce the performance overhead caused by insufficient resources, and ultimately improve the overall efficiency of the system and user experience. Moreover, in the dynamic batch processing of tasks, the task scheduling system will reasonably combine and schedule the tasks in the pre-filling stage (i.e., computing-intensive tasks) and the tasks in the decoding stage (i.e., storage and reading-intensive tasks), and dynamically balance the task allocation of the two stages according to the processing time data of the instance, so that the computing resources of each example, such as GPU (Graphics Processing Unit, image processor) resources, and storage and reading resources are efficiently utilized. And by dynamically adjusting the task allocation strategy, it is possible to avoid excessive occupation of resources in a certain stage,
[0143] When the task scheduling method provided by the present invention is applied in a cloud computing platform, in a cloud computing environment, efficient resource scheduling and task management can be provided, and the performance of the large language model server and the response speed to user requests can be improved. In the scenario of big data processing based on a large language model, the efficiency and stability of large-scale data processing can be improved by optimizing resource allocation. In the scenario of AI (Artificial Intelligence) reasoning and analysis services based on a large language model, the efficiency and timeliness of task processing are ensured, and the reliability and user experience of AI services are improved. In the scenario of high-performance computing based on a large language model, intelligent scheduling and resource allocation can be used to maximize the utilization of computing resources and improve overall computing performance.
[0144] Based on the same inventive concept as the above-mentioned task scheduling method, the present invention also provides a task scheduling device, which is applied to a central server in a task scheduling system, and the task scheduling system also includes multiple execution servers. Figure 3 , Figure 3 A structural diagram of a task scheduling device provided by an embodiment of the present invention. The device includes:
[0145] The acquisition module 301 is configured to acquire the number of text units to be pre-filled as a first number; wherein the text units to be pre-filled are text units that have not been pre-filled among the text units of the text to be processed; a text unit of a text represents an encoding result of characters in the text;
[0146] Prediction module 302 is configured to predict the number of text units of a target processing result as a predicted number; wherein the target processing result is obtained by pre-filling and decoding the text units of the to-be-processed text based on the large language model;
[0147] The number calculation module 303 is configured to calculate the difference between the predicted number and the number of generated text units to obtain a second number; wherein the generated text units are text units generated by decoding processing among the text units of the target processing result;
[0148] A determination module 304 is configured to determine, from among the execution servers, an execution server whose current remaining storage space satisfies the requirement for processing the text to be processed using the large language model, and select the server as a candidate execution server;
[0149] The duration calculation module 305 is configured to predict, for each candidate execution server, based on the first number and the second number, the duration required for the candidate execution server to obtain a processing result of the current specified text at the current processing stage using the large language model if the text to be processed is sent to the candidate execution server; wherein the current specified text is determined from text received from the candidate execution server; the current processing stage of any specified text represents a pre-filling stage or a decoding stage; the duration required for an execution server indicates the amount of target resources currently available to the execution server; the target resources are resources of the type required to obtain the processing result of the text to be processed at the current stage;
[0150] The sending module 306 is used to determine a target execution server from each alternative execution server based on the time required by each alternative execution server, and send the text to be processed to the target execution server, so that the target execution server processes the text to be processed based on the large language model using the target resources currently available to the target execution server.
[0151] Optionally, the determination module 304 is specifically used to calculate the first space size required for pre-filling the first number of text units through the large language model, calculate the second space size required for generating the second number of text units through the large language model, and calculate the sum of the first space size and the second space size as the total space size; from each execution server, determine the execution server whose current remaining storage space is not less than the total space size as an alternative execution server.
[0152] Optionally, there are multiple texts to be processed; the duration calculation module 305 is specifically used to determine, after each execution server has processed the texts received historically, the execution server currently used to process the texts to be processed from each execution server in accordance with a preset allocation order, as the current execution server; based on the first number and the second number of each text to be processed, from other texts to be processed except the target text, determine the text to be sent to the current execution server to be processed by the current execution server using the large language model as the target text corresponding to the current execution server; if all target texts corresponding to the current execution server do not meet the exception handling conditions, return to execute the first number of each text to be processed and a second number, determining, from other to-be-processed texts other than the target text, the step of sending the to-be-processed texts to the current execution server for processing by the current execution server using the large language model, until all target texts corresponding to the current execution server meet the exception handling condition, and obtaining the time required to obtain the processing results of all target texts corresponding to the current execution server using the large language model; wherein, multiple texts satisfying the exception handling condition means that: if the multiple texts are sent to the current execution server, an exception occurs in the processing process of the multiple texts by the current execution server; and returning to the step of determining, from each execution server in accordance with the preset allocation order, the execution server currently used to process the to-be-processed texts.
[0153] Optionally, the duration calculation module 305 is specifically used to, when the first number is not 0, predict for each alternative execution server, based on the first number, how long it will take for the alternative execution server to obtain the processing result of the current specified text for the pre-filling stage by using the large language model if the text to be processed is sent to the alternative execution server; and, when the first number is 0, predict for each alternative execution server, based on the second number, how long it will take for the alternative execution server to obtain the processing result of the current specified text for the decoding stage by using the large language model if the text to be processed is sent to the alternative execution server.
[0154] Optionally, the sending module 306 is specifically used to, when the first number is not 0, if the ratio of the first number to the second number is greater than the first ratio, determine the alternative execution server with the shortest required time as the target execution server; if the ratio of the first number to the second number is not greater than the first ratio, determine the alternative execution server with the longest required time as the target execution server; and when the first number is 0, determine the alternative execution server with the shortest required time as the target execution server.
[0155] Optionally, the device also includes: a screening module, which is used to determine the text that meets the preset processing conditions from the texts that have not been pre-filled as the first number before the acquisition module 301 executes the acquisition, as the text to be processed; wherein the preset processing condition indicates that the time taken to obtain the corresponding processing result is the shortest.
[0156] Optionally, the duration calculation module 305 is specifically used to, when the processing priority of the text to be processed is not less than the preset priority, predict, for each alternative execution server, based on the first number and the second number, the minimum duration required for the alternative execution server to obtain the processing result of the text to be processed for the current processing stage using the large language model if the text to be processed is sent to the alternative execution server; the sending module 306 is specifically used to take the alternative execution server with the shortest required duration as the target execution server, and send indication information and the text to be processed to the target execution server, so that the target execution server migrates the text to be migrated and the data to be migrated to the designated storage space based on the indication information, and then processes the text to be processed based on the large language model; wherein the text to be migrated includes: the text currently being processed by the target execution server, which is in the same processing stage as the text to be processed; and the data to be migrated includes: all data obtained in the process of processing the text to be migrated using the large language model.
[0157] Optionally, the sending module 306 is specifically used to send the text to be processed and the data to be reused to the target execution server when the text to be processed is obtained from any execution server, so that the target execution server continues to process the text to be processed using the data to be reused based on the large language model; wherein the data to be reused is: all data obtained in the process of processing the text to be processed using the large language model.
[0158] Optionally, the prediction module 302 is specifically used to, if the maximum number of text units that can be generated by the large language model has been set and the maximum number of units is not greater than a preset number threshold, determine that the maximum number of units is the number of text units of the target processing result, and obtain the predicted number; if the maximum number of units is not set, and the text to be processed carries a specified number of units indicating the number of text units of the target processing result, or if the maximum number of units is greater than the preset number threshold and the text to be processed carries the specified number of units, determine that the specified number of units is the predicted number; if the maximum number of units is not set and the text to be processed does not carry the specified number of units, or if the maximum number of units is greater than the preset number threshold and the text to be processed does not carry the specified number of units, input the text to be processed into a pre-trained output unit number prediction model to obtain the predicted number; wherein the output unit number prediction model is: trained based on sample text and sample labels indicating the number of text units of the sample processing result of the sample text.
[0159] The embodiment of the present invention also provides a central server, such as Figure 4 As shown, it includes a processor 401, a communication interface 402, a memory 403 and a communication bus 404, wherein the processor 401, the communication interface 402, and the memory 403 communicate with each other through the communication bus 404, and the memory 403 is used to store computer programs; the processor 401 is used to implement the steps of any of the above-mentioned task scheduling methods when executing the program stored in the memory 403.
[0160] The communication bus mentioned above for the central server can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. This communication bus can be divided into address buses, data buses, control buses, etc. For ease of illustration, the figure shows only one thick line, but this does not mean that there is only one bus or only one type of bus.
[0161] The communication interface is used for communication between the above central server and other devices.
[0162] The memory may include random access memory (RAM) or non-volatile memory, such as at least one disk storage. Alternatively, the memory may be at least one storage device located away from the processor.
[0163] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, and discrete hardware components.
[0164] In another embodiment of the present invention, a computer-readable storage medium is provided, wherein a computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor, the task scheduling method described in any one of the above embodiments is implemented.
[0165] In another embodiment of the present invention, a computer program product including instructions is provided. When the computer program product is run on a computer, the computer executes the task scheduling method described in any one of the above embodiments.
[0166] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).
[0167] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.
[0168] Each embodiment in this specification is described in a related manner. Similar portions between the various embodiments can be referenced to each other. Each embodiment focuses on the differences between the other embodiments. In particular, the device, central server, computer-readable storage medium, and computer program product embodiments are generally similar to the method embodiments, so their descriptions are relatively simplified. For related portions, reference can be made to the descriptions of the method embodiments.
[0169] The above description is only a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention are included in the scope of protection of the present invention.
Claims
1. A task scheduling method, characterized in that: The method comprises: Obtaining the number of text units to be pre-filled as a first number; wherein the text units to be pre-filled are: text units that have not been pre-filled among the text units of the text to be processed; a text unit of a text represents the encoding result of the characters in the text; Predicting the number of text units of a target processing result as the predicted number; wherein the target processing result is obtained by pre-filling and decoding the text units of the to-be-processed text based on the large language model; Calculating the difference between the predicted number and the number of generated text units to obtain a second number; wherein the generated text units are: text units generated by decoding processing among the text units of the target processing result; Determine, from among the execution servers, an execution server whose current remaining storage space satisfies the requirement for processing the text to be processed using the large language model as a candidate execution server; For each candidate execution server, based on the first number and the second number, a prediction is made as to how long it would take for the candidate execution server to obtain a processing result of the current designated text at the current processing stage using the large language model if the text to be processed is sent to the candidate execution server; wherein the current designated text is determined from text received from the candidate execution server; the current processing stage of any designated text represents a pre-filling stage or a decoding stage; the time required for an execution server indicates the amount of target resources currently available to the execution server; the target resources are resources of the type required to obtain the processing result of the text to be processed at the current stage; Based on the time required by each alternative execution server, a target execution server is determined from the alternative execution servers, and the text to be processed is sent to the target execution server, so that the target execution server processes the text to be processed based on the large language model using the target resources currently available to the target execution server.
2. The method according to claim 1, characterized in that The step of determining, from among the execution servers, an execution server whose current remaining storage space satisfies the requirement for processing the to-be-processed text using the large language model as a candidate execution server includes: calculating a first space size required for pre-filling the first number of text units using the large language model, calculating a second space size required for generating the second number of text units using the large language model, and calculating a sum of the first space size and the second space size as a total space size; An execution server whose current remaining storage space is not less than the total space size is determined from among the execution servers as a candidate execution server.
3. The method according to claim 1, characterized in that There are multiple texts to be processed; The step of predicting, for each candidate execution server, based on the first number and the second number, a time required for the candidate execution server to obtain a processing result for the current specified text at the current processing stage using the large language model if the text to be processed is sent to the candidate execution server, includes: After each execution server has processed the texts received in the past, the execution server currently used to process the text to be processed is determined from among the execution servers according to the preset allocation order as the current execution server; Based on the first number and the second number of each to-be-processed text, determining, from the other to-be-processed texts except the target text, the to-be-processed text to be sent to the current execution server for processing by the current execution server using the large language model as the target text corresponding to the current execution server; If all target texts corresponding to the current execution server do not meet the exception handling condition, the step of returning to the step of determining, from the other to-be-processed texts other than the target text, the to-be-processed texts to be sent to the current execution server for processing by the current execution server using the large language model based on the first and second numbers of the to-be-processed texts, until all target texts corresponding to the current execution server meet the exception handling condition, and obtaining the time required to obtain the processing results of all target texts corresponding to the current execution server using the large language model; wherein, if multiple texts meet the exception handling condition, it means that if the multiple texts are sent to the current execution server, an exception occurs in the processing process of the multiple texts by the current execution server; Return to the step of determining, from among the execution servers, the execution server currently used to process the text to be processed according to the preset allocation order.
4. The method according to claim 1, wherein The step of predicting, for each candidate execution server, based on the first number and the second number, a time required for the candidate execution server to obtain a processing result for the current specified text at the current processing stage using the large language model if the text to be processed is sent to the candidate execution server, includes: If the first number is not zero, for each candidate execution server, based on the first number, predicting a time required for the candidate execution server to obtain a processing result of the current specified text in the pre-filling phase using the large language model if the text to be processed is sent to the candidate execution server; When the first number is 0, for each alternative execution server, based on the second number, it is predicted that if the text to be processed is sent to the alternative execution server, the alternative execution server will use the large language model to obtain the processing result of the current specified text for the decoding stage.
5. The method according to claim 4, characterized in that The determining of the target execution server from the candidate execution servers based on the time required by each candidate execution server includes: When the first number is not 0, if the ratio of the first number to the second number is greater than the first ratio, determining the candidate execution server with the shortest required execution time as the target execution server; If the ratio of the first number to the second number is not greater than the first ratio, determining the candidate execution server with the longest required execution time as the target execution server; When the first number is 0, the candidate execution server with the shortest required time is determined as the target execution server.
6. The method according to claim 1, wherein Before obtaining the number of text units to be pre-filled as the first number, the method further includes: From the texts that have not been pre-filled, texts that meet preset processing conditions are determined as texts to be processed; wherein the preset processing conditions indicate that the time consumed to obtain the corresponding processing results is the shortest.
7. The method according to claim 1, characterized in that The step of predicting, for each candidate execution server, based on the first number and the second number, a time required for the candidate execution server to obtain a processing result for the current specified text at the current processing stage using the large language model if the text to be processed is sent to the candidate execution server, includes: When the processing priority of the to-be-processed text is not less than a preset priority, predicting, for each candidate execution server, based on the first number and the second number, a minimum time required for the candidate execution server to obtain a processing result for the to-be-processed text at the current processing stage using the large language model if the to-be-processed text is sent to the candidate execution server; The method of determining a target execution server from the candidate execution servers based on the time required by each candidate execution server, and sending the to-be-processed text to the target execution server so that the target execution server processes the to-be-processed text based on the large language model and using target resources currently available to the target execution server, includes: The candidate execution server with the shortest required execution time is selected as the target execution server, and instruction information and the text to be processed are sent to the target execution server, so that the target execution server migrates the text to be migrated and the data to be migrated to the designated storage space based on the instruction information, and then processes the text to be processed based on the large language model; The text to be migrated includes: the text currently being processed by the target execution server, which is in the same processing stage as the text to be processed; the data to be migrated includes: all data obtained in the process of processing the text to be migrated using the large language model.
8. The method according to claim 1, characterized in that Sending the to-be-processed text to the target execution server so that the target execution server processes the to-be-processed text based on the large language model and using target resources currently available to the target execution server includes: When the text to be processed is obtained from any execution server, the text to be processed and the data to be reused are sent to the target execution server, so that the target execution server continues to process the text to be processed using the data to be reused based on the large language model; wherein the data to be reused is: all data obtained in the process of processing the text to be processed using the large language model.
9. The method according to claim 1, characterized in that The number of text units of the predicted target processing result, as the predicted number, includes: If a maximum number of text units that can be generated by the large language model has been set, and the maximum number of units is not greater than a preset number threshold, determining the maximum number of units as the number of text units of the target processing result, and obtaining a predicted number; If the maximum number of units is not set, and the text to be processed carries a specified number of units indicating the number of text units of the target processing result, or if the maximum number of units is greater than the preset number threshold, and the text to be processed carries the specified number of units, determining the specified number of units as the predicted number; If the maximum number of units is not set and the text to be processed does not carry the specified number of units, or if the maximum number of units is greater than the preset number threshold and the text to be processed does not carry the specified number of units, the text to be processed is input into a pre-trained output unit number prediction model to obtain the predicted number; wherein the output unit number prediction model is: trained based on sample text and sample labels indicating the number of text units of the sample processing results of the sample text.
10. A task scheduling device, characterized in that: The device comprises: An acquisition module is configured to acquire the number of text units to be pre-filled as a first number; wherein the text units to be pre-filled are text units that have not been pre-filled among the text units of the text to be processed; a text unit of a text represents an encoding result of characters in the text; A prediction module, configured to predict the number of text units of a target processing result as a predicted number; wherein the target processing result is obtained by pre-filling and decoding the text units of the to-be-processed text based on a large language model; a number calculation module, configured to calculate a difference between the predicted number and the number of generated text units to obtain a second number; wherein the generated text units are text units generated by decoding processing among the text units of the target processing result; A determination module is configured to determine, from among the execution servers, an execution server whose current remaining storage space satisfies the requirement for processing the text to be processed using the large language model, as a candidate execution server; A duration calculation module is configured to predict, for each candidate execution server, based on the first number and the second number, the duration required for the candidate execution server to obtain a processing result of the current specified text for the current processing stage using the large language model if the text to be processed is sent to the candidate execution server; wherein the current specified text is determined from text received from the candidate execution server; the current processing stage of any specified text represents a pre-filling stage or a decoding stage; the duration required for an execution server indicates the amount of target resources currently available to the execution server; the target resources are resources of the type required to obtain the processing result of the text to be processed for the current stage; The sending module is used to determine a target execution server from each alternative execution server based on the time required by each alternative execution server, and send the text to be processed to the target execution server, so that the target execution server processes the text to be processed based on the large language model using the target resources currently available to the target execution server.
11. A central server, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus; Memory for storing computer programs; A processor, configured to implement the method according to any one of claims 1 to 9 when executing a program stored in a memory.
12. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.