Task scheduling method, device and system based on large language model and electronic equipment

Through the task scheduling method based on the large language model, the task run time is predicted and task scheduling is optimized, which solves the problem of low resource utilization caused by resource fragmentation, and realizes efficient task scheduling and resource utilization.

CN120144248APending Publication Date: 2025-06-13GUANGZHOU HUYA TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510212444.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

In the prior art, due to resource fragmentation problems, cluster resource utilization is low, affecting the scheduling efficiency of tasks.

Method used

The task scheduling method based on the large language model is adopted, and the task information and actual operation time-consuming of the tasks to be run and their similar historical tasks are obtained. The large language model is used to predict the running time of the tasks to be run, and the target task running node is determined from the non-idle task running nodes to achieve efficient task scheduling.

Benefits of technology

It significantly improves the overall utilization rate of system resources, optimizes the execution efficiency of tasks, and avoids the low utilization rate caused by resource fragmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144248A_ABST
    Figure CN120144248A_ABST
Patent Text Reader

Abstract

The invention provides a task scheduling method, device and system based on a large language model and electronic equipment. The method comprises the steps of obtaining task information of a to-be-run task and a plurality of historical tasks similar to the to-be-run task and actual running time consumption of the historical tasks; predicting the operation time of the to-be-operated task by the large language model based on the task information and the actual operation time consumption; and determining a target task operation node from the non-idle task operation nodes based on the operation time, and scheduling the to-be-operated task to the target task operation node. The problem of resource fragmentation can be solved in the task scheduling process, and the resource utilization rate is increased.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular, to a task scheduling method, device, system, and electronic device based on a large language model. Background Art

[0002] With the rapid development of artificial intelligence and deep learning technologies, more and more enterprises and research institutions have begun to use these technologies to solve practical problems. The training of deep learning models usually requires a large amount of resources, especially high-performance GPUs. Therefore, in order to efficiently utilize these resources, enterprises often build dedicated clusters to support the training of deep learning tasks.

[0003] In the prior art, the resources on each machine are fixed, but the resource requirements and running times of different training tasks vary. For example, some small-scale training tasks may only require a small amount of GPU resources to complete, while large-scale training tasks require more resources and longer running times. Since the number and types of tasks running on each machine are different, it may lead to a situation where the resources on some machines are completely occupied, while there are some idle resources on other machines, that is, resource fragmentation. In addition, due to the different running times of tasks, when a task is completed, the resources it occupies may not be released immediately, but only after another task is completed. This further exacerbates the problem of resource fragmentation, resulting in low overall resource utilization and affecting the task scheduling efficiency.

[0004] Therefore, how to solve the problem of resource fragmentation and improve the utilization rate of cluster resources is a technical problem to be solved. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to provide a task scheduling method, device, system, and electronic device based on a large language model to solve the problem of resource fragmentation and improve the utilization rate of cluster resources. To achieve the above purpose, the technical solution adopted by the present invention is as follows:

[0006] In a first aspect, the present invention provides a task scheduling method based on a large language model, the method comprising: obtaining the task information of a to-be-run task and a plurality of historical tasks similar to the to-be-run task, and the actual running duration of the historical tasks; predicting the running time of the to-be-run task by a large language model based on the task information and the actual running duration; determining a target task running node from non-idle task running nodes based on the running time, and scheduling the to-be-run task to the target task running node.

[0007] In an alternative embodiment, the large language model predicts the running time of the to-be-run task based on the task information and the actual running time consumption, including: constructing a prompt text based on the task information and the actual running time consumption; inputting the prompt text into the large language model to obtain the running time of the to-be-run task.

[0008] In an alternative embodiment, the method further includes: training the large language model according to the actual running time consumption of the to-be-run task.

[0009] In an alternative embodiment, the method further includes: making a fine-tuning training set according to the task information and the actual running time consumption of historical tasks, and performing fine-tuning training on the large language model through the fine-tuning training set.

[0010] In an alternative embodiment, obtaining the task information of the to-be-run task and a plurality of historical tasks similar to the to-be-run task, and the actual running time consumption of the historical tasks, includes: obtaining the vector corresponding to the to-be-run task; retrieving a plurality of target vectors most similar to the vector from a preset vector database; wherein, the vector database contains vectors corresponding to each historical task; taking the historical tasks corresponding to the target vectors as the historical tasks similar to the to-be-run task.

[0011] In an alternative embodiment, the method further includes: obtaining the vectors of the to-be-run task and the historical tasks respectively through an encoding model.

[0012] In an alternative embodiment, determining a target task running node from non-idle task running nodes based on the running time includes: determining non-idle task running nodes that meet the resources required by the to-be-run task and the latest running end time on the non-idle task running nodes; taking the non-idle task running node corresponding to the latest running end time with the smallest time difference from the running time as the target task running node.

[0013] In a second aspect, the present invention provides a task scheduling device based on a large language model, including: an obtaining module, configured to obtain the task information of the to-be-run task and a plurality of historical tasks similar to the to-be-run task, and the actual running time consumption of the historical tasks; a prediction module, configured to predict the running time of the to-be-run task by the large language model based on the task information and the actual running time consumption; a scheduling module, configured to determine a target task running node from non-idle task running nodes based on the running time.

[0014] In a third aspect, the present invention provides a task scheduling system based on a large language model, including a scheduling node and a plurality of task running nodes; the task running nodes are used to run tasks using system resources; the scheduling node is used to execute the task scheduling method based on the large language model according to any one of the foregoing embodiments.

[0015] In a fourth aspect, the present invention provides an electronic device, including a processor and a memory, the memory stores machine-executable instructions that can be executed by the processor, and the processor can execute the machine-executable instructions to implement the task scheduling method based on the large language model according to any one of the foregoing embodiments.

[0016] The task scheduling method, device, system and electronic device based on the large language model provided by the present invention, the method includes: obtaining the task information of the task to be run and a plurality of historical tasks similar to the task to be run, and the actual running time of the historical tasks; predicting the running time of the task to be run by the large language model based on the task information and the actual running time; determining a target task running node from the non-idle task running nodes based on the running time, and scheduling the task to be run to the target task running node. Different from the prior art, the embodiment of the present invention preferentially searches for scheduling targets from non-idle nodes, makes full use of the fragmented resources on these nodes, and avoids the problem of low utilization rate caused by resource fragmentation. This method not only significantly improves the overall utilization rate of system resources, but also optimizes the execution efficiency of tasks.

[0017] To make the above objects, features and advantages of the present invention more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, makes a detailed description as follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0019] Figure 1 FIG. is a schematic diagram of an existing cluster architecture.

[0020] Figure 2 FIG. is a schematic flowchart of the task scheduling method based on the large language model provided by the embodiment of the present invention.

[0021] Figure 3 FIG. is an example diagram of a prompt text provided by the embodiment of the present invention.

[0022] Figure 4A schematic diagram of the principle of the task scheduling method based on a large language model provided by an embodiment of the present invention.

[0023] Figure 5 A schematic structural diagram of the task scheduling system based on a large language model provided by an embodiment of the present invention.

[0024] Figure 6 A functional module diagram of the task scheduling device based on a large language model provided by an embodiment of the present invention.

[0025] Figure 7 A structural block diagram of the electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0026] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Usually, the components of the embodiments of the present invention described and illustrated here can be arranged and designed in various different configurations.

[0027] Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the present invention claimed, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.

[0028] It should be noted that relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0029] Please refer to Figure 1 , Figure 1 which is a schematic diagram of an existing cluster architecture. In the Figure 1 cluster shown, there are several machines, and each machine is equipped with various resources, and these resources include but are not limited to at least one of GPU, CPU, memory and disk.

[0030] Taking the GPU as an example, assume that each machine in the cluster has 8 GPUs, and the tasks running on each task execution node occupy 7 GPUs. At this time, if a task that requires 2 GPUs on a single machine needs to be scheduled, although the overall resources of the cluster are sufficient, since there is no machine with 2 idle GPUs on a single machine, the task cannot be scheduled. This is the problem of resource fragmentation, and this phenomenon reduces the utilization rate of cluster resources and the task scheduling efficiency.

[0031] To solve the above technical problems, the embodiments of the present invention provide a task scheduling method based on a large language model, which can solve the problem of resource fragmentation and make full use of cluster resources to improve task scheduling efficiency.

[0032] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of the task scheduling method based on a large language model provided by the embodiments of the present invention. The execution subject of this method can be an electronic device, including steps S201 to S203, which are described as follows:

[0033] S201: Obtain the task information of the task to be run and multiple historical tasks similar to the task to be run, and the actual running time of the historical tasks.

[0034] In the embodiments of the present invention, the task information may include, but is not limited to: running instructions, running configurations, etc. Among them, the running configuration may include, but is not limited to: running code, code repository, code version, required resources, etc.

[0035] S201: The large language model predicts the running time of the task to be run based on the task information and the actual running time.

[0036] In the embodiments of the present invention, the large language model can directly use an open-source large language model, or it can be a large language model obtained by fine-tuning a basic large language model in the embodiments of the present invention.

[0037] S202: Determine the target task execution node from the non-idle task execution nodes based on the running time, and schedule the task to be run to the target task execution node.

[0038] The task scheduling method based on a large language model provided by an embodiment of the present invention. First, in step S201, the task information and actual running time of the task to be run and its similar historical tasks are first obtained, providing a rich data basis and reliable reference for subsequent prediction and scheduling. Subsequently, in step S202, the large language model is used to quickly and accurately predict the running time of the task to be run. Based on this prediction result, in step S203, a target task running node is intelligently selected from the non-idle task running nodes for task scheduling. Different from the prior art, the embodiment of the present invention preferentially searches for scheduling targets among non-idle nodes, making full use of the fragmented resources on these nodes and avoiding the problem of low utilization caused by resource fragmentation. This method not only significantly improves the overall utilization rate of system resources but also optimizes the execution efficiency of tasks.

[0039] Next, the embodiments of the present invention will introduce the above steps in detail and clearly in combination with relevant drawings.

[0040] In step S201, the electronic device first obtains a plurality of historical tasks similar to the task to be run. As an example, the embodiment of the present invention will evaluate the similarity between the task to be run and the historical tasks based on task information.

[0041] In order to quantify the similarity between the task to be run and each historical task, the embodiment of the present invention adopts an implementation manner of vectorizing information such as running instructions and running configurations in the task information. Relevant personnel can also adopt other quantification methods, which are not limited herein.

[0042] In an embodiment of the present invention, in order to improve the retrieval efficiency, the running instructions and running configurations of each historical task can be first converted into vectors and stored in a vector database. Then, when the electronic device obtains the vector corresponding to the task information of the task to be run, k target vectors most similar to the vector of the task to be run can be retrieved from the vector database, and thus a plurality of similar historical tasks can be obtained.

[0043] In an embodiment of the present invention, in addition to maintaining the vectors corresponding to each historical task, the vector database can also maintain the actual running time corresponding to each historical task. In this way, the most similar k historical tasks can be quickly obtained, improving the efficiency of predicting the running time of the task to be run subsequently.

[0044] Optionally, the value of k can be flexibly set by relevant personnel according to actual needs, which is not limited herein.

[0045] In an embodiment of the present invention, in order to quickly and accurately convert task information into vectors, the embodiment of the present invention can adopt a frozen encoding model to obtain the vectors of the task to be run and the historical tasks respectively through the encoding model.

[0046] In an alternative embodiment, the encoding model can be, but is not limited to, a Bert model.

[0047] In an embodiment of the present invention, in order to improve the computing speed, when converting the running instructions and running configurations into vectors, the vector dimension can be set, for example, 1024 dimensions, to reduce the vector length and lower the processing complexity and time consumption.

[0048] Based on the above, an embodiment of step S201 provided by the embodiments of the present invention may include the following steps a1 to a3:

[0049] Step a1: Obtain the vector corresponding to the task to be run;

[0050] Step a2: Retrieve multiple target vectors that are most similar to the vector from a preset vector database; wherein, the vector database contains the vectors corresponding to each historical task;

[0051] Step a3: Use the historical task corresponding to the target vector as the historical task similar to the task to be run.

[0052] Through the above embodiment, the embodiments of the present invention can quickly and accurately retrieve multiple historical tasks that are most similar to the task to be run. According to the task information and actual running time of these historical tasks, the embodiments of the present invention can accurately predict the running time of the task to be run through a large language model, see step 202.

[0053] In step S202, the large language model can be a fine-tuned basic large language model, which can be, but is not limited to, Qwen2.5.

[0054] In an embodiment of the present invention, a fine-tuning training set is made according to the task information and actual running time of historical tasks, and the large language model is fine-tuned through the fine-tuning training set, and then the fine-tuned large language model is applied to the embodiments of the present invention.

[0055] In order to achieve the effect of allowing the large language model to predict the task running time, the electronic device can first construct a prompt text, and the large language model completes the prediction function based on the prompt text. Therefore, step S202 in the embodiments of the present invention may include the following steps:

[0056] Step b1: Construct a prompt text based on the task information and actual running time;

[0057] Step b2: Input the prompt text into the large language model to obtain the running time of the task to be run.

[0058] In an alternative embodiment, when the task to be run and similar historical tasks use the same code repository, task code differences can also be considered during the process of constructing the prompt text.

[0059] Optionally, the content of the task code differences can be generated by git diff. Git diff is a command provided by Git for comparing differences between code files or directories. It shows the differences between two versions, including which lines have been modified, which lines have been deleted, and which lines have been added. By running the git diff command, Git outputs a diff report that details the changes between the codes. You can save this report or view it directly in the terminal.

[0060] For ease of understanding, please refer to Figure 3 , Figure 3 which is a sample diagram of the prompt text provided by an embodiment of the present invention. Among them, Task A is the task to be run, and Task B and Task C are the two historical tasks most similar to Task A. Based on the running instructions, running configurations, code differences of Tasks A, B, and C, and the differences between the codes of Tasks B and C and the code of Task A, the prompt text is constructed and then provided to the large language model, which can predict the running time of Task A based on the context information of Tasks B and C.

[0061] It should be understood that Figure 3 the text structure of the shown prompt text is merely an example and not a limitation on its structure or composition. Relevant personnel can flexibly design according to the actual requirements of the large language model, which is not limited here.

[0062] In the embodiment of the present invention, after obtaining the running time of the task to be run, through step S203, the embodiment of the present invention can preferentially find task running nodes that meet the resources required by the task to be run from non-idle task running nodes, reasonably utilize the fragmented resources on the non-idle task running nodes, and improve the resource utilization rate.

[0063] In step S203, in order to find the target task running node, the embodiment of the present invention provides the following embodiments:

[0064] Step c1: Determine non-idle task running nodes that meet the resources required by the task to be run and the latest running end time on the non-idle task running nodes;

[0065] Step c2: Take the non-idle task running node corresponding to the latest running end time with the smallest time difference from the running time as the target task running node.

[0066] In the embodiments of the present invention, the remaining resources of all non-idle task running nodes in the cluster can be obtained first, and then the non-idle task running nodes with remaining resources greater than or equal to the resources required by the to-be-run task can be filtered out. For the non-idle task running nodes that meet the resources required by the to-be-run task, each task running on them has a predicted end time. Find the latest running end time, and the task with the lowest difference between the latest running end time and the running time of the to-be-run task is the closest to the completion time of the to-be-run task, which can make the resources of the task running node where the task is located become available almost simultaneously for subsequent task scheduling, thus solving the problem of resource fragmentation and making full use of the cluster resources and improving the resource utilization rate.

[0067] After finding the target task running node through the above implementation manner, the task scheduling node in the cluster can be notified to schedule the to-be-run task to the target task running node for processing.

[0068] It can be understood that the embodiments of the present invention preferentially find the task running nodes that can process the to-be-run task from the non-idle task running nodes. If there is no such node, the to-be-run task is scheduled to any idle task running node for processing, which can avoid the fragmentation of resources in the non-idle task running nodes and ensure that the to-be-run task can be successfully processed at the same time.

[0069] In an embodiment of the present invention, after the target task running node processes the to-be-run task, the actual running duration of the to-be-run task can also be obtained. Combining with the task information of the to-be-run task, the currently used large language model can be trained online to improve the prediction accuracy of the large language model.

[0070] For the convenience of overall understanding of the task scheduling method based on the large language model provided by the embodiments of the present invention, please refer to Figure 4 , Figure 4 which is a schematic diagram of the principle of the task scheduling method based on the large language model provided by the embodiments of the present invention.

[0071] In Figure 4 , first, the running instructions and running configurations of historical tasks are converted into vectors through an encoding model, and the vectors and the actual running durations of historical tasks are stored in a vector database. For the to-be-run task, the corresponding running instructions and running configurations are converted into vectors through the encoding model, and then the top-k target vectors and their actual running durations are retrieved from the vector database. Based on the running instructions, running configurations, and actual running durations corresponding to the target vectors and the running instructions and running configurations of the to-be-run task, a prompt text is constructed and then input into the large language model, and the large language model outputs an answer, that is, the running time of the to-be-run task.

[0072] In summary, the task scheduling method based on a large language model provided by the embodiments of the present invention has the following advantages: First, through the vector retrieval method, the historical task most similar to the task to be run can be obtained. Using the task information and actual running time of the similar historical task, the large language model can accurately predict the actual running of the task to be run, and the prediction result is reliable. Second, the embodiments of the present invention use a large language model to predict the running time of the task to be run, and the large language model is fine-tuned based on the task information of historical tasks, so the prediction result has high accuracy. Through the embodiments of the present invention, the waiting time for users to submit tasks is significantly reduced, and the resource idle rate of the cluster is also significantly reduced.

[0073] Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of a task scheduling system based on a large language model provided by the embodiments of the present invention. The task scheduling system 50 based on a large language model includes: a plurality of task running nodes 501 and a scheduling node 502. The task running nodes 501 are used to run tasks using system resources; the scheduling node 502 is used to execute the task scheduling method based on a large language model provided by the embodiments of the present invention.

[0074] Based on the same inventive concept as Figure 2 the embodiments of the present invention also provide a task scheduling device 60 based on a large language model. Please refer to Figure 6 , Figure 6 which is a functional module diagram of the task scheduling device based on a large language model provided by the embodiments of the present invention, including: an obtaining module 601, a prediction module 602, and a scheduling module 603.

[0075] The obtaining module 601 is used to obtain the task information of the task to be run and multiple historical tasks similar to the task to be run, and the actual running time of the historical tasks;

[0076] The prediction module 602 is used to predict the running time of the task to be run by the large language model based on the task information and the actual running time;

[0077] The scheduling module 603 is used to determine the target task running node from the non-idle task running nodes based on the running time.

[0078] It can be understood that the obtaining module 601, the prediction module 602, and the scheduling module 603 can cooperate to execute Figure 2 each step in

[0079] to achieve the corresponding technical effects.

[0080] In an alternative embodiment, the task scheduling device based on the large language model may further include a training module for training the large language model according to the actual running time of the task to be run.

[0081] In an alternative embodiment, the training module is further configured to create a fine-tuning training set according to the task information and the actual running time of historical tasks, and perform fine-tuning training on the large language model through the fine-tuning training set.

[0082] In an alternative embodiment, the obtaining module 601 is specifically configured to obtain a vector corresponding to the task to be run; retrieve multiple target vectors most similar to the vector from a preset vector database, where the vector database contains vectors corresponding to each historical task; and use the historical tasks corresponding to the target vectors as the historical tasks similar to the task to be run.

[0083] In an alternative embodiment, the obtaining module 601 is further specifically configured to obtain the vectors of the task to be run and the historical tasks respectively through an encoding model.

[0084] In an alternative embodiment, the scheduling module 603 is specifically configured to determine non-idle task running nodes that meet the resources required for the task to be run and the latest running end time on the non-idle task running nodes; and use the non-idle task running node corresponding to the latest running end time with the smallest time difference from the running time as the target task running node.

[0085] It should be noted that the division of modules in the above embodiments of the present application is illustrative, only a logical function division, and there may be other division methods in actual implementation. In addition, in each embodiment of the present application, each functional unit may be integrated in a processing unit, may exist separately physically, or two or more units may be integrated in one unit. The above integrated units may be implemented in the form of hardware or in the form of software functional units.

[0086] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing an electronic device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods in various embodiments of this application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program executable code.

[0087] Embodiments of the present invention also provide an electronic device. Please refer to Figure 7 , Figure 7 which is a structural block diagram of the electronic device provided by the embodiments of the present invention, including: a memory 701, a processor 702, and a communication interface 703. The memory 701, the processor 702, and the communication interface 703 are electrically connected directly or indirectly to each other to achieve data transmission or interaction. For example, these components can be electrically connected to each other through one or more communication buses or signal lines.

[0088] Optionally, the bus 704 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 7 only a thick line is used to represent it in

[0089] In the embodiments of the present invention, the processor 702 can be a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field programmable gate array, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor can be a microprocessor or any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present invention can be directly embodied as being executed by a hardware processor, or can be executed by a combination of hardware and software modules in the processor. The software module can be located in the memory 701, and the processor 702 reads the program instructions in the memory 701 and combines its hardware to complete the steps of the above method.

[0090] In an embodiment of the present invention, the memory 701 may be a non-volatile memory, such as a hard disk drive (HDD) or a solid-state drive (SSD), etc., or may also be a volatile memory, such as RAM. The memory may also be any other medium that can be used to carry or store the desired program executable code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory in the embodiment of the present invention may also be a circuit or any other device capable of implementing a storage function, for storing instructions and / or data.

[0091] The memory 701 can be used to store software programs and modules, such as the instructions / modules of the task scheduling device 60 based on a large language model provided in the embodiment of the present invention. They can be stored in the memory 701 in the form of software or firmware, or solidified in the operating system (OS) of the electronic device 70. The processor 702 executes various functional applications and data processing by executing the software programs and modules stored in the memory 701. The communication interface 703 can be used for signaling or data communication with other node devices.

[0092] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described devices and units can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0093] It can be understood that Figure 7 The structure shown is only schematic, and the electronic device 70 may also include more or fewer components than those shown Figure 7 herein, or have a different configuration from that shown Figure 7 herein. Figure 7 Each component shown can be implemented by hardware, software, or a combination thereof.

[0094] The electronic device 70 may also include a network device and / or a user device. Among them, the network device includes, but is not limited to, a single network server, a server group composed of multiple network servers, or a cloud composed of a large number of hosts or network servers based on cloud computing.

[0095] Based on the above embodiments, the present application also provides a storage medium. A computer program is stored in the computer-readable storage medium. When the computer program is executed by the computer, the computer executes the task scheduling method based on a large language model provided in the above embodiments.

[0096] Based on the above embodiments, an embodiment of the present invention further provides a computer program. When the computer program runs on a computer, it causes the computer to execute the task scheduling method based on a large language model provided in the above embodiments.

[0097] Based on the above embodiments, an embodiment of the present invention further provides a chip. The chip is used to read the computer program stored in a memory and execute the task scheduling method based on a large language model provided in the above embodiments.

[0098] An embodiment of the present invention also provides a computer program product, including instructions. When the instructions run on a computer, they cause the computer to execute the task scheduling method based on a large language model provided in the above embodiments.

[0099] Embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by instructions. These instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.

[0100] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing devices to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.

[0101] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, such that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable devices provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.

[0102] The above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A task scheduling method based on a large language model, characterized in that: The method comprises: Obtaining task information of the task to be run and a plurality of historical tasks similar to the task to be run, and actual running time of the historical tasks; Predicting the running time of the task to be run by a large language model based on the task information and the actual running time; A target task running node is determined from non-idle task running nodes based on the running time, and the task to be run is scheduled to the target task running node.

2. The task scheduling method based on a large language model according to claim 1, characterized in that: Predicting the running time of the task to be run by the large language model based on the task information and the actual running time, including: Constructing a prompt text based on the task information and the actual running time; The prompt text is input into the large language model to obtain the running time of the task to be run.

3. The task scheduling method based on a large language model according to claim 1 or 2, characterized in that: The method further comprises: The large language model is trained according to the actual running time of the task to be run.

4. The task scheduling method based on a large language model according to claim 1 or 2, characterized in that: The method further comprises: A fine-tuning training set is prepared according to task information of historical tasks and actual running time, and fine-tuning training is performed on the large language model through the fine-tuning training set.

5. The task scheduling method based on a large language model according to claim 1, characterized in that: Obtaining task information of the task to be run and a plurality of historical tasks similar to the task to be run, and actual running time of the historical tasks, including: Obtaining a vector corresponding to the task to be run; Retrieving multiple target vectors most similar to the vector from a preset vector database; wherein the vector database contains vectors corresponding to each historical task; The historical task corresponding to the target vector is regarded as a historical task similar to the task to be executed.

6. The task scheduling method based on a large language model according to claim 5, characterized in that: The method further comprises: The vectors of the task to be executed and the historical task are obtained through the encoding model.

7. The task scheduling method based on a large language model according to claim 1, characterized in that: Determining a target task running node from non-idle task running nodes based on the running time includes: Determine a non-idle task running node that meets the resources required by the task to be run and the latest running end time on the non-idle task running node; The non-idle task running node corresponding to the latest running end time having the smallest time difference with the running time is used as the target task running node.

8. A task scheduling device based on a large language model, characterized in that: include: An acquisition module, used to obtain task information of the task to be run and a plurality of historical tasks similar to the task to be run, and actual running time of the historical tasks; A prediction module, configured to predict the running time of the task to be run based on the task information and the actual running time by using a large language model; The scheduling module is used to determine the target task running node from the non-idle task running nodes based on the running time.

9. A task scheduling system based on a large language model, characterized in that: It comprises a scheduling node and several task running nodes; the task running node is used to run tasks using system resources; the scheduling node is used to execute the task scheduling method based on a large language model as described in any one of claims 1 to 7.

10. An electronic device, characterized in that: It includes a processor and a memory, wherein the memory stores machine executable instructions that can be executed by the processor, and the processor can execute the machine executable instructions to implement the task scheduling method based on a large language model as described in any one of claims 1-7.

Citation Information

Cited By

  • Large language model application workload scheduling method, system and equipment

    CN120803669A