Processor resource scheduling method and device, electronic equipment and medium
By obtaining the execution sequence information of the task node, the available resource information of the processor cluster and the resource requirement information of the task node, flexible processor resource scheduling for multiple task nodes is achieved, and the problem of low processor resource utilization in the prior art is solved, and the resource utilization rate and processor idle time utilization is improved.
Patent Information
- Application Number
- CN202510246197.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-06-20
AI Technical Summary
The prior art is difficult to effectively utilize processor resources, especially in complex computing tasks, resulting in waste of resources and low utilization.
By obtaining the execution sequence information of the task node, the available resource information of the processor cluster, and the resource requirement information of the task node, flexible processor resource scheduling for multiple task nodes is realized.
It improves the resource utilization rate of the processor cluster, effectively utilizes the free time of the processor, and reduces resource waste.
Smart Images

Figure CN120179357A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and in particular, to the fields of data processing and cloud computing technologies. Specifically, the present disclosure relates to a method, an apparatus, an electronic device, a computer-readable storage medium, and a computer program product for scheduling processor resources. Background Art
[0002] With the development of artificial intelligence and large model technologies, an increasing number of complex computing tasks have greatly increased the demand for processor resources. Therefore, how to efficiently utilize limited processor resources has become an urgent problem to be solved.
[0003] Currently, for distributed tasks, they can be split into multiple subtasks, and different priorities can be set for each subtask, so as to perform resource scheduling in an orderly manner based on the priorities of the tasks.
[0004] The methods described in this section are not necessarily methods that have been previously conceived or adopted. Unless otherwise specified, any method described in this section should not be considered prior art merely because it is included in this section. Similarly, unless otherwise specified, the problems mentioned in this section should not be considered to have been recognized in any prior art. Summary of the Invention
[0005] The present disclosure provides a method, an apparatus, an electronic device, a computer-readable storage medium, and a computer program product for scheduling processor resources.
[0006] According to an aspect of the present disclosure, there is provided a method for scheduling processor resources, including: obtaining execution order information of a plurality of task nodes, where at least two task nodes among the plurality of task nodes need to be executed in parallel; obtaining available resource information of a processor cluster, where the available resource information indicates the resource margin of each processor among a plurality of processors included in the processor cluster; determining resource requirement information, where the resource requirement information indicates the required execution time and required computing resources of each task node among the plurality of task nodes; and performing processor resource scheduling for the plurality of task nodes according to the execution order information, the available resource information, and the resource requirement information.
[0007] According to another aspect of the present disclosure, there is provided a scheduling device for processor resources, including: a first acquisition module configured to acquire execution order information of a plurality of task nodes, wherein at least two of the plurality of task nodes need to be executed in parallel; a second acquisition module configured to acquire available resource information of a processor cluster, wherein the available resource information indicates the resource margin of each of the plurality of processors included in the processor cluster; a determination module configured to determine resource requirement information, wherein the resource requirement information indicates the required execution time and required computing resources of each of the plurality of task nodes; and a first scheduling module configured to perform processor resource scheduling for the plurality of task nodes according to the execution order information, the available resource information, and the resource requirement information.
[0008] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the above method.
[0009] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the above method.
[0010] According to another aspect of the present disclosure, there is provided a computer program product including a computer program, wherein when the computer program is executed by a processor, the above method is implemented.
[0011] According to one or more embodiments of the present disclosure, there is provided a scheduling method for processor resources. For a plurality of task nodes with an execution order, by introducing the specific execution time of each task node and combining the execution order of the plurality of task nodes, the computing resources required by each task node, and the resource margin of each processor, more flexible allocation of processor resources is achieved, so that the idle time of the processor can be effectively utilized, and the resource utilization rate of the processor cluster is improved.
[0012] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings
[0013] The accompanying drawings exemplarily illustrate embodiments and form part of the specification, and are used together with the written description of the specification to explain the exemplary embodiments of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to similar but not necessarily identical elements.
[0014] Figure 1 is a schematic diagram showing an example system in which various methods described herein can be implemented according to an exemplary embodiment;
[0015] Figure 2 shows a flowchart of a method for scheduling processor resources according to an embodiment of the present disclosure;
[0016] Figure 3 shows an exemplary schematic diagram of task nodes of a face fusion task workflow with a complex execution order;
[0017] Figure 4 shows a partial flowchart of another method for scheduling processor resources according to an embodiment of the present disclosure;
[0018] Figure 5 shows a partial flowchart of another method for scheduling processor resources according to an embodiment of the present disclosure;
[0019] Figure 6 shows a partial flowchart of another method for scheduling processor resources according to an embodiment of the present disclosure;
[0020] Figure 7 shows a structural block diagram of a device for scheduling processor resources according to an embodiment of the present disclosure; and
[0021] Figure 8 shows a structural block diagram of an exemplary electronic device that can be used to implement the embodiments of the present disclosure. Detailed Embodiments
[0022] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0023] In the present disclosure, unless otherwise specified, the use of terms such as "first" and "second" to describe various elements is not intended to limit the positional relationship, temporal relationship, or importance relationship of these elements. Such terms are only used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of the element, and in certain cases, based on the context description, they may also refer to different instances.
[0024] In the description of the various examples in the present disclosure, the terms used are only for the purpose of describing specific examples and are not intended to be limiting. Unless the context clearly indicates otherwise, if the number of elements is not specifically limited, the element may be one or more. In addition, the term "and / or" used in the present disclosure covers any one of the listed items and all possible combinations.
[0025] In the related art, for distributed tasks with an execution order, they can be split into multiple subtasks for resource scheduling respectively. However, this method cannot ensure that the resource scheduling of all subtasks can succeed simultaneously, and it is difficult to ensure the timing of subtask execution.
[0026] In the related art, the corresponding execution priority can be set for each subtask according to the execution order of the subtasks, and the timing of the subtasks can be ensured by synchronously scheduling resources for the subtasks with the same execution priority. However, this method does not consider the specific execution time of each subtask in the subtasks with the same execution priority. Therefore, it may occur that before scheduling resources for the subtasks of the next execution priority, some processors with the current execution priority need to idle and wait for other subtasks with the same execution priority to continue execution after completing the corresponding subtasks, resulting in a large amount of resource waste.
[0027] To solve the above problems, the present disclosure provides a method for scheduling processor resources. For multiple task nodes with an execution order, by introducing the specific execution time of each task node and combining the execution order of the multiple task nodes, the computing resources required by each task node, and the resource margin of each processor, more flexible processor resource allocation is achieved, so that the idle time of the processor can be effectively utilized and the resource utilization rate of the processor cluster can be improved.
[0028] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0029] Figure 1 FIG. shows a schematic diagram of an exemplary system 100 in which the various methods and apparatuses described herein can be implemented according to an embodiment of the present disclosure. Refer to Figure 1, the system 100 includes one or more client devices 101, 102, 103, 104, 105, and 106, a server 120, and one or more communication networks 110 that couple the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105, and 106 may be configured to execute one or more applications.
[0030] In an embodiment of the present disclosure, the server 120 may run one or more services or software applications that enable a scheduling method of processor resources to be executed.
[0031] In certain embodiments, the server 120 may also provide other services or software applications that may include a non-virtual environment and a virtual environment. In certain embodiments, these services may be provided as web-based services or cloud services, such as provided to users of the client devices 101, 102, 103, 104, 105, and / or 106 under a software as a service (SaaS) model.
[0032] In Figure 1 In the configuration shown, the server 120 may include one or more components that implement the functions performed by the server 120. These components may include software components, hardware components, or a combination thereof that may be executed by one or more processors. Users operating the client devices 101, 102, 103, 104, 105, and / or 106 may in turn utilize one or more client applications to interact with the server 120 to utilize the services provided by these components. It should be understood that various different system configurations are possible, which may be different from the system 100. Therefore, Figure 1 is an example of a system for implementing the methods described herein and is not intended to be limiting.
[0033] Users may use the client devices 101, 102, 103, 104, 105, and / or 106 to execute a scheduling method of processor resources. The client device may provide an interface that enables a user of the client device to interact with the client device. The client device may also output information to the user via the interface. Although Figure 1 only six client devices are depicted, those skilled in the art will be able to understand that the present disclosure may support any number of client devices.
[0034] Client devices 101, 102, 103, 104, 105, and / or 106 may include various types of computing devices such as portable handheld devices, general-purpose computers (such as personal computers and laptop computers), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices, etc. These computing devices may run various types and versions of software applications and operating systems such as MICROSOFT Windows, APPLE iOS, UNIX-like operating systems, Linux, or Linux-like operating systems (such as GOOGLE Chrome OS); or include various mobile operating systems such as MICROSOFT WindowsMobile OS, iOS, Windows Phone, Android. Portable handheld devices may include cellular phones, smartphones, tablets, personal digital assistants (PDAs), etc. Wearable devices may include head-mounted displays (such as smart glasses) and other devices. Gaming systems may include various handheld gaming devices, Internet-enabled gaming devices, etc. Client devices are capable of executing various different applications such as various Internet-related applications, communication applications (such as email applications), short message service (SMS) applications, and may use various communication protocols.
[0035] Network 110 may be any type of network known to those skilled in the art, which may support data communication using any one of a variety of available protocols (including but not limited to TCP / IP, SNA, IPX, etc.). By way of example only, one or more networks 110 may be a local area network (LAN), an Ethernet-based network, token ring, wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a public switched telephone network (PSTN), an infrared network, a wireless network (such as Bluetooth, WIFI), and / or any combination of these and / or other networks.
[0036] Server 120 may include one or more general-purpose computers, dedicated server computers (such as PC (personal computer) servers, UNIX servers, midrange servers), blade servers, mainframe computers, server clusters, or any other suitable arrangement and / or combination. Server 120 may include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization (such as one or more flexible pools of logical storage devices that may be virtualized to maintain virtual storage devices for the server). In various embodiments, server 120 may run one or more services or software applications that provide the functions described below.
[0037] The computing unit in server 120 can run one or more operating systems including any of the above-mentioned operating systems and any commercially available server operating systems. Server 120 can also run any one of a variety of additional server applications and / or middleware applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.
[0038] In some embodiments, server 120 can include one or more applications to analyze and merge data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105, and 106. Server 120 can also include one or more applications to display data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105, and 106.
[0039] In some embodiments, server 120 can be a server of a distributed system or a server incorporating a blockchain. Server 120 can also be a cloud server or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. A cloud server is a host product in the cloud computing service system to address the defects of difficult management and weak business scalability existing in traditional physical hosts and virtual private server (VPS) services.
[0040] System 100 can also include one or more databases 130. In certain embodiments, these databases can be used to store task node data and other information. For example, one or more of databases 130 can be used to store information such as audio files and video files. Databases 130 can reside in various locations. For example, the databases used by server 120 can be local to server 120 or can be remote from server 120 and can communicate with server 120 via a network-based or dedicated connection. Databases 130 can be of different types. In certain embodiments, the databases used by server 120 can be, for example, relational databases. One or more of these databases can store, update, and retrieve data to and from the databases in response to commands.
[0041] In certain embodiments, one or more of databases 130 can also be used by applications to store application data. The databases used by applications can be different types of databases, such as key-value repositories, object repositories, or conventional repositories supported by a file system.
[0042] Figure 1 System 100 can be configured and operated in various ways to enable the application of the various methods and apparatuses described in this disclosure.
[0043] Figure 2 The flowchart of a method for scheduling processor resources according to an embodiment of the present disclosure is shown.
[0044] As Figure 2 shown, the method 200 for scheduling processor resources includes:
[0045] Step 210, obtaining execution order information of multiple task nodes, where at least two task nodes among the multiple task nodes need to be executed in parallel;
[0046] Step 220, obtaining available resource information of a processor cluster, where the available resource information indicates the resource margin of each processor among the multiple processors included in the processor cluster;
[0047] Step 230, determining resource requirement information, where the resource requirement information indicates the required execution time and required computing resources of each task node among the multiple task nodes; and
[0048] Step 240, performing processor resource scheduling for the multiple task nodes according to the execution order information, available resource information, and resource requirement information.
[0049] Thus, for multiple task nodes with an execution order, by introducing the specific execution time of each task node, and combining the execution order of the multiple task nodes, the computing resources required by each task node, and the resource margin of each processor, a more flexible processor resource allocation is achieved, thereby being able to effectively utilize the idle time of the processor and improve the resource utilization rate of the processor cluster.
[0050] In particular, for multiple task nodes with a complex execution order, compared with simply setting execution priorities for the task nodes, the above method 200 can obtain more details about the execution order (i.e., the specific execution time of each task node), so as to better find a scheduling scheme with a higher resource utilization rate.
[0051] It should be noted that in the present disclosure, the terms "execution time" and "required execution time" both indicate a period of time required to execute a task node, for example, 1s, 5min, and 1h, etc.
[0052] Here, a specific example is used to explain the above situation of multiple task nodes with a complex execution order in detail.
[0053] Figure 3 An exemplary schematic diagram of task nodes of a face fusion task workflow with a complex execution order is shown.
[0054] As Figure 3As shown, the face fusion task workflow has task branch 1 (including the "background matting" task node and the "basic text-to-image generation" task node) and task branch 2 (including the "style reference" task node, the "LoRA training" task node, and the "style reference text-to-image generation" task node).
[0055] Between task branch 1 and task branch 2, the "background matting" task node of task branch 1, the "basic text-to-image generation" task node of task branch 1, and the "style reference text-to-image generation" task node of task branch 2 are in a sequential execution order, while the "background matting" task node of task branch 1 and the "style reference" task node and the "LoRA training" task node of task branch 2 are in a parallel execution order.
[0056] Based on the above relatively complex execution order relationship, if the "background matting" task node and the "style reference" task node are simply taken as the first priority, the "LoRA training" task node as the second priority, and the "basic text-to-image generation" task node and the "style reference text-to-image generation" task node as the third priority, then in the case where the execution time of the "style reference" task node is less than that of the "background matting", there will be a situation where after the processor 1 finishes executing the "style reference" task node, the processor 2 is still executing the "background matting" task node, resulting in a waste of the computing resources of processor 1.
[0057] In this case, if the "style reference" task node points to the "reference" direction (i.e., the "LoRA training" task node of the second priority needs to be executed subsequently), then before the processor 1 finishes executing the "background matting" task node, the idle processor 1 can be scheduled to continue processing the "LoRA training" task node of the next priority, so as to effectively improve the overall utilization rate of the processor resources.
[0058] According to one or more embodiments, for tasks that need to be executed in parallel, the output data of the task nodes that have been executed can be locked in advance by using a shared variable or a lock mechanism first, and then the corresponding processor can be scheduled to process other task nodes, so as to better ensure the correct execution order among the task nodes in the scenario of multi-threaded access.
[0059] Exemplarily, also referring to Figure 3 , if the "style reference" task node points to the "not reference" direction (i.e., the "basic text-to-image generation" task node is directly executed subsequently), then during the period when the processor 1 is rescheduled to execute other tasks after finishing the "style reference" task node, the first data of the "style reference" task node output by it can be locked first, so that after the processor 1 finishes executing the "background matting" task to obtain the output second data, the first data and the second data can be output to the "basic text-to-image generation" task node as inputs at the same time, so as to clarify the execution start time of the "basic text-to-image generation" task node.
[0060] According to one or more embodiments, for each task node, the input data features and corresponding computing structures of the task node can be cached, so that when the same task node is encountered subsequently, the corresponding cached data can be directly called to reduce unnecessary computations.
[0061] It should be understood that the above examples are only for illustrative purposes. In the specific implementation process, the number of task nodes and the complexity of the execution order relationship can be higher than Figure 3 the face fusion task workflow shown, and no specific restrictions are imposed thereon.
[0062] Figure 4 FIG. shows a partial flowchart of another method for scheduling processor resources according to an embodiment of the present disclosure.
[0063] According to some embodiments, step 210 includes:
[0064] Step 410, obtain the workflow to be processed;
[0065] Step 420, determine a plurality of task nodes according to the workflow to be processed;
[0066] Step 430, define a graph structure of the workflow to be processed based on the plurality of task nodes; and
[0067] Step 440, process the graph structure using a traversal search algorithm to determine the execution order information.
[0068] For a distributed deep learning scenario, in particular, for various model inference and hybrid training scenarios based on a large-scale processor cluster environment (such as a cloud computing platform, a computing power provider, and a hardware resource management department, etc.), the execution order of multiple task nodes in its associated workflow usually has a relatively complex execution order (refer to Figure 3 ), and by defining the workflow to be processed as a corresponding graph structure and using a traversal search algorithm to process the graph structure, a more complete execution order can be determined more accurately.
[0069] In step 420, each task node among the determined multiple task nodes includes the operations that the task node must execute and the execution order relationship between the task node and other task nodes. The above operations and execution order relationship of each task node can be defined through a programming language or a specific configuration file. Exemplarily, the configuration file can be at least one of a file in JSON (JavaScript Object Notation) format or a file in XML (eXtensible Markup Language) format.
[0070] In step 430, the graph structure of the workflow to be processed can be defined based on a directed acyclic graph.
[0071] In step 440, exemplarily, the traversal search algorithm can be at least one of a breadth-first search algorithm and a depth-first search algorithm to more comprehensively determine the execution order information of the task nodes of the workflow.
[0072] In step 220, exemplarily, for complex computing scenarios such as cloud computing, the processor cluster can be, for example, a GPU (Graphic Process Unit) cluster.
[0073] Exemplarily, a Kubernetes (an open-source system for automatically deploying, scaling, and managing containerized applications) cluster can be deployed and a GPU device plugin can be installed to identify and manage GPU devices and use them as allocable resources.
[0074] Exemplarily, a DCGM-exporter (a GPU metric exporter) can be deployed to obtain GPU monitoring data, and key metrics such as the real-time utilization rate and video memory usage of the GPU can be collected through Prometheus (a system monitoring and alert toolkit) to provide data support.
[0075] Figure 5 Shows a partial flowchart of another processor resource scheduling method according to an embodiment of the present disclosure.
[0076] According to some embodiments, as Figure 5 shown, step 230 includes:
[0077] Step 510, determining at least one processor type associated with the processor cluster; and
[0078] Step 520, for each task node, determining the required execution time and required computing resources of the task node for each processor type in at least one processor type, so as to obtain at least one required execution time and at least one required computing resource corresponding to the task node and at least one processor type and add them to the resource requirement information.
[0079] For a large-scale processor cluster, which may include various types and models of processors, the required execution time and required computing resources of different task nodes corresponding to different types of processors may also be different. By pre-determining the required execution time and required computing resources of each task node for different types of processors, resource scheduling can be more accurately achieved.
[0080] In an example, a reference test data set can be used to test the required execution time and required computing resources of each workflow node for different types of processors. The required computing resources can be, for example, the required GPU core frequency and video memory occupancy resources.
[0081] According to some embodiments, step 240 includes: using a linear programming algorithm with the execution order information, available resource information, and resource requirement information as constraint conditions and the resource utilization rate of the processor cluster as the objective function to perform processor resource scheduling for multiple task nodes.
[0082] Thereby, an optimal processor resource scheduling scheme (for example, a scheduling scheme with the highest resource utilization rate) can be found. Specifically, for the many-to-many optimization scenario of multiple task nodes and multiple processors, and in the case where the complex execution order of multiple task nodes is one of the constraint conditions, using the linear programming algorithm can more efficiently and accurately find the optimal solution for processor resource scheduling for resource utilization.
[0083] According to some embodiments, the above linear programming algorithm includes at least one of a genetic algorithm, an ant colony algorithm, a simulated annealing algorithm, and a greedy algorithm.
[0084] Thereby, through the different characteristics of different linear programming algorithms, a balance can be achieved between the optimization efficiency and the accuracy of the optimization result to further improve the scenario universality of the above method.
[0085] Exemplarily, a greedy algorithm can be used to find a local optimal solution to improve the optimization efficiency.
[0086] Exemplarily, at least one of a genetic algorithm, an ant colony algorithm, or a simulated annealing algorithm can be used to replace the greedy algorithm to improve the global search ability, avoid local optimal solutions, and improve the accuracy of the optimization result.
[0087] According to some embodiments, in addition to the above steps 210 to 240, method 200 further includes:
[0088] Step 250: Match the required computing resources of each task node with the resource margin of each processor to establish an association relationship between the task nodes and processors that meet the matching requirements, where the association relationship indicates that the required computing resources of the corresponding task node are less than or equal to the resource margin of the corresponding processor;
[0089] And wherein, step 240 includes: performing processor resource scheduling for multiple task nodes according to the association relationship, execution order information, available resource information, and resource requirement information of each task node.
[0090] Thus, it is possible to pre-screen processors with insufficient resource margins for each task node and exclude them, effectively reducing the computational complexity in subsequent resource scheduling and improving efficiency.
[0091] According to one or more embodiments, step 240 may further include: using a linear programming algorithm, with the execution order information, available resource information, and resource requirement information of each task node as constraint conditions and the resource utilization rate of the processor cluster as the objective function, performing processor resource scheduling for multiple task nodes.
[0092] The specific description content is as above and will not be elaborated here.
[0093] Figure 6 Shows a partial flowchart of another processor resource scheduling method according to an embodiment of the present disclosure.
[0094] According to some embodiments, as Figure 6 shown, in addition to the above steps 210 to 240, method 200 further includes:
[0095] Step 610, obtaining a reference time threshold, where the reference time threshold is associated with the transmission time required to transmit reference data to each processor;
[0096] Step 620, screening out at least one fragmented task node from multiple task nodes whose execution time is less than the reference time threshold; and
[0097] Step 630, performing processor resource scheduling for multiple task nodes according to at least one fragmented task node, at least one execution order information, available resource information, and resource requirement information, where each fragmented task node in the at least one fragmented task node will be preferentially assigned to be executed in parallel with the data transmission process of other task nodes in the multiple task nodes.
[0098] At the beginning and end of starting each task node using a processor, there will be a period of data transmission time (e.g., model loading time and I / O processing time, etc.). At this time, the processor will not perform specific computing tasks. Therefore, some fragmented task nodes with short required execution times can be screened out and inserted into such data transmission processes to make full use of the data transmission time of large tasks and further improve the overall utilization rate of processor resources.
[0099] It should be noted that the "time" in the above "data transmission time", "model loading time", and "I / O processing time" all indicates a period of time length. For example, the data transmission time can be 1s, the model loading time can be 3μs, and the I / O processing time can be 5ms, etc.
[0100] Exemplarily, in step 610, based on each fragmented task node, the corresponding model loading time and I / O processing time of the task node can be tested, and one of its average value, maximum value, and minimum value can be used as the reference time threshold.
[0101] Exemplarily, when inserting the fragmented task node into the data transmission time period of the target task node, the execution order relationship between the fragmented task node and the target task node also needs to be further considered to ensure the execution timeliness among multiple task nodes.
[0102] According to one or more embodiments, after determining the resource scheduling scheme for multiple task nodes of the workflow to be processed, it can be used as a training dataset to train the responsible prediction model, so as to use the trained load prediction model and combine with the load prediction technology to predict the required processor resources for the workflow to be processed at a future moment, and make a resource scheduling decision in advance, for example, reserve the corresponding processor resources in advance, etc.
[0103] According to one or more embodiments, in some cases, when there are newly added processor resources in the processor cluster, the above method 200 can be re-performed based on the unexecuted task nodes to effectively improve the resource utilization rate of the processor cluster and reduce the load pressure on the original processors.
[0104] In the technical solution of the present disclosure, the processing of collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0105] According to another aspect of the present disclosure, a scheduling device for processor resources is provided. As Figure 7 shown, the scheduling device 700 for processor resources includes: a first acquisition module 710 configured to acquire the execution order information of multiple task nodes, where at least two of the multiple task nodes need to be executed in parallel; a second acquisition module 720 configured to acquire the available resource information of the processor cluster, where the available resource information indicates the resource margin of each processor included in the processor cluster; a determination module 730 configured to determine the resource requirement information, where the resource requirement information indicates the required execution time and required computing resources of each task node among the multiple task nodes; and a first scheduling module 740 configured to perform processor resource scheduling for the multiple task nodes according to the execution order information, available resource information, and resource requirement information.
[0106] According to another aspect of the present disclosure, there is also provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the foregoing method.
[0107] According to another aspect of the present disclosure, there is also provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the foregoing method.
[0108] According to another aspect of the present disclosure, there is also provided a computer program product, including a computer program, wherein the computer program, when executed by a processor, implements the foregoing method.
[0109] As Figure 8 shown, the electronic device 800 includes a computing unit 801, which can execute various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0110] A plurality of components in the electronic device 800 are connected to the I / O interface 805, including: an input unit 806, an output unit 807, a storage unit 808, and a communication unit 809. The input unit 806 can be any type of device capable of inputting information into the electronic device 800. The input unit 806 can receive input digital or character information, and generate key signal inputs related to the user settings and / or function controls of the electronic device, and can include, but is not limited to, a mouse, a keyboard, a touch screen, a trackpad, a trackball, a joystick, a microphone, and / or a remote control. The output unit 807 can be any type of device capable of presenting information, and can include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 808 can include, but is not limited to, a magnetic disk, an optical disk. The communication unit 809 allows the electronic device 800 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks, and can include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth TM device, an 802.11 device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.
[0111] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above, such as the GPU-based matrix calculation method. For example, in some embodiments, the GPU-based matrix calculation method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the GPU-based matrix calculation method described above can be executed. Alternatively, in other embodiments, the computing unit 801 can be configured to execute the GPU-based matrix calculation method in any other suitable manner (e.g., by means of firmware).
[0112] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor, that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0113] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program code is executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0114] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0115] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).
[0116] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0117] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.
[0118] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitations are imposed herein.
[0119] Although embodiments or examples of the present disclosure have been described with reference to the accompanying drawings, it should be understood that the above methods, systems, and devices are merely exemplary embodiments or examples, and the scope of the present invention is not limited by these embodiments or examples, but is only defined by the authorized claims and their equivalent scope. Various elements in the embodiments or examples can be omitted or replaced by their equivalent elements. In addition, the steps can be executed in an order different from that described in the present disclosure. Further, the various elements in the embodiments or examples can be combined in various ways. Importantly, with the evolution of technology, many of the elements described herein can be replaced by equivalent elements that emerge after the present disclosure.
Claims
1. A method for scheduling processor resources, comprising: Acquire execution order information of multiple task nodes, wherein at least two task nodes among the multiple task nodes need to be executed in parallel; Acquire available resource information of a processor cluster, wherein the available resource information indicates a resource margin of each processor among a plurality of processors included in the processor cluster; determining resource requirement information, wherein the resource requirement information indicates a required execution time and required computing resources for each of the plurality of task nodes; and Processor resource scheduling is performed for the multiple task nodes according to the execution order information, the available resource information, and the resource requirement information.
2. The method according to claim 1, wherein: The performing processor resource scheduling for the plurality of task nodes according to the execution order information, the available resource information and the resource requirement information includes: The execution order information, the available resource information and the resource demand information are used as constraint conditions, the resource utilization rate of the processor cluster is used as an objective function, and a linear programming algorithm is used to perform processor resource scheduling for the multiple task nodes.
3. The method according to claim 2, wherein: The linear programming algorithm includes at least one of a genetic algorithm, an ant colony algorithm, a simulated annealing algorithm and a greedy algorithm.
4. The method according to any one of claims 1 to 3, wherein: The obtaining the execution order information of the plurality of task nodes includes: Get the pending workflow; Determine the multiple task nodes according to the workflow to be processed; Defining a graph structure of the to-be-processed workflow based on the multiple task nodes; and The graph structure is processed using a traversal search algorithm to determine the execution order information.
5. The method according to any one of claims 1 to 4, further comprising: Matching is performed according to the required computing resources of each task node and the resource margin of each processor to establish an association relationship between the task nodes and processors that meet the matching requirements, wherein the association relationship indicates that the required computing resources of the corresponding task node are less than or equal to the resource margin of the corresponding processor; And wherein, performing processor resource scheduling for the multiple task nodes according to the execution order information, the available resource information and the resource requirement information includes: Processor resource scheduling is performed for the multiple task nodes according to the association relationship of each task node, the execution order information, the available resource information and the resource demand information.
6. The method according to any one of claims 1 to 5, wherein: The determining of resource requirement information includes: determining at least one processor type associated with the processor cluster; and For each task node, determine the required execution time and required computing resources of the task node for each processor type of the at least one processor type, so as to obtain at least one required execution time and at least one required computing resource corresponding to the task node and the at least one processor type and add the resource requirement information.
7. The method according to any one of claims 1 to 6, further comprising: obtaining a reference time threshold, wherein the reference time threshold is associated with a transmission time required to transmit reference data to each of the processors; Filtering out at least one fragment task node whose execution time is less than a reference time threshold from the plurality of task nodes; and Performing processor resource scheduling for the plurality of task nodes according to the at least one fragment task node, the at least one execution order information, the available resource information and the resource demand information, Among them, each fragment task node among the at least one fragment task node will be preferentially allocated to be executed in parallel with the data transmission process of other task nodes among the multiple task nodes.
8. A processor resource scheduling device, comprising: A first acquisition module is configured to acquire execution order information of a plurality of task nodes, wherein at least two task nodes among the plurality of task nodes need to be executed in parallel; A second acquisition module is configured to acquire available resource information of a processor cluster, wherein the available resource information indicates a resource margin of each processor among a plurality of processors included in the processor cluster; a determination module configured to determine resource requirement information, wherein the resource requirement information indicates a required execution time and required computing resources of each of the plurality of task nodes; and The first scheduling module is configured to perform processor resource scheduling for the multiple task nodes according to the execution order information, the available resource information and the resource requirement information.
9. An electronic device, comprising: at least one processor; as well as a memory communicatively coupled to the at least one processor; in The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-7.
11. A computer program product comprising a computer program, wherein: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.