Processor resource scheduling method, device and system
Through the collaborative work of client, proxy node and server node, the queue management and resource allocation of task requests are realized, which solves the problem of task failure caused by multi-task scrambling for processor resources, and improves resource utilization and task execution efficiency.
Patent Information
- Application Number
- PCT/CN2024/139408
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-28
- Filing Date
- 2024-12-13
- Publication Date
- 2025-07-03
AI Technical Summary
When multitasks compete for processor resources, the problem of task request failure due to insufficient resources is especially after the introduction of K8S and customized GPU resource scheduling algorithms, the system complexity increases, resource utilization decreases, and scheduling complexity and debugging difficulty increase.
Through the collaborative work of the client, proxy node and server node, the queue management and resource allocation of task requests are realized, ensuring that task requests are scheduled in order, the container corresponds one by one with the client user, and task requests are allocated and executed when the processor resources are idle, and task relationships and time thresholds are taken into account when releasing resources.
It greatly reduces the probability of task request failure due to insufficient resources when multitasks compete for processor resources, improves resource utilization and task execution efficiency, and reduces frequent resource occupation and release operations.
Smart Images

Figure CN2024139408_03072025_PF_FP_ABST
Abstract
Description
Processor resource scheduling method, device and system
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to the Chinese patent application filed with the China Patent Office on December 28, 2023, with application number 202311825166.2, and application name “Processor Resource Scheduling Method, Device and System”, all contents of which are incorporated by reference into this application. Technical Field
[0003] The present application relates to the field of computer technology, and in particular to a processor resource scheduling method, device, and system. Background Art
[0004] With the development and advancement of the times, artificial intelligence (AI) technology has exploded, and a variety of AI accelerator cards have emerged on the market. Because a single accelerator card cannot run multiple tasks concurrently, client users often fail when requesting a graphics processing unit (GPU) to run a computing task, severely impacting the user experience. The current mainstream solution involves introducing Kubernetes (K8S) and then adding customized GPU resource scheduling algorithms. Introducing K8S increases the complexity of the GPU resource cluster system, incurs additional resource overhead, and reduces system performance. Furthermore, in practical applications, customized scheduling algorithms are often required to achieve true scheduling. The allocation and destruction of Pods (the smallest scheduling unit in Kubernetes) consumes resources, rendering them unusable. Frequent allocation and destruction of Pods significantly reduces effective data utilization. Furthermore, the introduction of Pods increases scheduling complexity, makes application debugging more difficult, and consumes more server resources. Summary of the Invention
[0005] The purpose of this application is to provide a processor resource scheduling method, device and system to solve the problem in the related art that when multiple tasks compete for processor resources, task request failures are often caused by insufficient resources.
[0006] The present application provides a processor resource scheduling system, comprising: a client, a proxy node, and multiple server nodes, each server node comprising multiple containers, a one-to-one correspondence between client users and containers, each container corresponding to a processor, and each processor corresponding to multiple containers;
[0007] The client is configured to send the current task request generated by the client user to the proxy node;
[0008] The proxy node is configured to forward the current task request based on the correspondence between the client user ID (Identity Document) and the container user ID, forward the current task request to the server node where the container corresponding to the client user is located, receive the task execution result returned by the server node, and forward the task execution result to the client user;
[0009] The server node is configured to add the current task request to the task request queue of the processor corresponding to the client user; when the processor resources are idle and the task request queue is not empty, the task request at the head of the task request queue is taken out, and the processor resources are allocated to the currently taken out task request so that the processor executes the task corresponding to the currently taken out task request; after the task execution is completed, the task execution result is returned to the proxy node and the processor resources are released.
[0010] According to a processor resource scheduling system provided by the present application, a server node is configured to obtain a container user ID corresponding to a currently running task in the processor, and when the container user ID corresponding to the currently running task is the same as the container user ID corresponding to the current task request, the current task request is added to the head of the task request queue; otherwise, the current task request is added to the tail of the task request queue. The server node is also configured to obtain the container user ID corresponding to the task request at the head of the queue in the task request queue after returning the task execution result to the proxy node after the task execution is completed and before releasing the processor resources, and when the container user ID corresponding to the task request at the head of the queue is the same as the container user ID corresponding to the current task request, the task request at the head of the queue is loaded into the processor for execution.
[0011] According to a processor resource scheduling system provided by the present application, the server node has a function for inserting the current task request after the previous task request corresponding to the client user in the task request queue; otherwise, the current task request is added to the end of the task request queue.
[0012] According to a processor resource scheduling system provided by the present application, the server node is configured to return the task execution result to the agent node after the task execution is completed, and release the processor resources when no new task request from the client user corresponding to the currently completed task is received within a preset time period; when a new task request from the client user corresponding to the currently completed task is received within a preset time period, the new task request is loaded into the processor for execution.
[0013] According to a processor resource scheduling system provided by the present application, a server node is configured to release processor resources when the execution time of a task corresponding to a currently retrieved task request by the processor exceeds a preset execution time threshold.
[0014] According to a processor resource scheduling system provided by the present application, the proxy node is further configured to send client user login information to the server node;
[0015] The server node receiving the client user login information is further configured to, if an empty container exists, establish a one-to-one correspondence between the client user and the empty container based on the client user login information, and return information indicating successful container allocation to the proxy node; if no empty container exists, return information indicating failed container allocation to the proxy node;
[0016] The proxy node is further configured to send the client user login information to the next server node in a polling manner when receiving information that the container allocation fails.
[0017] The present application also provides a processor resource scheduling method, which is applied to each server node in a server node cluster, and the method includes:
[0018] Receive the current task request from the client user corresponding to any container forwarded by the proxy node. The current task request is generated by the client user and sent to the proxy node in a unified manner. The proxy node is configured to forward the current task request based on the correspondence between the client user ID and the container user ID;
[0019] Add the current task request to the task request queue of the processor corresponding to the client user;
[0020] When the processor resources are idle, the task request at the head of the task request queue is taken out, and the processor resources are allocated to the currently taken out task request so that the processor executes the task corresponding to the currently taken out task request;
[0021] After the task execution is completed, the task execution result is returned to the proxy node, and the processor resources are released. If the task request queue is not empty, jump to the step of taking out the task request at the head of the task request queue when the processor resources are idle, and allocate the processor resources to the currently taken task request so that the processor executes the task corresponding to the currently taken task request.
[0022] According to a processor resource scheduling method provided by the present application, a current task request is added to a task request queue of a processor corresponding to a client user, including:
[0023] Get the container user ID corresponding to the currently running task in the processor. If the container user ID corresponding to the currently running task is the same as the container user ID corresponding to the current task request, add the current task request to the head of the task request queue; otherwise, add the current task request to the tail of the task request queue.
[0024] After the task execution is completed and the task execution result is returned to the agent node, but before the processor resources are released, the following steps are also performed:
[0025] Obtain the container user ID corresponding to the head task request in the task request queue. If the container user ID corresponding to the head task request is the same as the container user ID corresponding to the current task request, load the head task request to the processor for execution.
[0026] According to a processor resource scheduling method provided by the present application, a current task request is added to a task request queue of a processor corresponding to a client user, including:
[0027] If a previous task request corresponding to the current task request already exists in the task request queue, the current task request is inserted after the previous task request; otherwise, the current task request is added to the end of the task request queue.
[0028] According to a processor resource scheduling method provided by the present application, after the task execution is completed, the task execution result is returned to the agent node and the processor resources are released, including:
[0029] After the task is completed, the task execution result is returned to the proxy node. If no new task request from the client user corresponding to the currently completed task is received within the preset time period, the processor resources are released. If a new task request from the client user corresponding to the currently completed task is received within the preset time period, the new task request is loaded into the processor for execution.
[0030] According to a processor resource scheduling method provided by the present application, when a new task request from a client user corresponding to a currently completed task is received within a preset time period, the new task request is loaded into the processor for execution, including:
[0031] When the task is completed, get the container user ID corresponding to the currently completed task;
[0032] Obtain the container user ID of a new task request received within a preset time period. If the container user ID of the new task request is the same as the container user ID corresponding to the currently executed completed task, load the new task request into the processor for execution.
[0033] A processor resource scheduling method provided by the present application further includes: releasing processor resources when the processor executes a task corresponding to a currently retrieved task request for a time exceeding a preset execution time threshold.
[0034] A processor resource scheduling method provided in the present application also includes: when the processor executes the task corresponding to the currently retrieved task request for more than a preset execution time threshold, saving the current execution progress of the task, adding the currently retrieved task request back to the end of the task request queue, and releasing processor resources.
[0035] According to a processor resource scheduling method provided by the present application, before receiving a current task request from a client user corresponding to any container forwarded by a proxy node, the method further includes:
[0036] Receive the client user login information sent by the proxy node. If there is an empty container on the current server node, establish a one-to-one correspondence between the client user and the empty container based on the client user login information, and return the information that the container allocation is successful to the proxy node. If there is no empty container on the current server node, return the information that the container allocation failed to the proxy node to instruct the proxy node to send the client user login information to the next server node in a polling manner until the container is successfully allocated.
[0037] The present application also provides a processor resource scheduling method, which is applied to a proxy node in a server node cluster. The method includes:
[0038] Receive the current task request sent by the client user, and forward the current task request to the server node according to the correspondence between the client user ID and the container user ID of the container in the server node. The server node is the server node where the container corresponding to the client user is located. The current task request is used to instruct the server node to add the current task request to the task request queue of the processor corresponding to the client user, and when the processor resources are idle and the task request queue is not empty, take out the task request at the head of the task request queue, allocate the processor resources to the currently taken out task request, so that the processor executes the task corresponding to the currently taken out task request; after the task execution is completed, return the task execution result to the proxy node and release the processor resources;
[0039] Receive the task execution results returned by the server node and forward the task execution results to the corresponding client user.
[0040] According to a processor resource scheduling method provided by the present application, before receiving a current task request sent by a client user, the method further includes:
[0041] The client user login information is sent to the server node, so that when an empty container exists in the server node, the server node that receives the client user login information establishes a one-to-one correspondence between the client user and the empty container based on the client user login information, and returns a message that the container allocation is successful to the proxy node; if no empty container exists, the server node returns a message that the container allocation fails to the proxy node;
[0042] When receiving the information that the container allocation fails, the client user login information is sent to the next server node in a polling manner.
[0043] The present application also provides a processor resource scheduling device, which is applied to each server node in a server node cluster, and the device includes:
[0044] A request receiving module is configured to receive a current task request from a client user corresponding to any container and forwarded by a proxy node. The current task request is generated by the client user and sent uniformly to the proxy node. The proxy node is configured to forward the current task request based on the correspondence between the client user ID and the container user ID;
[0045] A request queue module is configured to add the current task request to the task request queue of the processor corresponding to the client user;
[0046] The resource allocation module is configured to, when the processor resources are idle, take out the task request at the head of the task request queue and allocate the processor resources to the currently taken out task request so that the processor executes the task corresponding to the currently taken out task request;
[0047] The result return module is configured to return the task execution result to the proxy node after the task execution is completed, release the processor resources, and jump to the resource allocation module if the task request queue is not empty.
[0048] The present application also provides a processor resource scheduling device, which is applied to a proxy node in a server node cluster, and the device includes:
[0049] The task request forwarding module is configured to receive a current task request sent by a client user, and forward the current task request to a server node according to a correspondence between a client user ID and a container user ID of a container in a server node. The server node is a server node where a container corresponding to the client user is located. The current task request is used to instruct the server node to add the current task request to a task request queue of a processor corresponding to the client user, and when processor resources are idle and the task request queue is not empty, to retrieve the task request at the head of the queue in the task request queue, and to allocate processor resources to the currently retrieved task request, so that the processor executes the task corresponding to the currently retrieved task request; after the task is executed, the task execution result is returned to the proxy node, and the processor resources are released;
[0050] The execution result forwarding module is configured to receive the task execution result returned by the server node and forward the task execution result to the corresponding client user.
[0051] The present application also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps of any of the above-described processor resource scheduling methods are implemented.
[0052] The present application also provides a computer non-volatile readable storage medium having a computer program stored thereon, which implements the steps of any of the above-mentioned processor resource scheduling methods when executed by a processor.
[0053] The processor resource scheduling method, device and system provided in the present application receive task requests from different client users hijacked and forwarded by the proxy node, add the task requests to the task request queue and schedule them in sequence. Since the container corresponds one-to-one with the client user who logs into the server node cluster, for the client user who is successfully bound to the container in the server node, his task request will be allocated to the processor resources and executed, which greatly reduces the probability of task request failure due to insufficient resources when multiple tasks compete for processor resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] In order to more clearly illustrate the technical solutions in the present application or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0055] FIG1 is an architecture diagram of a server node in the processor resource scheduling method provided by the present application;
[0056] FIG2 is a flowchart of a method for scheduling processor resources provided by the present application;
[0057] FIG3 is a second flow chart of the processor resource scheduling method provided by the present application;
[0058] FIG4 is a schematic diagram of a structure of a processor resource scheduling device provided by the present application;
[0059] FIG5 is a second structural diagram of the processor resource scheduling device provided by the present application;
[0060] FIG6 is a schematic diagram of the structure of the processor resource scheduling system provided by the present application;
[0061] FIG7 is a schematic structural diagram of the electronic device provided in this application. DETAILED DESCRIPTION
[0062] To make the objectives, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments of this application, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this application.
[0063] The processor resource scheduling method of the embodiment of the present application is applied to each server node in the server node cluster, and the server node includes at least one processor, and the processor can be a CPU (Central Processing Unit), a GPU, and other dedicated computing acceleration cards. Optionally, the server node architecture is shown in Figure 1, and each server node includes multiple containers (such as: docker (container)), and the container corresponds one-to-one with the client user who logs in to the server node cluster, that is, after each client user logs in, the client user ID and the container user ID are bound, each container corresponds to a processor, and each processor corresponds to multiple containers, that is, the task request of the client user corresponding to the processor corresponding to the container is executed by the processor, and the container is configured with a program for executing the task. Taking the AI computing accelerator card as an example, the main processing device in the AI computing accelerator card is a GPU, and its supporting complete AI software stack (AI training software stack and AI reasoning software stack) is configured in each container, that is, each client user has its own independent AI software stack.
[0064] The method flow is shown in Figure 2 and includes:
[0065] Step S210: Receive a current task request from a client user corresponding to any container, forwarded by a proxy node. The current task request is generated by the client user and uniformly sent to the proxy node. The proxy node is also a server node, and a server node can be selected from a server node cluster as a proxy node. When the client user calls the API (Application Programming Interface) of the processor resource to send a task request, the proxy node hijacks the task request and then forwards the task request to the server node where the container corresponding to the client user is located based on the correspondence between the client user ID and the container user ID.
[0066] Step S220: Add the current task request to the task request queue of the processor corresponding to the client user. Each processor maintains an independent task request queue. Normally, according to the first-in-first-out feature of the queue, the current task request will be added to the end of the task request queue and wait for execution in the queue order.
[0067] Step S230: When the processor resources are idle, take out the task request at the head of the task request queue, and allocate the processor resources to the currently taken out task request, so that the processor executes the task corresponding to the currently taken out task request. Optionally, obtain the container user ID, execution task name and execution parameters in the task request at the head of the queue, take out the program corresponding to the execution task name from the container of the corresponding container user ID, and the processor loads the program and execution parameters and executes them, thereby realizing resource allocation. Taking the AI computing accelerator card as an example, the task request at the head of the queue is a request from a client user to execute a certain AI large model training, wherein the execution parameters include a training data set for executing the AI large model training, and the GPU resources of the AI computing accelerator card are allocated to the AI large model and the training data set, that is, the GPU loads the AI large model and the training data set to execute the AI large model training.
[0068] Step S240: After the task is completed, the task execution result is returned to the proxy node and the processor resources are released. Of course, after receiving the task execution result, the proxy node returns the task execution result to the corresponding client user. Optionally, the task execution result also includes the container user ID. After receiving the task execution result, the proxy node returns the task execution result to the corresponding user based on the client user ID corresponding to the container user ID. For example: after the task of executing AI large model training is completed, the task execution result including the training completion message and the container user ID will be returned to the proxy node. For another example: after the task of executing AI large model reasoning is completed, the task execution result including the reasoning result and the container user ID will be returned to the proxy node.
[0069] Step S250: After releasing the processor resources, determine whether the task request queue is empty. If it is not empty, jump to step S230 and continue to execute the remaining tasks in the task request queue. Otherwise, execute step S260.
[0070] Step S260: Waiting for a task request, that is, waiting for the task request to enter the queue.
[0071] In the processor resource scheduling method of the embodiment of the present application, by receiving task requests from different client users hijacked and forwarded by the proxy node, the task requests are added to the task request queue and scheduled in sequence. Since the container corresponds one-to-one with the client user who logs into the server node cluster, for the client user who successfully binds to the container in the server node, its task request will be allocated to the processor resources and executed, which greatly reduces the probability of task request failure due to insufficient resources when multiple tasks compete for processor resources. In addition, in the embodiment of the present application, the container and the client user correspond one-to-one, and different task execution programs are stored in the container. For the task request of the client user, the task execution program in the corresponding container is loaded into the processor for execution. After the execution is completed, the resources are released, thereby realizing business isolation between different client users.
[0072] The processor resource scheduling method of the embodiment of the present application is particularly suitable for large AI model scenarios. The training or inference tasks of large AI models generally occupy more than 80% of GPU resources, which can cause the second large AI model task request to fail due to insufficient GPU resources. Using the processor resource scheduling method of the embodiment of the present application can greatly reduce the probability of the second large AI model task request failing.
[0073] In some optional embodiments, step S220 includes: obtaining the container user ID corresponding to the currently running task in the processor, and if the container user ID corresponding to the currently running task is the same as the container user ID corresponding to the current task request, adding the current task request to the head of the task request queue; otherwise, adding the current task request to the tail of the task request queue. Based on this, in step S240, before releasing the processor resources, it also includes:
[0074] Obtain the container user ID corresponding to the head task request in the task request queue. If the container user ID corresponding to the head task request is the same as the container user ID corresponding to the current task request, load the head task request to the processor for execution.
[0075] In an embodiment of the present application, the container user ID corresponding to the currently running task is the same as the container user ID corresponding to the current task request, indicating that the currently running task and the current task request belong to the same client user. Since the program for executing the task in the container corresponding to the client user has been loaded into the processor, in order to avoid frequent loading of programs and frequent release of resources, the current task request is added to the head of the task request queue, so that after the currently running task is completed, there is no need to release processor resources, and the task in the current task request of the client user is directly taken out from the head of the queue and executed.
[0076] In some optional embodiments, step S220 includes: if a previous task request of the client user corresponding to the current task request already exists in the task request queue, inserting the current task request after the previous task request; otherwise, adding the current task request to the end of the task request queue, that is, different task requests of the same client user are arranged next to each other in the task request queue, further avoiding frequent loading of task execution programs and frequent release of resources of different client users.
[0077] In some optional embodiments, step S240 includes: returning the task execution result to the proxy node after the task execution is completed, and releasing the processor resources when no new task request from the client user corresponding to the currently completed task is received within a preset time period; and loading the new task request into the processor for execution when a new task request from the client user corresponding to the currently completed task is received within a preset time period.
[0078] Optionally, when the task is completed, the container user ID corresponding to the currently completed task is obtained.
[0079] Obtain the container user ID of a new task request received within a preset time period. If the container user ID of the new task request is the same as the container user ID corresponding to the currently executed completed task, load the new task request into the processor for execution.
[0080] The preset time period is greater than or equal to 0. When it is equal to 0, it means that the processor resources will be released immediately after the task is completed. The preset time period can be set to different values according to the actual application scenario. For example: a client user needs to train the AI big model first. After the training is completed, the trained AI big model will be used to infer multiple sets of real-time data to obtain multiple sets of inference results. In this application scenario, the time interval between each task request is approximately 2 minutes, so the preset time period can be set to 3 minutes. The client user first initiates a task request for AI big model training. After the AI big model training task is completed, if the corresponding AI big model inference task request is received within 3 minutes, the AI big model inference task request will be directly loaded into the GPU for execution. If the AI big model inference task request is not received within 3 minutes, the GPU resources occupied by the client user will be released.
[0081] In the embodiment of the present application, different preset time periods are set for different scenarios, which avoids frequent loading of programs and frequent release of resources while ensuring smooth execution of tasks of other client users in the task request queue.
[0082] In some optional embodiments, the processor resource scheduling method further includes: releasing processor resources when the processor executes the task corresponding to the currently retrieved task request for more than a preset execution time threshold. The execution time threshold can be set according to the completion time of different tasks under normal execution. For example, if the AI large model training task can usually be completed within 10 minutes, the execution time threshold can be set to 10 minutes. If it cannot be completed within 10 minutes due to a large number of data training sets or other reasons, in order to ensure that subsequent task requests in the task request queue can be executed smoothly, the GPU resource occupation of the AI large model training task is released, and subsequent task requests in the task request queue are executed.
[0083] In some optional embodiments, the processor resource scheduling method further includes: when the processor executes the task corresponding to the currently retrieved task request for a time exceeding a preset execution time threshold, saving the current execution progress of the task, re-adding the currently retrieved task request to the end of the task request queue, and releasing processor resources. In the embodiment of the present application, to ensure that the timed-out task can also be ultimately executed, the task request is re-added to the end of the task request queue after the processor resources occupied by the task request are released.
[0084] In some optional embodiments, before receiving a current task request from a client user corresponding to any container forwarded by a proxy node, the process further includes: receiving client user login information sent by the proxy node; if an empty container exists on the current server node, establishing a one-to-one correspondence between the client user and the empty container based on the client user login information, and returning a successful container allocation message to the proxy node; if an empty container does not exist on the current server node, returning a container allocation failure message to the proxy node, instructing the proxy node to send the client user login information to the next server node in a round-robin manner until a container is successfully allocated. When a client user logs in to a server node cluster, the proxy node sends the login information including the client user ID to the current server node in a round-robin manner; if an empty container exists on the current server node (no one-to-one correspondence is established with any client user), establishing a one-to-one correspondence between the client user and the empty container, i.e., binding the client user ID and the container user ID, and returning a successful container allocation message to the proxy node after the binding is complete; if an empty container does not exist on the current server node, returning a container allocation failure message to the proxy node, and the proxy node then sends the client user login information to the next server node. In the embodiment of the present application, each client user corresponds to a container, which realizes the business isolation between different client users, and uses the container user ID to identify the task requests of different client users within the server node, thereby realizing the complete decoupling of the client user management and the server scheduling algorithm.
[0085] The present application also provides a processor resource scheduling method, which is applied to a proxy node in a server node cluster. The method is shown in FIG3 and includes:
[0086] Step S310: Receive the current task request sent by the client user, and forward the current task request to the server node according to the correspondence between the client user ID and the container user ID of the container in the server node. The server node is the server node where the container corresponding to the client user is located. The current task request is used to instruct the server node to add the current task request to the task request queue of the processor corresponding to the client user, and when the processor resources are idle and the task request queue is not empty, take out the task request at the head of the task request queue, allocate the processor resources to the currently taken out task request, so that the processor executes the task corresponding to the currently taken out task request; after the task is executed, return the task execution result to the proxy node, and release the processor resources, wherein each server node includes multiple containers, the client user and the container correspond one to one, each container corresponds to one processor, and each processor corresponds to multiple containers.
[0087] Step S320: Receive the task execution result returned by the server node, and forward the task execution result to the corresponding client user.
[0088] In the processor resource scheduling method of the embodiment of the present application, the task requests of different client users are hijacked and forwarded to the server node through the proxy node, and the server node adds the task requests to the task request queue and schedules them in sequence. Since the container corresponds one-to-one with the client user who logs in to the server node cluster, for the client user who is successfully bound to the container in the server node, his task request will be allocated to the processor resources and executed, which greatly reduces the probability of task request failure due to insufficient resources when multiple tasks compete for processor resources.
[0089] In some optional embodiments, before step S310, the process further includes: sending the client user login information to the server node, so that when an empty container exists therein, the server node that receives the client user login information establishes a one-to-one correspondence between the client user and the empty container based on the client user login information, and returns information that the container allocation is successful to the proxy node; when no empty container exists, the server node returns information that the container allocation fails to the proxy node.
[0090] When receiving the information that the container allocation fails, the client user login information is sent to the next server node in a polling manner.
[0091] In the embodiment of the present application, client users are bound to containers by proxy node polling, so that the number of containers bound to client users in each server node is substantially the same, thereby maintaining load balancing between server nodes.
[0092] The processor resource scheduling device provided in the present application is described below. The processor resource scheduling device described below and the processor resource scheduling method described above can be referenced to each other.
[0093] FIG4 is a schematic diagram of the structure of a processor resource scheduling device provided by the present application. As shown in FIG4 , the processor resource scheduling device is applied to each server node in a server node cluster. Each server node includes multiple containers. Client users correspond to containers one by one. Each container corresponds to a processor, and each processor corresponds to multiple containers. The device includes:
[0094] The request receiving module 410 is configured to receive a current task request from a client user corresponding to any container and forwarded by a proxy node. The current task request is generated by the client user and sent uniformly to the proxy node. The proxy node is configured to forward the current task request based on the correspondence between the client user ID and the container user ID.
[0095] The request queue module 420 is configured to add the current task request to the task request queue of the processor corresponding to the client user.
[0096] The resource allocation module 430 is configured to retrieve the task request at the head of the task request queue when the processor resources are idle, and allocate processor resources to the currently retrieved task request so that the processor executes the task corresponding to the currently retrieved task request.
[0097] The result returning module 440 is configured to return the task execution result to the proxy node after the task execution is completed, release the processor resources, and jump to execute the resource allocation module 430 if the task request queue is not empty.
[0098] FIG5 is a second structural diagram of the processor resource scheduling device provided by the present application. As shown in FIG5 , the processor resource scheduling device is applied to an agent node in a server node cluster, and the device includes:
[0099] The task request forwarding module 510 is configured to receive the current task request sent by the client user, and forward the current task request to the server node according to the correspondence between the client user ID and the container user ID of the container in the server node. The server node is the server node where the container corresponding to the client user is located. The current task request is used to instruct the server node to add the current task request to the task request queue of the processor corresponding to the client user, and when the processor resources are idle and the task request queue is not empty, take out the task request at the head of the task request queue, allocate the processor resources to the currently taken out task request, so that the processor executes the task corresponding to the currently taken out task request; after the task is executed, return the task execution result to the proxy node and release the processor resources, wherein each server node includes multiple containers, the client user and the container correspond one to one, each container corresponds to one processor, and each processor corresponds to multiple containers.
[0100] The execution result forwarding module 520 is configured to receive the task execution result returned by the server node and forward the task execution result to the corresponding client user.
[0101] The present application also provides a processor resource scheduling system based on multi-node and multi-tasking, as shown in Figure 6. The system includes: a client 610, an agent node 620 and multiple server nodes 630. Each server node 630 includes multiple containers. Client users and containers correspond one to one. Each container corresponds to a processor, and each processor corresponds to multiple containers.
[0102] The client 610 is configured to send a current task request generated by a client user to the proxy node 620 .
[0103] The proxy node 620 is configured to forward the current task request according to the correspondence between the client user ID and the container user ID, forward the current task request to the server node 630 where the container corresponding to the client user is located, receive the task execution result returned by the server node 630, and forward the task execution result to the client user.
[0104] The server node 630 is configured to add the current task request to the task request queue of the processor corresponding to the client user; when the processor resources are idle and the task request queue is not empty, the task request at the head of the task request queue is taken out, and the processor resources are allocated to the currently taken out task request so that the processor executes the task corresponding to the currently taken out task request; after the task execution is completed, the task execution result is returned to the proxy node 620 and the processor resources are released.
[0105] The processor resource scheduling system of the embodiment of the present application hijacks and forwards task requests of different client users to the server node 630 through the proxy node 620. The server node 630 adds the task requests to the task request queue and schedules them in sequence. Since the containers correspond one-to-one to the client users who log in to the server node cluster, for the client users who are successfully bound to the containers in the server node, their task requests will be allocated to the processor resources and executed, which greatly reduces the probability of task request failure due to insufficient resources when multiple tasks compete for processor resources.
[0106] In some optional embodiments, the server node 630 is configured to obtain the container user ID corresponding to the currently running task in the processor, and add the current task request to the head of the task request queue when the container user ID corresponding to the currently running task is the same as the container user ID corresponding to the current task request; otherwise, the current task request is added to the tail of the task request queue. The server node is also configured to return the task execution result to the proxy node 620 after the task execution is completed, and before releasing the processor resources, obtain the container user ID corresponding to the task request at the head of the queue in the task request queue, and load the task request at the head of the queue into the processor for execution when the container user ID corresponding to the task request at the head of the queue is the same as the container user ID corresponding to the current task request.
[0107] In some optional embodiments, the server node 630 is configured to insert the current task request after a previous task request of the client user corresponding to the current task request in the task request queue; otherwise, add the current task request to the end of the task request queue.
[0108] In some optional embodiments, the server node 630 is configured to return the task execution result to the proxy node 620 after the task execution is completed, and release the processor resources when no new task request from the client user corresponding to the currently completed task is received within a preset time period, and load the new task request into the processor for execution when a new task request from the client user corresponding to the currently completed task is received within a preset time period.
[0109] In some optional embodiments, the server node 630 is configured to release processor resources when the time taken by the processor to execute the task corresponding to the currently retrieved task request exceeds a preset execution time threshold.
[0110] In some optional embodiments, the proxy node 620 is further configured to send the client user login information to the server node.
[0111] The server node 630 that receives the client user login information is further configured to, if an empty container exists therein, establish a one-to-one correspondence between the client user and the empty container based on the client user login information, and return information indicating successful container allocation to the proxy node 620; if no empty container exists, return information indicating failed container allocation to the proxy node 620.
[0112] The proxy node 620 is further configured to send the client user login information to the next server node 630 in a polling manner when receiving the information that the container allocation fails.
[0113] FIG7 is a schematic diagram of the structure of an electronic device provided by the present application. As shown in FIG7 , the electronic device may include: a processor 710, a communication interface 720, a memory 730, and a communication bus 740. The processor 710, the communication interface 720, and the memory 730 communicate with each other via the communication bus 740. The processor 710 may call the logic instructions in the memory 730 to execute a processor resource scheduling method applied to each server node in the server node cluster. The method includes:
[0114] The receiving agent node forwards the current task request of the client user corresponding to any container. The current task request is generated by the client user and sent to the agent node uniformly. The agent node is configured to forward the current task request based on the correspondence between the client user ID and the container user ID.
[0115] Add the current task request to the task request queue of the processor corresponding to the client user.
[0116] When the processor resources are idle, the task request at the head of the task request queue is taken out, and the processor resources are allocated to the currently taken out task request so that the processor executes the task corresponding to the currently taken out task request.
[0117] After the task execution is completed, the task execution result is returned to the proxy node, and the processor resources are released. If the task request queue is not empty, jump to the step of taking out the task request at the head of the task request queue when the processor resources are idle, and allocate the processor resources to the currently taken task request so that the processor executes the task corresponding to the currently taken task request.
[0118] Alternatively, a processor resource scheduling method applied to an agent node in a server node cluster is executed, the method comprising:
[0119] Receive the current task request sent by the client user, and forward the current task request to the server node according to the correspondence between the client user ID and the container user ID of the container in the server node. The server node is the server node where the container corresponding to the client user is located. The current task request is used to instruct the server node to add the current task request to the task request queue of the processor corresponding to the client user, and when the processor resources are idle and the task request queue is not empty, take out the task request at the head of the task request queue, allocate the processor resources to the currently taken out task request, so that the processor executes the task corresponding to the currently taken out task request; after the task execution is completed, return the task execution result to the proxy node and release the processor resources, wherein each server node includes multiple containers, the client user and the container correspond one to one, each container corresponds to a processor, and each processor corresponds to multiple containers.
[0120] Receive the task execution results returned by the server node and forward the task execution results to the corresponding client user.
[0121] In addition, the logic instructions in the above-mentioned memory 730 can be implemented in the form of a software functional unit and can be stored in a computer non-volatile readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a non-volatile storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned non-volatile storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various non-volatile storage media that can store program code.
[0122] The present application also provides a computer program product, which includes a computer program stored on a non-volatile computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute a processor resource scheduling method applied to each server node in a server node cluster. The method includes:
[0123] The receiving agent node forwards the current task request of the client user corresponding to any container. The current task request is generated by the client user and sent to the agent node uniformly. The agent node is configured to forward the current task request based on the correspondence between the client user ID and the container user ID.
[0124] Add the current task request to the task request queue of the processor corresponding to the client user.
[0125] When the processor resources are idle, the task request at the head of the task request queue is taken out, and the processor resources are allocated to the currently taken out task request so that the processor executes the task corresponding to the currently taken out task request.
[0126] After the task execution is completed, the task execution result is returned to the proxy node, and the processor resources are released. If the task request queue is not empty, jump to the step of taking out the task request at the head of the task request queue when the processor resources are idle, and allocate the processor resources to the currently taken task request so that the processor executes the task corresponding to the currently taken task request.
[0127] Alternatively, a processor resource scheduling method applied to an agent node in a server node cluster is executed, the method comprising:
[0128] Receive the current task request sent by the client user, and forward the current task request to the server node according to the correspondence between the client user ID and the container user ID of the container in the server node. The server node is the server node where the container corresponding to the client user is located. The current task request is used to instruct the server node to add the current task request to the task request queue of the processor corresponding to the client user, and when the processor resources are idle and the task request queue is not empty, take out the task request at the head of the task request queue, allocate the processor resources to the currently taken out task request, so that the processor executes the task corresponding to the currently taken out task request; after the task execution is completed, return the task execution result to the proxy node and release the processor resources, wherein each server node includes multiple containers, the client user and the container correspond one to one, each container corresponds to a processor, and each processor corresponds to multiple containers.
[0129] Receive the task execution results returned by the server node and forward the task execution results to the corresponding client user.
[0130] The present application also provides a computer non-volatile readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method for scheduling processor resources applied to each server node in a server node cluster is implemented. The method includes:
[0131] The receiving agent node forwards the current task request of the client user corresponding to any container. The current task request is generated by the client user and sent to the agent node uniformly. The agent node is configured to forward the current task request based on the correspondence between the client user ID and the container user ID.
[0132] Add the current task request to the task request queue of the processor corresponding to the client user.
[0133] When the processor resources are idle, the task request at the head of the task request queue is taken out, and the processor resources are allocated to the currently taken out task request so that the processor executes the task corresponding to the currently taken out task request.
[0134] After the task execution is completed, the task execution result is returned to the proxy node, and the processor resources are released. If the task request queue is not empty, jump to the step of taking out the task request at the head of the task request queue when the processor resources are idle, and allocate the processor resources to the currently taken task request so that the processor executes the task corresponding to the currently taken task request.
[0135] Alternatively, a processor resource scheduling method applied to an agent node in a server node cluster is executed, the method comprising:
[0136] Receive the current task request sent by the client user, and forward the current task request to the server node according to the correspondence between the client user ID and the container user ID of the container in the server node. The server node is the server node where the container corresponding to the client user is located. The current task request is used to instruct the server node to add the current task request to the task request queue of the processor corresponding to the client user, and when the processor resources are idle and the task request queue is not empty, take out the task request at the head of the task request queue, allocate the processor resources to the currently taken out task request, so that the processor executes the task corresponding to the currently taken out task request; after the task execution is completed, return the task execution result to the proxy node and release the processor resources, wherein each server node includes multiple containers, the client user and the container correspond one to one, each container corresponds to a processor, and each processor corresponds to multiple containers.
[0137] Receive the task execution results returned by the server node and forward the task execution results to the corresponding client user.
[0138] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0139] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a non-volatile computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.
[0140] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A processor resource scheduling system, characterized in that, Comprising: A client, a proxy node, and multiple server nodes. Each server node includes multiple containers, with a one-to-one correspondence between client users and the containers. Each container corresponds to a processor, and each processor corresponds to multiple containers; The client is configured to send a current task request generated by a client user to the proxy node; The proxy node is configured to forward the current task request according to the correspondence between the client user identity number ID and the container user ID, forward the current task request to the server node where the container corresponding to the client user is located, and receive the task execution result returned by the server node, and forward the task execution result to the client user; The server node is configured to add the current task request to the task request queue of the processor corresponding to the client user; When the processor resources are idle and the task request queue is not empty, take out the task request at the head of the task request queue, and allocate the processor resources to the currently taken out task request, so that the processor executes the task corresponding to the currently taken out task request; After the task execution is completed, return the task execution result to the proxy node and release the processor resources.
2. The processor resource scheduling system according to claim 1, wherein The server node is configured to obtain the container user ID corresponding to the currently running task in the processor. When the container user ID corresponding to the currently running task is the same as the container user ID corresponding to the current task request, add the current task request to the head of the task request queue; otherwise, add the current task request to the end of the task request queue. The server node is also configured to, after returning the task execution result to the proxy node after the task execution is completed and before releasing the processor resources, obtain the container user ID corresponding to the task request at the head of the task request queue. When the container user ID corresponding to the task request at the head is the same as the container user ID corresponding to the current task request, load the task request at the head onto the processor for execution.
3. The processor resource scheduling system according to claim 2, wherein The server node has a function of inserting the current task request after the prior task request corresponding to the client user of the current task request if it already exists in the task request queue; Otherwise, add the current task request to the end of the task request queue.
4. The processor resource scheduling system according to claim 1, wherein The server node is configured to return the task execution result to the proxy node after the task execution is completed, and release the processor resources when no new task request from the client user corresponding to the currently completed task is received within a preset time period. When a new task request from the client user corresponding to the currently completed task is received within the preset time period, load the new task request onto the processor for execution.
5. The processor resource scheduling system according to claim 1, wherein The server node is configured to release the processor resources when the processor executes the task corresponding to the currently taken out task request exceeds a preset execution time threshold.
6. The processor resource scheduling system according to any one of claims 1 to 5, characterized in that The proxy node is also configured to send the client user login information to the server node; The server node that receives the client user login information is further configured to, in the case where there is an empty container, establish a one-to-one correspondence between the client user and the empty container based on the client user login information, and return the information indicating successful container allocation to the proxy node; in the case where there is no empty container, return the information indicating failed container allocation to the proxy node; The proxy node is further configured to, in the case where it receives the information indicating failed container allocation, send the client user login information to the next server node in a polling manner.
7. A processor resource scheduling method, characterized in that, Applied to each server node in the server node cluster, the method includes: Receiving a current task request of a client user corresponding to any container forwarded by the proxy node, where the current task request is generated by the client user and uniformly sent to the proxy node, and the proxy node is configured to forward the current task request according to the correspondence between the client user ID and the container user ID; Adding the current task request to the task request queue of the processor corresponding to the client user; In the case where the processor resources are idle, taking out the task request at the head of the task request queue, and allocating the processor resources to the currently taken out task request, so that the processor executes the task corresponding to the currently taken out task request; After the task execution is completed, returning the task execution result to the proxy node, releasing the processor resources, and in the case where the task request queue is not empty, jumping to the step of, in the case where the processor resources are idle, taking out the task request at the head of the task request queue, and allocating the processor resources to the currently taken out task request, so that the processor executes the task corresponding to the currently taken out task request.
8. The processor resource scheduling method according to claim 7, wherein Adding the current task request to the task request queue of the processor corresponding to the client user includes: Obtaining the container user ID corresponding to the currently running task in the processor, and in the case where the container user ID corresponding to the currently running task is the same as the container user ID corresponding to the current task request, adding the current task request to the head of the task request queue; otherwise, adding the current task request to the tail of the task request queue; After returning the task execution result to the proxy node after the task execution is completed and before releasing the processor resources, it further includes: Obtaining the container user ID corresponding to the task request at the head of the task request queue, and in the case where the container user ID corresponding to the task request at the head is the same as the container user ID corresponding to the current task request, loading the task request at the head into the processor for execution.
9. The processor resource scheduling method according to claim 8, wherein Adding the current task request to the task request queue of the processor corresponding to the client user includes: In the case where there is a prior task request of the client user corresponding to the current task request in the task request queue, inserting the current task request after the prior task request; otherwise, adding the current task request to the tail of the task request queue.
10. The processor resource scheduling method according to claim 7, wherein After the task execution is completed, returning the task execution result to the proxy node and releasing the processor resources includes: After the task execution is completed, return the task execution result to the proxy node. When no new task request from the client user corresponding to the currently completed task is received within a preset time period, release the processor resources. When a new task request from the client user corresponding to the currently completed task is received within the preset time period, load the new task request onto the processor for execution.
11. The processor resource scheduling method according to claim 10, wherein When a new task request from the client user corresponding to the currently completed task is received within the preset time period, loading the new task request onto the processor for execution includes: When the task execution is completed, obtain the container user ID corresponding to the currently executed and completed task; Obtain the container user ID of the new task request received within the preset time period. When the container user ID of the new task request is the same as the container user ID corresponding to the currently executed and completed task, load the new task request onto the processor for execution.
12. The processor resource scheduling method according to claim 7, wherein It further includes: When the processor executes the task corresponding to the currently fetched task request and exceeds the preset execution time threshold, release the processor resources.
13. The processor resource scheduling method according to claim 7, wherein It further includes: When the processor executes the task corresponding to the currently fetched task request and exceeds the preset execution time threshold, save the current execution progress of the task, readd the currently fetched task request to the end of the task request queue, and release the processor resources.
14. The processor resource scheduling method according to any one of claims 7 to 13, characterized in that, Before receiving the current task request of the client user corresponding to any container forwarded by the proxy node, it further includes: Receive the client user login information sent by the proxy node. When there is an empty container in the current server node, establish a one-to-one correspondence between the client user and the empty container based on the client user login information, and return the information indicating successful container allocation to the proxy node; when there is no empty container in the current server node, return the information indicating failed container allocation to the proxy node to instruct the proxy node to send the client user login information to the next server node in a round-robin manner until the container allocation is successful.
15. A processor resource scheduling method, characterized in that, Applied to the proxy node in the server node cluster, the method includes: Receive the current task request sent by the client user, and forward the current task request to the server node according to the correspondence between the client user ID and the container user ID of the container in the server node. The server node is the server node where the container corresponding to the client user is located. The current task request is used to instruct the server node to add the current task request to the task request queue of the processor corresponding to the client user, and when the processor resources are idle and the task request queue is not empty, fetch the task request at the head of the task request queue, allocate the processor resources to the currently fetched task request, so that the processor executes the task corresponding to the currently fetched task request; after the task execution is completed, return the task execution result to the proxy node and release the processor resources; Receive the task execution result returned by the server node, and forward the task execution result to the corresponding client user.
16. The processor resource scheduling method according to claim 15, wherein Before receiving the current task request sent by the client user, it further includes: Send the client user login information to the server node, so that when the server node that receives the client user login information has an empty container, establish a one-to-one correspondence between the client user and the empty container based on the client user login information, and return the information indicating successful container allocation to the proxy node; when there is no empty container, return the information indicating failed container allocation to the proxy node. In the case of receiving the information indicating failed container allocation, send the client user login information to the next server node in a polling manner.
17. A processor resource scheduling device, characterized in that, Applied to each server node in the server node cluster, the device includes: A request receiving module, configured to receive the current task request of the client user corresponding to any container forwarded by the proxy node. The current task request is generated by the client user and uniformly sent to the proxy node, and the proxy node is configured to forward the current task request according to the correspondence between the client user ID and the container user ID. A request enqueueing module, configured to add the current task request to the task request queue of the processor corresponding to the client user. A resource allocation module, configured to, when the processor resources are idle, take out the task request at the head of the task request queue, allocate the processor resources to the currently taken out task request, so that the processor executes the task corresponding to the currently taken out task request. A result returning module, configured to return the task execution result to the proxy node after the task execution is completed, release the processor resources, and execute the resource allocation module when the task request queue is not empty.
18. A processor resource scheduling device, characterized in that, Applied to the proxy node in the server node cluster, the device includes: A task request forwarding module, configured to receive the current task request sent by the client user, and forward the current task request to the server node according to the correspondence between the client user ID and the container user ID of the container in the server node. The server node is the server node where the container corresponding to the client user is located. The current task request is used to instruct the server node to add the current task request to the task request queue of the processor corresponding to the client user, and when the processor resources are idle and the task request queue is not empty, take out the task request at the head of the task request queue, allocate the processor resources to the currently taken out task request, so that the processor executes the task corresponding to the currently taken out task request; after the task execution is completed, return the task execution result to the proxy node, and release the processor resources. An execution result forwarding module, configured to receive the task execution result returned by the server node and forward the task execution result to the corresponding client user.
19. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps of the processor resource scheduling method according to any one of claims 7 to 14, or when the processor executes the program, it implements the steps of the processor resource scheduling method according to any one of claims 15 to 16.
20. A computer non-volatile readable storage medium, characterized in that, A computer program is stored thereon, and when the computer program is executed by a processor, it implements the steps of the processor resource scheduling method according to any one of claims 7 to 14, or implements the steps of the processor resource scheduling method according to any one of claims 15 to 16.
Citation Information
Patent Citations
GPU resource using method and device and storage medium
CN110888743A
Resource configuration method, data processing method and device, equipment and storage medium
CN114924888A
Task processing method, system and device, electronic equipment and storage medium
CN115509713A
Processor resource scheduling method, device and system
CN117493022A
System and method for scheduling in a computing system
US20220229695A1
Cited By
Task scheduling method and device based on Jenkins and K8s and medium
CN121900921A