Task allocation method, device and equipment and computer readable storage medium
By determining the server cluster based on the CPU and memory requirements of the model training task in the model training platform and allocating tasks in the cluster, the task startup failure caused by insufficient servers is solved, and the task startup success rate and execution reliability are improved.
Patent Information
- Application Number
- CN202510213235.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-27
AI Technical Summary
In a model training platform, when the number of types available for users is smaller than the number determined by users, the model training task is started, which reduces the success rate of the model training task.
By obtaining the CPU and memory number applied for by the model training task, at least one server cluster is determined, the available CPU and memory number of the cluster meets the task requirements, and a second server cluster is determined in the server cluster, and the model training task is allocated to the second server cluster for execution.
This method effectively avoids task startup failure due to insufficient number of servers, and improves the startup success rate and execution reliability of model training tasks.
Smart Images

Figure CN120045331A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of computer technology, and in particular, to a task allocation method, device, equipment, and computer-readable storage medium. Background Art
[0002] With the continuous development of computer technology, the functions of terminal devices are becoming more and more powerful. For example, a terminal device can execute a model training task.
[0003] In the related art, a model training platform runs on a terminal device, and multiple servers are deployed in the model training platform. When a user needs to execute a model training task on the model training platform, the user needs to determine the type and quantity of servers required to execute the model training task, and the model training platform allocates the model training task to a server cluster so that the server cluster executes the model training task. Among them, the type of servers included in the server cluster is the type determined by the user, and the quantity of servers included is the quantity determined by the user.
[0004] However, when the quantity of servers of the type determined by the user available in the model training platform is less than the quantity determined by the user, the model training task fails to start, thereby reducing the startup success rate of the model training task. Summary of the Invention
[0005] The embodiments of the present application provide a task allocation method, device, equipment, and computer-readable storage medium, which can be used to solve the problem that the model training task fails to start and the startup success rate of the model training task is reduced in the related art. The technical solutions are as follows:
[0006] On the one hand, the embodiments of the present application provide a task allocation method, and the method includes:
[0007] Obtain a first CPU quantity of a central processing unit (CPU) and a first memory quantity of a memory applied for the model training task;
[0008] Determine at least one first server cluster according to the first CPU quantity and the first memory quantity, where at least one server is included in the first server cluster, and the available CPU quantity of the first server cluster is not less than the first CPU quantity, and the available memory quantity is not less than the first memory quantity;
[0009] Determine a second server cluster from the at least one first server cluster;
[0010] Allocate the model training task to the second server cluster.
[0011] In a possible implementation, determining a second server cluster in the at least one first server cluster includes:
[0012] If the number of the first server clusters is one, determining the first server cluster as the second server cluster;
[0013] If the number of the first server clusters is multiple, determining the available CPU quantity and the available memory quantity of each first server cluster; determining the second server cluster according to at least one of the available CPU quantity and the available memory quantity of each first server cluster.
[0014] In a possible implementation, determining the second server cluster according to at least one of the available CPU quantity and the available memory quantity of each first server cluster includes:
[0015] Determining the first server cluster with the least available CPU quantity in each first server cluster as the second server cluster; or,
[0016] Determining the first server cluster with the least available memory quantity in each first server cluster as the second server cluster; or,
[0017] Determining the metric value of each first server cluster according to the available CPU quantity and the available memory quantity of each first server cluster, where the metric value of the first server cluster is used to indicate the performance of the first server cluster; determining the first server cluster with the lowest metric value in each first server cluster as the second server cluster.
[0018] In a possible implementation, after allocating the model training task to the second server cluster, the method further includes:
[0019] Obtaining a first start command;
[0020] Starting the model training task based on the first start command, and executing the model training task through the CPUs with the first CPU quantity and the memory with the first memory quantity in the second server cluster.
[0021] In a possible implementation, the method further includes:
[0022] Obtaining a first GPU model and a first GPU quantity of the graphics processing unit (GPU) applied for the model training task;
[0023] Determining a second CPU quantity and a second memory quantity allowed to be applied for the model training task according to the first GPU model and the first GPU quantity;
[0024] Determining at least one first server cluster according to the first CPU quantity and the first memory quantity, including:
[0025] In a case where the second CPU quantity is not less than the first CPU quantity and the second memory quantity is not less than the first memory quantity, determining at least one first server cluster according to the first GPU model, the first GPU quantity, the first CPU quantity, and the first memory quantity, where the GPUs included in the servers in the first server cluster are of the first GPU model, and the available GPU quantity of the first server cluster is not less than the first GPU quantity.
[0026] In a possible implementation, the method further includes:
[0027] In a case where the second CPU quantity is less than the first CPU quantity, and / or the second memory quantity is less than the first memory quantity, displaying a first prompt message for prompting to change the quantity of CPUs applied for the model training task, and / or the quantity of memory applied for the model training task.
[0028] In a possible implementation, the determining the second CPU quantity and the second memory quantity allowed to be applied for the model training task according to the first GPU model and the first GPU quantity includes:
[0029] Determining the CPU application quantity and the memory application quantity corresponding to the GPUs of the first GPU model;
[0030] Determining the second CPU quantity allowed to be applied for the model training task according to the first GPU quantity and the CPU application quantity;
[0031] Determining the second memory quantity allowed to be applied for the model training task according to the first GPU quantity and the memory application quantity.
[0032] In a possible implementation, the determining a second server cluster in the at least one first server cluster includes:
[0033] If the number of the first server clusters is one, determining the first server cluster as the second server cluster;
[0034] If the number of the first server clusters is multiple, determining the available GPU quantity, the available CPU quantity, and the available memory quantity of each first server cluster; determining the second server cluster according to at least one of the available GPU quantity, the available CPU quantity, and the available memory quantity of each first server cluster.
[0035] In a possible implementation, determining the second server cluster according to at least one of the available GPU quantity, available CPU quantity, and available memory quantity of each first server cluster includes:
[0036] Determining the first server cluster with the least available CPU quantity among each first server cluster as the second server cluster; or,
[0037] Determining the first server cluster with the least available memory quantity among each first server cluster as the second server cluster; or,
[0038] Determining the first server cluster with the least available GPU quantity among each first server cluster as the second server cluster; or,
[0039] According to the available GPU quantity, available CPU quantity, and available memory quantity of each first server cluster, determining the metric value of each first server cluster, where the metric value of the first server cluster is used to indicate the performance of the first server cluster; determining the first server cluster with the lowest metric value among each first server cluster as the second server cluster.
[0040] In a possible implementation, after allocating the model training task to the second server cluster, the method further includes:
[0041] Obtaining a second start command;
[0042] When the second start command meets the start requirements, starting the model training task based on the second start command, and executing the model training task through the GPUs with the first GPU quantity, CPUs with the first CPU quantity, and memory with the first memory quantity in the second server cluster.
[0043] In a possible implementation, the method further includes:
[0044] When the second start command does not meet the start requirements, displaying a second prompt message, where the second prompt message is used to indicate changing the start command.
[0045] On the other hand, an embodiment of the present application provides a task allocation device, and the device includes:
[0046] An obtaining module, configured to obtain the first CPU quantity of the central processing unit (CPU) and the first memory quantity of the memory applied for the model training task;
[0047] A determination module, configured to determine at least one first server cluster according to the first CPU quantity and the first memory quantity, where the first server cluster includes at least one server, and the available CPU quantity of the first server cluster is not less than the first CPU quantity, and the available memory quantity is not less than the first memory quantity;
[0048] The determination module is further configured to determine a second server cluster from the at least one first server cluster;
[0049] An allocation module, configured to allocate the model training task to the second server cluster.
[0050] In a possible implementation manner, the determination module is configured to, if the number of the first server clusters is one, determine the first server cluster as the second server cluster; if the number of the first server clusters is multiple, determine the available CPU quantity and the available memory quantity of each first server cluster; and determine the second server cluster according to at least one of the available CPU quantity and the available memory quantity of each first server cluster.
[0051] In a possible implementation manner, the determination module is configured to determine the first server cluster with the least available CPU quantity in each first server cluster as the second server cluster; or,
[0052] determine the first server cluster with the least available memory quantity in each first server cluster as the second server cluster; or,
[0053] Determine the metric value of each first server cluster according to the available CPU quantity and the available memory quantity of each first server cluster, where the metric value of the first server cluster is used to indicate the performance of the first server cluster; and determine the first server cluster with the lowest metric value in each first server cluster as the second server cluster.
[0054] In a possible implementation manner, the acquisition module is further configured to acquire a first startup command;
[0055] The apparatus further includes:
[0056] A startup module, configured to start the model training task based on the first startup command, and execute the model training task through CPUs with the first CPU quantity and memories with the first memory quantity in the second server cluster.
[0057] In a possible implementation manner, the acquisition module is further configured to acquire a first GPU model and a first GPU quantity of a graphics processing unit GPU applied for the model training task;
[0058] The determining module is further configured to determine, according to the first GPU model and the first GPU quantity, the second CPU quantity and the second memory quantity that the model training task is allowed to apply for;
[0059] The determining module is further configured to, when the second CPU quantity is not less than the first CPU quantity and the second memory quantity is not less than the first memory quantity, determine at least one first server cluster according to the first GPU model, the first GPU quantity, the first CPU quantity, and the first memory quantity. The servers in the first server cluster include GPUs of the first GPU model, and the available GPU quantity of the first server cluster is not less than the first GPU quantity.
[0060] In a possible implementation manner, the apparatus further includes:
[0061] A display module, configured to display a first prompt message when the second CPU quantity is less than the first CPU quantity, and / or the second memory quantity is less than the first memory quantity. The first prompt message is used to prompt to change the quantity of CPUs applied for the model training task, and / or the quantity of memory applied for the model training task.
[0062] In a possible implementation manner, the determining module is configured to determine the CPU application quantity and the memory application quantity corresponding to the GPUs of the first GPU model;
[0063] Determine the second CPU quantity that the model training task is allowed to apply for according to the first GPU quantity and the CPU application quantity;
[0064] Determine the second memory quantity that the model training task is allowed to apply for according to the first GPU quantity and the memory application quantity.
[0065] In a possible implementation manner, the determining module is configured to, if the number of the first server clusters is one, determine the first server cluster as the second server cluster;
[0066] If the number of the first server clusters is multiple, determine the available GPU quantity, the available CPU quantity, and the available memory quantity of each first server cluster; determine the second server cluster according to at least one of the available GPU quantity, the available CPU quantity, and the available memory quantity of each first server cluster.
[0067] In a possible implementation manner, the determining module is configured to determine the first server cluster with the least available CPU quantity among the first server clusters as the second server cluster; or,
[0068] Determine the first server cluster with the least available memory among the respective first server clusters as the second server cluster; or,
[0069] Determine the first server cluster with the least available GPUs among the respective first server clusters as the second server cluster; or,
[0070] Based on the available GPUs, available CPUs, and available memory of the respective first server clusters, determine the metric values of the respective first server clusters, where the metric value of a first server cluster is used to indicate the performance of the first server cluster; determine the first server cluster with the lowest metric value among the respective first server clusters as the second server cluster.
[0071] In a possible implementation, the obtaining module is further configured to obtain a second startup command;
[0072] The apparatus further includes:
[0073] A startup module, configured to, when the second startup command meets the startup requirements, start the model training task based on the second startup command, and execute the model training task using the GPUs with the number of the first GPUs, the CPUs with the number of the first CPUs, and the memory with the amount of the first memory in the second server cluster.
[0074] In a possible implementation, the apparatus further includes:
[0075] A display module, configured to, when the second startup command does not meet the startup requirements, display a second prompt message for indicating to change the startup command.
[0076] On the other hand, an embodiment of the present application provides a computer device, including a processor and a memory, where at least one program code is stored in the memory, and the at least one program code is loaded and executed by the processor to enable the computer device to implement any one of the above task allocation methods.
[0077] On the other hand, a computer-readable storage medium is further provided, where at least one program code is stored in the computer-readable storage medium, and the at least one program code is loaded and executed by a processor to enable a computer to implement any one of the above task allocation methods.
[0078] On the other hand, a computer program or a computer program product is further provided, where at least one computer instruction is stored in the computer program or the computer program product, and the at least one computer instruction is loaded and executed by a processor to enable a computer to implement any one of the above task allocation methods.
[0079] The technical solutions provided by the embodiments of the present application at least bring the following beneficial effects:
[0080] When applying for resources for a model training task in the embodiments of the present application, the applied resources are CPU and memory. In this way, based on the number of CPUs and the amount of memory applied, the first server cluster that meets the application requirements is determined, and the second server cluster is determined within the first server cluster, and the model training task is assigned to the second server cluster so that the model training task is executed in the second server cluster. Since the resources applied for by this method are CPU and memory rather than servers, the resources can be effectively and reasonably utilized. It avoids the problem that the model training task cannot be started due to a large number of applied servers, resulting in a low start success rate of the model training task, thereby being able to improve the start success rate of the model training task and enhance the execution reliability of the model training task. BRIEF DESCRIPTION OF THE DRAWINGS
[0081] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following described drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0082] Figure 1 is a schematic diagram of the implementation environment of a task allocation method provided by an embodiment of the present application;
[0083] Figure 2 is a flowchart of a task allocation method provided by an embodiment of the present application;
[0084] Figure 3 is a schematic diagram of the visualization interface of a model training platform provided by an embodiment of the present application;
[0085] Figure 4 is a schematic diagram of a cluster configuration interface provided by an embodiment of the present application;
[0086] Figure 5 is a schematic diagram of the display of a resource application page provided by an embodiment of the present application;
[0087] Figure 6 is a schematic diagram of the display of another resource application page provided by an embodiment of the present application;
[0088] Figure 7 is a schematic diagram of a system configuration page provided by an embodiment of the present application;
[0089] Figure 8It is a schematic diagram of a system configuration page provided by an embodiment of the present application;
[0090] Figure 9 It is a flowchart of a task allocation method provided by an embodiment of the present application;
[0091] Figure 10 It is a schematic structural diagram of a task allocation device provided by an embodiment of the present application;
[0092] Figure 11 It is a schematic structural diagram of a terminal device provided by an embodiment of the present application. Detailed implementation manners
[0093] To make the objectives, technical solutions, and advantages of the present application clearer, the embodiments of the present application will be further described in detail below in conjunction with the accompanying drawings.
[0094] It should be noted that the terms "first", "second", etc. in the present application are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0095] Figure 1 It is a schematic diagram of the implementation environment of a task allocation method provided by an embodiment of the present application. As Figure 1 shown, the implementation environment includes: a terminal device 101, and the task allocation method provided by the embodiment of the present application is executed by the terminal device 101.
[0096] Optionally, the terminal device 101 can be any electronic device product that can perform human-computer interaction with a user in one or more ways such as a keyboard, a touchpad, a remote control, voice interaction, or a handwriting device. For example, the terminal device 101 can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, a PC (Personal Computer), a mobile phone, a PDA (Personal Digital Assistant), a wearable device, a PPC (Pocket PC), a smart vehicle console, a smart TV, etc.
[0097] The terminal device 101 can generally refer to one of multiple terminal devices. In this embodiment, only the terminal device 101 is used as an example for illustration. Those skilled in the art can understand that the number of the above-mentioned terminal devices 101 can be more or less. For example, the above-mentioned terminal device 101 can be only one, or the above-mentioned terminal device 101 can be dozens or hundreds, or a larger number. The embodiments of the present application do not limit the number and device type of the terminal device 101.
[0098] Those skilled in the art should understand that the above-mentioned terminal device 101 is only for illustration. Other existing or future possible terminal devices, if applicable to the present application, should also be included within the protection scope of the present application and are hereby incorporated by reference.
[0099] The embodiment of the present application provides a task allocation method, which can be applied to the above Figure 1 shown implementation environment. Taking the flowchart of a task allocation method provided by the embodiment of the present application shown in Figure 2 as an example, this method can be executed by the terminal device 101 in Figure 1 . As shown in Figure 2 , this method includes the following steps 201 to step 204.
[0100] In step 201, obtain the first CPU quantity of the central processing unit (CPU) applied for the model training task and the first memory quantity of the memory.
[0101] In an exemplary embodiment of the present application, a model training platform runs in the terminal device. Multiple servers are deployed in the model training platform, and the model training platform is used to execute the model training task. The model training platform is a software system or a collection of tools for supporting machine learning and deep learning model training. Optionally, the model training platform running in the terminal device can be any model training platform capable of executing the model training task. The embodiments of the present application do not limit the type of the model training platform running in the terminal device.
[0102] In a possible implementation manner, the model training platform supports configuring each server belonging to the model training platform in the system configuration module. After the configuration is completed, the model training platform can regularly check the utilization rate of each GPU (Graphic Process Unit) in each server. The user can view the utilization rate of the GPUs in each server belonging to the model training platform through the visualization interface of the model training platform. The user can also view the tasks executed in the GPUs in each server belonging to the model training platform and the status detection time of the GPUs through the visualization interface of the model training platform.
[0103] As Figure 3It is a schematic diagram of the visualization interface of a model training platform provided by an embodiment of the present application. It shows the utilization rate of GPUs, the tasks executed in the GPUs, and the status detection time of the GPUs in each server belonging to the model training platform. For example, a server with an IP address of 10.10.177.11, a name of GPU-177-11, and a GPU model of v100 includes 8 GPUs. Among them, the name of the task executed in the first GPU is biyuefeng-v100-4-1705023191-train-v0013-m5477, the status detection time is 2024-01-12 01:33:11, and the utilization rate is 0-20. The name of the task executed in the second GPU is biyuefeng-v100-4-1705023191-train-v0013-m5477, the status detection time is 2024-01-12 01:33:11, and the utilization rate is 60-80. The utilization rate, the tasks executed, and the status detection time of the GPUs included in other servers are shown in Figure 3 as shown, and will not be elaborated here one by one.
[0104] In this implementation, by displaying the visualization interface of the model training platform, users can view the utilization rate of GPUs, the tasks executed, and the status detection time of the GPUs in each server deployed in the model training platform, so that users can understand the resource occupancy of each server deployed in the model training platform, thereby facilitating users to apply for resources for model training tasks.
[0105] Optionally, the administrator of the model training platform can configure the relevant information of each server deployed in the model training platform in the cluster configuration interface. The relevant information of any server includes at least one of the name of any server, the model of the included GPUs, whether it is configured with a dedicated inter-machine communication high-speed network (Infiniband, IB), the status, the number of used CPUs (Central Processing Unit), the total number of included CPUs, the number of used memories, the total number of included memories, the number of used GPUs, and the total number of included GPUs. Users can view the relevant information of each server deployed in the model training platform in the cluster configuration interface.
[0106] Such as Figure 4It is a schematic diagram of a cluster configuration interface provided by an embodiment of the present application. It shows relevant information of 9 servers deployed in the model training platform. The name of the first server is node-50, the included GPU model is a10, it is configured with IB, the status is ready, the total number of included CPUs is 64000m (millicores), the used number of included CPUs is 567m, the total number of included memories is 125.41Gi (gigabytes), the used number of included memories is 19.34Gi, the total number of included GPUs is 4, and the used number of included GPUs is 0. See the relevant information of other servers in Figure 4 As shown, the embodiments of the present application will not be elaborated one by one here.
[0107] In a possible implementation manner, when a user hopes to execute a model training task in the model training platform, the user needs to apply for resources for the model training task. Optionally, the user can apply for CPUs with a first CPU quantity and memories with a first memory quantity for the model training task. The embodiments of the present application do not limit the manner in which the user applies for CPUs with a first CPU quantity and memories with a first memory quantity for the model training task. Optionally, a resource application page is displayed. The resource application page displays a CPU quantity box and a memory quantity box. Among them, the CPU quantity box is used to obtain the quantity of CPUs applied for the model training task, and the memory quantity box is used to obtain the quantity of memories applied for the model training task. In response to an input operation on the CPU quantity box, it is determined that the content input in the CPU quantity box is the first CPU quantity of the CPUs applied for the model training task; in response to an input operation on the memory quantity box, it is determined that the content input in the memory quantity box is the first memory quantity of the memories applied for the model training task.
[0108] In this implementation manner, the user can input the quantity of CPUs and the quantity of memories applied for the model training task in the CPU quantity box and the memory quantity box, so that the user can independently determine the quantity of CPUs and the quantity of memories applied for the model training task, making the applied quantity of CPUs and the quantity of memories more flexible.
[0109] Such as Figure 5 It is a schematic diagram showing the display of a resource application page provided by an embodiment of the present application. Among them, the resource type "CPU" is selected, indicating that the model training task requires CPU resources. Figure 5 It shows a CPU quantity box 501 and a memory quantity box 502. Among them, the content input in the CPU quantity box 501 is 2, that is, the first CPU quantity of the CPUs applied for the model training task is 2, and the content input in the memory quantity box 502 is 4, that is, the first memory quantity of the memories applied for the model training task is 4GB (Gigabyte, billion bytes).
[0110] Optionally, when it is desired that the model training platform perform a model training task, it is also necessary to configure a training algorithm, a training image, and a training dataset for the model training task. Optionally, a training algorithm box 503, a training image box 504, and a training dataset box 505 are also displayed on the resource application page. In response to a selection operation on the training algorithm box, optional training algorithms are displayed; in response to a selection operation on any one of the optional training algorithms, it is determined that any one of the training algorithms is the training algorithm configured for the model training task. In response to a selection operation on the training image box, optional training images are displayed; in response to a selection operation on any one of the optional training images, it is determined that any one of the training images is the training image configured for the model training task. In response to a selection operation on the training dataset box, optional training datasets are displayed; in response to a selection operation on any one of the optional training datasets, it is determined that any one of the training datasets is the training dataset configured for the model training task.
[0111] Optionally, the user can also select the training method for the model training task. A single-machine training control 506 and a multi-machine multi-card control 507 are also displayed on the resource application page. When the user selects the single-machine training control, it means that the user desires that the model training task be executed separately by one server. When the user selects the multi-machine multi-card control, it means that the user desires that the model training task be executed jointly by multiple servers.
[0112] In a possible implementation, the user can also apply for a first number of GPUs of a first GPU model for the model training task. At this time, the user selects the resource type "GPU", and a GPU model box and a GPU number box are also displayed on the resource application page. Among them, the GPU model box is used to determine the model of the GPU applied for the model training task, and the GPU number box is used to determine the number of GPUs applied for the model training task. In response to a trigger operation on the GPU model box, optional GPU models are displayed; in response to a trigger operation on any one of the optional GPU models, it is determined that any one of the GPU models is the first GPU model of the GPU applied for the model training task. In response to an input operation on the GPU number box, it is determined that the content input in the GPU number box is the first number of GPUs of the GPU applied for the model training task.
[0113] In this implementation, the user can also determine the model of the GPU applied for the model training task through the GPU model box and determine the number of GPUs applied for the model training task through the GPU number box, so that the user can more flexibly determine the model and number of GPUs required for the model training task.
[0114] Such as Figure 6It is a schematic diagram showing another display of the resource application page provided by an embodiment of this application. Among them, the resource type "GPU" is selected, indicating that the model training task requires GPU resources. Figure 6 It shows a GPU model box 601, a GPU quantity box 602, a CPU quantity box 603, and a memory quantity box 604. Among them, the content displayed in the GPU model box 601 is a10, that is, the first GPU model of the GPU applied for the model training task is a10; the content input in the GPU quantity box 602 is 2, that is, the first GPU quantity applied for the model training task is 2; the content input in the CPU quantity box 603 is 3, that is, the first CPU quantity of the CPU applied for the model training task is 3, and the content input in the memory quantity box 604 is 4, that is, the first memory quantity of the memory applied for the model training task is 4GB.
[0115] In step 202, according to the first CPU quantity and the first memory quantity, at least one first server cluster is determined. The first server cluster includes at least one server, and the available CPU quantity of the first server cluster is not less than the first CPU quantity, and the available memory quantity is not less than the first memory quantity.
[0116] In a possible implementation manner, after determining the first CPU quantity of the CPU and the first memory quantity of the memory applied for the model training task in the above step 201, it is also necessary to determine whether the training method selected by the user is single-machine training or multi-machine multi-card. If the training method selected by the user is single-machine training, then the number of servers included in the first server cluster determined according to the first CPU quantity and the first memory quantity is one. If the training method selected by the user is multi-machine multi-card, then the number of servers included in the first server cluster determined according to the first CPU quantity and the first memory quantity is multiple.
[0117] Optionally, after determining the first CPU quantity of the CPU and the first memory quantity of the memory applied for the model training task, if the training method selected by the user is single-machine training, then traverse each server deployed in the model training platform to determine the available CPU quantity and the available memory quantity of each server; determine candidate servers in each server whose available CPU quantity is not less than the first CPU quantity and whose available memory quantity is not less than the first memory quantity, and determine one candidate server as a first server cluster, so as to obtain at least one first server cluster.
[0118] Optionally, after determining the first CPU quantity of the CPUs applied for the model training task and the first memory quantity of the memory, if the selected training method is multi-machine multi-card, traverse each server deployed in the model training platform to determine the available CPU quantity and available memory quantity of each server; according to the available CPU quantity, available memory quantity, first CPU quantity, and first memory quantity of each server, determine at least one first server cluster, where a first server cluster includes multiple servers, and the available CPU quantity of a first server cluster is not less than the first CPU quantity, and the available memory quantity is not less than the first memory quantity.
[0119] In a possible implementation manner, after determining the first GPU model of the GPUs applied for the model training task, the first GPU quantity of the GPUs, the first CPU quantity of the CPUs, and the first memory quantity of the memory in step 201 above, it is also necessary to determine the second CPU quantity and the second memory quantity allowed to be applied for the model training task according to the first GPU model and the first GPU quantity; in the case where the second CPU quantity is not less than the first CPU quantity and the second memory quantity is not less than the first memory quantity, determine at least one first server cluster according to the first GPU model, the first GPU quantity, the first CPU quantity, and the first memory quantity. The first server cluster includes at least one server, the model of the GPUs included in the server is the first GPU model, and the available GPU quantity of the first server cluster is not less than the first GPU quantity, the available CPU quantity is not less than the first CPU quantity, and the available memory quantity is not less than the first memory quantity.
[0120] Among them, the process of determining the second CPU quantity and the second memory quantity allowed to be applied for the model training task according to the first GPU model and the first GPU quantity includes: determining the CPU application quantity and memory application quantity corresponding to the GPUs of the first GPU model; according to the first GPU quantity and the CPU application quantity, determining the second CPU quantity allowed to be applied for the model training task; according to the first GPU quantity and the memory application quantity, determining the second memory quantity allowed to be applied for the model training task.
[0121] In a possible implementation manner, the terminal device stores the CPU application quantity and memory application quantity corresponding to each model of GPUs. After the terminal device obtains the first GPU model, according to the first GPU model, the CPU application quantity and memory application quantity corresponding to each model of GPUs, determine the CPU application quantity and memory application quantity corresponding to the GPUs of the first GPU model.
[0122] Optionally, the administrator of the model training platform configures the CPU application quantity and memory application quantity corresponding to each model of GPUs in the system configuration page. Figure 7It is a schematic diagram of a system configuration page provided by an embodiment of the present application. Among them, the number of CPU applications corresponding to a GPU of model a10 is 16, and the amount of memory application is 32GB. The number of CPU applications corresponding to a GPU of model 3090 is 16, and the amount of memory application is 32GB. The number of CPU applications corresponding to a GPU of model v100 is 4, and the amount of memory application is 40GB. The number of CPU applications corresponding to a GPU of model Tesla-K40m is 16, and the amount of memory application is 32GB.
[0123] In a possible implementation, the system configuration page also displays the configuration time of the number of CPU applications and the amount of memory application corresponding to each model of GPU. For example, the configuration time of the number of CPU applications and the amount of memory application corresponding to a GPU of model a10 is 2024-08-07 14:44:20. For the configuration time of the number of CPU applications and the amount of memory application corresponding to other models of GPU, see Figure 7 as shown, and details are not elaborated here.
[0124] The system configuration page also displays an edit control and a delete control corresponding to each model of GPU. The edit control corresponding to any model of GPU is used to edit the number of CPU applications and the amount of memory application corresponding to any model of GPU. The delete control corresponding to any model of GPU is used to delete the number of CPU applications and the amount of memory application corresponding to any model of GPU. For example, Figure 7 701 in is the edit control corresponding to a GPU of model a10, and 702 is the delete control corresponding to a GPU of model a19.
[0125] Optionally, the process of determining the second number of CPUs allowed to be applied for the model training task according to the first number of GPUs and the number of CPU applications includes: determining the product between the first number of GPUs and the number of CPU applications as the second number of CPUs allowed to be applied for the model training task.
[0126] Exemplarily, if the first number of GPUs is 2 and the number of CPU applications is 16, then the second number of CPUs allowed to be applied for the model training task is 2 * 16 = 32.
[0127] Optionally, the process of determining the second amount of memory allowed to be applied for the model training task according to the first number of GPUs and the amount of memory application includes: determining the product between the first number of GPUs and the amount of memory application as the second amount of memory allowed to be applied for the model training task.
[0128] Exemplarily, if the first number of GPUs is 2 and the amount of memory application is 32, then the second amount of memory allowed to be applied for the model training task is 2 * 32 = 64.
[0129] In a possible implementation, after determining the second number of CPUs and the second amount of memory that the model training task is allowed to apply for, it is also necessary to determine whether the second number of CPUs is not less than the first number of CPUs and whether the second amount of memory is not less than the first amount of memory. In the case where the second number of CPUs is less than the first number of CPUs and / or the second amount of memory is less than the first amount of memory, a first prompt message is displayed, and the first prompt message is used to prompt to change the number of CPUs applied for the model training task and / or the amount of memory applied for the model training task.
[0130] In this implementation, if the number of CPUs applied for by the user is greater than the number of CPUs allowed to be applied for and / or the amount of memory applied for by the user is greater than the amount of memory allowed to be applied for, the user is prompted to re-apply to avoid waste of resources caused by the user applying for a relatively large number of CPUs and / or memory, thus affecting the use of other tasks.
[0131] In the case where the second number of CPUs is not less than the first number of CPUs and the second amount of memory is not less than the first amount of memory, if the training method selected by the user is single-machine training, then each server deployed in the model training platform is traversed to determine the server whose included GPU model is the first GPU model, and determine the available number of GPUs, available number of CPUs, and available amount of memory of the server whose included GPU model is the first GPU model; among the servers whose included GPU model is the first GPU model, a reference server is determined where the available number of GPUs is not less than the first number of GPUs, the available number of CPUs is not less than the first number of CPUs, and the available amount of memory is not less than the first amount of memory, and one reference server is determined as one first server cluster, thus obtaining at least one first server cluster.
[0132] In the case where the second number of CPUs is not less than the first number of CPUs and the second amount of memory is not less than the first amount of memory, if the training method selected by the user is multi-machine multi-card, then each server deployed in the model training platform is traversed to determine the server whose included GPU model is the first GPU model; determine the available number of GPUs, available number of CPUs, and available amount of memory of the server whose included GPU model is the first GPU model; according to the available number of GPUs, available number of CPUs, available amount of memory, first number of GPUs, first number of CPUs, and first amount of memory of the server whose included GPU model is the first GPU model, at least one first server cluster is determined, where one first server cluster includes multiple servers, and the available number of GPUs in one first server cluster is not less than the first number of GPUs, the available number of CPUs is not less than the first number of CPUs, and the available amount of memory is not less than the first amount of memory.
[0133] In step 203, determine a second server cluster in at least one first server cluster.
[0134] In a first possible implementation manner, after determining at least one first server cluster in step 202 above, the process of determining a second server cluster in at least one first server cluster includes: based on the number of first server clusters being one, determining the first server cluster as the second server cluster; based on the number of first server clusters being multiple, determining the available CPU quantity and available memory quantity of each first server cluster; and determining the second server cluster according to at least one of the available CPU quantity and available memory quantity of each first server cluster.
[0135] Optionally, the process of determining the second server cluster according to at least one of the available CPU quantity and available memory quantity of each first server cluster includes: determining the first server cluster with the least available CPU quantity among each first server cluster as the second server cluster; or determining the first server cluster with the least available memory quantity among each first server cluster as the second server cluster; or according to the available CPU quantity and available memory quantity of each first server cluster, determining the metric value of each first server cluster, where the metric value of the first server cluster is used to indicate the performance of the first server cluster; and determining the first server cluster with the lowest metric value among each first server cluster as the second server cluster.
[0136] In this implementation manner, three ways of determining the second server cluster in the first server cluster are provided, making the determination method of the second server cluster diversified and more flexible. Moreover, determining the first server cluster with the least available CPU quantity, or the first server cluster with the least available memory quantity, or the first server cluster with the lowest metric value as the second server cluster can effectively avoid resource waste and make the resources of the server be reasonably utilized.
[0137] In a possible implementation manner, the process of determining the metric value of each first server cluster according to the available CPU quantity and available memory quantity of each first server cluster includes: determining the metric value of each first server cluster according to the available CPU quantity, available memory quantity, weight parameter corresponding to the CPU, and weight parameter corresponding to the memory of each first server cluster. Among them, the weight parameter corresponding to the CPU and the weight parameter corresponding to the memory are both set based on experience or adjusted according to the implementation environment, and the embodiments of the present application do not limit this. Optionally, the sum value of the weight parameter corresponding to the CPU and the weight parameter corresponding to the memory is 1. Exemplarily, the weight parameter corresponding to the CPU is 0.7, and the weight parameter corresponding to the memory is 0.3.
[0138] In a possible implementation, according to the available CPU quantity, available memory quantity, weight parameter corresponding to the CPU, and weight parameter corresponding to the memory of each first server cluster, the metric value of each first server cluster is determined according to the following formula (1).
[0139] S = A * α + B * β (1)
[0140] In the above formula (1), S is the metric value of any first server cluster, A is the available CPU quantity of any first server cluster, B is the available memory quantity of any first server cluster, α is the weight parameter corresponding to the CPU, and β is the weight parameter corresponding to the memory.
[0141] In a second possible implementation, after determining at least one first server cluster in step 202 above, the process of determining the second server cluster in the at least one first server cluster includes: determining the first server cluster as the second server cluster based on the number of first server clusters being one; based on the number of first server clusters being multiple, determining the available GPU quantity, available CPU quantity, and available memory quantity of each first server cluster; and determining the second server cluster according to at least one of the available GPU quantity, available CPU quantity, and available memory quantity of each first server cluster.
[0142] Among them, the process of determining the second server cluster according to at least one of the available GPU quantity, available CPU quantity, and available memory quantity of each first server cluster includes: determining the first server cluster with the least available CPU quantity among each first server cluster as the second server cluster; or determining the first server cluster with the least available memory quantity among each first server cluster as the second server cluster; or determining the first server cluster with the least available GPU quantity among each first server cluster as the second server cluster; or determining the metric value of each first server cluster according to the available GPU quantity, available CPU quantity, and available memory quantity of each first server cluster, where the metric value of the first server cluster is used to indicate the performance of the first server cluster, and determining the first server cluster with the lowest metric value among each first server cluster as the second server cluster.
[0143] In this implementation, four methods for determining the second server cluster in the first server cluster are provided, making the determination method of the second server cluster more diverse and more flexible. Moreover, determining the first server cluster with the least available CPU quantity, or the first server cluster with the least available memory quantity, or the first server cluster with the lowest metric value, or the first server cluster with the lowest available GPU quantity as the second server cluster can effectively avoid resource waste and make the resources of the server be reasonably utilized.
[0144] Optionally, the process of determining the metric values of each first server cluster according to the available number of GPUs, available number of CPUs, and available memory of each first server cluster includes: determining the metric values of each first server cluster according to the available number of GPUs, available number of CPUs, available memory, weight parameter corresponding to the GPU, weight parameter corresponding to the CPU, and weight parameter corresponding to the memory of each first server cluster. Among them, the weight parameter corresponding to the GPU is also set based on experience or adjusted according to the implementation environment, and the embodiments of the present application do not limit this. Optionally, the sum of the weight parameter corresponding to the GPU, the weight parameter corresponding to the CPU, and the weight parameter corresponding to the memory is 1. Exemplarily, the weight parameter corresponding to the GPU is 0.4, the weight parameter corresponding to the CPU is 0.4, and the weight parameter corresponding to the memory is 0.2.
[0145] In a possible implementation manner, according to the available number of GPUs, available number of CPUs, available memory, weight parameter corresponding to the GPU, weight parameter corresponding to the CPU, and weight parameter corresponding to the memory of each first server cluster, the metric value of each first server cluster is determined according to the following formula (2).
[0146] S = P * γ + A * α + B * β (2)
[0147] In the above formula (2), S is the metric value of any first server cluster, P is the available number of GPUs of any first server cluster, A is the available number of CPUs of any first server cluster, B is the available memory of any first server cluster, γ is the weight parameter corresponding to the GPU, α is the weight parameter corresponding to the CPU, and β is the weight parameter corresponding to the memory.
[0148] In step 204, the model training task is assigned to the second server cluster.
[0149] Optionally, after determining the second server cluster in the above step 203, the model training task is assigned to the second server cluster so that the second server cluster executes the model training task.
[0150] In a possible implementation manner, after the model training task is assigned to the second server cluster, a first start command can also be obtained, and the model training task is started based on the first start command, and the model training task is executed by the CPUs with the number of the first CPUs and the memory with the number of the first memory in the second server cluster.
[0151] Among them, the embodiments of the present application do not limit the manner of obtaining the first start command. Optionally, a start command box is also displayed on the resource application page, such as Figure 5Among them, 508 is the start command box. In response to an input operation in the start command box, it is determined that the content input in the start command box is the first start command. For example, Figure 5 if the content input in the start command box is top as in Figure 5 , it is determined that the first start command is top.
[0152] In another possible implementation, after allocating the model training task to the second server cluster, a second start command can also be obtained. When the second start command meets the start requirements, the model training task is started based on the second start command, and the model training task is executed by GPUs with the number of the first GPUs, CPUs with the number of the first CPUs, and memory with the number of the first memory in the second server cluster.
[0153] Among them, the method of obtaining the second start command is similar to the method of obtaining the first start command, and the embodiments of the present application will not elaborate here.
[0154] Since the resources applied for the model training task include GPUs, CPUs, and memory, after obtaining the second start command, it is necessary to determine whether the second start command occupies CPUs and GPUs. When the second start command occupies CPUs and GPUs, it is determined that the second start command meets the start requirements. When the second start command only occupies CPUs, it is determined that the second start command does not meet the start requirements.
[0155] Optionally, start commands that only occupy CPUs are stored in the terminal device. After the terminal device obtains the second start command, it is determined whether the start commands that only occupy CPUs include the second start command. If the start commands that only occupy CPUs do not include the second start command, it is determined that the second start command meets the start requirements; otherwise, if the start commands that only occupy CPUs include the second start command, it is determined that the second start command does not meet the start requirements.
[0156] Exemplarily, the start commands that only occupy CPUs include top, top - b, and tail - f. The second start command is top. Since the start commands that only occupy CPUs include the second start command, it is determined that the second start command does not meet the start requirements. On the contrary, the second start command is AAA. Since the start commands that only occupy CPUs do not include the second start command, it is determined that the second start command meets the start requirements.
[0157] In a possible implementation, the administrator of the model training platform configures the start commands that only occupy CPUs in the system configuration page. After the administrator of the model training platform configures the start commands that only occupy CPUs in the system configuration page, the terminal device stores the start commands that only occupy CPUs.
[0158] For example, Figure 8It is a schematic diagram of a system configuration page provided by an embodiment of the present application. It shows two startup commands that only occupy the CPU and relevant information about the startup commands that only occupy the CPU. The relevant information about the startup commands that only occupy the CPU includes the description information, configuration time, corresponding editing control, and corresponding deletion control of the startup commands that only occupy the CPU. The editing control corresponding to the startup command that only occupies the CPU is used to edit the startup command that only occupies the CPU, and the deletion control corresponding to the startup command that only occupies the CPU is used to delete the startup command that only occupies the CPU. For example, Figure 8 In Figure 8 , the description information of the startup command "top" that only occupies the CPU is "illegal command for occupying the graphics card", the configuration time is 2024-08-02 15:14:57, the corresponding editing control is 801, and the corresponding deletion control is 802. The description information of the startup command "top-b" that only occupies the CPU is "illegal command for occupying the graphics card", the configuration time is 2024-08-02 15:18:30, the corresponding editing control is 803, and the corresponding deletion control is 804.
[0159] In a possible implementation manner, after the second startup command is obtained, when the second startup command does not meet the startup requirements, a second prompt message is displayed, and the second prompt message is used to indicate to change the startup command. Exemplarily, the second prompt message is "contains illegal command". By displaying the second prompt message, the user is enabled to change the startup command, so that the model training task can be successfully started, and thus the startup success rate of the model training task can be improved.
[0160] When the above method applies for resources for the model training task, the CPU and memory are applied for. In this way, based on the number of CPUs and the amount of memory applied for, the first server cluster that meets the application requirements is determined, and the second server cluster is determined within the first server cluster, and the model training task is assigned to the second server cluster so that the model training task is executed in the second server cluster. Since the resources applied for by this method are the CPU and memory rather than the server, the resources can be effectively and reasonably utilized. It avoids the problem that the number of servers applied for is large, resulting in the model training task being unable to start and the startup success rate of the model training task being low, thereby being able to improve the startup success rate of the model training task and enhance the execution reliability of the model training task.
[0161] Moreover, the GPU can also be applied for the model training task to enable the model training task to execute better.
[0162] In addition, when applying for a GPU for a model training task, when the start command for starting the model training task is obtained, it is also determined whether the start command only occupies the CPU. If the start command only occupies the CPU, the start command is re-obtained until the start command occupies both the CPU and the GPU. This can prevent the applied GPU from not being used when executing the model training task, thereby avoiding the incorrect occupation of the applied GPU, which may cause other tasks that need to use the GPU to be unable to apply for the GPU and affect the execution of other tasks.
[0163] Figure 9 is a flowchart of a task allocation method provided by an embodiment of the present application. As Figure 9 shown, the method includes the following steps 901 to step 918.
[0164] Step 901: Obtain the type of resources applied for the model training task.
[0165] In a possible implementation manner, the process of obtaining the type of resources applied for the model training task has been described in step 201 above and will not be elaborated here.
[0166] Step 902: Determine whether the type is GPU.
[0167] Based on the type being GPU, step 903 is executed; based on the type being non-GPU, step 913 is executed.
[0168] Step 903: Obtain the first GPU model of the GPU applied for the model training task.
[0169] In a possible implementation manner, the process of obtaining the first GPU model of the GPU applied for the model training task has been described in step 201 above and will not be elaborated here.
[0170] Step 904: Obtain the first GPU quantity of the GPU applied for the model training task, the first CPU quantity of the CPU, and the first memory quantity of the memory.
[0171] In a possible implementation manner, the process of obtaining the first GPU quantity of the GPU applied for the model training task, the first CPU quantity of the CPU, and the first memory quantity of the memory has been described in step 201 above and will not be elaborated here.
[0172] Step 905: Determine the second CPU quantity and the second memory quantity allowed to be applied for the model training task according to the first GPU model and the first GPU quantity.
[0173] In a possible implementation, the process of determining the allowable number of second CPUs and the amount of second memory for the model training task according to the first GPU model and the number of first GPUs has been described in step 201 above and will not be elaborated here.
[0174] Step 906: Determine whether the number of second CPUs is not less than the number of first CPUs and whether the amount of second memory is not less than the amount of first memory.
[0175] Based on the number of second CPUs being not less than the number of first CPUs and the amount of second memory being not less than the amount of first memory, execute step 907. Based on the number of second CPUs being less than the number of first CPUs and / or the amount of second memory being less than the amount of first memory, execute step 904.
[0176] Step 907: Determine at least one first server cluster according to the first GPU model, the number of first GPUs, the number of first CPUs, and the amount of first memory.
[0177] In a possible implementation, the process of determining at least one first server cluster according to the first GPU model, the number of first GPUs, the number of first CPUs, and the amount of first memory has been described in step 202 above and will not be elaborated here.
[0178] Step 908: Determine a second server cluster from at least one first server cluster.
[0179] In a possible implementation, the process of determining a second server cluster from at least one first server cluster has been described in step 203 above and will not be elaborated here.
[0180] Step 909: Allocate the model training task to the second server cluster.
[0181] In a possible implementation, the process of allocating the model training task to the second server cluster has been described in step 204 above and will not be elaborated here.
[0182] Step 910: Obtain a second startup command.
[0183] In a possible implementation, the process of obtaining a second startup command has been described in step 204 above and will not be elaborated here.
[0184] Step 911: Determine whether the second startup command only occupies the CPU.
[0185] Based on the second startup command only occupying the CPU, execute step 910; based on the second startup command not only occupying the CPU, execute step 912.
[0186] Step 912: Start the model training task based on the second start command, and execute the model training task using the GPUs with the first number of GPUs, the CPUs with the first number of CPUs, and the memory with the first amount of memory in the second server cluster.
[0187] In a possible implementation, the process of starting the model training task based on the second start command and executing the model training task using the GPUs with the first number of GPUs, the CPUs with the first number of CPUs, and the memory with the first amount of memory in the second server cluster has been described in step 204 above and will not be elaborated here.
[0188] Step 913: Obtain the first number of CPUs of the CPUs and the first amount of memory of the memory applied for the model training task.
[0189] In a possible implementation, the process of obtaining the first number of CPUs of the CPUs and the first amount of memory of the memory applied for the model training task has been described in step 201 above and will not be elaborated here.
[0190] Step 914: Determine at least one first server cluster according to the first number of CPUs and the first amount of memory.
[0191] In a possible implementation, the process of determining at least one first server cluster according to the first number of CPUs and the first amount of memory has been described in step 202 above and will not be elaborated here.
[0192] Step 915: Determine the second server cluster from at least one first server cluster.
[0193] In a possible implementation, the process of determining the second server cluster from at least one first server cluster has been described in step 203 above and will not be elaborated here.
[0194] Step 916: Allocate the model training task to the second server cluster.
[0195] In a possible implementation, the process of allocating the model training task to the second server cluster has been described in step 204 above and will not be elaborated here.
[0196] Step 917: Obtain the first start command.
[0197] In a possible implementation, the process of obtaining the first start command has been described in step 204 above and will not be elaborated here.
[0198] Step 918: Start a model training task based on the first startup command, and execute the model training task using the CPUs with the first number of CPUs and the memory with the first amount of memory in the second server cluster.
[0199] In a possible implementation, the process of starting a model training task based on the first startup command and executing the model training task using the CPUs with the first number of CPUs and the memory with the first amount of memory in the second server cluster has been described in the above step 204, and will not be elaborated here.
[0200] Figure 10 The following shows a schematic structural diagram of a task allocation device provided by an embodiment of the present application. As Figure 10 shown, the device includes:
[0201] An acquisition module 1001, configured to acquire the first number of CPUs of the central processing unit (CPU) and the first amount of memory applied for the model training task.
[0202] A determination module 1002, configured to determine at least one first server cluster according to the first number of CPUs and the first amount of memory. The first server cluster includes at least one server, and the available number of CPUs in the first server cluster is not less than the first number of CPUs, and the available amount of memory is not less than the first amount of memory.
[0203] The determination module 1002 is further configured to determine a second server cluster from the at least one first server cluster.
[0204] An allocation module 1003, configured to allocate the model training task to the second server cluster.
[0205] In a possible implementation, the determination module 1002 is configured to, if the number of first server clusters is one, determine the first server cluster as the second server cluster; if the number of first server clusters is multiple, determine the available number of CPUs and the available amount of memory of each first server cluster; and determine the second server cluster according to at least one of the available number of CPUs and the available amount of memory of each first server cluster.
[0206] In a possible implementation, the determination module 1002 is configured to determine the first server cluster with the least available number of CPUs in each first server cluster as the second server cluster; or,
[0207] determine the first server cluster with the least available amount of memory in each first server cluster as the second server cluster; or,
[0208] Determine the metric values of each first server cluster according to the available CPU quantity and available memory quantity of each first server cluster. The metric value of the first server cluster is used to indicate the performance of the first server cluster. Determine the first server cluster with the lowest metric value among each first server cluster as the second server cluster.
[0209] In a possible implementation, the obtaining module 1001 is further configured to obtain a first startup command.
[0210] The apparatus further includes:
[0211] A startup module, configured to start a model training task based on the first startup command, and execute the model training task through the CPUs with the quantity of the first CPU and the memory with the quantity of the first memory in the second server cluster.
[0212] In a possible implementation, the obtaining module 1001 is further configured to obtain a first GPU model and a first GPU quantity of a graphics processing unit (GPU) applied for the model training task.
[0213] The determining module 1002 is further configured to determine a second CPU quantity and a second memory quantity allowed to be applied for the model training task according to the first GPU model and the first GPU quantity.
[0214] The determining module 1002 is further configured to, when the second CPU quantity is not less than the first CPU quantity and the second memory quantity is not less than the first memory quantity, determine at least one first server cluster according to the first GPU model, the first GPU quantity, the first CPU quantity, and the first memory quantity. The servers in the first server cluster include GPUs with the first GPU model, and the available GPU quantity of the first server cluster is not less than the first GPU quantity.
[0215] In a possible implementation, the apparatus further includes:
[0216] A display module, configured to display a first prompt message when the second CPU quantity is less than the first CPU quantity and / or the second memory quantity is less than the first memory quantity. The first prompt message is used to prompt to change the quantity of the CPU applied for the model training task and / or the quantity of the memory applied for the model training task.
[0217] In a possible implementation, the determining module 1002 is configured to determine the CPU application quantity and the memory application quantity corresponding to the GPU with the first GPU model.
[0218] Determine the second CPU quantity allowed to be applied for the model training task according to the first GPU quantity and the CPU application quantity.
[0219] Determine the second memory quantity that the model training task is allowed to apply for according to the quantity of the first GPUs and the quantity of the memory applications.
[0220] In a possible implementation manner, the determining module 1002 is configured to, if the number of the first server clusters is one, determine the first server cluster as the second server cluster;
[0221] if the number of the first server clusters is multiple, determine the available GPU quantity, the available CPU quantity, and the available memory quantity of each first server cluster; determine the second server cluster according to at least one of the available GPU quantity, the available CPU quantity, and the available memory quantity of each first server cluster.
[0222] In a possible implementation manner, the determining module 1002 is configured to determine the first server cluster with the least available CPU quantity among each first server cluster as the second server cluster; or,
[0223] determine the first server cluster with the least available memory quantity among each first server cluster as the second server cluster; or,
[0224] determine the first server cluster with the least available GPU quantity among each first server cluster as the second server cluster; or,
[0225] Determine the metric value of each first server cluster according to the available GPU quantity, the available CPU quantity, and the available memory quantity of each first server cluster, where the metric value of the first server cluster is used to indicate the performance of the first server cluster; determine the first server cluster with the lowest metric value among each first server cluster as the second server cluster.
[0226] In a possible implementation manner, the obtaining module 1001 is further configured to obtain a second start command;
[0227] The apparatus further includes:
[0228] A start module, configured to, when the second start command meets the start requirements, start the model training task based on the second start command, and execute the model training task through the GPUs with the quantity of the first GPUs, the CPUs with the quantity of the first CPUs, and the memory with the quantity of the first memory in the second server cluster.
[0229] In a possible implementation manner, the apparatus further includes:
[0230] A display module, configured to, when the second start command does not meet the start requirements, display a second prompt message, where the second prompt message is used to indicate to change the start command.
[0231] When the above device applies for resources for a model training task, it applies for CPU and memory. In this way, based on the number of CPUs and the amount of memory applied for, the first server cluster that meets the application requirements is determined, and the second server cluster is determined within the first server cluster. The model training task is then assigned to the second server cluster so that the model training task can be executed in the second server cluster. Since the resources applied for by this method are CPU and memory, rather than servers, the resources can be effectively and reasonably utilized. This avoids the problem that a large number of servers are applied for, resulting in the inability to start the model training task and a low success rate of starting the model training task. Therefore, the success rate of starting the model training task can be improved, and the execution reliability of the model training task can be enhanced.
[0232] It should be understood that when the above-provided device implements its functions, only the above division of each functional module is used as an example for illustration. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device provided in the above embodiment and the method embodiment belong to the same concept. For the specific implementation process, please refer to the method embodiment, which will not be elaborated here.
[0233] Figure 11 The block diagram of the terminal device 1100 provided by an exemplary embodiment of the present application is shown. The terminal device 1100 can be any electronic device product that can perform human-computer interaction with the user in one or more ways such as a keyboard, a touchpad, a remote control, voice interaction, or a handwriting device. For example, a PC (Personal Computer), a mobile phone, a smart phone, a PDA (Personal Digital Assistant), a wearable device, a PPC (Pocket PC), a tablet computer, a smart in-vehicle device, a smart TV, a smart speaker, a smart watch, etc.
[0234] Generally, the terminal device 1100 includes: a processor 1101 and a memory 1102.
[0235] The processor 1101 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 1101 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 1101 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 1101 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1101 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.
[0236] The memory 1102 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 1102 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1102 is used to store at least one instruction, and the at least one instruction is used to be executed by the processor 1101 to implement the task allocation method provided in the method embodiments of the present application.
[0237] In some embodiments, the terminal device 1100 may further optionally include: a peripheral device interface 1103 and at least one peripheral device. The processor 1101, the memory 1102, and the peripheral device interface 1103 may be connected by a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 1103 through a bus, signal lines, or a circuit board. Specifically, the peripheral device includes at least one of a radio frequency circuit 1104, a display screen 1105, a camera assembly 1106, an audio circuit 1107, and a power supply 1108.
[0238] The peripheral device interface 1103 can be used to connect at least one I / O (Input / Output) related peripheral device to the processor 1101 and the memory 1102. In some embodiments, the processor 1101, the memory 1102, and the peripheral device interface 1103 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1101, the memory 1102, and the peripheral device interface 1103 can be implemented on separate chips or circuit boards, and this embodiment does not limit this.
[0239] The radio frequency circuit 1104 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 1104 communicates with a communication network and other communication devices through electromagnetic signals. The radio frequency circuit 1104 converts an electrical signal into an electromagnetic signal for transmission, or converts the received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 1104 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and so on. The radio frequency circuit 1104 can communicate with other terminal devices through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: the World Wide Web, a metropolitan area network, an intranet, each generation of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 1104 may further include a circuit related to NFC (Near Field Communication), and this application does not limit this.
[0240] The display screen 1105 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 1105 is a touch display screen, the display screen 1105 also has the ability to collect touch signals on or above the surface of the display screen 1105. The touch signals can be input as control signals to the processor 1101 for processing. At this time, the display screen 1105 can also be used to provide virtual buttons and / or virtual keyboards, also known as soft buttons and / or soft keyboards. In some embodiments, there may be one display screen 1105, which is provided on the front panel of the terminal device 1100; in other embodiments, there may be at least two display screens 1105, which are respectively provided on different surfaces of the terminal device 1100 or are in a foldable design; in other embodiments, the display screen 1105 may be a flexible display screen, which is provided on the curved surface or the folding surface of the terminal device 1100. Even, the display screen 1105 can also be set to an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 1105 can be prepared from materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0241] The camera module 1106 is used to collect images or videos. Optionally, the camera module 1106 includes a front camera and a rear camera. Generally, the front camera is provided on the front panel of the terminal device 1100, and the rear camera is provided on the back of the terminal device 1100. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth-of-field camera, a wide-angle camera, and a telephoto camera, so as to realize the function of background blurring by fusing the main camera and the depth-of-field camera, the function of panoramic shooting by fusing the main camera and the wide-angle camera, and the VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera module 1106 may further include a flash. The flash can be a single-color-temperature flash or a dual-color-temperature flash. The dual-color-temperature flash refers to the combination of a warm-light flash and a cold-light flash, which can be used for light compensation under different color temperatures.
[0242] The audio circuit 1107 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 1101 for processing, or input to the radio frequency circuit 1104 to achieve voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the terminal device 1100. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signal from the processor 1101 or the radio frequency circuit 1104 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into sound waves audible to humans, but also convert the electrical signal into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 1107 may further include a headphone jack.
[0243] The power supply 1108 is used to supply power to each component in the terminal device 1100. The power supply 1108 may be alternating current, direct current, a primary battery or a rechargeable battery. When the power supply 1108 includes a rechargeable battery, the rechargeable battery may be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery charged through a wired line, and a wireless rechargeable battery is a battery charged through a wireless coil. The rechargeable battery may also be used to support fast charging technology.
[0244] In some embodiments, the terminal device 1100 further includes one or more sensors 1109. The one or more sensors 1109 include but are not limited to: an acceleration sensor 1110, a gyroscope sensor 1111, a pressure sensor 1112, an optical sensor 1113, and a proximity sensor 1114.
[0245] The acceleration sensor 1110 can detect the magnitude of acceleration on the three coordinate axes of the coordinate system established with the terminal device 1100. For example, the acceleration sensor 1110 can be used to detect the components of the gravitational acceleration on the three coordinate axes. The processor 1101 can control the display screen 1105 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor 1110. The acceleration sensor 1110 can also be used for collecting game or user's motion data.
[0246] The gyroscope sensor 1111 can detect the body direction and rotation angle of the terminal device 1100. The gyroscope sensor 1111 can cooperate with the acceleration sensor 1110 to collect the 3D actions of the user on the terminal device 1100. According to the data collected by the gyroscope sensor 1111, the processor 1101 can implement the following functions: motion sensing (such as changing the UI according to the user's tilting operation), image stabilization during shooting, game control, and inertial navigation.
[0247] The pressure sensor 1112 can be disposed on the side frame of the terminal device 1100 and / or the lower layer of the display screen 1105. When the pressure sensor 1112 is disposed on the side frame of the terminal device 1100, it can detect the holding signal of the user on the terminal device 1100, and the processor 1101 can perform left / right hand recognition or quick operation according to the holding signal collected by the pressure sensor 1112. When the pressure sensor 1112 is disposed on the lower layer of the display screen 1105, the processor 1101 can control the operable controls on the UI interface according to the pressure operation of the user on the display screen 1105. The operable controls include at least one of a button control, a scroll bar control, an icon control, and a menu control.
[0248] The optical sensor 1113 is used to collect the ambient light intensity. In one embodiment, the processor 1101 can control the display brightness of the display screen 1105 according to the ambient light intensity collected by the optical sensor 1113. Specifically, when the ambient light intensity is high, the display brightness of the display screen 1105 is increased; when the ambient light intensity is low, the display brightness of the display screen 1105 is decreased. In another embodiment, the processor 1101 can also dynamically adjust the shooting parameters of the camera assembly 1106 according to the ambient light intensity collected by the optical sensor 1113.
[0249] The proximity sensor 1114, also known as a distance sensor, is usually disposed on the front panel of the terminal device 1100. The proximity sensor 1114 is used to collect the distance between the user and the front of the terminal device 1100. In one embodiment, when the proximity sensor 1114 detects that the distance between the user and the front of the terminal device 1100 is gradually decreasing, the processor 1101 controls the display screen 1105 to switch from the lit state to the off state; when the proximity sensor 1114 detects that the distance between the user and the front of the terminal device 1100 is gradually increasing, the processor 1101 controls the display screen 1105 to switch from the off state to the lit state.
[0250] Those skilled in the art can understand that Figure 11 the structure shown in does not limit the terminal device 1100, and it may include more or fewer components than shown in the figure, or combine some components, or adopt different component arrangements.
[0251] In an exemplary embodiment, a computer-readable storage medium is further provided, and at least one program code is stored in the storage medium. The at least one program code is loaded and executed by a processor to enable a computer to implement any one of the above task allocation methods.
[0252] Optionally, the above computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, a floppy disk, an optical data storage device, or the like.
[0253] In an exemplary embodiment, there is also provided a computer program or a computer program product. At least one computer instruction is stored in the computer program or the computer program product. The at least one computer instruction is loaded and executed by a processor to enable a computer to implement any one of the above task allocation methods.
[0254] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.), and signals involved in this application are all authorized by the user or fully authorized by all parties. The collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions. For example, the available number of CPUs, the available number of GPUs, and the available amount of memory in the server cluster involved in this application are obtained under full authorization.
[0255] It should be understood that the term "a plurality of" as mentioned herein refers to two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0256] The above are only exemplary embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the principles of the present application shall be included in the protection scope of the present application.
Claims
1. A task allocation method, characterized in that: The method comprises: Get the first CPU quantity of the central processing unit CPU and the first memory quantity of the memory applied for the model training task; Determine at least one first server cluster according to the first number of CPUs and the first number of memories, where the first server cluster includes at least one server, and the number of available CPUs in the first server cluster is not less than the first number of CPUs, and the number of available memories in the first server cluster is not less than the first number of memories; determining a second server cluster in the at least one first server cluster; Allocate the model training task to the second server cluster.
2. The method according to claim 1, characterized in that The determining of the second server cluster in the at least one first server cluster comprises: If the number of the first server clusters is one, determining the first server cluster to be the second server cluster; If there are multiple first server clusters, determine the number of available CPUs and the amount of available memory of each first server cluster; and determine the second server cluster based on at least one of the number of available CPUs and the amount of available memory of each first server cluster.
3. The method according to claim 2, characterized in that The determining the second server cluster according to at least one of the available CPU quantity and the available memory quantity of each of the first server clusters includes: Determine that the first server cluster with the least number of available CPUs among the first server clusters is the second server cluster; or, Determine that the first server cluster with the least available memory among the first server clusters is the second server cluster; or, According to the number of available CPUs and the number of available memories of the first server clusters, the index values of the first server clusters are determined, where the index values of the first server clusters are used to indicate the performance of the first server clusters; and the first server cluster with the lowest index value among the first server clusters is determined as the second server cluster.
4. The method according to claim 1, characterized in that: After allocating the model training task to the second server cluster, the method further includes: Get the first startup command; The model training task is started based on the first startup command, and the model training task is executed by using a first number of CPUs and a first number of memories in the second server cluster.
5. The method according to any one of claims 1 to 4, characterized in that: The method further comprises: Obtain a first GPU model and a first GPU quantity of a graphics processing unit GPU applied for the model training task; Determine the number of second CPUs and the second amount of memory allowed to be applied for the model training task according to the first GPU model and the first GPU number; The determining at least one first server cluster according to the first number of CPUs and the first number of memories includes: When the second number of CPUs is not less than the first number of CPUs, and the second number of memories is not less than the first number of memories, at least one first server cluster is determined according to the first GPU model, the first number of GPUs, the first number of CPUs, and the first number of memories, the model of the GPU included in the servers in the first server cluster is the first GPU model, and the number of available GPUs in the first server cluster is not less than the first number of GPUs.
6. The method according to claim 5, characterized in that The method further comprises: When the second number of CPUs is less than the first number of CPUs, and / or the second amount of memory is less than the first amount of memory, a first prompt message is displayed, wherein the first prompt message is used to prompt a change to the number of CPUs applied for the model training task, and / or the amount of memory applied for the model training task.
7. The method according to claim 5, characterized in that The determining, according to the first GPU model and the first GPU quantity, the second CPU quantity and the second memory quantity allowed to be applied for by the model training task comprises: Determine the CPU application quantity and memory application quantity corresponding to the GPU of the first GPU model; Determine the second number of CPUs allowed to be applied for the model training task according to the first number of GPUs and the number of CPUs applied for; According to the first number of GPUs and the memory application amount, a second memory amount allowed to be applied for the model training task is determined.
8. The method according to claim 5, characterized in that The determining of the second server cluster in the at least one first server cluster comprises: If the number of the first server clusters is one, determining the first server cluster to be the second server cluster; If there are multiple first server clusters, determine the number of available GPUs, the number of available CPUs and the amount of available memory of each first server cluster; determine the second server cluster based on at least one of the number of available GPUs, the number of available CPUs and the amount of available memory of each first server cluster.
9. The method according to claim 8, characterized in that The determining the second server cluster according to at least one of the available GPU quantity, available CPU quantity and available memory quantity of each of the first server clusters comprises: Determine that the first server cluster with the least number of available CPUs among the first server clusters is the second server cluster; or, Determine that the first server cluster with the least available memory among the first server clusters is the second server cluster; or, Determine that the first server cluster with the least number of available GPUs among the first server clusters is the second server cluster; or, Determine an index value of each of the first server clusters based on the number of available GPUs, the number of available CPUs, and the amount of available memory of each of the first server clusters, where the index value of the first server cluster is used to indicate the performance of the first server cluster; and determine that the first server cluster with the lowest index value among the first server clusters is the second server cluster.
10. The method according to claim 5, characterized in that After assigning the model training task to the second server cluster, the method further includes: Get the second startup command; When the second startup command meets the startup requirements, the model training task is started based on the second startup command, and the model training task is executed by using the first number of GPUs, the first number of CPUs, and the first number of memories in the second server cluster.
11. The method according to claim 10, characterized in that The method further comprises: In the case where the second startup command does not meet the startup requirement, a second prompt message is displayed, where the second prompt message is used to instruct to change the startup command.
12. A task allocation device, characterized in that: The device comprises: An acquisition module, used to acquire a first CPU quantity of a central processing unit CPU and a first memory quantity of a memory applied for a model training task; A determination module, configured to determine at least one first server cluster according to the first number of CPUs and the first number of memories, wherein the first server cluster includes at least one server, and the number of available CPUs in the first server cluster is not less than the first number of CPUs, and the number of available memories is not less than the first number of memories; The determination module is further configured to determine a second server cluster in the at least one first server cluster; An allocation module is used to allocate the model training task to the second server cluster.
13. A terminal device, characterized in that: The terminal device includes a processor and a memory, wherein the memory stores at least one program code, and the at least one program code is loaded and executed by the processor, so that the terminal device implements the task allocation method according to any one of claims 1 to 11.
14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores at least one program code, and the at least one program code is loaded and executed by a processor so that a computer implements the task allocation method according to any one of claims 1 to 11.
15. A computer program product, characterized in that The computer program product stores at least one computer instruction, and the at least one computer instruction is loaded and executed by a processor so that a computer implements the task allocation method according to any one of claims 1 to 11.