Resource allocation method of distributed training model, computing device and computing system

Through built-in preset configuration rules and predesign calculation rules, the resource configuration for artificial intelligence model training is automatically determined, which solves the inaccurate and wasteful resource configuration caused by user manual configuration, and improves training efficiency and ease of use.

CN120469798APending Publication Date: 2025-08-12HENAN KUNLUN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510549693.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

During the training process of artificial intelligence model, users need to manually specify resource configuration information, resulting in inaccurate resource configuration, which may cause memory overflow or waste of computing resources.

Method used

Through built-in preset configuration rules and predesign calculation rules, the resource configuration information is automatically determined, including the number of calculation nodes and the number of training cards, suitable for common and uncommon models.

Benefits of technology

It improves the accuracy and training efficiency of resource allocation, reduces the complexity of user configuration, and avoids resource waste.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120469798A_ABST
    Figure CN120469798A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, in particular to a resource allocation method of a distributed training model, computing equipment and a computing system. The method comprises the steps of displaying a model configuration interface; the model configuration interface is used for acquiring a training model selected by a user; the training model belongs to a preset model; the preset model is a trained model; determining resource configuration information corresponding to the training model based on a preset configuration rule; the preset configuration rule comprises a preset model and corresponding resource configuration information; wherein the resource configuration information comprises the number of computational nodes and the number of training cards; the training card is used for training the model; creating a distributed training task based on the resource configuration information; the distributed training task is used for executing training of the training model. Through the preset configuration rule of the built-in common model, the resource configuration information required by the common model is automatically determined, so that manual configuration is replaced by automatic configuration, and the accuracy, the training efficiency and the training usability of training model resource configuration are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a resource configuration method, computing device, and computing system for a distributed training model. Background Art

[0002] Currently, when using artificial intelligence (AI) development platforms to train and / or fine-tune AI models, users (e.g., algorithm engineers) need to manually specify resource configuration information when submitting training jobs, including the number of computing nodes and training card specifications (e.g., the number of graphics processing unit (GPU) cards, the number of central processing unit (CPU) cores, and memory).

[0003] If too few resources are configured during user configuration, out of memory (OOM) and poor training performance will occur, requiring users to invest more time in operation and maintenance. If too many resources are configured, computing resources will be wasted. Summary of the Invention

[0004] The embodiments of the present application provide a resource configuration method, computing device, and computing system for a distributed training model, which improves the accuracy of resource configuration of the training model by replacing manual configuration by automated configuration.

[0005] According to one aspect of an embodiment of the present application, a resource configuration method for a distributed training model is provided, the method comprising: displaying a model configuration interface; the model configuration interface is used to obtain a training model selected by a user; the training model belongs to a preset model; the preset model is a trained model; based on preset configuration rules, resource configuration information corresponding to the training model is determined; the preset configuration rules include the preset model and the corresponding resource configuration information; wherein the resource configuration information includes: the number of computing nodes and the number of training cards; the training cards are used for training models; creating distributed training tasks based on the resource configuration information; and the distributed training tasks are used to execute training model training.

[0006] By using pre-set configuration rules for built-in common models, the resource configuration information required for common models can be automatically determined, thereby replacing manual configuration with automated configuration, improving the accuracy of training model resource configuration, training efficiency, and ease of use.

[0007] In one possible embodiment, before displaying the model configuration interface, the method includes: determining one or more preset models to be displayed in the model configuration interface based on preset configuration rules.

[0008] In one possible embodiment, the method further includes: when the training model does not belong to a preset model, obtaining training scale information input by the user through the model configuration interface; the training scale information is used to determine the resources of the training model; and determining the resource configuration information based on the training scale information and preset calculation rules.

[0009] Through built-in preset calculation rules, the required resource configuration information for uncommon models is automatically calculated, thereby realizing automated configuration instead of manual configuration, improving the accuracy of training model resource configuration, training efficiency and training ease of use.

[0010] In one possible approach, the training scale information includes the number of model parameters and the amount of training data.

[0011] In one possible embodiment, the method further includes: determining a type of a training card based on the model parameter quantity; the type of the training card includes a module form or a peripheral component interconnect (PCI) form.

[0012] In one possible approach, the type of the training card is determined based on the model parameter quantity, including: when the model parameter quantity is greater than or equal to the threshold parameter quantity, determining the type of the training card to be a module form; when the model parameter quantity is less than the threshold parameter quantity, determining the type of the training card to be a PCI form.

[0013] In one possible embodiment, the method further includes: determining the number of computing nodes and the number of training cards based on the type of training cards, the number of model parameters, the amount of data, and the recommended training time.

[0014] In one possible approach, creating a distributed training task based on resource configuration information includes: using the resource configuration information as resource configuration information required for a distributed training model.

[0015] According to another aspect of an embodiment of the present application, a computing device is provided, wherein the computing device includes a memory and a processor; the memory and the processor are coupled; the memory is used to store computer-readable instructions; and the processor is used to run the computer-readable instructions so that the computing device executes the resource configuration method as described above.

[0016] According to another aspect of an embodiment of the present application, a computing system is provided, wherein the computing system includes a management node and a computing node; wherein the management node is a computing device as described above; the management node and the computing node are connected; and the computing node is used for distributed training of a training model.

[0017] According to another aspect of the embodiment of the present application, a resource configuration device for a distributed training model is provided, the device comprising: a display module for displaying a model configuration interface; the model configuration interface is used to obtain a training model selected by a user; the training model belongs to a preset model; the preset model is a trained model; a determination module is used to determine the resource configuration information corresponding to the training model based on a preset configuration rule; the preset configuration rule includes the preset model and the corresponding resource configuration information; wherein the resource configuration information includes: the number of computing nodes and the number of training cards; the training card is used to train the model; a creation module is used to create a distributed training task based on the resource configuration information; the distributed training task is used to perform training model training. According to another aspect of the embodiment of the present application, a non-transitory computer-readable storage medium is provided for storing computer-readable instructions, which, when the computer-readable instructions are executed by a processor, causes the processor to execute the resource configuration method described above.

[0018] According to another aspect of the embodiments of the present application, a computer program product is provided, including a computer program, wherein when the computer program is executed by a processor, the resource configuration method described above is implemented. It should be understood that both the foregoing general description and the following detailed description are exemplary and are intended to provide further explanation of the claimed technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The above and other purposes, features, and advantages of the present application will become more apparent through a more detailed description of the embodiments of the present application in conjunction with the accompanying drawings. The accompanying drawings are intended to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the drawings, the same reference numerals generally represent the same components or steps.

[0020] Figure 1 It is a scenario diagram illustrating an application scenario of a resource configuration method for a distributed training model according to an embodiment of the present application.

[0021] Figure 2 It is a method flow chart illustrating a resource configuration method for a distributed training model according to an embodiment of the present application.

[0022] Figure 3 It is a flow chart illustrating the overall process of the resource configuration method of the distributed training model according to an embodiment of the present application.

[0023] Figure 4 It is a flow chart that further illustrates the overall process of the resource configuration method of the distributed training model according to an embodiment of the present application.

[0024] Figure 52 is a schematic diagram illustrating a web portal interface according to an embodiment of the present application.

[0025] Figure 6 2 is a schematic diagram illustrating a resource configuration device for a distributed training model according to an embodiment of the present application.

[0026] Figure 7 is a hardware block diagram illustrating a computing device according to an embodiment of the present application.

[0027] Figure 8 is a schematic diagram illustrating a computer program product according to an embodiment of the present application. DETAILED DESCRIPTION

[0028] In order to make the purpose, technical solutions and advantages of this application more apparent, the following exemplary embodiments of this application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of this application, rather than all the embodiments of this application, and it should be understood that this application is not limited to the exemplary embodiments described herein.

[0029] First, in order to make the description of the following embodiments clear and concise, a brief introduction to the terms involved in this application is first given.

[0030] A server cluster is a group of loosely or tightly connected servers working together, typically used to execute large-scale jobs. Clustered servers typically offer higher computing efficiency than a single server of comparable speed or availability. By running parallel computations across multiple servers and utilizing diverse computing resources to solve problems, the cluster system's computing and processing speeds can be increased, boosting its overall performance. Each server in a cluster is connected via a network and runs its own operating system.

[0031] Containers enable different applications to run in relatively isolated, secure environments, isolating them from each other and the external environment. Containers are a lightweight virtualization technology with advantages such as fast startup, easy deployment and migration, and excellent security and scalability. They are widely used in fields such as AI and cloud computing.

[0032] A distributed task involves breaking down a large computing task into multiple subtasks. These subtasks are then assigned to individual compute nodes (e.g., servers) within a server cluster. Each compute node independently executes its assigned subtask, while nodes can exchange data and collaborate over the network. A typical distributed task is the distributed training of AI models. Training datasets can be divided into multiple batches, with each compute node training each batch in parallel, improving training efficiency.

[0033] Below, we will refer to Figure 1 The application scenarios according to the embodiments of the present application are summarized.

[0034] Figure 1 1 is a schematic diagram illustrating an application scenario of a resource configuration method for a distributed training model according to an embodiment of the present application. Figure 1 As shown, the application scenario includes at least one computing system, including: a management node 11 and a computing node 12 (for example: computing node 1, computing node 2, ... computing node n), the management node 11 and the computing node 12 are connected, and an AI development platform is deployed on the management node 11. Figure 7 The management node 11 is further described in .

[0035] Among them, the computing node 12 refers to the node in the distributed system that is responsible for executing the subtasks assigned by the management node 11. These subtasks are executed in parallel on different computing nodes 12, and the entire distributed task (such as the task of the distributed training model in the embodiment of the present application) is completed through mutual cooperation.

[0036] Based on the different functions implemented by different types of AI models, the computing nodes 12 can be different types of servers, such as CPU servers, GPU servers, neural network processing unit (NPU) servers, etc. Among them, CPU servers are more versatile and suitable for deep learning fields such as pattern recognition, target detection, and image processing; GPU servers can be used to specifically process image computing tasks and accelerate graphics rendering; NPU servers can be used to accelerate the training and inference calculations of artificial neural networks and can efficiently perform large-scale neural network calculations.

[0037] In one embodiment of the present application, the computing node 12 may be an NPU server, and the NPU server may include multiple NPU acceleration cards (eg, NPU1 to NPU8).

[0038] The management node 11 is a node in a distributed system that is responsible for coordinating and managing the execution of the entire distributed task. It can dispatch tasks to the computing nodes 12 that meet the resource configuration information when the user 10 (e.g., an algorithm engineer) issues a training task.

[0039] An AI development platform is a comprehensive platform that provides a series of tools, libraries, frameworks, and services to help developers design, train, deploy, and manage AI models more quickly and efficiently.

[0040] Figure 2 1 is a flowchart illustrating a method for configuring resources of a distributed training model according to an embodiment of the present application. Figure 2 As shown, the resource configuration method may include at least the following steps.

[0041] In step S201, a model configuration interface is displayed; the model configuration interface is used to obtain a training model selected by a user; the training model belongs to a preset model; the preset model is a trained model.

[0042] That is, in one embodiment of the present application, the management node 11 (or AI development platform) provides a model configuration interface to the user 10. The interface displays one or more preset models for the user 10 to select. These preset models can be understood as common models, which are all trained. Figure 5 Further description of .

[0043] In step S202, based on the preset configuration rules, the resource configuration information corresponding to the training model is determined; the preset configuration rules include the preset model and the corresponding resource configuration information; wherein the resource configuration information includes: the number of computing nodes and the number of training cards; the training cards are used for training the model.

[0044] In one embodiment of the present application, the preset configuration rules can be understood as a rule library of the above-mentioned common models, which may include one or more preset models and resource configuration information corresponding to each preset model.

[0045] That is to say, each preset model displayed to the user 10 in step S201 is included in the preset configuration rules, so the management node 11 can automatically determine the resource configuration information corresponding to the training model based on the preset configuration rules and the training model obtained in step S201, so as to achieve the technical effect of replacing user manual configuration with automated configuration and improving the accuracy of resource configuration of the training model.

[0046] In step S203, a distributed training task is created based on the resource configuration information; the distributed training task is used to perform training of the training model. That is, the management node 11 creates a distributed training task for the computing node 12 to perform training based on the resource configuration information obtained in step S202. Specifically, the resource configuration information obtained in step S202 is used as the resource configuration information required for the distributed training model. A more specific resource configuration method for the distributed training model will be combined with Figure 3-Figure 5 A further detailed description is given.

[0047] Figure 3 FIG is a flow chart illustrating the overall process of the resource configuration method of the distributed training model according to an embodiment of the present application. Figure 3 As shown, the overall process of the resource configuration method of the embodiment of the present application can include at least the following steps.

[0048] Step S1, determine the preset model set. Specifically, based on the preset configuration rules, determine one or more preset models displayed in the model configuration interface. That is, the management node 11 has built-in preset configuration rules, which may include one or more preset models (i.e., common models), and resource configuration information corresponding to each preset model. The management node 11 forms a preset model set based on the preset configuration rules, and presents the preset model set to the user 10 through the model configuration interface to facilitate the user 10 to select. It can be understood that this step can be understood as a prerequisite step for the above-mentioned step S201.

[0049] In one embodiment of the present application, the preset model set may be presented by setting a drop-down menu on the model configuration interface, and the drop-down menu includes one or more preset models mentioned above.

[0050] In one embodiment of the present application, the preset configuration rules may be as shown in Table 1.

[0051]

[0052] Table 1

[0053] It should be noted that the preset model set can be updated iteratively on a regular basis. That is, as training tasks accumulate, more preset models may be regarded as regular models and written into the preset configuration rules to further achieve automated configuration.

[0054] Step S2: Display the model configuration interface and obtain the training model or training scale information. Specifically, the management node 11 displays the model configuration interface to the user 10 and obtains the relevant information of the model that the user 10 needs to train through the model configuration interface.

[0055] If the model that user 10 wants to train is in the above drop-down menu (that is, the training model belongs to the preset model), user 10 can directly select it. In this case, the management node 11 obtains the training model; if the model that user 10 wants to train is not in the drop-down menu (that is, the training model does not belong to the preset model), in order to enable the management node 11 to more accurately match it with appropriate resources, user 10 needs to enter some scale information of the model to be trained on the model configuration interface. In this case, the management node 11 obtains the training scale information. It can be understood that this step includes step S201, or in other words, step S201 is part of the solution described in this step. Specifically, the model configuration interface can be as follows Figure 5 shown.

[0056] In one embodiment of the present application, the training scale information may include: the amount of model parameters and the amount of training data.

[0057] In step S3a, when the training model is obtained, resources are automatically configured. Specifically, if the training model is a preset model, in response to the training model selected by the user 10 obtained through the model configuration interface in step S2, resource configuration information corresponding to the training model is determined based on the preset configuration rules.

[0058] As described above, if a training model is received in step S2, the training model is considered a common model. Management node 11 can automatically determine the resource configuration information corresponding to the common model based on preset configuration rules. This avoids manual configuration by user 10, improving resource configuration accuracy, training efficiency, and training usability. It is understood that this step is the same as step S202 described above.

[0059] In step S3b, if the training scale information is obtained, resources are automatically calculated. Specifically, if the training model does not belong to the preset model, the training scale information input by the user 10 is obtained through the model configuration interface; the training scale information is used to determine the resources for the training model; and resource configuration information is determined based on the training scale information and preset calculation rules. The training scale information includes the number of model parameters and the amount of training data.

[0060] That is, the management node 11 also has built-in preset calculation rules. Specifically, the preset calculation rules may include preset thresholds and calculation formulas. When the training scale information received in step S2 is considered to be an uncommon model, the management node 11 will automatically calculate the resource configuration information corresponding to the uncommon model according to the preset calculation rules, thereby avoiding manual configuration by the user 10 and improving the accuracy of resource configuration, training efficiency, and training usability. The automatic calculation method may include the following steps:

[0061] (1) Determine the type of the training card based on the model parameter quantity; the training card type includes a module form or a peripheral component interconnect (PCI) form. Specifically, the management node 11 can automatically determine the type of the training card based on the model parameter quantity input by the user 10 and a preset threshold value in a preset calculation rule. If the model parameter quantity is greater than or equal to the threshold parameter quantity, the training card type is determined to be a module form; if the model parameter quantity is less than the threshold parameter quantity, the training card type is determined to be the PCI form.

[0062] In one embodiment of the present application, the preset threshold may be 30B (ie, billion). When the model parameter amount is ≥30B, a module-type accelerator card is used; when the model parameter amount is <30B, a PCI-type accelerator card is used.

[0063] Among them, modular accelerator cards are a type of accelerator card with a higher level of integration and a specific architectural design. They can provide more powerful computing power and more efficient parallel processing capabilities, making them suitable for handling large-scale computing tasks and effectively improving the efficiency of training and inference. PCI accelerator cards are also a type of accelerator card. This accelerator card can be connected to the computer motherboard through the PCI interface. Compared with modular accelerator cards, this type of accelerator card has weaker computing power and parallel processing capabilities, but it is also lower in cost and can meet computing needs under 30B, offering a higher cost-effectiveness.

[0064] (2) Determine the number of computing nodes and training cards based on the type of training card, the amount of model parameters, the amount of data, and the recommended training duration. Specifically, the management node 11 can automatically calculate the required number of computing nodes 12 and the required number of training cards based on the accelerator card type determined in (1) and the received information such as the amount of model parameters, the amount of training data, and the recommended training duration.

[0065] In one embodiment of the present application, the calculation formula for the number of accelerator cards may be:

[0066] Number of accelerator cards = 8 * amount of training data * number of model parameters / (training time * utilization rate), where the utilization rate is usually 50%.

[0067] In summary, the management node 11 automatically determines the resource configuration information required for common models through the built-in preset configuration rules of common models; and automatically calculates the resource configuration information required for uncommon models through the built-in preset calculation rules. The combination of the two covers all training models, thereby fully realizing automated configuration instead of manual configuration, and improving the accuracy of training model resource configuration, training efficiency and training ease of use.

[0068] Step S4: Create a distributed training task. That is, the management node 11 can use the resource configuration information determined in step S3a or S3b as the resource configuration information required for the distributed training model to create a training task. Figure 4 It is a flow chart that further illustrates the overall process of the resource configuration method of the distributed training model according to an embodiment of the present application. Figure 4 This can be understood as an application example of the resource allocation method of this application. Figure 4 As shown, the management node 11 may include multiple modules, which are used to manage various resources when the training task is issued, as follows:

[0069] The web portal module 111 can be understood as a centralized entry point that integrates various information, services, and applications. Specifically, the web portal module 111 may include a web user interface (webUI). The web portal module 111 displays the model configuration interface to the user 10 through the webUI, and the user 10 accesses various functions and services within the web portal module 111 through the webUI.

[0070] Specifically, the webUI may include one or more drop-down menus of preset models for the user 10 to select, and one or more training scale information for the user 10 to input or fill in.

[0071] The model management module 112 manages the entire lifecycle of AI models, including model creation, version control, storage, deployment, evaluation, and optimization. Through the model management module 112, users can conveniently manage AI models at different stages, versions, and types, track model performance, and continuously improve and optimize the models.

[0072] The data management module 113 is used to manage data during the AI development process. This includes data collection, cleaning, annotation, storage, retrieval, and sharing. This ensures data quality and consistency, providing high-quality datasets for model training. It also categorizes and manages data, allowing users to quickly access required data based on different tasks and needs.

[0073] Image management module 114 manages container images, which include all the files and configuration information required to run AI applications and their dependencies. Image management module 114 can create, store, distribute, and update images, ensuring that AI applications can be deployed quickly and accurately in different environments, achieving environmental consistency and repeatability.

[0074] Resource management module 115 is used to uniformly manage and allocate various resources within the AI development platform, such as computing resources (CPU, GPU, memory, etc.), storage resources, and network resources. It can rationally allocate resources based on different task requirements and priorities, improve resource utilization, avoid resource waste and conflicts, and ensure that each module and task can operate normally under limited resource conditions.

[0075] The training job module 116 is used to organize and manage AI model training tasks. This includes creating training jobs, configuring training parameters (such as learning rate and number of iterations), monitoring the training process (such as viewing training logs and performance metrics in real time), handling training task interruptions and resumptions, and coordinating dependencies between multiple training tasks to ensure efficient and stable training tasks.

[0076] The scheduler module 117 is used to dynamically schedule and dispatch tasks such as training jobs to appropriate computing nodes 12 based on factors such as system resource conditions and task priorities. It can monitor resource usage and task queue status, and intelligently determine when and where to start tasks to optimize the performance and efficiency of the entire system and ensure that tasks can be completed in a timely and efficient manner.

[0077] The resource configuration library module 118 may include at least one common model and its corresponding resource configuration information, and is used to assist the resource management module 115 in automatically configuring the resources required by the common model.

[0078] The resource calculation module 119 is used to automatically calculate the resources required for uncommon models and can be regarded as a supplement to the resource configuration library module 118. When the training model is not included in the resource configuration library module 118, in order to avoid problems caused by manual configuration, the resource calculation module 119 can automatically calculate the resource configuration information required for the training task based on preset calculation rules.

[0079] Here’s how:

[0080] In steps 1-4, user 10 selects information such as a model, image, dataset, and training hyperparameters. Specifically, web portal module 111 obtains the model list, dataset list, and image list from model management module 112, data management module 113, and image management module 114, respectively, and displays them to user 10 via the web UI. User 10 makes selections based on the displayed information and enters the training hyperparameters.

[0081] When issuing a task, the web portal module 111 can provide the user 10 with the option to manually input or fill in the resource configuration information (hereinafter referred to as manual configuration) or to automatically determine the resource configuration information (hereinafter referred to as system automatic recommendation). For details, see Figure 5 .

[0082] Figure 5 Schematic diagram of the web portal interface according to an embodiment of the present application. Figure 5 As shown, in an application embodiment of the present application, the manually configured webUI can be as follows Figure 5 As shown in (A), please refer to steps 5a-6a for specific steps; the system automatically recommends webUI as follows Figure 5 As shown in (B), please see steps 5b-6c for specific steps.

[0083] Step 5a: When the user 10 selects manual configuration, the web portal module 111 obtains a resource configuration information list (eg Figure 5 In (A)), the user 10 manually selects and / or fills in the form.

[0084] Step 6a: The user 10 creates a training task. The web portal module 111 sends a request to create the training task to the training job module 116 , instructing it to create the training task according to the resource configuration information filled in and / or selected by the user 10 .

[0085] Step 5b: When the user 10 selects the system automatic recommendation, the web portal module 111 obtains the preset configuration rules from the resource configuration library module 118 via the resource management module 115, and displays them to the user 10 after processing. Specifically, it may include a preset model list (e.g. Figure 5 The "Model Name" shown in (B) and its drop-down menu, which includes one or more common models).

[0086] Step 6b. If the training model is included in the preset model list in the above drop-down menu, the training model is considered to be a common model. The user 10 selects the model and creates a training task. The web portal module 111 obtains the resource configuration information corresponding to the model in the preset configuration rules from the resource configuration library module 118 via the resource management module 115. Then the webportal module 111 sends the request to create the training task to the training job module 116, notifying it to create a distributed training task based on the resource configuration information obtained from the resource configuration library module 118.

[0087] Step 5c: When the user 10 selects the system to automatically recommend, if there is no training model in the preset model list in the drop-down menu, the training model is considered to be an uncommon model, and the user 10 needs to fill in the training scale information displayed on the webUI (e.g. Figure 5 The number of model parameters and training tokens shown in (B)).

[0088] In step 6c, user 10 creates a training task. Web portal module 111 sends the training scale information entered by user 10 in step 5c to resource management module 115. Resource management module 115 then sends this training scale information to resource calculation module 119 for automatic calculation based on preset calculation rules. Resource calculation module 119 then feeds back its calculation results (i.e., recommended resource configuration information) to resource management module 115.

[0089] Optionally, the resource management module 115 may feed back the recommended resource configuration information to the web portal module 111 so as to be displayed to the user 10. Furthermore, the user 10 may adjust the recommended resource configuration information on the web UI.

[0090] Then, the web portal module 111 sends a request for creating a training task to the training operation module 116 , instructing it to create a training task according to the resource configuration information obtained from the resource calculation module 119 .

[0091] Step 7: The training job module 116 sends a request to create a distributed task to the scheduler module 117, which is responsible for the actual creation of the distributed task.

[0092] Figure 6 1 is a schematic diagram illustrating a resource configuration device for a distributed training model according to an embodiment of the present application. Figure 6 As shown, the resource configuration device 600 for the distributed training model may include at least the following modules.

[0093] Display module 601 is used to display the model configuration interface; the model configuration interface is used to obtain the training model selected by the user; the training model belongs to the preset model; the preset model is a trained model.

[0094] Determination module 602 is used to determine the resource configuration information corresponding to the training model based on preset configuration rules; the preset configuration rules include the preset model and the corresponding resource configuration information; wherein the resource configuration information includes: the number of computing nodes and the number of training cards; the training cards are used for training the model.

[0095] Specifically, the determining module 602 may further include:

[0096] The creation module 603 is used to create a distributed training task based on the resource configuration information; the distributed training task is used to perform training of the training model. In this case, the resource configuration information is used as the resource configuration information required by the distributed training model.

[0097] In addition, the resource configuration device 600 for the distributed training model may further include:

[0098] Automatic calculation unit 604 is configured to obtain training scale information input by the user through the model configuration interface if the training model does not belong to the preset model. The training scale information is used to determine resources for the training model. Resource configuration information is determined based on the training scale information and preset calculation rules. The training scale information includes the number of model parameters and the amount of training data.

[0099] Specifically, the automatic calculation unit 604 may further include:

[0100] Type subunit 6041 is configured to determine the type of the training card based on the model parameter values. The training card type includes a module form factor or a Peripheral Component Interconnect (PCI) form factor. If the model parameter values are greater than or equal to a threshold parameter value, the training card type is determined to be a module form factor. If the model parameter values are less than the threshold parameter value, the training card type is determined to be a PCI form factor.

[0101] The quantity subunit 6042 is used to determine the number of computing nodes and the number of training cards based on the type of training cards, the number of model parameters, the amount of data and the recommended training time.

[0102] The list module 605 is used to determine one or more preset models to be displayed in the model configuration interface based on preset configuration rules before displaying the model configuration interface.

[0103] Figure 7 is a hardware block diagram illustrating a computing device according to an embodiment of the present application. A computing device 700 according to an embodiment of the present application includes at least a processor 701 and a memory 702; the processor 701 may be a central processing unit (CPU) or other similar processing unit, which is coupled to the memory 702 (e.g., interconnected via a bus); the memory 702 may be used to store computer-readable instructions. When the computer-readable instructions are loaded and executed by the processor 701, the computing device 700 executes the resource configuration method for the distributed training model described above.

[0104] Figure 8 Schematic diagram of a computer program product according to an embodiment of the present application. Figure 8 As shown, a computer program product 800 according to an embodiment of the present application has a computer program 801 stored thereon. When the computer program 801 is executed by a processor, the resource configuration method of the distributed training model described with reference to the above figures is executed. The computer program product includes, but is not limited to, for example, volatile memory and / or non-volatile memory. The volatile memory may, for example, include random access memory (RAM) and / or cache memory (cache), etc. The non-volatile memory may, for example, include read-only memory (ROM), hard disk, flash memory, optical disk, magnetic disk, etc.

[0105] Above, with reference to the accompanying drawings, a resource configuration method, computing device and computing system for a distributed training model according to an embodiment of the present application are described. According to the resource configuration method for a distributed training model according to an embodiment of the present application, the resource configuration information required for common models is automatically determined through the built-in preset configuration rules of common models; and the required resource configuration information for uncommon models is automatically calculated through the built-in preset calculation rules. The combination of the two covers all training models, thereby fully realizing automated configuration instead of manual configuration, and improving the accuracy of training model resource configuration, training efficiency and training ease of use.

[0106] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0107] The basic principles of the present application have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, and effects mentioned in this application are merely illustrative and not restrictive, and it should not be assumed that these advantages, strengths, and effects are required of each embodiment of this application. In addition, the specific details disclosed above are merely illustrative and facilitating understanding, and are not restrictive. The above details do not limit this application to necessarily being implemented using the above specific details.

[0108] The block diagrams of the devices, devices, equipment, and systems involved in this application are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As will be appreciated by those skilled in the art, these devices, devices, equipment, and systems can be connected, arranged, or configured in any manner. Words such as "include," "comprise," "have," and the like are open-ended words, meaning "including but not limited to," and can be used interchangeably therewith. The words "or" and "and" used herein refer to the words "and / or" and can be used interchangeably therewith, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to," and can be used interchangeably therewith.

[0109] Additionally, as used herein, "or" used in a list of items beginning with "at least one" indicates a separate list, so that, for example, a list of "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). Furthermore, the word "exemplary" does not mean that the example described is preferred or better than other examples.

[0110] It should also be noted that in the system and method of the present application, each component or each step can be decomposed and / or recombined. Such decomposition and / or recombination should be regarded as equivalent solutions of the present application.

[0111] Various changes, substitutions, and modifications of the technology described herein may be made without departing from the teachings defined by the appended claims. Moreover, the scope of the claims herein is not limited to the specific aspects of the processes, machines, manufactures, compositions of things, means, methods, and actions described above. Currently existing or later developed processes, machines, manufactures, compositions of things, means, methods, or actions that perform substantially the same functions or achieve substantially the same results as the corresponding aspects described herein may be utilized. Accordingly, the appended claims include within their scope such processes, machines, manufactures, compositions of things, means, methods, or actions.

[0112] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present application. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the present application. Therefore, the present application is not intended to be limited to the aspects shown herein, but rather to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0113] The above description has been provided for the purpose of illustration and description. Furthermore, this description is not intended to limit the embodiments of the present application to the forms disclosed herein. Although a number of example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.

Claims

1. A resource configuration method for a distributed training model, characterized in that: The method comprises: Display the model configuration interface; the model configuration interface is used to obtain the training model selected by the user; the training model belongs to the preset model; the preset model is a trained model; Determine resource configuration information corresponding to the training model based on preset configuration rules; the preset configuration rules include a preset model and corresponding resource configuration information; wherein the resource configuration information includes: the number of computing nodes and the number of training cards; the training cards are used to train the model; A distributed training task is created based on the resource configuration information; the distributed training task is used to execute the training model training.

2. The resource allocation method according to claim 1, wherein: Before displaying the model configuration interface, the method includes: Based on the preset configuration rule, one or more preset models are determined to be displayed in the model configuration interface.

3. The resource allocation method according to claim 2, wherein: The method further comprises: In the case where the training model does not belong to the preset model, obtaining training scale information input by the user through the model configuration interface; the training scale information is used to determine the resources described in the training model; The resource configuration information is determined according to the training scale information and preset calculation rules.

4. The resource allocation method according to claim 3, wherein: The training scale information includes the amount of model parameters and the amount of training data.

5. The resource allocation method according to claim 4, characterized in that: The method further comprises: Based on the model parameter quantity, the type of the training card is determined; the type of the training card includes a module form or a peripheral component interconnect (PCI) form.

6. The resource allocation method according to claim 5, characterized in that: The determining the type of the training card based on the model parameter amount includes: When the model parameter amount is greater than or equal to the threshold parameter amount, determining that the type of the training card is a module form; When the model parameter amount is less than the threshold parameter amount, it is determined that the type of the training card is the PCI form.

7. The resource allocation method according to claim 5 or 6, characterized in that: The method further comprises: The number of computing nodes and the number of training cards are determined according to the type of training card, the amount of model parameters, the amount of training data and the recommended training duration.

8. The resource allocation method according to any one of claims 1 to 7, characterized in that: The creating a distributed training task based on the resource configuration information includes: The resource configuration information is used as the resource configuration information required for the distributed training model.

9. A computing device, characterized in that The computing device includes a memory and a processor; the memory and the processor are coupled; The memory is used to store computer-readable instructions; The processor is configured to execute the computer-readable instructions so that the computing device executes the resource configuration method according to any one of claims 1 to 8.

10. A computing system, characterized in that: The computing system includes a management node and a computing node; wherein the management node is a computing device as described in claim 9; the management node and the computing node are connected; and the computing node is used for distributed training of the training model.