Model training method, device and computer readable storage medium

Through the management node, the model training task is directly allocated to the computing node with free resources and training data, and the training efficiency problem caused by resource and data redistribution in the existing technology is solved, and more efficient intelligent model training is achieved.

CN113849295BActive Publication Date: 2025-05-06HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010600109.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-06-28
Publication Date
2025-05-06
Estimated Expiration
2040-06-28

AI Technical Summary

Technical Problem

In the process of intelligent model training, when a smart model is configured, resources need to be reassigned and training data are obtained each time an intelligent model is configured, resulting in increased time-consuming and reduced model training efficiency.

Method used

By scheduling model training tasks by managing nodes, identifying computing nodes with free resources and training data, and directly sending training requests to avoid reallocation of resources and data.

Benefits of technology

Save time in resource allocation and data acquisition, and improve the efficiency of intelligent model training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113849295B_ABST
    Figure CN113849295B_ABST
Patent Text Reader

Abstract

The present application discloses a method, device and computer-readable storage medium for model training, belonging to the field of communications. The method includes: a management node schedules a first model training task, the first model training task includes a job identifier of a first intelligent model and a first parameter adjustment job, the first intelligent model is obtained by configuring the algorithm corresponding to the first parameter adjustment job based on a first parameter value set; a first computing node is determined according to the job identifier, the first computing node has at least one of the first training data and the idle first resource, the first resource is a resource required for processing the first parameter adjustment job, and the first training data is the training data required for training the intelligent model of the first parameter adjustment job; a first training request is sent to the first computing node, and the first training request is used for the first computing node to train the first intelligent model according to at least one of the first resource and the first training data. The present application can improve the efficiency of model training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of communications, and in particular to a method, device and computer-readable storage medium for model training. Background Art

[0002] By training intelligent algorithms such as deep learning, we can obtain intelligent models with specific functions, such as image recognition, speech recognition and synthesis, or natural language processing. Training intelligent algorithms means constantly adjusting the values ​​of super parameters and common parameters of intelligent algorithms to make them intelligent models with specific functions. Super parameters are used to define the structure and training process of intelligent models, while common parameters are used to define the functions implemented by intelligent models.

[0003] Currently, computing clusters can be used to train intelligent algorithms, and cloud storage systems can be used to store training samples required for training intelligent algorithms. When training an intelligent model, the user configures the intelligent algorithm and at least one super parameter in the computing cluster. The computing cluster initializes the initial value of each super parameter, configures the intelligent algorithm according to the initial value of each super parameter, and obtains a first intelligent model. Resources are allocated to the first intelligent model, and training data is retrieved from the cloud storage system, and the first intelligent model is trained using the training data and the allocated resources. The computing cluster continuously adjusts the values ​​of the common parameters of the first intelligent model during the training of the first intelligent model until the first intelligent model converges or fails to converge successfully, or the training is stopped when the number of times the first intelligent model is trained reaches a specified number of times.

[0004] When the training is stopped, the computing cluster obtains the training results of the first intelligent model. If the training results do not meet the specified conditions, the new value of each super parameter is configured according to the current value of each super parameter and the training results, and the intelligent algorithm is configured according to the new value of each super parameter to obtain the second intelligent model. Resources are allocated to the second intelligent model, and training data is retrieved from the cloud storage system. The training data and the allocated resources are used to train the second intelligent model. The process of training the second intelligent model is also to continuously adjust the values ​​of the common parameters of the second intelligent model until the second intelligent model converges or fails to converge successfully, and then the training is stopped, or the number of times the second intelligent model is trained reaches the specified number of times.

[0005] When the training of the second intelligent model stops, the computing cluster still obtains the training results of the second intelligent model. If the training results of the second intelligent model do not meet the specified conditions, the above process of obtaining the second intelligent model and training the second intelligent model is repeated. If the training results of the second intelligent model meet the specified conditions, the second intelligent model is the final trained model with specific functions.

[0006] In the process of implementing this application, the inventor found that the prior art has at least the following problems:

[0007] In the above process, each time an intelligent model is configured, it is necessary to reallocate resources for the intelligent model and retrieve training data from the cloud storage system, which increases time consumption and reduces the efficiency of model training. Summary of the invention

[0008] The present application provides a method, device and computer-readable storage medium for model training to improve the efficiency of model training. The technical solution is as follows:

[0009] In the first aspect, the present application provides a method for model training, in which a management node schedules a first model training task, the first model training task includes a job identifier of a first intelligent model and a first parameter adjustment job, the first intelligent model is obtained by configuring the algorithm corresponding to the first parameter adjustment job based on a first parameter value set, and the first parameter value set includes the first parameter value of each super parameter in at least one super parameter corresponding to the first parameter adjustment job. The management node determines a first computing node from a node cluster according to the job identifier, the first computing node has at least one of first training data and idle first resources, the first resource is a resource required for processing the model training task of the first parameter adjustment job, and the first training data is training data required for training the intelligent model corresponding to the first parameter adjustment job. The management node sends a first training request to the first computing node, the first training request includes a first model training task, and the first training request is used for the first computing node to train the first intelligent model according to at least one of the first resource and the first training data.

[0010] Among them, the first computing node determined by the management node has at least one of the first training data and the idle first resource. In this way, after the first computing node receives the first training request including the first model training task, it is not necessary to allocate the first resource for the first model training task and / or obtain the first training data, thereby saving time for allocating the first resource and / or obtaining the first training data, and improving the efficiency of training the first intelligent model.

[0011] In a possible implementation, the management node determines the first computing node from the node cluster according to the resource correspondence, the data correspondence and the job identifier. Among them, any record in the resource correspondence includes the job identifier of the parameter adjustment job, the node identifier of the computing node in the node cluster, the resource identifier and the resource status, and the resource identifier is used to identify the resources required for the model training task of the parameter adjustment job included in the computing node, and the resource status is used to describe whether the resource is currently idle. Any record in the data correspondence includes the job identifier of the parameter adjustment job, the node identifier of the computing node in the node cluster and the data identifier, and the data identifier is used to identify the training data required for the intelligent model corresponding to the parameter adjustment job included in the computing node. Since the context information of training the first parameter adjustment job can be recorded through the resource correspondence and the data correspondence, when the computing node is assigned to the first model training task of the first parameter adjustment job, the computing node including the first resource and / or the first training data can be accurately determined.

[0012] In another possible implementation, the management node determines N computing nodes in the node cluster that include the first training data and / or the first resource based on the resource correspondence, data correspondence and job identifier, where N is an integer greater than 0. When there is at least one target node among the N computing nodes, the management node selects a target node from the at least one target node as the first computing node. Among them, since the target node includes the idle first resource, or the target node includes the first training data and the idle first resource, or the target node includes the first training data and the size of the unprotected resources included in the target node exceeds the size of the resources required to process the first model training task, the unprotected resources are other resources in the target node except the protected resources, and the protected resources are resources that have been allocated to the parameter adjustment job and the protection time period corresponding to the protected resources has not yet ended. Therefore, selecting the first computing node from the target node can ensure that the first computing node has the first training data and sufficient resources to train the first intelligent model, thereby improving the training success rate.

[0013] In another possible implementation, the management node determines at least one target node according to the job identifier, and selects one target node from each target node as the first computing node according to the load information and / or node attribute information of each target node in the at least one target node. Wherein, the target node includes the first idle resource, or the target node includes the first training data and the idle first resource, or the target node includes the first training data and the size of the unprotected resources included in the target node exceeds the size of the resources required to process the first model training task, the unprotected resources are other resources in the target node except the protected resources, and the protected resources are resources that have been allocated to the parameter adjustment job and the protection time period corresponding to the protected resources has not yet ended. Due to the load information and / or node attribute information of each target node, one target node is selected from each target node as the first computing node, which can meet one or more requirements. For example, selecting the first computing node according to the load information of each computing node can meet the requirements of load balancing or energy saving.

[0014] In another possible implementation, when there is no target node among the N computing nodes, the management node detects whether any computing node among the N computing nodes becomes a target node within a first time period, the start time of the first time period is the time for scheduling the first model training task, the time length of the first time period is a first threshold, and the N computing nodes are computing nodes that include the first training data and / or the first resource. The management node detects that a computing node becomes a target node within the first time period, and determines the detected target node as the first computing node. Among them, the target node includes the idle first resource, or the target node includes the first training data and the idle first resource, or the target node includes the first training data and the size of the unprotected resources included in the target node exceeds the size of the resources required to process the first model training task, the unprotected resources are other resources in the target node except the protected resources, and the protected resources are resources that have been allocated to the parameter adjustment job and the protection time period corresponding to the protected resources has not yet ended.

[0015] When there is no target node among the N computing nodes, a computing node is not immediately allocated to the first model training task. Instead, the process waits to see whether a computing node becomes a target node within a first time period. If so, the target node is allocated to the first model training task. The first time period is often short, so when the target node processes the first model training task, it does not need to allocate the first resource and / or obtain the first training data, thereby improving the efficiency of model training.

[0016] In another possible implementation, any record in the resource correspondence also includes the resource size of the resource identified by the resource identifier. The management node detects that no computing node becomes a target node within the first time period, and after the first time period ends, determines a second computing node from the node cluster based on the resource correspondence, and the size of unprotected resources included in the second computing node is greater than the size of resources required to process the first model training task. The management node sends a second training request to the second computing node, the second training request includes the first model training task, and the second training request is used for the second computing node to train the first intelligent model.

[0017] In another possible implementation, the management node receives a first deletion request, the first deletion request includes a node identifier of the computing node and a resource identifier of the first resource, the first deletion request is sent by the first computing node after the first protection time period ends, the start time of the first protection time period is the time when the first resource was last used, and the time length of the first protection time period is the second threshold. The management node deletes the record including the node identifier of the first computing node and the resource identifier of the first resource from the resource correspondence. In this way, when the first computing node releases the first resource, the resource correspondence is updated in time, thereby ensuring the accuracy of the content stored in the resource correspondence.

[0018] In another possible implementation, the management node receives a second deletion request, the second deletion request includes the node identifier of the first computing node and the data identifier of the first training data, the second deletion request is sent by the first computing node after the second protection time period ends, the start time of the second protection time period is the time when the first training data was last used, and the time length of the second protection time period is a third threshold. The management node deletes the record including the node identifier of the first computing node and the data identifier of the first training data from the data correspondence relationship. In this way, when the first computing node deletes the first training data, the data correspondence relationship is updated in time, thereby ensuring the accuracy of the content stored in the data correspondence relationship.

[0019] In another possible implementation, the management node sends a third training request to the first computing node, the third training request includes a second model training task, the second model training task includes a job identifier of a second intelligent model and a first parameter adjustment job, the second model training task is a model training task included in the first batch of tasks corresponding to the first parameter adjustment job, the second intelligent model is obtained by configuring the algorithm based on a second parameter value set, the second parameter value set includes a second parameter value of each super parameter, and the third training request is used by the first computing node to allocate a first resource for training the second intelligent model and obtain first training data for training the second intelligent model. The management node receives a storage request sent by the first computing node, the storage request includes a data identifier of the first training data, a resource identifier of the first resource, and a resource status. The management node saves the corresponding relationship between the job identifier, the node identifier of the first computing node, the resource identifier of the first resource, and the resource status in the resource corresponding relationship; and saves the corresponding relationship between the job identifier, the node identifier of the first computing node, and the data identifier of the first training data in the data corresponding relationship. In this way, the context information of the first parameter adjustment job is saved to ensure that when the i-th batch of tasks of the first parameter adjustment job is trained, i=2, 3, ..., the i-th batch of tasks can be allocated to the computing nodes that have the resources and / or training data required to process the first parameter adjustment job.

[0020] In a second aspect, the present application provides a method for model training, in which a computing node receives a first training request sent by a management node, the first training request includes a first model training task, the first model training task includes a job identifier of a first intelligent model and a first parameter adjustment job, the first intelligent model is obtained by configuring an algorithm corresponding to the first parameter adjustment job based on a first parameter value set, the first parameter value set includes a first parameter value of each super parameter in at least one super parameter corresponding to the first parameter adjustment job, and the computing node has at least one of a first resource and a first training data bound to the first parameter adjustment job. The computing node obtains at least one of the first resource and the first training data according to the job identifier; and trains the first intelligent model according to at least one of the first resource and the first training data.

[0021] Among them, since the computing node has the first resources and / or first training data required for processing the first model training task, when the computing node receives the first model training task, it can save time in allocating the first resources and / or time in acquiring the first training data, thereby improving the efficiency of training the intelligent model.

[0022] In a possible implementation, the computing node receives a third training request, the third training request includes a second model training task, the second model training task includes a job identifier of a second intelligent model and a first parameter adjustment job, the second model training task is a model training task included in the first batch of tasks corresponding to the first parameter adjustment job, the second intelligent model is obtained by configuring the algorithm based on a second parameter value set, and the second parameter value set includes a second parameter value of each super parameter. The computing node allocates a first resource for training the second intelligent model from an unprotected resource, and obtains first training data for training the second intelligent model, the unprotected resource is other resources in the computing node except the protected resource, the protected resource is a resource that has been allocated to the parameter adjustment job and the protection time period corresponding to the protected resource has not yet ended. The computing node trains the second intelligent model according to the first resource and the first training data. Since the computing node allocates the first resource for training the second intelligent model from the unprotected resource, it is ensured that the protected resource is not occupied, and the protected resource is a resource used to train other model training tasks, which ensures that when the computing node receives the other model training task, it does not need to allocate resources for the other model training task, thereby improving the efficiency of processing other model training tasks.

[0023] In another possible implementation, the computing node sends a storage request, the storage request includes the data identifier of the first training data, the resource identifier of the first resource, and the resource state, and the storage request is used for the management node to save the corresponding relationship between the job identifier, the node identifier of the first computing node, the resource identifier of the first resource, and the resource state in the resource corresponding relationship, and save the corresponding relationship between the job identifier, the node identifier of the first computing node, and the data identifier of the first training data in the data corresponding relationship. In this way, it can be ensured that the management node saves the context information of the training first parameter adjustment parameter.

[0024] In another possible implementation, the computing node sends a first deletion request after the first protection time period ends, the first deletion request includes the node identifier of the first computing node and the resource identifier of the first resource, the start time of the first protection time period is the time when the computing node last used the first resource, the time length of the first protection time period is the second threshold, and the first deletion request is used by the management node to delete the record including the node identifier of the computing node and the resource identifier of the first resource from the resource correspondence relationship. In this way, when the computing node deletes the first training data, the data correspondence relationship in the management node can be updated in time, thereby ensuring the accuracy of the content stored in the data correspondence relationship.

[0025] In another possible implementation, the computing node sends a second deletion request after the second protection time period ends, the second deletion request includes the node identifier of the computing node and the data identifier of the first training data, the start time of the second protection time period is the time when the computing node last used the first training data, the time length of the second protection time period is the third threshold, and the second deletion request is used by the management node to delete the record including the node identifier of the computing node and the data identifier of the first training data from the data correspondence relationship. In this way, when the computing node releases the first resource, the resource correspondence relationship in the management node is updated in time, thereby ensuring the accuracy of the content stored in the resource correspondence relationship.

[0026] In a third aspect, the present application provides a model training device for executing the method in the first aspect or any possible implementation of the first aspect. Specifically, the device includes a unit for executing the method in the first aspect or any possible implementation of the first aspect.

[0027] In a fourth aspect, the present application provides a model training device for executing the method in the second aspect or any possible implementation of the second aspect. Specifically, the device includes a unit for executing the method in the second aspect or any possible implementation of the second aspect.

[0028] In a fifth aspect, the present application provides a device for model training, the device comprising: a processor, a memory, and a network interface. The processor, the memory, and the network interface may be connected via a bus system. The memory is used to store one or more programs, and the processor is used to execute one or more programs in the memory, so that the device completes the method in the first aspect or any possible implementation of the first aspect.

[0029] In a sixth aspect, the present application provides a device for model training, the device comprising: a processor, a memory, and a network interface. The processor, the memory, and the network interface may be connected via a bus system. The memory is used to store one or more programs, and the processor is used to execute one or more programs in the memory, so that the device completes the method in the second aspect or any possible implementation of the second aspect.

[0030] In the seventh aspect, the present application provides a computer-readable storage medium, which stores program code. When the computer-readable storage medium is run on a computer, it enables the computer to execute the method in the above-mentioned first aspect, second aspect, any possible implementation of the first aspect, or any possible implementation of the second aspect.

[0031] In an eighth aspect, the present application provides a computer program product comprising program code, which, when executed on a computer, enables the computer to execute the method in the above-mentioned first aspect, second aspect, any possible implementation of the first aspect, or any possible implementation of the second aspect.

[0032] In the ninth aspect, the present application provides a computer-readable storage medium, characterized in that the computer-readable storage medium stores a program, and the program is used to implement the method in the above-mentioned first aspect, second aspect, any possible implementation of the first aspect, or any possible implementation of the second aspect.

[0033] In the tenth aspect, the present application provides a system for model training, the system comprising the device described in the third aspect and the device described in the fourth aspect, or comprising the device described in the fifth aspect and the device described in the sixth aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 is a schematic diagram of a network architecture provided in an embodiment of the present application;

[0035] Figure 2 It is a flow chart of a model training method provided in an embodiment of the present application;

[0036] Figure 3 is a flow chart of another model training method provided in an embodiment of the present application;

[0037] Figure 4 It is a schematic diagram of the structure of a model training device provided in an embodiment of the present application;

[0038] Figure 5 It is a schematic diagram of the structure of another device for model training provided in an embodiment of the present application;

[0039] Figure 6 It is a schematic diagram of the structure of another device for model training provided in an embodiment of the present application;

[0040] Figure 7 It is a schematic diagram of the structure of another device for model training provided in an embodiment of the present application;

[0041] Figure 8 It is a schematic diagram of the system structure of a model training provided in an embodiment of the present application. DETAILED DESCRIPTION

[0042] See also Figure 1The embodiment of the present application provides a system for model training, which includes a management node, a node cluster and a storage system. The node cluster includes at least one computing node, each computing node includes resources for training an intelligent model, and the resource may be one or more of a central processing unit (CPU), a graphics processing unit (GPU), and a memory. The storage system is used to store a training data set required for training an intelligent model.

[0043] Optionally, the training data set includes multiple training samples.

[0044] Optionally, a network connection is established between the management node and each computing node in the node cluster, and a network connection is established between each computing node in the node cluster and the storage system.

[0045] Optionally, when the user needs the system to train an intelligent model, a parameter adjustment job can be submitted to the management node. For the sake of convenience, the parameter adjustment job is called the first parameter adjustment job, and the first parameter adjustment job includes at least one super parameter, an algorithm, a resource name and resource size required for the first parameter adjustment job, and a storage location of a training data set required for the first parameter adjustment job, and the training data set includes multiple training samples.

[0046] The management node receives the first parameter adjustment job, initializes the parameter value of each super parameter in the at least one super parameter, configures the algorithm according to the parameter value of each super parameter to obtain an intelligent model, and divides the training data set into multiple training data, each of which includes at least one training sample. Generate the first batch of tasks corresponding to the first parameter adjustment job, the first batch of tasks includes at least one model training task, and for any model training task, the model training task includes the intelligent model, the job identifier of the first parameter adjustment job, the resource name and resource size, the storage location of the training data set, the offset of a piece of training data in the training data set, and the size of the training data. Assign a computing node to each model training task included in the first batch of tasks, and send each model training task to the computing node corresponding to each model training task.

[0047] Optionally, the size of the training data may be the number of training samples included in the training data.

[0048] Optionally, the parameter value of each super parameter in the at least one super parameter is used to define the structure and training process of the intelligent model, etc.

[0049] For any computing node, the computing node receives a model training task, allocates resources required for the first parameter adjustment job according to the resource name and resource size included in the model training task, and obtains the training data from the storage system according to the storage location of the training data set included in the model training task and the offset and size of a training data, that is, obtains the training data required for the first parameter adjustment job. According to the training data, the intelligent model included in the model training task is trained through the allocated resources.

[0050] The computing node also sends the job identifier of the first parameter adjustment job, the data identifier of the training data, the resource identifier of the resource, and the resource status to the management node, wherein the resource status is a usage status. The data identifier is used to identify the training data in the computing node, and the resource identifier is used to identify the resource in the computing node.

[0051] It should be noted that: after the computing node allocates the resources required for the first parameter adjustment job, it determines the protection time period of the resources, and will not release the resources during the protection time period, nor will it allocate the resources to model training tasks included in other parameter adjustment jobs other than the first parameter adjustment job. Also, after the computing node obtains the training data required for the first parameter adjustment job, it determines the protection time period of the training data, and will not delete the training data during the protection time period.

[0052] The management node receives the job identifier of the first parameter adjustment job, the data identifier of the training data, the resource identifier and the resource status of the resource, and saves the corresponding relationship between the job identifier of the first parameter adjustment job, the node identifier of the computing node, the resource identifier and the resource status of the resource in the resource correspondence relationship, and saves the corresponding relationship between the job identifier of the first parameter adjustment job, the node identifier of the computing node and the data identifier of the training data in the data correspondence relationship.

[0053] Optionally, after training the intelligent model, the computing node sends the training result of the intelligent model, the resource identifier and the resource status of the resource to the management node, and the resource status is an idle state. The management node receives the training result, the resource identifier and the resource status of the resource, obtains the resource status of the resource from the resource correspondence according to the node identifier of the computing node and the resource identifier of the resource, and updates the resource status of the resource to an idle state.

[0054] The management node can receive the training results sent by at least one computing node assigned to the model training task, that is, receive at least one training result. If the at least one training result does not meet the specified conditions, the new parameter value of each super parameter is configured according to the current value of each super parameter and the at least one training result and other information, and the algorithm is configured according to the new parameter value of each super parameter to obtain a new intelligent model. Generate a second batch of tasks corresponding to the first parameter adjustment job, the second batch of tasks includes at least one model training task, and for any model training task, the model training task includes the new intelligent model, the job identifier of the first parameter adjustment job, the resource name and resource size, the storage location of the training data set, the offset of a piece of training data in the training data set, and the size of the training data.

[0055] The management node allocates a computing node with resources and / or training data required for the first parameter adjustment operation to each model training task included in the second batch of tasks according to the resource correspondence and data correspondence, and the resource state of the resource is idle, and sends a model training task to the computing node. Since the computing node has the resources and / or training data required for the first parameter adjustment operation, after receiving the model training task, the computing node can train the intelligent model included in the model training task according to the resources and / or training data required for the first parameter adjustment operation, thereby improving the efficiency of training the model.

[0056] When the management node generates the i-th batch of tasks for the first parameter adjustment job, i=3, 4, ..., the management node allocates computing nodes to each model training task included in the i-th batch of tasks according to the processing method of the second batch of tasks. The detailed implementation process will be described later. Figure 3 The illustrated embodiment is described in detail.

[0057] For the sake of convenience, each model training task in the i-th batch of tasks is called a first model training task, and the intelligent model in the first model training task is called a first intelligent model. Also, each model training task in the first batch of tasks is called a second model training task, and the intelligent model in the second model training task is called a second intelligent model.

[0058] See also Figure 2 , the embodiment of the present application provides a method for model training, the intelligent model trained by the method is the intelligent model included in each model training task in the first batch of tasks corresponding to the parameter adjustment job. The method can be applied to Figure 1 The system shown, the method includes:

[0059] Step 201: The management node receives a first parameter adjustment job, which includes at least one super parameter, an algorithm, a resource name and resource size required for the first parameter adjustment job, and a storage location of a training data set required for the first parameter adjustment job.

[0060] When a user needs to train an intelligent model, he can configure a first parameter adjustment job in his corresponding terminal and send the first parameter adjustment job to the management node.

[0061] Optionally, the algorithm may be a machine learning algorithm, for example, a neural network algorithm.

[0062] The resource name and resource size required for the first parameter adjustment job may be the resource name and resource size required for processing a model training task.

[0063] The training data set required for the first parameter adjustment operation includes a plurality of training samples. The training data set required for the first parameter adjustment operation can be stored in a storage system.

[0064] Step 202: The management node generates a first batch of tasks corresponding to the first parameter adjustment job, the first batch of tasks including at least one second model training task. For any second model training task in the first batch of tasks, the second model training task includes the second intelligent model, the job identifier of the first parameter adjustment job, the resource name and resource size, the storage location of the training data set, the offset of a piece of training data in the training data set, and the size of the training data.

[0065] In this step, for each super parameter of the at least one super parameter, the management node configures M parameter values ​​of each super parameter, where M is an integer greater than 0, to obtain M second parameter value sets. The algorithm is configured according to each second parameter value set to obtain M second intelligent models, and the training data set is divided into multiple training data, each of which includes at least one training sample.

[0066] Generate M model training jobs corresponding to the first parameter adjustment job, each model training job may include Y second model training tasks, where Y is an integer greater than 1. The second model training tasks included in the M model training jobs constitute the first batch of tasks for the first parameter adjustment business, that is, the first batch of tasks may include M*Y second model training tasks, where * is a multiplication operation. For any second model training task, the second model training task includes a second intelligent model, the job identifier of the first parameter adjustment job, the resource name and resource size, the storage location of the training data set, the offset of a piece of training data in the training data set, and the size of the piece of training data.

[0067] Optionally, for any second intelligent model among the M second intelligent models, the management node may generate Y second model training tasks for the second intelligent model to obtain a model training job, and the Y second model training tasks include the second intelligent model. Therefore, the management node generates Y second model training tasks for each second intelligent model to obtain the second model training tasks included in the M model training jobs, and the second model training tasks included in the M model training jobs are the M*Y second model training tasks included in the first batch of tasks.

[0068] Optionally, the management node divides the training data set into multiple pieces of training data. For any piece of training data, the size of the piece of training data is the number of training samples included in the piece of training data.

[0069] Optionally, the number of training samples included in each set of training data may be equal or unequal.

[0070] Optionally, the management node may save the M model training jobs in a scheduling queue.

[0071] Step 203: The management node allocates a computing node to each second model training task included in the first batch of tasks in the node cluster, and sends a training request to the computing node corresponding to any second model training task, where the training request includes the second model training task.

[0072] In this step, the management node can obtain the resource size of the unprotected resources included in each computing node in the computing cluster. Among them, the unprotected resources in the computing node are other resources in the computing node except the protected resources, and the protected resources in the computing node are resources that the computing node has allocated to the parameter adjustment job and the protection time period corresponding to the protected resources has not yet ended. For any second model training task included in the first batch of tasks, the management node selects a computing node from the node cluster whose resource size of unprotected resources is greater than or equal to the resource size included in the second model task, and sends a training request to the selected computing node, and the training request includes the second model training task.

[0073] Optionally, the management node may query each computing node in the node cluster for the resource size of the unprotected resources included in each computing node.

[0074] Optionally, for the implementation of this step, an example is listed below for detailed description, and the example is:

[0075] The management node schedules a model training job from the scheduling queue, and schedules a second model training task from the second model training tasks included in the model training job. The management node selects a computing node from the node cluster whose resource size of unprotected resources is greater than or equal to the resource size included in the scheduled second model task, and sends a training request to the selected computing node, where the training request includes the scheduled second model training task. The management node continues to schedule other second model training tasks included in the model training job until all other second model training tasks included in the model training job are scheduled.

[0076] Then, the management node schedules another model training job from the scheduling queue, and schedules the second model training task included in the other model training job in the above manner. The management node repeats the above operation until all model training tasks included in the model training jobs corresponding to the first parameter adjustment job are scheduled.

[0077] Step 204: The computing node receives the training request, which includes a second model training task, and obtains resources and training data required for processing the second model training task.

[0078] In this step, the computing node receives the training request, which includes a second model training task, and the second model training task includes a second intelligent model, a job identifier of the first parameter adjustment job, the resource name and resource size, the storage location of the training data set, the offset of a piece of training data in the training data set, and the size of the training data.

[0079] The computing node allocates resources required to process the second model training task according to the resource name and resource size included in the second model training task, and obtains the training data set from the storage system according to the storage location of the training data set included in the second model training task; and obtains a piece of training data required to process the second model training task from the training data set according to the offset of a piece of training data corresponding to the second model training task in the training data set and the size of the training data.

[0080] Optionally, the computing node further allocates a resource identifier for the resource required to process the second model training task, and the resource identifier identifies the resource in the computing node. The job identifier of the first parameter adjustment job and the resource identifier are correspondingly stored in the corresponding relationship between the job identifier and the resource identifier. And,

[0081] Optionally, the computing node further allocates a data identifier for the training data required for processing the second model training task, and the data identifier identifies the training data in the computing node. The job identifier of the first parameter adjustment job and the data identifier are correspondingly stored in the corresponding relationship between the job identifier and the data identifier.

[0082] Optionally, the computing node further allocates a first protection time period for the resource, the start time of the first protection time period is the time when the resource is used, and the time length of the first protection time period is the second threshold. Since after the resource is allocated, the operation of step 205 below will be performed to use the resource, the start time of the first protection time period here is equal to the time when the resource is allocated.

[0083] Optionally, the computing node further allocates a second protection time period for the training data, the start time of the second protection time period is the time when the training data is used, and the time length of the second protection time period is a third threshold. Since after the training data is obtained, the following step 205 will be performed to use the training data, the start time of the second protection time period here is equal to the time when the training data is obtained.

[0084] Optionally, the computing node also sends a storage request to the management node, where the storage request includes a job identifier of the first parameter adjustment job, a data identifier of the training data, a resource identifier and a resource status of the resource, where the resource status is a usage status.

[0085] The management node receives the storage request, combines the job identifier of the first parameter adjustment job, the node identifier of the computing node, the resource identifier and the resource status into a record and saves the record in the resource correspondence relationship; and combines the job identifier of the first parameter adjustment job, the node identifier of the computing node and the data identifier of the training data into a record and saves the record in the data correspondence relationship.

[0086] Optionally, the storage request may also include the task identifier of the second model training task. Correspondingly, the record saved in the resource correspondence relationship also includes the task identifier, and the record saved in the data correspondence relationship also includes the task identifier.

[0087] For example, assuming that the first batch of tasks includes second model training tasks 1, 2, 3, and 4, the management node allocates computing nodes 1, 2, 3, and 4 to the second model training tasks 1, 2, 3, and 4, respectively. The management node sends training request 1 to computing node 1, which includes second model training task 1; sends training request 2 to computing node 2, which includes second model training task 2; sends training request 3 to computing node 3, which includes second model training task 3; sends training request 4 to computing node 4, which includes second model training task 4.

[0088] Computing node 1 receives training request 1 including second model training task 1, allocates resources according to the resource name and resource size included in the second model training task 1, and obtains a copy of training data according to the storage location of the training data set included in the second model training task 1 and the offset and size of a copy of training data corresponding to the second model training task 1; sends storage request 1 to the management node, and the storage request 1 includes the job identifier IZ1 of the first parameter adjustment job, the data identifier ID1 of the copy of training data, the resource identifier IR1 of the resource, and the resource status is a usage status.

[0089] Computing node 2 receives training request 2 including second model training task 2, allocates resources according to the resource name and resource size included in the second model training task 2, and obtains a copy of training data according to the storage location of the training data set included in the second model training task 2 and the offset and size of a copy of training data corresponding to the second model training task 2; sends storage request 2 to the management node, and the storage request 2 includes the job identifier IZ1 of the first parameter adjustment job, the data identifier ID2 of the copy of training data, the resource identifier IR2 of the resource, and the resource status is a usage status.

[0090] The computing node 3 receives a training request 3 including a second model training task 3, allocates resources according to the resource name and resource size included in the second model training task 3, and obtains a copy of training data according to the storage location of the training data set included in the second model training task 3 and the offset and size of a copy of training data corresponding to the second model training task 3; and sends a storage request 3 to the management node, which includes the job identifier IZ1 of the first parameter adjustment job, the data identifier ID3 of the copy of training data, the resource identifier IR3 of the resource, and the resource status, which is a usage status.

[0091] The computing node 4 receives a training request 4 including a second model training task 4, allocates resources according to the resource name and resource size included in the second model training task 4, and obtains a copy of training data according to the storage location of the training data set included in the second model training task 4 and the offset and size of a copy of training data corresponding to the second model training task 4; and sends a storage request 4 to the management node, the storage request 4 including the job identifier IZ1 of the first parameter adjustment job, the data identifier ID4 of the copy of training data, the resource identifier IR4 of the resource and the resource status, and the resource status is a usage status.

[0092] The management node receives storage request 1, and combines the node identifier IN1 of computing node 1, the job identifier IZ1 of the first parameter adjustment job included in the storage request 1, the resource identifier IR1, and the resource status into a record and saves it in the resource correspondence shown in Table 1 below. The management node receives storage request 2, and combines the node identifier IN2 of computing node 2, the job identifier IZ1 of the first parameter adjustment job included in the storage request 2, the resource identifier IR2, and the resource status into a record and saves it in the resource correspondence shown in Table 1 below. The management node receives storage request 3, and combines the node identifier IN3 of computing node 3, the job identifier IZ1 of the first parameter adjustment job included in the storage request 3, the resource identifier IR3, and the resource status into a record and saves it in the resource correspondence shown in Table 1 below. And, the management node receives storage request 4, and combines the node identifier IN4 of computing node 4, the job identifier IZ1 of the first parameter adjustment job included in the storage request 4, the resource identifier IR4, and the resource status into a record and saves it in the resource correspondence shown in Table 1 below.

[0093] Table 1

[0094] Job ID Node ID Resource Identifier Resource Status IIZ1 IN1 IR1 Use Status IZ1 IN2 IR2 Use Status IZ1 IN3 IR3 Use Status IZ1 IN4 IR4 Use Status

[0095] The management node also forms a record with the node identifier IN1 of computing node 1, the job identifier IZ1 of the first parameter adjustment job included in the storage request 1, and the data identifier ID1, and saves it in the data correspondence shown in Table 2 below. The node identifier IN2 of computing node 2, the job identifier IZ1 of the first parameter adjustment job included in the storage request 2, and the data identifier ID2, and saves it in the data correspondence shown in Table 2 below. The node identifier IN3 of computing node 3, the job identifier IZ1 of the first parameter adjustment job included in the storage request 3, and the data identifier ID3, and save it in the data correspondence shown in Table 2 below. And, the node identifier IN4 of computing node 4, the job identifier IZ1 of the first parameter adjustment job included in the storage request 4, and the data identifier ID4, and save it in the resource correspondence shown in Table 2 below.

[0096] Table 2

[0097] Job ID Node ID Data Identification IZ1 IN1 ID1 IZ1 IN2 ID2 IZ1 IN3 ID3 IZ1 IN4 ID4

[0098] Optionally, the user needs to query the computing nodes, resources, or training data corresponding to the first parameter adjustment job in the management node, and may input the job identifier of the first parameter adjustment job in the management node.

[0099] The management node queries the node identification, resource identification, resource status, and other information of the computing node corresponding to the first parameter adjustment job from the resource correspondence according to the job identification of the first parameter adjustment job, and displays the queried information. And / or, the management node queries the node identification, data identification, and other information of the computing node corresponding to the first parameter adjustment job from the data correspondence according to the job identification of the first parameter adjustment job, and displays the queried information.

[0100] Step 205: The computing node trains the second intelligent model through the resource according to the training data.

[0101] The computing node continuously adjusts the parameter values ​​of the common parameters of the second intelligent model during the training of the second intelligent model. The common parameters of the second intelligent model are used to determine the functions of the second intelligent model. For example, assuming that an intelligent model for speech recognition needs to be trained, the parameter values ​​of the common parameters of the second intelligent model can be continuously adjusted through training data so that the second intelligent model has speech recognition function.

[0102] The computing node continuously adjusts the parameter values ​​of the common parameters of the second intelligent model during the training of the second intelligent model until the second intelligent model converges or fails to converge successfully and stops training, or the number of times the second intelligent model is trained reaches a specified number of times. The computing node obtains the training result of the second intelligent model training and sends a notification message to the management node, the notification message including the training result and the job identifier of the first parameter adjustment job.

[0103] Among them, any computing node that receives the second model training task trains the second intelligent model according to the above operations 204 and 205, and sends a notification message including the training results and the job identifier of the first parameter adjustment job to the management node after the training is completed.

[0104] The management node receives notification messages sent by each computing node. If the training result included in each notification message does not meet the specified condition, the current parameter value of each super parameter in the at least one super parameter corresponding to the first parameter adjustment job is obtained, and the X parameter values ​​of each super parameter are reconfigured according to the current parameter value of each super parameter and the training result included in each notification message, where X is an integer greater than 0, to obtain X first parameter value sets. According to each first parameter value set, the algorithm corresponding to the first parameter adjustment job is configured to obtain X first intelligent models, and X model training jobs corresponding to the first parameter adjustment job are generated, and each model training job may include Y first model training tasks. The first model training tasks included in the X model training jobs constitute the second batch of tasks for the first parameter adjustment business, that is, the second batch of tasks includes X*Y first model training tasks. For any first model training task, the first model training task includes a first intelligent model, a job identifier of the first parameter adjustment job, the resource name and resource size, the storage location of the training data set, the offset of a training data in the training data set, and the size of the training data. Next, the management node sends each first model training task included in the second batch of tasks to the computing nodes included in the node cluster to train the first intelligent model included in each first model training task. For detailed implementation process, see the following Figure 3 The illustrated embodiment will not be described in detail here.

[0105] Optionally, the management node saves the X model training jobs in a scheduling queue.

[0106] Optionally, for the above computing node, when the computing node stops training the second intelligent model, the computing node sends an update request to the management node, the update request includes the job identifier of the first parameter adjustment job, the resource identifier of the resource, and the resource status, and the resource status is an idle state. The management node receives the update request, and according to the job identifier of the first parameter adjustment job, the resource identifier of the resource, and the node identifier of the computing node included in the update request, sets the resource status of the resource in the resource correspondence to an idle state.

[0107] For example, for the above-mentioned computing node 1, when computing node 1 stops training the second intelligent model, it sends an update request 1 to the management node, and the update request 1 includes the job identifier IZ1 of the first parameter adjustment job, the resource identifier IR1 of the resource, and the resource state, and the resource state is an idle state. The management node receives the update request 1, and according to the job identifier IZ1 of the first parameter adjustment job, the resource identifier IR1 of the resource, and the node identifier IN1 of the computing node included in the update request 1, sets the resource state of the resource in the resource correspondence shown in Table 1 to an idle state, as shown in the following Table 3.

[0108] Similarly, when computing node 2 stops training the second intelligent model, it sends an update request 2 to the management node, and the update request 2 includes the job identifier IZ1 of the first parameter adjustment job, the resource identifier IR2 of the resource, and the resource state, and the resource state is an idle state. The management node receives the update request 2, and according to the job identifier IZ1 of the first parameter adjustment job, the resource identifier IR2 of the resource, and the node identifier IN2 of the computing node included in the update request 2, sets the resource state of the resource in the resource correspondence shown in Table 1 to an idle state, as shown in Table 3 below.

[0109] When the computing node 3 stops training the second intelligent model, it sends an update request 3 to the management node, and the update request 3 includes the job identifier IZ1 of the first parameter adjustment job, the resource identifier IR3 of the resource, and the resource state, and the resource state is an idle state. The management node receives the update request 3, and according to the job identifier IZ1 of the first parameter adjustment job, the resource identifier IR3 of the resource, and the node identifier IN3 of the computing node included in the update request 3, sets the resource state of the resource in the resource correspondence shown in Table 1 to an idle state, as shown in the following Table 3.

[0110] When the computing node 4 stops training the second intelligent model, it sends an update request 4 to the management node, and the update request 4 includes the job identifier IZ1 of the first parameter adjustment job, the resource identifier IR4 of the resource, and the resource state, and the resource state is an idle state. The management node receives the update request 4, and according to the job identifier IZ1 of the first parameter adjustment job, the resource identifier IR4 of the resource, and the node identifier IN4 of the computing node included in the update request 4, sets the resource state of the resource in the resource correspondence shown in Table 1 to an idle state, as shown in Table 3 below.

[0111] Table 3

[0112] Job ID Node ID Resource Identifier Resource Status IIZ1 IN1 IR1 Idle state IZ1 IN2 IR2 Idle state IZ1 IN3 IR3 Idle state IZ1 IN4 IR4 Idle state

[0113] In an embodiment of the present application, the management node allocates a computing node to each second model training task included in the first batch of tasks corresponding to the first parameter adjustment job, and the computing node obtains the resources and training data shown in the second model training task, and sends a storage request to the management node, the storage request including the job identifier of the first parameter adjustment job, the data identifier of the training data, the resource identifier of the resource, and the resource status. The management node forms a record with the job identifier, the node identifier of the computing node, the resource identifier, and the resource status and saves it in the resource correspondence relationship, and forms a record with the job identifier, the node identifier of the computing node, and the data identifier and saves it in the data correspondence relationship table. In this way, when the management node allocates computing nodes for the model training tasks included in the i-th batch of tasks corresponding to the first parameter adjustment job, i=2, 3, ..., and preferentially allocates computing nodes including the resources and / or training data required for the model training tasks for processing the first parameter adjustment job, the computing node does not need to obtain resources and / or training data when processing the model training tasks included in the i-th batch of tasks, thereby reducing the time consumption of model training and improving the efficiency of model training.

[0114] See also Figure 3 , the embodiment of the present application provides a method for model training, the intelligent model trained by the method is the intelligent model included in each model training task in the i-th batch of tasks corresponding to the parameter adjustment job, i=2, 3, .... This method can be applied to Figure 1 The system shown, the method includes:

[0115] Step 301: The management node schedules a first model training task, where the first model training task includes a job identifier of a first intelligent model and a first parameter adjustment job.

[0116] Optionally, the first model training task also includes information such as the resource name and resource size required to process the first model training task, the storage location of the training data set corresponding to the first parameter adjustment job, the offset of a piece of training data required to process the first model training task in the training data set, and the size of the training data.

[0117] Optionally, the scheduling queue of the management node includes a model training job corresponding to the first parameter adjustment job, each model training job includes at least one first model training task, and the first model training tasks included in each model training job constitute the i-th batch of tasks of the first parameter adjustment job.

[0118] In this step, the management node schedules a model training job from the scheduling queue, and schedules a first model training task from the first model training tasks included in the model training job.

[0119] For the model training job in the scheduling queue, the model training job is obtained in the following way:

[0120] The management node receives notification messages sent by each computing node, and includes a training result obtained by training an i-1th batch of tasks of the first parameter adjustment job according to each notification message. If the training result included in each notification message does not meet the specified condition, the current parameter value of each super parameter in at least one super parameter corresponding to the first parameter adjustment job is obtained, and X parameter values ​​of each super parameter are reconfigured according to the current parameter value of each super parameter and the training result included in each notification message, where X is an integer greater than 0, to obtain X first parameter value sets, and for any first parameter value set in the X first parameter value sets, the first parameter value set includes a parameter value of each super parameter. According to each first parameter value set, the algorithm corresponding to the first parameter adjustment job is configured to obtain X first intelligent models, and X model training jobs corresponding to the first parameter adjustment job are generated, and each model training job may include Y first model training tasks. The first model training tasks included in the X model training jobs constitute the i-th batch of tasks of the first parameter adjustment business, that is, the i-th batch of tasks includes X*Y first model training tasks. For any first model training task, the first model training task includes a first intelligent model, a job identifier of a first parameter adjustment job, the resource name and resource size, the storage location of the training data set, the offset of a piece of training data in the training data set and the size of the training data, etc. The management node saves the X model training jobs to the scheduling queue.

[0121] Step 302: The management node determines a first computing node from the node cluster according to the job identifier of the first parameter adjustment job, the first computing node includes first training data and at least one of idle first resources, the first resource is a resource required for processing the model training task of the first parameter adjustment job, and the first training data is training data required for training the intelligent model corresponding to the first parameter adjustment job.

[0122] Optionally, the management node determines the first computing node from the node cluster according to the resource correspondence, the data correspondence, and the job identifier of the first parameter adjustment job. In implementation, this can be achieved through the following operations 3021 to 3022, which are respectively:

[0123] 3021: The management node determines N computing nodes in the node cluster that include the first training data and / or the first resource according to the resource correspondence, the data correspondence, and the job identifier of the first parameter adjustment job, where N is an integer greater than 0.

[0124] Optionally, the management node adjusts the job identifier of the job according to the first parameter, obtains the corresponding node identifier of each computing node, the resource identifier and resource status of the first resource on each computing node from the resource correspondence relationship; and, according to the first parameter, adjusts the job identifier of the job, obtains the corresponding node identifier of each computing node and the data identifier of the first training data on each computing node from the data correspondence relationship. Assuming that the number of node identifiers of the computing nodes obtained twice is N, N computing nodes including the first training data and / or the first resource are determined.

[0125] For example, suppose a first model training task is scheduled, which includes a first intelligent model, a job identifier IZ1 of a first parameter adjustment job, a resource name and resource size, a storage location of a training data set, an offset of a piece of training data in the training data set, and the size of the training data.

[0126] The management node adjusts the job identifier IZ1 of the job according to the first parameter, and obtains the corresponding node identifier IN1 of computing node 1, the resource identifier IR1 of the first resource on computing node 1, and the resource status (idle state), the node identifier IN2 of computing node 2, the resource identifier IR2 of the first resource on computing node 2, and the resource status (idle state), the node identifier IN3 of computing node 3, the resource identifier IR3 of the first resource on computing node 3, and the resource status (idle state), the node identifier IN4 of computing node 4, the resource identifier IR4 of the first resource on computing node 4, and the resource status (idle state). And,

[0127] The management node adjusts the job identifier IZ1 of the job according to the first parameter, and obtains the corresponding node identifier IN1 of computing node 1 and the data identifier ID1 of the first training data on computing node 1, the node identifier IN2 of computing node 2 and the data identifier ID2 of the first training data on computing node 2, the node identifier IN3 of computing node 3 and the data identifier ID3 of the first training data on computing node 3, and the node identifier IN4 of computing node 4 and the data identifier ID4 of the first training data on computing node 4 from the data correspondence shown in Table 2. The node identifiers of the four computing nodes are obtained twice, that is, the computing nodes 1, 2, 3, and 4 including the first training data and / or the first resource are determined.

[0128] 3022: When there is at least one target node among the N computing nodes, the management node selects a target node from the at least one target node as a first computing node.

[0129] Among them, the target node includes an idle first resource, or the target node includes the first training data and the idle first resource, or the target node includes the first training data and the size of the unprotected resources included in the target node exceeds the size of resources required to process the first model training task, the unprotected resources are other resources in the target node except the protected resources, and the protected resources are resources that have been allocated to the parameter adjustment job and the protection time period corresponding to the protected resources has not yet ended.

[0130] For example, computing nodes 1, 2, 3, and 4 include first training data and / or idle first resources, and the management node may select computing node 1 from computing nodes 1, 2, 3, and 4 as the first computing node.

[0131] In this operation, the management node selects a target node from each target node as the first computing node according to the load information and / or node attribute information of each target node in the at least one target node.

[0132] Optionally, the management node scores each target node according to the load information and / or node attribute information of each target node, and selects a target node with the highest score from each target node as the first computing node, or selects a target node with a score exceeding a score threshold as the first computing node.

[0133] Optionally, the management node scores each target node according to a specified rule. The specified rule corresponds to a requirement, and different requirements have different scoring rules.

[0134] For example, if you want to balance the load of each computing node in the node cluster, the specified rule defines that the lighter the load of the computing node, the higher the score given to the computing node, and the higher the load of the computing node, the lower the score given to the computing node.

[0135] For another example, it is hoped that the load in the node cluster is concentrated on one or more nodes so as to shut down the nodes without load to achieve the purpose of energy saving. The specified rule defines that the heavier the load on the computing node, the higher the score given to the computing node, and the lighter the load on the computing node, the lower the score given to the computing node.

[0136] Optionally, the management node has a delayed scheduling function, so that when there is no target node among the N computing nodes, the computing node is not immediately allocated to the first model training task from the entire node cluster. Instead, when there is no target node among the N computing nodes, it is detected whether there is a computing node among the N computing nodes that becomes the target node within a first time period. The start time of the first time period is the time for scheduling the first model training task, and the time length of the first time period is the first threshold. If it is detected that a computing node becomes the target node within the first time period, the detected target node is determined as the first computing node. In this way, the computing node with idle first resources and / or first training data is still allocated to the first training task, so as to save the time spent by the computing node in allocating resources for the first model training task and / or the time spent in acquiring training data, thereby improving the efficiency of model training.

[0137] If it is detected that no computing node becomes the target node within the first time period, after the first time period ends, a second computing node is determined from the node cluster, and the size of unprotected resources included in the second computing node is greater than the size of resources required to process the first model training task.

[0138] Optionally, the first threshold may be configured in advance by an administrator in the management node.

[0139] Optionally, the management node schedules the next first model training task from the first model training tasks included in the model training job, and repeatedly executes this step until all the first model training tasks included in the model training job are scheduled. Then, the following step 303 is executed.

[0140] Step 303: The management node sends a first training request to the first computing node, where the first training request includes a first model training task.

[0141] Optionally, when the first computing node includes an idle first resource, the first training request further includes a resource identifier of the first resource. When the first computing node includes the first training data and the idle first resource, the first training request further includes a resource identifier of the first resource and a data identifier of the first training data. When the first computing node includes the first training data, the first training request further includes a data identifier of the first training data.

[0142] Optionally, in step 302, the management node determines a first computing node for each first model training task included in the model training job, so in the step, for any first model training task included in the model training job, the management node sends a first training request to the first computing node corresponding to the first model training task, and the first training request includes the first model training task. According to the operation of this step, the management node sends a first training request to the first computing node corresponding to each first model training task included in the model training job.

[0143] Then, the management node schedules the next model training job from the scheduling queue. The management node processes each first model training task included in the next model training job according to the above steps 301 to 303 until all the first model training tasks included in each model training job in the scheduling queue are scheduled.

[0144] Step 304: The first computing node trains the first intelligent model through the first resource according to the first training data.

[0145] The first computing node may be any one of the following three situations: first, the first computing node includes idle first resources; second, the first computing node includes first training data and idle first resources; third, the first computing node includes first training data and the size of unprotected resources included in the first computing node exceeds the size of resources required to process the first model training task.

[0146] For the first case mentioned above, the first computing node includes an idle first resource. In this step, the first computing node obtains the local first resource, obtains the training data set from the storage system according to the storage location of the training data set included in the first model training task, obtains a training data set corresponding to the first model training task from the training data set according to the offset and size of the training data set corresponding to the first model training task, that is, obtains the first training data, and trains the first intelligent model according to the first training data through the first resource.

[0147] Optionally, the first computing node further allocates a data identifier of the first training data, and sends a storage request to the management node, the storage request including the job identifier of the first parameter adjustment job and the data identifier of the first training data. The management node receives the storage request, and combines the node identifier of the first computing node, the job identifier of the first parameter adjustment job included in the storage request, and the data identifier of the first training data into a record and saves it in the data correspondence relationship.

[0148] Optionally, the first computing node further allocates a first protection time period for the first resource, where a start time of the first protection time period is a time when use of the first resource begins, and a time length of the first protection time period is a second threshold.

[0149] Optionally, the first computing node further allocates a second protection time period for the first training data, the start time of the second protection time period is the time when the first training data is acquired, and the time length of the second protection time period is a third threshold.

[0150] For the second case mentioned above, the first computing node includes first training data and idle first resources. In this step, the first computing node obtains the local first training data and first resources, and trains the first intelligent model through the first resources according to the first training data.

[0151] Optionally, the first computing node further allocates a first protection time period for the first resource, where a start time of the first protection time period is a time when use of the first resource begins, and a time length of the first protection time period is a second threshold.

[0152] Optionally, the first computing node further allocates a second protection time period for the first training data, the start time of the second protection time period is the time when the first training data starts to be used, and the time length of the second protection time period is a third threshold.

[0153] Optionally, the second threshold or the third threshold may be configured in advance by an administrator in each computing node in the node cluster.

[0154] In the first and second cases above, the first computing node sends an update request to the management node, the update request including the job identifier of the first parameter adjustment job, the resource identifier and resource status of the first resource, and the resource status is a usage status.

[0155] The management node receives the update request, adjusts the job identifier of the job, the node identifier of the first computing node, and the resource identifier of the first resource according to the first parameter, and updates the resource state of the first resource stored in the resource correspondence to a use state.

[0156] For the third case above, the first computing node includes the first training data and the size of the unprotected resources included in the first computing node exceeds the size of the resources required to process the first model training task. In this step, the first computing node obtains the first local training data, allocates the first resource from the unprotected resources included in the first computing node according to the resource name and resource size required to process the first model training task included in the first model training task, and trains the first intelligent model through the first resource according to the first training data. Among them, the unprotected resources are other resources in the first computing node except the protected resources, and the protected resources are the resources that the first computing node has allocated to the parameter adjustment job and the protection time period corresponding to the protected resources has not yet ended.

[0157] Optionally, the first computing node also allocates a resource identifier of the first resource, and sends a storage request to the management node, the storage request including the job identifier of the first parameter adjustment job, the resource identifier of the first resource, and the resource status, where the resource status is a usage status. The management node receives the storage request, and combines the node identifier of the first computing node, the job identifier of the first parameter adjustment job included in the storage request, the node identifier of the first computing node, the resource identifier of the first resource, and the resource status into a record and saves it in the resource correspondence.

[0158] Optionally, the first computing node further allocates a first protection time period for the first resource, the start time of the first protection time period is the time when the first resource is allocated, and the time length of the first protection time period is the second threshold.

[0159] Optionally, the first computing node further allocates a second protection time period for the first training data, the start time of the second protection time period is the time when the first training data starts to be used, and the time length of the second protection time period is a third threshold.

[0160] Optionally, the process for the first computing node to obtain the first training data locally may be: when the first training request includes the data identifier of the first training data, the first training data is obtained locally according to the data identifier of the first training data. Alternatively, when the first training request does not include the data identifier of the first training data, the job identifier of the first parameter adjustment job included in the first model training task is obtained, the data identifier of the first training data is obtained from the corresponding relationship between the job identifier and the data identifier, and the first training data is obtained locally according to the data identifier of the first training data.

[0161] Optionally, the process of obtaining the first local resource for the first computing node may be: when the first training request includes the resource identifier of the first resource, the first resource is obtained locally according to the resource identifier of the first resource. Alternatively, when the first training request does not include the resource identifier of the first resource, the job identifier of the job is adjusted according to the first parameter included in the first model training task, the resource identifier of the first resource is obtained from the correspondence between the job identifier and the resource identifier, and the local first resource is obtained according to the resource identifier of the first resource.

[0162] The first computing node continuously adjusts the parameter values ​​of the common parameters of the first intelligent model during the training of the first intelligent model until the first intelligent model converges or fails to converge successfully and stops training, or the number of times the first intelligent model is trained reaches a specified number of times. The first computing node obtains the training result of the first intelligent model training and sends a notification message to the management node, the notification message including the training result and the job identifier of the first parameter adjustment job.

[0163] Optionally, in the case where the management node allocates a second computing node to the first model training task, the management node sends a second training request to the second computing node, the second training request includes the first model training task. The second computing node receives the second training request, the second training request includes the first model training task, the first model training task includes the first intelligent model, the job identifier of the first parameter adjustment job, the resource name and resource size, the storage location of the training data set, the offset of a piece of training data in the training data set, and the size of the training data, and other information.

[0164] The second computing node allocates the first resource required for processing the first model training task according to the resource name and resource size included in the first model training task, and obtains the training data set from the storage system according to the storage location of the training data set included in the first model training task; according to the offset of a piece of training data corresponding to the first model training task in the training data set and the size of the training data set, obtains a piece of training data required for processing the first model training task from the training data set to obtain the first training data. According to the first training data, the first intelligent model included in the first model training task is trained through the first resource until the first intelligent model converges or fails to converge successfully, or the training is stopped when the number of times the first intelligent model is trained reaches a specified number. The second computing node obtains the training result of the first intelligent model training and sends a notification message to the management node, which includes the training result and the job identifier of the first parameter adjustment job.

[0165] Optionally, the second computing node further allocates a resource identifier for the first resource required to process the first model training task, and the resource identifier identifies the first resource in the computing node. The job identifier of the first parameter adjustment job and the resource identifier are correspondingly saved in the corresponding relationship between the job identifier and the resource identifier. And,

[0166] Optionally, the second computing node further allocates a data identifier for the first training data required for processing the first model training task, the data identifier identifying the first training data in the computing node. The job identifier of the first parameter adjustment job and the data identifier are correspondingly stored in the corresponding relationship between the job identifier and the data identifier.

[0167] Optionally, the second computing node further allocates a first protection time period for the first resource, the start time of the first protection time period is the time when the first resource is used, and the time length of the first protection time period is the second threshold. Since the first resource will be used after it is allocated, the start time of the first protection time period here is equal to the time when the first resource is allocated.

[0168] Optionally, the second computing node further allocates a second protection time period for the second training data, the start time of the second protection time period is the time when the second training data is used, and the time length of the second protection time period is a third threshold. Since the second training data will be used after it is acquired, the start time of the second protection time period here is equal to the time when the second training data is acquired.

[0169] Optionally, the second computing node also sends a storage request to the management node, where the storage request includes a job identifier of the first parameter adjustment job, a data identifier of the first training data, a resource identifier and a resource status of the first resource, where the resource status is a usage status.

[0170] The management node receives the storage request, combines the job identifier of the first parameter adjustment job, the node identifier of the second computing node, the resource identifier and the resource status into a record and saves the record in the resource correspondence relationship; and combines the job identifier of the first parameter adjustment job, the node identifier of the second computing node and the data identifier of the training data into a record and saves the record in the data correspondence relationship.

[0171] Optionally, the management node may receive notification messages sent by different computing nodes, and when the training results included in each notification message do not meet the specified conditions, the management node obtains the i+1th batch of tasks corresponding to the first parameter adjustment job, and then starts execution from step 301. When the training results included in each notification message meet the specified conditions, stop training the intelligent model of the first parameter adjustment job.

[0172] After stopping training the intelligent model of the first parameter adjustment job, for the above-mentioned first computing node, the first resource and / or the first training data in the first computing node will not be used.

[0173] Optionally, after the first protection time period corresponding to the first resource in the first computing node ends, the first resource may be released or may not be released. When releasing the first resource, the first computing node sends a first deletion request to the management node, and the first deletion request includes a node identifier of the first computing node and a resource identifier of the first resource. The management node receives the first deletion request and deletes the record including the node identifier of the first computing node and the resource identifier of the first resource from the resource correspondence.

[0174] Optionally, after the second protection period corresponding to the first training data in the first computing node ends, the first training data may be deleted or not deleted. When deleting the first training data, the first computing node sends a second deletion request to the management node, and the second deletion request includes the node identifier of the first computing node and the data identifier of the first training data. The management node receives the second deletion request and deletes the record including the node identifier of the first computing node and the data identifier of the first training data from the data correspondence relationship.

[0175] In an embodiment of the present application, the management node allocates a computing node to each second model training task included in the first batch of tasks corresponding to the first parameter adjustment job, and the computing node obtains the resources and training data shown in the second model training task, and sends a storage request to the management node, the storage request including the job identifier of the first parameter adjustment job, the data identifier of the training data, the resource identifier of the resource, and the resource status. The management node forms a record with the job identifier, the node identifier of the computing node, the resource identifier, and the resource status and saves it in the resource correspondence relationship, and forms a record with the job identifier, the node identifier of the computing node, and the data identifier and saves it in the data correspondence relationship table. In this way, when the management node allocates computing nodes for the model training tasks included in the i-th batch of tasks corresponding to the first parameter adjustment job, i=2, 3, ..., and preferentially allocates computing nodes including the resources and / or training data required for the model training tasks for processing the first parameter adjustment job, the computing node does not need to obtain resources and / or training data when processing the model training tasks included in the i-th batch of tasks, thereby reducing the time consumption of model training and improving the efficiency of model training.

[0176] See also Figure 4 The present application embodiment provides a model training device 400, which is deployed in Figure 1 , Figure 2 or Figure 3 The management node in the illustrated embodiment includes:

[0177] A processing unit 401 is used to schedule a first model training task, where the first model training task includes a first intelligent model and a job identifier of a first parameter adjustment job, where the first intelligent model is obtained by configuring an algorithm corresponding to the first parameter adjustment job based on a first parameter value set, and the first parameter value set includes a first parameter value of each super parameter in at least one super parameter corresponding to the first parameter adjustment job;

[0178] The processing unit 401 is further used to determine a first computing node from the node cluster according to the job identifier, the first computing node having at least one of first training data and idle first resources, the first resource being a resource required for processing a model training task of a first parameter adjustment job, and the first training data being training data required for training an intelligent model corresponding to the first parameter adjustment job;

[0179] The transceiver unit 402 is used to send a first training request to the first computing node, where the first training request includes a first model training task, and the first training request is used for the first computing node to train the first intelligent model based on at least one of the first resource and the first training data.

[0180] Optionally, the detailed implementation process of the processing unit 401 determining the first computing node can be found in Figure 3 The relevant contents in step 302 of the illustrated embodiment will not be described in detail here.

[0181] Optionally, the processing unit 401 is configured to:

[0182] Determine a first computing node from the node cluster according to the resource correspondence, the data correspondence, and the job identifier;

[0183] Among them, any record in the resource correspondence relationship includes the job identifier of the parameter adjustment job, the node identifier of the computing node in the node cluster, the resource identifier and the resource status, wherein the resource identifier is used to identify the resources required for the model training task included in the computing node for processing the parameter adjustment job, and the resource status is used to describe whether the resource is currently idle;

[0184] Any record in the data correspondence relationship includes the job identifier of the parameter adjustment job, the node identifier of the computing node in the node cluster, and the data identifier, where the data identifier is used to identify the training data included in the computing node for training the intelligent model corresponding to the parameter adjustment job.

[0185] Optionally, the processing unit 401 is configured to:

[0186] Determine, according to the resource correspondence, the data correspondence, and the job identifier, N computing nodes in the node cluster that include the first training data and / or the first resource, where N is an integer greater than 0;

[0187] When there is at least one target node among the N computing nodes, selecting a target node from the at least one target node as a first computing node;

[0188] Among them, the target node includes an idle first resource, or the target node includes the first training data and the idle first resource, or the target node includes the first training data and the size of the unprotected resources included in the target node exceeds the size of resources required to process the first model training task, the unprotected resources are other resources in the target node except the protected resources, and the protected resources are resources that have been allocated to the parameter adjustment job and the protection time period corresponding to the protected resources has not yet ended.

[0189] Optionally, the detailed implementation process of the processing unit 401 determining N computing nodes can be found in Figure 3 The relevant contents in step 3021 of the illustrated embodiment will not be described in detail here.

[0190] Optionally, the processing unit 401 is configured to:

[0191] Determine at least one target node according to the job identifier, and select a target node from each target node as a first computing node according to load information and / or node attribute information of each target node in the at least one target node;

[0192] Among them, the target node includes an idle first resource, or the target node includes the first training data and the idle first resource, or the target node includes the first training data and the size of the unprotected resources included in the target node exceeds the size of resources required to process the first model training task, the unprotected resources are other resources in the target node except the protected resources, and the protected resources are resources that have been allocated to the parameter adjustment job and the protection time period corresponding to the protected resources has not yet ended.

[0193] Optionally, the detailed implementation process of the processing unit 401 selecting a target node as the first computing node can be found in Figure 3 The relevant contents in step 3022 of the illustrated embodiment will not be described in detail again.

[0194] Optionally, the processing unit 401 is further configured to:

[0195] When there is no target node among the N computing nodes, detecting whether a computing node among the N computing nodes becomes a target node within a first time period, where the start time of the first time period is the time for scheduling the first model training task, the time length of the first time period is a first threshold, and the N computing nodes are computing nodes including the first training data and / or the first resource;

[0196] In a first time period, it is detected that a computing node becomes a target node, and the detected target node is determined as a first computing node;

[0197] Among them, the target node includes an idle first resource, or the target node includes the first training data and the idle first resource, or the target node includes the first training data and the size of the unprotected resources included in the target node exceeds the size of resources required to process the first model training task, the unprotected resources are other resources in the target node except the protected resources, and the protected resources are resources that have been allocated to the parameter adjustment job and the protection time period corresponding to the protected resources has not yet ended.

[0198] Optionally, any record in the resource correspondence relationship also includes the resource size of the resource identified by the resource identifier.

[0199] The processing unit 401 is further configured to detect that no computing node becomes a target node within the first time period, and after the first time period ends, determine a second computing node from the node cluster according to the resource correspondence, where the size of unprotected resources included in the second computing node is greater than the size of resources required to process the first model training task;

[0200] The transceiver unit 402 is also used to send a second training request to the second computing node, where the second training request includes the first model training task, and the second training request is used for the second computing node to train the first intelligent model.

[0201] Optionally, the transceiver unit 402 is further configured to receive a first deletion request, where the first deletion request includes a node identifier of the computing node and a resource identifier of the first resource, where the first deletion request is sent by the first computing node after the first protection time period ends, where the start time of the first protection time period is the time when the first resource was last used, and where the length of the first protection time period is the second threshold;

[0202] The processing unit 401 is further configured to delete the record including the node identifier of the first computing node and the resource identifier of the first resource from the resource correspondence relationship.

[0203] Optionally, the transceiver unit 402 is further configured to receive a second deletion request, where the second deletion request includes a node identifier of the first computing node and a data identifier of the first training data, where the second deletion request is sent by the first computing node after the second protection time period ends, where the start time of the second protection time period is the time when the first training data was last used, and where the time length of the second protection time period is a third threshold;

[0204] The processing unit 401 is further configured to delete the record including the node identifier of the first computing node and the data identifier of the first training data from the data correspondence relationship.

[0205] Optionally, the transceiver unit 402 is further used to send a third training request to the first computing node, the third training request including a second model training task, the second model training task including a job identifier of a second intelligent model and a first parameter adjustment job, the second model training task is a model training task included in the first batch of tasks corresponding to the first parameter adjustment job, the second intelligent model is obtained by configuring the algorithm based on a second parameter value set, the second parameter value set includes a second parameter value of each super parameter, and the third training request is used for the first computing node to allocate a first resource for training the second intelligent model and obtain first training data for training the second intelligent model; receive a storage request sent by the first computing node, the storage request including a data identifier of the first training data, a resource identifier of the first resource, and a resource status;

[0206] The processing unit 401 is also used to save the correspondence between the job identifier, the node identifier of the first computing node, the resource identifier of the first resource and the resource status in the resource correspondence; and to save the correspondence between the job identifier, the node identifier of the first computing node and the data identifier of the first training data in the data correspondence.

[0207] Optionally, the detailed implementation process of the transceiver unit 401 sending the third training request can be found in Figure 2 The relevant contents in step 203 of the illustrated embodiment will not be described in detail here.

[0208] Optionally, the detailed implementation process of the processing unit 401 storing the content in the resource correspondence relationship and the data correspondence relationship can be found in Figure 2 The relevant contents in step 204 of the illustrated embodiment will not be described in detail here.

[0209] In an embodiment of the present application, a processing unit schedules a first model training task, the first model training task includes a job identifier of a first intelligent model and a first parameter adjustment job, the first intelligent model is obtained by configuring an algorithm corresponding to the first parameter adjustment job based on a first parameter value set, and the first parameter value set includes a first parameter value of each super parameter in at least one super parameter corresponding to the first parameter adjustment job. A first computing node is determined from a node cluster according to the job identifier, the first computing node has at least one of first training data and idle first resources, the first resource is a resource required for processing the model training task of the first parameter adjustment job, and the first training data is training data required for training the intelligent model corresponding to the first parameter adjustment job. The transceiver unit sends a first training request to the first computing node, the first training request includes a first model training task, and the first training request is used for the first computing node to train the first intelligent model according to at least one of the first resource and the first training data. Among them, since the first computing node determined by the processing unit has at least one of the first training data and the idle first resource, after the first computing node receives the first training request including the first model training task, it is not necessary to allocate the first resource for the first model training task and / or obtain the first training data, thereby saving time for allocating the first resource and / or obtaining the first training data, and improving the efficiency of training the first intelligent model.

[0210] See also Figure 5 , the embodiment of the present application provides a model training device 500, the device 500 is deployed in Figure 1 , Figure 2 or Figure 3 The computing node in the embodiment shown includes:

[0211] The transceiver unit 501 is used to receive a first training request sent by the management node, the first training request includes a first model training task, the first model training task includes a first intelligent model and a job identifier of a first parameter adjustment job, the first intelligent model is obtained by configuring an algorithm corresponding to the first parameter adjustment job based on a first parameter value set, the first parameter value set includes a first parameter value of each super parameter in at least one super parameter corresponding to the first parameter adjustment job, and the device 500 has at least one of a first resource and first training data bound to the first parameter adjustment job;

[0212] The processing unit 502 is used to obtain at least one of the first resource and the first training data according to the job identifier; and train the first intelligent model according to at least one of the first resource and the first training data.

[0213] Optionally, the detailed implementation process of the processing unit 502 training the first intelligent model can be found in Figure 3The relevant contents in step 304 in the illustrated embodiment will not be described in detail here.

[0214] Optionally, the transceiver unit 501 is further used to receive a third training request, where the third training request includes a second model training task, where the second model training task includes a second intelligent model and a job identifier of a first parameter adjustment job, where the second model training task is a model training task included in a first batch of tasks corresponding to the first parameter adjustment job, where the second intelligent model is obtained by configuring the algorithm based on a second parameter value set, where the second parameter value set includes a second parameter value of each super parameter;

[0215] The processing unit 502 is also used to allocate a first resource for training a second intelligent model from unprotected resources, and to obtain first training data for training the second intelligent model, where the unprotected resources are resources other than protected resources in the device 500, and the protected resources are resources that have been allocated to the parameter adjustment operation and the protection time period corresponding to the protected resources has not yet ended; the second intelligent model is trained based on the first resources and the first training data.

[0216] Optionally, the detailed implementation process of the processing unit 502 allocating the first resource, obtaining the first training data and training the second intelligent model can be found in Figure 2 The relevant contents in steps 204 and 205 in the illustrated embodiment will not be described in detail here.

[0217] Optionally, the transceiver unit 501 is also used to send a storage request, which includes a data identifier of the first training data, a resource identifier and a resource status of the first resource. The storage request is used for the management node to save the correspondence between the job identifier, the node identifier of the device 500, the resource identifier and the resource status of the first resource in a resource correspondence relationship, and to save the correspondence between the job identifier, the node identifier of the device 500 and the data identifier of the first training data in a data correspondence relationship.

[0218] Optionally, the transceiver unit 501 is also used to send a first deletion request after the first protection time period ends, the first deletion request includes the node identifier of the device 500 and the resource identifier of the first resource, the start time of the first protection time period is the time when the device 500 last used the first resource, the time length of the first protection time period is the second threshold, and the first deletion request is used by the management node to delete the record including the node identifier of the device 500 and the resource identifier of the first resource from the resource correspondence relationship.

[0219] Optionally, the transceiver unit 501 is also used to send a second deletion request after the second protection time period ends, the second deletion request includes the node identifier of the device 500 and the data identifier of the first training data, the start time of the second protection time period is the time when the device 500 last used the first training data, the time length of the second protection time period is a third threshold, and the second deletion request is used for the management node to delete the record including the node identifier of the device 500 and the data identifier of the first training data from the data correspondence relationship.

[0220] In an embodiment of the present application, the transceiver unit receives a first training request sent by a management node, the first training request includes a first model training task, the first model training task includes a job identifier of a first intelligent model and a first parameter adjustment job, the first intelligent model is obtained by configuring an algorithm corresponding to the first parameter adjustment job based on a first parameter value set, and the first parameter value set includes a first parameter value of each super parameter in at least one super parameter corresponding to the first parameter adjustment job. Since the device locally has at least one of the first resource and the first training data bound to the first parameter adjustment job. Therefore, the processing unit can obtain at least one of the first resource and the first training data according to the job identifier; and train the first intelligent model according to at least one of the first resource and the first training data. In this way, when the processing unit receives the first model training task, it can save time for allocating the first resource and / or time for acquiring the first training data, thereby improving the efficiency of training the intelligent model.

[0221] See also Figure 6 , an embodiment of the present application provides a schematic diagram of a device 600 for model training. The device 600 may be a management node in any of the above embodiments. The device 600 includes at least one processor 601, a bus system 602, a memory 603, and at least one network interface 604.

[0222] The device 600 is a hardware structure device that can be used to implement Figure 4 The functional modules in the device 400 are as follows. For example, those skilled in the art may think of Figure 4 The processing unit 401 in the device 400 shown can be implemented by the at least one processor 601 calling the code in the memory 603. Figure 4 The transceiver unit 402 in the device 400 shown can be implemented through the network interface 604 .

[0223] Optionally, the processor 601 may be a general-purpose central processing unit (CPU), a network processor (NP), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the present application.

[0224] The bus system 602 may include a path for transmitting information between the components.

[0225] The network interface 604 is used to communicate with other devices or communication networks.

[0226] The memory 603 may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory may exist independently and be connected to the processor via a bus. The memory may also be integrated with the processor.

[0227] The memory 603 is used to store the application code for executing the solution of the present application, and the execution is controlled by the processor 601. The processor 601 is used to execute the application code stored in the memory 603, so as to realize the functions in the method of the present patent.

[0228] In a specific implementation, as an embodiment, the processor 601 may include one or more CPUs, such as Figure 6 CPU0 and CPU1 in.

[0229] In a specific implementation, as an embodiment, the device 600 may include multiple processors, such as Figure 6601 and processor 607 in FIG. Each of these processors may be a single-CPU processor or a multi-CPU processor. A processor herein may refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).

[0230] See also Figure 7 , an embodiment of the present application provides a schematic diagram of a communication device 700 for a PLC system. The device 700 may be a computing node in any of the above embodiments. The device 700 includes at least one processor 701, a bus system 702, a memory 703, and at least one network interface 704.

[0231] The device 700 is a hardware structure device that can be used to implement Figure 5 The functional modules in the device 500 are as follows. For example, those skilled in the art may think of Figure 5 The processing unit 502 in the device 500 shown can be implemented by the at least one processor 701 calling the code in the memory 703. Figure 5 The transceiver unit 501 in the device 500 shown can be implemented through the network interface 704 .

[0232] Optionally, the processor 701 may be a general-purpose central processing unit (CPU), a network processor (NP), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the present application.

[0233] The bus system 702 may include a path for transmitting information between the components.

[0234] The network interface 704 is used to communicate with other devices or communication networks.

[0235] The memory 703 may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory may exist independently and be connected to the processor via a bus. The memory may also be integrated with the processor.

[0236] The memory 703 is used to store the application code for executing the solution of the present application, and the execution is controlled by the processor 701. The processor 701 is used to execute the application code stored in the memory 703, thereby realizing the functions in the method of the present patent.

[0237] In a specific implementation, as an embodiment, the processor 701 may include one or more CPUs, such as Figure 7 CPU0 and CPU1 in.

[0238] In a specific implementation, as an embodiment, the device 700 may include multiple processors, such as Figure 7 701 and processor 707 in the embodiment of the present invention. Each of these processors may be a single-CPU processor or a multi-CPU processor. The processor herein may refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).

[0239] The present application embodiment provides a system for model training, including: Figure 4 The device 400 provided in the embodiment shown and Figure 5 The device 500 provided in the embodiment shown, or, includes Figure 6 The device 600 provided in the embodiment shown and Figure 7 The illustrated embodiment provides an apparatus 700 .

[0240] See also Figure 8 ,like Figure 4The device 400 provided in the embodiment shown or as Figure 6 The device 600 provided in the illustrated embodiment is a management node 801, such as Figure 5 The device 500 provided in the embodiment shown or as Figure 7 The device 700 provided in the illustrated embodiment is a computing node 802 .

[0241] A person skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware or by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.

[0242] The above description is only an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the principles of the present application shall be included in the protection scope of the present application.

Claims

1. A model training method, characterized in that: The method comprises: The management node schedules a first model training task, the first model training task includes a first intelligent model and a job identifier of a first parameter adjustment job, the first intelligent model is obtained by configuring an algorithm corresponding to the first parameter adjustment job based on a first parameter value set, and the first parameter value set includes a first parameter value of each super parameter in at least one super parameter corresponding to the first parameter adjustment job; The management node determines a first computing node from the node cluster according to the resource correspondence, the data correspondence and the job identifier, the first computing node having at least one of first training data and idle first resources, the first resource being a resource required for processing the model training task of the first parameter adjustment job, the first training data being training data required for training the intelligent model corresponding to the first parameter adjustment job, any one of the records in the resource correspondence includes the job identifier of the parameter adjustment job, the node identifier of the computing node in the node cluster, the resource identifier and the resource status, the resource identifier being used to identify the resource required for processing the model training task of the parameter adjustment job included in the computing node, and the resource status being used to describe whether the resource is currently idle; any one of the records in the data correspondence includes the job identifier of the parameter adjustment job, the node identifier of the computing node in the node cluster and the data identifier, the data identifier being used to identify the training data required for training the intelligent model corresponding to the parameter adjustment job included in the computing node; The management node sends a first training request to the first computing node, where the first training request includes the first model training task, and the first training request is used by the first computing node to train the first intelligent model based on at least one of the first resource and the first training data.

2. The method according to claim 1, characterized in that The management node determines a first computing node from the node cluster according to the resource correspondence, the data correspondence, and the job identifier, including: The management node determines, according to the resource correspondence, the data correspondence, and the job identifier, N computing nodes in the node cluster that include the first training data and / or the first resource, where N is an integer greater than 0; When there is at least one target node among the N computing nodes, the management node selects a target node from the at least one target node as a first computing node; Among them, the target node includes the idle first resource, or the target node includes the first training data and the idle first resource, or the target node includes the first training data and the size of the unprotected resources included in the target node exceeds the resource size required to process the first model training task, the unprotected resources are other resources in the target node except the protected resources, and the protected resources are resources allocated to the parameter adjustment job and the protection time period corresponding to the protected resources has not yet ended.

3. The method according to claim 1 or 2, characterized in that The management node determines a first computing node from a node cluster according to the job identifier, including: The management node determines at least one target node according to the job identifier, and selects a target node from each target node as a first computing node according to load information and / or node attribute information of each target node in the at least one target node; Among them, the target node includes the idle first resource, or the target node includes the first training data and the idle first resource, or the target node includes the first training data and the size of the unprotected resources included in the target node exceeds the resource size required to process the first model training task, the unprotected resources are other resources in the target node except the protected resources, and the protected resources are resources allocated to the parameter adjustment job and the protection time period corresponding to the protected resources has not yet ended.

4. The method according to claim 1 or 2, characterized in that: The method further comprises: When there is no target node among the N computing nodes, the management node detects whether a computing node among the N computing nodes becomes a target node within a first time period, where the start time of the first time period is the time for scheduling the first model training task, the time length of the first time period is a first threshold, and the N computing nodes are computing nodes that include the first training data and / or the first resource; The management node detects that a computing node becomes a target node within the first time period, and determines the detected target node as a first computing node; Among them, the target node includes the idle first resource, or the target node includes the first training data and the idle first resource, or the target node includes the first training data and the size of the unprotected resources included in the target node exceeds the resource size required to process the first model training task, the unprotected resources are other resources in the target node except the protected resources, and the protected resources are resources allocated to the parameter adjustment job and the protection time period corresponding to the protected resources has not yet ended.

5. The method according to claim 1 or 2, characterized in that: The method further comprises: The management node receives a first deletion request, where the first deletion request includes a node identifier of the computing node and a resource identifier of the first resource, where the first deletion request is sent by the first computing node after a first protection time period ends, where the start time of the first protection time period is the time when the first resource was last used, and the length of the first protection time period is a second threshold; The management node deletes the record including the node identifier of the first computing node and the resource identifier of the first resource from the resource correspondence.

6. The method according to claim 1 or 2, characterized in that: The method further comprises: The management node receives a second deletion request, where the second deletion request includes a node identifier of the first computing node and a data identifier of the first training data, the second deletion request is sent by the first computing node after a second protection time period ends, the start time of the second protection time period is the time when the first training data was last used, and the length of the second protection time period is a third threshold; The management node deletes the record including the node identifier of the first computing node and the data identifier of the first training data from the data correspondence relationship.

7. A method for model training, characterized in that: The method comprises: The computing node receives a first training request sent by the management node, the first training request includes a first model training task, the first model training task includes a first intelligent model and a job identifier of a first parameter adjustment job, the first intelligent model is obtained by configuring an algorithm corresponding to the first parameter adjustment job based on a first parameter value set, the first parameter value set includes a first parameter value of each super parameter in at least one super parameter corresponding to the first parameter adjustment job, and the computing node has at least one of a first resource and first training data bound to the first parameter adjustment job; The computing node acquires at least one of the first resource and the first training data according to the job identifier; The computing node trains the first intelligent model based on at least one of the first resource and the first training data.

8. The method according to claim 7, characterized in that The method further comprises: The computing node sends a first deletion request after the first protection time period ends, the first deletion request includes the node identifier of the computing node and the resource identifier of the first resource, the start time of the first protection time period is the last time the computing node used the first resource, the time length of the first protection time period is a second threshold, the first deletion request is used by the management node to delete the record including the node identifier of the computing node and the resource identifier of the first resource from the resource correspondence relationship, and the resource correspondence relationship is used to save the correspondence between the job identifier of the parameter adjustment job, the node identifier of the computing node, the resource identifier and the resource status.

9. The method according to claim 7 or 8, characterized in that The method further comprises: The computing node sends a second deletion request after the second protection time period ends, the second deletion request includes the node identifier of the computing node and the data identifier of the first training data, the start time of the second protection time period is the time when the computing node last used the first training data, the time length of the second protection time period is a third threshold, and the second deletion request is used by the management node to delete the record including the node identifier of the computing node and the data identifier of the first training data from the data correspondence relationship, and the data correspondence relationship is used to save the correspondence between the job identifier of the parameter adjustment job, the node identifier of the computing node, and the data identifier.

10. A device for model training, characterized in that: The device comprises: a processing unit, configured to schedule a first model training task, the first model training task including a first intelligent model and a job identifier of a first parameter adjustment job, the first intelligent model being obtained by configuring an algorithm corresponding to the first parameter adjustment job based on a first parameter value set, the first parameter value set including a first parameter value of each super parameter in at least one super parameter corresponding to the first parameter adjustment job; The processing unit is further used to determine a first computing node from the node cluster according to the resource correspondence, the data correspondence and the job identifier, the first computing node having at least one of first training data and idle first resources, the first resource being a resource required for processing the model training task of the first parameter adjustment job, the first training data being training data required for training the intelligent model corresponding to the first parameter adjustment job, any one of the records in the resource correspondence includes the job identifier of the parameter adjustment job, the node identifier of the computing node in the node cluster, the resource identifier and the resource status, the resource identifier is used to identify the resource required for processing the model training task of the parameter adjustment job included in the computing node, and the resource status is used to describe whether the resource is currently idle; any one of the records in the data correspondence includes the job identifier of the parameter adjustment job, the node identifier of the computing node in the node cluster and the data identifier, the data identifier is used to identify the training data required for training the intelligent model corresponding to the parameter adjustment job included in the computing node; A transceiver unit is used to send a first training request to the first computing node, where the first training request includes the first model training task, and the first training request is used for the first computing node to train the first intelligent model based on at least one of the first resource and the first training data.

11. The device according to claim 10, characterized in that The processing unit is used for: Determine, according to the resource correspondence, the data correspondence, and the job identifier, N computing nodes in the node cluster that include the first training data and / or the first resource, where N is an integer greater than 0; When there is at least one target node among the N computing nodes, selecting a target node from the at least one target node as a first computing node; Among them, the target node includes the idle first resource, or the target node includes the first training data and the idle first resource, or the target node includes the first training data and the size of the unprotected resources included in the target node exceeds the resource size required to process the first model training task, the unprotected resources are other resources in the target node except the protected resources, and the protected resources are resources allocated to the parameter adjustment job and the protection time period corresponding to the protected resources has not yet ended.

12. The device according to claim 10 or 11, characterized in that The processing unit is used for: Determine at least one target node according to the job identifier, and select a target node from each target node as a first computing node according to load information and / or node attribute information of each target node in the at least one target node; Among them, the target node includes the idle first resource, or the target node includes the first training data and the idle first resource, or the target node includes the first training data and the size of the unprotected resources included in the target node exceeds the resource size required to process the first model training task, the unprotected resources are other resources in the target node except the protected resources, and the protected resources are resources allocated to the parameter adjustment job and the protection time period corresponding to the protected resources has not yet ended.

13. The device according to claim 10 or 11, characterized in that The processing unit is further used for: When there is no target node among the N computing nodes, detecting whether a computing node among the N computing nodes becomes a target node within a first time period, where the start time of the first time period is the time for scheduling the first model training task, the time length of the first time period is a first threshold, and the N computing nodes are computing nodes that include the first training data and / or the first resource; Detecting that a computing node becomes a target node within the first time period, and determining the detected target node as a first computing node; Among them, the target node includes the idle first resource, or the target node includes the first training data and the idle first resource, or the target node includes the first training data and the size of the unprotected resources included in the target node exceeds the resource size required to process the first model training task, the unprotected resources are other resources in the target node except the protected resources, and the protected resources are resources allocated to the parameter adjustment job and the protection time period corresponding to the protected resources has not yet ended.

14. The device according to claim 10 or 11, characterized in that The transceiver unit is further configured to receive a first deletion request, the first deletion request including a node identifier of the computing node and a resource identifier of the first resource, the first deletion request being sent by the first computing node after a first protection time period ends, the start time of the first protection time period being the time when the first resource was last used, and the time length of the first protection time period being a second threshold; The processing unit is further configured to delete, from the resource correspondence, a record including the node identifier of the first computing node and the resource identifier of the first resource.

15. The device according to claim 10 or 11, characterized in that The transceiver unit is further configured to receive a second deletion request, the second deletion request including a node identifier of the first computing node and a data identifier of the first training data, the second deletion request being sent by the first computing node after a second protection time period ends, the start time of the second protection time period being the time when the first training data was last used, and the time length of the second protection time period being a third threshold; The processing unit is further configured to delete, from the data correspondence, a record including a node identifier of the first computing node and a data identifier of the first training data.

16. A device for model training, characterized in that: The device comprises: a transceiver unit, configured to receive a first training request sent by a management node, wherein the first training request includes a first model training task, the first model training task includes a first intelligent model and a job identifier of a first parameter adjustment job, the first intelligent model is obtained by configuring an algorithm corresponding to the first parameter adjustment job based on a first parameter value set, the first parameter value set includes a first parameter value of each super parameter in at least one super parameter corresponding to the first parameter adjustment job, and the device has at least one of a first resource and first training data bound to the first parameter adjustment job; A processing unit is used to obtain at least one of the first resource and the first training data according to the job identifier; and train the first intelligent model according to at least one of the first resource and the first training data.

17. The device according to claim 16, characterized in that The transceiver unit is also used to send a first deletion request after the first protection time period ends, the first deletion request includes the node identifier of the device and the resource identifier of the first resource, the start time of the first protection time period is the time when the device last used the first resource, and the time length of the first protection time period is a second threshold value, the first deletion request is used by the management node to delete the record including the node identifier of the device and the resource identifier of the first resource from the resource correspondence relationship, and the resource correspondence relationship is used to save the correspondence between the job identifier of the parameter adjustment job, the node identifier of the computing node, the resource identifier and the resource status.

18. The device according to claim 16 or 17, characterized in that The transceiver unit is further used to send a second deletion request after the second protection time period ends, the second deletion request includes the node identifier of the device and the data identifier of the first training data, the start time of the second protection time period is the time when the device last used the first training data, and the time length of the second protection time period is a third threshold value, and the second deletion request is used by the management node to delete the record including the node identifier of the device and the data identifier of the first training data from the data correspondence relationship, and the data correspondence relationship is used to save the correspondence between the job identifier of the parameter adjustment job, the node identifier of the computing node, and the data identifier.

19. A model training device, characterized in that: The device comprises a processor and a memory, and the processor executes a program in the memory, so that the device performs the method according to any one of claims 1 to 9.

20. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a program for implementing the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Hyper-parameter tuning method and device, server, client and medium

    CN110324185A

  • Model training method and device and cluster system

    CN111327692A