A model training control method, device and electronic device

By obtaining the attribute information of the model training task and the remaining time of the user, and using the training time prediction model to control the execution of the model training task, the problem of user timeout usage in the existing technology is solved, and more refined resource management and better user experience is achieved.

CN115310556BActive Publication Date: 2025-07-25BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211072303.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-02
Publication Date
2025-07-25
Estimated Expiration
2042-09-02

AI Technical Summary

Technical Problem

The billing scheme of the existing machine learning platform cannot effectively control the execution of model training tasks when users time out of the training resources, resulting in poor user experience and the possibility of overuse.

Method used

By obtaining the attribute information of the model training task and the remaining training time of the user, the training time is used to predict the required time of the model, and determining whether the model training task is allowed based on the prediction time and the remaining time, and the task is only executed if allowed.

Benefits of technology

Effectively control the execution of model training, reduce the possibility of users exceeding the training time, improve the control of model training, and improve user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115310556B_ABST
    Figure CN115310556B_ABST
Patent Text Reader

Abstract

The present disclosure provides a model training control method, apparatus, and electronic device, relating to the field of artificial intelligence technologies, and particularly to the field of deep learning technologies. The specific implementation solution is as follows: Obtain the attribute information of a first model training task initiated by a user, and obtain the remaining training duration of the user, where the attribute information includes the model attributes corresponding to the first model training task; input the attribute information into a training duration prediction model to obtain the predicted duration required for the first model training task; based on the predicted duration and the remaining training duration, determine whether to allow the execution of the first model training task; and in the case where it is determined to allow the execution of the first model training task, execute the first model training task according to the attribute information. The present disclosure realizes the control of model training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, and particularly to the field of deep learning technology. Background Art

[0002] In related technologies, artificial intelligence models can basically be divided into two types: classical machine learning and deep learning. The training and development of various models involve multiple aspects of work such as data processing, model management, model inference, and service deployment, thus having certain requirements for the development environment. Based on this, machine learning platforms have emerged, aiming to provide developers with full-life-cycle management services for model development. Among them, the public cloud provided by machine learning platforms is one of the most common service methods, that is, providing third-party users with the required model training-related resources based on cloud hardware resources. Summary of the Invention

[0003] This disclosure provides a model training control method, apparatus, and electronic device.

[0004] According to one aspect of this disclosure, a model training control method is provided, including:

[0005] Obtain the attribute information of the first model training task initiated by the user, and obtain the remaining training duration of the user, where the attribute information includes the model attributes corresponding to the first model training task;

[0006] Input the attribute information into a training duration prediction model to obtain the predicted duration required for the first model training task;

[0007] Based on the predicted duration and the remaining training duration, determine whether to allow the execution of the first model training task;

[0008] In the case of determining that the execution of the first model training task is allowed, execute the first model training task according to the attribute information.

[0009] According to another aspect of this disclosure, a training method for a training duration prediction model is provided, including:

[0010] Obtain historical model training logs, where the historical model training logs include the attribute information and execution duration of the model training tasks executed historically;

[0011] Input the attribute information into a training duration prediction model to obtain the training predicted duration of the historically executed model training tasks;

[0012] Compare the training predicted duration and the execution duration to obtain the loss of the training duration prediction model;

[0013] Adjust the parameters of the training duration prediction model based on the loss until the training is completed, and obtain the target training duration prediction model.

[0014] According to another aspect of the present disclosure, there is provided a model training control device, including:

[0015] An information acquisition module, configured to acquire the attribute information of the first model training task initiated by the user, and acquire the remaining training duration of the user, where the attribute information includes the model attribute corresponding to the first model training task;

[0016] A duration prediction module, configured to input the attribute information into the training duration prediction model to obtain the predicted duration required for the first model training task;

[0017] An execution judgment module, configured to judge whether to allow the execution of the first model training task based on the predicted duration and the remaining training duration;

[0018] A task execution module, configured to execute the first model training task according to the attribute information when it is determined that the execution of the first model training task is allowed.

[0019] According to still another aspect of the present disclosure, there is provided a training device for a training duration prediction model, including:

[0020] A log acquisition module, configured to acquire historical model training logs, where the historical model training logs include the attribute information and execution duration of the historically executed model training tasks;

[0021] A duration acquisition module, configured to input the attribute information into the training duration prediction model to obtain the training predicted duration of the historically executed model training task;

[0022] A loss acquisition module, configured to compare the training predicted duration and the execution duration to obtain the loss of the training duration prediction model;

[0023] A parameter adjustment module, configured to adjust the parameters of the training duration prediction model based on the loss until the training is completed, and obtain the target training duration prediction model.

[0024] The model training control method provided by the present disclosure obtains the attribute information of the first model training task initiated by the user and the remaining training duration of the user. Among them, the attribute information includes the model attributes corresponding to the first model training task; inputs the attribute information into the training duration prediction model to obtain the predicted duration required for the first model training task; based on the predicted duration and the remaining training duration, determines whether to allow the execution of the first model training task; and in the case of determining that the execution of the first model training task is allowed, executes the first model training task according to the attribute information. Thus, the control of model training is achieved.

[0025] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Brief Description of the Drawings

[0026] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0027] Figure 1 is a flowchart of the first model training control method provided by the present disclosure;

[0028] Figure 2a is a possible implementation manner of step S11 provided by the present disclosure;

[0029] Figure 2b is a flowchart of the second model training control method provided by the present disclosure;

[0030] Figure 3 is a flowchart of the training method of a training duration prediction model provided by the present disclosure;

[0031] Figure 4 is an example diagram of the training method of a training duration prediction model provided by the present disclosure;

[0032] Figure 5a is an example diagram of applying a model training control method to a model training control system provided by the present disclosure;

[0033] Figure 5b is an example diagram of the application of a Web platform provided by the present disclosure;

[0034] Figure 5c is an example diagram of the application of a billing module provided by the present disclosure;

[0035] Figure 5d is an example diagram of the application of a training module provided by the present disclosure;

[0036] Figure 5e It is an application example diagram of a billing duration prediction module provided according to the present disclosure;

[0037] Figure 6 It is a schematic structural diagram of a model training control device provided according to the present disclosure;

[0038] Figure 7 It is a schematic structural diagram of a training device for a training duration prediction model provided according to the present disclosure;

[0039] Figure 8 It is a block diagram of an electronic device for implementing the model training control method and the training method of the training duration prediction model in the embodiments of the present disclosure. Detailed implementation manners

[0040] The following makes an explanation of the exemplary embodiments of the present disclosure in conjunction with the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted below.

[0041] In the related art, the operation modes of machine learning platforms are mainly divided into pay-per-use and pay-by-period. The former means billing according to short-period quantities such as minutes and hours of third-party users using training models, and the latter means billing according to long-period quantities such as weeks, months, and years (which can also be called prepaid, resource package payment, etc.). For users with infrequent model training and long idle time of resources, obviously the billing method based on short-period quantities is more suitable.

[0042] The short-period billing scheme mainly includes the following aspects: First, different duration quotas are set for different types of training resources; second, the user is prompted to pay attention to the training duration in the document (for example, display the current training duration); third, before the model training task starts, it is verified whether the user is in arrears (whether there is still a training duration quota), and this training is allowed to be initiated when the user is not in arrears. The billing will stop when the model training task is in the end state. To prevent the influence of too long model training time, etc., an automatic stop option is set so that the user can choose to stop training after a fixed number of hours of running; fourth, the time point for deducting fees for the model training task is after the model training ends (including manual suspension of training or automatic stop of training). If the model training task fails to run due to system exceptions, the duration of the corresponding training task will not be involved in the billing.

[0043] However, in actual applications, there will be a situation where during the running of a training task, the user's time limit has been exhausted, but the user can still initiate other training tasks. This is because the time limit in the user's account will not be updated until the training task ends, resulting in the conclusion that other training tasks do not owe fees when verifying the user's time limit. In the prior art, in order to reduce the losses caused by users' overuse of training resources, the time limit of the user will be counted when any one of the user's training tasks ends. When it is found that the time limit is exhausted, flow control will be triggered to stop all training tasks that use the model resources corresponding to the time limit.

[0044] Although such improvements alleviate certain losses, there are still certain problems: First, most users cannot grasp the training duration of the model and cannot predict the applicable range of the time limit in the account. As a result, they may initiate multiple model training tasks simultaneously. When any one of the model training tasks exhausts the time limit, all other model training tasks that use the same type of resources will stop, resulting in great resistance to the user's model training and poor user experience. Second, there are users who do not understand the training resources, resulting in difficulty in determining the time selection corresponding to the resources required for model training, or situations where parameter settings cause the training duration to exceed expectations. Third, there is still a high probability of overuse. For example, in visual modeling, a canvas may have multiple training components. Users can create different canvases, and the number of components running simultaneously is large, and the overused time will also be more.

[0045] To solve at least one of the above problems, the present disclosure provides a model training control method, including:

[0046] Obtaining the attribute information of the first model training task initiated by the user, and obtaining the remaining training time of the user, where the attribute information includes the model attribute corresponding to the first model training task;

[0047] Inputting the attribute information into a training duration prediction model to obtain the predicted duration required for the first model training task;

[0048] Based on the predicted duration and the remaining training time, determining whether to allow the execution of the first model training task;

[0049] In the case of determining that the execution of the first model training task is allowed, executing the first model training task according to the attribute information.

[0050] As can be seen from the above, for the model training control method provided by the present disclosure, when a model training task initiated by a user is obtained, first, the attribute information of the model training task and the remaining training duration of the user are obtained, and the predicted duration of the model training task is predicted based on the attribute information. Then, based on the predicted duration and the remaining training duration of the user, it is determined whether the remaining training duration of the user allows the execution of the model training task. Only when it is determined that execution is allowed, the model training task is executed. Thus, it is possible to effectively control whether the model training is executed when the user initiates the model training, reduce the possibility that the user overuses the training duration, and improve the control degree of the model training. And it can also enable the user to formulate more fine-grained strategies for model training subsequently, such as selecting a more suitable training model, etc., allowing the user to anticipate the training duration required for the model training task, and thereby improving the user experience.

[0051] The model training control method provided by the present disclosure will be described in detail below through specific embodiments.

[0052] The method of the embodiments of the present disclosure is applied to an intelligent terminal and can be implemented through the intelligent terminal. In actual use, the intelligent terminal can be a computer, a server, a data center, etc.

[0053] See Figure 1 , Figure 1 which is a schematic flowchart of the first model training control method provided by the present disclosure, including:

[0054] Step S11: Obtain the attribute information of the first model training task initiated by the user, and obtain the remaining training duration of the user.

[0055] Among them, the attribute information includes the model attributes corresponding to the first model training task. Optionally, when the user initiates the first model training task, the user can input the attribute information of the model corresponding to the first model training task. For example, model framework, model network, number of training dataset samples, batch size, number of iterations, training model, etc.

[0056] The first model training task initiated by the user represents the model training currently requested by the user. In an embodiment of the present disclosure, the first model training task is used for image classification.

[0057] When receiving the first model training task initiated by the user, first, for the model training requested by the user, obtain its corresponding attribute information, and the attribute information includes the model attributes corresponding to the model training requested by the user.

[0058] In one embodiment of the present disclosure, the attribute information includes a model framework (framework) required for the first model training task, for example, paddle, tensorflow, pytorch (all are existing model frameworks); a model network (network), for example, RNN, YOLO, BERT (all are existing model networks); at least one of the number of training data set samples (dataSize), batch size (batchSize), number of iterations (epochNum), and training machine type (vmType).

[0059] In addition, the user's remaining training time is also obtained. Specifically, the remaining training time indicates the time the user can currently perform model training, which can be calculated based on the user's account balance.

[0060] Step S12: inputting the attribute information into a training duration prediction model to obtain the predicted duration required for the first model training task.

[0061] After obtaining the attribute information of the first model training task requested by the user, the attribute information is input into the training duration prediction model, and the predicted duration required for the first model training task is obtained based on the training duration prediction model. Specifically, the predicted duration indicates the duration required for the execution process of the first model training task predicted in advance by the training duration prediction model, that is, the duration required from the start to the completion of the model training requested by the user.

[0062] Step S13: Based on the predicted duration and the remaining training duration, determine whether to allow the first model training task to be executed.

[0063] After obtaining the predicted duration of the first model training task and the remaining training time of the user, it is determined whether the first model training task is allowed to be executed. In one example, when the remaining training time meets the predicted duration, it can be determined that the first model training task is allowed to be executed. Specifically, when the predicted duration does not exceed the remaining training time of the user, it means that the user's current remaining training time is still sufficient to execute the first model training task, and it can be determined that the first model training task is allowed to be executed. In one example, it can also be determined whether the predicted duration does not exceed a certain proportion of the remaining training time, for example, the predicted duration does not exceed one-half, two-thirds, etc. of the remaining training time. The specific proportion value can be set according to actual needs. At this time, it means that the remaining training time is sufficient to allow the first model training task to be executed smoothly, and it can be determined that the first model training task is allowed to be executed.

[0064] Step S14: when it is determined that the first model training task is allowed to be executed, the first model training task is executed according to the attribute information.

[0065] When it is determined that the first model training task is allowed to be executed, the first model training task is executed according to the attribute information of the first model training task.

[0066] In one example, when it is determined that the first model training task is allowed to be executed, an instruction indicating that the first model training task is allowed to be executed can also be sent to the user. The instruction may include the predicted duration obtained based on the training prediction duration model described above, so that the user can determine whether to adjust the attribute information of the first model training task based on the predicted duration. If the user adjusts the attribute information, after receiving the adjusted attribute information of the first model training task sent by the user, the predicted duration is obtained again based on the adjusted attribute information and fed back to the user until an instruction to execute the first model training task sent by the user is received, and the first model training task is executed according to the current attribute information of the first model training task. Based on this, it is possible to provide the user with a prediction of the execution duration of the training task, reduce the user's uncertainty panic about the model training duration, and at the same time provide the predicted duration to the user, which can also facilitate the user to select more appropriate attribute information resources for the model training task.

[0067] In one embodiment of the present disclosure,

[0068] In the case where it is determined that the first model training task is not allowed to be executed, an instruction indicating that the initiation of the first model training task has failed is sent to the user.

[0069] In one example, a prompt message of the failure reason can also be sent to the user. For example, a prompt message indicating that the remaining training duration does not meet the predicted duration is sent to the user.

[0070] As can be seen from the above, in the model training control method provided by the present disclosure, when a model training task initiated by the user is obtained, first, the attribute information of the model training task and the remaining training duration of the user are obtained, and the predicted duration of the model training task is predicted based on the attribute information. Then, based on the predicted duration and the remaining training duration of the user, it is determined whether the remaining training duration of the user allows the execution of the model training task. Only when it is determined that the execution is allowed, the model training task is executed. Thus, it is possible to effectively control whether the model training is executed when the user initiates the model training, reduce the possibility of the user overusing the training duration, and improve the control degree of the model training. And it can also enable the user to formulate a more fine-grained model training strategy in the future, such as selecting a more appropriate training model, etc., so that the user can have a pre-judgment on the training duration required for the model training task, thereby improving the user experience.

[0071] In a possible implementation manner, as Figure 2a shown, the above step S11 of obtaining the remaining training duration of the user includes:

[0072] Step S21: Obtain the first resource type required for the first model training task;

[0073] Step S22: Obtain the remaining training duration corresponding to the first resource type according to the preset correspondence between the resource type and the remaining training duration of the user.

[0074] It can be understood that for different types of model training tasks, the required resource types are different. For example, when the first model training task is an image classification task, the required dataset is an image dataset; when the first model training task is a speech recognition task, the required dataset is a speech dataset.

[0075] In the embodiments of the present disclosure, there is a correspondence between the resource type and the remaining training duration of the user, that is, for different resource types, the user may have different remaining training durations. For example, for the resource type related to image classification, the remaining training duration of the user is 24 hours; for the resource type related to speech recognition, the remaining training duration of the user is 48 hours. Specifically, the correspondence between the resource type and the remaining training duration of the user is preset, which can be arranged by the user himself when setting the remaining training duration in advance, or can be arranged by the intelligent terminal implementing the model training control method provided by the present disclosure based on the execution duration of the historical model training task.

[0076] Therefore, when obtaining the remaining training duration of the user, first obtain the first resource type required for the first model training task initiated by the user, and then obtain the remaining training duration corresponding to the first resource type according to the preset correspondence between the resource type and the remaining training duration of the user. The remaining training duration obtained at this time is the duration that can be used to execute the first model training task.

[0077] After the execution of the first model training task ends, as Figure 2b shown, a flowchart of the second model training control method is provided. The above method further includes:

[0078] Step S23: Obtain the execution duration of the first model training task;

[0079] Step S24: Update the remaining training duration corresponding to the first resource type based on the execution duration;

[0080] Step S25: When the remaining training duration corresponding to the first resource type is exhausted, send an instruction to the user to indicate stopping the second model training task.

[0081] Wherein, the second model training task is a model training task whose required resource type is the first resource type.

[0082] After the execution of the first model training task ends, the actual execution duration of the first model training task is also obtained, that is, the duration from the actual start to the end of the first model training task. Then, based on the execution duration, the remaining training duration corresponding to the first resource type is updated. Specifically, it can be obtained by subtracting the execution duration of the first model training task from the remaining training duration of the user obtained above, so as to obtain the updated remaining training duration corresponding to the first resource type. It can be understood that during the execution of the first model training task, the user may also initiate a second model training task, and the required resource type is also the first resource type. Then, when the remaining training duration corresponding to the first resource type is exhausted, an instruction to stop the second model training task is sent to the user.

[0083] As can be seen from the above, the model training control method provided by the present disclosure pre-sets the corresponding relationship between the remaining training duration of the user and the resource type. For different resource types, the user can have different remaining training durations. After the first model training task ends and the remaining training duration is updated, if the remaining training duration is exhausted, only the second model training task corresponding to this resource type is stopped, without involving other model training tasks, thereby realizing more fine-grained control of model training.

[0084] In an embodiment of the present disclosure, after the execution of the first model training task ends, it further includes:

[0085] Collect the training log of the first model training task and store it as a historical model training log, where the training log includes the attribute information and execution duration of the first model training task.

[0086] In an embodiment of the present disclosure, based on the historical model training log, the training duration prediction model is trained and updated.

[0087] After the execution of the first model training task ends, the training log of the first model training task is also collected and stored as a historical model training log. Based on the historical model training log, the training duration prediction model is trained and updated. Specifically, at preset time intervals, such as every three days, one week, one month, etc., the updated historical model training log is used to train and update the training duration prediction model.

[0088] As can be seen from the above, the model training control method provided by the present disclosure, after the execution of the first model training task ends, also collects the training log and stores it as a historical model training log, and trains and updates the training duration prediction model based on this, so that the training duration prediction model can be updated frequently and better meet the current needs, thereby better realizing the control of model training.

[0089] See Figure 3 ,Figure 3 A flowchart showing a method for training a training duration prediction model provided by the present disclosure, including:

[0090] Step S31: Obtain historical model training logs.

[0091] Among them, the historical model training logs include attribute information and execution duration of historical model training tasks;

[0092] Step S32: Input the attribute information into the training duration prediction model to obtain the training prediction duration of the historical model training task;

[0093] Step S33: Compare the training prediction duration and the execution duration to obtain the loss of the training duration prediction model;

[0094] Step S34: Adjust the parameters of the training duration prediction model based on the loss until the training is completed to obtain the target training duration prediction model.

[0095] In an embodiment of the present disclosure, the attribute information includes at least one of the model framework, model network, number of training dataset samples, batch size, number of epochs, and training model type applied in the historical model training task.

[0096] In one example, the training duration prediction model can be trained using the following model:

[0097] F = {f|Y = f θ (X), θ ∈ R n}

[0098] The loss function used to calculate the loss of the training duration prediction model is:

[0099] L(Y, f(X))

[0100] Where X is the input sample data, obtained by appropriately processing the attribute information of the historical model training task. Y is the predicted duration. f is a regression model, specifically, it can be linear regression, Logistic regression (a regression function), etc. Y = f(X) is the trained model, and F is the set of models that conform to Y = f(X). θ is the model parameter vector, R is the value range of the model parameter vector, which is a set of natural numbers, and n represents the dimension of the model parameter vector.

[0101] In one example, the model training control method provided by the present disclosure can be applied to a model training system. Specifically, the model training system can be any machine learning platform, such as Figure 4As shown, it includes a training module for performing model training tasks. The present disclosure adds a billing duration prediction module to the model training system for placing a training duration prediction model, and based on this, obtains the predicted duration of the model training task. When implementing the training method of the training duration prediction model, the training module (there are training module 1, training module 2... training module N, etc.) sends historical model training logs (including log 1, log 2... log N, etc.) to the billing duration prediction module, enabling the billing duration prediction module to perform offline training on the training duration prediction model, and querying and predicting the model training task based on the target training duration prediction model obtained after training to obtain the corresponding predicted duration.

[0102] In one example, the present disclosure also provides a specific example diagram of an application of a model training control method to a model training system, such as Figure 5a shown, where the Web platform represents a platform for receiving model training tasks initiated by users, the intelligent cloud is used to store resources required for model training tasks, and the billing system is used to calculate the execution duration of the resources applied to the model training task to complete the billing for users.

[0103] The present disclosure also provides an application example of a model training control method. Among them, the Web platform is as Figure 5b shown, and is used to receive model training tasks initiated by users, including attribute information of the model training task (set data set, selected model framework, model network, etc., set training data set sample number, batch size, number of iterations, training model type, etc. training parameters), then calls the training duration prediction model to obtain the predicted duration, and based on this, predicts whether the remaining training duration is sufficient. If it is sufficient, submit the training and start executing the model training task; if it is not sufficient, end the process, and then an instruction indicating a failure in initiation can be sent to the user.

[0104] The billing module is as Figure 5c shown, and is used to parse the billing information and update the user resource quota after the model training task is executed, that is, update the remaining training duration of the user for the corresponding resource type based on the execution duration and the remaining training duration. When the remaining training duration for the corresponding resource type is exhausted, that is, the quota for this resource type is exhausted, the flow control is notified to notify the training module to stop executing the model training task, and at the same time, an instruction indicating to stop the model training task using the same type of resource is sent to the user.

[0105] The training module is as Figure 5dAs shown, it is used to perform a model training task. When a flow control notification is received during the execution, if the training has been completed at this time, that is, the model training task has ended, a billing message is pushed to the billing module. The billing message is the resource type of the ended model training task and the corresponding execution duration, and the training log is pushed and stored as a historical model training log. If a flow control notification is received during the execution and the training has not been completed at this time, that is, the model training task has not ended, the training is stopped.

[0106] The billing duration prediction module is as Figure 5e shown. After collecting the training logs and completing the training of the training duration prediction model based on the training logs, the obtained target training duration prediction model provides a duration prediction service for users.

[0107] Among them, Start indicates the start of a step, End indicates the end of a step, Y indicates yes, and N indicates no.

[0108] As can be seen from the above, the training method of the training duration prediction model provided by the present disclosure trains the training duration prediction model based on the model framework, model network, number of training dataset samples, batch size, number of iterations, training model type, and execution duration applied in the historical model training task, which can fully consider various influencing factors of the model training duration, thereby obtaining a more accurate prediction result and providing a more reliable duration reference.

[0109] See Figure 6 , the present disclosure also provides a structural schematic diagram of a model training control device, including:

[0110] An information acquisition module 601, configured to acquire the attribute information of the first model training task initiated by the user, and acquire the remaining training duration of the user, where the attribute information includes the model attribute corresponding to the first model training task;

[0111] A duration prediction module 602, configured to input the attribute information into the training duration prediction model to obtain the predicted duration required for the first model training task;

[0112] An execution judgment module 603, configured to judge whether to allow the execution of the first model training task based on the predicted duration and the remaining training duration;

[0113] A task execution module 604, configured to execute the first model training task according to the attribute information when it is determined that the execution of the first model training task is allowed.

[0114] As can be seen from the above, when the model training control device provided by the present disclosure obtains a model training task initiated by a user, it first obtains the attribute information of the model training task and the remaining training duration of the user, predicts the predicted duration of the model training task based on the attribute information, and determines whether the remaining training duration of the user allows the execution of the model training task based on the predicted duration and the remaining training duration of the user. Only when it is determined that execution is allowed, the model training task is executed. Thereby, it can effectively control whether to execute model training when the user initiates model training, reduce the possibility of the user overusing the training duration, and improve the control degree of model training. And it can also enable the user to formulate a more fine-grained model training strategy subsequently, such as selecting a more suitable training model, etc., allowing the user to have a pre-judgment on the training duration required for the model training task, thereby improving the user experience.

[0115] In one embodiment of the present disclosure, it further includes:

[0116] A failure instruction sending module, configured to send an instruction indicating the failure of the initiation of the first model training task to the user when it is determined that the execution of the first model training task is not allowed.

[0117] In one embodiment of the present disclosure, wherein, the attribute information includes at least one of the model framework, model network, number of training dataset samples, batch size, number of iterations, and training model required for the first model training task.

[0118] In one embodiment of the present disclosure, wherein, the first model training task is for image classification.

[0119] In one embodiment of the present disclosure, the information acquisition module 601 is specifically configured to:

[0120] Obtain the first resource type required for the first model training task;

[0121] According to the preset correspondence between the resource type and the remaining training duration of the user, obtain the remaining training duration corresponding to the first resource type;

[0122] The device further includes:

[0123] An execution duration acquisition module, configured to acquire the execution duration of the first model training task;

[0124] A duration update module, configured to update the remaining training duration corresponding to the first resource type based on the execution duration;

[0125] A stop instruction sending module, configured to send an instruction to the user to stop the second model training task when the remaining training duration corresponding to the first resource type is exhausted, where the second model training task is a model training task that requires the first resource type.

[0126] As can be seen from the above, the model training control device provided by the present disclosure pre-sets the correspondence between the remaining training duration of the user and the resource type. For different resource types, the user can have different remaining training durations. After the first model training task ends and the remaining training duration is updated, if the remaining training duration is exhausted, only the second model training task corresponding to this resource type is stopped, without involving other model training tasks, thus realizing more fine-grained control of model training.

[0127] In one embodiment of the present disclosure, it further includes:

[0128] A log storage module, configured to collect the training logs of the first model training task and store them as historical model training logs, where the training logs include the attribute information and execution duration of the first model training task.

[0129] In one embodiment of the present disclosure, it further includes:

[0130] A model update module, configured to train and update the training duration prediction model based on the historical model training logs.

[0131] As can be seen from the above, the model training control device provided by the present disclosure, after the first model training task is executed, also collects training logs and stores them as historical model training logs, and trains and updates the training duration prediction model based on this, so that the training duration prediction model can be updated frequently and better meet the current needs, thus better realizing the control of model training.

[0132] See Figure 7 , the present disclosure also provides a structural schematic diagram of a training device for a training duration prediction model, including:

[0133] A log acquisition module 701, configured to acquire historical model training logs, where the historical model training logs include the attribute information and execution duration of the historically executed model training tasks;

[0134] A duration acquisition module 702, configured to input the attribute information into the training duration prediction model to obtain the training prediction duration of the historically executed model training task;

[0135] A loss acquisition module 703, configured to compare the training prediction duration and the execution duration to obtain the loss of the training duration prediction model;

[0136] A parameter adjustment module 704 is configured to adjust the parameters of the training duration prediction model based on the loss until the training is completed, so as to obtain a target training duration prediction model.

[0137] As can be seen from the above, the training device for the training duration prediction model provided by the present disclosure trains the training duration prediction model based on the model framework, model network, number of training dataset samples, batch size, number of iterations, training model type, and execution duration applied in the historical model training task, and can fully consider various influencing factors of the model training duration, thereby obtaining a more accurate prediction result and providing a more reliable duration reference.

[0138] In one embodiment of the present disclosure, the attribute information includes at least one of the model framework, model network, number of training dataset samples, batch size, number of iterations, and training model type applied in the historical model training task.

[0139] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information and other processing are all in compliance with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0140] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0141] Figure 8 FIG. shows a schematic block diagram of an exemplary electronic device 800 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0142] As Figure 8 shown, the device 800 includes a computing unit 801 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0143] Multiple components in device 800 are connected to I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a disk, an optical disc, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0144] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above, such as the model training control method and the training method of the training duration prediction model. For example, in some embodiments, the model training control method and the training method of the training duration prediction model can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the model training control method and the training method of the training duration prediction model described above can be executed. Alternatively, in other embodiments, the computing unit 801 can be configured to execute the model training control method and the training method of the training duration prediction model in any other suitable manner (e.g., by means of firmware).

[0145] The various embodiments of the systems and technologies described above in this article can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on a chip (SOC) systems, complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, where the programmable processor can be a special or general programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0146] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.

[0147] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0148] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0149] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected with each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0150] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, can also be a server of a distributed system, or a server incorporating a blockchain.

[0151] It should be understood that various forms of the flow shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution disclosed in this disclosure can be achieved, and no limitation is imposed herein.

[0152] The above specific implementation manners do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A model training control method, comprising: Obtaining attribute information of a first model training task initiated by a user, where the attribute information includes a model attribute corresponding to the first model training task; Obtaining a first resource type required for the first model training task; Obtaining a remaining training duration corresponding to the first resource type according to a preset correspondence between the resource type and the remaining training duration of the user; Inputting the attribute information into a training duration prediction model to obtain a predicted duration required for the first model training task; Based on the predicted duration and the remaining training duration, determining whether to allow the execution of the first model training task; In the case of determining that the execution of the first model training task is allowed, executing the first model training task according to the attribute information; Obtaining an execution duration of the first model training task; Updating the remaining training duration corresponding to the first resource type based on the execution duration; In the case where the remaining training duration corresponding to the first resource type is exhausted, sending an instruction to the user to stop a second model training task, where the second model training task is a model training task whose required resource type is the first resource type.

2. The method according to claim 1, further comprising: In the case of determining that the execution of the first model training task is not allowed, sending an instruction to the user indicating that the initiation of the first model training task fails.

3. The method according to claim 1, wherein, The attribute information includes at least one of a model framework, a model network, a number of training dataset samples, a batch size, a number of iteration rounds, and a training model type required for the first model training task.

4. The method according to claim 1, wherein The first model training task is used for image classification.

5. The method according to any one of claims 1-4, after the execution of the first model training task is completed, further comprising: Collecting training logs of the first model training task and storing them as historical model training logs, where the training logs include the attribute information and the execution duration of the first model training task.

6. The method according to claim 5, further comprising: Training and updating the training duration prediction model based on the historical model training logs.

7. A training method for a training duration prediction model, comprising: Obtaining historical model training logs, where the historical model training logs include attribute information and execution duration of a historically executed model training task; Inputting the attribute information into a training duration prediction model to obtain a training predicted duration of the historically executed model training task; Comparing the training predicted duration and the execution duration to obtain a loss of the training duration prediction model; Adjusting parameters of the training duration prediction model based on the loss until the training is completed to obtain a target training duration prediction model, where the target training duration prediction model is applied to the model training control method described in any one of claims 1-6.

8. The method according to claim 7, wherein, The attribute information includes at least one of a model framework, a model network, a number of training dataset samples, a batch size, a number of iteration rounds, and a training model type applied to the historical model training task.

9. A model training control device, comprising: An information acquisition module, configured to acquire the attribute information of a first model training task initiated by a user, where the attribute information includes the model attribute corresponding to the first model training task; acquire the first resource type required for the first model training task; and obtain the remaining training duration corresponding to the first resource type according to a preset correspondence between the resource type and the remaining training duration of the user. A duration prediction module, configured to input the attribute information into a training duration prediction model to obtain the predicted duration required for the first model training task. An execution judgment module, configured to judge whether to allow the execution of the first model training task based on the predicted duration and the remaining training duration. A task execution module, configured to execute the first model training task according to the attribute information when it is determined that the execution of the first model training task is allowed. An execution duration acquisition module, configured to acquire the execution duration of the first model training task. A duration update module, configured to update the remaining training duration corresponding to the first resource type based on the execution duration. A stop instruction sending module, configured to send an instruction indicating to stop a second model training task to the user when the remaining training duration corresponding to the first resource type is exhausted, where the second model training task is a model training task whose required resource type is the first resource type.

10. The apparatus according to claim 9, further comprising: A failure instruction sending module, configured to send an instruction indicating that the initiation of the first model training task fails to the user when it is determined that the execution of the first model training task is not allowed.

11. The device according to claim 9, wherein, The attribute information includes at least one of a model framework, a model network, the number of training data set samples, a batch size, the number of iteration rounds, and a training model type required for the first model training task.

12. The apparatus according to claim 9, wherein, The first model training task is used for image classification.

13. The apparatus according to any one of claims 9-12, further comprising: A log storage module, configured to collect the training log of the first model training task and store it as a historical model training log, where the training log includes the attribute information and the execution duration of the first model training task.

14. The apparatus according to claim 13, further comprising: A model update module, configured to train and update the training duration prediction model based on the historical model training log.

15. A training apparatus for a training duration prediction model, comprising: A log acquisition module, configured to acquire a historical model training log, where the historical model training log includes the attribute information and the execution duration of a historically executed model training task. A duration obtaining module, configured to input the attribute information into a training duration prediction model to obtain the training predicted duration of the historically executed model training task. A loss obtaining module, configured to compare the training predicted duration and the execution duration to obtain the loss of the training duration prediction model. A parameter adjustment module, configured to adjust parameters of the training duration prediction model based on the loss until the training is completed, so as to obtain a target training duration prediction model, where the target training duration prediction model is applied to the model training control device described in any one of claims 9-14.

16. The device according to claim 15, wherein, The attribute information includes at least one of a model framework, a model network, the number of training data set samples, a batch size, the number of iterations, and a training machine type applied in a historical model training task.

17. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method described in any one of claims 1-8.

18. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method described in any one of claims 1-8.

19. A computer program product, comprising a computer program which, when executed by a processor, implements the method described in any one of claims 1-8.

Citation Information

Patent Citations

  • Resource management method, apparatus, and computer-readable storage medium

    CN109324890A

  • Big data resource processing method and device, terminal and storage medium

    CN111198767A