Model training resource allocation method and device, electronic equipment and readable storage medium
By dynamically allocating resource information during the training phase, the problem of low resource utilization in machine learning platforms is solved, achieving efficient resource utilization and cost optimization.
Patent Information
- Application Number
- CN202511054641.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-11-14
AI Technical Summary
The low resource utilization rate in existing machine learning platforms is mainly due to the failure of existing resource allocation methods to dynamically adjust according to the differences in resource requirements at each training stage, resulting in resource waste.
By receiving training tasks, the system determines the resource information required for each training stage and dynamically allocates central processing unit and memory resources based on this information. It also uses the training database and pre-trained models to match historical training tasks and adjusts resource allocation in real time by combining identification information and output logs to ensure that each stage receives appropriate resources.
It improved resource utilization, avoided resource waste, optimized task operation efficiency, and reduced operating costs.
Smart Images

Figure CN120950244A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a method, apparatus, electronic device, and readable storage medium for allocating model training resources. Background Technology
[0002] In existing machine learning platforms, the training process for artificial intelligence models typically includes multiple training stages, such as data collection, data preprocessing, model selection, model training, model validation, and tuning. Each training stage has different resource requirements for the Central Processing Unit (CPU) and memory.
[0003] Currently, in order to ensure the smooth completion of the entire training task, existing technologies usually allocate CPU and memory according to the resource requirements of the stage with the highest resource consumption in the model training phase. However, this setting will cause a certain degree of resource waste in the stage with low resource demand, resulting in the problem of low resource utilization. Summary of the Invention
[0004] This application provides a method, apparatus, electronic device, and readable storage medium for allocating model training resources to solve the problem of low resource utilization in the prior art.
[0005] To solve the above problems, this application is implemented as follows:
[0006] In a first aspect, embodiments of this application provide a method for allocating model training resources, the method comprising:
[0007] Receive training tasks, which are used to train the target artificial intelligence model. The training tasks include N training phases, where N is a positive integer.
[0008] N primary resource information points are determined, and each of the N primary resource information points corresponds one-to-one with a training stage. The primary resource information points indicate the training resources required for the corresponding training stage.
[0009] Each of the N training phases that perform the training task;
[0010] In each of the N training stages of the training task, corresponding training resources are allocated based on the first resource information corresponding to each training stage.
[0011] Optionally, the steps for determining N pieces of first resource information include:
[0012] Matching is performed on the training database based on the training task to determine at least one target historical training task. The training database includes multiple historical training tasks, and each historical training task includes at least one training phase.
[0013] Determine each second resource information corresponding to each training stage of at least one target historical training task, wherein the second resource information indicates the training resources required for the corresponding training stage in at least one target historical training task.
[0014] N first resource information are determined based on each second resource information.
[0015] Optionally, matching is performed in the training database based on the training task to determine at least one target historical training task, including:
[0016] Determine the model type and training requirements corresponding to the target artificial intelligence model;
[0017] Based on model type and training requirements, a matching process is performed in the training database to identify at least one target historical training task, the model type of at least one target historical training task matches the target artificial intelligence model, and the training requirements of at least one target historical training task match the training requirements of the target artificial intelligence model.
[0018] Optionally, the steps for each of the N training phases in performing the training task include:
[0019] Set the identification information corresponding to the current training phase;
[0020] Execute the current training phase;
[0021] Detect the identification information corresponding to the current training phase;
[0022] The current training stage is determined from among N training stages based on the identification information corresponding to the current training stage.
[0023] Optionally, the steps for each of the N training phases in performing the training task include:
[0024] Get the output logs of the current training phase in real time;
[0025] Match the current training stage in the output log based on at least one of the keywords, regular expressions, or specific formats corresponding to the current training stage, in order to determine the corresponding current training stage among N training stages.
[0026] Optionally, training resources include central processing unit (CPU) resources and memory resources. During each of the N training stages of the training task, based on the first resource information corresponding to the current training stage, the corresponding training resources are allocated, including:
[0027] Based on the current training phase, determine the corresponding primary resource information;
[0028] Based on the first resource information corresponding to the current training stage, determine the first value and the second value corresponding to the current training stage. The first value is the value of the central processing unit resources, and the second value is the value of the memory resources.
[0029] Adjust the upper limit of central processing unit resources to the first value, and adjust the upper limit of memory resources to the second value;
[0030] Based on the upper limits of CPU resources and memory resources, allocate corresponding CPU resources and memory resources for the current training phase.
[0031] Optionally, before each step of the N training phases of the training task, the method further includes:
[0032] Determine the upper limit resource information from N first resource information, where the upper limit resource information is the maximum value among the N first resource information;
[0033] Allocate at least one processing container for the training task based on the upper limit resource information.
[0034] Secondly, embodiments of this application also provide a model training resource allocation device, comprising:
[0035] The receiving module is used to receive training tasks, which are used to train the target artificial intelligence model. The training task includes N training stages, where N is a positive integer.
[0036] The determination module is used to determine N first resource information, which correspond one-to-one with N training stages. The first resource information indicates the training resources required for the corresponding training stage.
[0037] The execution module is used to execute each of the N training phases of the training task;
[0038] In each of the N training stages of the training task, corresponding training resources are allocated based on the first resource information corresponding to each training stage.
[0039] Thirdly, embodiments of this application also provide an electronic device, including: a memory, a processor, and a program stored in the memory and executable on the processor; the processor is configured to read the program in the memory to implement the steps of the method described in the first aspect above.
[0040] Fourthly, embodiments of this application also provide a readable storage medium for storing a program, which, when executed by a processor, implements the steps of the method described in the first aspect above.
[0041] Fifthly, embodiments of this application also provide a computer program product, which is stored in a storage medium and executed by at least one processor to implement the steps in the method described in the first aspect above. Attached Figure Description
[0042] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0043] Figure 1 One of the flowcharts illustrating a method for allocating model training resources provided in an embodiment of this application;
[0044] Figure 2 A second schematic flowchart illustrating a method for allocating model training resources according to an embodiment of this application;
[0045] Figure 3 A third schematic flowchart illustrating a method for allocating model training resources according to an embodiment of this application;
[0046] Figure 4 A fourth flowchart illustrating a method for allocating model training resources as provided in an embodiment of this application;
[0047] Figure 5 Fifth flowchart illustrating a method for allocating model training resources according to an embodiment of this application;
[0048] Figure 6 A flowchart illustrating a method for allocating model training resources according to an embodiment of this application is shown in Figure 6.
[0049] Figure 7 The seventh flowchart illustrates a method for allocating model training resources according to an embodiment of this application.
[0050] Figure 8 This is the eighth flowchart illustrating a method for allocating model training resources according to an embodiment of this application.
[0051] Figure 9 A schematic diagram of a model training resource allocation device provided in an embodiment of this application;
[0052] Figure 10 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0053] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0054] The terms "first," "second," etc., used in the embodiments of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or devices. Additionally, the use of "and / or" in this application indicates at least one of the connected objects, such as A and / or B and / or C, representing seven possibilities: including A alone, B alone, C alone, and the presence of both A and B, both B and C, both A and C, and the presence of A, B, and C.
[0055] See Figure 1 , Figure 1 This is one of the flowcharts illustrating the method for allocating model training resources provided in the embodiments of this application. Figure 1 The method for allocating model training resources shown can be executed by a computer, such as a server. Figure 1 As shown, the method for allocating model training resources may include the following steps:
[0056] Step 101: Receive the training task. The training task is used to train the target artificial intelligence model. The training task includes N training stages, where N is a positive integer.
[0057] In this embodiment, the training task can be a model training request input by the user. The training task is used to perform a specified training on the target artificial intelligence model or to achieve a specified training effect. In this embodiment, no specific limitation is made.
[0058] It should be noted that, in this embodiment, the training process of the target artificial intelligence model includes N training stages, where N is a positive integer. For example, the training stages may include data collection, data preprocessing, model selection, model training, model validation, and model tuning. The training resources required for each training stage may be the same or different; for example, training resources may include CPU utilization and memory usage. The N training stages can be performed simultaneously or sequentially, without specific limitations in this embodiment. Specifically, the server in this application can be illustrated using Kubernetes as an example. Kubernetes is an open-source system for container cluster management, providing mechanisms for application deployment, maintenance, and expansion. When a user needs to train the target artificial intelligence model, the training task needs to be input into Kubernetes for processing, ultimately resulting in the trained target artificial intelligence model.
[0059] Step 102: Determine N first resource information items. The N first resource information items correspond one-to-one with the N training stages. The first resource information items indicate the training resources required for the corresponding training stage.
[0060] In this embodiment, the training process of the target artificial intelligence model includes N training stages. Each training stage corresponds to a first resource information, which represents the training resources required for the corresponding training stage, such as the CPU and memory resources required for a certain training stage. It should be noted that the first resource information is generally different for different training stages because their training content is different.
[0061] Specifically, N first resource information items can be determined based on historical training data. For example, if the target artificial intelligence model is an image recognition model, then the completed image recognition model records can be queried in the server, and the training resources used for the N training stages of the image recognition model can be extracted from these records. For example, the N resource information items corresponding to these N training stages can be determined as N first resource information items. It should be noted that in this embodiment, the historical training data must include the same training stages as the target artificial intelligence model training stages in this embodiment, that is, N represents the same number, and the model trained in the historical training data must be the same type as the target artificial intelligence model in this embodiment; otherwise, the corresponding N first resource information items cannot be accurately determined.
[0062] Step 103: Perform each of the N training phases of the training task;
[0063] In each of the N training stages of the training task, corresponding training resources are allocated based on the first resource information corresponding to each training stage.
[0064] In this embodiment, after acquiring N pieces of first resource information, each of the N training stages of the training task is executed to train the target artificial intelligence model. Specifically, during the training process, it is necessary to determine which of the N training stages the current training stage belongs to, in order to determine the resource information corresponding to the current training stage.
[0065] It should be noted that in this embodiment, the current training stage can be accurately determined by analyzing relevant data. For example, data can be analyzed at preset time intervals. When data changes, it indicates a change in the training stage, thus identifying the current training stage. The preset time interval can be set by the user to ensure timely awareness of data changes. Alternatively, the time period required for each training stage can be determined based on historical training information or user input. After resources are allocated to the first training stage and it begins, data changes are periodically monitored until the execution time reaches the minimum time period corresponding to that stage to determine whether to proceed to the second training stage, and so on. This allows for data change monitoring only when the current stage is nearing completion, saving resources. For instance, each training stage typically contains training data, and the corresponding training data is generally different across different training stages. By acquiring and analyzing the training data of the current training stage, the current training stage can be determined.
[0066] In this embodiment, after determining the current training stage, the first resource information corresponding to the current training stage can be accurately determined by the one-to-one correspondence between N training stages and N first resource information. Therefore, the corresponding training resources are allocated to the current training stage according to the training resources indicated by the first resource information.
[0067] During the training process of the target artificial intelligence model, the above method can be used to determine the first resource information corresponding to each training stage in each of the N training stages. This allows for the allocation of appropriate training resources to each training stage, ensuring that each training stage can use sufficient training resources while avoiding waste of training resources.
[0068] This application receives a training task comprising N training stages and determines N first resource information corresponding to the N training stages. Thus, when executing each training stage, resource allocation is performed based on the first resource information corresponding to each training stage, thereby improving resource utilization by allocating training resources that match the training stage in different training stages.
[0069] In some feasible implementations, optionally, step 102 involves determining N first resource information items, where each of the N first resource information items corresponds one-to-one with one of the N training stages. The first resource information items indicate the training resources required for the corresponding training stage, including:
[0070] Step 1021: Match the training task in the training database to determine at least one target historical training task. The training database includes multiple historical training tasks, and each historical training task includes at least one training phase.
[0071] Step 1022: Determine the second resource information corresponding to each training stage of at least one target historical training task, wherein the second resource information indicates the training resources required for the corresponding training stage in at least one target historical training task.
[0072] Step 1023: Determine N first resource information based on each second resource information.
[0073] In this embodiment, as Figure 2 As shown, Figure 2 This is the second flowchart illustrating the model training resource allocation method provided in this application embodiment. When determining the N first resource information corresponding to N training stages, matching can be performed in a pre-set training database to determine at least one target historical training task. Specifically, the training database includes multiple historical training tasks, i.e., training tasks and training processes of multiple different models. By matching the training tasks in the training database, at least one target historical training task matching the training task can be determined. The target historical training task includes at least one training stage identical to the received training task, and the model type trained by the target historical training task is the same as the model type of the target human intelligence model in this embodiment. It should be noted that the target historical training task can be a part of the training stages of a complete training task, and since the target historical training task matches the received training task, resource information from the historical task can be obtained as resource information for a part of the received training task. That is, multiple target historical training tasks can be combined to obtain all training stages of the training task.
[0074] After identifying at least one target historical training task, at least one second resource information corresponding to at least one training stage for the target historical training task is determined. This second resource information indicates the training resources required by the training stage corresponding to the target historical training task. Since the target historical training task has the same model type as the received training task and has at least one identical training stage, at least one second resource information can be used as a recommendation value to generate N first resource information items.
[0075] It should be noted that when obtaining recommended resource information allocation through at least one target historical training task, it can be determined by the matching degree between at least one target historical training task and the training task. For example, the historical training task with the highest matching degree can be selected as the target historical training task. If this target historical training task has the same model type and includes the same training stages as the training model corresponding to the received training task, then the second resource information corresponding to each training stage of this target historical training task can be used as the first resource information corresponding to each training stage of the received training task. In some embodiments, if multiple historical training tasks with the same model type and training stages as the training model corresponding to the received training task are matched, all of these historical training tasks can be determined as target historical training tasks, and the average of the resource values in the second resource information required for each training stage of the multiple target historical training tasks can be used as the recommended first resource information for each stage of the received training task. In other embodiments, historical training tasks can be filtered according to the model type of the received training task to obtain historical training tasks with the same model type as the received training task as target historical training tasks. Then, each of the N training stages corresponding to the received training task is matched with the filtered target historical training tasks in turn to obtain at least one target historical training task including the training stage. The average value of the resource values in the second resource information corresponding to the stage in the target historical training task including the training stage is used as the first resource information of the stage in the received training task, and so on to obtain N first resource information corresponding to N training stages. In other embodiments, historical training tasks can be filtered according to the model type of the received training task to obtain historical training tasks with the same model type as the received training task as target historical training tasks. Then, the weight values of each target historical training task are set according to the degree of matching between the training stages included in the target historical training task and the N training stages included in the received training task. The more training stages that the target historical training task and the received training task have in common, the larger the weight value. For each of the N training stages, the first resource information corresponding to that stage in the received training task is obtained by using the second resource information of that stage in the target historical training task that includes that training stage and its corresponding weight value.
[0076] In this embodiment, by using a pre-set training database, at least one target historical training task that matches the training task can be accurately determined. Then, at least one second resource information corresponding to at least one target historical training task can be used to determine N first resource information, thereby achieving the effect of accurately generating N first resource information. That is, by quickly and accurately determining the resource information corresponding to each training stage, the efficiency of resource allocation and resource utilization are improved.
[0077] Optionally, step 1021 involves matching the training task against the training database to determine at least one target historical training task. The training database includes multiple historical training tasks, and each historical training task includes at least one training phase, including:
[0078] Step 10211: Determine the model type and training requirements corresponding to the target artificial intelligence model.
[0079] Step 10212: Match the model type and training requirements in the training database to determine at least one target historical training task, the model type of at least one target historical training task is matched with the target artificial intelligence model, and the training requirements of at least one target historical training task are matched with the training requirements of the target artificial intelligence model.
[0080] In other words, based on the above embodiments, this embodiment not only matches the model types corresponding to the historical training tasks and the received training tasks, but also matches the training requirements corresponding to both, so as to obtain more similar historical training tasks.
[0081] In this embodiment, as Figure 3 As shown, Figure 3 This is the third flowchart illustrating the model training resource allocation method provided in this application embodiment. Based on the training task, the model type and training requirements corresponding to the target artificial intelligence model can be determined. The model type indicates the type of the target artificial intelligence model, i.e., the specific function of the target artificial intelligence model, such as an image recognition model, a model training resource allocation model, a text classification model, etc. The training requirements indicate the training objectives to be achieved when training the target artificial intelligence model. These objectives may include, for example, the required training accuracy of the target artificial intelligence model, or training using a set number of training samples.
[0082] By identifying the model type and training requirements corresponding to the target artificial intelligence model, a matching process is performed in the training database based on the model type and training requirements to identify at least one target historical training task that has the same or similar model type and training requirements as the target artificial intelligence model.
[0083] It should be further explained that during the matching process, when the training requirements of all historical training tasks in the training database differ from the training requirements of the target artificial intelligence model, it is necessary to identify target historical training tasks with similar training requirements, such as target historical training tasks with the smallest difference in training requirements or with differences within a certain range. For example, the target historical training task can be identified by setting a difference function or loss function for calculation.
[0084] In this embodiment, the target artificial intelligence model is matched in the training database according to the model type and training requirements, thereby improving the matching accuracy of the model.
[0085] Optionally, step 102 involves determining N pieces of first resource information, each corresponding one-to-one with one of the N training stages. The first resource information indicates the training resources required for the corresponding training stage, including:
[0086] Step 102': Input the training task into the pre-trained artificial intelligence model for matching, and output N first resource information.
[0087] In this embodiment, as Figure 4 As shown, Figure 4 This is the fourth flowchart of the model training resource allocation method provided in this application embodiment. The pre-trained artificial intelligence model can be a pre-trained deep learning model. Specifically, the pre-trained artificial intelligence model is used to output N first resource information that have the same training stage as the training task according to the input training task. That is, the matching process is entirely executed by the pre-trained artificial intelligence model.
[0088] It should be noted that the training samples for the pre-trained AI model can be information related to various AI models, and the labels corresponding to these training samples can be resource information used in multiple training stages for each AI model. After the pre-trained AI model has been trained, inputting a training task can quickly match N primary resource information corresponding to N training stages, thereby improving matching efficiency.
[0089] Optionally, step 103, each of the N training phases of the training task includes the following steps:
[0090] Step 1031: Set the identification information corresponding to the current training phase.
[0091] Step 1032: Execute the current training phase.
[0092] Step 1033: Detect the identification information corresponding to the current training phase.
[0093] Step 1034: Determine the corresponding current training stage among the N training stages based on the identification information corresponding to the current training stage.
[0094] In this embodiment, as Figure 5 As shown, Figure 5 This is the fifth flowchart illustrating the method for allocating model training resources provided in this application embodiment. Before training each stage of the target artificial intelligence model, it is necessary to set an identifier for each training stage to generate corresponding identifier information. Through this identifier information, the current training stage of the target artificial intelligence model can be quickly determined.
[0095] Specifically, the identifier information corresponding to the current training stage can be in the form of a string, etc., and this embodiment does not impose specific limitations. What needs to be limited is that the identifier information corresponding to each training stage needs to be set differently, so as to distinguish between different training stages. For example, N training stages would require N identifier information settings. For example, before the training task starts, the environment variable TRAINING_STAGE is set, such as export TRAINING_STAGE = DATA_PREPROCESSING. By setting different values for DATA, different training stages can be represented. The environment variable identifies the current training stage, which facilitates the system's identification and resource adjustment.
[0096] It should be noted that environment variables are variables used by the operating system to store configuration information required for system and application runtime. Environment variables exist in the form of key-value pairs and can affect program behavior and the system's operating environment. In this embodiment, environment variables are generally set using the `export` command. In other embodiments, the identification information may take other forms, which are not specifically limited here.
[0097] During the training process of the target artificial intelligence model, the current training stage of the target artificial intelligence model is determined by detecting the identification information corresponding to the current training stage and matching it among N identification information.
[0098] By identifying different training stages using labeling information, the current training stage of the model can be quickly determined during the training process, improving the efficiency of identifying the current training stage and thus improving the efficiency of training resource allocation.
[0099] Optionally, step 103, each of the N training phases of the training task includes the following steps:
[0100] Step 1035: Obtain the output logs of the current training phase in real time.
[0101] Step 1036: Match the current training stage in the output log based on at least one of the keywords, regular expressions, and specific formats corresponding to the current training stage, so as to determine the corresponding current training stage among N training stages.
[0102] In this embodiment, as Figure 6 As shown, Figure 6 This is the sixth flowchart illustrating the model training resource allocation method provided in this application embodiment. During the training process of each training task, a training output log is output. Specifically, the output log in the training model (usually called the training log or training output) refers to the recorded information generated during model training. Generally, the output log can include training progress, loss value, accuracy, learning rate, etc. In this embodiment, the output log is used to represent the current training content and the current training resource utilization, that is, the training content corresponding to the current training stage and the resource usage of the current training.
[0103] After obtaining the output logs, the current training content and resource utilization are determined. Analysis is then performed to identify which of the N training stages the current training stage belongs to. Specifically, matching can be performed within the current training stage based on at least one of the following: corresponding keywords, regular expressions, or specific formats. The corresponding keywords are those specific to the current training stage and do not exist in other training stages. Regular expressions are patterns used to match character combinations in a string; in this embodiment, the regular expressions differ for different training stages. Specific formats refer to special formats included in the current training stage but not in other training stages.
[0104] For example, the output logs may also include specific processes or threads started at different stages of the training process, so that the current training stage of the target artificial intelligence model can be determined based on the output logs.
[0105] By identifying different training stages through output logs, the current training stage of the model can be quickly determined by detecting the corresponding output logs during the model training process. This improves the efficiency of identifying the current training stage and thus improves the efficiency of training resource allocation.
[0106] Optionally, training resources include central processing unit resources and memory resources. Step 103: During each of the N training stages of the training task, based on the first resource information corresponding to the current training stage, allocate the corresponding training resources, including:
[0107] Step 1031': Based on the current training phase, determine the corresponding first resource information.
[0108] Step 1032': Based on the first resource information corresponding to the current training stage, determine the first value and the second value corresponding to the current training stage. The first value is the value of the central processing unit resources, and the second value is the value of the memory resources.
[0109] Step 1033': Adjust the upper limit of the central processing unit resources to the first value, and adjust the upper limit of the memory resources to the second value.
[0110] Step 1034': Based on the upper limits of CPU resources and memory resources, allocate corresponding CPU resources and memory resources for the current training phase.
[0111] In this embodiment, as Figure 7 As shown, Figure 7 This is the seventh flowchart illustrating the model training resource allocation method provided in this application embodiment. After determining the current training stage, corresponding first resource information can be determined based on the current training stage. Then, based on the first resource information corresponding to the current training stage, a first value and a second value corresponding to the current training stage are determined. Specifically, the first value is the value of the central processing unit resource, i.e., CPU utilization. The second value is the value of the memory resource, i.e., memory occupancy. It should be noted that CPU utilization and memory occupancy are the most important aspects of the training resources. Other training resources may be included in other embodiments, but this embodiment does not impose specific limitations.
[0112] In this embodiment, control groups (cgroups) can be used to adjust the upper limit of CPU resources to a first value and the upper limit of memory resources to a second value. Control groups are a feature of the Linux kernel that allows processes to be organized into hierarchical structures and their resource usage to be restricted, monitored, and isolated. Cgroups can control the CPU, memory, disk I / O, and network bandwidth used by processes, helping system administrators better manage and optimize system resources. In this embodiment, cgroups can be used to set the upper limits of CPU and memory resources on the server.
[0113] Specifically, after determining the current training phase, the upper limits of CPU resources and memory resources are set to the first and second values corresponding to the current training phase through cgroups, thereby allocating the CPU resources and memory resources in the server corresponding to the first and second values to the current training phase.
[0114] In this embodiment, after determining the first resource information corresponding to the current training stage, the control group adjusts the pre-set resource limit of the server so that each training stage is allocated appropriate training resources, avoiding resource waste and insufficient resources, thereby improving resource utilization.
[0115] Optionally, before step 103, which involves performing each of the N training phases of the training task, the method further includes:
[0116] Step 104: Determine the upper limit resource information from the N first resource information. The upper limit resource information is the maximum value among the N first resource information.
[0117] Step 105: Allocate at least one processing container for the training task based on the upper limit resource information.
[0118] In this embodiment, as Figure 8 As shown, Figure 8 This is the eighth flowchart illustrating the model training resource allocation method provided in this application embodiment. After receiving a training task, the training stage with the highest demand for training resources among N training stages can be determined through N first resource information. Based on this upper limit resource information, at least one processing container (Pod) is allocated to the training task. Taking Kubernetes as an example, a Pod is the smallest deployment unit in Kubernetes, consisting of one or more containers with shared network and storage. In this embodiment, the target artificial intelligence model is trained using Pods to complete the training task.
[0119] In this embodiment, by combining the above implementation methods, the resource limits of the pod can be dynamically adjusted, making full use of system resources and improving resource utilization. Furthermore, by dynamically adjusting the resource limits, the resource waste caused by setting excessively high resource limits for peak resource demands is avoided, thereby achieving refined resource management and saving computing resources and energy consumption.
[0120] It should be noted that the method provided in this embodiment can also be applied to machine learning platforms and cloud computing platforms, aiming to optimize task running efficiency, reduce operating costs, and improve user satisfaction.
[0121] This application receives a training task comprising N training stages and determines N first resource information corresponding to the N training stages. Thus, when executing each training stage, resource allocation is performed based on the first resource information corresponding to each training stage, thereby improving resource utilization by allocating training resources that match the training stage in different training stages.
[0122] See Figure 9 , Figure 9This is a structural diagram of the model training resource allocation device provided in an embodiment of this application. For example... Figure 9 As shown, the model training resource allocation device 900 includes:
[0123] The receiving module 910 is used to receive training tasks, which are used to train the target artificial intelligence model. The training tasks include N training stages, where N is a positive integer.
[0124] The determination module 920 is used to determine N first resource information, which correspond one-to-one with N training stages. The first resource information indicates the training resources required for the corresponding training stage.
[0125] Execution module 930 is used to execute each of the N training phases of the training task;
[0126] In each of the N training stages of the training task, corresponding training resources are allocated based on the first resource information corresponding to each training stage.
[0127] Optionally, the determining module 920 includes:
[0128] The first matching submodule is used to perform matching in the training database based on the training task to determine at least one target historical training task. The training database includes multiple historical training tasks, and each historical training task includes at least one training phase.
[0129] The first determining submodule is used to determine each second resource information corresponding to each training stage of at least one target historical training task, wherein the second resource information indicates the training resources required for the corresponding training stage in at least one target historical training task.
[0130] The second determination submodule is used to determine N first resource information based on each second resource information.
[0131] Optionally, the first matching submodule includes:
[0132] The determination unit is used to determine the model type and training requirements corresponding to the target artificial intelligence model;
[0133] The matching unit is used to perform matching in the training database based on model type and training requirements to determine at least one target historical training task, the model type of at least one target historical training task is matched with the target artificial intelligence model, and the training requirements of at least one target historical training task are matched with the training requirements of the target artificial intelligence model.
[0134] Optionally, execution module 930 includes:
[0135] The settings submodule is used to set the identification information corresponding to the current training phase;
[0136] The execution submodule is used to execute the current training phase;
[0137] The detection submodule is used to detect the identification information corresponding to the current training phase;
[0138] The third determination submodule is used to determine the corresponding current training stage among N training stages based on the identification information corresponding to the current training stage.
[0139] Optionally, execution module 930 includes:
[0140] The `get` submodule is used to obtain the output logs of the current training phase in real time.
[0141] The second matching submodule is used to match the output log based on at least one of the keywords, regular expressions, and specific formats corresponding to the current training stage, so as to determine the corresponding current training stage among N training stages.
[0142] Optionally, training resources include central processing unit resources and memory resources, and during each of the N training phases of the training task, they also include:
[0143] The fourth determination submodule is used to determine the corresponding first resource information based on the current training stage;
[0144] The fifth determination submodule is used to determine the first value and the second value corresponding to the current training stage based on the first resource information corresponding to the current training stage. The first value is the value of the central processing unit resources, and the second value is the value of the memory resources.
[0145] The adjustment module is used to adjust the upper limit of the central processing unit resources to the first value and the upper limit of the memory resources to the second value;
[0146] The allocation submodule is used to allocate corresponding CPU and memory resources for the current training phase based on the upper limits of CPU and memory resources.
[0147] Optional, also includes:
[0148] The resource determination module is used to determine the upper limit resource information from N first resource information, where the upper limit resource information is the maximum value among the N first resource information;
[0149] The container allocation module is used to allocate at least one processing container to the training task based on the upper limit resource information.
[0150] This application receives a training task comprising N training stages and determines N first resource information corresponding to the N training stages. Thus, when executing each training stage, resource allocation is performed based on the first resource information corresponding to each training stage, thereby improving resource utilization by allocating training resources that match the training stage in different training stages.
[0151] This application also provides an electronic device. Please refer to [link to relevant documentation]. Figure 10 The electronic device may include a processor 1001, a memory 1002, and a program 10021 stored in the memory 1002 and executable on the processor 1001.
[0152] When program 10021 is executed by processor 1001, it can achieve the following: Figure 1 Any step in the corresponding method embodiment:
[0153] Receive training tasks, which are used to train the target artificial intelligence model. The training tasks include N training phases, where N is a positive integer.
[0154] N primary resource information points are determined, and each of the N primary resource information points corresponds one-to-one with a training stage. The primary resource information points indicate the training resources required for the corresponding training stage.
[0155] Each of the N training phases that perform the training task;
[0156] In each of the N training stages of the training task, corresponding training resources are allocated based on the first resource information corresponding to each training stage.
[0157] Optionally, the steps for determining N pieces of first resource information include:
[0158] Matching is performed on the training database based on the training task to determine at least one target historical training task. The training database includes multiple historical training tasks, and each historical training task includes at least one training phase.
[0159] Determine each second resource information corresponding to each training stage of at least one target historical training task, wherein the second resource information indicates the training resources required for the corresponding training stage in at least one target historical training task.
[0160] N first resource information are determined based on each second resource information.
[0161] Optionally, matching is performed in the training database based on the training task to determine at least one target historical training task, including:
[0162] Determine the model type and training requirements corresponding to the target artificial intelligence model;
[0163] Based on model type and training requirements, a matching process is performed in the training database to identify at least one target historical training task, the model type of at least one target historical training task matches the target artificial intelligence model, and the training requirements of at least one target historical training task match the training requirements of the target artificial intelligence model.
[0164] Optionally, the steps for each of the N training phases in performing the training task include:
[0165] Set the identification information corresponding to the current training phase;
[0166] Execute the current training phase;
[0167] Detect the identification information corresponding to the current training phase;
[0168] The current training stage is determined from among N training stages based on the identification information corresponding to the current training stage.
[0169] Optionally, the steps for each of the N training phases in performing the training task include:
[0170] Get the output logs of the current training phase in real time;
[0171] Match the current training stage in the output log based on at least one of the keywords, regular expressions, or specific formats corresponding to the current training stage, in order to determine the corresponding current training stage among N training stages.
[0172] Optionally, training resources include central processing unit (CPU) resources and memory resources. During each of the N training stages of the training task, based on the first resource information corresponding to the current training stage, the corresponding training resources are allocated, including:
[0173] Based on the current training phase, determine the corresponding primary resource information;
[0174] Based on the first resource information corresponding to the current training stage, determine the first value and the second value corresponding to the current training stage. The first value is the value of the central processing unit resources, and the second value is the value of the memory resources.
[0175] Adjust the upper limit of central processing unit resources to the first value, and adjust the upper limit of memory resources to the second value;
[0176] Based on the upper limits of CPU resources and memory resources, allocate corresponding CPU resources and memory resources for the current training phase.
[0177] Optionally, before each step of the N training phases of the training task, the method further includes:
[0178] Determine the upper limit resource information from N first resource information, where the upper limit resource information is the maximum value among the N first resource information;
[0179] Allocate at least one processing container for the training task based on the upper limit resource information.
[0180] This application receives a training task comprising N training stages and determines N first resource information corresponding to the N training stages. Thus, when executing each training stage, resource allocation is performed based on the first resource information corresponding to each training stage, thereby improving resource utilization by allocating training resources that match the training stage in different training stages.
[0181] This application also provides a computer-readable storage medium storing a computer program. When executed by a processor, this computer program implements the various processes of the above-described model training resource allocation method embodiments and achieves the same technical effects. To avoid repetition, it will not be described again here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0182] This application also provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above-described model training resource allocation method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0183] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0184] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of the various embodiments of this application.
[0185] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A method for allocating model training resources, characterized in that, The method includes: Receive a training task, which is used to train a target artificial intelligence model. The training task includes N training stages, where N is a positive integer. N first resource information are determined, and the N first resource information corresponds one-to-one with the N training stages. The first resource information indicates the training resources required for the corresponding training stage. Perform each of the N training phases of the training task; In each of the N training stages of the training task, corresponding training resources are allocated based on the first resource information corresponding to each training stage.
2. The method according to claim 1, characterized in that, The step of determining N pieces of first resource information includes: Based on the training task, a matching is performed in the training database to determine at least one target historical training task. The training database includes multiple historical training tasks, and each historical training task includes at least one training stage. Determine each second resource information corresponding to each training stage of the at least one target historical training task, wherein the second resource information indicates the training resources required for the corresponding training stage in the at least one target historical training task. The N first resource information are determined based on each of the second resource information.
3. The method according to claim 2, characterized in that, The step of matching the training task in the training database to determine at least one target historical training task includes: Determine the model type and training requirements corresponding to the target artificial intelligence model; Based on the model type and the training requirements, a matching is performed in the training database to determine at least one target historical training task. The model type of the at least one target historical training task matches the target artificial intelligence model, and the training requirements of the at least one target historical training task match the training requirements of the target artificial intelligence model.
4. The method according to claim 1, characterized in that, The steps for each of the N training phases in performing the training task include: Set the identification information corresponding to the current training phase; Execute the current training phase; Detect the identification information corresponding to the current training phase; The current training stage is determined from among the N training stages based on the identification information corresponding to the current training stage.
5. The method according to claim 1, characterized in that, The steps for each of the N training phases in performing the training task include: Get the output logs of the current training phase in real time; Matching is performed in the output log based on at least one of the keywords, regular expressions, and specific formats corresponding to the current training stage, in order to determine the corresponding current training stage among the N training stages.
6. The method according to claim 1, characterized in that, The training resources include central processing unit resources and memory resources. During each of the N training stages of the training task, based on the first resource information corresponding to the current training stage, the corresponding training resources are allocated, including: Based on the current training phase, determine the corresponding primary resource information; Based on the first resource information corresponding to the current training stage, determine the first value and the second value corresponding to the current training stage, wherein the first value is the value of the central processing unit resources and the second value is the value of the memory resources; The upper limit of the central processing unit resources is adjusted to the first value, and the upper limit of the memory resources is adjusted to the second value; Based on the upper limits of the central processing unit (CPU) resources and the memory resources, corresponding CPU and memory resources are allocated for the current training phase.
7. The method according to claim 1, characterized in that, Before the step of performing each of the N training phases of the training task, the method further includes: Determine the upper limit resource information from the N first resource information, wherein the upper limit resource information is the maximum value among the N first resource information; At least one processing container is allocated to the training task based on the upper limit resource information.
8. A device for allocating model training resources, characterized in that, The device includes: A receiving module is used to receive training tasks, which are used to train a target artificial intelligence model. The training tasks include N training stages, where N is a positive integer. The determination module is used to determine N first resource information, wherein the N first resource information corresponds one-to-one with the N training stages, and the first resource information indicates the training resources required for the corresponding training stage; An execution module is used to execute each of the N training stages of the training task; In each of the N training stages of the training task, corresponding training resources are allocated based on the first resource information corresponding to each training stage.
9. An electronic device, comprising: A memory, a processor, and a program stored in the memory and executable on the processor; characterized in that the processor is configured to read the program from the memory to implement the steps in the method for allocating model training resources as described in any one of claims 1 to 7.
10. A readable storage medium for storing a program, characterized in that, When the program is executed by the processor, it implements the steps in the method for allocating model training resources as described in any one of claims 1 to 7.