Model training method, device and equipment and computer readable storage medium

By cached sample data on multiple training nodes of the local training platform and performed training tasks, the low efficiency and safety of model training caused by the large amount of sample data in the field of intelligent driving are solved, and efficient and safe model training is achieved.

CN120339748APending Publication Date: 2025-07-18NINGBO LOTUS ROBOTICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510396231.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the field of intelligent driving, the existing technology has problems such as low model training efficiency, long data transmission time and data security caused by large sample data. Especially when relying on third-party platforms for model training, it cannot meet personalized needs and data privacy protection.

Method used

Using a local training platform, multiple training nodes are used to cache sample data. By generating training tasks and performing training tasks on nodes that cache sample data, avoiding data transmission and improving training efficiency and security.

Benefits of technology

By locally caching sample data, the data transmission time is reduced, the efficiency and security of model training are improved, personalized needs are met, and the dependence on third-party platforms is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339748A_ABST
    Figure CN120339748A_ABST
Patent Text Reader

Abstract

The invention relates to the field of machine learning, in particular to a model training method, device and equipment and a computer readable storage medium. The model training method is applied to a local training platform comprising a plurality of training nodes, and comprises the following steps: in response to a training instruction for a to-be-trained model, generating a training task, and determining sample data corresponding to the training task; and for the training task, at least under the condition that any training node caches the sample data corresponding to the training task, executing the training task by using the sample data by the any training node.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of machine learning, and particularly to a model training method, apparatus, device, and computer-readable storage medium. Background Art

[0002] In the related art, training a large data model can rely on an online third-party model training platform. In the entire model training process, it is necessary to upload the sample data set to be trained to the third-party platform and then configure the relevant parameters of the model to perform model training.

[0003] However, in the field of intelligent driving, the intelligent driving model to be trained requires a large amount of sensor data, resulting in a very large data volume of the sample data set to be trained (for example, it can reach the PB level). This makes the data transmission process of the sample data take too long and affects the efficiency of model training. And due to possible data security issues, the third-party platform usually does not save the sample data uploaded by users. Therefore, if the same user needs to use the same sample data set to train the model again, it is necessary to re-upload the sample data set. Summary of the Invention

[0004] To overcome the problems in the related art, the present disclosure provides a model training method, apparatus, device, and computer-readable storage medium, which can solve the above problems.

[0005] According to a first aspect of an embodiment of the present disclosure, there is provided a model training method, which is applied to a local training platform including multiple training nodes. The method includes:

[0006] Responding to a training instruction for a model to be trained, generating a training task, and determining the sample data corresponding to the training task;

[0007] For the training task, at least when the sample data corresponding to the training task is cached in any one of the training nodes, the training task is executed by the any one of the training nodes using the sample data.

[0008] According to a second aspect of an embodiment of the present disclosure, there is provided a model training apparatus, which is applied to a local training platform including multiple training nodes. The apparatus includes:

[0009] A response unit, configured to respond to a training instruction for a model to be trained, generate a training task, and determine the sample data corresponding to the training task;

[0010] An execution unit, configured to, for the training task, at least when the sample data corresponding to the training task is cached in any one of the training nodes, execute the training task by the any one of the training nodes using the sample data.

[0011] According to a third aspect of the embodiments of the present disclosure, an electronic device is provided, including: a processor and a memory;

[0012] The memory is used to store a computer program;

[0013] The processor is configured to execute the model training method as described in the first aspect by calling the computer program.

[0014] According to a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the model training method as described in the first aspect is implemented.

[0015] According to a fifth aspect of the embodiments of the present disclosure, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the method as described in the first aspect is implemented.

[0016] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects:

[0017] The present disclosure proposes a model training method, which is applied to a local training platform. In response to a training instruction for model training, a training task can be generated, and the sample data corresponding to the training task can be determined. Since it is a local training platform, multiple training nodes in this training platform can store the data set in the local storage space, and the data security is higher. For the generated training task, a node that caches the sample data required for the training task can be determined from multiple training nodes of the local training platform, and the training task is assigned to the determined node, so that the node uses the cached sample data to execute the training task.

[0018] Based on the method of the present disclosure, the training nodes of the local training platform are equipped with the condition of caching sample data. After the training task is generated, the training task can be preferentially assigned to the training node that caches the sample data required for the task, thereby saving the transmission time of the sample data and accelerating the efficiency of model training.

[0019] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The accompanying drawings herein are incorporated into the specification and constitute a part of the present disclosure, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure.

[0021] Figure 1 is a schematic flowchart of a model training method shown according to an exemplary embodiment of the present disclosure.

[0022] Figure 2It is a schematic flowchart of a model training method shown according to an exemplary embodiment of the present disclosure.

[0023] Figure 3 It is a schematic diagram of a method for allocating training tasks shown according to an exemplary embodiment of the present disclosure.

[0024] Figure 4 It is a block diagram of a model training device shown according to an exemplary embodiment of the present disclosure.

[0025] Figure 5 It is a schematic block diagram of a device for a model training device shown according to an exemplary embodiment of the present disclosure. Detailed implementation manners

[0026] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0027] The terms used in the present disclosure are only for the purpose of describing specific embodiments and are not intended to limit the present disclosure. The singular forms "a", "the", and "said" used in the present disclosure and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0028] It should be understood that although terms such as first, second, and third may be used in the present disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present disclosure, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".

[0029] With the popularization of big data technology, machine learning and artificial intelligence technologies have become the core components of some industries. To improve the training efficiency and performance of big data models, a model training platform can be used to manage and optimize the model training process.

[0030] In the field of intelligent driving, the perception of the surrounding environment by vehicles and the prediction algorithms all rely on big data models. Therefore, the training and tuning of big data models are relatively frequent in the version iteration process of intelligent driving strategies, and there are a large number of model training requirements.

[0031] Related technologies typically use a third - party training platform for model training. However, this approach has some drawbacks. For example, the cost is relatively high, there are data privacy issues, the model training is highly dependent on the third - party platform, it cannot meet personalized needs, and the data transmission takes a long time.

[0032] In the field of intelligent driving, model training relies on a large amount of sample data, such as sensor data generated by sensors. For example, the training of a perception model requires TB or even PB - level images and point cloud data, and these data are all generated by a local data annotation platform and need to be uploaded to a third - party training platform, which poses challenges to the data storage of the third - party training platform and incurs a relatively high transmission cost.

[0033] To solve the above - mentioned technical problems, the present disclosure proposes a model training method.

[0034] Figure 1 FIG. is a schematic flow chart of a model training method shown according to an embodiment of the present disclosure. The model training method is applied to a local training platform including multiple training nodes.

[0035] As Figure 1 shown, the model training method includes:

[0036] In step S101, in response to a training instruction for a model to be trained, a training task is generated, and the sample data corresponding to the training task is determined.

[0037] In step S102, for the training task, at least when the sample data corresponding to the training task is cached at any one of the training nodes, the training task is executed by the any one of the training nodes using the sample data.

[0038] In some embodiments, the local training platform may include multiple training nodes.

[0039] The local training platform may be a cluster composed of multiple training nodes. Each training node may include one or more processing units. Exemplarily, when the processing unit is a GPU (Graphics Processing Unit), the training node may be the physical machine where the GPU is located. Here, according to the actual situations such as the size, heat dissipation, and network resources of the physical machine, any physical machine may install only one GPU or may also install multiple GPUs, and this specification does not limit this.

[0040] In some embodiments, the local training platform may provide an interactive web interface for users.

[0041] Based on this web interface, users can intuitively view and adjust the steps of model training and clearly understand the execution process of each training step. The model training steps may include data preparation, model construction, parameter tuning, etc.

[0042] Based on the interactive web interface, it enables users to more easily and intuitively understand the entire model training process and helps users better monitor and manage the training process.

[0043] In some embodiments, in response to a training instruction for a model to be trained, a training task is generated and the sample data corresponding to the training task is determined.

[0044] Users can initiate a training instruction for a model to be trained based on a local training platform. In response to this training instruction, the local training platform can generate a training task, and the execution result of this training task is the model to be trained. The local training platform can determine the sample data corresponding to the training task, and the sample data is used for the training task to generate the corresponding model to be trained.

[0045] In some embodiments, the generated training task needs to be executed by a training node of the local training platform.

[0046] The training node can execute this training task to train a model to be trained that meets the requirements using the sample data.

[0047] However, since the local training platform includes multiple training nodes, it is necessary to allocate a suitable training node for each training task to improve the training efficiency of the model.

[0048] In some embodiments, for the training task, at least when the sample data corresponding to the training task is cached in any one training node, the any one training node uses the sample data to execute the training task.

[0049] Before executing the training task, the training node needs to first cache the sample data of the training task in the storage space of the training node.

[0050] Therefore, for the training task, the sample data corresponding to the training task can be determined first, and the training node whose storage space contains the sample data is used to process this training task. In this case, since the storage space of the training node contains the sample data, data transmission can be avoided and the required sample data can be directly indicated to the training node and the training task can be allocated for the training node to execute.

[0051] Based on this embodiment, the time consumed by data transmission can be saved, data transmission can be avoided, and thus the training efficiency can be improved.

[0052] In some embodiments, the training node executing the training task also needs to be in an idle state.

[0053] If a training node has been assigned a training task, even if the storage space of the training node contains sample data of another training task to be assigned, it is necessary to wait until the training node completes the training task being executed before assigning another training task to be assigned to the training node.

[0054] In some embodiments, the training node executing the training task also needs to consider the computing resource requirements.

[0055] It is necessary to assign the training task to the node for execution when the training node meets the computing resource requirements of the training task and caches the sample data.

[0056] The model training method proposed by the present disclosure relies on a local training platform including multiple training nodes for model training. Since the training nodes are local training nodes, sample data can be pre-stored in their storage spaces. When performing the model training task, the training task is preferentially executed by the training node that caches the sample data corresponding to the training task, thereby avoiding the transmission of sample data, saving the time for sample data transmission, and thus improving the model training efficiency.

[0057] In some embodiments, the local training platform can be associated with an internal data platform.

[0058] The local training platform can obtain sample data from the internal data platform and pre-cache the sample data in the storage spaces of each node.

[0059] Since the types of training tasks faced by the local training platform are relatively concentrated, there is an intersection in the sample data of some training tasks. In this case, the sample data of the training task can be cached in the storage space of the node for subsequent model training based on the same or similar sample data.

[0060] In some embodiments, the sample data includes sensor data.

[0061] The sensor data may include but is not limited to: point cloud data, labeled pictures.

[0062] When the local training platform is associated with the internal data platform, the user can manage the sample data in the internal data platform through the local training platform. For example, the sample data can be pre-processed, including screening, filtering, cropping, and classification. These pre-processing functions can help the user better prepare the sample data and improve the effect and accuracy of model training.

[0063] In some embodiments, the training node can use a distributed storage system to store the sample data.

[0064] The distributed storage system has a relatively fast storage speed and can meet the bandwidth and IOPS (Input / Output Operations Per Second) requirements for loading sample data. This disclosure can pre-load some sample data onto the training nodes, and have the training nodes execute the training tasks for the corresponding cached sample data, thereby alleviating the pressure on the central storage (Ceph) during the model training process and improving the speed of model training.

[0065] In some embodiments, the information of the sample data stored in the distributed storage system can be maintained and managed on the local training platform.

[0066] The information of the sample data includes but is not limited to: sample data ID, description information of the sample data, size of the sample data, version of the sample data, storage directory path of the sample data in the storage system, update time of the sample data, creation time of the sample data, etc.

[0067] Based on the maintenance and management of the information of the sample data, the local training platform can determine the information of the sample data cached by each training node, and then determine the nodes to execute the training tasks.

[0068] In some embodiments, after determining the training node to execute the training task, the local training platform can update the information of the sample data corresponding to the training task cached by the training node.

[0069] For example, the update time of the sample data can be updated, and / or part of the data content in the sample data can be updated to ensure the timeliness of the sample data cached in the training node.

[0070] In some embodiments, the local training platform may include a task scheduling module, and the task scheduling module is used to assign training tasks to target nodes, and the target nodes execute the training tasks.

[0071] The task scheduler can determine the sample data corresponding to the training task and the information of the sample data cached by each training node. If the ID of the sample data corresponding to the training task is the same as the ID of the sample data cached by any training node, it can be considered that the any training node caches the sample data required for the training task. The task scheduler can, when the any training node is idle, assign the training task to the any training node for the any training node to execute.

[0072] In some embodiments, the method further includes: in the case where no training node caches the sample data corresponding to the training task, determining a target node that meets the training requirements of the training task from the multiple training nodes, and sending the sample data corresponding to the training task to the target node to execute the training task.

[0073] If none of the training nodes caches the sample data required for the training task, it is necessary to determine a target node from multiple training nodes and send the sample data corresponding to the training task to the target node. The target node can cache the received sample data in the storage space and then execute the training task.

[0074] The following introduces the method for determining the target node in conjunction with the illustrations.

[0075] Figure 2 It is a schematic flowchart of a model training method shown according to an embodiment of the present disclosure.

[0076] In some embodiments, after generating a training task and determining the sample data corresponding to the training task, it is possible to determine whether there is a training node that caches the sample data.

[0077] As Figure 2 shown, if so, step S102 can be executed, and the training task is executed by the training node that caches the sample data; if not, step S201 can be executed.

[0078] In step S201, it is possible to determine whether there is a training node with sufficient remaining storage space to cache the sample data; if so, execute step S202A, if not, execute step S202B;

[0079] Step S202A, determine the target node based on the remaining storage space;

[0080] Step S202B, clear the cached sample data and then determine the target node.

[0081] In some embodiments, the target node includes a node with sufficient remaining storage space to cache the sample data corresponding to the training task.

[0082] Step S202A may specifically include determining the node with sufficient remaining storage space to cache the sample data corresponding to the training task as the target node.

[0083] If the remaining storage space of the training node is sufficient to cache the sample data, it may not be necessary to perform additional processing on the training node. Send the sample data corresponding to the training task to the target node, and then the target node executes the training task.

[0084] In some embodiments, when there are multiple nodes with sufficient remaining storage space to cache the sample data corresponding to the training task, determine the node with the largest remaining storage space as the target node.

[0085] In the case where the remaining storage space of multiple training nodes is sufficient to cache the sample data corresponding to the training task, the node with the largest remaining storage space can be determined as the target node.

[0086] This embodiment can balance the sizes of the sample data in the storage spaces of each training node, avoiding that the storage spaces of some training nodes cache more sample data while the remaining storage spaces of some training nodes are larger. The more sample data cached in the storage space of a training node, the easier it is to match the sample data required for subsequent training tasks, and thus more subsequent training tasks can be executed. Therefore, balancing the sizes of the sample data cached in each training node can balance the number of training tasks executed by each node, avoiding that some training nodes are hardly used to execute training tasks while some training nodes execute training tasks too frequently, resulting in damage to the hardware or a decrease in the processing speed of the nodes, affecting the training efficiency.

[0087] In some embodiments, the method further includes: in the case where the remaining storage space of any training node is not sufficient to cache the sample data corresponding to the training task, cleaning the sample data cached in the multiple training nodes according to the expiration date until the remaining storage space of any training node is sufficient to cache the sample data corresponding to the training task; determining the any training node as the target node.

[0088] The expiration date of the sample data can be determined based on the update time, creation time, etc. of the sample data.

[0089] Since each time a training task is executed, the training node can update the information of the sample data of the training task that has been cached, therefore, the older the update time, the longer it can be considered that the corresponding sample data has not been used for a long time period, and thus the possibility of subsequent use of the sample data is considered to be smaller.

[0090] Therefore, in the case where the remaining storage space of any training node is not sufficient to cache the sample data of the training task, the sample data cached in each training node can be cleaned according to the update time. The sample data can be sorted according to the order of the update time, and the sample data with a later update time can be cleaned one by one, so that the remaining storage space in the training node where the sample data is cleaned increases until there is any training node whose remaining storage space can cache the sample data of the training task to be executed, and this training node is determined as the target node.

[0091] In some embodiments, the LRU (Least Recently Used, cache eviction) policy can be used for cleaning.

[0092] Since the update time (i.e., the most recent usage time) of the sample data of each training node is maintained in the local training platform, the LRU strategy can be adopted to clean the cache, and the cleaning of the sample data can be continuously carried out until any training node can meet the storage requirements of the sample data for the training task to be executed.

[0093] In some embodiments, the training task includes a set of multiple subtasks, each subtask corresponds to sample data respectively, and each of the subtasks is executed by a corresponding training node.

[0094] Any training task corresponds to a model. A training task may also include a set of multiple subtasks, each subtask needs to be executed by a training node respectively, and each subtask may correspond to a copy of sample data respectively.

[0095] For example, a training task R can be divided into multiple subtasks (R1 to Rn), each subtask is executed by a training node respectively, and each subtask can also correspond to sample data (S1 to Sm). It should be noted that both n and m are positive integers, and m can be less than or equal to n, that is, the subtasks and the sample data can be in one-to-one correspondence, or there may be some subtasks that are other links in the training task and do not require sample data. Whether the training task is executed by one training node or multiple training nodes, it finally obtains a large data model and satisfies the model training method proposed in any of the above embodiments of the present disclosure.

[0096] Figure 3 It is a schematic diagram of a method for allocating a training task shown according to an embodiment of the present disclosure.

[0097] In some embodiments, based on the number of nodes required to execute the training task, the training task is divided into single-node tasks and multi-node tasks. The method further includes: arranging the single-node tasks in a single-node task queue in chronological order, and arranging the multi-node tasks in a multi-node task queue; when the idle resources of the multiple training nodes support the tasks in any queue, allocating the training tasks in the any queue to at least one training node for execution by the at least one training node; when the idle resources of the training nodes support the tasks in both queues, determining the start times of the tasks in the two queues respectively, and allocating the training task with the earlier start time to the at least one training node for execution by the at least one training node.

[0098] As Figure 3 shown, after generating a training task, according to whether the training task is a single-node task or a multi-node task, the training task can be divided into different task queues.

[0099] The task queue can be a first-in-first-out queue. After a training task enters the queue, in addition to the parameters required to execute the training task, the attributes maintained by each training task can also include parameters such as the initiation time and priority.

[0100] The task scheduler can retrieve the tasks at the head of the two queues at a certain frequency (for example, every 5 seconds) and determine the status of each training node, including whether the training node is available, the number of resources of the training node, the computing resources already used by the training node, and the unused computing resources.

[0101] In the case where the idle resources of the multiple training nodes support the tasks in either queue but not the tasks in the other queue, allocate the training tasks in the either queue to at least one training node for execution by the at least one training node.

[0102] Generally, a single-node task requires less computing resources, so it is easier to meet the conditions compared to a multi-node task. Therefore, when the training nodes are sufficient to support single-node tasks but not sufficient to support multi-node tasks, single-node tasks can be allocated to the training nodes first.

[0103] In the case where the idle resources of the training nodes support the tasks in both queues, determine the initiation times of the tasks in the two queues respectively, and allocate the training task with the earlier initiation time to the at least one training node for execution by the at least one training node.

[0104] If the training nodes can support both single-node tasks and multi-node tasks simultaneously, determine the initiation times of these two tasks respectively, and allocate the training task with the earlier initiation time, that is, the longer waiting time, to the training nodes for execution by the training nodes.

[0105] In this embodiment, two queues are adopted to separate single-node tasks and multi-node tasks into different queues and then allocate them to the training nodes for execution. Since the number of resources (number of training nodes) required for multi-node tasks is relatively large, it is difficult for the resources of multiple training nodes on the local training platform to meet the requirements of multi-node tasks. Therefore, if a single task queue is used, it may cause the multi-node task at the head of the queue to be shelved due to insufficient resources, and the subsequent single-node tasks that meet the resources are suspended and cannot be allocated to the nodes, resulting in the idle computing resources of the training nodes and a decrease in the model training efficiency.

[0106] By using two queues, when the resources of the training nodes do not meet the requirements of multi-node tasks but meet the requirements of single-node tasks, single-node tasks can be executed, thus making full use of the resources of the training nodes and improving the training efficiency.

[0107] In some embodiments, the waiting duration of the first multi-node task in the multi-node task queue can be determined. When the waiting duration is greater than the first moment, the allocation of the single-node task queue is suspended until the multi-node task is allocated to the training node.

[0108] Since the multi-node task requires more resources while the single-node task requires fewer resources, it may cause that as soon as there is idle resource on a training node, it is allocated with a single-node task, resulting in the multi-node task never having sufficient resources to execute.

[0109] Therefore, in this embodiment, when it is determined that the waiting duration of the multi-node task is too long, the allocation of the single-node task can be prohibited, so as to wait for more training nodes to become idle after completing the training tasks, accumulate idle computing resources until they meet the multi-node task, and then allocate the multi-node task with too long waiting duration to the training node.

[0110] In some embodiments, multiple training nodes can adopt a distributed training scheme for the multi-node task.

[0111] For example, a one-master-multi-slave scheme can be adopted. If a multi-node task requires a total of m training resources and each training node includes n training resources, the task scheduling module can split the multi-node task into m / n containers (pods), initiate one pod as the master pod, and set the status of the other several training nodes to be occupied for the multi-node task to prevent other tasks from occupying and conflicting. After the master pod is initiated, it can obtain its own IP and port number and call them back to the task scheduling module as parameters. The task scheduling module uses the IP and port number of the master pod as parameters to initiate the remaining (m / n - 1) pods.

[0112] In some embodiments, the local training platform also supports users to independently develop data preprocessing algorithms.

[0113] Based on the data preprocessing algorithm, users can make the preparation of sample data more personalized and automated. Users can customize the processing flow of sample data according to their needs to meet different data processing requirements.

[0114] In some embodiments, the sample data can be pre-stored in the storage space of the training node.

[0115] To improve the data access speed and prevent the storage system from being overloaded when there are too many training tasks, some or all of the sample data can be obtained from the storage system before the training node starts training or when the pressure is relatively small and stored in the storage space of the training node, so as to ensure that the training efficiency is not affected by the reading and transmission of the sample data.

[0116] In some embodiments, based on the local training platform, the user can, through training instructions, select the sample data for the training task and determine the base model, base configuration, and hyperparameters.

[0117] The user can, through the local training platform, select a suitable base model from the base models pre-set in the platform and adjust the base configuration and hyperparameters, so as to implement model training for the sample data selected by the user.

[0118] The local training platform can also predict the remaining training duration of the training task and feedback it to the user. The user can also log in to the inside of the container of the running k8s (Kubernetes) through the platform, so as to view the training log of the model.

[0119] In some embodiments, after the training task is completed on the local training platform, the obtained target model can be packaged and uploaded.

[0120] The local training platform can also upload the specific training log. The user can obtain the target model and the training log and conduct analysis based on the training log, which helps the user understand the progress and results of the model training, and adjust and improve the model, so that the user can select more appropriate parameters for the next model training.

[0121] In some embodiments, the training instructions also include the required number of resources.

[0122] The number of resources is specified by the user. If the number of resources exceeds the resources available on a single training node, the platform can convert the training task into a multi-node training task and allocate multiple training nodes to execute the task.

[0123] In some embodiments, after the training task is completed, the number of resources actually used by the training task can also be analyzed, and a resource occupancy table can be issued.

[0124] Since the number of resources required for each training task is specified by the user, but the user may not be able to accurately judge the actual resources required for the training task. Therefore, after the training task is completed, the number of resources actually used in each stage of the training task can be analyzed, and the maximum resources actually used by the training task can be analyzed, and the resource utilization rate can be displayed, so as to improve the user's awareness of resource conservation, help the user select appropriate resources for the next training task, and improve the resource utilization rate.

[0125] In some embodiments, after the training task is completed, the local training platform can also issue a model evaluation report.

[0126] The platform can evaluate and analyze the completed training tasks, and conduct multi - angle automated evaluation on the results of model training. For example, it can include but is not limited to model accuracy, inference speed, model size, memory usage, etc. This helps users more comprehensively understand the advantages and disadvantages between different models, so as to adjust or optimize the models.

[0127] For some specific models, the platform also supports automated model optimization, such as pruning the model. By removing redundant parameters, neurons or connections, the complexity of the model size is reduced, and the inference speed is increased and resource occupancy is reduced without affecting the model accuracy. Thus, intelligent decision - making support is provided to users to help them make more optimized models.

[0128] In some embodiments, the model generated by executing the training task can be sent to a specified downstream platform.

[0129] For the model generated by training, the local training platform can automatically compile algorithm execution files of each version, and perform continuous integration and deployment on the model. Sending the model to a specified downstream platform can improve development efficiency, improve stability, and be more efficient and reliable.

[0130] Corresponding to the embodiment of the model training method of the present disclosure, the present disclosure also provides an embodiment of a corresponding model training device.

[0131] Please refer to Figure 4 , Figure 4 which is a block diagram of a model training device in an embodiment of the present disclosure. As Figure 4 shown, the model training device includes:

[0132] A response unit 410, configured to generate a training task in response to a training instruction for a model to be trained, and determine sample data corresponding to the training task;

[0133] An execution unit 420, configured to, for the training task, when at least any one training node caches the sample data corresponding to the training task, use the sample data by the any one training node to execute the training task.

[0134] In some embodiments, the device is further configured to: when no training node caches the sample data corresponding to the training task, determine a target node that meets the training requirements of the training task from the multiple training nodes, and send the sample data corresponding to the training task to the target node to execute the training task.

[0135] In some embodiments, the target node includes a node with sufficient remaining storage space to cache the sample data corresponding to the training task.

[0136] In some embodiments, when there are multiple nodes whose remaining storage space is sufficient to cache the sample data corresponding to the training task, the node with the largest remaining storage space is determined as the target node.

[0137] In some embodiments, the apparatus is further configured to: when the remaining storage space of any training node is not sufficient to cache the sample data corresponding to the training task, clean the sample data cached by the multiple training nodes according to the expiration date until the remaining storage space of any training node is sufficient to cache the sample data corresponding to the training task; and determine the any training node as the target node.

[0138] In some embodiments, the training task includes a set of multiple subtasks, each subtask corresponds to sample data, and each subtask is executed by a corresponding training node.

[0139] In some embodiments, based on the number of nodes required to execute the training task, the training task is divided into a single-node task and a multi-node task. The apparatus is further configured to: arrange the single-node tasks in a single-node task queue in chronological order, and arrange the multi-node tasks in a multi-node task queue; when the idle resources of the multiple training nodes support the tasks in any queue, allocate the training tasks in the any queue to at least one training node for execution by the at least one training node; when the idle resources of the training nodes support the tasks in both queues, determine the initiation times of the tasks in the two queues respectively, and allocate the training task with the earlier initiation time to the at least one training node for execution by the at least one training node.

[0140] The implementation processes of the functions and roles of the various units in the above apparatus are specifically detailed in the implementation processes of the corresponding steps in the above method, and will not be elaborated here.

[0141] An embodiment of the present disclosure also provides an electronic device, including: a processor, a memory; the memory is used to store a computer program; the processor is used to execute the model training method according to any one of the above embodiments by calling the computer program.

[0142] An embodiment of the present disclosure also provides a computer-readable storage medium, on which a computer program is stored, and characterized in that the program, when executed by a processor, implements the model training method according to any one of the above embodiments.

[0143] An embodiment of the present disclosure also provides a computer program product, including a computer program, and the computer program, when executed by a processor, implements the method described in any one of the above embodiments.

[0144] Figure 5FIG. 0 is a schematic block diagram of a model training apparatus 500 shown according to an embodiment of the present disclosure. For example, the apparatus 500 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0145] Referring Figure 5 to FIG., the apparatus 500 may include one or more of the following components: a processing component 502, a memory 504, a power component 506, a multimedia component 508, an audio component 510, an input / output (I / O) interface 512, a sensor component 514, and a communication component 516.

[0146] The processing component 502 generally controls the overall operation of the apparatus 500, such as operations associated with display, telephone calls, data communications, camera operations, and recording operations. The processing component 502 may include one or more processors 520 to execute instructions to complete all or part of the steps of the above-described model training method. In addition, the processing component 502 may include one or more modules to facilitate interaction between the processing component 502 and other components. For example, the processing component 502 may include a multimedia module to facilitate interaction between the multimedia component 508 and the processing component 502.

[0147] The memory 504 is configured to store various types of data to support the operation of the apparatus 500. Examples of such data include instructions for any application or method operating on the apparatus 500, contact data, phone book data, messages, pictures, videos, etc. The memory 504 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a disk, or an optical disk.

[0148] The power component 506 provides power to the various components of the apparatus 500. The power component 506 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the apparatus 500.

[0149] The multimedia component 508 includes a screen that provides an output interface between the device 500 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of the touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 508 includes a front camera and / or a rear camera. When the device 500 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have focal length and optical zoom capabilities.

[0150] The audio component 510 is configured to output and / or input audio signals. For example, the audio component 510 includes a microphone (MIC) that is configured to receive external audio signals when the device 500 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 504 or transmitted via the communication component 516. In some embodiments, the audio component 510 further includes a speaker for outputting audio signals.

[0151] The I / O interface 512 provides an interface between the processing component 502 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power button, and a lock button.

[0152] The sensor component 514 includes one or more sensors for providing an assessment of the various aspects of the state of the device 500. For example, the sensor component 514 can detect the on / off state of the device 500, the relative positioning of components, such as the display and the keypad of the device 500. The sensor component 514 can also detect a change in the position of the device 500 or a component of the device 500, the presence or absence of user contact with the device 500, the orientation or acceleration / deceleration of the device 500, and the temperature change of the device 500. The sensor component 514 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component 514 can also include a light sensor, such as a CMOS or a CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 514 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0153] The communication component 516 is configured to facilitate communication between the device 500 and other devices in a wired or wireless manner. The device 500 can access a communication standard-based wireless network, such as WiFi, 2G, 3G, 4G LTE, 5G NR, or a combination thereof. In an exemplary embodiment, the communication component 516 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 516 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0154] In an exemplary embodiment, the device 500 can be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above-described model training method.

[0155] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as the memory 504 including instructions, and the above instructions can be executed by the processor 520 of the device 500 to complete the above-described model training method. For example, the non-transitory computer-readable storage medium can be a ROM, Random Access Memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0156] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the disclosure herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed herein. The specification and embodiments are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.

[0157] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

[0158] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. The term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.

[0159] The methods and devices provided by the embodiments of the present disclosure have been introduced in detail above. Specific examples are used herein to elaborate on the principles and implementation manners of the present disclosure. The description of the above embodiments is only used to help understand the method and its core idea of the present disclosure; at the same time, for those of ordinary skill in the art, according to the idea of the present disclosure, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present disclosure.

Claims

1. A model training method, characterized in that, Applied to a local training platform including multiple training nodes, the method includes: In response to a training instruction for a model to be trained, generating a training task and determining the sample data corresponding to the training task; For the training task, at least when any one of the training nodes caches the sample data corresponding to the training task, using the sample data by the any one of the training nodes to execute the training task.

2. The method according to claim 1, characterized in that The method further includes: When no training node caches the sample data corresponding to the training task, determining a target node that meets the training requirements of the training task from the multiple training nodes, and sending the sample data corresponding to the training task to the target node to execute the training task.

3. The method according to claim 2, characterized in that, The target node includes a node with sufficient remaining storage space to cache the sample data corresponding to the training task.

4. The method according to claim 3, wherein When there are multiple nodes with sufficient remaining storage space to cache the sample data corresponding to the training task, determining the node with the largest remaining storage space among them as the target node.

5. The method according to claim 2, wherein The method further includes: When the remaining storage space of any one of the training nodes is not sufficient to cache the sample data corresponding to the training task, cleaning the sample data cached by the multiple training nodes according to the expiration date until the remaining storage space of any one of the training nodes is sufficient to cache the sample data corresponding to the training task; Determining the any one of the training nodes as the target node.

6. The method according to claim 1, wherein The training task includes a group of multiple subtasks, each subtask corresponds to sample data respectively, and each of the subtasks is executed by a corresponding one of the training nodes.

7. The method according to claim 6, wherein Based on the number of nodes required to execute the training task, dividing the training task into a single-node task and a multi-node task, the method further includes: Arranging the single-node tasks in a single-node task queue in chronological order, and arranging the multi-node tasks in a multi-node task queue; When the idle resources of the multiple training nodes support the tasks in any one of the queues, allocating the training tasks in the any one of the queues to at least one of the training nodes for execution by the at least one of the training nodes; When the idle resources of the training nodes support the tasks in both queues, respectively determining the start times of the tasks in the two queues, and allocating the training task with the earlier start time to the at least one of the training nodes for execution by the at least one of the training nodes.

8. A model training device, characterized in that, Applied to a local training platform including multiple training nodes, the device includes: A response unit configured to generate a training task and determine the sample data corresponding to the training task in response to a training instruction for a model to be trained; An execution unit configured to, for the training task, at least when any one of the training nodes caches the sample data corresponding to the training task, use the sample data by the any one of the training nodes to execute the training task.

9. An electronic device, characterized in that, Including: A processor and a memory; The memory is used to store a computer program; The processor is configured to execute the model training method according to any one of claims 1-7 by calling the computer program.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the model training method according to any one of claims 1-7.