Method and device for carrying out model training through computing power of intelligent computing center
Through the computing power of the intelligent computing center, the target checkpoint information is obtained when the model training is interrupted, and the problem of inefficient model training is solved, and efficient model training and resource utilization is achieved.
Patent Information
- Application Number
- CN202510352590.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-06-27
AI Technical Summary
The current model is inefficient in training, especially when it needs to be retrained after training interruptions, resulting in waste of resources and inefficiency.
Through the computing power of the intelligent computing center, when a model training interrupt is detected, the target checkpoint information is obtained, and the training information of the model is obtained based on this information, and the training information is continued without retraining the model.
This greatly improves the training efficiency of the model, avoids the waste of computing resources, and reduces training time and cost.
Smart Images

Figure CN120216261A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of intelligent computing centers, intelligent computing centers, and computing power infrastructure, and particularly relates to a method and device for model training through the computing power of an intelligent computing center. Background Art
[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "intelligent computing centers" have emerged as the times require.
[0003] An "intelligent computing center" refers to a facility that provides the required computing power, data, and algorithms for artificial intelligence applications (such as scenarios of artificial intelligence deep learning model development, model training, and model inference) by using large-scale heterogeneous computing power resources, including general computing power and intelligent computing power. The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.
[0004] The "intelligent computing center" includes, but is not limited to, the "intelligent computing center".
[0005] An "intelligent computing center", that is, an artificial intelligence computing center, is a type of computing power infrastructure that provides computing power services, data services, and algorithm services required for artificial intelligence applications based on artificial intelligence theory and using an artificial intelligence computing architecture.
[0006] "Computing power" is the core of "intelligent computing centers" and "intelligent computing centers". It is the ability of computer devices or computing / data centers to process information, the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement, the computing ability to achieve the output of the target result by processing information data, and a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity, which is mainly provided to society through computing power infrastructure.
[0007] Currently, models play an increasingly important role in people's lives. However, during the model training process, due to various reasons, when the model training is interrupted, the model usually needs to be retrained. It can be seen that the current model training efficiency is very low. Summary of the Invention
[0008] The present invention provides a method and device for model training through the computing power of an intelligent computing center to solve the problem of very low current model training efficiency.
[0009] To solve the above problems, the present invention is implemented as follows:
[0010] In a first aspect, the present invention provides a method for model training through the computing power of an intelligent computing center, including:
[0011] Step S1: When it is detected that the training of the model is interrupted, obtain the target checkpoint information through a first function, where the target checkpoint information is the checkpoint information with the smallest time difference between the generation time of the checkpoint information and the current time, and the first function is a pre-set function for obtaining checkpoint information;
[0012] Step S2: Obtain the training information of the model according to the target checkpoint information, and the training information of the model corresponds one-to-one with the target checkpoint information;
[0013] Step S3: Train the model based on the training information of the model.
[0014] Optionally, before step S1, the method further includes:
[0015] Step S4: Generate a second function, where the second function is a function for generating checkpoint information;
[0016] Step S5: Generate the first function according to the second function, and the first function corresponds to the second function.
[0017] Optionally, step S1 includes:
[0018] Step S11: When it is detected that the Nth round of training of the model is interrupted, obtain the target checkpoint information through the first function, where N is an integer greater than 1;
[0019] Wherein, the target checkpoint information is the checkpoint information generated by the model during the (N - 1)th round of training, and the training information of the model is the training information of the model after the (N - 1)th round of training.
[0020] Optionally, before step S1, the method further includes:
[0021] Step S6: Train the model through a target computing power training platform;
[0022] Step S7: When it is detected that the target computing power training platform meets a preset condition, determine that the training of the model is interrupted;
[0023] The preset condition includes at least one of the following:
[0024] The application program of the target computing power training platform stops running, and the application program is used to train the model;
[0025] The target computing power training platform restarts;
[0026] The target computing power training platform receives target information for stopping the training of the model.
[0027] Optionally, step S3 includes:
[0028] Step S31: Calculate the acquisition speed value of the target Checkpoint information;
[0029] Step S32: When the acquisition speed value of the target Checkpoint information is within a preset numerical range, train the model based on the training information of the model.
[0030] Optionally, the training information of the model includes at least one of the following:
[0031] The weight information of the model;
[0032] The status information of the optimizer corresponding to the model;
[0033] The training round number information corresponding to the model.
[0034] In a second aspect, the present invention provides a model training device using the computing power of an intelligent computing center, including:
[0035] A first acquisition module, configured to obtain target Checkpoint information through a first function when detecting an interruption in the training of the model, where the target Checkpoint information is the Checkpoint information with the smallest time difference between the Checkpoint information generation time and the current time, and the first function is a pre-set function for obtaining Checkpoint information;
[0036] A second acquisition module, configured to obtain the training information of the model according to the target Checkpoint information, where the training information of the model corresponds to the target Checkpoint information one by one;
[0037] A first training module, configured to train the model based on the training information of the model.
[0038] In a third aspect, the present invention provides an electronic device, including: a processor, a memory, and a program stored on the memory and executable on the processor, and when the program is executed by the processor, it implements the steps of the model training method using the computing power of an intelligent computing center as described in the first aspect above.
[0039] Fourthly, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method for model training by means of the computing power of an intelligent computing center as described in the first aspect above are implemented.
[0040] Fifthly, the present invention provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, the steps of the method for model training by means of the computing power of an intelligent computing center as described in the first aspect above are implemented.
[0041] In the present invention, in the case where the training of the model is interrupted, target Checkpoint information is obtained through a first function. The target Checkpoint information is the Checkpoint information with the smallest time difference between the generation time of the Checkpoint information and the current time. The first function is a pre-set function for obtaining Checkpoint information; training information of the model is obtained according to the target Checkpoint information, and the training information of the model corresponds one-to-one with the target Checkpoint information; the model is trained based on the training information of the model.
[0042] In this way, when the training of the model is interrupted, the target Checkpoint information with the smallest time difference between the generation time of the Checkpoint information and the current time can be obtained through the first function, the training information of the model can be obtained according to the target Checkpoint information, and the model can be trained based on the training information of the model, that is, the model is trained based on the training information of the model, without re-training the model, and the computing power of the intelligent computing center is added, so that the training efficiency of the model can be greatly improved, and furthermore, a large amount of waste of computing power resources can be avoided. Description of the Drawings
[0043] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0044] Figure 1 is a flowchart of a method for model training by means of the computing power of an intelligent computing center provided by the present invention;
[0045] Figure 2 is a structural schematic diagram of a device for model training by means of the computing power of an intelligent computing center provided by the present invention;
[0046] Figure 3Schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners
[0047] The technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described content is part of the present invention, rather than all of the content. Based on the content in the present invention, all other content obtained by those of ordinary skill in the art without creative efforts belongs to the scope of protection of the present invention. The "computing power" referred to in the present invention means: the ability of a computer device or a computing / data center to process information, the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement, the computing ability to output a target result by processing information data, and a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity, which mainly provides services to society through computing power infrastructure.
[0048] The "computational power" (Computational Power, CP) referred to in the present invention means: the ability of a data center server to process data and output results, a comprehensive index to measure the computing ability of a data center, including general computing ability, supercomputing ability, and intelligent computing ability. The commonly used measurement unit is the number of floating-point operations per second (FLOPS, 1 EFLOPS = 10^18 FLOPS), and the larger the value, the stronger the comprehensive computing ability. It is estimated that 1 EFLOPS is approximately the computing power output of 5 Tianhe 2A or 500,000 mainstream server CPUs or 2 million mainstream laptops. The calculation formula is: CP = CP 通用 + CP 智能 + CP 超级 .
[0049] The "carrying capacity" (Network Power, NP) referred to in the present invention means: the performance of the data transmission ability of computing power facilities, a comprehensive ability including network architecture, network bandwidth, transmission delay, intelligent management and scheduling, etc., involving network transmission inside and between data centers, and a comprehensive index to measure the network transmission scheduling ability.
[0050] The "Storage Power (SP)" in the present invention refers to the comprehensive ability of a data center in four aspects: data storage capacity, performance, security and reliability, and green and low-carbon. It is a comprehensive indicator for measuring the data storage ability of a data center, including external storage devices such as storage arrays and built-in storage devices of servers. The commonly used measurement unit for storage capacity is exabyte (EB, 1EB = 2^60 bytes), the commonly used measurement unit for performance is the number of read and write operations per second per unit capacity (IOPS / TB, Input / Output Operations Per Second / TB), and the disaster recovery ratio is an important manifestation of security and reliability.
[0051] The "computing power infrastructure" in the present invention refers to a new type of information infrastructure that integrates information computing power, network carrying capacity, and data storage power, and can realize the centralized computing, storage, transmission, and application of information.
[0052] The "new type of information infrastructure" in the present invention mainly includes network infrastructures such as 5G networks, fiber broadband networks, backbone networks, international communication networks, and satellite Internet, computing power infrastructures such as data centers, general computing power centers, intelligent computing centers, and supercomputing centers, and new technology facilities such as artificial intelligence, blockchain, and quantum computing.
[0053] The "computing power" in the present invention includes general computing power, intelligent computing power, and super computing power.
[0054] The "general computing power" in the present invention refers to the computing ability provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.
[0055] The "intelligent computing power" in the present invention refers to a computing platform that is deployed on a large scale based on dedicated chips such as GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), and ASIC (Application Specific Integrated Circuit) for various artificial intelligence innovation applications, such as natural language processing and machine vision.
[0056] The "super computing power" in the present invention mainly refers to the computing ability provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of a multi-computer system working in parallel and processes extremely complex or data-intensive problems through a dedicated operating system. It is mainly used for computing in cutting-edge scientific fields, such as planetary simulation, drug molecule design, and gene analysis.
[0057] The "Intelligent Computing Center" as described in the present invention refers to a facility that uses large-scale heterogeneous computing power resources, including general computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.), and mainly provides the required computing power, data, and algorithms for artificial intelligence applications (such as scenarios like artificial intelligence deep learning model development, model training, and model inference). The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enablement.
[0058] The "Intelligent Computing Center" as described in the present invention includes, but is not limited to, the "Intelligent Computing Center".
[0059] The "Intelligent Computing Center" as described in the present invention, namely the artificial intelligence computing center, is a type of computing power infrastructure based on artificial intelligence theory, adopting an artificial intelligence computing architecture, and providing computing power services, data services, and algorithm services required for artificial intelligence applications.
[0060] The "Computing Power Center" as described in the present invention refers to a facility mainly composed of infrastructure such as wind, fire, water, and electricity and IT software and hardware devices, and having computing power, carrying capacity, and storage capacity, including general data centers, intelligent computing centers, supercomputing centers, etc.
[0061] The "Supercomputing Center" as described in the present invention, namely the supercomputing data center, is a data center based on supercomputers or large-scale computing clusters, capable of providing functions such as large-scale computing, storage, and network services, and is widely used in application scenarios such as aerospace, national defense, oil exploration, climate modeling, and genome sequencing.
[0062] The "Computing Power Resources" as described in the present invention refers to technologies and facilities with information computing, transmission, storage, and application capabilities required for the development of the digital society, including but not limited to computing resources such as CPU and GPU, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and support and guarantee resources such as wind, fire, water, and electricity.
[0063] The "Model" as described in the present invention includes, but is not limited to, the "Large Language Model" and the "Multimodal Large Model".
[0064] The "Large Language Model" as described in the present invention refers to a large-scale language model (LLM), which is a language model with a relatively large number of parameters, aiming to understand and generate human language, trained through a large amount of text data, and can perform a wide range of tasks including text summarization, translation, sentiment analysis, etc.
[0065] The "Multimodal Large Models" described in the present invention refers to a model that jointly trains multimodal information such as text, images, videos, and audio, including but not limited to multimodal large language models.
[0066] Please refer to Figure 1 , Figure 1 which is a flowchart of a method for model training using the computing power of an intelligent computing center provided by the present invention. As Figure 1 shown, it includes the following steps:
[0067] Step S1: When it is detected that the training of the model is interrupted, obtain the target Checkpoint information through a first function. The target Checkpoint information is the Checkpoint information with the smallest time difference between the generation time of the Checkpoint information and the current time. The first function is a pre-set function for obtaining Checkpoint information.
[0068] Among them, the model can be a neural network model or other models. The specific type of the model is not limited here. In the initial training stage of the model, the model and the optimizer can be initialized first, and then the model training can be started.
[0069] Among them, the Checkpoint information is the intermediate state and result information of the model during the training process, and it is a key technology for fault tolerance in the training process of the large language model. By saving the intermediate state and result information of the model as checkpoint information to persistent storage, in this way, when the training process of the model is interrupted, the training of the model can be resumed by loading the latest checkpoint.
[0070] It should be noted that the above intermediate state and result information of the model can also be understood as the training information of the model.
[0071] Among them, the first function is a pre-set function for obtaining Checkpoint information, and the specific setting method of the first function is not limited here. Optionally, the first function can be the function with the fastest acquisition speed and the highest accuracy for obtaining Checkpoint information selected from multiple functions; alternatively, the first function can also be the function corresponding to the information acquisition model pre-trained according to the Checkpoint information, that is, the information acquisition model is pre-trained and is used to obtain the Checkpoint information model. The Checkpoint information and the model can correspond, so that the first function can more accurately identify the Checkpoint information of the model, and thus obtain the target Checkpoint information more accurately and quickly.
[0072] Among them, during the training process of the model, a Checkpoint message can be generated every preset period, or a Checkpoint message can be generated after each round of iterative training. In this way, multiple Checkpoint messages can be generated, and the generation times of the multiple Checkpoint messages are different. According to the actual situation, it can be known that among the above multiple Checkpoint messages, the difference between the Checkpoint message with the smallest difference between the generation time of the Checkpoint message and the current time and the training information of the model obtained at the current time is also the smallest. Therefore, in the case where the training of the model is interrupted, the Checkpoint message with the smallest difference between the generation time of the Checkpoint message and the current time among the above multiple Checkpoint messages can be obtained through the first function, and this Checkpoint message can be determined as the target Checkpoint message. Then, the training information of the model is obtained based on the target Checkpoint message, and the model is trained based on the obtained Checkpoint message of the model. In this way, the training time and training resources of the model can be saved as much as possible, the training efficiency of the model can be improved, and the training cost of the model can be reduced.
[0073] Step S2: Obtain the training information of the model according to the target Checkpoint message, and the training information of the model corresponds one-to-one with the target Checkpoint message;
[0074] Among them, the training information of the model can be understood as the intermediate state and result information of the above model. Optionally, the target Checkpoint message can be used to store the training information of the model. In this way, when the target Checkpoint message is obtained, the training information of the model can be obtained.
[0075] Among them, the training information of the model corresponds one-to-one with the target Checkpoint message, that is, if the target Checkpoint message is different, the training information of the model is different. For example: when the target Checkpoint message is the first Checkpoint message, the first Checkpoint message can store the training information of the third round of training of the model. When the target Checkpoint message is the second Checkpoint message, the second Checkpoint message can store the training information of the fourth round of training of the model, etc.
[0076] Step S3: Train the model based on the training information of the model.
[0077] Among them, training the model based on the training information of the model can be understood as: training the model based on the training information of the model. For example, when the training information of the model is the fourth-round training information of the model, the model can be trained based on the fourth-round training information of the model. In this way, it is possible to avoid training the model from the first round, saving the training time and training resources of the model, improving the training efficiency of the model, and saving the training cost of the model.
[0078] In the present invention, through steps S1 to S3, when the training of the model is interrupted, the target Checkpoint information with the smallest time difference between the generation time of the Checkpoint information and the current time can be obtained through the first function, and the training information of the model can be obtained based on the target Checkpoint information, and the model can be trained based on the training information of the model, that is, training the model based on the training information of the model, without retraining the model, and adding the computing power of the intelligent computing center, thereby greatly improving the training efficiency of the model and avoiding a large waste of computing power resources.
[0079] Optionally, before the step S1, the method further includes:
[0080] Step S4: Generate a second function, where the second function is a function for generating Checkpoint information;
[0081] Step S5: Generate the first function according to the second function, and the first function corresponds to the second function.
[0082] Among them, the second function is a function for generating Checkpoint information, and the first function is a function for obtaining Checkpoint information. That is, the execution steps corresponding to the first function and the second function can be understood as opposite steps. Therefore, optionally, the first function and the second function can be inverse functions, so that the target Checkpoint information can be obtained more accurately.
[0083] In addition, the specific manner of generating the first function according to the second function is not limited herein. Optionally, when the execution steps corresponding to the first function and the second function can be understood as opposite steps, the second function can be inversely transformed to generate the first function; alternatively, the second function can be corrected according to the training scenario parameters of the model to obtain the first function.
[0084] In the present invention, the first function is generated according to the second function, and the first function corresponds to the second function, so that the generated first function can have higher accuracy.
[0085] Optionally, the step S1 includes:
[0086] Step S11: When it is detected that the Nth round of training of the model is interrupted, obtain the target Checkpoint information through the first function, where N is an integer greater than 1;
[0087] Wherein, the target Checkpoint information is the Checkpoint information generated by the model during the (N - 1)th round of training, and the training information of the model is the training information of the model after the (N - 1)th round of training.
[0088] In the present invention, when the Nth round of training of the model is interrupted, since the Nth round of training of the model is not completed, the target Checkpoint information generated by the model during the (N - 1)th round of training can be obtained, and the (N - 1)th round of training of the model must have been completed. In this way, by obtaining the target Checkpoint information generated by the model during the (N - 1)th round of training, it can be ensured that the integrity of the training information of the model obtained according to the target Checkpoint information is relatively high. Compared with the method of obtaining the target Checkpoint information generated by the model during the Nth round of training, the phenomenon that the model training goes wrong due to the relatively low integrity of the training information of the model can be reduced.
[0089] Optionally, before step S1, the method further includes:
[0090] Step S6: Train the model through a target computing power training platform;
[0091] Step S7: When it is detected that the target computing power training platform meets a preset condition, determine that the training of the model is interrupted;
[0092] The preset condition includes at least one of the following:
[0093] The application program of the target computing power training platform stops running, and the application program is used to train the model;
[0094] The target computing power training platform restarts;
[0095] The target computing power training platform receives target information, and the target information is used to stop the training of the model.
[0096] Among them, the application program of the target computing power training platform stops running and the target computing power training platform restarts can be understood as that the target computing power training platform fails, resulting in the passive interruption of the model training.
[0097] Among them, the target computing power training platform receives target information can be understood as actively interrupting the training of the model.
[0098] It should be noted that the target computing power training platform can be understood as a dedicated platform for the model. Training the model through the target computing power training platform can improve the training efficiency and accuracy of the model.
[0099] In the present invention, when it is detected that the target computing power training platform meets the preset conditions, it can be determined that the training of the model is interrupted. In this way, the accuracy of the determined training interruption result of the model can be improved.
[0100] Optionally, step S3 includes:
[0101] Step S31: Calculate the acquisition speed value of the target Checkpoint information;
[0102] Step S32: When the acquisition speed value of the target Checkpoint information is within the preset numerical range, train the model based on the training information of the model.
[0103] Among them, the specific type of the acquisition speed value of the target Checkpoint information is not limited here. Optionally, the acquisition speed value of the target Checkpoint information may refer to the acquisition speed value of a certain target Checkpoint information, or the acquisition speed value of the target Checkpoint information may refer to the average value of the acquisition speeds of multiple target Checkpoint information. Or, the acquisition speed value of the target Checkpoint information may include the average value of multiple speed values, the maximum speed value among multiple speed values, and the minimum average value among multiple speed values.
[0104] Optionally, the acquisition speed value of the target Checkpoint information may include the average value of multiple speed values and the floating range value of multiple speed values, and the floating range value of multiple speed values can be determined according to the difference between the maximum speed value and the minimum average value among multiple speed values.
[0105] In the present invention, when the acquisition speed value of the target Checkpoint information is within the preset numerical range, train the model based on the training information of the model. In this way, the training accuracy of the model can be ensured, and it can be ensured that the computing resources of the training environment of the model are relatively rich, thereby further improving the training efficiency of the model.
[0106] Optionally, the training information of the model includes at least one of the following:
[0107] The weight information of the model;
[0108] The status information of the optimizer corresponding to the model;
[0109] The training round information corresponding to the model.
[0110] Among them, the weight information of the model can refer to the parameter information of the model. The state information of the optimizer corresponding to the model can refer to the state information of the optimizer used to optimize the model during the training process of the model. The training round information corresponding to the model can refer to the information of the number of training rounds that the model has completed.
[0111] Optionally, the training information of the model may further include at least one of the following: the state information of the model, the training progress information of the model, the learning rate metric information of the model, the performance metric information of the model, the random seed information of the model, the data location information of the model, and the hardware state information of the model.
[0112] It should be noted that when training the model based on the above training information of the model, the state information of the model can be used to strictly match the network structure of the model, and the state information of the optimizer corresponding to the model can include optimization variable information such as momentum and second moment. The training progress information can be used to adjust the position of the data loader corresponding to the model. The learning rate metric information can include warmup and decay stage information. The performance metric information can be used for early stopping judgment of model training and selection of the specific type of the model. The random seed information can be used to ensure the consistency of data augmentation. The data location information can be used to restore the breakpoint position information of the data loader. The hardware state information can be saved separately during multi-card training of the model.
[0113] In the present invention, the training information of the model includes: the weight information of the model, the state information of the optimizer corresponding to the model, and the training round information corresponding to the model. In this way, the diversity and flexibility of the content of the training information of the model are increased. When the types of the content of the training information of the model are more, the training efficiency of the model is also higher.
[0114] See Figure 2 , Figure 2 is a schematic structural diagram of a model training device using the computing power of an intelligent computing center provided by the present invention. As Figure 2 shown, the model training device 200 using the computing power of the intelligent computing center includes:
[0115] The first acquisition module 201 is configured to, when detecting an interruption in the training of the model, obtain target Checkpoint information through a first function, where the target Checkpoint information is the Checkpoint information with the smallest time difference between the generation time of the Checkpoint information and the current time, and the first function is a pre-set function for obtaining Checkpoint information;
[0116] The second acquisition module 202 is configured to acquire the training information of the model according to the target Checkpoint information, and the training information of the model corresponds one-to-one with the target Checkpoint information;
[0117] The first training module 203 is configured to train the model based on the training information of the model.
[0118] Optionally, the model training apparatus 200 using the computing power of the intelligent computing center further includes:
[0119] The first generation module is configured to generate a second function, and the second function is a function for generating Checkpoint information;
[0120] The second generation module is configured to generate the first function according to the second function, and the first function corresponds to the second function.
[0121] Optionally, the first acquisition module 201 is further configured to, when detecting that the Nth round of training of the model is interrupted, acquire the target Checkpoint information through the first function, where N is an integer greater than 1;
[0122] Wherein, the target Checkpoint information is the Checkpoint information generated by the model during the (N - 1)th round of training, and the training information of the model is the training information of the model after the (N - 1)th round of training.
[0123] Optionally, the model training apparatus 200 using the computing power of the intelligent computing center further includes:
[0124] The second training module is configured to train the model through the target computing power training platform;
[0125] The determination module is configured to determine that the training of the model is interrupted when detecting that the target computing power training platform meets a preset condition;
[0126] The preset condition includes at least one of the following:
[0127] The application program of the target computing power training platform stops running, and the application program is used to train the model;
[0128] The target computing power training platform restarts;
[0129] The target computing power training platform receives target information, and the target information is used to stop the training of the model.
[0130] Optionally, the first training module 203 includes:
[0131] A calculation sub-module for calculating the acquisition speed value of the target Checkpoint information;
[0132] A training sub-module for training the model based on the training information of the model when the acquisition speed value of the target Checkpoint information is within a preset numerical range.
[0133] The model training device 200 provided by the present invention through the computing power of the intelligent computing center can execute each step in the above-mentioned model training method through the computing power of the intelligent computing center, and thus has the same beneficial technical effects as the above-mentioned model training method through the computing power of the intelligent computing center, which will not be elaborated here specifically.
[0134] Please refer to Figure 3 , the present invention also provides an electronic device 30, including a processor 31, a memory 32, and a computer program stored on the memory 32 and executable on the processor 31. When the computer program is executed by the processor 31, it realizes each process shown in the above-mentioned model training method through the computing power of the intelligent computing center, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0135] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it realizes each process shown in the above-mentioned model training method through the computing power of the intelligent computing center, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here. Among them, the computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.
[0136] The present invention also provides a computer program product, including computer instructions, which when executed by a processor, realize each process shown in the above-mentioned Figure 1 model training method through the computing power of the intelligent computing center shown, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0137] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such a process, method, article or device. Without further limitations, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or device including that element.
[0138] Through the description of the above embodiments, those skilled in the art can clearly understand that the method provided by the above invention can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the various methods provided by the present invention.
[0139] The present invention has been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the purpose of the present invention and the scope protected by the claims, and all of them belong to the protection scope of the present invention.
Claims
1. A method for model training using the computing power of an intelligent computing center, characterized in that: include: Step S1: when the training interruption of the model is detected, the target checkpoint information is obtained through a first function, the target checkpoint information is the checkpoint information with the smallest time difference between the checkpoint information generation time and the current time, and the first function is a preset function for obtaining the checkpoint information; Step S2: acquiring the training information of the model according to the target Checkpoint information, wherein the training information of the model corresponds one-to-one to the target Checkpoint information; Step S3: training the model based on the training information of the model.
2. The method according to claim 1, characterized in that Before step S1, the method further includes: Step S4: Generate a second function, where the second function is a function for generating Checkpoint information; Step S5: Generate the first function according to the second function, wherein the first function corresponds to the second function.
3. The method according to claim 1, characterized in that The step S1 comprises: Step S11: when it is detected that the Nth round of training of the model is interrupted, obtaining the target Checkpoint information through the first function, where N is an integer greater than 1; The target Checkpoint information is the Checkpoint information generated by the model during the N-1th round of training, and the training information of the model is the training information of the model after the N-1th round of training.
4. The method according to claim 1, characterized in that: Before step S1, the method further includes: Step S6: training the model through a target computing power training platform; Step S7: when it is detected that the target computing power training platform meets the preset conditions, determining to interrupt the training of the model; The preset condition includes at least one of the following: The application of the target computing power training platform stops running, and the application is used to train the model; The target computing power training platform is restarted; The target computing power training platform receives target information, and the target information is used to stop the training of the model.
5. The method according to any one of claims 1 to 4, characterized in that The step S3 comprises: Step S31: Calculate the acquisition speed value of the target Checkpoint information; Step S32: When the acquisition speed value of the target Checkpoint information is within a preset value range, training the model based on the training information of the model.
6. The method according to any one of claims 1 to 4, characterized in that The training information of the model includes at least one of the following: weight information of the model; Status information of the optimizer corresponding to the model; The number of training rounds corresponding to the model.
7. A model training device using the computing power of an intelligent computing center, characterized in that: include: A first acquisition module is used to acquire target Checkpoint information through a first function when a training interruption of the model is detected, wherein the target Checkpoint information is the Checkpoint information with the smallest time difference between the Checkpoint information generation time and the current time, and the first function is a preset function for acquiring Checkpoint information; A second acquisition module, used to acquire the training information of the model according to the target Checkpoint information, where the training information of the model corresponds to the target Checkpoint information one by one; The first training module is used to train the model based on the training information of the model.
8. An electronic device, characterized in that: include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of a method for model training using the computing power of an intelligent computing center as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of a method for model training using the computing power of an intelligent computing center as described in any one of claims 1 to 6.
10. A computer program product, characterized in that It includes computer instructions, which, when executed by a processor, implement the steps of the method for model training through the computing power of an intelligent computing center as described in any one of claims 1 to 6.