Model training method, device and system, storage medium and program product
By writing the model state snapshot to the paired training server's snapshot memory during the training process of large language model, the problems of long storage time and fault restart lost state are solved, and the effect of improving training efficiency and stability is achieved.
Patent Information
- Application Number
- CN202510276901.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-27
AI Technical Summary
During the training of large language models, the model state snapshot is stored for a long time, resulting in low training efficiency and easy to lose model state snapshots when the fault restarts.
During the model training process, training is paused when the preset model state saving condition is detected to meet the preset model state saving condition, obtain the snapshot memory address of the paired training server, and write the model state snapshot to the snapshot memory of the paired training server, shortening the saving path and time.
It shortens the saving time of model state snapshots, improves the efficiency of large model training, and avoids the loss of model state snapshots caused by failure restart.
Smart Images

Figure CN120217325A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to a model training method, device, system, storage medium, and program product. Background Art
[0002] The training of large language models is a complex, time-consuming, and error-prone process. Since the training of large language models requires a large amount of computing resources, thousands, tens of thousands, or even hundreds of thousands of artificial intelligence acceleration cards are often needed for synchronous computing and training, and it is not completed in a short time. It often takes several weeks or even months to train. During the training process, the training is often interrupted due to various abnormalities or faults, so it is necessary to use a model state snapshot to persist the training state to storage during the training process. When an abnormality occurs and the training cannot continue, the training does not need to start from scratch. As long as the model state snapshot is imported, the training can continue from the saved iteration.
[0003] Currently, when saving a model state snapshot, the data is usually first transferred from the memory of the artificial intelligence accelerator to the local memory and then saved to an external shared storage. To ensure the consistency of the training state, all artificial intelligence accelerators need to pause training when saving the model state snapshot. However, in the case of low storage bandwidth and large training parameters, hundreds of gigabytes or even terabytes of data need to be saved. If the model state snapshot is saved frequently, the time for saving data is very long, seriously affecting the training efficiency. And if the saving interval is too long, then after re-training, the training time wasted on re-iterating is very long, which is equivalent to wasting the training time. Summary of the Invention
[0004] The present invention provides a model training method, device, system, storage medium, and program product, which can shorten the time spent on training large models and improve the training efficiency of large models.
[0005] According to one aspect of the present invention, there is provided a model training method, including:
[0006] During the model training process, when it is detected that a preset model state saving condition is satisfied, the training is paused;
[0007] Obtain the paired training server corresponding to the current training server, and the memory addresses of the snapshot memories in the paired training server corresponding to each artificial intelligence accelerator in the current training server;
[0008] Through the training process, obtain the model state snapshots corresponding to each of the artificial intelligence accelerators, and write the model state snapshots corresponding to each of the artificial intelligence accelerators into the snapshot memory of the paired training server according to the memory addresses corresponding to each of the artificial intelligence accelerators, and continue the training.
[0009] According to another aspect of the present invention, there is provided a model training device, including:
[0010] A training pause module, configured to pause training when it is detected that a preset model state saving condition is satisfied during the model training process;
[0011] An address acquisition module, configured to acquire the paired training server corresponding to the current training server, and the memory addresses of the snapshot memories in the paired training server corresponding to each artificial intelligence accelerator in the current training server;
[0012] A snapshot writing module, configured to obtain the model state snapshots corresponding to each of the artificial intelligence accelerators through the training process, and write the model state snapshots corresponding to each of the artificial intelligence accelerators into the snapshot memory of the paired training server according to the memory addresses corresponding to each of the artificial intelligence accelerators, and continue the training.
[0013] According to another aspect of the present invention, there is provided a model training system, including a plurality of training servers, the training servers are paired in pairs, each training server includes at least one processor, and a local memory, a snapshot memory and an artificial intelligence accelerator that are communicatively connected to the at least one processor; wherein,
[0014] The local memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the model training method according to any embodiment of the present invention;
[0015] The local memory is used to store the model state snapshots corresponding to the artificial intelligence accelerators in the current training server;
[0016] The snapshot memory is used to store the model state snapshots corresponding to the artificial intelligence accelerators in the paired training server.
[0017] According to another aspect of the present invention, there is provided a computer-readable storage medium, the computer-readable storage medium stores a computer program, and the computer program is used to enable a processor to implement the model training method according to any embodiment of the present invention when executed.
[0018] According to another aspect of the present invention, there is provided a computer program product, including a computer program, and the computer program implements the model training method according to any embodiment of the present invention when executed by a processor.
[0019] In the technical solution of the embodiment of the present invention, during the model training process, when it is detected that the preset model state saving condition is met, the training is paused; the paired training server corresponding to the current training server is obtained, and the memory addresses of the snapshot memories in the paired training server corresponding to each artificial intelligence accelerator in the current training server are obtained; through the training process, the model state snapshots corresponding to each artificial intelligence accelerator are obtained, and according to the memory addresses corresponding to each artificial intelligence accelerator, the model state snapshots corresponding to each artificial intelligence accelerator are written into the snapshot memory of the paired training server, and the training is continued; by storing the model state snapshots in the snapshot memory of the paired training server, the saving path of the model state snapshots can be shortened, the time spent on saving the model state snapshots can be shortened, so that the time spent on large model training can be shortened, the training efficiency of the large model can be improved, and at the same time, the loss of the model state snapshots caused by the failure restart of this training server can be avoided.
[0020] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0022] Figure 1 is a flowchart of a model training method provided in Embodiment 1 of the present invention;
[0023] Figure 2 is a flowchart of saving a model state snapshot provided in Embodiment 1 of the present invention;
[0024] Figure 3 is a flowchart of a model training method provided in Embodiment 2 of the present invention;
[0025] Figure 4 is a flowchart of another model training method provided in Embodiment 2 of the present invention;
[0026] Figure 5 is a schematic structural diagram of a model training device provided in Embodiment 3 of the present invention;
[0027] Figure 6 is a schematic structural diagram of a model training system provided in Embodiment 4 of the present invention. Detailed implementation manners
[0028] In order to enable those skilled in the art of the present technology to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.
[0029] It should be noted that the terms "first", "second", "target", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0030] Embodiment 1
[0031] Figure 1 The following is a flowchart of a model training method provided for Embodiment 1 of the present invention. This embodiment is applicable to the situation of saving and reloading model state snapshots during the training of large language models. This method can be executed by a model training device, which can be implemented in the form of hardware and / or software. Typically, this model training device can be configured in the training server of a model training system. As Figure 1 shown, this method includes:
[0032] S110. During the model training process, when it is detected that a preset model state saving condition is satisfied, pause the training.
[0033] It should be noted that the saving process of the model state snapshot can be as Figure 2As shown in the figure. During the training process of a large language model, the model goes through multiple iterations or epochs. At the end of each epoch, the model's parameters (such as weights and biases, etc.) are updated. A model state snapshot (Checkpoint) is the process of saving the current state of the model (including parameters, optimizer state, etc.) at a certain point in time during these epochs. If the training process is interrupted, the training can be resumed through the Checkpoint, continuing from the last saved state, avoiding starting the training from scratch and saving time and computing resources.
[0034] In this embodiment, multiple training servers are formed into a server cluster to jointly train the large language model. Among them, each training server includes multiple artificial intelligence (AI) accelerators. An AI accelerator is a type of specialized hardware accelerator or computer system, which is a special-purpose processor mainly applied to artificial intelligence, artificial neural networks, machine vision, or machine learning. Moreover, the training servers are paired in pairs, that is, two training servers form a pair, and each training server is the paired training server of the other. Secondly, in each training server, in addition to the original local memory, an additional large memory is configured as a snapshot memory, which is specifically responsible for storing the model state snapshots generated by each AI accelerator in the paired training server. Data transmission between the snapshot memory and the AI accelerator can be achieved based on the Remote direct memory access (RDMA) technology.
[0035] Among them, the preset model state saving condition can be the preset condition information for Checkpoint saving. For example, it can be reaching a preset training iteration or epoch, reaching a specified time, etc. In a specific example, during the training process of the large language model, if it is detected that the specified time point of the preset training epoch is reached, it means that a Checkpoint needs to be saved. At this time, the model training of all training servers can be paused.
[0036] S120. Obtain the paired training server corresponding to the current training server, and the memory addresses of the snapshot memories in the paired training server corresponding to each AI accelerator in the current training server.
[0037] It should be noted that in each training server, the identity information (such as server identification, etc.) of the paired training server corresponding to the training server can be pre-configured, as well as the memory address of the snapshot memory corresponding to each artificial intelligence accelerator in the paired training server, and the above configuration information can be pre-stored in the local memory. In this embodiment, the snapshot memory can be evenly divided into regions according to the number of artificial intelligence accelerators in the paired training server to obtain the memory region corresponding to each artificial intelligence accelerator, and the physical address of the memory region can be used as the memory address corresponding to the artificial intelligence accelerator.
[0038] Thus, after suspending the training, the preset configuration information can be read from the local memory through the management process or the main management process to obtain the paired training server corresponding to the current training server, and the memory addresses of the snapshot memories corresponding to each artificial intelligence accelerator in the paired training server.
[0039] Optionally, obtaining the paired training server corresponding to the current training server, and the memory addresses of the snapshot memories corresponding to each artificial intelligence accelerator in the current training server in the paired training server may include:
[0040] Obtaining the identity information of the current training server;
[0041] According to the identity information of the current training server, if it is determined that the current training server is the main training server, the paired training server corresponding to the current training server, and the memory addresses of the snapshot memories corresponding to each of the artificial intelligence accelerators in the paired training server are directly obtained through the main management process of the current training server.
[0042] In an alternative embodiment, in the model training system, the identity information corresponding to each training server can be pre-configured and stored in the local memory. Typically, the identity information can include the main training server and the slave training servers. Among all the training servers, one training server can be designated as the main training server, and the remaining other training servers are all slave training servers. For example, if the model training system includes training servers node1 to nodeN, the identity information of node1 can be designated as the main training server, and correspondingly, the identity information of node2 to nodeN is the slave training server.
[0043] Secondly, run the main management process in the main training server, which is responsible for exchanging information with the management processes of all slave training servers, obtaining the real-time status information of each slave training server, and starting the training tasks on the local machine. In this embodiment, the pairing relationship between training servers can be configured only in the main training server, and the memory addresses of the peer snapshot memory are allocated to the paired training servers. The above configuration information can be transmitted by the main management process to the management process of each slave training server to enable all training servers to obtain the configuration information. At the same time, a management process runs in the slave training server, which is responsible for exchanging information with the main management process, obtaining the identity information of the corresponding paired training server through the main management process, and the memory address where the Checkpoint is stored in the snapshot memory of the paired training server.
[0044] The advantage of the above setting is that it can improve the information configuration efficiency and make the configuration effective for all training servers with one configuration.
[0045] Specifically, if it is determined that the identity information of the current training server is the main training server, the current training server can directly obtain the identity information of the corresponding paired training server and the memory address corresponding to the snapshot memory of each artificial intelligence accelerator in the paired training server through the main management process according to the preset configuration information.
[0046] Optionally, after obtaining the identity information of the current training server, it may further include:
[0047] According to the identity information of the current training server, if it is determined that the current training server is a slave training server, an address acquisition request is sent to the main management process of the main training server through the management process of the current training server, and the paired training server corresponding to the current training server and the memory addresses of the snapshot memories of the artificial intelligence accelerators in the paired training server fed back by the main management process are received.
[0048] Secondly, if the identity information of the current training server is a slave training server, the current training server can communicate with the main management process of the main training server through the management process to read the identity information of the corresponding paired training server and the memory address corresponding to the snapshot memory of each artificial intelligence accelerator in the paired training server from the main training server.
[0049] S130. Through the training process, obtain the model state snapshots corresponding to the artificial intelligence accelerators, and write the model state snapshots corresponding to the artificial intelligence accelerators into the snapshot memory of the paired training server according to the memory addresses corresponding to the artificial intelligence accelerators, and continue the training.
[0050] Among them, the training process is responsible for the training task of the large language model. During the model training process, it can write the Checkpoint into the corresponding snapshot memory or local memory, and can also continue the model training after reading the Checkpoint.
[0051] In this embodiment, through the training process, each artificial intelligence accelerator can be required to save its own training parameters according to the specified rules to generate a Checkpoint. Then, the corresponding memory address can be sent to each artificial intelligence accelerator so that the snapshot memory of the paired training server can be accessed by the artificial intelligence accelerator. Further, the artificial intelligence accelerator can write the generated Checkpoint into the corresponding memory area in the snapshot memory to achieve the saving of the Checkpoint. After the writing is completed, the training process can continue the training of the large language model. In this embodiment, the trained large language model can be used in application fields such as intelligent question answering and text translation.
[0052] It can be understood that since writing to a memory file is much faster than writing to shared storage, the model training can continue after the saving is completed. Therefore, storing the Checkpoint in the snapshot memory of the paired training server can greatly shorten the time required for saving the Checkpoint, improve the model training efficiency, and avoid the loss of the Checkpoint caused by the restart of the training server, thus improving the stability of the model training.
[0053] Optionally, the technical solution of this embodiment may further include:
[0054] If a server - type failure occurs, after the current training server is successfully restarted, obtain the memory addresses of the snapshot memory in the paired training server corresponding to each of the artificial intelligence accelerators;
[0055] Through the training process, according to the memory addresses corresponding to each of the artificial intelligence accelerators, read the model state snapshots corresponding to each of the artificial intelligence accelerators from the snapshot memory of the paired training server, and continue the model training based on the model state snapshots corresponding to each of the artificial intelligence accelerators.
[0056] Specifically, if a server - type failure occurs in the current training server, after the current training server is successfully restarted, the memory addresses of the snapshot memory in the paired training server corresponding to each artificial intelligence accelerator can be obtained through the main management process or the management process. Then, the training process can send the memory addresses to the corresponding artificial intelligence accelerators so that the artificial intelligence accelerators can directly read the corresponding Checkpoint from the corresponding memory area of the snapshot memory to continue the model training.
[0057] It should be noted that the larger the model, the longer the training duration, the more errors occur, and the more times of retraining after abnormal interruption, resulting in a greater overall time consumption. In the prior art, when retraining after an exception occurs, thousands of artificial intelligence accelerators simultaneously read the Checkpoint from the shared storage. It is necessary to first read the Checkpoint from the shared storage to the local memory and then from the local memory to the memory of the artificial intelligence accelerator. The data loading path is very long, directly leading to the bottleneck of the storage bandwidth and blocking the restart of model training. In the solution of this embodiment, when performing fault recovery, the artificial intelligence accelerator directly reads the Checkpoint from the snapshot memory of the paired training server, shortening the data loading path, improving the data reading speed, and achieving fast recovery after training interruption.
[0058] Optionally, if a server - type failure occurs and the current training server fails to restart, the standby training server of the current training server is enabled. The standby training server has the same number of artificial intelligence accelerators as the current training server. Then, the management process of the standby training server interacts with the main management process of the main training server to inform the failure status of the current training server and report its own identity information and the memory address of the snapshot memory. The main management process can update the existing configuration information according to the reported information of the standby training server. For example, the pairing information between training servers, the memory addresses corresponding to each artificial intelligence accelerator in the paired training server corresponding to the current training server, etc. For example, if the current training server is node3, the paired training server is node4, and the standby training server is node10, the pairing information [node3, node4] can be updated to [node10, node4].
[0059] Secondly, the standby training server can read the model state snapshot originally corresponding to the current training server with an exception from the shared storage or the snapshot memory of the paired training server to continue training. When the Checkpoint needs to be saved again in the future, the standby training server and the paired training server can store the Checkpoint of each other through their respective snapshot memories.
[0060] In the technical solution of the embodiment of the present invention, during the model training process, when it is detected that the preset model state saving condition is met, the training is paused; the paired training server corresponding to the current training server is obtained, and the memory addresses of the snapshot memories in the paired training server corresponding to each artificial intelligence accelerator in the current training server are obtained; through the training process, the model state snapshots corresponding to each artificial intelligence accelerator are obtained, and according to the memory addresses corresponding to each artificial intelligence accelerator, the model state snapshots corresponding to each artificial intelligence accelerator are written into the snapshot memory of the paired training server, and the training is continued; by storing the model state snapshots in the snapshot memory of the paired training server, the saving path of the model state snapshots can be shortened, the time spent on saving the model state snapshots can be shortened, so that the time spent on large model training can be shortened, the training efficiency of the large model can be improved, and at the same time, the loss of the model state snapshots caused by the failure restart of this training server can be avoided.
[0061] Embodiment 2
[0062] Figure 3 FIG. is a flowchart of a model training method provided by Embodiment 2 of the present invention. This embodiment further refines the above technical solution, and the technical solution in this embodiment can be combined with one or more of the above embodiments. As Figure 3 shown, the method includes:
[0063] S210. During the model training process, when it is detected that the preset model state saving condition is met, the training is paused.
[0064] S220. Obtain the paired training server corresponding to the current training server, and the memory addresses of the snapshot memories in the paired training server corresponding to each artificial intelligence accelerator in the current training server.
[0065] S230. Through the training process, obtain the model state snapshots corresponding to each artificial intelligence accelerator, and according to the memory addresses corresponding to each artificial intelligence accelerator, write the model state snapshots corresponding to each artificial intelligence accelerator into the snapshot memory of the paired training server.
[0066] S240. Through the training process, write the model state snapshots corresponding to each artificial intelligence accelerator into the local memory, send a write completion message to the write shared storage process, and continue the training.
[0067] In this embodiment, while writing the model state snapshot to the snapshot memory of the paired training server, it can also be written to the local memory of the current training server. After the writing is completed, a writing completion message can be sent to the write shared storage process to inform the write shared storage process and continue the training. The write shared storage process is mainly responsible for writing data to the shared storage, and the shared storage can be accessed by all training servers.
[0068] S250. After continuing the training, according to the writing completion message, the write shared storage process reads the model state snapshots corresponding to the artificial intelligence accelerators from the local memory and writes the model state snapshots corresponding to the artificial intelligence accelerators to the shared storage.
[0069] Specifically, after receiving the writing completion message, the write shared storage process writes the model state snapshot to the shared storage and writes a completion flag after the data writing is completed. It should be noted that during the period when the write shared storage process writes data to the shared storage, the model training is not blocked at all, that is, the model training does not need to wait for the shared storage to complete writing data.
[0070] Optionally, the technical solution of this embodiment may further include:
[0071] If a non-server type failure occurs, after the training process restarts, the training process reads the model state snapshots corresponding to the artificial intelligence accelerators from the local memory and continues the model training based on the model state snapshots corresponding to the artificial intelligence accelerators.
[0072] Specifically, when a non-server type failure occurs, after the training process restarts, it directly reads the latest Checkpoint from the local memory and then loads it into the corresponding artificial intelligence accelerator, and continues the model training after the loading is completed.
[0073] The advantage of the above setting is that it can further improve the reading efficiency of the Checkpoint when a non-server type failure occurs, thereby further improving the recovery speed after the training is interrupted.
[0074] In the technical solution of the embodiment of the present invention, through the training process, while writing the model state snapshots corresponding to each artificial intelligence accelerator into the snapshot memory of the paired training server according to the memory addresses corresponding to each artificial intelligence accelerator, the model state snapshots corresponding to each artificial intelligence accelerator are written into the local memory, and a write completion message is sent to the write shared storage process, and training continues; after continuing the training, the write shared storage process reads the model state snapshots corresponding to each artificial intelligence accelerator from the local memory according to the write completion message, and writes the model state snapshots corresponding to each artificial intelligence accelerator into the shared storage, which can avoid the loss of model state snapshots caused by the abnormality of the paired training server, can achieve the persistent storage of the model state snapshots, and can improve the stability of model training.
[0075] In a specific implementation manner of this embodiment, the flow of the model training method can be as Figure 4 shown. The model training system includes training servers node1 to nodeN, where node1 is the main training server running the main management process, and node2 to nodeN are slave training servers running the management process, and each management process communicates with the main management process. The training servers are paired in pairs in the order of the identifiers, that is, node1 is paired with node2, and nodeN-1 is paired with nodeN. When a Checkpoint needs to be saved, the Checkpoint of each AI accelerator is loaded into the snapshot memory of the paired training server while being transmitted to the local memory. After the data writing is completed, the model training can continue. Then, the write shared storage process reads the Checkpoint from the local memory and writes the Checkpoint into the shared storage.
[0076] When a training server abnormality occurs, if it is a non-server type failure, and the local memory data is not lost, only the training process needs to be restarted, and the Checkpoint is directly read from the memory of this server without obtaining it from the shared storage, and the data loading is completed independently and quickly, and the training continues. If it is a server type failure and the local memory data is lost, there is no need for multiple transfers of importing from the shared storage to the local memory and then to the AI accelerator. The Checkpoint in the snapshot memory of the paired training server is directly loaded into the AI accelerator, directly avoiding multiple data transfers in the entire link, achieving the purpose of acceleration, and enabling the model training task to be quickly restored.
[0077] Embodiment III
[0078] Figure 5 is a schematic structural diagram of a model training device provided in Embodiment III of the present invention. As Figure 5As shown in the figure, the device includes: a training pause module 310, an address acquisition module 320, and a snapshot writing module 330; where
[0079] The training pause module 310 is used to pause the training when it is detected that a preset model state saving condition is met during the model training process;
[0080] The address acquisition module 320 is used to acquire the paired training server corresponding to the current training server, and the memory addresses of the snapshot memories in the paired training server corresponding to each artificial intelligence accelerator in the current training server;
[0081] The snapshot writing module 330 is used to obtain the model state snapshots corresponding to each artificial intelligence accelerator through the training process, and write the model state snapshots corresponding to each artificial intelligence accelerator into the snapshot memory of the paired training server according to the memory addresses corresponding to each artificial intelligence accelerator, and then continue the training.
[0082] In the technical solution of the embodiment of the present invention, during the model training process, when it is detected that a preset model state saving condition is met, the training is paused; the paired training server corresponding to the current training server is acquired, and the memory addresses of the snapshot memories in the paired training server corresponding to each artificial intelligence accelerator in the current training server are acquired; through the training process, the model state snapshots corresponding to each artificial intelligence accelerator are acquired, and the model state snapshots corresponding to each artificial intelligence accelerator are written into the snapshot memory of the paired training server according to the memory addresses corresponding to each artificial intelligence accelerator, and then the training is continued; by storing the model state snapshots in the snapshot memory of the paired training server, the saving path of the model state snapshots can be shortened, the time taken to save the model state snapshots can be shortened, thereby the time taken for large model training can be shortened, the training efficiency of the large model can be improved, and at the same time, the loss of the model state snapshots caused by the failure restart of the current training server can be avoided.
[0083] Optionally, the address acquisition module 320 is specifically used to acquire the identity information of the current training server;
[0084] According to the identity information of the current training server, if it is determined that the current training server is the main training server, the paired training server corresponding to the current training server and the memory addresses of the snapshot memories in the paired training server corresponding to each artificial intelligence accelerator are directly acquired through the main management process of the current training server.
[0085] Optionally, the address acquisition module 320 is further configured to, according to the identity information of the current training server, if it is determined that the current training server is a slave training server, send an address acquisition request to the master management process of the master training server through the management process of the current training server, and receive the paired training server corresponding to the current training server and the memory addresses of the snapshot memories in the paired training server corresponding to each of the artificial intelligence accelerators fed back by the master management process.
[0086] Optionally, the model training device further includes:
[0087] A server restart module, configured to, if a server - type failure occurs, after the current training server is successfully restarted, acquire the memory addresses of the snapshot memories in the paired training server corresponding to each of the artificial intelligence accelerators;
[0088] A first snapshot reading module, configured to, through the training process, according to the memory addresses corresponding to each of the artificial intelligence accelerators, read the model state snapshots corresponding to each of the artificial intelligence accelerators from the snapshot memory of the paired training server, and continue model training based on the model state snapshots corresponding to each of the artificial intelligence accelerators.
[0089] Optionally, the snapshot writing module 330 is further configured to, through the training process, write the model state snapshots corresponding to each of the artificial intelligence accelerators into the local memory, and send a writing completion message to the write - shared storage process;
[0090] After continuing training, through the write - shared storage process, according to the writing completion message, read the model state snapshots corresponding to each of the artificial intelligence accelerators from the local memory, and write the model state snapshots corresponding to each of the artificial intelligence accelerators into the shared storage.
[0091] Optionally, the model training device further includes:
[0092] A second snapshot reading module, configured to, if a non - server - type failure occurs, after the training process is restarted, read the model state snapshots corresponding to each of the artificial intelligence accelerators from the local memory through the training process, and continue model training based on the model state snapshots corresponding to each of the artificial intelligence accelerators.
[0093] The model training device provided by the embodiments of the present invention can execute the model training method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.
[0094] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information and other processes all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0095] Embodiment 4
[0096] Figure 6 FIG. 4 is a schematic structural diagram of a model training system provided in Embodiment 4 of the present invention. The model training system 4 includes a plurality of training servers 40, the training servers 40 are paired in pairs, and each training server 40 includes at least one processor 41, and a local memory 42, a snapshot memory 43, and an artificial intelligence accelerator 44 that are communicatively connected to the at least one processor 41; wherein, in this embodiment, the type and quantity of the artificial intelligence accelerator 44 may not be specifically limited. The snapshot memory 43 may be a dedicated memory newly added in the training server 40, and the type thereof may not be specifically limited in this embodiment.
[0097] The local memory 42 stores a computer program executable by the at least one processor 41. The computer program is executed by the at least one processor 41 so that the at least one processor 41 can execute the model training method described in any embodiment of the present invention; the local memory 42 is used to store a model state snapshot corresponding to the artificial intelligence accelerator 44 in the current training server; the snapshot memory 43 is used to store a model state snapshot corresponding to the artificial intelligence accelerator 44 in the paired training server.
[0098] It should be noted that the components shown in this article, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present invention described and / or claimed in this article.
[0099] Among them, the local memory 42 may be a read-only memory (ROM), a random access memory (RAM), etc. The local memory 42 stores a computer program executable by at least one processor 41. The processor 41 can execute various appropriate actions and processes according to the computer program stored in the read-only memory or the computer program loaded from the storage unit into the random access memory. In the RAM, various programs and data required for the operation of the training server 40 can also be stored. The processor 41, ROM, and RAM are connected to each other through a bus. An input / output (I / O) interface is also connected to the bus.
[0100] A plurality of components in the training server 40 are connected to the I / O interface, including: an input unit, such as a keyboard, a mouse, etc.; an output unit, such as various types of displays, speakers, etc.; a storage unit, such as a disk, an optical disc, etc.; and a communication unit, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit allows the training server 40 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0101] The processor 41 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 41 include, but are not limited to, a central processing unit, a graphics processing unit, various dedicated artificial intelligence computing chips, various processors running machine learning model algorithms, a digital signal processor, and any suitable processor, controller, microcontroller, etc. The processor 41 executes the various methods and processes described above, such as the model training method.
[0102] In some embodiments, the model training method can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as the storage unit. In some embodiments, part or all of the computer program can be loaded and / or installed onto the training server 40 via the ROM and / or the communication unit. When the computer program is loaded into the RAM and executed by the processor 41, one or more steps of the model training method described above can be executed. Alternatively, in other embodiments, the processor 41 can be configured to execute the model training method by any other suitable means (e.g., by means of firmware).
[0103] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays, application-specific integrated circuits, application-specific standard products, systems-on-a-chip, programmable logic devices with load, computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor, receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0104] The computer programs for implementing the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer programs are executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer programs can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0105] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory, an optical fiber, a portable compact disk read-only memory, an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0106] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected to each other by any form or medium of digital data communication (e.g., a communication network). Examples of the communication network include: a local area network, a wide area network, a blockchain network, and the Internet.
[0107] The computing system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The relationship between the client and the server is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server.
[0108] This embodiment may also include a computer program product that includes a computer program which, when executed by a processor, implements the model training method provided in any embodiment of the present invention.
[0109] It should be understood that various forms of the flow shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is imposed herein.
[0110] The above specific embodiments do not constitute a limitation to the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A model training method, characterized in that: include: During the model training process, when it is detected that the preset model state saving conditions are met, the training is suspended; Obtain a paired training server corresponding to the current training server, and a memory address of a snapshot memory in the paired training server corresponding to each artificial intelligence accelerator in the current training server; Through the training process, the model state snapshot corresponding to each artificial intelligence accelerator is obtained, and according to the memory address corresponding to each artificial intelligence accelerator, the model state snapshot corresponding to each artificial intelligence accelerator is written to the snapshot memory of the paired training server, and the training is continued.
2. The method according to claim 1, characterized in that Obtaining a paired training server corresponding to the current training server, and a memory address of a snapshot memory in the paired training server corresponding to each artificial intelligence accelerator in the current training server, including: Get the identity information of the current training server; According to the identity information of the current training server, if it is determined that the current training server is the main training server, the paired training server corresponding to the current training server and the memory address of the snapshot memory corresponding to each artificial intelligence accelerator in the paired training server are directly obtained through the main management process of the current training server.
3. The method according to claim 2, characterized in that After obtaining the identity information of the current training server, it also includes: According to the identity information of the current training server, if it is determined that the current training server is a slave training server, an address acquisition request is sent to the main management process of the main training server through the management process of the current training server, and the paired training server corresponding to the current training server fed back by the main management process is received, as well as the memory address of the snapshot memory of each artificial intelligence accelerator in the paired training server.
4. The method according to claim 1, characterized in that: Also includes: If a server failure occurs, after the current training server is successfully restarted, the memory address of the snapshot memory in the paired training server corresponding to each of the artificial intelligence accelerators is obtained; Through the training process, according to the memory address corresponding to each artificial intelligence accelerator, the model state snapshot corresponding to each artificial intelligence accelerator is read from the snapshot memory of the paired training server, and model training is continued based on the model state snapshot corresponding to each artificial intelligence accelerator.
5. The method according to claim 1, characterized in that After obtaining the model state snapshot corresponding to each of the artificial intelligence accelerators through the training process, it also includes: Through the training process, the model state snapshot corresponding to each artificial intelligence accelerator is written to the local memory, and the write completion information is sent to the write shared storage process; After continuing the training, the model state snapshot corresponding to each artificial intelligence accelerator is read from the local memory according to the write completion information through the write shared storage process, and the model state snapshot corresponding to each artificial intelligence accelerator is written to the shared storage.
6. The method according to claim 5, characterized in that Also includes: If a non-server failure occurs, after the training process is restarted, the model state snapshot corresponding to each of the artificial intelligence accelerators is read from the local memory through the training process, and model training is continued based on the model state snapshot corresponding to each of the artificial intelligence accelerators.
7. A model training device, characterized in that: include: The training pause module is used to pause the training during the model training process when it is detected that the preset model state saving conditions are met; An address acquisition module, used to acquire a paired training server corresponding to the current training server, and a memory address of a snapshot memory in the paired training server corresponding to each artificial intelligence accelerator in the current training server; The snapshot writing module is used to obtain the model state snapshot corresponding to each artificial intelligence accelerator through the training process, and write the model state snapshot corresponding to each artificial intelligence accelerator into the snapshot memory of the paired training server according to the memory address corresponding to each artificial intelligence accelerator, and continue training.
8. A model training system, characterized in that: The system comprises a plurality of training servers, the training servers are paired in pairs, each training server comprises at least one processor, and a local memory, a snapshot memory and an artificial intelligence accelerator that are communicatively connected to the at least one processor; wherein, The local memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the model training method described in any one of claims 1 to 6; The local memory is used to store a model state snapshot corresponding to the artificial intelligence accelerator in the current training server; The snapshot memory is used to store the model state snapshot corresponding to the artificial intelligence accelerator in the paired training server.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is used to implement the model training method described in any one of claims 1 to 6 when executed by a processor.
10. A computer program product, characterized in that It includes a computer program, which implements the model training method described in any one of claims 1 to 6 when executed by a processor.