Data processing method and device, electronic equipment and storage medium

By splitting and merging the optimizer state parameters and adding weight parameters in the model checkpoint file in multimodal training in deep learning, the problem of how to effectively integrate the new and old model parameters is solved, and dynamic update and integration of model parameters and optimizer states is realized, improving the flexibility and efficiency of model training.

CN119940474APending Publication Date: 2025-05-06BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411844781.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-13
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In the multimodal training of deep learning, how to effectively integrate the trained model parameters and optimizer states, ensure that the new and old parameters work together, and maintain the correctness and effectiveness of model training is a technical problem.

Method used

By obtaining the original checkpoint file of the current training model, the splitting optimizer state parameters generates a first checkpoint file, and the added weight parameters and their corresponding optimizer state parameters are divided into the first checkpoint file according to the number of parameters to form a second checkpoint file to achieve dynamic update and integration of model parameters and optimizer state.

Benefits of technology

It realizes dynamic expansion of model parameters and precise management of optimizer state, ensuring that the model can seamlessly continue hot-start training after introducing new parameters, improving the flexibility and efficiency of model training, and enhancing the model's ability to process multimodal data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940474A_ABST
    Figure CN119940474A_ABST
Patent Text Reader

Abstract

The invention provides a data processing method and device, electronic equipment and a storage medium, and relates to the technical field of deep learning, in particular to the technical field of distributed large model training optimization. According to the specific implementation scheme, an original check point file corresponding to a current training model is obtained, wherein the original check point file comprises weight parameters and optimizer state parameters; segmenting optimizer state parameters in the original check point file to obtain a first check point file; and adding the newly added weight parameters into the first check point file, segmenting and merging the optimizer state parameters corresponding to the newly added weight parameters into the first check point file according to the number of the parameters to obtain a second check point file, and performing hot start training through the parameters in the second check point file. The multi-modal data processing capability of the model is remarkably enhanced, meanwhile, time and resource consumption caused by training from the beginning is avoided, and the continuity of model training and the stability of model performance are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of deep learning technology, specifically to the field of distributed large model training optimization technology, and especially to data processing methods, devices, electronic devices and storage media. Background Art

[0002] In the field of deep learning, multimodal training is a technique that integrates multiple types of data (such as text, images, audio, etc.) to improve model performance. This training method enables the model to learn richer feature representations, so that it can perform better in various application scenarios. However, in practical applications, we often face a challenge: how to effectively integrate the trained model parameters and optimizer states during the training process to enhance the model's ability to process new modal data. Specifically, when introducing new parameters to the model being trained, how to ensure that these new and old parameters can work together to maintain the correctness and effectiveness of model training is a technical problem. Traditional solutions often fail to meet this requirement because they may not properly handle the interaction between new and old parameters, or introduce too much computational overhead during parameter integration. Therefore, developing a method that can dynamically add new parameters in multimodal training and keep the model stable is of great significance for improving model performance and adaptability. Summary of the invention

[0003] The present disclosure provides a data processing method, device, electronic device and storage medium.

[0004] According to one aspect of the present disclosure, a data processing method is provided, the method comprising:

[0005] Obtain an original checkpoint file corresponding to the current training model, wherein the original checkpoint file includes weight parameters and optimizer state parameters;

[0006] Splitting the optimizer state parameters in the original checkpoint file to obtain a first checkpoint file;

[0007] Add the newly added weight parameters to the first checkpoint file, and split and merge the optimizer state parameters corresponding to the newly added weight parameters into the first checkpoint file according to the number of parameters to obtain a second checkpoint file, so as to perform hot start training with the parameters in the second checkpoint file.

[0008] According to another aspect of the present disclosure, there is provided a data processing device, comprising:

[0009] An acquisition module is used to acquire an original checkpoint file corresponding to the current training model, wherein the original checkpoint file includes weight parameters and optimizer state parameters;

[0010] A splitting module, used for splitting the optimizer state parameters in the original checkpoint file to obtain a first checkpoint file;

[0011] A merging module is used to add the newly added weight parameters to the first checkpoint file, and to split and merge the optimizer state parameters corresponding to the newly added weight parameters into the first checkpoint file according to the number of parameters to obtain a second checkpoint file, so as to perform hot start training with the parameters in the second checkpoint file.

[0012] According to a third aspect of the present disclosure, there is provided an electronic device, including:

[0013] at least one processor; and

[0014] a memory communicatively connected to the at least one processor; wherein,

[0015] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute any method in any of the above technical solutions.

[0016] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute any one of the methods described in the above technical solutions.

[0017] According to a fifth aspect of the present disclosure, a computer program product is provided, including a computer program, wherein the computer program implements any one of the methods described in the above technical solutions when executed by a processor.

[0018] The present disclosure provides a data processing method, apparatus, device and storage medium. The present disclosure obtains the original checkpoint file of the current training model, and segments the optimizer state parameters to generate a first checkpoint file, and then integrates the newly added weight parameters and their corresponding optimizer state parameters into the first checkpoint file after segmentation by number to form a second checkpoint file. This solution realizes the correct integration of the newly added parameters and their optimizer states into the existing model framework, ensures the consistency of the model structure and parameters, avoids potential conflicts and inconsistencies, and thus ensures that the model can still be trained correctly after adding new parameters, thereby realizing the dynamic expansion of model parameters and the precise management of optimizer states. This process not only improves the flexibility and efficiency of model training, but also ensures that hot-start training can be seamlessly continued after adding new parameters. In addition, this method significantly enhances the model's ability to process multimodal data, while avoiding the time and resource consumption caused by training from scratch, and improves the continuity of model training and the stability of model performance.

[0019] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.

[0021] Figure 1 is a schematic diagram of steps of a data processing method in one embodiment of the present disclosure;

[0022] Figure 2 yes Figure 1 Schematic diagram of the process of step S102;

[0023] Figure 3 is a schematic diagram of steps of a data processing method in another embodiment of the present disclosure;

[0024] Figure 4 is a schematic diagram of steps of a data processing method in yet another embodiment of the present disclosure;

[0025] Figure 5 yes Figure 1 Schematic diagram of the process of step S103;

[0026] Figure 6 A principle block diagram of a data processing device in an embodiment of the present disclosure;

[0027] Figure 7 It is a block diagram of an electronic device used to implement the data processing method of the embodiment of the present disclosure. DETAILED DESCRIPTION

[0028] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0029] The present disclosure provides a data processing method, see Figure 1 As shown, Figure 1 : is a schematic diagram of the steps of the data processing method in an embodiment of the present disclosure, which can be applied to the server side, and includes:

[0030] Step S101, obtaining the original checkpoint file corresponding to the current training model, the original checkpoint file includes weight parameters and optimizer state parameters.

[0031] Specifically, in the deep learning training process, the "original checkpoint file" refers to a file containing the current state of the model, which saves the "weight parameters" and "optimizer state parameters" of the model. The checkpoint file can also be called a "checkpoint". The weight parameters are the values ​​used to predict the output in each layer of the network that constitutes the model, while the optimizer state parameters record some key information of the optimizer (such as SGD, Adam, etc.) during the training process, such as momentum and gradient accumulation values. This information is crucial for updating model parameters. The process of obtaining the original checkpoint file containing the weight parameters and optimizer state parameters is usually completed at a certain stage of model training to ensure that the current state of the model can be saved and restored, so as to facilitate subsequent training, evaluation or migration of model parameters.

[0032] In this way, by obtaining the original checkpoint file corresponding to the current training model, which contains the model's weight parameters and optimizer state parameters, the continuity and recoverability of model training can be ensured. The preservation of weight parameters enables the model to continue learning from the exact state of the last training, while the preservation of optimizer state parameters ensures that the optimizer can maintain the continuity of its update strategy, so that training can be seamlessly resumed after training is interrupted, or training consistency can be maintained when migrating models between different computing environments. This capability is particularly important for long-running training tasks because it allows rapid recovery after interruptions caused by hardware failures, planned maintenance, or resource competition, while reducing the time and computing resources wasted due to repeated calculations, improving the efficiency and reliability of the training process.

[0033] Step S102, splitting the optimizer state parameters in the original checkpoint file to obtain a first checkpoint file.

[0034] Specifically, "optimizer state parameters" refer to the parameters maintained by the optimizer (such as SGD, Adam, etc.) during the training process. These parameters include but are not limited to momentum, gradient accumulation values, etc., which are crucial for updating the model weights. The first checkpoint file is generated by splitting these optimizer state parameters in the original checkpoint file. This process involves identifying the optimizer state parameters in the original checkpoint file, and then splitting these parameters into several parts according to a specific splitting strategy (which may be based on the number of parameters, size or other criteria), and saving the split results as a new checkpoint file, namely the first checkpoint file. Such an implementation process helps to balance the computing load of each processor in a distributed training environment, optimize resource utilization, and ensure that the model can resume training from the interruption point, thereby improving the flexibility and efficiency of training.

[0035] In this way, by splitting the optimizer state parameters in the original checkpoint file to obtain the first checkpoint file, this splitting enables the optimizer state parameters to be more evenly distributed to different processors or devices, thereby balancing the computing load, reducing the risk of overload of a single processor, and optimizing resource usage. In addition, this splitting also helps to achieve a more efficient model recovery and fault recovery mechanism in a distributed training environment, because it allows each processor to independently load and maintain its own share of the optimizer state, thereby speeding up the speed of resuming model training from the interruption point and improving the robustness of the overall training process.

[0036] Step S103, adding the newly added weight parameters to the first checkpoint file, and splitting and merging the optimizer state parameters corresponding to the newly added weight parameters into the first checkpoint file according to the number of parameters, to obtain a second checkpoint file, so as to perform hot start training using the parameters in the second checkpoint file.

[0037] Specifically, the "first checkpoint file" refers to a file containing the current state of the model, and the "newly added weight parameters" refer to the newly added parameters to expand the model capabilities. After obtaining the new weight parameters, these new weight parameters are added to the first checkpoint file. Subsequently, for these newly added weight parameters, the corresponding optimizer state parameters also need to be updated accordingly. These optimizer state parameters, such as momentum and gradient accumulation values, are divided according to the number of parameters, which means that the optimizer state parameters are allocated or updated accordingly according to the number of new weight parameters, and these divided optimizer state parameters are merged back into the first checkpoint file to form a second checkpoint file. This second checkpoint file contains the complete, updated model parameters and optimizer state, so that the model can quickly start hot training from this new state, that is, continue training from where it was interrupted last time, instead of starting from scratch, which not only saves training time, but also improves the flexibility and efficiency of the training process.

[0038] In this way, by integrating the newly added weight parameters into the first checkpoint file, and splitting and merging the corresponding optimizer state parameters according to the number of parameters to form a second checkpoint file, the dynamic update of the model parameters and the optimizer state is achieved. The scheme disclosed in the present invention enables the model to seamlessly continue hot-start training from the interruption point after the introduction of new parameters, thereby maintaining the continuity of training. This not only improves the flexibility of model training, allowing the model to adapt to new tasks or data, but also reduces the time and computing resources wasted due to repeated calculations, thereby improving training efficiency. In addition, this method also enhances the stability of the model training process, because it ensures that even if new parameters are introduced during the training process, the model can recover from a consistent state, reducing the risk of training failure.

[0039] The present disclosure provides a data processing method, apparatus, device and storage medium. The present disclosure obtains the original checkpoint file of the current training model, and segments the optimizer state parameters to generate a first checkpoint file, and then integrates the newly added weight parameters and their corresponding optimizer state parameters into the first checkpoint file after segmentation by number to form a second checkpoint file. This solution realizes the correct integration of the newly added parameters and their optimizer states into the existing model framework, ensures the consistency of the model structure and parameters, avoids potential conflicts and inconsistencies, and thus ensures that the model can still be trained correctly after adding new parameters, thereby realizing the dynamic expansion of model parameters and the precise management of optimizer states. This process not only improves the flexibility and efficiency of model training, but also ensures that hot-start training can be seamlessly continued after adding new parameters. In addition, this method significantly enhances the model's ability to process multimodal data, while avoiding the time and resource consumption caused by training from scratch, and improves the continuity of model training and the stability of model performance.

[0040] In some optional embodiments, see Figure 2 , Figure 2 yes Figure 1 Flow diagram of step S102. Step S102, splitting the optimizer state parameters in the original checkpoint file to obtain a first checkpoint file, including:

[0041] Step S201, splitting the optimizer state parameters in the original checkpoint file according to a first splitting method to obtain sub-checkpoint files.

[0042] Specifically, "first splitting mode" refers to a specific parameter allocation strategy for evenly distributing optimizer state parameters to multiple processors or devices. The implementation process of this solution involves splitting these optimizer state parameters in the original checkpoint file according to the first splitting mode, that is, storing them in different sub-checkpoint files. The purpose of this is to balance the computing load of each processor in a distributed training environment, optimize resource utilization, and ensure that the model can resume training from the interruption point, thereby improving the flexibility and efficiency of training.

[0043] Step S202, loading the sub-checkpoint file, and splitting the optimizer state parameters in the sub-checkpoint file according to the second splitting method to obtain a first checkpoint file.

[0044] Specifically, a "sub-checkpoint file" refers to a file containing some optimizer state parameters, which are key information for the optimizer during the training process and are used to guide the update of model weights. The "second splitting method" refers to another parameter allocation strategy used to further subdivide the optimizer state parameters. The first splitting method is different from the second splitting method. The main difference between the first splitting method and the second splitting method is that they are used for the allocation of optimizer state parameters at different stages. The first splitting method is used to preliminarily allocate the parameters in the original checkpoint file to the sub-checkpoint files, while the second splitting method is used to further refine the allocation of parameters in these sub-checkpoint files to achieve more accurate load balancing and optimize the distributed training process.

[0045] After obtaining the sub-checkpoint files, first load these sub-checkpoint files, and then split the optimizer state parameters again according to the second splitting method, so as to adjust the parameter distribution more finely and ensure that they can be more evenly distributed to each processor. The purpose of this step is to optimize the distributed training process of the model, improve the training efficiency and flexibility of model recovery by accurately controlling the distribution of the optimizer state parameters, and finally integrate the split parameters into the "first checkpoint file" to prepare for the subsequent training or recovery of the model.

[0046] In this way, by splitting the optimizer state parameters in the original checkpoint file according to the first splitting method to obtain sub-checkpoint files, and then loading these sub-checkpoint files and further splitting them according to the second splitting method, this method realizes the refined management and load balancing distribution of the optimizer state parameters. This step-by-step splitting method not only improves the scalability and flexibility of the model in a distributed training environment, but also ensures that each processor can efficiently process the appropriate amount of parameters, thereby optimizing the use of computing resources and accelerating the training process. In addition, this method helps to maintain the stability and consistency of model training by precisely controlling the distribution of optimizer states, allowing the model to transition to a new training stage more smoothly, improving training efficiency and model performance.

[0047] In some optional embodiments, the optimizer state parameters in the original checkpoint file are segmented according to the first segmentation method to obtain sub-checkpoint files, including:

[0048] The optimizer state parameters in the original checkpoint file are split according to the fusion balance splitting method to obtain sub-checkpoint files.

[0049] Specifically, the present solution involves a key step in deep learning model training, namely the management of optimizer state. "Fusion balanced splitting method" is a specific parameter allocation strategy that evenly distributes optimizer state parameters to multiple processors or devices to ensure that each processor can bear an equal computing load. The specific implementation process of the disclosed embodiment includes: first extracting the optimizer state parameters from the original checkpoint file, and then dividing these parameters into several subsets according to the fusion balanced splitting method, each subset corresponding to a sub-checkpoint file. In this way, each sub-checkpoint file contains the optimizer state parameters assigned to a specific processor, providing a basis for subsequent distributed training or model recovery. This process helps to improve the efficiency and scalability of model training, especially in large-scale parallel training scenarios.

[0050] In this way, by adopting the fusion balanced splitting method to split the optimizer state parameters in the original checkpoint file and obtain the sub-checkpoint file, the balanced distribution of the optimizer state parameters among multiple processors is achieved. This balanced distribution ensures that each processor can bear an appropriate amount of computing load, thereby improving the efficiency and scalability of model training. Since the optimizer state parameters are evenly distributed to each processor, the risk of overload of a single processor can be reduced, while the overall training speed can be accelerated, which is particularly important for large-scale parallel training. In addition, this method also improves the robustness of model training, because even if a processor has a problem, it will not affect the continuation of the entire training process, because each processor is only responsible for a part of the optimizer state parameters. Therefore, this splitting method not only optimizes resource utilization, but also enhances the stability and reliability of the training process.

[0051] In some optional embodiments, see Figure 3 , Figure 3 The optimizer state parameters in the original checkpoint file are split according to the fusion balance splitting method to obtain a sub-checkpoint file, including the following steps:

[0052] Step S301, identifying the optimizer state parameters in the original checkpoint file, and summarizing to obtain the total amount of all optimizer state parameters.

[0053] Specifically, by identifying and extracting these optimizer state parameters stored in the original checkpoint file, and then summarizing them, the total amount of all optimizer state parameters is obtained. This process involves scanning the checkpoint file, collecting and accumulating the optimizer state parameters, providing an accurate data basis for subsequent parameter management and training strategy adjustment, and ensuring that the optimizer state can be accurately reconstructed when the model is restored or training continues.

[0054] Step S302: Calculate the optimizer state parameter share allocated to each processor according to the total amount of the aggregated optimizer state parameters and the number of processors to obtain the optimizer state parameter share.

[0055] Specifically, the "number of processors" refers to the number of computing units (such as GPUs or CPUs) involved in distributed training. First, the total amount of all optimizer state parameters is summarized, and then the share of optimizer state parameters that should be allocated to each processor is calculated based on the total amount of these parameters and the number of available processors. This process involves distributing the total amount of optimizer state parameters to each processor evenly or on demand to ensure load balancing and computational efficiency. In this way, each processor can obtain an appropriate amount of optimizer state parameters to perform training tasks, thereby optimizing the efficiency and scalability of the entire training process. Ultimately, this calculated share of optimizer state parameters for each processor will be used to guide subsequent model training and parameter updates.

[0056] Step S303 , according to the optimizer state parameter share, the corresponding optimizer state parameter is split and saved into each corresponding processor to obtain a sub-checkpoint file.

[0057] Specifically, after calculating the optimizer state parameter share, the corresponding parameters are split from the overall optimizer state according to the pre-calculated optimizer state parameter share of each processor, and saved to each corresponding processor respectively. This step ensures that each processor obtains the corresponding optimizer state parameters that need to be processed, thereby forming a "sub-checkpoint file". These sub-checkpoint files will be used to restore and continue training the model on a specific processor in a distributed training environment, so that the training process can be carried out efficiently and in a balanced manner.

[0058] In this way, by identifying the optimizer state parameters in the original checkpoint file and summarizing the total amount of all related parameters, the optimizer state parameter share that should be allocated to each processor is calculated based on the total amount of these parameters and the number of processors, and finally the corresponding optimizer state parameters are split and saved to each corresponding processor to form a sub-checkpoint file process. This solution realizes efficient management and load balancing distribution of optimizer state parameters. This method not only optimizes the use of computing resources, improves the efficiency and scalability of model training, but also ensures that each processor can bear an appropriate amount of workload in a distributed training environment, thereby reducing the risk of overloading a single processor, speeding up the overall training speed, and enhancing the stability and reliability of the training process.

[0059] In some optional embodiments, the optimizer state parameters in the sub-checkpoint file are segmented according to the second segmentation method to obtain a first checkpoint file, including:

[0060] The optimizer state parameters in the sub-checkpoint file are divided according to the number of parameters to obtain the first checkpoint file.

[0061] Specifically, the method of dividing by the number of parameters refers to dividing by the specific number of each parameter to ensure that each processor or device can obtain the appropriate number of parameters for processing. The implementation process of this solution involves dividing the optimizer state parameters in these sub-checkpoint files again according to the number of parameters so as to distribute them to each processor in a balanced manner, and then saving the results of the division to form a "first checkpoint file". This first checkpoint file contains the reallocated optimizer state parameters, providing a newer and more efficient parameter management method for subsequent model training, thereby supporting large-scale distributed training and rapid recovery of models.

[0062] In this way, by finely dividing the optimizer state parameters in the sub-checkpoint file according to the number of parameters, this process ensures that each processor can evenly bear the computing load, thereby obtaining the first checkpoint file. This segmentation method optimizes resource allocation in a distributed training environment and improves the efficiency and scalability of model training. Since each processor processes roughly the same number of parameters, this reduces processing bottlenecks caused by uneven parameters, speeds up training, and improves overall training throughput. At the same time, this method also helps maintain stability during training, because uniform parameter distribution reduces the risk of failure of a single processor due to overload, thereby enhancing the robustness of model training.

[0063] In some optional embodiments, see Figure 4 , Figure 4 The optimizer state parameters in the sub-checkpoint file are divided according to the number of parameters to obtain a first checkpoint file, including the following steps:

[0064] Step S401, identifying the optimizer state parameters in the original checkpoint file, and obtaining the total number of parameters of the optimizer state parameters.

[0065] Specifically, the total number of optimizer state parameters refers to the total number of these state parameters of the optimizer, which is specifically achieved by identifying and extracting the optimizer state parameters stored in the original checkpoint file, and then counting the total number of these parameters. This process involves parsing the checkpoint file, extracting all relevant optimizer state information, and calculating their total number, providing an accurate data basis for subsequent parameter management and training strategy adjustment. This step is a prerequisite for the segmentation and distribution of optimizer state parameters, ensuring that the optimizer state can be accurately reconstructed when the model is restored or training continues.

[0066] Step S402: Calculate the optimizer state parameter allocated to each processor according to the total number of optimizer state parameter parameters and the number of processors to obtain the optimizer state parameter share.

[0067] Specifically, the "number of processors" refers to the number of computing units (such as GPUs or CPUs) participating in distributed training. After obtaining the total number of optimizer state parameters, the number of optimizer state parameters that should be allocated to each processor is calculated based on the total number of optimizer state parameters and the number of processors. This process is called "optimizer state parameter share". Specifically, the optimizer state parameters are evenly or on demand distributed to each processor to ensure that each processor can obtain an appropriate amount of parameters for processing, thereby achieving load balancing and improving training efficiency. In this way, each processor can obtain its corresponding optimizer state parameter share, providing accurate data support for subsequent model training and parameter updates.

[0068] Step S403 : According to the optimizer state parameter share, the corresponding optimizer state parameter is split and saved into each corresponding processor to obtain a sub-checkpoint file.

[0069] Specifically, after calculating the optimizer state parameter share, the corresponding optimizer state parameters in the original checkpoint file are accurately divided according to the optimizer state parameter share of each processor, and they are saved to each processor separately. This step ensures that each processor has the optimizer state information required to perform its training task, thereby generating a "sub-checkpoint file". These sub-checkpoint files will be used to restore and continue training models on specific processors in a distributed training environment, so that the training process can be carried out efficiently and balanced, and each processor can independently load and maintain its own optimizer state parameters, thereby optimizing the overall training process and resource allocation.

[0070] In this way, by identifying the optimizer state parameters in the original checkpoint file and calculating the total number of parameters, and then accurately calculating the share of optimizer state parameters that should be allocated to each processor based on the total number of parameters and the number of processors, a balanced distribution of optimizer state parameters is achieved. This process not only optimizes resource utilization and ensures that each processor can efficiently process the appropriate amount of parameters, but also forms a sub-checkpoint file by splitting and saving the corresponding optimizer state parameters to the corresponding processor, thereby improving the scalability and flexibility of model training. This method enables each processor to train independently and efficiently in a distributed training environment, reducing performance bottlenecks caused by uneven parameter processing, speeding up training, and improving the overall stability and reliability of the training process.

[0071] In some optional embodiments, see Figure 5 , Figure 5 yes Figure 1 Flow diagram of step S103. Step S103, adding the newly added weight parameters to the first checkpoint file, and merging the optimizer state parameters corresponding to the newly added weight parameters into the first checkpoint file according to the number of parameters, to obtain a second checkpoint file, so as to perform hot start training with the parameters in the second checkpoint file, including:

[0072] Step S501, adding the newly added weight parameter to the first checkpoint file to obtain an updated file.

[0073] Specifically, "new weight parameters" refer to additional parameters introduced to expand or improve model performance. The "first checkpoint file" refers to a file that contains the current state of the model, including existing weight parameters and optimizer status. The implementation process of this solution involves integrating these new weight parameters into the first checkpoint file. This step is to update and expand the parameter set of the model so that the model can learn new functions or adapt to new data. By adding the new weight parameters to the first checkpoint file, an updated file is obtained, called the "updated file". This updated file now contains the new and old parameters of the model, providing a basis for further training and fine-tuning of the model, so that the model can continue to be trained on this basis to achieve more complex functions or improve its performance on specific tasks.

[0074] Step S502 , dividing the optimizer state parameters corresponding to the newly added weight parameters according to the number of parameters to obtain divided optimizer state parameters.

[0075] Specifically, when "new weight parameters" are introduced to expand the model, each new parameter requires a corresponding "optimizer state parameter" to guide its update process. These state parameters include momentum, gradient accumulation, etc., which are key information maintained by optimizers (such as SGD, Adam, etc.) during training. The implementation process of this solution involves splitting the optimizer state parameters corresponding to the new weight parameters according to the number of parameters, which means that the optimizer state share corresponding to each parameter needs to be determined based on the number of new weight parameters. In this way, it can be ensured that each new weight parameter has a corresponding optimizer state parameter corresponding to it, so that the optimizer can correctly update these newly introduced parameters. After splitting, a set of updated optimizer state parameters can be obtained, which will be used in subsequent training processes to ensure the correctness and effectiveness of model parameter updates.

[0076] Step S503: merge the updated file and the segmented optimizer state parameters to obtain a second checkpoint file, so as to perform hot start training using the parameters in the second checkpoint file.

[0077] Specifically, the "updated file" refers to the updated file obtained after the newly added weight parameters have been added, which contains the current weight information of the model. The "split optimizer state parameters" refers to the optimizer state information that is split accordingly according to the number of newly added weight parameters. This information is crucial for how the optimizer adjusts the weights. The implementation process of this solution involves merging these two parts, the weight parameters in the updated file and the split optimizer state parameters, into a complete "second checkpoint file". This second checkpoint file integrates all the necessary information of the model, including weights and optimizer status, so that the model can quickly hot-start training from this updated state, that is, continue training from where it was last interrupted instead of starting from scratch. This can save training time, improve training efficiency, and ensure the continuity of the training process and the stability of model performance.

[0078] In this way, by integrating the newly added weight parameters into the first checkpoint file to form an updated file, and then dividing the optimizer state parameters corresponding to these newly added weight parameters according to the number of parameters, and finally merging these divided optimizer state parameters with the updated file to form a second checkpoint file, this solution realizes the dynamic update and expansion of model parameters and optimizer state. This process not only ensures that the newly added parameters can be correctly managed and updated by the optimizer, but also ensures the continuity and consistency of the model during hot-start training, so that the model can continue training seamlessly, effectively avoiding training interruptions or performance losses caused by parameter updates. In addition, this method also improves the flexibility and adaptability of model training, allowing the model to adapt to new tasks or data more easily and enhancing the generalization ability of the model.

[0079] The following describes an embodiment of the device of the present application, which can be used to execute the data processing method in the above embodiment of the present application. For details not disclosed in the embodiment of the device of the present application, please refer to the embodiment of the data processing method in the above embodiment of the present application.

[0080] The present disclosure also provides a data processing device 600, such as Figure 6 As shown, including:

[0081] An acquisition module 601 is used to acquire an original checkpoint file corresponding to the current training model, where the original checkpoint file includes weight parameters and optimizer state parameters;

[0082] A splitting module 602 is used to split the optimizer state parameters in the original checkpoint file to obtain a first checkpoint file;

[0083] The merging module 603 is used to add the newly added weight parameters to the first checkpoint file, and to split and merge the optimizer state parameters corresponding to the newly added weight parameters into the first checkpoint file according to the number of parameters to obtain a second checkpoint file, so as to perform hot start training using the parameters in the second checkpoint file.

[0084] In some optional embodiments, the segmentation module 602 segments the optimizer state parameters in the original checkpoint file to obtain a first checkpoint file, including:

[0085] Splitting the optimizer state parameters in the original checkpoint file according to the first splitting method to obtain sub-checkpoint files;

[0086] The sub-checkpoint file is loaded, and the optimizer state parameters in the sub-checkpoint file are split according to the second splitting method to obtain a first checkpoint file.

[0087] In some optional embodiments, the segmentation module 602 segments the optimizer state parameters in the original checkpoint file according to the first segmentation method to obtain sub-checkpoint files, including:

[0088] The optimizer state parameters in the original checkpoint file are split according to the fusion balance splitting method to obtain sub-checkpoint files.

[0089] In some optional embodiments, the splitting module 602 splits the optimizer state parameters in the original checkpoint file according to the fusion balance splitting method to obtain sub-checkpoint files, including:

[0090] Identify the optimizer state parameters in the original checkpoint file and summarize them to get the total amount of all optimizer state parameters;

[0091] According to the total amount of the summarized optimizer state parameters and the number of processors, the optimizer state parameter share allocated to each processor is calculated to obtain the optimizer state parameter share;

[0092] According to the optimizer state parameter share, the corresponding optimizer state parameter is split and saved to each corresponding processor to obtain a sub-checkpoint file.

[0093] In some optional embodiments, the segmentation module 602 segments the optimizer state parameters in the sub-checkpoint file according to the second segmentation method to obtain a first checkpoint file, including:

[0094] The optimizer state parameters in the sub-checkpoint file are divided according to the number of parameters to obtain the first checkpoint file.

[0095] In some optional embodiments, the segmentation module 602 segments the optimizer state parameters in the sub-checkpoint file according to the number of parameters to obtain a first checkpoint file, including:

[0096] Identify the optimizer state parameters in the original checkpoint file and obtain the total number of parameters of the optimizer state parameters;

[0097] According to the total number of optimizer state parameter parameters and the number of processors, the optimizer state parameter allocated to each processor is calculated to obtain the optimizer state parameter share;

[0098] According to the optimizer state parameter share, the corresponding optimizer state parameter is split and saved to each corresponding processor to obtain a sub-checkpoint file.

[0099] In some optional embodiments, the merging module 603 adds the newly added weight parameters to the first checkpoint file, and divides and merges the optimizer state parameters corresponding to the newly added weight parameters into the first checkpoint file according to the number of parameters to obtain a second checkpoint file, so as to perform hot start training through the parameters in the second checkpoint file, including:

[0100] Add the newly added weight parameters to the first checkpoint file to obtain an updated file;

[0101] The optimizer state parameters corresponding to the newly added weight parameters are divided according to the number of parameters to obtain the divided optimizer state parameters;

[0102] The updated file and the segmented optimizer state parameters are merged to obtain a second checkpoint file, so as to perform hot start training using the parameters in the second checkpoint file.

[0103] In the technical solution disclosed herein, the acquisition, storage and application of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0104] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.

[0105] Figure 7A schematic block diagram of an example electronic device 700 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0106] like Figure 7 As shown, the electronic device 700 includes a computing unit 701, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the device 700 can also be stored. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0107] A number of components in the device 700 are connected to the I / O interface 705, including: an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a disk, an optical disk, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the device 700 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0108] The computing unit 701 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 701 performs the various methods and processes described above, such as data processing methods. For example, in some embodiments, the data processing method may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 708. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 700 via ROM 702 and / or communication unit 709. When the computer program is loaded into RAM 703 and executed by the computing unit 701, one or more steps of the applet distribution described above may be performed. Alternatively, in other embodiments, the computing unit 701 may be configured to perform the data processing method in any other appropriate manner (e.g., by means of firmware).

[0109] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0110] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0111] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0112] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0113] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0114] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0115] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.

[0116] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A data processing method, the method comprising: Obtain an original checkpoint file corresponding to the current training model, wherein the original checkpoint file includes weight parameters and optimizer state parameters; Splitting the optimizer state parameters in the original checkpoint file to obtain a first checkpoint file; Add the newly added weight parameters to the first checkpoint file, and split and merge the optimizer state parameters corresponding to the newly added weight parameters into the first checkpoint file according to the number of parameters to obtain a second checkpoint file, so as to perform hot start training with the parameters in the second checkpoint file.

2. The method according to claim 1, wherein: The step of segmenting the optimizer state parameters in the original checkpoint file to obtain a first checkpoint file includes: Splitting the optimizer state parameters in the original checkpoint file according to the first splitting method to obtain sub-checkpoint files; The sub-checkpoint file is loaded, and the optimizer state parameters in the sub-checkpoint file are segmented according to the second segmentation method to obtain a first checkpoint file.

3. The method according to claim 2, wherein: The optimizer state parameters in the original checkpoint file are split according to the first splitting method to obtain sub-checkpoint files, including: The optimizer state parameters in the original checkpoint file are split according to a fusion balance splitting method to obtain sub-checkpoint files.

4. The method according to claim 3, wherein: The optimizer state parameters in the original checkpoint file are split according to the fusion balance splitting method to obtain sub-checkpoint files, including: Identify the optimizer state parameters in the original checkpoint file, and summarize to obtain the total amount of all optimizer state parameters; According to the total amount of the summarized optimizer state parameters and the number of processors, the optimizer state parameter share allocated to each processor is calculated to obtain the optimizer state parameter share; According to the optimizer state parameter share, the corresponding optimizer state parameter is segmented and saved into each corresponding processor to obtain a sub-checkpoint file.

5. The method according to claim 2, wherein: The step of dividing the optimizer state parameters in the sub-checkpoint file according to the second dividing method to obtain the first checkpoint file includes: The optimizer state parameters in the sub-checkpoint file are divided according to the number of parameters to obtain a first checkpoint file.

6. The method according to claim 5, wherein: The optimizer state parameters in the sub-checkpoint file are divided according to the number of parameters to obtain a first checkpoint file, including: Identify the optimizer state parameters in the original checkpoint file and obtain the total number of parameters of the optimizer state parameters; According to the total number of optimizer state parameter parameters and the number of processors, the optimizer state parameter allocated to each processor is calculated to obtain the optimizer state parameter share; According to the optimizer state parameter share, the corresponding optimizer state parameter is segmented and saved into each corresponding processor to obtain a sub-checkpoint file.

7. The method according to any one of claims 1 to 6, wherein: The adding of the newly added weight parameters to the first checkpoint file, and merging the optimizer state parameters corresponding to the newly added weight parameters into the first checkpoint file according to the number of parameters to obtain a second checkpoint file, so as to perform hot start training through the parameters in the second checkpoint file, include: Adding the newly added weight parameter to the first checkpoint file to obtain an updated file; Dividing the optimizer state parameters corresponding to the newly added weight parameters according to the number of parameters to obtain the divided optimizer state parameters; The updated file and the segmented optimizer state parameters are merged to obtain a second checkpoint file, so as to perform hot start training with the parameters in the second checkpoint file.

8. A data processing device, comprising: An acquisition module is used to acquire an original checkpoint file corresponding to the current training model, wherein the original checkpoint file includes weight parameters and optimizer state parameters; A splitting module, used for splitting the optimizer state parameters in the original checkpoint file to obtain a first checkpoint file; A merging module is used to add the newly added weight parameters to the first checkpoint file, and to split and merge the optimizer state parameters corresponding to the newly added weight parameters into the first checkpoint file according to the number of parameters to obtain a second checkpoint file, so as to perform hot start training with the parameters in the second checkpoint file.

9. The device according to claim 8, wherein: The segmentation module segments the optimizer state parameters in the original checkpoint file to obtain a first checkpoint file, including: Splitting the optimizer state parameters in the original checkpoint file according to the first splitting method to obtain sub-checkpoint files; The sub-checkpoint file is loaded, and the optimizer state parameters in the sub-checkpoint file are segmented according to the second segmentation method to obtain a first checkpoint file.

10. The device according to claim 9, wherein: The segmentation module segments the optimizer state parameters in the original checkpoint file according to a first segmentation method to obtain sub-checkpoint files, including: The optimizer state parameters in the original checkpoint file are split according to a fusion balance splitting method to obtain sub-checkpoint files.

11. The device according to claim 10, wherein: The segmentation module segments the optimizer state parameters in the original checkpoint file according to the fusion balance segmentation method to obtain sub-checkpoint files, including: Identify the optimizer state parameters in the original checkpoint file, and summarize to obtain the total amount of all optimizer state parameters; According to the total amount of the summarized optimizer state parameters and the number of processors, the optimizer state parameter share allocated to each processor is calculated to obtain the optimizer state parameter share; According to the optimizer state parameter share, the corresponding optimizer state parameter is segmented and saved into each corresponding processor to obtain a sub-checkpoint file.

12. The device according to claim 9, wherein: The segmentation module segments the optimizer state parameters in the sub-checkpoint file according to the second segmentation method to obtain a first checkpoint file, including: The optimizer state parameters in the sub-checkpoint file are divided according to the number of parameters to obtain a first checkpoint file.

13. The device according to claim 12, wherein: The segmentation module segments the optimizer state parameters in the sub-checkpoint file according to the number of parameters to obtain a first checkpoint file, including: Identify the optimizer state parameters in the original checkpoint file and obtain the total number of parameters of the optimizer state parameters; According to the total number of optimizer state parameter parameters and the number of processors, the optimizer state parameter allocated to each processor is calculated to obtain the optimizer state parameter share; According to the optimizer state parameter share, the corresponding optimizer state parameter is segmented and saved into each corresponding processor to obtain a sub-checkpoint file.

14. The device according to any one of claims 8 to 13, wherein: The merging module adds the newly added weight parameters to the first checkpoint file, and divides and merges the optimizer state parameters corresponding to the newly added weight parameters into the first checkpoint file according to the number of parameters to obtain a second checkpoint file, so as to perform hot start training through the parameters in the second checkpoint file, including: Adding the newly added weight parameter to the first checkpoint file to obtain an updated file; Dividing the optimizer state parameters corresponding to the newly added weight parameters according to the number of parameters to obtain the divided optimizer state parameters; The updated file and the segmented optimizer state parameters are merged to obtain a second checkpoint file, so as to perform hot start training with the parameters in the second checkpoint file.

15. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.

16. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-7.

17. A computer program product, comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Large model training check point storage method and device and storage medium

    CN121681149A