Data storage method and equipment

By splitting the checkpoint data of the large model into multiple data blocks and saving and loading it using CXL memory modules, the problem of low loading efficiency of checkpoint data in large model training is solved, and more efficient data loading and storage is achieved.

CN119937933APending Publication Date: 2025-05-06LENOVO (BEIJING) LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510088189.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

During the training of large-scale models, in the prior art, the full amount of data needs to be loaded every time the checkpoint data is loaded, resulting in low loading efficiency.

Method used

By splitting the checkpoint data to multiple data chunks, and using the CXL memory module to save and load these data chunks, loading or writing only parts that are different from the saved data chunks in the CXL memory module.

Benefits of technology

Improve the loading efficiency of checkpoint data, avoid the need to load full data, and save storage space and loading time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119937933A_ABST
    Figure CN119937933A_ABST
Patent Text Reader

Abstract

The invention discloses a data storage method and equipment, and the method comprises the steps: obtaining the check point data of a target large model, and the check point data at least comprises the model parameter data, state data and training progress data of the target large model; creating a check point data structure corresponding to the check point data, wherein the check point data structure is used for storing data blocks of the check point data and block identifiers of the data blocks; and writing data blocks contained in the check point data into the CXL memory module according to the check point data structure, wherein the data blocks are obtained by splitting the check point data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of data storage, and in particular to a data storage method and device. Background Art

[0002] During the training of a large model, it is usually necessary to regularly save the checkpoint data of the large model and load the previously saved checkpoint data on demand. Therefore, the efficiency of saving and loading checkpoint data directly affects the training performance of the large model.

[0003] The checkpoint data of a large model contains a large amount of data. In related technologies, each load requires loading the entire checkpoint data, resulting in low loading efficiency. Summary of the invention

[0004] To this end, this application discloses the following technical solutions:

[0005] The first aspect of the present application provides a data storage method, comprising:

[0006] Obtaining checkpoint data of a target large model, wherein the checkpoint data includes at least model parameter data, state data, and training progress data of the target large model;

[0007] Creating a checkpoint data structure corresponding to the checkpoint data, wherein the checkpoint data structure is used to store data blocks and block identifiers of the data blocks of the checkpoint data;

[0008] The data blocks contained in the checkpoint data are written into the CXL memory module according to the checkpoint data structure, wherein the data blocks are obtained by splitting the checkpoint data.

[0009] Optionally, the method of splitting the checkpoint data to obtain data blocks includes:

[0010] According to the data dimension of the checkpoint data and the preset data block size, the checkpoint data is split into a plurality of data blocks.

[0011] Optionally, writing the data blocks contained in the checkpoint data to the CXL memory module according to the checkpoint data structure includes:

[0012] Determining a first data block from the data blocks included in the checkpoint data, wherein the first data block is a data block different from the data blocks stored in the CXL memory module;

[0013] A first data block of the checkpoint data is written to the CXL memory module according to the checkpoint data structure.

[0014] Optionally, also include:

[0015] Obtaining a storage address and a block identifier of a second data block contained in the checkpoint data, wherein the second data block is the same data block as the data block stored in the CXL memory module;

[0016] The storage address and block identifier of the second data block contained in the checkpoint data are recorded in the checkpoint data structure.

[0017] Optionally, writing the first data block of the checkpoint data to the CXL memory module according to the checkpoint data structure includes:

[0018] Obtaining storage space of a first data block in the CXL memory module for storing the checkpoint data;

[0019] Recording the storage address corresponding to the obtained storage space and the block identifier of the first data block of the checkpoint data in the checkpoint data structure;

[0020] The first data block of the checkpoint data is written into the obtained storage space.

[0021] Optionally, determining the first data block from the data blocks included in the checkpoint data includes:

[0022] Obtaining a first checksum value of a data block contained in the checkpoint data;

[0023] The first check value and the second check value are compared to determine the first data block based on the comparison result, wherein the second check value is a check value of the data block stored in the memory module.

[0024] Optionally, at least one of the following is also included:

[0025] Recording metadata of the data blocks of the checkpoint data in the checkpoint data structure, where the metadata of the data blocks includes at least one of the name, size, and quantity of the data blocks;

[0026] The metadata of the checkpoint data is recorded in the checkpoint data structure, wherein the metadata of the checkpoint data includes at least one of a timestamp, a timestamp size, and a version of the checkpoint data.

[0027] Optionally, also include:

[0028] Creating an address data structure corresponding to the checkpoint data;

[0029] The storage address of the data block in the checkpoint data and the checkpoint identifier corresponding to the checkpoint data are stored in the checkpoint address data structure.

[0030] Optionally, at least one of the following is also included:

[0031] According to the target checkpoint identifier and the address data structure, read the data blocks contained in the target checkpoint data corresponding to the target checkpoint identifier from the CXL memory module;

[0032] According to the target block identifier and the checkpoint data structure, a target data block corresponding to the target block identifier is read from the CXL memory module.

[0033] A second aspect of the present application provides a data storage device, comprising:

[0034] A first processor, a second processor and a CXL memory module;

[0035] The first processor is used to process data according to the target large model;

[0036] The second processor is used for:

[0037] Obtain checkpoint data of the target large model from the processor memory of the first processor, wherein the checkpoint data at least includes model parameter data, state data, and training progress data of the target large model;

[0038] Creating a checkpoint data structure corresponding to the checkpoint data, wherein the checkpoint data structure is used to store data blocks and block identifiers of the data blocks of the checkpoint data;

[0039] The data blocks contained in the checkpoint data are written to the CXL memory module according to the checkpoint data structure, wherein the data blocks are obtained by splitting the checkpoint data. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0041] Figure 1 is a flow chart of a data storage method provided in an embodiment of the present application;

[0042] Figure 2 is a flow chart of another data storage method provided by an embodiment of the present application;

[0043] Figure 3 is a schematic diagram of an address data structure provided in an embodiment of the present application;

[0044] Figure 4It is a schematic diagram of the association relationship between data blocks and checkpoint data provided by an embodiment of the present application;

[0045] Figure 5 It is a structural schematic diagram of a data storage device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0046] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0047] This embodiment provides a data storage method, see Figure 1 , the method may include the following steps.

[0048] S101, obtaining checkpoint data of the target large model, the checkpoint data at least including model parameter data, status data and training progress data of the target large model.

[0049] The method provided in this embodiment can be executed during the process of training the target large model. For example, during the training of the target large model, a set of checkpoint data of the target large model can be obtained at regular intervals and saved according to the method of this embodiment, or a set of checkpoint data of the target large model can be obtained after each update of the model parameters of the target large model and saved according to the method of this embodiment.

[0050] The target large model can be a large language model (LLM) or other large models without limitation.

[0051] The checkpoint data of the target large model may be acquired by a processor in an electronic device for processing a training task of the target large model, such as a graphics processing unit (GPU).

[0052] A checkpoint data (also referred to as Checkpoint) of a target large model obtained at any time may include various data related to the target large model at that time, such as model parameter data, state data, and training progress data. Model parameter data may include various model weight parameters of the target large model at the corresponding time, state data may include optimizer state data and scheduler state data (including but not limited to learning rate scheduler state data) of the target large model at the corresponding time, and training progress data may include learning rate, loss value, number of iterations (epoch), global number of steps, loss of the target large model in the validation set, and other data of the target large model at the corresponding time. In addition, checkpoint data may also include mixed precision training state data, momentum and gradient cache data, model configuration data, distributed training configuration data, etc. of the target large model at the corresponding time, without limitation.

[0053] By saving the checkpoint data of the target large model, the state of the target large model at the corresponding time can be completely saved. Correspondingly, by loading a copy of the checkpoint data, the target large model can be converted from the current state to the state at the corresponding time of the checkpoint data, and training can continue from the state at the corresponding time of the checkpoint data.

[0054] The saving method provided in this embodiment can be executed by a processor in an electronic device for processing data storage and loading tasks, for example, by a central processing unit (CPU) of the electronic device. The GPU can provide the obtained checkpoint data to the CPU, and the CPU saves it according to the method of this embodiment.

[0055] S102, creating a checkpoint data structure corresponding to the checkpoint data, where the checkpoint data structure is used to store data blocks and block identifiers of the data blocks of the checkpoint data.

[0056] The form of the checkpoint data structure is not limited. As an example, the checkpoint data structure may be a data structure defined in the following form.

[0057] typedef struct {

[0058] ModelParam *model_params; / / Model parameter list

[0059] size_t model_params_count; / / Number of model parameters

[0060] OptimizerState *optimizer_state; / / Optimizer state

[0061] size_t optimizer_state_count; / / Number of optimizers

[0062] SchedulerState *scheduler_state; / / Learning rate scheduler state

[0063] size_t scheduler_state_count; / / Number of learning rate schedulers

[0064] TrainingProgress progress; / / Training progress information (epoch, global_step, loss, etc.)

[0065] MixedPrecisionState mixed_precision_state; / / Mixed precision training state

[0066] MomentumAndGradientState momentum_and_gradient; / / Momentum and gradient cache

[0067] ModelConfig config; / / Model configuration (including architecture, hyperparameters, etc.)

[0068] DistributedTrainingConfig distributed_config; / / Distributed training configuration information

[0069] char *timestamp; / / Timestamp when saving checkpoint

[0070] size_t timestamp_size; / / Timestamp size

[0071] int version; / / version information

[0072] } Checkpoint;

[0073] Among them, the model parameter list model_params can include multiple model parameter structures (ModelParam), the model parameter structure is used to save the data blocks and block identifiers corresponding to the model parameter data, the optimizer state structure OptimizerState and the learning rate scheduler state structure SchedulerState are used to save the data blocks and block identifiers corresponding to the optimizer state data and the scheduler state data respectively, optimizer_state and scheduler_state are pointers to the above two structures respectively, the training progress structure TrainingProgress is used to save the data blocks and block identifiers corresponding to the training progress data, and progress is a pointer to the training progress structure. MixedPrecisionState, MomentumAndGradientState, ModelConfig, DistributedTrainingConfig are respectively used to save the mixed precision training state data, momentum and gradient cache data, model configuration data, distributed training configuration data The corresponding data block structure, mixed_precision_state, momentum_and_gradient, config, distributed_config are respectively pointers to the above structures.

[0074] The block identifier of a data block may be the name of the data contained in the data block, or the number of the data block, or other identifiers that can distinguish the data block from other data blocks in the checkpoint data, without limitation.

[0075] S103, writing data blocks contained in the checkpoint data into the CXL memory module according to the checkpoint data structure, where the data blocks are obtained by splitting the checkpoint data.

[0076] Compute Express Link (CXL) is a high-speed serial protocol developed by relevant manufacturers. CXL memory module refers to a memory module that supports the CXL protocol.

[0077] In the checkpoint data structure, the structure for storing data blocks may include pointers to corresponding data blocks. In step S103, the data blocks obtained by splitting the checkpoint data may be stored in the CXL memory module based on these pointers.

[0078] As some examples, the ModelParam structure may be defined as follows.

[0079] typedef struct {

[0080] char *parameter_name; / / parameter name (e.g. "layer1.weight")

[0081] void *parameter_data; / / Parameter data pointer, usually a binary data block (such as floating array, matrix, etc.)

[0082] size_t data_size; / / parameter data size (in bytes)

[0083] } ModelParam;

[0084] In this structure, parameter_data is a pointer to the data block of the model parameters.

[0085] The OptimizerState structure can be defined as follows.

[0086] typedef struct {

[0087] char *optimizer_name; / / Optimizer name (e.g. "Adam")

[0088] void *state_data; / / Optimizer state data, which may include momentum, gradient, etc.

[0089] size_t data_size; / / Size of status data

[0090] }OptimizerState;

[0091] In this structure, state_data is a pointer to the data block of the optimizer state data.

[0092] The beneficial effect of this embodiment is that the data blocks obtained by splitting the checkpoint data are saved in the CXL memory module through the checkpoint data structure. Based on the above saving method, when the checkpoint data needs to be loaded, the specific data blocks can be directly loaded through the corresponding checkpoint data structure without loading the complete checkpoint data, thereby improving the loading efficiency.

[0093] In this embodiment, before executing S103, the obtained checkpoint data may be split to obtain a plurality of data blocks constituting the checkpoint data. The splitting method may be:

[0094] According to the data dimension of the checkpoint data and the preset data block size, the checkpoint data is split into multiple data blocks.

[0095] In different embodiments, a checkpoint data may be divided into different data dimensions. For example, the checkpoint data may be divided into data dimensions such as model parameters, optimizer status, scheduler status, and training progress.

[0096] The data block size is used to limit the amount of data in each data block, and its value can be set as needed. For example, the preset data block size can be 100 megabytes (MB), 200 MB, etc., without limitation.

[0097] When splitting in the above manner, the checkpoint data can be first divided into multiple parts according to the data dimension, and then for the data of each data dimension, if the total amount of data in this dimension exceeds the preset data block size, the data of this dimension is further split according to the data block size to obtain multiple data blocks in which the amount of data in this dimension does not exceed the preset data block size; if the total amount of data in this dimension is less than or equal to the preset data block size, all the data in this dimension can be used as one data block.

[0098] Exemplarily, the checkpoint data can be split into model parameter data, optimizer status data, scheduler status data and training progress data. The total amount of model parameter data is greater than the preset data block size of 200MB. The model parameter data can be split according to 200MB to obtain multiple data blocks of model parameter data, each block contains a portion of the model parameter data and the data amount is less than or equal to 200MB. The total amount of optimizer status data, scheduler status data and training progress data is no more than 200MB, so all optimizer state data is taken as a data block, all scheduler status data is taken as a data block, and all training progress data is taken as a data block.

[0099] During the training of the target large model, the checkpoint data obtained this time may have duplicate parts with the previous checkpoint data. In this case, the data blocks of the checkpoint data can be written to the CXL memory module as follows:

[0100] Determine a first data block from the data blocks included in the checkpoint data, the first data block being a data block different from the data blocks stored in the CXL memory module;

[0101] A first data block of the checkpoint data is written to the CXL memory module according to the checkpoint data structure.

[0102] In this embodiment, after splitting the data blocks of the current checkpoint data, the processor can compare these data blocks with the data blocks already saved in the CXL memory module. If it is found after comparison that a data block of the current checkpoint data is the same as a data block already saved in the CXL memory module, then this data block in the current checkpoint data is determined to be the second data block; if it is found after comparison that a data block of the current checkpoint data is different from the data blocks already saved in the CXL memory module, then this data block in the current checkpoint data is determined to be the first data block.

[0103] When writing the data blocks of the current checkpoint data to the CXL memory module, only the first data block of the current checkpoint data may be written, and the second data block of the current checkpoint data may not be written.

[0104] Exemplarily, assume that three data blocks of model parameters are split from the currently obtained checkpoint data, namely data block 1, data block 2 and data block 3. After comparison, it is found that data blocks 1 and 2 are the same as the data blocks saved in the CXL memory module, and data block 3 is different from the data blocks saved in the CXL memory module. Then, it is determined that data block 1 and data block 2 are the second data blocks, and data block 3 is the first data block. Therefore, when writing data blocks to the CXL memory module, data blocks 1 and 2 are not written, but data block 3 is written.

[0105] The advantage of screening the first data block and writing the first data block to the CXL memory module is that it can avoid storing duplicate data blocks in the CXl memory module, thereby saving storage space of the CXL memory module.

[0106] The first data block may be determined by the processor comparing the content of the data block of the checkpoint data with the content of the data block stored in the CXL memory module to determine the first data block.

[0107] The first data block may also be determined in the following manner:

[0108] Obtaining a first checksum value of a data block contained in the checkpoint data;

[0109] The first check value and the second check value are compared to determine the first data block based on the comparison result, wherein the second check value is a check value of the data block stored in the memory module.

[0110] See also Figure 2 In this embodiment, the processor can process a data block based on the check value generation algorithm each time it writes a data block to the CXL memory module, obtain a check value corresponding to the data block, and write the check value into the check value library as the second check value corresponding to the data block.

[0111] The check value library may be stored in the CXL memory module or in other storage modules of the electronic device without limitation.

[0112] After obtaining a checkpoint data and splitting the checkpoint data into multiple data blocks, the processor can use the above-mentioned checkpoint value generation algorithm to process each data block of the checkpoint data, and obtain the first check value corresponding to each data block of the checkpoint data, which is recorded as [Hd1, Hd2, ...Hdn], where Hd1 to Hdn respectively represent the first check values ​​of data block 1 to data block n contained in the checkpoint data, and n is the total number of data blocks split from the currently obtained checkpoint data.

[0113] Then, for each first check value Hdi, the processor compares Hdi with each second check value in the check value library. If Hdi is the same as a second check value in the check value library, it is determined that the data block corresponding to Hdi is the second data block. If Hdi is different from each second check value in the check value library, it is determined that the data block corresponding to Hdi is the first data block. i is an integer from 1 to n.

[0114] The check value generation algorithm used in this embodiment can be any algorithm that meets the following uniqueness conditions:

[0115] For any two data blocks, if the contents of the two data blocks are different, the check values ​​of the two data blocks generated by the algorithm are different; if the contents of the two data blocks are the same, the check values ​​of the two data blocks generated by the algorithm are the same.

[0116] As some examples, the check value generation algorithm may be any hash algorithm in the related art, and the check value corresponding to the data block may be a hash value of the data block obtained by processing the data block based on the hash algorithm.

[0117] By determining the first data block by comparing the check value, the time spent on comparison can be shortened and the efficiency of determining the first data block can be improved.

[0118] In some embodiments, the first data block of the checkpoint data may be written to the CXL memory module according to the checkpoint data structure by:

[0119] Obtaining storage space of a first data block in a CXL memory module for storing checkpoint data;

[0120] Recording the storage address corresponding to the obtained storage space and the block identifier of the first data block of the checkpoint data in the checkpoint data structure;

[0121] A first data block of the checkpoint data is written to the obtained storage space.

[0122] In this embodiment, for any first data block, the processor can find the structure for storing the first data block in the checkpoint data structure according to the dimension to which the first data block belongs, allocate storage space for storing the first data block in the CXL memory module, and then record the storage address corresponding to the storage space and the block identifier of the data block in the structure for storing the first data block by assignment, and then write the first data block to the storage space.

[0123] If there are multiple structures corresponding to a dimension in the checkpoint data structure, any free structure can be selected to record the storage address and block identifier. For example, the model parameter list model_params corresponding to the model parameter dimension contains multiple model parameter structures. When recording the storage address and block identifier corresponding to a first data block, an free model parameter structure can be selected for recording.

[0124] As an example, assuming that a first data block is a data block corresponding to the model parameter data (recorded as data block 1), and its block identifier is layer1.weight, the processor can allocate a storage space in the CXL memory module according to the data amount of data block 1, and the storage address of the storage space is represented by 0x72fd500;

[0125] Then, the processor can obtain an idle model parameter structure in the model parameter list model_params of the checkpoint data structure, assign the storage address 0x72fd500 to the parameter data pointer of the model parameter structure, and the assignment method can be ModelParam.parameter_data=0x72fd500, and record the block identifier layer1.weight of the data block 1 in the parameter name of the model parameter structure, and the recording method can be *ModelParam.parameter_name="layer1.weight";

[0126] Finally, the processor can write data block 1 into the storage space corresponding to 0x72fd500 in the CXL memory module according to the parameter data pointer ModelParam.parameter_data of the model parameter structure.

[0127] The idle model parameter structure refers to a model parameter structure in which the parameter data pointer and parameter name contained therein have not been assigned values ​​in the above manner, that is, a model parameter structure in which a specific data block in the checkpoint data has not been saved.

[0128] As another example, assuming that a first data block is a data block corresponding to the optimizer state (recorded as data block 10), and its block identifier is Adam, the processor can allocate a storage space in the CXL memory module according to the data amount of data block 10, and the storage address of the storage space is represented by 0x72fd300;

[0129] Then, the processor may obtain the optimizer state structure of the checkpoint data structure, assign the storage address 0x72fd300 to the pointer state_data for storing the optimizer state data in the optimizer state structure, and the assignment method may be OptimizerState.state_data=0x72fd300, and record the block identifier Adam of the data block 10 in the optimizer name of the optimizer state structure, and the recording method may be *OptimizerState.optimizer_name=“Adam”;

[0130] Finally, the processor may write data block 10 into the storage space corresponding to 0x72fd300 in the CXL memory module according to the pointer OptimizerState.state_data storing the optimizer state data in the optimizer state structure.

[0131] The writing method of other first data blocks is similar and will not be repeated here.

[0132] In some embodiments, the first data block may be written by assigning each data item included in the first data block to a corresponding variable in the checkpoint data structure, and these variables may be stored in a CXL memory module.

[0133] As an example, the training progress structure TrainingProgress can be defined as follows:

[0134] typedef struct {

[0135] uint32_t epoch; / / Current training epoch

[0136] uint32_t global_step; / / The global step number of current training

[0137] float loss; / / Current loss value

[0138] float validation_loss; / / validation set loss (if any)

[0139] float learning_rate; / / Current learning rate

[0140] } TrainingProgress;

[0141] When the first data block to be written belongs to the training progress data, the data representing the training progress in the block, such as the number of iterations (epochs) of the current training, the global number of steps, the loss value, the validation set loss, and the learning rate, can be directly assigned to the corresponding variables of the above training progress structure.

[0142] As another example, data chunks belonging to model configuration data may be saved via the ModelConfig structure of the checkpoint data structure, which may be defined as follows:

[0143] typedef struct {

[0144] char *model_architecture; / / Model architecture (e.g. "GPT-3", "BERT")

[0145] uint32_t num_layers; / / Number of model layers

[0146] uint32_t hidden_size; / / Hidden layer size

[0147] uint32_t num_attention_heads; / / Number of attention heads

[0148] uint32_t vocab_size; / / vocabulary size (e.g. vocabulary size of GPT-2)

[0149] uint32_t max_position_embeddings; / / Maximum number of position embeddings (such as the maximum sequence length of BERT / GPT)

[0150] char *activation_function; / / Activation function (e.g. "ReLU", "GELU")

[0151] char *tokenizer_config; / / tokenizer configuration (such as BPE)

[0152] uint32_t batch_size; / / Batch size

[0153] float weight_decay; / / weight decay

[0154] float dropout; / / Dropout ratio

[0155] } ModelConfig;

[0156] As another example, the data blocks belonging to the model configuration data may include data such as model architecture, number of model layers, hidden layer size, number of attention heads, vocabulary size, maximum number of position embeddings, activation function, tokenizer configuration, batch size, weight decay and dropout rate. These data can be saved through the pointers or corresponding variables contained in the above structure. For example, the model architecture, activation function and tokenizer configuration are saved respectively through the pointers model_architecture, activation_function and tokenizer_config, and other data are saved through the corresponding variables.

[0157] The data chunks belonging to the distributed training configuration data can be saved through the DistributedTrainingConfig structure of the checkpoint data, which can be defined as follows:

[0158] typedef struct {

[0159] uint32_t rank; / / Current device number (rank)

[0160] uint32_t world_size; / / Total number of devices (total number of devices in the training cluster)

[0161] uint32_t local_rank; / / Local device number (number on the same machine)

[0162] uint32_t global_rank; / / Global device number (cross-machine number)

[0163] char *master_addr; / / master node address

[0164] uint32_t master_port; / / Master node port number

[0165] } DistributedTrainingConfig;

[0166] The data blocks belonging to the distributed training configuration data may include the number of the current device, the total number of devices, the local device number, the global device number, the address of the master node, the master node port number and other data. The address of the master node can be saved through the pointer master_addr of the DistributedTrainingConfig structure, and other data can be saved through the corresponding variables in the structure.

[0167] In some optional embodiments, for the second data block determined after the comparison, the processor may not write the second data block to the CXL memory module, but record the relevant information in the checkpoint data structure as follows:

[0168] Obtaining a storage address and a block identifier of a second data block contained in the checkpoint data, where the second data block is the same data block as the data block stored in the CXL memory module;

[0169] The storage address and block identifier of the second data block contained in the checkpoint data are recorded in the checkpoint data structure.

[0170] The storage address of the second data block refers to the storage address of the data block in the CXL memory module that is the same as the second data block.

[0171] Exemplarily, data block 13 in the currently obtained checkpoint data is the second data block, and a data block 3 identical to data block 13 has been stored in the CXL memory module. In this case, the processor can obtain the storage address of data block 3 in the CXL memory module as the storage address of data block 13.

[0172] In this embodiment, for any second data block, the processor can find the structure used to save the second data block in the checkpoint data structure according to the dimension to which the second data block belongs, and then record the storage address of the second data block and the block identifier of the second data block into the structure used to save the second data block by assignment.

[0173] If there are multiple structures corresponding to a dimension in the checkpoint data structure, any free structure can be selected to record the storage address and block identifier. For example, the model parameter list model_params corresponding to the model parameter dimension contains multiple model parameter structures. When recording the storage address and block identifier corresponding to a second data block, an free model parameter structure can be selected for recording.

[0174] In combination with the above example, assume that data block 13 is the data block corresponding to the model parameter data, and its block identifier is layer5.weight. Assume that the storage address of data block 3 in the CXL memory module is 0x61fd500;

[0175] The processor can obtain an idle model parameter structure in the model parameter list model_params of the checkpoint data structure, assign the storage address 0x61fd500 to the parameter data pointer of the model parameter structure, and the assignment method can be ModelParam.parameter_data=0x61fd500. In addition, the block identifier layer5.weight of the data block 13 is recorded in the parameter name of the model parameter structure, and the recording method can be *ModelParam.parameter_name=“layer5.weight”. At this point, the processor has recorded the storage address and block identifier of the data block 13 in the checkpoint data structure.

[0176] The storage addresses and block identifiers of other second data blocks are recorded in a similar manner and will not be described in detail.

[0177] In some optional embodiments, metadata of each data block and metadata of the checkpoint data may also be recorded in the checkpoint data structure in at least one of the following ways:

[0178] Recording metadata of a data block of the checkpoint data in a checkpoint data structure, wherein the metadata of the data block includes at least one of a name, a size, and a quantity of the data block;

[0179] The metadata of the checkpoint data is recorded in the checkpoint data structure, where the metadata of the checkpoint data includes at least one of a timestamp, a timestamp size, and a version of the checkpoint data.

[0180] Among them, the metadata of the data block belonging to the model parameter data may include the block identifier and the size of the data block. The recording method of the block identifier is mentioned above. The size of the data block can be recorded by directly assigning a value to the corresponding variable data_size in the model parameter structure. For example, if the size of a data block belonging to the model parameter data is 1000 bytes, the data_size variable of the model parameter structure used to store the data block can be assigned 1000, that is, data_size=1000, thereby recording the size of the data block.

[0181] Optionally, the metadata of the data blocks belonging to the model parameter data may also include the number of data blocks of the model parameter data, that is, the total number of data blocks belonging to the model parameter data.

[0182] The metadata of the data block belonging to the optimizer state data may include the block identifier of the data block and the size of the data block. The size of the data block may be recorded in the data_size variable of the optimizer state structure. The recording method is as mentioned above and will not be repeated here.

[0183] The data blocks belonging to the mixed precision training state data can be saved by the MixedPrecisionState structure contained in the checkpoint data structure, which can be defined as follows:

[0184] typedef struct {

[0185] uint32_t enabled; / / Whether to enable mixed precision training (1 means enabled, 0 means disabled)

[0186] void *scaled_loss; / / Scaled loss, suitable for mixed precision

[0187] size_t data_size; / / Scaling loss size

[0188] } MixedPrecisionState;

[0189] Among them, the data block belonging to the mixed precision training state data may include the scaling loss data, which can be saved in the storage space pointed to by the pointer scaled_loss in the CXL memory module. Enabled and data_size are used to record the metadata of the data block. Enabled is used to record whether mixed precision training is enabled, and data_size is used to record the size of the data block.

[0190] The data blocks belonging to the momentum and gradient cache data can be saved by the MomentumAndGradientState structure of the checkpoint data structure, which can be defined as follows:

[0191] typedef struct {

[0192] void *momentum; / / Momentum cache

[0193] size_t momentum_size; / / Momentum cache size

[0194] void *gradient; / / gradient cache

[0195] size_t gradient_size; / / gradient cache size

[0196] } MomentumAndGradientState;

[0197] Among them, the two pointers momentum and gradient record the storage addresses of the data blocks stored in the CXL memory module. The metadata of the data blocks belonging to the momentum and gradient cache data may include the size of the momentum cache data (i.e., the data volume) and the size of the gradient cache data. These metadata can be recorded in the variables momentum_size and gradient_size, respectively.

[0198] In the metadata of the checkpoint data, the timestamp can be recorded in the form of a pointer in the checkpoint data structure, and the timestamp size and version can be recorded in the variables in the checkpoint data structure.

[0199] Among them, the timestamp of the checkpoint data indicates when the checkpoint data is obtained, and the timestamp can be saved through the pointer timestamp of the aforementioned checkpoint data structure. The timestamp size indicates the number of bytes contained in the timestamp of the checkpoint data, and the data can be saved using the variable timestamp_size of the checkpoint data structure. The version of the checkpoint data can be saved through the variable version of the aforementioned checkpoint data structure.

[0200] The purpose of storing the metadata of data blocks in the checkpoint data structure is to:

[0201] When loading data blocks, metadata can be used as a search condition to search for the structure corresponding to the data block with specific metadata, and then the corresponding data block is loaded from the CXL memory module through the structure, so that the specified data block can be loaded quickly, further improving the loading efficiency.

[0202] In some optional embodiments, see Figure 3 , the method of this embodiment may further include:

[0203] Create the address data structure corresponding to the checkpoint data;

[0204] The storage address of the data block in the checkpoint data and the checkpoint identifier corresponding to the checkpoint data are stored in the checkpoint address data structure.

[0205] The address data structure may include a checkpoint identifier and the storage address of each data block contained in the checkpoint data. For example, the address data structure corresponding to the checkpoint data may be recorded as CheckpointAddr, which may be defined in the following form:

[0206] typedef struct {

[0207] char checkpoint_name

[256] ; / / Checkpoint name (up to 255 characters)

[0208] CheckpointMemoryPtr memoryptr; / / Memory addresses of various states and configurations

[0209] } CheckpointAddr;

[0210] Among them, checkpoint_name

[256] is used to record the checkpoint identifier, that is, the name of the checkpoint data, and memoryptr is a structure containing the storage address of each data block. The structure can be defined as follows:

[0211] typedef struct {

[0212] void *model_params_addr; / / Model parameter memory address

[0213] void *optimizer_state_addr; / / Optimizer state memory address

[0214] void *scheduler_state_addr; / / Learning rate scheduler state memory address

[0215] void *training_progress_addr; / / Training progress status memory address

[0216] void *mixed_precision_state_addr; / / Mixed precision training state memory address

[0217] void *momentum_and_gradient_cache_addr; / / Momentum and gradient cache memory address

[0218] void *model_config_addr; / / Model configuration memory address

[0219] void *distributed_training_config_addr; / / Distributed training configuration memory address

[0220] } CheckpointMemoryPtr

[0221] by Figure 3For example, in the address data structure corresponding to checkpoint data 1, the model parameter address model_params_addr can be 0x72fd700, indicating that the data block of the model parameter data in checkpoint data 1 is stored at 0x72fd700 of the CXL memory module, the scheduler state address optimizer_state_addr can be 0x72fd500, indicating that the data block of the scheduler state data in checkpoint data 1 is stored at 0x72fd500 of the CXL memory module, and the training progress address scheduler_state_addr can be 0x72fd300, indicating that the training progress in checkpoint data 1 is The data block of the training progress data is stored at 0x72fd300 of the CXL memory module. The momentum and gradient cache address momentum_and_gradient_cache_addr can be 0x72fc200, indicating that the data block of the momentum and gradient cache data in the checkpoint data 1 is stored at 0x72fc200 of the CXL memory module. The optimizer state address optimizer_state_addr can be 0x72fc400, indicating that the data block of the optimizer state data 1 in the checkpoint data 1 is stored at 0x72fc400 of the CXL memory module. The meanings of other addresses are similar and will not be repeated here.

[0222] In some embodiments, the storage addresses of data blocks in the checkpoint data and the checkpoint identifiers corresponding to the checkpoint data may also be recorded in other forms, not limited to the form of the above-mentioned address data structure. For example, the storage addresses of data blocks in the checkpoint data and the checkpoint identifiers corresponding to the checkpoint data may also be recorded in the form of a table.

[0223] The data blocks of different checkpoint data may be the same, so the address data structures of different checkpoint data may include the same storage address.

[0224] by Figure 3 For example, the data block of the scheduler status data in checkpoint data 1 is the same as the data block of the scheduler status data in checkpoint data 2. Therefore, when obtaining checkpoint data 2, the processor does not write the data block of the scheduler status data in checkpoint data 2 to the CXL memory module, and records the storage address of the data block of the scheduler status data in checkpoint data 1 in the checkpoint data structure and address data structure corresponding to checkpoint data 2. Therefore, the scheduler status addresses of checkpoint data 1 and checkpoint data 2 are both 0x72fd500.

[0225] Similarly, the data blocks of the training progress data of checkpoint data 1 and checkpoint data 2 are the same, so the training progress addresses of both are 0x72fd300, and the data blocks of the optimizer state data of checkpoint data 1 and checkpoint data 2 are the same, so the optimizer state addresses of both are 0x72fc400. The corresponding addresses of other identical data blocks are also the same and will not be repeated here.

[0226] In contrast, the data blocks of the model parameter data of the two are different. When saving checkpoint data 2, the processor allocates other storage space in the CXL memory module to save the data blocks of the model parameter data in checkpoint data 2, so the model parameter addresses of checkpoint data 1 and checkpoint data 2 are different.

[0227] After the storage address of each data block in the checkpoint data and the checkpoint identifier of the checkpoint data are saved through the address data structure, when a certain checkpoint data needs to be loaded, the corresponding address data structure can be quickly found based on the checkpoint identifier corresponding to the checkpoint data, and then each data block can be loaded from the CXL memory module according to the storage address recorded in the address data structure, thereby improving loading efficiency.

[0228] In some embodiments, the processor may further read the data block from the CXL memory module based on the obtained load requirement in at least one of the following ways:

[0229] According to the target checkpoint identifier and the address data structure, read the data blocks contained in the target checkpoint data corresponding to the target checkpoint identifier from the CXL memory module;

[0230] According to the target block identifier and the checkpoint data structure, the target data block corresponding to the target block identifier is read from the CXL memory module.

[0231] On the one hand, see Figure 2 and Figure 4 Through the address data structure corresponding to the checkpoint data, the association relationship between the checkpoint data and each data block in the CXL memory module can be recorded. During the processing of the training task, any target checkpoint identifier determined in the training task can be obtained, and then based on the target checkpoint identifier and the address data structure of each pre-saved checkpoint data, the storage address of each data block belonging to the target checkpoint data in the CXL memory module is obtained, and then the target checkpoint data is read from the CXL memory module according to these storage addresses. The target checkpoint data refers to the checkpoint data corresponding to the target checkpoint identifier.

[0232] On the other hand, the processor can also obtain the target checkpoint identifier and the target block identifier, determine the checkpoint data structure corresponding to the target checkpoint data based on the target checkpoint identifier, and then find the structure for storing the target data block in the checkpoint data structure based on the target block identifier, and then read the target data block from the CXL memory module through the structure, wherein the target data block refers to the data block corresponding to the target block identifier.

[0233] This embodiment also provides a data storage device, see Figure 5 , the device may include:

[0234] A first processor 501, a second processor 502 and a CXL memory module 503;

[0235] The first processor 501 is used to process data according to the target large model;

[0236] The second processor 502 is used for:

[0237] Obtain checkpoint data of the target large model from the processor memory of the first processor, the checkpoint data including at least model parameter data, state data and training progress data of the target large model;

[0238] Create a checkpoint data structure corresponding to the checkpoint data, where the checkpoint data structure is used to store data blocks and block identifiers of the data blocks of the checkpoint data;

[0239] The data blocks contained in the checkpoint data are written to the CXL memory module 503 according to the checkpoint data structure, where the data blocks are obtained by splitting the checkpoint data.

[0240] Optionally, the second processor 502 splits the checkpoint data to obtain data blocks, including:

[0241] According to the data dimension of the checkpoint data and the preset data block size, the checkpoint data is split into multiple data blocks.

[0242] Optionally, when the second processor 502 writes the data blocks included in the checkpoint data to the CXL memory module according to the checkpoint data structure, it is configured to:

[0243] Determine a first data block from the data blocks included in the checkpoint data, the first data block being a data block different from the data blocks stored in the CXL memory module;

[0244] A first data block of the checkpoint data is written to the CXL memory module according to the checkpoint data structure.

[0245] Optionally, the second processor 502 is further configured to:

[0246] Obtaining a storage address and a block identifier of a second data block contained in the checkpoint data, where the second data block is the same data block as the data block stored in the CXL memory module;

[0247] The storage address and block identifier of the second data block contained in the checkpoint data are recorded in the checkpoint data structure.

[0248] Optionally, when the second processor 502 writes the first data block of the checkpoint data to the CXL memory module according to the checkpoint data structure, it is configured to:

[0249] Obtaining storage space of a first data block in a CXL memory module for storing checkpoint data;

[0250] Recording the storage address corresponding to the obtained storage space and the block identifier of the first data block of the checkpoint data in the checkpoint data structure;

[0251] A first data block of the checkpoint data is written to the obtained storage space.

[0252] Optionally, when the second processor 502 determines the first data block from the data blocks included in the checkpoint data, it is configured to:

[0253] Obtaining a first checksum value of a data block contained in the checkpoint data;

[0254] The first check value and the second check value are compared to determine the first data block based on the comparison result, wherein the second check value is a check value of the data block stored in the memory module.

[0255] Optionally, the second processor 502 is further configured to:

[0256] Recording metadata of a data block of the checkpoint data in a checkpoint data structure, wherein the metadata of the data block includes at least one of a name, a size, and a quantity of the data block;

[0257] The metadata of the checkpoint data is recorded in the checkpoint data structure, where the metadata of the checkpoint data includes at least one of a timestamp, a timestamp size, and a version of the checkpoint data.

[0258] Optionally, the second processor 502 is further configured to:

[0259] Create the address data structure corresponding to the checkpoint data;

[0260] The storage address of the data block in the checkpoint data and the checkpoint identifier corresponding to the checkpoint data are stored in the checkpoint address data structure.

[0261] Optionally, the second processor 502 is further configured to:

[0262] According to the target checkpoint identifier and the address data structure, read the data blocks contained in the target checkpoint data corresponding to the target checkpoint identifier from the CXL memory module;

[0263] According to the target block identifier and the checkpoint data structure, the target data block corresponding to the target block identifier is read from the CXL memory module.

[0264] The working principle of the data storage device of this embodiment can be found in the relevant equipment of the data storage method provided in any embodiment of the present application, and will not be repeated here.

[0265] It should be noted that the various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0266] For the convenience of description, the above system or device is described by dividing it into various modules or units according to its functions. Of course, when implementing the present application, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0267] It can be known from the description of the above implementation methods that those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present application or certain parts of the embodiments.

[0268] Finally, it should be noted that, in this article, relational terms such as first, second, third and fourth are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the statement "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.

[0269] The above is only a preferred implementation of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A data storage method, comprising: Obtaining checkpoint data of a target large model, wherein the checkpoint data includes at least model parameter data, state data, and training progress data of the target large model; Creating a checkpoint data structure corresponding to the checkpoint data, wherein the checkpoint data structure is used to store data blocks and block identifiers of the data blocks of the checkpoint data; The data blocks contained in the checkpoint data are written into the CXL memory module according to the checkpoint data structure, wherein the data blocks are obtained by splitting the checkpoint data.

2. According to the method of claim 1, the method of splitting the checkpoint data to obtain data blocks comprises: According to the data dimension of the checkpoint data and the preset data block size, the checkpoint data is split into a plurality of data blocks.

3. The method according to claim 1, wherein writing the data blocks contained in the checkpoint data to the CXL memory module according to the checkpoint data structure comprises: Determining a first data block from the data blocks included in the checkpoint data, wherein the first data block is a data block different from the data blocks stored in the CXL memory module; A first data block of the checkpoint data is written to the CXL memory module according to the checkpoint data structure.

4. The method according to claim 3, further comprising: Obtaining a storage address and a block identifier of a second data block contained in the checkpoint data, wherein the second data block is the same data block as the data block stored in the CXL memory module; The storage address and block identifier of the second data block contained in the checkpoint data are recorded in the checkpoint data structure.

5. The method of claim 3, wherein writing the first data block of the checkpoint data to the CXL memory module according to the checkpoint data structure comprises: Obtaining storage space of a first data block in the CXL memory module for storing the checkpoint data; Recording the storage address corresponding to the obtained storage space and the block identifier of the first data block of the checkpoint data in the checkpoint data structure; The first data block of the checkpoint data is written into the obtained storage space.

6. The method according to claim 3, wherein determining the first data block from the data blocks included in the checkpoint data comprises: Obtaining a first check value of a data block contained in the checkpoint data; The first check value and the second check value are compared to determine the first data block based on the comparison result, wherein the second check value is a check value of the data block stored in the memory module.

7. The method according to claim 1, further comprising at least one of the following: Recording metadata of the data blocks of the checkpoint data in the checkpoint data structure, where the metadata of the data blocks includes at least one of the name, size, and quantity of the data blocks; The metadata of the checkpoint data is recorded in the checkpoint data structure, wherein the metadata of the checkpoint data includes at least one of a timestamp, a timestamp size, and a version of the checkpoint data.

8. The method according to claim 1, further comprising: Creating an address data structure corresponding to the checkpoint data; The storage address of the data block in the checkpoint data and the checkpoint identifier corresponding to the checkpoint data are stored in the checkpoint address data structure.

9. The method according to claim 8, further comprising at least one of the following: According to the target checkpoint identifier and the address data structure, read the data blocks contained in the target checkpoint data corresponding to the target checkpoint identifier from the CXL memory module; According to the target block identifier and the checkpoint data structure, a target data block corresponding to the target block identifier is read from the CXL memory module.

10. A data storage device, comprising: A first processor, a second processor and a CXL memory module; The first processor is used to process data according to the target large model; The second processor is used for: Obtain checkpoint data of the target large model from the processor memory of the first processor, wherein the checkpoint data at least includes model parameter data, state data, and training progress data of the target large model; Creating a checkpoint data structure corresponding to the checkpoint data, wherein the checkpoint data structure is used to store data blocks and block identifiers of the data blocks of the checkpoint data; The data blocks contained in the checkpoint data are written to the CXL memory module according to the checkpoint data structure, wherein the data blocks are obtained by splitting the checkpoint data.