Model training data loading method and device

By dividing the original data file into multiple data blocks and generating metadata, quickly locate and loading the target training data, the time-consuming and training interrupt problems caused by traditional data loading methods are solved, and the efficiency of large model training is improved.

CN120045747APending Publication Date: 2025-05-27BEIJING CO WHEELS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311595890.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-27
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

In large model training, traditional data loading methods make data loading take a long time, especially in distributed training scenarios, which can easily cause interruptions in model training and affect efficiency.

Method used

By dividing the original data file into multiple data blocks and generating the first metadata and the second metadata of each data block, the target data block is determined for each training node by querying the second metadata, and the target training data is quickly positioned and loaded.

Benefits of technology

This method reduces the time-consuming data loading, improves the efficiency of model training, and avoids the problem of model training interruption caused by excessive data loading in distributed training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045747A_ABST
    Figure CN120045747A_ABST
Patent Text Reader

Abstract

The invention provides a data loading method and device for large model training, and relates to the technical field of data processing. The method comprises the following steps: acquiring an original data file, and dividing data in the original data file to obtain a plurality of data blocks corresponding to the original data file; generating first metadata corresponding to each data block in the plurality of data blocks and second metadata corresponding to the plurality of data blocks; wherein the first metadata is used for describing the data information of the data blocks corresponding to the first metadata, and the second metadata is used for describing the data information of the plurality of data blocks; for each training node, determining a target data block corresponding to the training node by querying second metadata; and determining target training data in the target data block through the first metadata corresponding to the target data block, so as to input the target training data corresponding to the training node into the training node for model training, thereby solving the problem of low training efficiency during large model training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing. Specifically, it relates to a method and device for loading model training data. Background Art

[0002] With the rapid development of deep learning, the maturity of models, model parameters, etc. are increasing, and the amount of data required for model training is also growing exponentially. Model training has entered the era of large model distributed training.

[0003] When performing model training, traditional data loading methods include: loading the data required for training into memory, preprocessing it to form a data set, and the training nodes completing model training by traversing the data set in memory. However, when the amount of data required for model training is huge, the time-consuming for loading data based on this traditional method is relatively large, and in a distributed training scenario, it is easy to cause the interruption of the model training process, affecting the efficiency of model training. In a distributed training scenario, each training node usually obtains the corresponding training data through a file handle. However, when the training data of a training node is in the middle of the data set, it is necessary to calculate the starting row number of the training data and skip the previous rows to obtain the corresponding training data, which is time-consuming and affects the efficiency of large model training.

[0004] Therefore, when the training data set is large, how to improve the training efficiency of large model training is an urgent problem to be solved. Summary of the Invention

[0005] To solve the problem of low training efficiency during large model training, embodiments of this application provide a method and device for loading data for large model training, an electronic device, and a computer-readable storage medium.

[0006] In a first aspect, embodiments of this application provide a method for loading data for large model training, including:

[0007] Obtain an original data file, and divide the data in the original data file to obtain a plurality of data blocks corresponding to the original data file;

[0008] Generate first metadata corresponding to each of the plurality of data blocks, and second metadata corresponding to the plurality of data blocks; wherein, the first metadata is used to describe the data information of the data block corresponding to the first metadata, and the second metadata is used to describe the data information of the plurality of data blocks;

[0009] For each training node, determine the target data block corresponding to the training node by querying the second metadata;

[0010] Determine target training data in the target data block through the first metadata corresponding to the target data block, so as to input the target training data corresponding to the training node into the training node for model training.

[0011] As an optional implementation manner of an embodiment of the present application, the determining the target data block corresponding to the training node by querying the second metadata includes:

[0012] Construct a data block index based on the data information of the multiple data blocks described in the second metadata;

[0013] Obtain the target data block corresponding to the training node by querying the data block index.

[0014] As an optional implementation manner of an embodiment of the present application, the determining the target data block corresponding to the training node by querying the second metadata includes:

[0015] Obtain the process position of the training process corresponding to the training node in the training process of the large model training, and calculate the data area where the training data corresponding to the training node is located according to the process position;

[0016] Determine at least one data block where the training data in the data area is located by querying the second metadata, and determine the at least one data block as the target data block corresponding to the training node.

[0017] As an optional implementation manner of an embodiment of the present application, after determining the target training data in the target data block through the first metadata corresponding to the target data block, the method further includes:

[0018] Determine the prefetch data volume according to the training parameters of the training node, and load the target training data of the prefetch data volume into the preset cache space corresponding to the training node;

[0019] Read the target training data from the preset cache space and input it into the training node for model training.

[0020] As an optional implementation manner of an embodiment of the present application, the method further includes:

[0021] After all the target training data in the preset cache space is read, clear the preset cache space;

[0022] Read the target training data of the prefetch data volume from the target training data again and load it into the preset cache space, so that the training node reads the target training data from the preset cache space to continue model training.

[0023] As an optional implementation manner of an embodiment of the present application, the dividing the data in the original data file to obtain multiple data blocks corresponding to the original data file includes:

[0024] Performing serialization processing on the original data file to obtain serialized data, where the serialized data is data recognizable during large model training;

[0025] Performing chunking processing on the serialized data according to the data chunks of the original data file to obtain multiple data blocks corresponding to the original data file.

[0026] As an optional implementation manner of an embodiment of the present application, it is characterized in that the first metadata includes one or more of the file format, file name, file version, and data volume in the data block; the second metadata includes one or more of the name of the original data file, the version of the original data file, the number of data blocks, and the identifier of the data block.

[0027] In a second aspect, an embodiment of the present application provides a data loading device for large model training, including:

[0028] An acquisition module, configured to acquire an original data file and divide the data in the original data file to obtain multiple data blocks corresponding to the original data file;

[0029] A processing module, configured to generate first metadata corresponding to each of the multiple data blocks and second metadata corresponding to the multiple data blocks; wherein, the first metadata is used to describe the data information of the data block corresponding to the first metadata, and the second metadata is used to describe the data information of the multiple data blocks;

[0030] A query module, for each training node, configured to determine a target data block corresponding to the training node by querying the second metadata;

[0031] A loading module, configured to determine target training data in the target data block through the first metadata corresponding to the target data block, so as to input the target training data into the training node for model training.

[0032] As an optional implementation manner of an embodiment of the present application, the query module is specifically configured to construct a data block index based on the data information of the multiple data blocks described by the second metadata;

[0033] By querying the data block index, obtaining the target data block corresponding to the training node.

[0034] As an optional implementation manner of an embodiment of the present application, the query module is specifically configured to obtain the process position of the training process corresponding to the training node in the training process of the large model training, and calculate the data area where the training data corresponding to the training node is located according to the process position;

[0035] By querying the second metadata, determine at least one data block where the training data in the data area is located, and determine the at least one data block as the target data block corresponding to the training node.

[0036] As an optional implementation manner of an embodiment of the present application, the device further includes:

[0037] A prefetch module, configured to, after obtaining target training data in the target data block by querying the first metadata corresponding to the target data block, determine a prefetch data volume according to the training parameters of the training node, and load the target training data of the prefetch data volume into a preset cache space corresponding to the training node;

[0038] The loading module is specifically configured to read the target training data from the preset cache space and input it into the training node for model training.

[0039] As an optional implementation manner of an embodiment of the present application, the device further includes:

[0040] An iteration module, configured to clear the preset cache space after all the target training data in the preset cache space is read;

[0041] Read the target training data of the prefetch data volume from the target training data again and load it into the preset cache space, so that the training node reads the target training data from the preset cache space to continue model training.

[0042] As an optional implementation manner of an embodiment of the present application, the acquisition module is specifically configured to perform serialization processing on the original data file to obtain serialized data, and the serialized data is data recognizable during large model training;

[0043] Perform block processing on the serialized data according to the data blocks of the original data file to obtain multiple data blocks corresponding to the original data file.

[0044] As an optional implementation manner of an embodiment of the present application, the first metadata includes one or more of the file format, file name, file version, and data volume in the data block; the second metadata includes one or more of the name of the original data file, the version of the original data file, the number of data blocks, and the identifier of the data block.

[0045] In a third aspect, an embodiment of the present application provides an electronic device, including: a memory and a processor. The memory is used to store a computer program, and the processor is used to execute the data loading method for large model training according to the first aspect or any optional implementation manner of the first aspect when calling the computer program.

[0046] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the data loading method for large model training according to the first aspect or any optional implementation manner of the first aspect.

[0047] The technical solutions provided by the embodiments of the present application have the following advantages compared with the prior art:

[0048] The data loading method for large model training provided by the embodiments of the application includes: obtaining an original data file, and processing the original data file to obtain a plurality of data blocks corresponding to the original data file; generating first metadata respectively corresponding to each of the plurality of data blocks, and second metadata corresponding to the plurality of data blocks; wherein, the first metadata is used to describe the data information of the data block corresponding to the first metadata, and the second metadata is used to describe the data information of the plurality of data blocks; by querying the first metadata and the second metadata, obtaining target training data respectively corresponding to each training node in the large model training; for each training node, inputting the target training data corresponding to the training node into the training node for model training. In the embodiments of the present application, by dividing the original data file into a plurality of data blocks, and querying the first metadata respectively corresponding to the plurality of data blocks and the second metadata corresponding to the original data file, the target training data corresponding to each training node can be obtained, so that the data area to be read can be quickly located, the time consumption of loading data is reduced, and the training efficiency of model training is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] The drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0050] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0051] Figure 1 It is a flowchart of the data loading method for large model training according to one or more embodiments of the present application;

[0052] Figure 2 Schematic diagram of the data structure of multiple data blocks provided according to one or more embodiments of the present application;

[0053] Figure 3 Schematic diagram of the logic for traversing data blocks provided according to one or more embodiments of the present application;

[0054] Figure 4 Flowchart of the data loading method for large model training provided according to one or more embodiments of the present application;

[0055] Figure 5 Structural block diagram of the data loading device for large model training provided according to one or more embodiments of the present application;

[0056] Figure 6 Structural block diagram of the data loading device for large model training provided according to one or more embodiments of the present application;

[0057] Figure 7 Internal structure diagram of the in-vehicle terminal provided according to one or more embodiments of the present application. Detailed implementation manners

[0058] To make the purpose, implementation manners and advantages of the present application clearer, the following will clearly and completely describe the exemplary implementation manners of the present application in conjunction with the accompanying drawings in the exemplary embodiments of the present application. Obviously, the described exemplary embodiments are only a part of the embodiments of the present application, rather than all the embodiments.

[0059] Based on the exemplary embodiments described in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope protected by the appended claims of the present application. In addition, although the disclosed content in the present application is introduced according to exemplary one or several examples, it should be understood that each aspect of these disclosed contents can also be independently constituted as a complete implementation manner. It should be noted that the brief description of the terms in the present application is only for the convenience of understanding the subsequent described implementation manners, rather than intending to limit the implementation manners of the present application. Unless otherwise specified, these terms should be understood in their ordinary and general meanings.

[0060] First, the application scenario of the embodiments of the present application is described: With the rapid development of deep learning, model training has entered the era of large models. The scale of the model itself, such as the number of model layers and the number of model parameters, is increasing, and the amount of data for model training is also increasing. Model training first requires loading the training data. The traditional data loading methods mainly include: First, load all the original files into memory, perform data preprocessing in memory to generate a training dataset, and the training model traverses the dataset in memory to perform model training. Second, for the original file, by maintaining a file handle and multiple counting variables, when reading data, control information such as the number of lines to skip, the number of lines already read, and the current line number, and gradually load to complete model training. However, when the file size is large, the preloading time is too long when loading the file into memory in advance, that is, the time from starting the training program to starting training is too long. In distributed training, the probability of cluster communication timeout may be relatively large, which easily leads to training failure; a single host will start multiple training processes, and each training process needs to perform data loading, doubling the memory requirement, which easily causes the host OOM Killer (Linux Out of Memory killer linux) to kill the training process and cause training interruption. By maintaining the file handle and multiple counting variables, the time-consuming for controlling the number of lines to skip is also related to the file size, which easily leads to cluster timeout and affects the efficiency of model training.

[0061] Based on this, the embodiments of the present application provide a data loading method for large model training. By dividing the original data file into multiple data blocks and querying the first metadata corresponding to each data block and the second metadata corresponding to the original data file, the target training data corresponding to each training node can be obtained, so as to quickly locate the data to be read, reduce the time-consuming for loading data, and improve the training efficiency of model training.

[0062] The following will illustrate the data loading method for large model training provided by the present application through specific exemplary embodiments.

[0063] Figure 1 For the flowchart of the data loading method for large model training provided by the embodiments of the present application, refer to Figure 1 As shown, the data loading method for large model training provided in this embodiment includes the following steps S11 to S14:

[0064] S11. Obtain the original data file, and divide the data in the original data file to obtain multiple data blocks corresponding to the original data file.

[0065] Exemplarily, the original file can be a json file, and each line of data in the json file can be used as one piece of data. The original data file can also be a picture, voice, etc.

[0066] After obtaining the original data file, the data in the original file can be processed as follows:

[0067] Perform serialization processing on the original data file to obtain serialized data, which is data recognizable during large model training. After obtaining the serialized data, the serialized data needs to be partitioned, that is, block processed, to obtain multiple data blocks corresponding to the serialized data, that is, multiple data blocks corresponding to the above original data file.

[0068] Among them, as a way of block processing: according to the data blocks of the original data file, the serialized data is block processed to obtain multiple data blocks corresponding to the original data file, that is, if the original data file has a certain number of data blocks, the serialized data is processed into the corresponding number of data blocks, and the data included in each processed data block corresponds to the data in each data block of the original data file.

[0069] As another way of block processing: according to the number of data entries in the original data file, the serialized data is divided into multiple data blocks, that is, each divided data block contains a preset number of data entries.

[0070] Among them, multiple processing processes can be used to perform serialization processing on the original data file. Exemplarily, when the original data file is a json file, by traversing the json file directory, the number of json files is determined, and the number of files is used as a parallelism parameter to start multiple processing processes.

[0071] Exemplarily, for each processing process, traverse the json file line by line, convert the json text into sequence data recognizable by the model, and save it as a separate file. Binary serialization can be used, and the serialization method can be configured, and the file content can be encrypted as needed.

[0072] In the embodiments of the present application, the format of the data block supports extensible design, that is, the data block can be stored in multiple different formats. For example, the data block can be stored in the way of pickling the in-memory type into binary data. This method is simple and practical and is supported by a large number of tools. However, in the case of other extremely large amounts of data, this format may use too much storage space and provide low query and loading performance. Therefore, a plug-in multi-format support can be designed, and the data generation party can select the current high-performance or high-compression storage format.

[0073] In the embodiments of the present application, after the training data is divided into data blocks, when modifying the training data set, only local modification is required, that is, modification is performed on specific data blocks, or data blocks are added or reduced, which greatly reduces the time required for generating the data set and saves computing and storage resources.

[0074] S12. Generate first metadata corresponding to each of the multiple data blocks, and second metadata corresponding to the multiple data blocks.

[0075] Wherein, the first metadata is used to describe the data information of the data block corresponding to the first metadata, and the second metadata is used to describe the data information of the multiple data blocks.

[0076] That is, the second metadata is used to describe the data information of all data blocks, such as the number of data blocks, the data block identifier of each data block, the data block version, the dataset name, etc. Exemplarily, referring to Figure 2 as shown Figure 2 FIG. is a schematic diagram of the data structure of multiple data blocks in an embodiment of the present application. The data information in each data block is described by the first metadata corresponding to the data block, and the data information of all data blocks and the information of all first metadata are described by the second metadata. For example, data block 1 corresponds to first metadata A, data block 2 corresponds to first metadata B, and data block M corresponds to first metadata N.

[0077] Wherein, the first metadata may include one or more of: file format, file name, file version, data volume in the data block; the second metadata includes one or more of: the name of the original data file, the version of the original data file, the number of data blocks, the identifier of the data block.

[0078] That is, the first metadata describes the file format, file name, file version, data volume (number of data entries), etc. of the data block corresponding to the first metadata, and attributes can be extended as needed to facilitate quickly obtaining information such as the internal storage format of the file block without reading the file block. The second metadata summarizes multiple first metadata, combines the scattered metadata blocks to form an overall dataset, that is, serialized data, and describes the name of the serialized data, the number of data blocks, the identification information of the data blocks, the file version of the serialized data, the data volume of the serialized data, the number of entries in each data block, etc.

[0079] The multiple data blocks can adopt a microkernel architecture to form a block dataset operation library, providing capabilities such as length query, appending file blocks, and fast moving access index. It can support more file formats through extension, improve the applicable scenarios of the block dataset, and solve the problems of sequential reading, sequential production, and the need to regenerate when modifying common datasets. For example, a training dataset can be added to the training dataset by adding a data block at the end. Since only the metadata is modified, it can be considered that the time consumption is almost zero, realizing the convenience of reducing the dataset. Combined with supporting tools or scripts, the training dataset can be very flexibly combined.

[0080] S13. For each training node, determine the target data block corresponding to the training node by querying the second metadata.

[0081] Exemplarily, as an implementable manner of the above step S13, it includes: constructing a data block index based on the data information of the multiple data blocks described in the second metadata; obtaining the target data block corresponding to the training node by querying the data block index.

[0082] Exemplarily, by obtaining the process position of the training process corresponding to the training node in the training process of the large model training, calculate the data area where the training data corresponding to the training node is located according to the process position; by querying the second metadata or the data block index, determine at least one data block where the training data in the data area is located, and determine the at least one data block as the target data block corresponding to the training node. For example, query the second metadata through the info query description information, and determine the target data block where the training data is located through the identification information of the data block in the second metadata. The training data may be located in the same data block or in different data blocks. Use the move_to() method of the chunk data loader to quickly skip the data that does not need to be read and locate to the data area responsible for this process.

[0083] S14. Determine the target training data in the target data block through the first metadata corresponding to the target data block, so as to input the target training data into the training node for model training.

[0084] Among them, after determining the target data, the move_to(index) function can be used to quickly locate to the specific data entry to start traversing and read the target training data. That is, since the serialized data is divided into multiple data blocks, and the data volume in the database is much smaller than the total data volume of the serialized data, after quickly locating to the specific target data block through the data block index, traverse the specified number of rows in the target data block, reducing the loading time of the training data. In addition, when the structure of the data block is a simple structure such as pickle that appends by row, data block formats such as arrow or lmdb that support columnar or B-tree can be used, which can further compress the loading time of the training data.

[0085] Exemplarily, in a distributed training scenario, start the distributed large model training, schedule the training program, configure the training environment, pull the training code, and specify the environment variables, where the environment variables describe information such as the node numbers of each training node and the total number of nodes in the cluster.

[0086] For each training node, at least one training process is started according to the number of graphics cards. For each training process, the training cluster information is initialized, that is, communication is established with the master node in the training cluster. After the training node joins the training cluster, a chunk data loader is started to read the target training data item by item for model training.

[0087] When a super-large dataset is loaded and used for training, in the embodiments of the present application, since multiple training nodes obtain the target training data in their respective target databases, the situation of excessive data loading time caused by the growth of the data scale and failure due to excessive memory occupied by the data is completely avoided.

[0088] To ensure the least modification to the training code and minimize the impact of the chunk structure of the training dataset on the training code, an operation library for chunk datasets corresponding to multiple data chunks can be designed to implement a doubly linked list for indexing and traversing data. Refer to Figure 3 As shown, it is a logical schematic diagram for traversing data chunks provided by an embodiment of the present application, where the head and the tail are used to indicate the order of reading data chunks.

[0089] The data loading method for large model training provided by the embodiments of the present application includes: obtaining an original data file, and partitioning the data in the original data file to obtain multiple data chunks corresponding to the original data file; generating first metadata respectively corresponding to each of the multiple data chunks, and second metadata corresponding to the multiple data chunks; where the first metadata is used to describe the data information of the data chunk corresponding to the first metadata, and the second metadata is used to describe the data information of the multiple data chunks; for each training node, by querying the second metadata, determining the target data chunk corresponding to the training node; determining the target training data in the target data chunk through the first metadata corresponding to the target data chunk, so as to input the target training data corresponding to the training node into the training node for model training. In the embodiments of the present application, by partitioning the original data file into multiple data chunks and querying the first metadata corresponding to each of the multiple data chunks and the second metadata corresponding to the original data file, the target training data corresponding to each training node can be obtained, so that the data area to be read can be quickly located, reducing the time-consuming of data loading and improving the training efficiency of model training.

[0090] Figure 4 It is a flowchart of a fault handling method provided by another embodiment of the present application. Refer to Figure 4 As shown, on the basis of the embodiment shown in Figure 1 After step S14, the following steps S41 to S42 are further included.

[0091] S41. Determine the prefetch data volume according to the training parameters of the training node, and load the target training data of the prefetch data volume into the preset cache space corresponding to the training node.

[0092] Among them, the prefetch data volume can be specified by the set buffersize. A remote data cache can be set. When the data set is large and stored in low-performance storage, such as a common hard disk or an FTP / NFS server, the concurrent read pressure is relatively large. For scenarios where there are many small data blocks and frequent repeated accesses, the local SSD disk can be set as the preset cache space to improve the access speed of duplicate data.

[0093] For each training node, step S52 can be an implementation manner of step S14.

[0094] S42. Read the target training data from the preset cache space and input it into the training node for model training.

[0095] The training code can use an iterator to iterate through the data one by one from the preset cache space for training, which can effectively avoid the huge memory consumption caused by loading all data into the memory.

[0096] After all the target training data in the preset cache space has been read, clear the preset cache space, and then read the target training data of the prefetch data volume from the target training data again and load it into the preset cache space, so that the training node can continue to perform model training by reading the target training data from the preset cache space. That is, maintain a prefetch data pointer in the memory. When the prefetch cache is read out, the memory is automatically released and the next part of the data is prefetched. This continuous reading of data can make full use of disk characteristics, accelerate reading, and reduce the number of IOs.

[0097] Based on the same inventive concept, as an implementation of the above method, the embodiments of the present application also provide a data loading device for large model training that executes the above embodiments. The device embodiments correspond to the foregoing method embodiments. For the convenience of reading, the details in the foregoing method embodiments will not be described one by one in the device embodiments of the present application. However, it should be clear that the data loading method device for large model training in this embodiment can correspondingly implement all the content in the foregoing method embodiments.

[0098] Figure 5 Shown in the structure diagram of the data loading device for large model training provided by an embodiment of the present application, as Figure 5 shown, the data loading device 500 for large model training provided by this embodiment includes:

[0099] An acquisition module 510, configured to acquire an original data file and partition the data in the original data file to obtain a plurality of data blocks corresponding to the original data file.

[0100] A processing module 520, configured to generate first metadata corresponding to each of the plurality of data blocks and second metadata corresponding to the plurality of data blocks; wherein, the first metadata is used to describe the data information of the data block corresponding to the first metadata, and the second metadata is used to describe the data information of the plurality of data blocks.

[0101] A query module 530, configured to, for each training node, determine a target data block corresponding to the training node by querying the second metadata.

[0102] A loading module 540, configured to determine target training data in the target data block through the first metadata corresponding to the target data block, and input the target training data into the training node for model training.

[0103] As an optional implementation manner of an embodiment of the present application, the query module 530 is specifically configured to construct a data block index based on the data information of the plurality of data blocks described by the second metadata; and obtain the target data block corresponding to the training node by querying the data block index.

[0104] As an optional implementation manner of an embodiment of the present application, the query module 530 is specifically configured to obtain the process position of the training process corresponding to the training node in the training process of the large model training, calculate the data area where the training data corresponding to the training node is located according to the process position; determine at least one data block where the training data in the data area is located by querying the second metadata, and determine the at least one data block as the target data block corresponding to the training node.

[0105] Figure 6 The structure diagram of the data loading device for large model training provided by another embodiment of the present application. On the basis of the Figure 5 shown device, it further includes:

[0106] A prefetch module 610, configured to, after determining the target training data in the target data block through the first metadata corresponding to the target data block, determine the prefetch data volume according to the training parameters of the training node, and load the target training data of the prefetch data volume into the preset cache space corresponding to the training node.

[0107] The loading module 540 is specifically configured to read the target training data from the preset cache space and input it into the training node for model training.

[0108] As an alternative implementation manner of the embodiment of the present application, the apparatus further includes:

[0109] An iteration module 620, configured to clear the preset cache space after all the target training data in the preset cache space is read; and read the target training data with a prefetch data volume from the target training data again and load it into the preset cache space, so that the training node reads the target training data from the preset cache space to continue model training.

[0110] As an alternative implementation manner of the embodiment of the present application, the obtaining module 610 is specifically configured to perform serialization processing on the original data file to obtain serialized data, where the serialized data is data recognizable during large model training; and perform chunking processing on the serialized data according to the data blocks of the original data file to obtain a plurality of data blocks corresponding to the original data file.

[0111] The data loading apparatus for large model training provided by the embodiment of the present application can execute the data loading method for large model training provided by the above method embodiment, and its implementation principle and technical effect are similar, which will not be elaborated here. Each module in the above data loading apparatus for large model training can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor in the computer device in hardware form or independent of it, or stored in the memory in the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0112] In one embodiment, an electronic device is provided, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of any one of the data loading methods for large model training described in the above method embodiment.

[0113] Exemplarily, Figure 7 is a schematic structural diagram of an in-vehicle terminal provided by the embodiment of the present application. As Figure 7 shown, the electronic device provided in this embodiment includes: a memory 71 and a processor 72. The memory 71 is used to store a computer program; the processor 72 is used to execute the steps in the data loading method for large model training provided by the above method embodiment when calling the computer program, and its implementation principle and technical effect are similar, which will not be elaborated here. Those skilled in the art can understand that Figure 7 the structure shown in

[0114] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the data loading method for large model training described in any one of the above method embodiments are implemented.

[0115] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, database, or other medium used in the various embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static random access memory (SRAM) and dynamic random access memory (DRAM), etc.

[0116] It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0117] For the sake of convenience of explanation, the above description has been made in conjunction with specific embodiments. However, the above discussion in some embodiments is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. According to the above teachings, various modifications and variations can be obtained. The selection and description of the above embodiments are for better explaining the principles and actual applications, so that those skilled in the art can better use the embodiments and various different variations of the embodiments suitable for specific use considerations.

Claims

1. A data loading method for large model training, It is characterized in that include: Acquire an original data file, and divide the data in the original data file to obtain a plurality of data blocks corresponding to the original data file; Generate first metadata corresponding to each data block of the multiple data blocks, and second metadata corresponding to the multiple data blocks; wherein the first metadata is used to describe data information of the data block corresponding to the first metadata, and the second metadata is used to describe data information of the multiple data blocks; For each training node, determining a target data block corresponding to the training node by querying the second metadata; Target training data is determined in the target data block by using the first metadata corresponding to the target data block, so as to input the target training data into the training node for model training.

2. The method according to claim 1, It is characterized in that The step of determining the target data block corresponding to the training node by querying the second metadata includes: Building a data block index based on the data information of the plurality of data blocks described by the second metadata; By querying the data block index, the target data block corresponding to the training node is obtained.

3. The method according to claim 1, It is characterized in that The step of determining the target data block corresponding to the training node by querying the second metadata includes: Obtaining a process position of the training process corresponding to the training node in the training process of the large model training, and calculating a data area where the training data corresponding to the training node is located according to the process position; By querying the second metadata, at least one data block where the training data of the data area is located is determined, and the at least one data block is determined as a target data block corresponding to the training node.

4. The method according to claim 1, It is characterized in that After determining the target training data in the target data block by using the first metadata corresponding to the target data block, the method further includes: Determine the amount of pre-fetched data according to the training parameters of the training node, and load the target training data of the pre-fetched data amount into a preset cache space corresponding to the training node; The target training data is read from the preset cache space and input into the training node for model training.

5. The method according to claim 4, It is characterized in that The method further comprises: After all the target training data in the preset cache space is read, clearing the preset cache space; The target training data of the pre-fetched data amount is read again from the target training data and loaded into the preset cache space, so that the training node reads the target training data from the preset cache space to continue model training.

6. The method according to claim 1, It is characterized in that The dividing the data in the original data file to obtain a plurality of data blocks corresponding to the original data file includes: Serializing the original data file to obtain serialized data, where the serialized data is identifiable during large model training; The serialized data is divided into blocks according to the data blocks of the original data file to obtain a plurality of data blocks corresponding to the original data file.

7. The method according to any one of claims 1 to 6, It is characterized in that The first metadata includes: one or more of the file format, file name, file version, and data volume in the data block; the second metadata includes: one or more of the name of the original data file, the version of the original data file, the number of data blocks, and the identification of the data blocks.

8. A data loading device for large model training, It is characterized in that include: An acquisition module, used for acquiring an original data file, and dividing the data in the original data file to obtain a plurality of data blocks corresponding to the original data file; a processing module, configured to generate first metadata corresponding to each data block of the plurality of data blocks, and second metadata corresponding to the plurality of data blocks; wherein the first metadata is used to describe data information of the data block corresponding to the first metadata, and the second metadata is used to describe data information of the plurality of data blocks; A query module, for each training node, used to determine a target data block corresponding to the training node by querying the second metadata; A loading module is used to determine target training data in the target data block through the first metadata corresponding to the target data block, so as to input the target training data into the training node for model training.

9. An electronic device, include: A memory and a processor, wherein the memory stores a computer program, and wherein the processor implements the data loading method for large model training described in any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium having a computer program stored thereon, It is characterized in that When the computer program is executed by a processor, the data loading method for large model training described in any one of claims 1 to 7 is implemented.

11. A vehicle, It is characterized in that The vehicle is equipped with a data loading device for large model training as described in claim 8, or an electronic device as described in claim 9, or a storage medium as described in claim 10.