Large model training acceleration method and device, equipment and storage medium
Through the cache system, the cache system performs layered cache according to the access frequency of data blocks, the time-consuming download problem in large model training is solved, and efficient data processing and training performance is improved.
Patent Information
- Application Number
- CN202510483660.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-08
AI Technical Summary
During the training of large-scale models, downloading models and data sets from remote storage pools takes time and consumes a lot of network bandwidth, affecting training efficiency.
The cache system determines the access frequency of data blocks, and hierarchical cache strategy, caches small data blocks with high access frequency to memory, and caches large data blocks with low access frequency to disk, and optimizes data access using distributed consistency protocols and cache management tools.
Reduces remote download time, improves data processing efficiency and overall training performance of large-scale model training, and improves the stability and scalability of the cache system.
Smart Images

Figure CN120276680A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and particularly to a method, device, equipment, and storage medium for accelerating large model training. Background Art
[0002] During the large model training process, it is usually necessary to download the model and dataset from remote storage pools such as Amazon S3 (Amazon Simple Storage Service), Alibaba Cloud OSS (Object Storage Service), and Hadoop HDFS (the core distributed file system of the Hadoop ecosystem) to the local GPU (Graphics Processing Unit) server for training. Since the model and dataset are usually large in scale, the download process is both time-consuming and occupies a large amount of network bandwidth, which has a significant impact on the training efficiency. Therefore, how to accelerate the large model training process is an urgent problem to be solved currently. Summary of the Invention
[0003] In view of this, the purpose of the present invention is to provide a method, device, equipment, and storage medium for accelerating large model training, which can accelerate the large model training process. The specific solutions are as follows:
[0004] In a first aspect, the present application discloses a method for accelerating large model training, which is applied to a caching system and includes:
[0005] Determine the access frequency corresponding to each target data block, and classify the target data blocks based on the access frequency to determine whether the target data block is target type data; the access frequency corresponding to the target type data is higher than that corresponding to other type data, and the data block size of the target type data is smaller than a preset data block size threshold;
[0006] If the target data block is target type data, cache the target data block into a preset memory space so as to read the target data block from the preset memory space to train the target large model;
[0007] If the target data block is not target type data, cache the target data block into a preset disk space, and use a cache management tool to manage the target data block cached in the preset disk space so as to read the target data block from the preset disk space to train the target large model.
[0008] Optionally, before determining the access frequency corresponding to each target data block, it further includes:
[0009] Download the target large model and the target data set from the remote storage pool to the local GPU server;
[0010] Determine the target training tasks corresponding to the target nodes of the local GPU server through load balancing;
[0011] Split the target data set based on the target training tasks to obtain each target data block;
[0012] Use the data access pattern and historical access records, and / or use the target hot data prediction algorithm to determine the target type data corresponding to each target training task from all the target data blocks;
[0013] Based on the correspondence between the target nodes and the target training tasks, cache the target type data corresponding to each target training task into the preset memory space corresponding to the respective target nodes;
[0014] Cache the target data set into the preset disk space corresponding to each target node;
[0015] Wherein, the target node is a node on the local GPU server for training the target large model.
[0016] Optionally, if the target data block is target type data, cache the target data block into the preset memory space so as to read the target data block from the preset memory space to train the target large model, including:
[0017] If the target data block is target type data on the target node, cache the target data block into the preset memory space corresponding to the target node so as to read the target data block from the preset memory space corresponding to the target node when using the target node to train the target large model;
[0018] Correspondingly, if the target data block is not target type data, cache the target data block into the preset disk space and use a cache management tool to manage the target data block cached in the preset disk space so as to read the target data block from the preset disk space to train the target large model, including:
[0019] If the target data block is not of the target type data on the target node, cache the target data block in the preset disk space corresponding to the target node, and use a cache management tool to manage the target data block cached in the preset disk space, so that when training the target large model using the target node, read the target data block from the preset disk space corresponding to the target node.
[0020] Optionally, the large model training acceleration method further includes:
[0021] If there is data content in the target type data corresponding to the target node that has been modified, use a preset distributed consistency protocol to synchronize the modified data content to the target node that meets the modification synchronization condition.
[0022] Optionally, determining the access frequency corresponding to each target data block includes:
[0023] Use a preset frequency determination algorithm to track the access frequency corresponding to each target data block in real time;
[0024] Wherein, the preset frequency determination algorithm includes the least recently used algorithm and the least frequently used algorithm.
[0025] Optionally, downloading the target large model and target dataset from the remote storage pool to the local GPU server includes:
[0026] Use a preset application programming interface to connect the cache system with each remote storage pool;
[0027] Download the target large model and target dataset from each remote storage pool to the local GPU server;
[0028] Perform format conversion on the target dataset in the local GPU server based on the target format, so as to perform segmentation on the format-converted target dataset based on the target training task to obtain each target data block.
[0029] Optionally, if the target data block is of the target type data, caching the target data block in the preset memory space includes:
[0030] If the target data block is of the target type data, use a first preset data compression algorithm to cache the target data block in the preset memory space;
[0031] Correspondingly, if the target data block is not of the target type data, caching the target data block in the preset disk space includes:
[0032] If the target data block is not of the target type, the target data block is cached in a preset disk space by using a second preset data compression algorithm.
[0033] In a second aspect, the present application discloses a large model training acceleration device, which is applied to a caching system and includes:
[0034] A data type determination module, configured to determine the access frequency corresponding to each target data block, and classify the target data blocks based on the access frequency to determine whether the target data block is of the target type; the access frequency corresponding to the target type data is higher than the access frequencies corresponding to other types of data, and the data block size of the target type data is smaller than a preset data block size threshold;
[0035] A first data caching module, configured to cache the target data block in a preset memory space if the target data block is of the target type, so as to read the target data block from the preset memory space to train a target large model;
[0036] A second data caching module, configured to cache the target data block in a preset disk space if the target data block is not of the target type, and manage the target data block cached in the preset disk space by using a cache management tool, so as to read the target data block from the preset disk space to train the target large model.
[0037] In a third aspect, the present application discloses an electronic device, including:
[0038] A memory, configured to store a computer program;
[0039] A processor, configured to execute the computer program to implement the foregoing large model training acceleration method.
[0040] In a fourth aspect, the present application discloses a computer-readable storage medium, configured to store a computer program, wherein the computer program, when executed by a processor, implements the foregoing large model training acceleration method.
[0041] In this application, when training a large model, the caching system determines the access frequency corresponding to each target data block, and classifies the target data blocks based on the access frequency to determine whether the target data blocks are target type data; the access frequency corresponding to the target type data is higher than that of other type data, and the data block size of the target type data is smaller than a preset data block size threshold; if the target data block is target type data, the target data block is cached in a preset memory space so as to read the target data block from the preset memory space to train the target large model; if the target data block is not target type data, the target data block is cached in a preset disk space, and a cache management tool is used to manage the target data block cached in the preset disk space so as to read the target data block from the preset disk space to train the target large model. It can be seen that after determining the access frequency corresponding to each target data block, the caching system in this application determines the target type data, that is, small-scale hot data, based on the access frequency and the size of each target data block. Then, by combining the hierarchical caching strategy of small-scale data and large-scale data, the target type data is stored in the preset disk space, and other type data is stored in the preset disk space. Therefore, when training the large model, if the data access pattern changes, the caching system can automatically transfer the frequently accessed target data blocks from the disk cache to the memory. The data in the memory can be quickly called, and there is no need to repeat the download during each training, reducing the remote download time, significantly improving the data processing efficiency and overall training performance during the large-scale model training process, and enhancing the stability and scalability of the caching system. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0043] Figure 1 It is a flowchart of a method for accelerating large model training disclosed in this application;
[0044] Figure 2 It is a schematic diagram of the architecture of a caching system disclosed in this application;
[0045] Figure 3 It is a schematic diagram of the structure of a device for accelerating large model training disclosed in this application;
[0046] Figure 4 It is a structural diagram of an electronic device disclosed in this application. Detailed Implementation Modes
[0047] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0048] During the training process of large models, it is usually necessary to download the model and the dataset from a remote storage pool (such as Amazon S3, Alibaba Cloud OSS, Hadoop HDFS, etc.) to a local GPU server for training. Since the model and the dataset are usually huge in size, the download process is both time-consuming and occupies a large amount of network bandwidth, which has a significant impact on the training efficiency. To solve the above technical problems, this application discloses a large model training acceleration method, which can accelerate the process of large model training.
[0049] See Figure 1 As shown, an embodiment of the present invention discloses a large model training acceleration method, which is applied to a cache system and includes:
[0050] Step S11: Determine the access frequency corresponding to each target data block, and classify the target data blocks based on the access frequency to determine whether the target data blocks are target type data; the access frequency corresponding to the target type data is higher than the access frequency corresponding to other type data, and the data block size of the target type data is smaller than a preset data block size threshold.
[0051] In this embodiment, the target large model can be a large-scale model for realizing market trend prediction. By accelerating the training iteration process of the target large model, a large amount of market data can be analyzed to support the adjustment of real-time marketing strategies; it can also be a large-scale model for urban governance and public services, such as a public security early warning large model. By improving the training efficiency of this large model, the deployment of an abnormal event recognition system can be accelerated. In addition, the target large model can also be a large-scale model for drug research and development, such as a protein structure prediction model. By improving the training efficiency of the target large model and shortening the training cycle, the acceleration of the new drug development process can be achieved.
[0052] In this embodiment, considering that the memory space is small and large-scale data usually cannot be fully loaded into the memory, the NVMe disk can be used as a secondary cache (i.e., disk space), and the target data set with a large data scale is divided into multiple target data blocks, and each target data block is stored according to the access frequency. Before determining the access frequency corresponding to each target data block for the first time, it is also necessary to download the target large model and the target data set from the remote storage pool to the local GPU server, and determine the target training tasks corresponding to each target node of the local GPU server through load balancing; then the target data set is divided based on the target training tasks to obtain each target data block, and the target type data corresponding to each target training task is determined from all target data blocks by using the data access pattern and historical access records, and / or by using the target hot data prediction algorithm; finally, based on the correspondence between the target node and the target training task, the target type data corresponding to each target training task is cached to the preset memory space corresponding to the corresponding target node, and the target data set is cached to the preset disk space corresponding to each target node, as Figure 2 shown. Among them, the target node is the node on the local GPU server used to train the target large model. That is to say, after obtaining the target large model and the target data set from the remote storage pool, it is possible to predict the data blocks with the highest access frequency that may be required by the training tasks corresponding to each target node during the training process, and load the data blocks with the highest access frequency corresponding to each target node into the preset memory space corresponding to the target node in advance. This process can be achieved by analyzing the data access pattern and historical access records, or by using the target hot data prediction algorithm, so as to realize cache prefetching and reduce data request latency. When allocating the memory space of the local GPU server, it can be carried out according to the cache policy, so as to allocate an appropriate preset memory space for each target node, so as to load the hot data (target type data) corresponding to each target node into the preset memory space corresponding to the target node.
[0053] It should be noted that in this embodiment, the caching system can significantly improve the training efficiency by implementing distributed caching on each target node in the GPU server cluster. Specifically, install a cache management service on each target node in the cluster, configure the cache synchronization mechanism between the target nodes, and use a load balancing algorithm to allocate the training data according to the computing power and storage capacity of each node. The preset memory space corresponding to each node is responsible for caching the target type data corresponding to the target node, so as to ensure that the cache load of each target node is evenly distributed and avoid some nodes becoming bottlenecks. At the same time, in order to ensure the consistency of the data in each cache node, distributed consistency protocols (such as Paxos and Raft) can be used to ensure the consistency of the cache data. When the data changes, the caches of all relevant nodes are updated in real time to ensure data consistency. That is to say, if there is data content that has been changed in the target type data corresponding to the target node, the changed data content is synchronized to the target node that meets the change synchronization condition by using the preset distributed consistency protocol. In a specific implementation manner, the target type data corresponding to target node A and target node B is target data block 1, and the target type data corresponding to target node C is target data block 2. When the target data block 1 on target node A changes, only this data change will be synchronized to target node B; when the target type data corresponding to target node C also becomes target data block 1, the changes that occur in target data block 1 on target node A and target node B will also be synchronized to target node C, so as to achieve efficient data sharing and access. In addition, the caching system can be used to periodically update the cache content to maintain the timeliness and accuracy of the data. The use of load balancing can also establish a fault tolerance mechanism to handle node failures and data loss problems.
[0054] In this embodiment, the caching system can be adapted to multiple storage systems (remote storage pools), such as OSS, HDFS, Amazon S3, etc. To achieve the adaptation of the caching system to multiple remote storage pools, a general preset application programming interface can be designed, and these preset application programming interfaces are used to dock the caching system with each remote storage pool. After that, the target large model and target dataset can be downloaded from each remote storage pool to the local GPU server. In addition, it is also necessary to perform format conversion on the target dataset in the local GPU server based on the target format, and convert the data downloaded from the remote storage pool into the target format suitable for the caching system, so as to split the format-converted target dataset based on the target training task to obtain each target data block. At the same time, by implementing an adaptation layer in the caching system, it is possible to support data read and write operations of different storage systems, be compatible with multiple underlying storage protocols, and support multiple usage scenarios.
[0055] It can be understood that to improve the data transmission rate, when downloading the target dataset and the target large model from the remote storage pool, the caching system can use parallel processing technology to download multiple data blocks from the remote storage system simultaneously. Specifically, the data transmission rate can be improved by adjusting the number of download threads, optimizing the network bandwidth utilization, and reducing the waiting time of a single request. In addition, in this embodiment, the caching system can also adopt asynchronous I / O (Input / Output) technology to ensure that data reading and writing operations do not block the main training process, thereby improving the data flow rate and the system's response speed, and enhancing the efficiency of data download and processing.
[0056] In this embodiment, to determine the access frequency corresponding to each target data block, the caching system can use a preset frequency determination algorithm to track the access frequency corresponding to each target data block in real time. The preset frequency determination algorithm includes the Least Recently Used (LRU) algorithm and the Least Frequently Used (LFU) algorithm. After obtaining the real-time access frequency corresponding to each target data block, the target data blocks can be classified based on the access frequency to determine whether the target data blocks are target type data. By using the preset frequency determination algorithm to track the access frequency corresponding to each target data block, it is possible to track the data access frequency in real time according to the actual data access situation during the training process, identify which data becomes hot, and dynamically adjust the data allocation in the preset memory space and the NVMe disk, that is, dynamically adjust the caching policy according to the real-time access pattern. For example, when the data access pattern changes, the system automatically transfers the frequently accessed data from the disk cache to the memory. Thus, it can be determined which data blocks are target type data, and the access frequency corresponding to these target type data is higher than that of other types of data, and the data block size of these target type data is smaller than the preset data block size threshold. By using the algorithm to mark the hot and cold data, the hot data (i.e., the target type data) resides in the local memory or disk of the GPU server and does not need to be redownloaded every time training is performed, thereby reducing the data download time. In addition, the caching system can cache the target dataset in the local idle disk or memory without incurring additional costs.
[0057] Step S12: If the target data block is target type data, cache the target data block into the preset memory space so as to read the target data block from the preset memory space to train the target large model.
[0058] In this embodiment, if the target data block is of the target type, the target data block can be cached in a preset memory space, and then the target type data can be quickly read from the preset memory space when training the target large model, thereby accelerating the large model training process. Specifically, if the target data block is of the target type on the target node, the target data block needs to be cached in the preset memory space corresponding to the target node, so that when using the target node to train the target large model, the target data block can be read from the preset memory space corresponding to the target node. At the same time, to improve the effectiveness of caching, if the target data block is of the target type, the target data block is cached in the preset memory space by using the first preset data compression algorithm, so that more data blocks can be stored in the same memory and disk space, reducing data storage occupancy, transmission bandwidth occupancy, and transmission time.
[0059] Step S13: If the target data block is not of the target type, cache the target data block in a preset disk space, and use a cache management tool to manage the target data block cached in the preset disk space, so as to read the target data block from the preset disk space to train the target large model.
[0060] In this embodiment, if the target data block is not of the target type, the target data block can be cached in a preset disk space, and a cache management tool is used to manage the target data block cached in the preset disk space, so as to read the target data block from the preset disk space to train the target large model. Specifically, if the target data block is not of the target type on the target node, the target data block is cached in the preset disk space corresponding to the target node, and a cache management tool is used to manage the target data block cached in the preset disk space, so that when using the target node to train the target large model, the target data block can be read from the preset disk space corresponding to the target node. At the same time, when caching non-target type data in the preset disk space, the target data block can also be cached in the preset disk space by using the second preset data compression algorithm, so that more data blocks can be stored in the same memory and disk space, reducing data storage occupancy and improving the effectiveness of caching. The use of a cache management tool (such as Redis, Memcached) can ensure the fast reading and writing of disk data.
[0061] It can be seen that after the caching system in the present application determines the access frequency corresponding to each target data block, it determines the target type data, that is, small-scale hot data, based on the access frequency and the size of each target data block. Then, in combination with the hierarchical caching strategy for small-scale data and large-scale data, the target type data is stored in a preset disk space, and other type data is stored in a preset disk space. Thus, when training a large model, if the data access pattern changes, the caching system can automatically transfer the frequently accessed target data blocks from the disk cache to the memory. The data in the memory can be quickly called, and there is no need to repeat the download during each training, reducing the remote download time, significantly improving the data processing efficiency and overall training performance during the large-scale model training process, and enhancing the stability and scalability of the caching system.
[0062] See Figure 3 As shown, the present application discloses an acceleration device for large model training, which is applied to a caching system and includes:
[0063] A data type determination module 11, configured to determine the access frequency corresponding to each target data block, and classify the target data block based on the access frequency to determine whether the target data block is target type data; the access frequency corresponding to the target type data is higher than the access frequency corresponding to other type data, and the data block size of the target type data is less than a preset data block size threshold;
[0064] A first data caching module 12, configured to cache the target data block into a preset memory space if the target data block is target type data, so as to read the target data block from the preset memory space to train a target large model;
[0065] A second data caching module 13, configured to cache the target data block into a preset disk space if the target data block is not target type data, and manage the target data block cached in the preset disk space by using a cache management tool, so as to read the target data block from the preset disk space to train the target large model.
[0066] It can be seen that after the cache system in this application determines the access frequency corresponding to each target data block, it determines the target type data, that is, small-scale hot data, based on the access frequency and the size of each target data block. Then, by combining the hierarchical caching strategy for small-scale data and large-scale data, the target type data is stored in a preset disk space, and other type data is stored in a preset disk space. Thus, when training a large model, if the data access pattern changes, the cache system can automatically transfer the frequently accessed target data blocks from the disk cache to the memory. The data in the memory can be quickly called, and there is no need to repeat the download during each training, reducing the remote download time, significantly improving the data processing efficiency and overall training performance during the large-scale model training process, and enhancing the stability and scalability of the cache system.
[0067] In a specific embodiment, the device may further include:
[0068] A data download module, configured to download the target large model and the target data set from a remote storage pool to a local GPU server;
[0069] A training task determination module, configured to determine the target training tasks corresponding to the target nodes of the local GPU server through load balancing;
[0070] A data set splitting module, configured to split the target data set based on the target training tasks to obtain each of the target data blocks;
[0071] A data classification module, configured to determine the target type data corresponding to each of the target training tasks from all the target data blocks by using the data access pattern and historical access records, and / or by using a target hot data prediction algorithm;
[0072] A third data caching module, configured to cache the target type data corresponding to each of the target training tasks to the preset memory space corresponding to the corresponding target node based on the correspondence between the target node and the target training task;
[0073] A data set caching module, configured to cache the target data set to the preset disk space corresponding to each of the target nodes;
[0074] Wherein, the target node is a node on the local GPU server for training the target large model.
[0075] In a specific embodiment, the first data caching module 12 may specifically include:
[0076] A first data cache unit, configured to cache the target data block into a preset memory space corresponding to the target node if the target data block is of a target type on the target node, so that when training the target large model using the target node, the target data block can be read from the preset memory space corresponding to the target node;
[0077] Correspondingly, the second data cache module 13 may specifically include:
[0078] A second data cache unit, configured to cache the target data block into a preset disk space corresponding to the target node if the target data block is not of a target type on the target node, and manage the target data block cached in the preset disk space using a cache management tool, so that when training the target large model using the target node, the target data block can be read from the preset disk space corresponding to the target node.
[0079] In a specific embodiment, the apparatus may further include:
[0080] A data content synchronization module, configured to synchronize the data content with changes to the target nodes that meet the change synchronization conditions using a preset distributed consistency protocol if there is data content with changes in the target type data corresponding to the target node
[0081] In a specific embodiment, the data type determination module 11 may specifically include:
[0082] An access frequency determination unit, configured to track the access frequency corresponding to each target data block in real time using a preset frequency determination algorithm;
[0083] Wherein, the preset frequency determination algorithm includes the least recently used algorithm and the least frequently used algorithm.
[0084] In a specific embodiment, the data download module may specifically include:
[0085] A docking unit, configured to dock the cache system with each remote storage pool using a preset application programming interface;
[0086] A data download unit, configured to download the target large model and the target data set from each remote storage pool to the local GPU server;
[0087] A format conversion unit, configured to perform format conversion on the target data set in the local GPU server based on a target format, so as to perform segmentation on the format-converted target data set based on the target training task to obtain each target data block.
[0088] In a specific embodiment, the first data cache module 12 may specifically include:
[0089] A third data cache unit, configured to cache the target data block into a preset memory space by using a first preset data compression algorithm if the target data block is of a target type of data;
[0090] Correspondingly, the second data cache module 13 may specifically include:
[0091] A fourth data cache unit, configured to cache the target data block into a preset disk space by using a second preset data compression algorithm if the target data block is not of a target type of data.
[0092] Furthermore, an embodiment of the present application also discloses an electronic device. Figure 4 It is a structural diagram of an electronic device 20 shown according to an exemplary embodiment. The content in the figure should not be considered as any limitation on the scope of use of the present application.
[0093] Figure 4 It is a schematic structural diagram of an electronic device 20 provided by an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the large model training acceleration method disclosed in any of the foregoing embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0094] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows is any communication protocol applicable to the technical solution of the present application, and no specific limitation is imposed on it here; the input / output interface 25 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application needs, and no specific limitation is made here.
[0095] In addition, as a carrier for resource storage, the memory 22 may be a read-only memory, a random access memory, a disk, or an optical disc, etc. The resources stored thereon may include an operating system 221, a computer program 222, etc., and the storage method may be short-term storage or permanent storage.
[0096] Among them, the operating system 221 is used to manage and control each hardware device and computer program 222 on the electronic device 20, and it can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program that can be used to complete the large model training acceleration method executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 222 may further include computer programs that can be used to complete other specific tasks.
[0097] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the large model training acceleration method disclosed above is implemented. For the specific steps of this method, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be elaborated herein.
[0098] In this specification, each embodiment is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and reference can be made to the description of the method part for related parts.
[0099] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0100] The steps of the method or algorithm described in combination with the embodiments disclosed in this article can be directly implemented by hardware, a software module executed by a processor, or a combination of both. The software module can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.
[0101] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising said element.
[0102] The technical solutions provided in this application have been introduced in detail above. Specific examples are used in this text to elaborate on the principles and implementation manners of this application. The descriptions of the above embodiments are only used to help understand the method and its core idea of this application. At the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.
Claims
1. A method for accelerating large model training, characterized in that, Applied to a cache system, including: Determine the access frequency corresponding to each target data block, and classify the target data blocks based on the access frequency to determine whether the target data blocks are target type data; the access frequency corresponding to the target type data is higher than that corresponding to other type data, and the data block size of the target type data is smaller than a preset data block size threshold; If the target data block is target type data, cache the target data block in a preset memory space so as to read the target data block from the preset memory space to train a target large model; If the target data block is not target type data, cache the target data block in a preset disk space, and use a cache management tool to manage the target data block cached in the preset disk space so as to read the target data block from the preset disk space to train the target large model.
2. The large model training acceleration method according to claim 1, wherein Before determining the access frequency corresponding to each target data block, further include: Download the target large model and target data set from a remote storage pool to a local GPU server; Determine the target training tasks corresponding to each target node of the local GPU server through load balancing; Based on the target training tasks, split the target data set to obtain each of the target data blocks; Use a data access pattern and historical access records, and / or use a target hot data prediction algorithm to determine the target type data corresponding to each of the target training tasks from all the target data blocks; Based on the corresponding relationship between the target node and the target training task, cache the target type data corresponding to each of the target training tasks in the preset memory space corresponding to the corresponding target node; Cache the target data set in the preset disk space corresponding to each of the target nodes; Wherein, the target node is a node on the local GPU server for training the target large model.
3. The large model training acceleration method according to claim 2, wherein The step of if the target data block is target type data, cache the target data block in a preset memory space so as to read the target data block from the preset memory space to train a target large model, includes: If the target data block is target type data on the target node, cache the target data block in the preset memory space corresponding to the target node so as to read the target data block from the preset memory space corresponding to the target node when using the target node to train the target large model; Correspondingly, the step of if the target data block is not target type data, cache the target data block in a preset disk space, and use a cache management tool to manage the target data block cached in the preset disk space so as to read the target data block from the preset disk space to train the target large model, includes: If the target data block is not of the target type on the target node, cache the target data block in the preset disk space corresponding to the target node, and use a cache management tool to manage the target data block cached in the preset disk space, so that when using the target node to train the target large model, read the target data block from the preset disk space corresponding to the target node.
4. The large model training acceleration method according to claim 3, wherein It further includes: If there is data content in the target type data corresponding to the target node that has been modified, use a preset distributed consistency protocol to synchronize the modified data content to the target node that meets the modification synchronization condition.
5. The large model training acceleration method according to claim 1, wherein The determining the access frequency corresponding to each target data block includes: Using a preset frequency determination algorithm to track the access frequency corresponding to each target data block in real time; Wherein, the preset frequency determination algorithm includes the least recently used algorithm and the least frequently used algorithm.
6. The large model training acceleration method according to claim 2, wherein The downloading the target large model and the target data set from the remote storage pool to the local GPU server includes: Using a preset application programming interface to dock the cache system with each remote storage pool; Downloading the target large model and the target data set from each remote storage pool to the local GPU server; Performing format conversion on the target data set in the local GPU server based on the target format, so as to perform segmentation on the format-converted target data set based on the target training task to obtain each target data block.
7. The large model training acceleration method according to any one of claims 1 to 6, characterized in that The caching the target data block in the preset memory space if the target data block is of the target type includes: If the target data block is of the target type, use a first preset data compression algorithm to cache the target data block in the preset memory space; Correspondingly, the caching the target data block in the preset disk space if the target data block is not of the target type includes: If the target data block is not of the target type, use a second preset data compression algorithm to cache the target data block in the preset disk space.
8. An apparatus for accelerating large model training, characterized in that, Applied to a cache system, it includes: A data type determination module, configured to determine the access frequency corresponding to each target data block, and perform data classification on the target data block based on the access frequency to determine whether the target data block is of the target type; the access frequency corresponding to the target type data is higher than the access frequency corresponding to other type data, and the data block size of the target type data is smaller than a preset data block size threshold; A first data caching module, configured to cache the target data block in the preset memory space if the target data block is of the target type, so as to read the target data block from the preset memory space to train the target large model; A second data caching module, configured to cache the target data block in the preset disk space if the target data block is not of the target type, and use a cache management tool to manage the target data block cached in the preset disk space, so as to read the target data block from the preset disk space to train the target large model.
9. An electronic device, characterized in that, It includes: A memory for storing a computer program; A processor for executing the computer program to implement the large model training acceleration method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, For storing a computer program, wherein when the computer program is executed by a processor, the large model training acceleration method according to any one of claims 1 to 7 is implemented.