Large model data distributed management method and device, equipment and storage medium
Through sliding window technology and long and short-term memory network model, data block access frequency is predicted, combined with genetic algorithms to optimize resource allocation and distributed snapshot algorithm, the problems of resource idleness and delay in hyper-large-scale model training are solved, and efficient and stable data management and task allocation are achieved.
Patent Information
- Application Number
- CN202510704394.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-08-29
AI Technical Summary
In hyper-large-scale model training, traditional scheduling managers cannot perceive model parallel strategies, resulting in high resource idle rate and high latency of remote storage slowing down training efficiency, seriously affecting cluster efficiency and task stability.
Through sliding window technology and long and short-term memory network model, data block access frequency is predicted, combined with genetic algorithm-optimized hybrid strategy and distributed snapshot algorithm, dynamically manage resource allocation and data storage, realize the classification cache of data blocks and efficient allocation of tasks, and perform periodic snapshot operations.
Effectively reduce resource idle rate, improve task stability, reduce data loading delay, improve training efficiency, and save resource costs.
Smart Images

Figure CN120560809A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a large-scale model data distributed management method, device, equipment and storage medium. Background Art
[0002] In ultra-large-scale model training, computing power requirements are growing exponentially. Single-machine computing power is insufficient to meet these requirements, necessitating the collaboration of distributed GPU (Graphics Processing Unit) clusters. However, traditional schedulers allocate tasks based on static resources and fail to understand the explicit requirements of model parallelization strategies, resulting in excessively high resource idleness. Furthermore, model training requires frequent loading of massive amounts of data from storage systems, and the high latency of remote storage slows down training efficiency, severely limiting cluster efficiency and task stability.
[0003] As can be seen from the above, how to reduce resource idleness and improve task stability is an urgent problem to be solved. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a method, apparatus, device, and storage medium for distributed management of large-scale model data, which can reduce resource idleness and improve task stability. The specific solution is as follows:
[0005] In a first aspect, the present application provides a method for distributed management of large model data, comprising:
[0006] The large model training data is segmented to obtain a number of data blocks, the historical access frequency of the data blocks is determined using a sliding window technique, and the predicted access frequency of the data blocks within a preset time period is determined based on the long short-term memory network model and the historical access frequency;
[0007] Classifying the data blocks using the predicted access frequencies to obtain classification results, and caching the large model training data corresponding to the data blocks in a corresponding data cache layer based on the classification results;
[0008] A resource profile is constructed using the static properties and dynamic indicators of the GPU node. Based on the predicted access frequency of the data block and the corresponding storage location in the data cache layer, the large model training task is allocated to the corresponding target GPU node using a preset colocation strategy and the resource profile; the preset colocation strategy is a strategy optimized based on a genetic algorithm.
[0009] When the execution of the large model training task on the target GPU node is monitored, a distributed snapshot algorithm is used to perform periodic snapshot operations on the large model training task, and the obtained complete data status is saved to a distributed storage center to complete the distributed management operation of the large model training data.
[0010] Optionally, the large model training data is segmented to obtain a plurality of data blocks, a sliding window technique is used to determine a historical access frequency of the data blocks, and a predicted access frequency of the data blocks within a preset time period is determined based on a long short-term memory network model and the historical access frequency, including:
[0011] The large model training data is segmented to obtain a number of data blocks, and the number of accesses of each data block in each window is determined based on a sliding window technique and an access counter, and the number of historical accesses corresponding to the data block is determined using each access count;
[0012] Determining a historical access interval based on the historical access records of the data block, and determining a mean and a variance corresponding to the historical access interval;
[0013] An initial access frequency prediction model is constructed using a long short-term memory network model, and the initial access frequency prediction model is trained based on the historical number of accesses, the mean, and the variance to obtain a target access frequency prediction model;
[0014] The access frequency of the data block within a preset time period is predicted based on the target access frequency prediction model to obtain a predicted access frequency corresponding to the data block.
[0015] Optionally, the classifying the data blocks using the predicted access frequencies to obtain classification results, and caching the large model training data corresponding to the data blocks to a corresponding data cache layer based on the classification results, includes:
[0016] If the predicted access frequency corresponding to the data block is greater than a first probability threshold, the data block is determined to be hot data, and the hot data is cached in the memory layer and managed using large page technology;
[0017] If the predicted access frequency corresponding to the data block is less than the first probability threshold and greater than the second probability threshold, the data block is determined as warm data and the warm data is cached to a preset disk;
[0018] If the predicted access frequency corresponding to the data block is less than the second probability threshold, the data block is determined as cold data, and the cold data is cached in a remote storage layer.
[0019] Optionally, the method of constructing a resource profile using the static attributes and dynamic indicators of the GPU node, allocating the large model training task to the corresponding target GPU node based on the predicted access frequency of the data block and the corresponding storage location in the data cache layer and using a preset colocation strategy and the resource profile, includes:
[0020] Collect dynamic indicators including GPU utilization, remaining video memory, and network throughput;
[0021] Utilize static attributes including GPU model, video memory capacity, NVLink bandwidth, and network port, as well as the dynamic indicators, to build a resource profile reflecting the operating status of each GPU node;
[0022] Based on the GPU utilization, the remaining amount of video memory, and the network throughput, the genetic algorithm is optimized using a fitness function to obtain a preset colocation strategy;
[0023] Based on the predicted access frequency corresponding to the data block and the corresponding storage location in the data cache layer and using the preset mixed deployment strategy and the resource profile, the large model training task is allocated to the corresponding target GPU node.
[0024] Optionally, after allocating the large model training task to the corresponding target GPU node based on the predicted access frequency corresponding to the data block and the storage location corresponding to the data cache layer and using a preset colocation strategy and the resource profile, the method further includes:
[0025] Monitoring the GPU utilization corresponding to each GPU node based on a preset monitoring frequency, and adding the GPU node if a monitoring result indicates that the GPU utilization is greater than a first utilization threshold and meets a first preset duration;
[0026] If the monitoring result indicates that the GPU utilization is less than a second utilization threshold and exceeds a second preset duration, an idle GPU node is determined based on the number of preset core GPU nodes.
[0027] Optionally, when the execution of the large model training task on the target GPU node is monitored, a distributed snapshot algorithm is used to perform periodic snapshot operations on the large model training task, and the obtained complete data state is saved to a distributed storage center, including:
[0028] When the execution of the large model training task on the target GPU node is monitored, a full snapshot is taken of the model parameters, optimizer state, and data location of the large model training data corresponding to the large model training task based on a first preset interval time to obtain a full snapshot result;
[0029] During the execution process, performing incremental snapshots of the model parameters, the optimizer state, and the parameters that have changed in the data position based on a second preset interval time and the initial incremental result to obtain incremental snapshot results;
[0030] The complete data status of the large model training task is determined based on the full snapshot result and the incremental snapshot result, and the complete data status is saved to a distributed storage center.
[0031] Optionally, after saving the obtained complete data state to the distributed storage center, the following steps are included:
[0032] Determining an initial fault condition of the GPU node based on a heartbeat packet sent by the GPU node;
[0033] Determining a health score of the GPU node using the initial fault condition in combination with an ECC error rate, a GPU temperature, and a video memory bandwidth of the GPU node;
[0034] If the health score is less than the target health threshold, the large model training task on the GPU node is migrated in batches, and the GPU node is determined to be a faulty node, and then the faulty node is recovered using erasure coding technology.
[0035] In a second aspect, the present application provides a large-scale model data distributed management device, comprising:
[0036] An access frequency prediction module is used to segment the large model training data to obtain a number of data blocks, determine the historical access frequency of the data blocks using a sliding window technique, and determine the predicted access frequency of the data blocks within a preset time period based on the long short-term memory network model and the historical access frequency;
[0037] a data block cache module, configured to classify the data blocks using the predicted access frequencies to obtain classification results, and cache the large model training data corresponding to the data blocks to a corresponding data cache layer based on the classification results;
[0038] A task allocation module is configured to construct a resource profile using the static properties and dynamic indicators of the GPU node, and to allocate the large model training task to the corresponding target GPU node based on the predicted access frequency of the data block and the corresponding storage location in the data cache layer, using a preset colocation strategy and the resource profile; the preset colocation strategy is a strategy optimized based on a genetic algorithm;
[0039] The task snapshot module is used to perform periodic snapshot operations on the large model training task using a distributed snapshot algorithm when monitoring the execution of the large model training task on the target GPU node, and save the obtained complete data status to a distributed storage center to complete the distributed management operation of the large model training data.
[0040] In a third aspect, the present application provides an electronic device, comprising:
[0041] Memory, used to store computer programs;
[0042] The processor is used to execute the computer program to implement the aforementioned large model data distributed management method.
[0043] In a fourth aspect, the present application provides a computer-readable storage medium for storing a computer program, wherein the computer program implements the aforementioned large model data distributed management method when executed by a processor.
[0044] This application first divides the large model training data to obtain several data blocks, uses sliding window technology to determine the historical access frequency of the data blocks, and determines the predicted access frequency of the data blocks within a preset time period based on the long short-term memory network model and the historical access frequency; uses the predicted access frequency to classify the data blocks to obtain classification results, and caches the large model training data corresponding to the data blocks to the corresponding data cache layer based on the classification results; uses the static properties and dynamic indicators of the GPU nodes to construct a resource portrait, and based on the predicted access frequency corresponding to the data blocks and the storage location corresponding to the data cache layer and using a preset mixed deployment strategy and the resource portrait, allocates the large model training task to the corresponding target GPU node; the preset mixed deployment strategy is a strategy obtained by optimization based on a genetic algorithm; when the execution of the large model training task on the target GPU node is monitored, a distributed snapshot algorithm is used to perform periodic snapshot operations on the large model training task, and the obtained complete data state is saved to a distributed storage center to complete the distributed management operation of the large model training data.
[0045] As can be seen from the above, this application uses sliding window technology and long short-term memory network models to predict the access frequency of data blocks within a preset time period, and then classifies and stores the data blocks based on the predicted access frequency, which can effectively reduce data loading time and alleviate network congestion; then, the static properties and dynamic indicators of the GPU nodes are used to build resource portraits, and based on the resource portraits, the predicted access frequency and the preset mixed deployment strategy, the large model training tasks are assigned to the corresponding target GPU nodes to improve node utilization. In this way, periodic snapshot operations are performed on the executed large model training tasks to ensure the stability of model training. Even in the face of ultra-large-scale model training, data loading delays can be avoided, which not only saves resource costs, but also greatly improves training efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0047] Figure 1 This is a flow chart of a large model data distributed management method disclosed in this application;
[0048] Figure 2 A flow chart of a specific large-scale model data distributed management method provided in this application;
[0049] Figure 3 A flowchart of a multi-level caching mechanism method provided by this application;
[0050] Figure 4 A fault-tolerant recovery timing diagram provided for this application;
[0051] Figure 5 This is a schematic diagram of the structure of a large-scale model data distributed management device disclosed in this application;
[0052] Figure 6 This is a structural diagram of an electronic device disclosed in this application. DETAILED DESCRIPTION
[0053] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0054] At present, in ultra-large-scale model training, traditional scheduling managers cannot perceive the explicit requirements of model parallel strategies through static resource allocation tasks, resulting in excessively high resource idle rates. In addition, when training models, massive amounts of data need to be frequently loaded from the storage system. The high latency of remote storage slows down training efficiency, seriously restricting cluster efficiency and task stability. To this end, this application provides a large-model data distributed management method that performs periodic snapshot operations on the executed large-model training tasks to ensure the stability of model training. Even in the face of ultra-large-scale model training, data loading delays can be avoided, which not only saves resource costs but also greatly improves training efficiency.
[0055] See also Figure 1 As shown, the embodiment of the present invention discloses a large model data distributed management method, including:
[0056] Step S11: split the large model training data to obtain several data blocks, use the sliding window technology to determine the historical access frequency of the data blocks, and determine the predicted access frequency of the data blocks within a preset time period based on the long short-term memory network model and the historical access frequency.
[0057] In this embodiment, the large model training data is segmented into several data blocks of a fixed size, and each data block is assigned a unique identifier. The large model training data can be training data from a BERT (Bidirectional Encoder Representations from Transformers) text corpus. The number of accesses to each data block within each window is then determined using a sliding window technique and an access counter. When a data block within a time window is accessed, the access count is incremented by 1 to update the access count, and expired access records are cleared. The historical access count corresponding to the data block is then determined using each access count. The historical access interval is then determined based on the historical access records of the data block, along with the mean and variance corresponding to the historical access interval. An initial access frequency prediction model is constructed using a long short-term memory network model. The initial access frequency prediction model is trained based on the historical access count, the mean, and the variance to obtain a target access frequency prediction model. After obtaining the target access frequency prediction model, the access frequency of the data block within a preset time period is predicted based on the target access frequency prediction model to obtain a predicted access frequency corresponding to the data block.
[0058] Specifically, the large model training data is divided to obtain several data blocks, the historical access frequency of the data blocks is determined using a sliding window technology, and the predicted access frequency of the data blocks within a preset time period is determined based on a long short-term memory network model and the historical access frequency, including: dividing the large model training data to obtain several data blocks, determining the number of accesses to the data blocks in each window based on a sliding window technology and an access counter, and determining the historical access number corresponding to the data block using each access number; determining the historical access interval based on the historical access records of the data blocks, and determining the mean and variance corresponding to the historical access interval; constructing an initial access frequency prediction model using a long short-term memory network model, training the initial access frequency prediction model based on the historical access number, the mean and the variance to obtain a target access frequency prediction model; predicting the access frequency of the data block within a preset time period based on the target access frequency prediction model to obtain the predicted access frequency corresponding to the data block.
[0059] Step S12: classify the data blocks using the predicted access frequencies to obtain classification results, and cache the large model training data corresponding to the data blocks to the corresponding data cache layer based on the classification results.
[0060] In this embodiment, after obtaining the preset access frequency, the data blocks are classified based on the preset access frequency to divide the data blocks into hot data, warm data, and cold data, and the data blocks are saved in the corresponding data cache layer. In a specific embodiment, if the predicted access frequency corresponding to the data block is greater than 0.8, the data block is determined to be hot data, and the hot data is cached to the memory layer, and the hot data is managed using large page technology; if the predicted access frequency corresponding to the data block is less than or equal to 0.8 and greater than 0.4, the data block is determined to be warm data, and the warm data is cached to the NVME (Non-Volatile Memory Express, i.e., non-volatile memory host controller interface specification) disk; if the predicted access frequency corresponding to the data block is less than or equal to 0.4, the data block is determined to be cold data, and the cold data is cached to the remote storage layer.
[0061] Specifically, the data blocks are classified using the predicted access frequency to obtain the classified results, and the large model training data corresponding to the data blocks are cached to the corresponding data cache layer based on the classified results, including: if the predicted access frequency corresponding to the data block is greater than the first probability threshold, the data block is determined as hot data, and the hot data is cached to the memory layer, and the hot data is managed using large page technology; if the predicted access frequency corresponding to the data block is less than the first probability threshold and greater than the second probability threshold, the data block is determined as warm data, and the warm data is cached to the preset disk; if the predicted access frequency corresponding to the data block is less than the second probability threshold, the data block is determined as cold data, and the cold data is cached to the remote storage layer. It is worth mentioning that the first probability threshold, the second probability threshold and the third probability threshold can be adjusted according to actual conditions.
[0062] It is understandable that the hot data is cached in the memory layer, and the memory layer retains the K most recent access records of the hot data based on the LRU-K algorithm (i.e., cache elimination algorithm), that is, by comprehensively considering the access frequency and timeliness of the hot data to determine the retention of the data block, priority is given to retaining those hot data that have been accessed multiple times recently and have a relatively recent access time, and replacing those hot data that have been accessed less frequently and have a longer time. The warm data is cached in the NVME disk, and the NVME disk uses the block device direct write (Bypass file system) method to bypass the traditional file system and directly perform read and write operations on the disk block device, reducing the metadata overhead of the file system, thereby improving the read and write performance of the NVMe disk. The cold data is cached in the remote storage layer, and the storage layer is compatible with multiple remote storage protocols such as S3 / OSS / HDFS through the FUSE (Filesystem in Userspace, i.e., user space file system) module, and provides a unified remote storage access interface.
[0063] Step S13: Build a resource profile using the static properties and dynamic indicators of the GPU node, and assign the large model training task to the corresponding target GPU node based on the predicted access frequency corresponding to the data block and the corresponding storage location in the data cache layer and using the preset mixed deployment strategy and the resource profile; the preset mixed deployment strategy is a strategy obtained based on genetic algorithm optimization.
[0064] In this embodiment, static properties including GPU model, video memory capacity, NVLink (a bus and its communication protocol) bandwidth and network port are obtained through a system configuration file or a hardware detection tool, and dynamic indicators including GPU utilization, remaining video memory and network throughput are collected. A resource portrait reflecting the operating status of each GPU node is constructed based on the static properties and the dynamic indicators. Then, based on the GPU utilization, the remaining video memory and the network throughput, the genetic algorithm is optimized using a fitness function to obtain a preset hybrid strategy. In a specific embodiment, the chromosome is represented as a two-dimensional matrix, with rows representing GPU nodes and columns representing time window task lists. Each element represents a large model training task assigned to a specific GPU node within a specific time window. The memory utilization is then determined based on the remaining video memory and the video memory capacity, and the fitness function value is determined using the GPU utilization, the memory utilization and the network throughput. The corresponding formula is as follows:
[0065] ;
[0066] in, is the fitness function value; is the GPU utilization; is the memory usage; is the network throughput; 、 、 are the weights corresponding to the GPU utilization, the memory usage, and the network throughput, respectively. , , . Genetic operations include single-point crossover mutation of exchanging the node allocation sequence of the parent chromosome and directed crossover mutation of migrating tasks of low-utilization nodes to high-utilization nodes. Through these genetic operations, the chromosomes are continuously optimized and the fitness function value is improved to obtain a better preset mixed deployment strategy. Then, based on the predicted access frequency corresponding to the data block and the corresponding storage location in the data cache layer, and using the preset mixed deployment strategy and the resource profile, the large model training task is allocated to the corresponding target GPU node.
[0067] Specifically, the static properties and dynamic indicators of the GPU node are used to construct a resource profile, and based on the predicted access frequency corresponding to the data block and the corresponding storage location in the data cache layer and using the preset mixed deployment strategy and the resource profile, the large model training task is allocated to the corresponding target GPU node, including: collecting dynamic indicators including GPU utilization, remaining video memory, and network throughput; using static properties including GPU model, video memory capacity, NVLink bandwidth and network port, and the dynamic indicators to construct a resource profile reflecting the operating status of each GPU node; based on the GPU utilization, the remaining video memory and network throughput, and using the fitness function to optimize the genetic algorithm to obtain a preset mixed deployment strategy; based on the predicted access frequency corresponding to the data block and the corresponding storage location in the data cache layer and using the preset mixed deployment strategy and the resource profile, the large model training task is allocated to the corresponding target GPU node.
[0068] In this embodiment, when assigning a large model training task to a corresponding target GPU node, by determining the memory requirements of the large model training task, video memory resources can be reasonably allocated to avoid task failure due to insufficient video memory. Specifically, the peak video memory is determined based on model parameters, gradients, and optimizer status, and the memory requirements are determined by the floating-point operations corresponding to the large model training task and the peak video memory. Then, a directed acyclic graph is used to represent the various stages of the large model training task and the data dependency method, ensuring that tasks are executed in the correct order and that resources are reasonably utilized.
[0069] As you can see, in addition to optimizing the colocation strategy based on a genetic algorithm for task allocation, a preemptive scheduling process is also incorporated to rationally allocate resources based on the task's QoS (Quality of Service) level. When high-priority tasks require resources, they can promptly preempt those of lower-priority tasks. Automatic checkpoint and breakpoint recovery mechanisms ensure task continuity and stability. Furthermore, an elastic scaling mechanism dynamically starts and stops nodes based on cluster load, reducing hardware costs while ensuring cluster performance. When GPU utilization is excessively high, additional nodes are promptly added to meet task demands. When GPU utilization is low, idle nodes are released to improve resource utilization.
[0070] In this embodiment, in the process of allocating large model training tasks, the hierarchical storage characteristics of the data cache layer are fully considered. For hot data cached in the memory layer, it is preferentially allocated to training tasks that have high requirements for computing speed and real-time data access, such as the key iterative training stage of the deep learning model, to ensure that the task can efficiently read data from the memory and reduce data loading delays. For warm data stored in the NVMe disk layer, it is allocated to tasks that have certain requirements for data access speed but are not extremely demanding, such as some data loading tasks in the pre-training stage of the model, and the high-speed read and write characteristics of the NVMe disk are used to meet the task requirements. As for the cold data retained in the remote storage layer, when the corresponding large model training task needs to be accessed, the data can be quickly located and loaded through multi-protocol adaptation and metadata caching mechanism to avoid long waits caused by remote data access.
[0071] As you can understand, the distributed cache proxy design enables cross-node cache coordination. When a data block needs to be stored, its hash key is calculated and then mapped to the nearest node on the ring for storage. This approach ensures relatively even data distribution across nodes, and when nodes are added or removed, the affected data is relatively small, improving system scalability and stability. Each physical node can correspond to 1,000 virtual nodes, ensuring even distribution and load balancing of data blocks across the cluster. Furthermore, Gossip (a distributed information dissemination mechanism) is used to periodically broadcast node liveness and cache lists, ensuring cache index synchronization within the cluster. This ensures accurate data block location within the cluster when task allocation occurs. Furthermore, to optimize RDMA (Remote Direct Memory Access) transmission, a zero-copy mechanism and bandwidth control algorithm are employed to reduce CPU involvement and network congestion during data transmission. Specifically, memory buffers are registered through the InfiniBand Verbs API (an interface for high-performance data transmission), allowing remote nodes to directly read and write to local memory without requiring the CPU to copy data. Avoid network congestion by determining the number of simultaneous RDMA unilateral operations. The formula for the number of simultaneous RDMA unilateral operations is as follows:
[0072] ;
[0073] in, The number of RDMA unilateral operations performed simultaneously, that is, the number of concurrency; It is the bandwidth resource available for data transmission in the network, that is, the available bandwidth; The average size of the data block transferred for each RDMA operation; The time it takes from sending a data packet to receiving the corresponding acknowledgment packet.
[0074] Furthermore, after the large model training task is assigned to the corresponding target GPU node, the GPU utilization corresponding to each GPU node is monitored based on a preset monitoring frequency. In a specific embodiment, the cluster controller monitors the GPU utilization corresponding to each GPU node every 5 minutes. When it is monitored that the average GPU utilization is greater than 85% and lasts for 10 minutes, the operation of adding a new GPU node is triggered, and the upper limit of the new GPU node is 200% of the current scale; when it is monitored that the average GPU utilization is less than 30% and lasts for more than 30 minutes, the idle GPU nodes are determined based on the number of preset core GPU nodes, the minimum number of core nodes is retained, and the idle nodes are released.
[0075] Specifically, after allocating the large model training task to the corresponding target GPU node based on the predicted access frequency corresponding to the data block and the storage location corresponding to the data cache layer and using the preset mixed deployment strategy and the resource portrait, it also includes: monitoring the GPU utilization corresponding to each of the GPU nodes based on the preset monitoring frequency, if the monitoring result indicates that the GPU utilization is greater than the first utilization threshold and meets the first preset duration, then adding the GPU node; if the monitoring result indicates that the GPU utilization is less than the second utilization threshold and exceeds the second preset duration, then determining the idle GPU node based on the number of preset core GPU nodes. It is worth mentioning that the preset monitoring frequency, the first utilization, the second utilization, the first preset duration and the second preset duration can be adjusted accordingly according to actual conditions and are not specifically limited here.
[0076] Step S14: When the execution of the large model training task on the target GPU node is monitored, a distributed snapshot algorithm is used to perform periodic snapshot operations on the large model training task, and the obtained complete data status is saved to a distributed storage center to complete the distributed management operation of the large model training data.
[0077] In this embodiment, when the execution of the large model training task on the target GPU node is monitored, a full snapshot is taken of the model parameters, optimizer status, and data location of the large model training data corresponding to the large model training task based on a first preset interval and using a distributed snapshot algorithm to obtain a full snapshot result. During the execution process, incremental snapshots are taken of the model parameters, optimizer status, and parameters that have changed in the data location based on a second preset interval and the initial incremental result and using a distributed snapshot algorithm, and the difference in parameters is recorded to obtain an incremental snapshot result. In a specific embodiment, a full snapshot is taken every hour and an incremental snapshot is taken every 5 minutes. The first preset interval and the second preset interval can also be adjusted according to actual conditions. After obtaining the full snapshot result and the incremental snapshot result, the complete data status of the large model training task is determined based on the full snapshot result and the incremental snapshot result, and the complete data status is saved to Ceph (i.e., a distributed file system).
[0078] Specifically, when the execution of the large model training task on the target GPU node is monitored, a distributed snapshot algorithm is used to perform periodic snapshot operations on the large model training task, and the obtained complete data state is saved to a distributed storage center, including: when the execution of the large model training task on the target GPU node is monitored, a full snapshot is performed on the model parameters, optimizer status and data location of the large model training data corresponding to the large model training task based on a first preset interval time to obtain a full snapshot result; during the execution process, an incremental snapshot is performed on the model parameters, the optimizer status and the parameters that have changed in the data location based on a second preset interval time and an initial incremental result to obtain an incremental snapshot result; the complete data state of the large model training task is determined based on the full snapshot result and the incremental snapshot result, and the complete data state is saved to a distributed storage center.
[0079] It is understandable that after the complete data state is saved to the distributed storage center, the initial fault condition of the GPU node is determined based on the heartbeat packet sent by the GPU node. For example, the GPU node sends a heartbeat packet to the monitoring center every 10 seconds. If it is not received for three consecutive times, the initial fault condition is a suspected fault. Then, the heartbeat detection technology is used to determine the corresponding node heartbeat information, and the hardware performance counter is used to determine the ECC (Error Correction Code) error rate of the GPU node. The health score of the GPU node is determined based on the initial fault condition and combined with the ECC error rate of the GPU node, GPU temperature, and memory bandwidth. The formula for the health score is as follows:
[0080] ;
[0081] in, scoring said health; is the ECC error rate; The heartbeat delay information; is the GPU utilization rate. Determine whether the health score is less than the target health threshold; if the health score is less than the target health threshold, migrate the large model training task on the GPU node in batches, determine the GPU node as a faulty node, and then use erasure coding technology to recover the faulty node. In a specific embodiment, if the health score is less than 0.5, task migration is triggered, and rolling migration is performed based on a load balancing algorithm (such as Power of Two Choices) to avoid performance jitter, and actively migrate the large model training task to other non-faulty nodes to ensure the continuity of the training task. At the data level, Reed-Solomon (a polynomial error correction code) (12,4) encoding is used, where the parameter (12,4) indicates that the code length of the encoding is 12 symbols and the information length is 4 symbols. 12 data blocks and 4 parity blocks are read from the distributed storage center, and attempts are made to recover the remaining 8 blocks through any 8 blocks, allowing any 4 blocks to be lost and then recovered. At least three copies of key data are retained in the cluster, and erasure coding technology is used to improve fault tolerance and ensure data security and availability.
[0082] Specifically, after the obtained complete data state is saved to the distributed storage center, it includes: determining the initial fault condition of the GPU node based on the heartbeat packet sent by the GPU node; using the initial fault condition and combining the ECC error rate, GPU temperature and memory bandwidth of the GPU node to determine the health score of the GPU node; if the health score is less than the target health threshold, migrating the large model training task on the GPU node in batches, and determining the GPU node as a faulty node, and then using erasure code technology to recover the faulty node.
[0083] In a specific embodiment, if the GPT-3 model is trained on a GPU cluster consisting of 8 nodes, a full snapshot is generated every hour and an incremental snapshot is generated every 5 minutes. During the training process, a problem occurs in node 3, and the health score corresponding to the node is less than the target health threshold. In order to ensure the continuity of training, the system will trigger the task migration mechanism. Assuming that the target node is selected as node 7, the specific migration process is as follows: node 7 loads the latest full snapshot from the distributed storage center. After the full snapshot is loaded, node 7 also needs to superimpose the three most recent incremental snapshots to restore to the latest task status.
[0084] In addition, performance optimization is performed based on the characteristics of the aforementioned large-model training tasks. In terms of pipeline parallel optimization, the model computation graph is automatically split, subgraphs are allocated across multiple GPUs, and the communication topology is optimized. A 1F1B (One Forward One Backward) scheduling strategy is used to balance video memory usage and throughput. Hybrid-precision training acceleration technology is utilized to dynamically select FP16 (16-bit floating point) / FP32 (32-bit floating point) precision modes, combined with gradient scaling technology to prevent numerical overflow. At the same time, the packet size and communication frequency of the AllReduce algorithm (an algorithm used to synchronize data between multiple processing nodes) are adjusted based on real-time network status to achieve network bandwidth-aware transmission, further improving the overall performance of the cluster and ensuring that large-model training tasks can be executed efficiently and stably.
[0085] As can be seen from the above, this application uses sliding window technology and long short-term memory network models to predict the access frequency of data blocks within a preset time period, and then classifies and stores the data blocks based on the predicted access frequency, which can effectively reduce data loading time and alleviate network congestion; then, the static properties and dynamic indicators of the GPU nodes are used to build resource portraits, and based on the resource portraits, the predicted access frequency and the preset mixed deployment strategy, the large model training tasks are assigned to the corresponding target GPU nodes to improve node utilization. In this way, periodic snapshot operations are performed on the executed large model training tasks to ensure the stability of model training. Even in the face of ultra-large-scale model training, data loading delays can be avoided, which not only saves resource costs, but also greatly improves training efficiency.
[0086] As can be seen from the above embodiments, the present application monitors the node status to determine the faulty node to achieve task migration and recovery execution. Therefore, the process of determining the faulty node based on monitoring the node status is described.
[0087] See also Figure 2 As shown, the embodiment of the present invention discloses a specific large model data distributed management method, including:
[0088] In this embodiment, the user terminal submits a large model training task through a task and specifies the relevant parameters of the task, such as model type, data set path, computing resource requirements, etc. The operating status of the cluster is displayed in real time through the monitoring dashboard, including GPU utilization, memory usage, network throughput, task execution progress and other information, which is convenient for users to monitor and manage tasks. The large model training task submitted by the user segment is then received through the scheduling controller, and the task scheduling strategy is determined based on the resource requirements of the large model training task and the current status of the cluster, such as which GPU node to choose to execute the task on; the resource manager maintains the cluster's resource information, including the usage and availability of resources such as GPU, memory, and storage, to provide a decision-making basis for the scheduling controller. Then, according to the load of each GPU node, the task allocation is dynamically adjusted to ensure that the load of each node is relatively balanced, avoiding the situation where some nodes are overloaded while other nodes are idle.
[0089] Figure 3A multi-level caching mechanism flow chart is provided for this embodiment. First, the large model training data corresponding to the large model training task is divided to obtain several data blocks. The sliding window technology is used to count the number of accesses to the data block in each window, and the target access frequency prediction model constructed based on the long short-term memory network model is used to predict the access frequency of the data block in a preset time period to obtain the predicted access frequency corresponding to the data block. If the predicted access frequency is greater than a first probability threshold, the data block is determined as hot data, and the hot data is cached to the memory layer; if the predicted access frequency is less than the first probability threshold and greater than the second probability threshold, the data block is determined as warm data, and the warm data is cached to the preset disk; if the predicted access frequency is less than the second probability threshold, the data block is determined as cold data, and the cold data is cached to the remote storage layer.
[0090] Then, a logical ring is constructed based on the consistent hashing algorithm, and the node IDs are sorted by hash values. The data block hash key is mapped to the nearest node storage. When data needs to be read, the node where the data block is stored is determined through a consistent hash index query. The memory buffer is registered through the InfiniBand Verbs API, and the remote node directly reads and writes the local memory, reducing CPU participation and improving data transmission efficiency. If the data block is stored in the cache of node B, the data block is read from the memory of node B using remote direct memory access technology. It should be noted that the version number of the data block is increased by 1 each time it is modified. The consistency of the version number needs to be verified when reading to ensure that the data read is the latest and correct data.
[0091] It can be understood that the genetic algorithm is optimized to obtain a preset mixed deployment strategy, and then the static properties and dynamic indicators of the GPU node are used to build a resource portrait. Based on the predicted access frequency corresponding to the data block and the corresponding storage location in the data cache layer, and using the preset mixed deployment strategy and the resource portrait, the large model training task is assigned to the corresponding target GPU node. Figure 4 A fault-tolerant recovery timing diagram is provided for this embodiment, in which the status of each GPU node is continuously monitored by a controller, the heartbeat information of the GPU node is determined based on the heartbeat packet sent by the GPU node, and the ECC error rate of the GPU node is determined using a hardware performance counter. The fault score of the GPU node is determined based on the heartbeat information and the ECC error rate. If the fault score is greater than the target fault threshold, the GPU node is considered to be a faulty node, and the task migration mechanism is started, that is, the controller initiates a task migration operation to migrate the large model training task on the faulty node to the backup node.
[0092] Furthermore, during the task migration process, it is necessary to load the checkpoint data of the large model training task; the checkpoint data contains the status information of the large model training task at a certain point in time; after the checkpoint data is loaded, some snapshot information may be returned; the snapshot information contains the key parameters required during the task execution process to ensure that the large model training task can be correctly restored and continued on the backup node. When the task is successfully migrated and resumed on the backup node, a task migration completion notification will be sent to the controller, informing the controller that the task has been successfully migrated and continues to run. The faulty node will be isolated offline to prevent it from continuing to participate in cluster work and avoid affecting the cluster.
[0093] As can be seen from the above, this application can detect node failures in a timely manner through the heartbeat timeout and ECC alarm mechanism. If a faulty node occurs, task migration is initiated. By loading checkpoints and returning snapshot information, large model tasks can be accurately restored and continued on the backup node, ensuring task continuity and data integrity. The faulty node is isolated offline to prevent it from continuing to participate in cluster work, thus avoiding the negative impact of the faulty node on the overall performance and stability of the system. In this way, by timely detecting faults, effectively migrating tasks, and isolating faulty nodes, the high availability, data integrity, and system stability of the distributed system are significantly improved, while automated management and good scalability are achieved, thereby improving model training efficiency.
[0094] Accordingly, see Figure 5 As shown, the present application also provides a large model data distributed management device, including:
[0095] An access frequency prediction module 11 is configured to segment the large model training data to obtain a plurality of data blocks, determine the historical access frequency of the data blocks using a sliding window technique, and determine the predicted access frequency of the data blocks within a preset time period based on the long short-term memory network model and the historical access frequency;
[0096] A data block cache module 12 is configured to classify the data blocks using the predicted access frequencies to obtain classification results, and cache the large model training data corresponding to the data blocks to a corresponding data cache layer based on the classification results;
[0097] The task allocation module 13 is configured to construct a resource profile using the static attributes and dynamic indicators of the GPU node, and allocate the large model training task to the corresponding target GPU node based on the predicted access frequency of the data block and the storage location of the data cache layer, using a preset colocation strategy and the resource profile; the preset colocation strategy is a strategy optimized based on a genetic algorithm;
[0098] The task snapshot module 14 is used to perform periodic snapshot operations on the large model training task using a distributed snapshot algorithm when monitoring the execution of the large model training task on the target GPU node, and save the obtained complete data status to a distributed storage center to complete the distributed management operation of the large model training data.
[0099] As can be seen from the above, this application uses sliding window technology and long short-term memory network models to predict the access frequency of data blocks within a preset time period, and then classifies and stores the data blocks based on the predicted access frequency, which can effectively reduce data loading time and alleviate network congestion; then, the static properties and dynamic indicators of the GPU nodes are used to build resource portraits, and based on the resource portraits, the predicted access frequency and the preset mixed deployment strategy, the large model training tasks are assigned to the corresponding target GPU nodes to improve node utilization. In this way, periodic snapshot operations are performed on the executed large model training tasks to ensure the stability of model training. Even in the face of ultra-large-scale model training, data loading delays can be avoided, which not only saves resource costs, but also greatly improves training efficiency.
[0100] In some specific implementations, the access frequency prediction module 11 may specifically include:
[0101] A historical access count determination unit is configured to segment the large model training data to obtain a plurality of data blocks, determine the access count of each data block within each window based on a sliding window technique and an access counter, and determine the historical access count corresponding to the data block using each access count;
[0102] a historical access interval determining unit, configured to determine a historical access interval based on the historical access records of the data block, and determine a mean and a variance corresponding to the historical access interval;
[0103] A prediction model training unit, configured to construct an initial access frequency prediction model using a long short-term memory network model, and train the initial access frequency prediction model based on the historical number of accesses, the mean, and the variance to obtain a target access frequency prediction model;
[0104] The access frequency prediction completing unit is configured to predict the access frequency of the data block within a preset time period based on the target access frequency prediction model to obtain a predicted access frequency corresponding to the data block.
[0105] In some specific implementations, the data block cache module 12 may specifically include:
[0106] a hot data determination unit, configured to determine the data block as hot data if the predicted access frequency corresponding to the data block is greater than a first probability threshold, cache the hot data in a memory layer, and manage the hot data using a large page technology;
[0107] a warm data determining unit, configured to determine the data block as warm data if the predicted access frequency corresponding to the data block is less than the first probability threshold and greater than a second probability threshold, and cache the warm data to a preset disk;
[0108] A cold data determination unit is configured to determine the data block as cold data if the predicted access frequency corresponding to the data block is less than the second probability threshold, and cache the cold data in a remote storage layer.
[0109] In some specific implementations, the task allocation module 13 may specifically include:
[0110] Dynamic indicator collection unit, used to collect dynamic indicators including GPU utilization, remaining video memory, and network throughput;
[0111] A resource profile construction unit, configured to construct a resource profile reflecting the operating status of each GPU node using static attributes including GPU model, video memory capacity, NVLink bandwidth, and network port, as well as the dynamic indicators;
[0112] A genetic algorithm optimization unit, configured to optimize a genetic algorithm based on the GPU utilization, the remaining amount of video memory, and the network throughput using a fitness function to obtain a preset colocation strategy;
[0113] A training task allocation unit is used to allocate large model training tasks to corresponding target GPU nodes based on the predicted access frequency corresponding to the data block and the corresponding storage location in the data cache layer and using a preset mixed deployment strategy and the resource profile.
[0114] In some specific implementations, the large model data distributed management device may further include:
[0115] a utilization monitoring unit, configured to monitor the GPU utilization corresponding to each of the GPU nodes based on a preset monitoring frequency, and add the GPU node if a monitoring result indicates that the GPU utilization is greater than a first utilization threshold and satisfies a first preset duration;
[0116] The idle node determining unit is configured to determine an idle GPU node based on the number of preset core GPU nodes if the monitoring result indicates that the GPU utilization is less than a second utilization threshold and exceeds a second preset duration.
[0117] In some specific implementations, the task snapshot module 14 may specifically include:
[0118] a full snapshot unit, configured to, when monitoring the execution of the large model training task on the target GPU node, take a full snapshot of the model parameters, optimizer state, and data location of the large model training data corresponding to the large model training task based on a first preset interval, so as to obtain a full snapshot result;
[0119] an incremental snapshot unit, configured to take incremental snapshots of the model parameters, the optimizer state, and the parameters that have changed in the data location based on a second preset interval time and an initial incremental result during execution to obtain an incremental snapshot result;
[0120] A data state saving unit is used to determine the complete data state of the large model training task based on the full snapshot result and the incremental snapshot result, and save the complete data state to a distributed storage center.
[0121] In some specific implementations, the large model data distributed management device may further include:
[0122] an initial fault condition determining unit, configured to determine an initial fault condition of the GPU node based on a heartbeat packet sent by the GPU node;
[0123] a health score determination unit, configured to determine a health score of the GPU node using the initial fault condition in combination with an ECC error rate, a GPU temperature, and a video memory bandwidth of the GPU node;
[0124] The task migration unit is used to migrate the large model training task on the GPU node in batches if the health score is less than the target health threshold, determine the GPU node as a faulty node, and then use erasure coding technology to recover the faulty node.
[0125] Furthermore, the embodiment of the present application also discloses an electronic device, Figure 6 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content in the diagram should not be considered as any limitation on the scope of use of this application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 is used to store a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the large model data distributed management method disclosed in any of the aforementioned embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0126] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and the external device. The communication protocol it follows is any communication protocol that can be applied to the technical solution of this application and is not specifically limited here; the input and output interface 25 is used to obtain external input data or output data to the outside world. Its specific interface type can be selected according to specific application needs and is not specifically limited here.
[0127] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or CD, etc. The resources stored thereon can include an operating system 221, a computer program 222, etc., and the storage method can be temporary storage or permanent storage.
[0128] The operating system 221 is used to manage and control the hardware devices on the electronic device 20 and the computer program 222, and can be Windows Server, NetWare, Unix, Linux, etc. In addition to including computer programs capable of implementing the large model data distributed management method performed by the electronic device 20 disclosed in any of the aforementioned embodiments, the computer program 222 can further include computer programs capable of performing other specific tasks.
[0129] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when executed by a processor, the computer program implements the aforementioned method for distributed management of large model data. The specific steps of this method can be found in the corresponding contents disclosed in the aforementioned embodiments and will not be further described here.
[0130] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from the other embodiments. Reference can be made to the descriptions of the identical or similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and the relevant parts can be referred to the descriptions of the methods.
[0131] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0132] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0133] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0134] The above is a detailed introduction to the technical solution provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for those skilled in the art, according to the ideas of the present application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A large-scale model data distributed management method, characterized in that: include: The large model training data is segmented to obtain a number of data blocks, the historical access frequency of the data blocks is determined using a sliding window technique, and the predicted access frequency of the data blocks within a preset time period is determined based on the long short-term memory network model and the historical access frequency; Classifying the data blocks using the predicted access frequencies to obtain classification results, and caching the large model training data corresponding to the data blocks in a corresponding data cache layer based on the classification results; A resource profile is constructed using the static properties and dynamic indicators of the GPU node. Based on the predicted access frequency of the data block and the corresponding storage location in the data cache layer, the large model training task is allocated to the corresponding target GPU node using a preset colocation strategy and the resource profile; the preset colocation strategy is a strategy optimized based on a genetic algorithm. When the execution of the large model training task on the target GPU node is monitored, a distributed snapshot algorithm is used to perform periodic snapshot operations on the large model training task, and the obtained complete data status is saved to a distributed storage center to complete the distributed management operation of the large model training data.
2. The large-scale model data distributed management method according to claim 1, characterized in that: The large model training data is segmented to obtain a plurality of data blocks, a sliding window technique is used to determine the historical access frequency of the data blocks, and a predicted access frequency of the data blocks within a preset time period is determined based on a long short-term memory network model and the historical access frequency, including: The large model training data is segmented to obtain a number of data blocks, and the number of accesses of each data block in each window is determined based on a sliding window technique and an access counter, and the number of historical accesses corresponding to the data block is determined using each access count; Determining a historical access interval based on the historical access records of the data block, and determining a mean and a variance corresponding to the historical access interval; An initial access frequency prediction model is constructed using a long short-term memory network model, and the initial access frequency prediction model is trained based on the historical number of accesses, the mean, and the variance to obtain a target access frequency prediction model; The access frequency of the data block within a preset time period is predicted based on the target access frequency prediction model to obtain a predicted access frequency corresponding to the data block.
3. The large-scale model data distributed management method according to claim 1, characterized in that: The method of classifying the data blocks using the predicted access frequencies to obtain classification results, and caching the large model training data corresponding to the data blocks to a corresponding data cache layer based on the classification results, includes: If the predicted access frequency corresponding to the data block is greater than a first probability threshold, the data block is determined to be hot data, and the hot data is cached in the memory layer and managed using large page technology; If the predicted access frequency corresponding to the data block is less than the first probability threshold and greater than the second probability threshold, the data block is determined as warm data and the warm data is cached to a preset disk; If the predicted access frequency corresponding to the data block is less than the second probability threshold, the data block is determined as cold data, and the cold data is cached in a remote storage layer.
4. The large-scale model data distributed management method according to claim 1, characterized in that: The method uses the static attributes and dynamic indicators of the GPU node to build a resource profile, allocates the large model training task to the corresponding target GPU node based on the predicted access frequency of the data block and the storage location corresponding to the data cache layer, and uses the preset colocation strategy and the resource profile, including: Collect dynamic indicators including GPU utilization, remaining video memory, and network throughput; Utilize static attributes including GPU model, video memory capacity, NVLink bandwidth, and network port, as well as the dynamic indicators, to build a resource profile reflecting the operating status of each GPU node; Based on the GPU utilization, the remaining amount of video memory, and the network throughput, the genetic algorithm is optimized using a fitness function to obtain a preset colocation strategy; Based on the predicted access frequency corresponding to the data block and the corresponding storage location in the data cache layer and using the preset mixed deployment strategy and the resource profile, the large model training task is allocated to the corresponding target GPU node.
5. The large-scale model data distributed management method according to claim 4, characterized in that: After allocating the large model training task to the corresponding target GPU node based on the predicted access frequency corresponding to the data block and the storage location corresponding to the data cache layer and using the preset colocation strategy and the resource profile, the method further includes: Monitoring the GPU utilization corresponding to each GPU node based on a preset monitoring frequency, and adding the GPU node if a monitoring result indicates that the GPU utilization is greater than a first utilization threshold and meets a first preset duration; If the monitoring result indicates that the GPU utilization is less than a second utilization threshold and exceeds a second preset duration, an idle GPU node is determined based on the number of preset core GPU nodes.
6. The large-scale model data distributed management method according to claim 1, characterized in that: When the large model training task on the target GPU node is monitored, a distributed snapshot algorithm is used to perform periodic snapshot operations on the large model training task, and the obtained complete data state is saved to a distributed storage center, including: When the execution of the large model training task on the target GPU node is monitored, a full snapshot is taken of the model parameters, optimizer state, and data location of the large model training data corresponding to the large model training task based on a first preset interval time to obtain a full snapshot result; During the execution process, performing incremental snapshots of the model parameters, the optimizer state, and the parameters that have changed in the data position based on a second preset interval time and the initial incremental result to obtain incremental snapshot results; The complete data status of the large model training task is determined based on the full snapshot result and the incremental snapshot result, and the complete data status is saved to a distributed storage center.
7. The large model data distributed management method according to any one of claims 1 to 6, characterized in that: After the obtained complete data state is saved to the distributed storage center, it includes: Determining an initial fault condition of the GPU node based on a heartbeat packet sent by the GPU node; Determine a health score of the GPU node using the initial fault condition in combination with an ECC error rate, a GPU temperature, and a video memory bandwidth of the GPU node; If the health score is less than the target health threshold, the large model training task on the GPU node is migrated in batches, and the GPU node is determined to be a faulty node, and then the faulty node is recovered using erasure coding technology.
8. A large-scale model data distributed management device, characterized in that: include: An access frequency prediction module is used to segment the large model training data to obtain a number of data blocks, determine the historical access frequency of the data blocks using a sliding window technique, and determine the predicted access frequency of the data blocks within a preset time period based on the long short-term memory network model and the historical access frequency; a data block cache module, configured to classify the data blocks using the predicted access frequencies to obtain classification results, and cache the large model training data corresponding to the data blocks to a corresponding data cache layer based on the classification results; A task allocation module is configured to construct a resource profile using the static properties and dynamic indicators of the GPU node, and to allocate the large model training task to the corresponding target GPU node based on the predicted access frequency of the data block and the corresponding storage location in the data cache layer, using a preset colocation strategy and the resource profile; the preset colocation strategy is a strategy optimized based on a genetic algorithm; The task snapshot module is used to perform periodic snapshot operations on the large model training task using a distributed snapshot algorithm when monitoring the execution of the large model training task on the target GPU node, and save the obtained complete data status to a distributed storage center to complete the distributed management operation of the large model training data.
9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the large model data distributed management method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that Used to store a computer program, wherein when the computer program is executed by a processor, the large model data distributed management method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Cache space adjustment method and electronic equipment
CN120743968A
Data storage method
CN120950010A
Data storage method
CN120950010B
Electronic component performance data monitoring method and system
CN121166510A
Data sending method and device
CN121255710A