Dynamic setting method and device of check points and storage medium
Through the method of dynamically setting checkpoints, the resource occupation status of the large model training cluster in the target period is predicted, which solves the high overhead and low efficiency problems caused by checkpoint settings in the prior art, and achieves more efficient and stable model training.
Patent Information
- Application Number
- CN202510571845.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-05-06
AI Technical Summary
During large-scale model training, the periodic checkpoint setting of the prior art results in large-scale overhead and failure recovery losses, affecting training efficiency and stability.
The method of dynamically setting checkpoints is adopted, and the resource usage data of the large model training cluster is collected, the system resource occupation status is predicted within the target period, whether there will be a failure, and the checkpoint is set based on the prediction results.
It improves the effectiveness of checkpoints, reduces the fixed overhead and fault recovery losses of periodic checkpoints, and improves the efficiency and stability of model training.
Smart Images

Figure CN120087413A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of artificial intelligence technology, and particularly to a method, device, and storage medium for dynamically setting checkpoints. Background Art
[0002] With the rapid development of artificial intelligence technology, natural language processing (NLP) technology represented by large language models (LLMs) has made significant breakthroughs. These models perform excellently in tasks such as intelligent question answering, text summarization, and machine translation, greatly expanding the boundaries of artificial intelligence applications. Taking OpenAI's GPT series models as an example, from GPT-1 to GPT-4, the scale of model parameters has been continuously climbing, developing from hundreds of millions of parameters initially to the current scale of trillions of parameters, demonstrating the potential for performance improvement brought by the expansion of model scale.
[0003] However, the continuous expansion of model scale has also brought unprecedented challenges. To process massive amounts of data and perform complex computational tasks, it is usually necessary to rely on large-scale distributed clusters composed of dedicated accelerators such as GPUs or TPUs. These clusters are interconnected through high-speed networks to achieve parallel computing and accelerate the training process. But as the scale of the cluster expands, especially when the scale of the training cluster reaches tens of thousands of GPUs, a series of problems begin to emerge, including load balancing between computing nodes, data synchronization, and communication overhead, which seriously affect the efficiency and stability of training. As a result, in such a large computing cluster, software and hardware failures are often inevitable, especially in ultra-large-scale training scenarios. Taking the training of GPT-4 as an example, it relies on approximately 25,000 A100 GPUs and runs continuously for 90 to 100 days. However, the training system efficiency (MFU) is only 32% to 36%, and this extremely low utilization rate is due to frequent software or hardware failures in the cluster and the resulting training interruptions.
[0004] To address the problem of training interruptions caused by software and hardware failures, existing training systems usually save the intermediate state of model training periodically. However, this periodic checkpoint setting method is too simple, resulting in relatively large fixed overheads of checkpoints and losses in fault recovery.
[0005] Therefore, how to improve the effectiveness of checkpoints configured during the training of large-scale models to reduce the fixed overheads of periodic checkpoints and losses in fault recovery is an urgent problem to be solved. Summary of the Invention
[0006] This specification provides a method, device, and storage medium for dynamically setting checkpoints to partially solve the above problems existing in the prior art.
[0007] This specification adopts the following technical solutions: This specification provides a method for dynamically setting checkpoints. The method is applied to a management device of a large model training cluster, and the method includes: Collect resource usage data of the large model training cluster during the model training process. The resource usage data is used to characterize the system resource occupancy status of the large model training cluster at each moment in the historical period. According to the resource usage data, predict the first dynamic relationship data of the system resource occupancy status of the large model training cluster changing with time during the target period. Determine at least part of the dynamic relationship data that matches the target period in the global dynamic relationship data as the second dynamic relationship data. According to whether the difference between the first dynamic relationship data and the second dynamic relationship data meets a preset warning condition, determine the prediction result of whether the large model training cluster will fail during the target period, and set checkpoints according to the prediction result; the global dynamic relationship data is the expected dynamic relationship of the entire model training process without failure.
[0008] Optionally, predicting the first dynamic relationship data of the large model training cluster during the target period according to the resource usage data specifically includes: Extract at least one periodic component data from the resource usage data. The periodic component data refers to at least part of the data used to characterize the regular and repetitive changes in the resource usage data, and the regularities corresponding to different periodic component data are different. According to the periodic component data, predict the first dynamic relationship data of the large model training cluster during the target period.
[0009] Optionally, predicting the first dynamic relationship data of the large model training cluster during the target period according to the periodic component data specifically includes: For each periodic component data, fill the length of the periodic component data to a specified length to obtain adjusted periodic component data. Fuse the adjusted periodic component data to predict the first dynamic relationship data of the large model training cluster during the target period.
[0010] Optionally, fusing the adjusted periodic component data to predict the first dynamic relationship data of the large model training cluster during the target period specifically includes: Fuse each adjusted periodic component data with the preset residual data to predict the first dynamic relationship data of the large model training cluster within the target time period; the residual data is obtained by modeling the residual data used when determining the historical first dynamic relationship data of different historical time periods through an autoregressive model.
[0011] Optionally, the method further includes: Determine the difference between the historical first dynamic relationship data of the historical time period and at least part of the dynamic relationship data in the global dynamic relationship data that matches the historical time period as the reference difference; When it is determined that the difference is greater than the reference difference, adjust the parameters of the autoregressive model to obtain an adjusted autoregressive model.
[0012] Optionally, extracting at least one periodic component data from the resource usage data specifically includes: Determine at least part of the data in the resource usage data that has regular and repetitive changes as periodic resource usage data segments; According to the predicted key point extraction strategy, extract each periodic key point from the periodic resource usage data segments, and obtain the periodic component data according to each periodic key point, where the key point extraction strategy is used to screen each data point included in the periodic resource usage data segments.
[0013] Optionally, the method further includes: Determine the difference between the historical first dynamic relationship data of the historical time period and at least part of the dynamic relationship data in the global dynamic relationship data that matches the historical time period as the reference difference; When it is determined that the difference is greater than the reference difference, add a limiting condition for screening each data point to the key point extraction strategy to obtain an adjusted key point extraction strategy; According to the adjusted key point extraction strategy, re-extract the periodic component data.
[0014] This specification provides a dynamic checkpoint setting device, including: An acquisition module, configured to acquire the resource usage data of the large model training cluster during the model training process, where the resource usage data is used to characterize the system resource occupancy status of the large model training cluster at each moment in the historical time period; A prediction module, configured to predict the first dynamic relationship data of the system resource occupancy status of the large model training cluster changing with time within the target time period according to the resource usage data; A determination module, configured to determine at least part of the dynamic relationship data in the global dynamic relationship data that matches the target time period as the second dynamic relationship data; A setting module, configured to determine a prediction result of whether a failure will occur in the large model training cluster during a target period according to whether the difference between the first dynamic relationship data and the second dynamic relationship data meets a preset warning condition, and set a checkpoint according to the prediction result; the global dynamic relationship data is the expected dynamic relationship of the entire model training process without failure.
[0015] This specification provides a management device for dynamic setting of checkpoints. The management device includes: a resource monitoring module, a mode prediction module, and a checkpoint setting module; The resource monitoring module is configured to collect resource usage data of the large model training cluster during the model training process and transmit it to the mode prediction module. The resource usage data is used to characterize the system resource occupancy status of the large model training cluster at each moment in the historical period. The mode prediction module is configured to predict first dynamic relationship data of the large model training cluster during the target period according to the resource usage data and transmit it to the checkpoint setting module; the first dynamic relationship data is used to characterize the dynamic relationship of the system resource occupancy status of the large model training cluster changing with time during the target period. The checkpoint setting module is configured to determine at least part of the dynamic relationship data matching the target period in the global dynamic relationship data as second dynamic relationship data, and determine a prediction result of whether a failure will occur in the large model training cluster during the target period according to whether the difference between the first dynamic relationship data and the second dynamic relationship data meets a preset warning condition, and set a checkpoint according to the prediction result; the global dynamic relationship data is the expected dynamic relationship of the entire model training process without failure.
[0016] This specification provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the above method for dynamic setting of checkpoints is implemented.
[0017] At least one of the above technical solutions adopted in this specification can achieve the following beneficial effects: In the method for dynamically setting checkpoints provided in this specification, first, resource usage data of the large model training cluster during the model training process is collected. The resource usage data is used to characterize the system resource occupancy status of the large model training cluster at each moment in the historical period. According to the resource usage data, first dynamic relationship data representing the change of the system resource occupancy status over time within the target period of the large model training cluster is predicted. At least part of the dynamic relationship data that matches the target period in the global dynamic relationship data is determined as the second dynamic relationship data. According to whether the difference between the first dynamic relationship data and the second dynamic relationship data meets the preset warning condition, a prediction result of whether a failure will occur in the large model training cluster within the target period is determined, and a checkpoint is set according to the prediction result. The global dynamic relationship data is the expected dynamic relationship of the entire model training process in the case of no failure.
[0018] As can be seen from the above method, before model training is performed through the large model training cluster, based on the historical model training process, the dynamic relationship representing the change of the expected system resource occupancy status over time during the entire process in the case of no abnormality during the training of the large model can be determined as the global dynamic relationship data. During the training process, the resource usage data of the large model training cluster in the previous historical period is periodically collected, so as to predict the first dynamic relationship data representing the change of the system resource occupancy status over time within the future target period based on the collected resource usage data. Furthermore, by comparing the first dynamic relationship data with the global dynamic relationship data, a prediction result of whether a failure may occur in the large model training cluster within the target period can be determined, and a checkpoint can be configured according to the prediction result, thereby improving the effectiveness of the checkpoints configured during the training of large-scale models and reducing the fixed overhead of periodic checkpoints and the loss of fault recovery. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The drawings described herein are used to provide a further understanding of this specification and form a part of this specification. The illustrative embodiments of this specification and their descriptions are used to explain this specification and do not constitute an improper limitation of this specification. In the drawings: Figure 1 It is a schematic flowchart of a method for dynamically setting checkpoints provided in this specification; Figure 2 It is a schematic diagram of the process of a method for dynamically setting checkpoints provided in this specification; Figure 3 It is a schematic diagram of the adjustment process provided in this specification; Figure 4 It is a schematic diagram of a device for dynamically setting checkpoints provided in this specification; Figure 5Schematic diagram of a management device for dynamic setting of checkpoints provided in this specification. Detailed implementation manners
[0020] To make the objectives, technical solutions, and advantages of this specification clearer, the technical solutions of this specification will be clearly and completely described below in conjunction with specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this specification.
[0021] The following details the technical solutions provided in each embodiment of this specification with reference to the drawings.
[0022] Currently, in order to address the problem of training interruption caused by software and hardware failures, existing training systems usually periodically save the intermediate state of model training for use in rolling back the training task to the nearest checkpoint in the event of a failure. However, due to problems such as the settings of checkpoints: checkpoints need to save the complete state of the model, including model weights, optimizer state, random number seeds, etc. As the model scale increases, the storage requirements for these data will also increase significantly (for example, a model with trillions of parameters may require hundreds of terabytes of storage space to save checkpoints), and the process of saving checkpoints will block the training process, resulting in training suspension. This time overhead depends on the size of the model and the performance of the storage system (for example, saving the checkpoint of a large model may take several hours). As a result, this periodically set checkpoint method will lead to relatively large fixed overheads of checkpoints and losses in fault recovery.
[0023] Figure 1 Schematic flowchart of a method for dynamic setting of checkpoints provided in this specification, including the following steps: S101: Collect resource usage data of the large model training cluster during the model training process, where the resource usage data is used to characterize the system resource occupancy status of the large model training cluster at each moment in the historical period.
[0024] In this specification, the above method for dynamic setting of checkpoints is applied to the management device of the large model training cluster. That is, the method for dynamic setting of checkpoints can be executed by the management device to dynamically manage the checkpoints to be set during the process of model training through the large model training cluster.
[0025] Among them, the above management device can be a specified node device selected from the large model training cluster, or a device deployed with a CPU, operating system, and task dispatcher for executing the method for dynamic setting of checkpoints.
[0026] Furthermore, during the process of model training by the large model training cluster, the management device can collect the resource usage data of the large model training cluster at specified time intervals.
[0027] Among them, the above-mentioned resource usage data is used to characterize the dynamic relationship of the system resource occupancy status of the large model training cluster changing over time within a historical period. Here, the resource usage data can be time series data composed of system resource occupancy data at different moments included in the historical period.
[0028] The above-mentioned system resource occupancy data can be data used to reflect the occupancy of system resources such as CPU, memory, and storage. For example: CPU occupancy rate, memory usage rate, disk occupancy rate, etc.
[0029] The above-mentioned historical period can be set according to actual needs. For example: a specified length of time period before the current moment is used as the historical period. Here, the specified length can be determined according to actual needs. For example: the time length required for the model to perform one round of iterative training is used as the specified length. Another example: each hour is used as the specified length.
[0030] S102: Predict the first dynamic relationship data of the system resource occupancy status of the large model training cluster changing over time within the target period according to the resource usage data.
[0031] In this specification, after collecting the resource usage data, the management device can perform time series analysis according to the resource usage data to predict the first dynamic relationship data of the system resource occupancy status of the large model training cluster changing over time within the target period.
[0032] Specifically, the management device can input the resource usage data into a pre-trained prediction model to predict the first dynamic relationship data of the large model training cluster within the target period through the prediction model according to the resource usage data.
[0033] It should be noted that during the process of model training by the large model training cluster, due to factors such as task scheduling strategies, resource allocation strategies, system administrator configurations for the large model training cluster, system administrator optimization habits for the large model training cluster, training stage characteristics of the model training itself, training data loading or processing conditions, etc., there may be different sub-data with regular and repetitive changes in the resource usage data during the process of model training by the large model training cluster.
[0034] For example, large model training clusters usually adopt specific task scheduling strategies, such as round-robin, priority scheduling, etc. These strategies will regularly allocate model training tasks to different nodes according to the priorities and resource requirements of the model training tasks. For example, a batch of high-priority tasks may be centrally scheduled during a certain period of each day, resulting in a fixed peak in the occupancy rates of CPU, memory, and hard disk resources during that period.
[0035] The different resource allocation strategies adopted by large model training clusters will cause the large model training clusters to reallocate resources at specific time points every day to optimize resource utilization. This may lead to short-term fluctuations in resource occupancy rates after reallocation.
[0036] If the system administrator configures the large model training cluster such that the node configurations in the cluster are unbalanced, resulting in some nodes possibly undertaking more computing tasks during specific time periods, it will lead to periodic changes in resource occupancy rates.
[0037] The optimization habits of the system administrator for the large model training cluster may cause the system administrator to regularly optimize the cluster, such as adjusting resource allocation parameters, upgrading software versions, etc. These optimization operations may change the pattern of resource occupancy rates in the short term, resulting in periodic changes.
[0038] Large model training is usually divided into multiple stages, such as data loading, model training, parameter update, etc. Different stages have different resource requirements, which may lead to periodic changes in resource occupancy rates. For example: The data loading stage may require a large number of I / O operations, resulting in a high occupancy rate of hard disk resources, while the model training stage may rely more on CPU and memory resources.
[0039] The size and complexity of the training data will also affect the resource occupancy rate. If a large amount of training data is loaded or processed during a certain period of each day, it may lead to periodic peaks in resource occupancy rates.
[0040] Based on this, in order to improve the accuracy of the determined first dynamic relationship data, the server can also extract at least one periodic component data from the resource usage data, and predict the first dynamic relationship data of the large model training cluster during the target period according to the periodic component data.
[0041] In the above content, the periodic component data refers to at least part of the data used to characterize the regular and repetitive changes in the resource usage data, and the regularities corresponding to different periodic component data are different.
[0042] For example: If the length of the historical period is the past 72 hours, and the resource usage data of the historical period is a time series composed of the system resource occupancy data for each hour in the past 72 hours. At this time, due to the fact that large model training is usually divided into multiple stages, and the impact of each stage on the system occupancy resource data is different, the resource usage data within every 48 hours in the resource usage data of the historical period shows a periodic fluctuation. At this time, a periodic component data extracted from the resource usage data of the historical period can be determined based on the resource usage data within every 48 hours in the resource usage data of the historical period.
[0043] For another example: If the length of the historical period is the past 72 hours, and the resource usage data of the historical period is a time series composed of the system resource occupancy data for each hour in the past 72 hours. At this time, due to the habit of the system administrator regularly optimizing the cluster, the resource usage data within every 24 hours in the resource usage data of the historical period shows a periodic fluctuation. At this time, a periodic component data extracted from the resource usage data of the historical period can be determined based on the resource usage data within every 24 hours in the resource usage data of the historical period.
[0044] Specifically, the management device can determine at least part of the data with regular and repetitive changes in the resource usage data as the periodic resource usage data segments. Then, according to the predicted key point extraction strategy, each periodic key point can be extracted from the periodic resource usage data segments, and the periodic component data can be obtained based on each periodic key point.
[0045] In the above content, the characteristic vector data used to represent each periodic key point can be determined according to the value of each periodic key point as the periodic component data, where the value of each dimension in the characteristic vector data is the value of each periodic key point.
[0046] In the above content, the key point extraction strategy is used to screen each data point included in the periodic resource usage data segments. The above key point extraction strategy can be, for example: starting from the first data point among each data point included in the periodic resource usage data segments, every specified number of data points (such as every 50 data points) is used as the selected key point.
[0047] For another example: The peak points, valley points, turning points (i.e., points where the data trend changes significantly, such as changing from an upward trend to a downward trend, or from a downward trend to an upward trend), mutation points (points where the data value suddenly changes greatly), and inflection points (points where the second derivative is zero, usually indicating a change in the curvature of the data) in each data point included in the periodic resource usage data segments are used as the selected key points.
[0048] Further, after extracting at least one periodic component data from the resource usage data, the management device can, for each periodic component data, pad the length of the periodic component data to a specified length to obtain adjusted periodic component data, and fuse the adjusted periodic component data to calculate the first dynamic relationship data of the large model training cluster within the target time period.
[0049] For example: If the length of a periodic component data is 24 and the specified length is 48, then the periodic component data can be copied and padded to a feature vector data with a length of 48 to obtain the adjusted periodic component data.
[0050] To make the prediction result more accurate, the management device can also fuse each adjusted periodic component data with preset residual data to calculate the first dynamic relationship data of the large model training cluster within the target time period. Specifically, the following formula can be referred to:
[0051] In the above formula, is the value of the i-th dimension in the calculated first dynamic relationship data of the large model training cluster within the target time period (that is, the feature vector data used to characterize the change of the system resource occupancy state of the large model training cluster over time determined according to each adjusted periodic component data), is the value of the i-th dimension included in the j-th adjusted periodic component data, is the value of the i-th dimension included in the residual data.
[0052] It can be seen from the above formula that the first dynamic relationship data of the large model training cluster within the target time period can be modeled based on the system resource occupancy state of the large model training cluster at each moment in the historical time period, so as to dynamically set the checkpoint according to the modeling result.
[0053] Among them, the above residual data is obtained by modeling the residual data used to determine the historical first dynamic relationship data of different historical time periods through an autoregressive model. Specifically, the following formula can be referred to:
[0054] In the above formula, is the error term, and are two parameters that can be adjusted according to actual needs. Of course, they can be prior data.
[0055] In addition, the method for the management device to calculate the first dynamic relationship data of the large model training cluster during the target period based on the periodic component data can also be to input the above periodic component data into a preset prediction model, so as to predict the dynamic relationship of the system resource occupancy state of the large model training cluster changing with time during the target period through the preset prediction model according to the periodic component data.
[0056] S103: Determine at least part of the dynamic relationship data that matches the target period in the global dynamic relationship data as the second dynamic relationship data.
[0057] S104: Determine the prediction result of whether a fault will occur in the large model training cluster during the target period according to whether the difference between the first dynamic relationship data and the second dynamic relationship data meets a preset warning condition, and set a checkpoint according to the prediction result; the global dynamic relationship data is the expected dynamic relationship of the entire model training process without faults.
[0058] In this specification, the management device can determine at least part of the dynamic relationship data that matches the target period in the global dynamic relationship data as the second dynamic relationship data. Furthermore, it can determine the prediction result of whether a fault will occur in the large model training cluster during the target period according to whether the difference between the first dynamic relationship data and the above second dynamic relationship data meets a preset warning condition. Furthermore, it can perform a checkpoint operation according to the prediction result to configure a checkpoint. Specifically, as Figure 2 shown.
[0059] Figure 2 is a schematic diagram of the process of a method for dynamically setting a checkpoint provided in this specification.
[0060] Combined with Figure 2 it can be seen that after each round of model training is completed, the management device can use the previous round of model training period as the historical period. Furthermore, it can determine the prediction result of whether a fault will occur in the large model training cluster during the next round of model training period, that is, the target period, according to the resource usage data during the historical period. Thus, it can save a checkpoint according to the prediction result, which can greatly reduce the fixed overhead and recovery loss of the checkpoint mechanism and improve the model training efficiency on the basis of recording the intermediate training state of the model and avoiding the impact of system faults on training.
[0061] It should be noted that the above global dynamic relationship data is the expected dynamic relationship of the entire model training process under normal circumstances, determined based on the historical model training process before model training. It can be understood that it is obtained by initially fusing the actual resource usage data at each moment except for failures collected during the entire model training process of the historical model training tasks executed by the large model training cluster.
[0062] The above warning conditions can be set according to actual needs. For example, if the difference between at least part of the dynamic relationship data in the above first dynamic relationship data and the global dynamic relationship data preset and matching the target time period is greater than the preset difference threshold, it can be considered that the preset warning conditions are met.
[0063] In the actual application scenario, in order to improve the accuracy of the prediction results, the management device can also adjust the parameters and key point extraction strategies used in the process of calculating the first dynamic relationship data of the large model training cluster within the target time period according to the change trend between the difference between the first dynamic relationship data and the second dynamic relationship data and the reference difference, specifically as Figure 3 shown.
[0064] Figure 3 This is the schematic diagram of the adjustment process provided in this specification.
[0065] Combined with Figure 3 it can be seen that the management device can determine the difference between the historical first dynamic relationship data in the historical time period and at least part of the dynamic relationship data in the global dynamic relationship data that matches the historical time period as the reference difference. When it is determined that the difference between the first dynamic relationship data and the second dynamic relationship data is greater than the reference difference, the parameters of the autoregressive model (i.e., at least one of the , and in the above formula) are adjusted to obtain the adjusted autoregressive model.
[0066] In addition, when the management device determines that the difference between the first dynamic relationship data and at least part of the dynamic relationship data in the preset global dynamic relationship data that matches the target time period is greater than the reference difference, it can add a limiting condition for screening each data point to the key point extraction strategy to obtain the adjusted key point extraction strategy, and re-extract the periodic component data according to the adjusted key point extraction strategy.
[0067] The above-mentioned restrictive conditions can be set according to actual requirements. For example, if the above-mentioned key point extraction strategy takes the data points every 10s in the segmented periodic resource usage data as the selected periodic key points. Then the above-mentioned newly added restrictive condition can be that the data points every 5s in the segmented periodic resource usage data are taken as the selected periodic key points.
[0068] For another example, if the above-mentioned key point extraction strategy takes the data points corresponding to the peaks or valleys in the segmented periodic resource usage data as the selected periodic key points. Then the above-mentioned newly added restrictive condition can be that the data points corresponding to the turning points in the segmented periodic resource usage data are taken as the selected periodic key points.
[0069] It should be noted that in the actual application scenario, if the difference between at least part of the dynamic relationship data that matches the target time period in the above-mentioned first dynamic relationship data and the preset global dynamic relationship data is greater than the reference difference, it indicates that there may be an error in the first dynamic relationship data of the large model training cluster calculated by the above method during the target time period.
[0070] Therefore, in order to avoid the adverse effects of errors, when the management device determines that the difference between at least part of the dynamic relationship data that matches the target time period in the above-mentioned first dynamic relationship data and the preset global dynamic relationship data is greater than the reference difference, it can directly set a checkpoint without considering the prediction result to prevent the possible inaccurate prediction result from having an adverse impact on the model training process. At the same time, the parameters and key point extraction strategy used in the process of calculating the first dynamic relationship data of the large model training cluster during the target time period are adjusted through the above adjustment method.
[0071] It can be seen from the above method that before training the model through the large model training cluster, the dynamic relationship of the expected system resource occupancy state changing with time during the entire process without anomalies during the training of the large model can be determined according to the historical model training process as the global dynamic relationship data. During the training process, the resource usage data of the large model training cluster in the previous historical time period is periodically collected to predict the first dynamic relationship data used to characterize the dynamic relationship of the system resource occupancy state changing with time during the future target time period based on the collected resource usage data. Furthermore, the first dynamic relationship data can be compared with the global dynamic relationship data to determine the prediction result of whether the large model training cluster may have a failure during the target time period, and the checkpoint can be configured according to the prediction result, thereby improving the effectiveness of the checkpoint configured during the training of the large-scale model to reduce the fixed overhead of the periodic checkpoint and the loss of fault recovery.
[0072] The above is a method for dynamically setting one or more implementation checkpoints of this specification. Based on the same idea, this specification also provides a corresponding device for dynamically setting checkpoints, as Figure 4 shown.
[0073] Figure 4 is a schematic diagram of a device for dynamically setting checkpoints provided by this specification, including: An acquisition module 401, configured to acquire resource usage data of a large model training cluster during the model training process, where the resource usage data is used to characterize the system resource occupancy status of the large model training cluster at each moment in a historical period; A prediction module 402, configured to predict first dynamic relationship data indicating how the system resource occupancy status of the large model training cluster changes over time within a target period based on the resource usage data; A determination module 403, configured to determine at least part of the dynamic relationship data that matches the target period in the global dynamic relationship data as second dynamic relationship data; A setting module 404, configured to determine a prediction result indicating whether a failure will occur in the large model training cluster within the target period based on whether the difference between the first dynamic relationship data and the second dynamic relationship data meets a preset warning condition, and set a checkpoint according to the prediction result; the global dynamic relationship data is the expected dynamic relationship of the entire model training process without failure.
[0074] Optionally, the prediction module 402 is specifically configured to extract at least one periodic component data from the resource usage data, where the periodic component data refers to at least part of the data used to characterize the regular and repetitive changes in the resource usage data, and the regularities corresponding to different periodic component data are different; and predict the first dynamic relationship data of the large model training cluster within the target period according to the periodic component data.
[0075] Optionally, the prediction module 402 is specifically configured to, for each periodic component data, fill the length of the periodic component data to a specified length to obtain adjusted periodic component data; and fuse the adjusted periodic component data to predict the first dynamic relationship data of the large model training cluster within the target period.
[0076] Optionally, the prediction module 402 is specifically configured to fuse each adjusted periodic component data with preset residual data to predict the first dynamic relationship data of the large model training cluster within the target period; the residual data is obtained by modeling the residual data used when determining the historical first dynamic relationship data of different historical periods through an autoregressive model.
[0077] Optionally, the prediction module 402 is further configured to determine a reference difference between the historical first dynamic relationship data of the historical period and at least part of the dynamic relationship data in the global dynamic relationship data that matches the historical period; when it is determined that the difference is greater than the reference difference, adjust the parameters of the autoregressive model to obtain an adjusted autoregressive model.
[0078] Optionally, the prediction module 402 is specifically configured to determine at least part of the data with regular and repetitive changes in the resource usage data as periodic resource usage data segments; according to a predicted key point extraction strategy, extract each periodic key point from the periodic resource usage data segments, and obtain periodic component data according to each periodic key point, where the key point extraction strategy is used to screen each data point included in the periodic resource usage data segments.
[0079] Optionally, the prediction module 402 is specifically configured to determine a reference difference between the historical first dynamic relationship data of the historical period and at least part of the dynamic relationship data in the global dynamic relationship data that matches the historical period; when it is determined that the difference is greater than the reference difference, add a constraint condition for screening each data point to the key point extraction strategy to obtain an adjusted key point extraction strategy; according to the adjusted key point extraction strategy, re-extract the periodic component data.
[0080] For ease of understanding, this specification also provides a dynamic checkpoint setting system, specifically as Figure 5 shown.
[0081] Figure 5 It is a schematic diagram of a management device for dynamic checkpoint setting provided in this specification.
[0082] Combined Figure 5 It can be seen that the above management device includes: a resource monitoring module, a mode prediction module, and a checkpoint setting module.
[0083] Among them, the resource monitoring module is configured to collect resource usage data of the large model training cluster during the model training process and transmit it to the mode prediction module. The resource usage data here is used to characterize the system resource occupancy status of the large model training cluster at each moment in the historical period.
[0084] The mode prediction module is configured to calculate the first dynamic relationship data of the large model training cluster within the target period according to the resource usage data and transmit it to the checkpoint setting module. The first dynamic relationship data here is used to characterize the dynamic relationship of the system resource occupancy status of the large model training cluster changing with time within the target period.
[0085] The checkpoint setting module is used to determine at least part of the dynamic relationship data that matches the target time period in the global dynamic relationship data as the second dynamic relationship data, and determine the prediction result of whether the large model training cluster will fail during the target time period according to whether the difference between the first dynamic relationship data and the second dynamic relationship data meets the preset warning condition, and set the checkpoint according to the prediction result.
[0086] The above-mentioned global dynamic relationship data is determined according to the historical model training process before model training.
[0087] It should be noted that the above-mentioned resource monitoring module and pattern prediction module include routines, programs, objects, components, data structures, etc. that execute the above-mentioned dynamic setting method of the checkpoint. In the large model training environment, the added resource monitoring module, pattern prediction module, and checkpoint setting module responsible for recording the checkpoints of the large model training execute the above-mentioned dynamic setting method of the checkpoint. In the large model training environment, the checkpoint setting module is usually located in the nodes included in the large model training cluster.
[0088] It can be seen from the above that before model training through the large model training cluster, the dynamic relationship of the expected system resource occupancy status changing with time during the entire process without anomalies during the training of the large model can be determined according to the historical model training process as the global dynamic relationship data. During the training process, the resource usage data of the large model training cluster in the previous historical time period can be periodically collected through the resource monitoring module, pattern prediction module, and checkpoint setting module included in the management device, so as to predict the first dynamic relationship data used to represent the dynamic relationship of the system resource occupancy status changing with time during the future target time period. Furthermore, by comparing the first dynamic relationship data with the global dynamic relationship data, the prediction result of whether the large model training cluster may fail during the target time period can be determined, and the checkpoint can be configured according to the prediction result, thereby improving the accuracy of the checkpoint configured during the training of the large-scale model, and improving the training efficiency of the model and the utilization rate of system resources.
[0089] This specification also provides a computer-readable storage medium, which stores a computer program that can be used to execute the above Figure 1 provided dynamic setting method of a checkpoint.
[0090] Those skilled in the art should understand that the embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, this specification can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0091] This specification can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0092] The embodiments in this specification are described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and reference can be made to the corresponding parts of the method embodiments for the relevant content.
[0093] The above description is only for the embodiments of this specification and is not intended to limit this specification. For those skilled in the art, various changes and modifications can be made to this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification should be included within the scope of the claims of this specification.
Claims
1. A method for dynamically setting a checkpoint, characterized in that: The method is applied to a management device of a large model training cluster, and the method comprises: Collect resource usage data of the large model training cluster during the model training process, where the resource usage data is used to characterize the system resource occupancy status of the large model training cluster at each moment in a historical period; Predicting first dynamic relationship data of changes in system resource occupancy status of the large model training cluster over time within a target period of time based on the resource usage data; Determine at least a portion of the dynamic relationship data in the global dynamic relationship data that matches the target time period as second dynamic relationship data; According to whether the difference between the first dynamic relationship data and the second dynamic relationship data meets the preset warning condition, the prediction result of whether the large model training cluster will fail within the target period is determined, and checkpoints are set according to the prediction result; the global dynamic relationship data is the expected dynamic relationship of the entire model training process when no failure occurs.
2. The method according to claim 1, characterized in that Predicting the first dynamic relationship data of the large model training cluster within a target period according to the resource usage data specifically includes: Extracting at least one periodic component data from the resource usage data, wherein the periodic component data refers to at least part of the data used to characterize regular and repetitive changes in the resource usage data, and different periodic component data correspond to different regularities; Based on the periodic component data, the first dynamic relationship data of the large model training cluster within the target time period is predicted.
3. The method according to claim 2, characterized in that Predicting the first dynamic relationship data of the large model training cluster within a target period according to the periodic component data specifically includes: For each periodic component data, fill the length of the periodic component data to a specified length to obtain adjusted periodic component data; The adjusted period component data are fused to predict the first dynamic relationship data of the large model training cluster within the target time period.
4. The method according to claim 3, characterized in that The adjusted period component data are integrated to predict the first dynamic relationship data of the large model training cluster within the target period, specifically including: The adjusted period component data are fused with the preset residual data to predict the first dynamic relationship data of the large model training cluster within the target period; the residual data is obtained by modeling the residual data used to determine the historical first dynamic relationship data of different historical periods through an autoregressive model.
5. The method according to claim 4, characterized in that The method further comprises: Determine a difference between the historical first dynamic relationship data of the historical period and at least a portion of dynamic relationship data in the global dynamic relationship data that matches the historical period as a reference difference; When it is determined that the difference is greater than the reference difference, the parameters of the autoregressive model are adjusted to obtain an adjusted autoregressive model.
6. The method according to claim 2, characterized in that Extracting at least one periodic component data from the resource usage data specifically includes: Determine at least a portion of the resource usage data that has regular and repetitive changes as a periodic resource usage data segment; According to the predicted key point extraction strategy, each periodic key point is extracted from the periodic resource usage data segment, and based on the each periodic key point, the periodic component data is obtained. The key point extraction strategy is used to screen the data points contained in the periodic resource usage data segment.
7. The method according to claim 6, characterized in that The method further comprises: Determine a difference between the historical first dynamic relationship data of the historical period and at least a portion of dynamic relationship data in the global dynamic relationship data that matches the historical period as a reference difference; When it is determined that the difference is greater than the reference difference, adding a restriction condition for screening the data points in the key point extraction strategy to obtain an adjusted key point extraction strategy; According to the adjusted key point extraction strategy, the periodic component data is re-extracted.
8. A device for dynamically setting a checkpoint, characterized in that: include: A collection module, used to collect resource usage data of a large model training cluster during model training, wherein the resource usage data is used to characterize the system resource occupancy status of the large model training cluster at each moment in a historical period; A prediction module, configured to predict first dynamic relationship data of changes in the system resource occupancy state of the large model training cluster over time within a target period of time based on the resource usage data; A determination module, configured to determine at least a portion of the dynamic relationship data in the global dynamic relationship data that matches the target time period as second dynamic relationship data; a setting module, configured to determine a prediction result of whether the large model training cluster will fail within a target period according to whether the difference between the first dynamic relationship data and the second dynamic relationship data meets a preset warning condition, and to set a checkpoint according to the prediction result; The global dynamic relationship data is the expected dynamic relationship of the entire model training process when no failure occurs.
9. A management device for dynamic setting of checkpoints, characterized in that: The management device includes: a resource monitoring module, a mode prediction module, and a checkpoint setting module; The resource monitoring module is used to collect resource usage data of the large model training cluster during the model training process and transmit it to the pattern prediction module, wherein the resource usage data is used to characterize the system resource occupancy status of the large model training cluster at each moment in the historical period; The pattern prediction module is used to predict the first dynamic relationship data of the large model training cluster in the target time period according to the resource usage data and transmit it to the checkpoint setting module; the first dynamic relationship data is used to characterize the dynamic relationship of the system resource occupancy state of the large model training cluster in the target time period changing with time; The checkpoint setting module is used to determine at least part of the dynamic relationship data in the global dynamic relationship data that matches the target time period as the second dynamic relationship data, and determine the prediction result of whether the large model training cluster will fail within the target time period based on whether the difference between the first dynamic relationship data and the second dynamic relationship data meets the preset warning condition, and set a checkpoint based on the prediction result; the global dynamic relationship data is the expected dynamic relationship of the entire model training process when no failure occurs.
10. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the method described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Detecting method and server
CN103902437A
Virtual resource scheduling method and device, electronic equipment, storage medium and product
CN119336450A
Large model training fault recovery method and device based on distributed memory management
CN119473732A
Traffic flow prediction and prediction model training method and device
CN119851460A
Resource scheduling method and model training method of large model
CN119883615A