An intelligent monitoring and operation and maintenance guarantee method for IBM middleware software running state
Patent Information
- Application Number
- CN202611123292.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-28
- Publication Date
- 2026-09-29
AI Technical Summary
[0008]为此,本发明提供一种IBM中间件软件运行状态智能监测与运维保障方法,用以克服现有技术中无法结合理论计算法和实测统计法进行IBM大模型训练中间件软件的训练GPU显存的精准预测,进而无法实现提高其运维保障的GPU集群显存资源利用率的问题
[0019]与现有技术相比,本发明的有益效果在于,本发明解决了传统理论计算方法中动态激活值估算缺失、硬件环境差异未考虑的缺陷,通过轻量残差全连接网络对理论计算值与实测值的偏差进行自适应学习,进一步将最终预测显存下限的平均相对误差降低,基于高精度的显存预测结果,能够为每个大模型训练任务推荐最优的GPU配置组合,避免了经验式配置中常见的显存不足导致训练崩溃和过度配置造成计算资源浪费的问题。
Smart Images

Figure CN122838221A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of operation and maintenance support, and in particular to a method for intelligent monitoring and operation and maintenance support of IBM middleware software running status. Background Technology
[0002] With the widespread application of large-scale deep learning models in fields such as natural language processing and computer vision, IBM has launched middleware software platforms for training large models, such as IBMwatsonx.ai and IBMWatson MachineLearning Accelerator, providing enterprise users with distributed training, model tuning, and deployment services. When these middleware platforms handle large model training tasks, they heavily rely on GPU accelerators, and the efficient operation and maintenance of GPU memory resources directly determines the cost of the training task.
[0003] During the operation of IBM middleware software, large models need to be trained and fine-tuned. Before the formal deployment of the training and fine-tuning task begins, operations and maintenance personnel need to pre-assess the minimum amount of GPU memory required for the task to run, i.e., the lower limit of GPU memory, in order to allocate appropriate GPU configurations for the target job. If the GPU configuration is too high, it will lead to the waste of expensive computing resources. Conversely, if the GPU configuration is insufficient, it is very easy to cause GPU memory overflow failures in the IBM middleware software, causing unexpected interruptions in the training process, seriously disrupting business continuity and increasing the operational burden.
[0004] Currently, the industry's assessment of the GPU memory requirements for large model training tasks mainly relies on empirical theoretical formulas. A common approach is to use the model's structural parameters and training hyperparameters to estimate the static GPU memory usage of model parameters, optimizer states, etc., and to roughly estimate the dynamic activation value GPU memory during training based on batch size, sequence length, etc.
[0005] However, in complex distributed training scenarios managed by IBM's large model training middleware, theoretical estimation methods only model standard data parallelism modes and fail to fully incorporate GPU card partitioning parameters and advanced strategies such as tensor parallelism and pipeline parallelism. This results in a serious lack of accuracy in calculating dynamic activation values during training, and a significant deviation between the calculated theoretical lower limit of video memory and the actual resource requirements.
[0006] In actual training environments, memory usage is also affected by a combination of factors such as runtime framework implementation details, memory fragmentation, temporary buffer allocation, and aggregate communication overhead. This results in a non-linear difference between the measured lower limit of memory usage and the theoretical calculated value under the same configuration, which is difficult to describe with simple analytical rules.
[0007] Therefore, how to combine theoretical calculation methods and experimental statistical methods to accurately predict the training GPU memory of IBM's large model training middleware software, so as to improve the utilization rate of GPU cluster memory resources for its operation and maintenance, is a technical problem that needs to be solved. Summary of the Invention
[0008] To address this, the present invention provides an intelligent monitoring and maintenance support method for IBM middleware software operation status, which overcomes the problem in the prior art that it is impossible to combine theoretical calculation methods and experimental statistical methods to accurately predict the training GPU memory of IBM large model training middleware software, and thus cannot improve the utilization rate of GPU cluster memory resources for its maintenance support.
[0009] To achieve the above objectives, this invention proposes a method for intelligent monitoring and maintenance assurance of IBM middleware software runtime status, comprising: The runtime parameters of IBM's large model training middleware software are converted to obtain static memory values; The structural parameters of the IBM large model training middleware software and the GPU card partitioning parameters are converted into tensor parallel activation values to calculate the training dynamic activation values, and the theoretical lower limit of video memory is calculated based on the static video memory values and the training dynamic activation values. The IBM large model training middleware software was used to obtain the measured lower limit of video memory through a cluster monitoring algorithm. After using the difference between the theoretical memory lower limit and the measured memory lower limit of the IBM large model training middleware software to train a lightweight residual fully connected network, the newly uploaded structural training parameters and theoretical memory lower limit of the IBM large model training middleware software are passed through the lightweight residual fully connected network to generate a predicted memory lower limit. The predicted lower bound of GPU memory based on the candidate GPU cluster configuration combinations is used to determine the recommended GPU configuration for operational support of IBM's large model training middleware software.
[0010] Furthermore, the process of training a lightweight residual fully connected network includes: The theoretical lower bound of video memory and the structural training parameters of IBM's large model training middleware software are passed through a lightweight residual fully connected network to generate the original theoretical lower bound of prediction. The logarithm of the ratio of the original theoretical lower bound to the measured lower bound is used to train a lightweight residual fully connected network through a quantile loss function.
[0011] Furthermore, the process of generating the predicted lower bound of video memory includes: The structural training parameters and theoretical lower limit of the GPU memory of the IBM large model training middleware software are passed through the first fully connected layer to generate the original GPU memory features. The original memory features are passed through a second fully connected layer to generate higher-order memory features; The higher-order memory features are passed through the output layer to generate a predicted lower bound for memory. The lightweight residual fully connected network includes a first fully connected layer, a second fully connected layer, and an output layer.
[0012] Furthermore, the process of obtaining the static memory value includes: The sum of the corresponding byte amounts calculated according to the running parameter categories is then multiplied by the total amount of the running parameters to obtain the static video memory value. The operating parameters include operating parameter categories.
[0013] Furthermore, the categories of operating parameters include the training accuracy type of the forward parameters, whether the master parameter copy is enabled for mixed accuracy, the forward accuracy of the gradient, and the optimizer type.
[0014] Furthermore, the process of calculating the dynamic activation values during training includes: Based on the structure type and activation tensor parameter quantity of the IBM large model training middleware software, the single micro-batch activation value is calculated. Based on the single micro-batch activation value and the number of GPU cards to be divided, the single card activation value is obtained; The training dynamic activation value is obtained by multiplying the number of single-card calculation layers, the average number of micro-batches in residence, and the single-card activation value. The GPU card segmentation parameters include the number of GPU cards to be segmented, the number of computing layers per card, and the average number of resident micro-batches.
[0015] Furthermore, the process of obtaining the measured lower limit of video memory includes: The IBM large model training middleware software uses a training framework callback algorithm to obtain multiple memory distribution values over multiple iterations, and calculates the steady-state memory value based on these multiple memory distribution values. The IBM large model training middleware software is monitored using a cluster monitoring algorithm to obtain the net increase in video memory usage. The measured lower limit of video memory is calculated by summing the steady-state video memory value and the video memory usage.
[0016] Furthermore, the process of calculating the steady-state memory value includes: The steady-state memory value is obtained by subtracting the memory distribution value corresponding to the set quantile from the peak value of the memory distribution value within the steady-state window.
[0017] Furthermore, the process of determining the recommended GPU configuration for operational support includes: If the physical video memory of the candidate GPU cluster configuration combination is greater than the predicted lower limit of video memory, then the candidate GPU cluster configuration combination will be used as the recommended GPU configuration for operation and maintenance assurance. If the physical video memory of the candidate GPU cluster configuration combination is less than or equal to the predicted lower limit of video memory, then the parallelism of the candidate GPU cluster configuration combination is increased, and the predicted lower limit of video memory is recalculated.
[0018] Furthermore, the structure training parameters include the total number of trainable parameters after normalization, the number of model layers, the hidden dimension, the number of attention heads, the vocabulary size, the maximum sequence length used for training, the forward computation precision, whether mixed precision is enabled, the optimizer type, the GPU micro-batch size, the single GPU physical memory capacity, and the number of configured nodes.
[0019] Compared with the prior art, the beneficial effects of the present invention are that it solves the defects of the traditional theoretical calculation method, such as the lack of dynamic activation value estimation and the failure to consider the differences in hardware environment. By using a lightweight residual fully connected network to adaptively learn the deviation between the theoretical calculation value and the measured value, the average relative error of the final prediction of the lower limit of video memory is further reduced. Based on the high-precision video memory prediction results, the optimal GPU configuration combination can be recommended for each large model training task, avoiding the problems of insufficient video memory leading to training crashes and over-configuration causing waste of computing resources that are common in empirical configuration.
[0020] In particular, this invention first calculates the activation value of a single micro-batch based on the model structure type and the number of activation tensor parameters, then obtains the activation value of a single GPU based on the number of GPUs split, and finally introduces two key parameters, the number of computational layers per GPU and the average number of micro-batches residing, to calculate the final dynamic activation value for training. This fully considers the distribution characteristics of activation values and the residing characteristics of the computational pipeline in the tensor parallel implementation of IBM's large model training middleware, avoiding the serious underestimation of dynamic activation values caused by traditional methods that do not consider tensor parallel splitting and pipeline residing. It also avoids the common problems in empirical configurations, such as insufficient GPU memory leading to training crashes and over-configuration causing waste of computational resources. Attached Figure Description
[0021] Figure 1 This is a flowchart illustrating the IBM middleware software runtime status intelligent monitoring and operation and maintenance assurance method according to an embodiment of the present invention; Figure 2 This is a flowchart illustrating the calculation of the measured lower limit of video memory in the IBM middleware software runtime status intelligent monitoring and operation and maintenance assurance method according to an embodiment of the present invention. Figure 3 This is a flowchart illustrating the calculation of dynamic activation values for the IBM middleware software runtime status intelligent monitoring and maintenance assurance method according to an embodiment of the present invention. Figure 4This is a flowchart illustrating the IBM large model training middleware software used in the IBM middleware software intelligent monitoring and maintenance assurance method for the running status of IBM middleware software, as described in an embodiment of the present invention. Detailed Implementation
[0022] To make the objectives and advantages of the present invention clearer, the present invention will be further described below with reference to embodiments; it should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.
[0023] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.
[0024] It should be noted that in the description of this invention, the terms "upper", "lower", "left", "right", "inner", "outer", etc., which indicate directions or positional relationships, are based on the directions or positional relationships shown in the accompanying drawings. This is only for the convenience of description and is not intended to indicate or imply that the device or element must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation of this invention.
[0025] Furthermore, it should be noted that, in the description of this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0026] like Figures 1 to 4 As shown, this invention provides an intelligent monitoring and maintenance assurance method for the running status of IBM middleware software, which overcomes the problem in the prior art that it is impossible to combine theoretical calculation methods and experimental statistical methods to accurately predict the training GPU memory of IBM large model training middleware software, and thus cannot improve the utilization rate of GPU cluster memory resources for its maintenance assurance.
[0027] like Figure 1 and 2 As shown, this embodiment proposes a method for intelligent monitoring and operation and maintenance assurance of IBM middleware software runtime status, including: The runtime parameters of IBM's large model training middleware software are converted to obtain static memory values; The structural parameters of the IBM large model training middleware software and the GPU card partitioning parameters are converted into tensor parallel activation values to calculate the training dynamic activation values, and the theoretical lower limit of video memory is calculated based on the static video memory values and the training dynamic activation values. The IBM large model training middleware software was used to obtain the measured lower limit of video memory through a cluster monitoring algorithm. After using the difference between the theoretical memory lower limit and the measured memory lower limit of the IBM large model training middleware software to train a lightweight residual fully connected network, the newly uploaded structural training parameters and theoretical memory lower limit of the IBM large model training middleware software are passed through the lightweight residual fully connected network to generate a predicted memory lower limit. The predicted lower bound of GPU memory based on the candidate GPU cluster configuration combinations is used to determine the recommended GPU configuration for operational support of IBM's large model training middleware software.
[0028] Specifically, the IBM large model training middleware software is built on IBM's Watsonx.ai TrainingMiddleware v2.3. The above process is used to determine the predicted lower bound of the GPU cluster configuration combination of the IBM large model training middleware software during the fine-tuning training process through the TuningStudio component, and to determine the recommended GPU combination in the candidate GPU cluster configuration combination based on the predicted lower bound, thereby optimizing the utilization of computing resources for the fine-tuning training operation and maintenance of the IBM large model training middleware software.
[0029] Understandably, the theoretical lower bound of video memory is handled by analytical formulas, capturing deterministic, linear physical laws. Feeding these parameters into a lightweight residual fully connected network allows it to focus solely on learning nonlinear biases that are difficult to describe with formulas and are strongly correlated with specific scenarios. For example, the network might learn a rule: when the training precision type of the parameters is 16-bit floating-point, the GPU batch size is 8, and the single GPU physical memory capacity is 40GB, the theoretical value needs to be increased by 6%; while when the training precision type is 32-bit floating-point, the GPU batch size is 1, and the single GPU physical memory capacity is 80GB, only a 2% increase is needed. This achieves a gray-box division of labor between the physical baseline and scene correction to generate the predicted lower bound of video memory. The newly uploaded IBM large model training middleware software can obtain the lower bound of the prediction memory by static parsing or directly reading the configuration without actually running the code for training. For new models or new parallel combinations that the system has never seen before, as long as their feature vectors fall in the space spanned by this set of parameters, the residual network can use the patterns obtained from training with historical samples to give reasonable bias predictions, instead of needing to retrain with massive amounts of samples for each new configuration, as is the case with pure black-box models.
[0030] Furthermore, the process of obtaining the static memory value includes: The sum of the corresponding byte amounts calculated according to the running parameter categories is then multiplied by the total amount of the running parameters to obtain the static video memory value. The operating parameters include operating parameter categories.
[0031] Furthermore, the categories of operating parameters include the training accuracy type of the forward parameters, whether the master parameter copy is enabled for mixed accuracy, the forward accuracy of the gradient, and the optimizer type.
[0032] Specifically, during training, the tensors bound to each trainable parameter and never released contain multiple categories. Based on the training configuration, the size of a single parameter in bytes is determined for each category. Specifically, the forward parameters are copies of the parameters actually involved in forward and backward computation during training in the IBM large model training middleware software. Their corresponding code is `params_fwd`, and their size in bytes is determined by the training precision type, including: 4 bytes for 32-bit floating-point training precision, 2 bytes for 16-bit floating-point training precision, and 1 byte for 8-bit floating-point training precision. Therefore, this describes the total static workspace required for training this IBM large model training middleware software from the perspective of forward parameters.
[0033] The master parameter copy is an additional high-precision copy of the parameters retained by the optimizer during mixed-precision training in the IBM Large Model Training Middleware software. Its corresponding code is `params_master`, and its size in bytes is determined by whether mixed precision is enabled and the forward precision, including: 4 bytes for mixed precision enabled and the forward precision is not 32-bit floating-point; 0 bytes for mixed precision enabled and the forward precision is 32-bit floating-point; and 0 bytes for not enabling mixed precision. Therefore, this describes the total static workspace required for training this IBM Large Model Training Middleware software from the perspective of the master parameter copy.
[0034] The gradient is calculated via backpropagation during training in the IBM large model training middleware software. Its corresponding code is `grads`, and its size in bytes typically matches the precision of the forward parameters: 4 bytes for 32-bit floating-point precision, 2 bytes for 16-bit floating-point precision, and 1 byte for 8-bit floating-point precision. Therefore, this implementation describes the total static workspace required to train this IBM large model training middleware software from the perspective of gradients.
[0035] The optimizer type refers to the additional state variables maintained by the optimizer for each parameter. Its corresponding code is `opt_state`, and its size in bytes is determined by the optimizer type, the number of maintained variables, and the precision of the state. These include: 0 bytes for optimizer type SGD with no maintained variables; 4 bytes for optimizer type SGD with maintained variables; 8 bytes for optimizer type Adam or AdamW; 8 bytes for optimizer type LAMB; 4 bytes for optimizer type AdaGrad; and 2 bytes for optimizer type 8-bitAdam. Therefore, this describes the total static workspace required to train this large IBM model training middleware software from the perspective of optimizer type.
[0036] Therefore, the various tensors that must reside persistently in GPU memory during training are converted to bytes based on the average number of bytes per parameter. After summing these, they are multiplied by the total number of running parameters to obtain the full static GPU memory value. This allows for a fast estimation of static GPU memory based on the summation of parameter category coefficients. This static GPU memory value serves as the theoretical static GPU memory baseline. Subsequently, it will be divided and converted according to the parallel strategies of TP and PP to finally obtain the theoretical lower limit of static GPU memory for a single card. Through the above-mentioned dimensional division, the entire gray-box scheduling method can be more accurate than pure formulas or pure black-box models because the theoretical static GPU memory value ensures that the prediction results do not violate physical laws, while the residual network focuses on correcting the specific biases brought about by the real world.
[0037] like Figure 2 As shown, the process of calculating the training dynamic activation values further includes: Based on the structure type and activation tensor parameter quantity of the IBM large model training middleware software, the single micro-batch activation value is calculated. Based on the single micro-batch activation value and the number of GPU cards to be divided, the single card activation value is obtained; The training dynamic activation value is obtained by multiplying the number of single-card calculation layers, the average number of micro-batches in residence, and the single-card activation value. The GPU card segmentation parameters include the number of GPU cards to be segmented, the number of computing layers per card, and the average number of resident micro-batches.
[0038] Specifically, the process of obtaining the single-micro-batch activation value is as follows: If the structure type is an attention mechanism, the single-layer single-micro-batch activation value is the activation tensor parameter quantity output by the linear transformation, plus the activation tensor parameter quantity output by the attention score matrix, plus the activation tensor parameter quantity output by the attention. The activation tensor parameter quantity is calculated based on the batch size, sequence length, and hidden layer dimension. If the structure type is a feedforward network, the single-layer single-micro-batch activation value is the batch size of the activation tensor parameter quantity output by the feedforward network multiplied by the sequence length multiplied by the hidden layer dimension. Based on the structure of the historical IBM large model training middleware software, that is, the number of attention mechanisms and feedforward networks set in it multiplied by the corresponding single-layer single-micro-batch activation value and summed, the single-micro-batch activation value is obtained. Therefore, using the activation tensors that must be retained in the forward propagation of the historical IBM large model training middleware software, through static code analysis, the number of attention layers and feedforward network layers in the model is obtained, multiplied by their respective single-layer activation values and summed, to obtain the total number of activation bytes that the entire model theoretically needs to retain when processing a micro-batch.
[0039] Specifically, the single-micro-batch activation value is divided by the number of GPUs to be split, resulting in the single-GPU activation value. The number of GPUs to be split refers to the number of GPUs on which a layer of the model is computed; this process is executed via NVLink's high-speed interconnect. Therefore, the core idea of using IBM's large model training middleware software Tensor Parallel Processing (TP) is to split the parameters and computations within a layer along a specific dimension across multiple GPUs. A high-precision single-GPU activation baseline can be obtained using division, eliminating the need for recalculating layer by layer.
[0040] Specifically, in the Tensor Parallel (TP) configuration, layer L is evenly divided into p pipeline stages according to depth. Each GPU is only responsible for calculating the single-card computational layer number L / p. The single-card computational layer number is then multiplied by the average number of resident micro-batches, and then by the single-card activation value to obtain the training dynamic activation value. The average number of resident micro-batches is the minimum of the pipeline depth and the total number of micro-batches. Therefore, by utilizing the core idea of the historical IBM large model training middleware software Pipeline Parallel (PP), the model is divided according to depth, and each device is only responsible for calculating the single-card computational layer number. Thus, the activations that a single GPU needs to retain only come from that single-card computational layer number. This aligns with the requirement that to maintain full pipeline load, GPU devices must process multiple micro-batches simultaneously. In the stable phase, multiple micro-batches of activation will reside on a single GPU device simultaneously. This average number of resident micro-batches is determined by the scheduling algorithm and the total number of micro-batches.
[0041] Specifically, the static memory value and the training dynamic activation value are added together to obtain the theoretical lower limit of memory. This theoretical lower limit of memory is the derived minimum memory budget of the GPU, which is the ideal minimum memory requirement for a single card.
[0042] like Figure 3 As shown, the process of obtaining the measured lower limit of video memory further includes: The IBM large model training middleware software uses a training framework callback algorithm to obtain multiple memory distribution values over multiple iterations, and calculates the steady-state memory value based on these multiple memory distribution values. The IBM large model training middleware software is monitored using a cluster monitoring algorithm to obtain the net increase in video memory usage. The measured lower limit of video memory is calculated by summing the steady-state video memory value and the video memory usage.
[0043] Specifically, the cluster monitoring algorithm takes the GPU's memory usage for a short period before the job starts as the environmental baseline, takes the maximum memory usage within the job time window, and subtracts the environmental baseline from the maximum memory usage to obtain the net increase in memory usage.
[0044] Furthermore, the process of calculating the steady-state memory value includes: The steady-state memory value is obtained by subtracting the memory distribution value corresponding to the set quantile from the peak value of the memory distribution value within the steady-state window.
[0045] Specifically, by inserting monitoring points during the forward and backward propagation process of the IBM large model training middleware software, monitoring data from multiple consecutive iterations of the framework callback algorithm are collected. Only when the monitoring data tends to stabilize, i.e., the standard deviation of three adjacent monitoring data is less than a threshold, can a steady-state window be obtained.
[0046] Understandably, the peak value of the memory distribution within the steady-state window represents the highest level of memory usage during the forward, backward, and update cycles of a complete historical IBM large model training middleware software training process, when memory is occupied by the cumulative amounts of activation values, gradients, optimizer states, etc. The preferred quantile is the 1st percentile, meaning that 1% of the sampled values are below or equal to this value, while 99% of the sampled values are above it. The 1st percentile memory distribution value represents the lowest level of memory usage within an iteration cycle, at which point gradient synchronization is complete, activation values are released, and only persistent data such as running parameters and optimizer states remain.
[0047] Specifically, the process of obtaining the theoretical lower bound of memory includes: The difference between the net increase in video memory usage and the video memory distribution value at the same moment is used to obtain the cluster hardware offset value based on the median of multiple differences within the steady-state window. The steady-state video memory value, video memory usage, and cluster hardware offset value are summed to obtain the theoretical lower limit of video memory.
[0048] Specifically, the memory usage refers to the dynamic peak increment of memory usage that fluctuates periodically during training. The difference between the net increase in memory usage and the memory distribution value at any given moment represents a fixed hardware overhead introduced by the driver, NCCL communication library, GPU context, etc., that cannot be directly observed by the framework. Taking the median of the differences at multiple moments, rather than the mean, is to eliminate occasional transient spikes.
[0049] Therefore, the steady-state memory value in the theoretical lower bound reflects the parameters, gradients, and optimizer state, and persists continuously within the process. The memory usage reflects the activation value and temporary buffer, which fluctuates periodically with iteration. The cluster hardware offset reflects the out-of-process overhead such as drivers, communication libraries, and context. These three levels are physically independent, so their usage can be directly linearly added together to form a complete picture of memory usage, thus achieving a complete acquisition of the theoretical lower bound.
[0050] Furthermore, the process of training a lightweight residual fully connected network includes: The theoretical lower bound of video memory and the structural training parameters of IBM's large model training middleware software are passed through a lightweight residual fully connected network to generate the original theoretical lower bound of prediction. The logarithm of the ratio of the original theoretical lower bound to the measured lower bound is used to train a lightweight residual fully connected network through a quantile loss function.
[0051] Specifically, the target quantile position of the quantile loss function is 0.9, the penalty weight for underestimation is 0.9, and the penalty weight for overestimation is 0.1. This forces the network to learn the upper bound estimate within the safety margin, which is in line with the improvement of avoiding OOM. The training objective is changed from predicting the mean of the residuals to predicting the conservative quantile of the residuals, which can proactively reserve a safety margin for fragment fluctuations.
[0052] like Figure 4 As shown, the process of generating the predicted lower bound of video memory further includes: The structural training parameters and theoretical lower limit of the GPU memory of the IBM large model training middleware software are passed through the first fully connected layer to generate the original GPU memory features. The original memory features are passed through a second fully connected layer to generate higher-order memory features; The higher-order memory features are passed through the output layer to generate a predicted lower bound for memory. The lightweight residual fully connected network includes a first fully connected layer, a second fully connected layer, and an output layer.
[0053] Specifically, the first fully connected layer uses the ReLU activation function and the Dropout probability is 0.1, the second fully connected layer uses the ReLU activation function and the Dropout probability is 0.1, and the output layer has no activation function.
[0054] Furthermore, the process of determining the recommended GPU configuration for operational support includes: If the physical video memory of the candidate GPU cluster configuration combination is greater than the predicted lower limit of video memory, then the candidate GPU cluster configuration combination will be used as the recommended GPU configuration for operation and maintenance assurance. If the physical video memory of the candidate GPU cluster configuration combination is less than or equal to the predicted lower limit of video memory, then the parallelism of the candidate GPU cluster configuration combination is increased, and the predicted lower limit of video memory is recalculated.
[0055] Specifically, for candidate combinations with insufficient physical video memory, attempts are made to salvage the situation by adjusting the parallel strategy parameters without changing the number or model of GPUs. This involves increasing tensor parallelism to distribute a single layer across more GPUs to reduce the activation value per GPU, or increasing pipeline parallelism to reduce the number of large model layers handled by a single GPU. For each set of parameters adjusted, the linked correction process is re-invoked to generate a new predicted lower limit for video memory. If a feasible strategy exists that lowers the new predicted lower limit for video memory below the physical video memory limit, then this combination, along with the adjusted strategy, is re-marked as the recommended GPU configuration for operational support, allowing it to enter the subsequent total cost of ownership assessment. Therefore, the above process prioritizes strategy combinations that minimize the loss in training throughput.
[0056] Furthermore, the structure training parameters include the total number of trainable parameters after normalization, the number of model layers, the hidden dimension, the number of attention heads, the vocabulary size, the maximum sequence length used for training, the forward computation precision, whether mixed precision is enabled, the optimizer type, the GPU micro-batch size, the single GPU physical memory capacity, and the number of configured nodes.
[0057] Specifically, the IBM large model training middleware software is IBM Watsonx.ai Training Middleware v2.3, and its IBM large model is the ViT-L / 14 model for image classification. The training configuration set is constructed based on mixed precision mode, parallel strategy, and micro-batch size. The following memory prediction methods are selected as comparison benchmarks: Comparison Method 1 calculates only the static memory usage of model weights, gradients, and optimizer states; Comparison Method 2 adds basic activation value calculation to the static memory usage; Comparison Method 3 uses the memory performance analysis tool provided by PyTorch Profiler; and Comparison Method 4 uses only a 3-layer fully connected network for memory prediction. The mean absolute error of memory prediction accuracy for Comparison Method 1, Comparison Method 2, Comparison Method 3, Comparison Method 4, and the methods described in this embodiment, in GB, are 18.72, 10.36, 7.85, 4.21, and 1.28, respectively. As can be seen, the average relative error of the method of this invention is only 2.17%, which is much lower than all the comparison methods and is 70.5% higher than the best performing comparison method 4, thus realizing the effectiveness of the intelligent monitoring and operation and maintenance guarantee method for the running status of IBM large model training middleware.
[0058] Therefore, the total number of trainable parameters, the number of model layers, the hidden dimension, the number of attention heads, and the vocabulary size collectively determine the shape, quantity, and distribution of the model tensor. Forward computation precision, whether mixed precision is enabled, and the optimizer type directly affect the byte coefficients of each component of static memory. The GPU micro-batch size and the number of configured nodes directly determine the size of the dynamic activation value. Larger micro-batches cause PyTorch's cache allocator to reserve larger contiguous memory blocks, thereby changing the fragmentation pattern. The maximum sequence length is the most important amplifier for dynamic activation values, causing the memory usage of the attention matrix to spike non-linearly. Thus, the above parameters cover all dimensions from model microstructure, training hyperparameters, parallel strategies to hardware physical boundaries, enabling the residual network to identify the source of memory bias in a given scenario. The network will activate specific neurons based on these feature combinations, outputting the typical fragmentation correction amount for that scenario.
[0059] This embodiment addresses the shortcomings of traditional theoretical calculation methods, such as the lack of dynamic activation value estimation and the failure to consider hardware environment differences. It adaptively learns the deviation between theoretical and measured values using a lightweight residual fully connected network, further reducing the average relative error of the final predicted lower limit of GPU memory. Based on high-precision GPU memory prediction results, it can recommend the optimal GPU configuration combination for each large model training task, avoiding the problems of insufficient GPU memory leading to training crashes and over-configuration causing wasted computing resources, common in empirical configurations. First, the activation value of a single micro-batch is calculated based on the model structure type and the number of activation tensor parameters. Then, the activation value per GPU is obtained based on the number of GPUs split. Finally, two key parameters, the number of computational layers per GPU and the average number of resident micro-batches, are introduced to calculate the final dynamic activation value for training. This fully considers the distribution characteristics of activation values and the resident characteristics of the computational pipeline in the tensor parallel implementation of IBM's large model training middleware, avoiding the severe underestimation of dynamic activation values caused by traditional methods that do not consider tensor parallel splitting and pipeline resident characteristics. This further avoids the problems of insufficient GPU memory leading to training crashes and over-configuration causing wasted computing resources, common in empirical configurations.
[0060] Those skilled in the art will recognize that the modules and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0061] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.
[0062] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for intelligent monitoring and maintenance assurance of IBM middleware software runtime status, characterized in that, include: The runtime parameters of IBM's large model training middleware software are converted to obtain static memory values; The structural parameters of the IBM large model training middleware software and the GPU card partitioning parameters are converted into tensor parallel activation values to calculate the training dynamic activation values, and the theoretical lower limit of video memory is calculated based on the static video memory values and the training dynamic activation values. The IBM large model training middleware software was used to obtain the measured lower limit of video memory through a cluster monitoring algorithm. After using the difference between the theoretical memory lower limit and the measured memory lower limit of the IBM large model training middleware software to train a lightweight residual fully connected network, the newly uploaded structural training parameters and theoretical memory lower limit of the IBM large model training middleware software are passed through the lightweight residual fully connected network to generate a predicted memory lower limit. The predicted lower bound of GPU memory based on the candidate GPU cluster configuration combinations is used to determine the recommended GPU configuration for operational support of IBM's large model training middleware software.
2. The method for intelligent monitoring and maintenance assurance of IBM middleware software runtime status according to claim 1, characterized in that, The process of training a lightweight residual fully connected network includes: The theoretical lower bound of video memory and the structural training parameters of IBM's large model training middleware software are passed through a lightweight residual fully connected network to generate the original theoretical lower bound of prediction. The logarithm of the ratio of the original theoretical lower bound to the measured lower bound is used to train a lightweight residual fully connected network through a quantile loss function.
3. The method for intelligent monitoring and maintenance assurance of IBM middleware software runtime status according to claim 1, characterized in that, The process of generating the predicted lower bound of video memory includes: The structural training parameters and theoretical lower limit of the GPU memory of the IBM large model training middleware software are passed through the first fully connected layer to generate the original GPU memory features. The original memory features are passed through a second fully connected layer to generate higher-order memory features; The higher-order memory features are passed through the output layer to generate a predicted lower bound for memory. The lightweight residual fully connected network includes a first fully connected layer, a second fully connected layer, and an output layer.
4. The method for intelligent monitoring and maintenance assurance of IBM middleware software runtime status according to claim 1, characterized in that, The process of obtaining static video memory values includes: The sum of the corresponding byte amounts calculated according to the running parameter categories is then multiplied by the total amount of the running parameters to obtain the static video memory value. The operating parameters include operating parameter categories.
5. The method for intelligent monitoring and maintenance assurance of IBM middleware software runtime status according to claim 4, characterized in that, The categories of operating parameters include the training accuracy type of the forward parameters, whether mixed accuracy is enabled for the master parameter replicas, the forward accuracy of the gradients, and the optimizer type.
6. The method for intelligent monitoring and maintenance assurance of IBM middleware software runtime status according to claim 1, characterized in that, The process of calculating training dynamic activation values includes: Based on the structure type and activation tensor parameter quantity of the IBM large model training middleware software, the single micro-batch activation value is calculated. Based on the single micro-batch activation value and the number of GPU cards to be divided, the single card activation value is obtained; The training dynamic activation value is obtained by multiplying the number of single-card calculation layers, the average number of micro-batches in residence, and the single-card activation value. The GPU card segmentation parameters include the number of GPU cards to be segmented, the number of computing layers per card, and the average number of resident micro-batches.
7. The method for intelligent monitoring and maintenance assurance of IBM middleware software runtime status according to claim 1, characterized in that, The process of obtaining the measured lower limit of video memory includes: The IBM large model training middleware software uses a training framework callback algorithm to obtain multiple memory distribution values over multiple iterations, and calculates the steady-state memory value based on these multiple memory distribution values. The IBM large model training middleware software is monitored using a cluster monitoring algorithm to obtain the net increase in video memory usage. The measured lower limit of video memory is calculated by summing the steady-state video memory value and the video memory usage.
8. The method for intelligent monitoring and maintenance assurance of IBM middleware software runtime status according to claim 7, characterized in that, The process of calculating steady-state memory values includes: The steady-state memory value is obtained by subtracting the memory distribution value corresponding to the set quantile from the peak value of the memory distribution value within the steady-state window.
9. The method for intelligent monitoring and maintenance assurance of IBM middleware software runtime status according to claim 7, characterized in that, The process of determining the recommended GPU configuration for operational support includes: If the physical video memory of the candidate GPU cluster configuration combination is greater than the predicted lower limit of video memory, then the candidate GPU cluster configuration combination will be used as the recommended GPU configuration for operation and maintenance assurance. If the physical video memory of the candidate GPU cluster configuration combination is less than or equal to the predicted lower limit of video memory, then the parallelism of the candidate GPU cluster configuration combination is increased, and the predicted lower limit of video memory is recalculated.
10. The method for intelligent monitoring and maintenance assurance of IBM middleware software runtime status according to any one of claims 1 to 9, characterized in that, The structure training parameters include the total number of trainable parameters after normalization, the number of model layers, the hidden dimension, the number of attention heads, the vocabulary size, the maximum sequence length used for training, the forward computation precision, whether mixed precision is enabled, the optimizer type, the GPU micro-batch size, the single GPU physical memory capacity, and the number of configured nodes.