Anomaly management method for artificial intelligence (AI) model training, and related system
Patent Information
- Application Number
- PCT/CN2025/132493
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-28
- Filing Date
- 2025-11-04
- Publication Date
- 2026-10-01
Smart Images

Figure CN2025132493_01102026_PF_FP_ABST
Abstract
Description
An anomaly management method and related system for AI model training
[0001] This application claims priority to Chinese Patent Application No. 202510392402.9, filed with the State Intellectual Property Office of China on March 28, 2025, entitled "An Anomaly Management Method and Related System for Artificial Intelligence AI Model Training", the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of artificial intelligence (AI) technology, and in particular to an anomaly management method for AI model training, an anomaly management system, a computer-readable storage medium, and a computer program product. Background Technology
[0003] With the continuous development of artificial intelligence (AI) technology, the parameter scale of AI models is constantly increasing. For example, the parameter scale of large language models (LLMs), which are widely used in various fields, is usually in the hundreds of billions or trillions. Large-scale AI models are more prone to anomalies such as gradient spikes, loss spikes, loss steps, and loss runaways during the training process. These anomalies are also known as training instability.
[0004] During the training of AI models, training instability such as gradient spikes, loss spikes, loss steps, and loss runaway is a type of anomaly that cannot be ignored and urgently needs to be addressed. How to manage these anomalies has become an urgent problem to solve. Summary of the Invention
[0005] This application provides an anomaly management method for AI model training. This method introduces interpretable, lightweight, and observable feature metrics, such as attention entropy, for anomaly detection across various AI model training scenarios. It can cover a wide range of unstable training scenarios, has low detection costs, and can even achieve non-destructive detection, exhibiting high availability. This application also provides an anomaly management system, computing device cluster, computer-readable storage medium, and computer program products corresponding to the above method.
[0006] Firstly, this application provides an anomaly management method for AI model training. This method can be applied to an anomaly management system. The anomaly management system may include a software system, which can be an independent software system or dependent on other software systems. For example, the anomaly management system may include a monitoring tool (monitor) for detecting training anomalies such as loss spikes and loss divergence. The monitoring tool may include plug-in monitoring code, which can be imported into the training script to detect training anomalies such as training instability during AI model operation, or to perform anomaly localization and recovery. In some possible implementations, the anomaly management system may also include a hardware system, such as a computing device cluster with anomaly detection, anomaly localization, or anomaly recovery capabilities. The computing device cluster executes the AI model training anomaly management method of this application during runtime.
[0007] The training of the AI model is performed by at least one computing card. This at least one computing card includes a first computing card. Specifically, during the training process of the AI model, the anomaly management system can identify the first module currently being executed by the first computing card. This first module includes one or more modules within the AI model. When the first module's module type is a target type, such as an attention module, the anomaly management system can obtain observable feature indicators when the first computing card executes the first module. These observable feature indicators include attention entropy. Attention entropy, also simply called attention entropy, represents the degree of clustering of correlations in the input data of the first module. The anomaly management system can then perform anomaly detection based on these observable feature indicators to obtain anomaly detection results for the AI model's training process.
[0008] This method introduces interpretable, lightweight, and observable feature metrics for various AI model training scenarios. For example, it incorporates attention entropy for anomaly detection, covering a wide range of unstable training scenarios with low detection costs and even lossless detection capabilities, resulting in high availability. Furthermore, this method enables online observation of the AI model's loss function convergence, reducing the time cost of retraining due to inappropriate hyperparameter settings (such as interruption or rollback time).
[0009] In some possible implementations, the anomaly management system can input observable feature indicators into a predictive model for anomaly detection to obtain anomaly detection results. The predictive model is obtained by training a logistic regression model using training data, which includes historical observable feature indicators when the loss function converges and when the loss function fails to converge during AI model training.
[0010] This method constructs training data by using historical observable feature indicators when the loss function converges and when the loss function fails to converge during AI model training. The prediction model constructed based on this training data can not only predict the current training state (whether training instability or other anomalies occur), but also predict the future training state, that is, predict whether training anomalies will occur in the future (in the next training step). This enables early detection of training instability and other anomalies, timely warning, and timely intervention to avoid AI model performance degradation.
[0011] In some possible implementations, the anomaly management system can also match observable feature indicators with anomaly detection rules to obtain anomaly detection results. These anomaly detection rules can be experience-based rules used to detect training anomalies such as training instability.
[0012] This method simplifies the detection process and improves efficiency by using rule matching for anomaly detection. Furthermore, combining prediction models with anomaly detection rules can achieve a balance between detection accuracy and efficiency.
[0013] In some possible implementations, the first module includes the attention module of the target layer in the AI model. Accordingly, when the attention entropy of the target layer's attention module first decays below a first threshold and then recovers to above the first threshold within Q training steps, the anomaly management system can obtain an anomaly detection result. The anomaly detection result indicates that a short-term anomaly occurred during the training process of the AI model and has since returned to normal, where Q is a positive integer.
[0014] In this method, the anomaly detection rule may include a single threshold judgment rule. By comparing the attention entropy of the attention module of the target layer (such as the attention entropy of the shallow layer) with the first threshold, it is possible to quickly detect whether training anomalies have occurred during the training process of the AI model, thereby improving the efficiency of anomaly detection.
[0015] In some possible implementations, the first module includes the attention module of the target layer in the AI model. When the attention entropy of the target layer's attention module is greater than a first threshold, and the difference between the attention entropy of the target layer's attention module and the historical attention entropy of the previous training step is greater than a second threshold, the anomaly management system can obtain an anomaly detection result. The anomaly detection result indicates that an anomaly has occurred during the training process of the AI model.
[0016] In this method, the anomaly detection rules can include multi-threshold judgment rules. By comparing the attention entropy of the target layer's attention module (such as the attention entropy of a shallow layer) with a first threshold, and comparing the difference in attention entropy between different training steps with a second threshold, training anomalies can be quickly detected during the AI model's training process, improving anomaly detection efficiency. Furthermore, the multi-threshold judgment mechanism can improve detection accuracy.
[0017] In some possible implementations, the first module includes the attention module of the target layer in the AI model. When the attention entropy of the target layer's attention module decays below a first threshold, and the difference between the attention entropy of the target layer's attention module and the historical attention entropy of the previous training step is greater than a second threshold, the anomaly management system can obtain an anomaly detection result. The anomaly detection result indicates that an anomaly has occurred in the training process of the AI model.
[0018] In this method, the anomaly detection rules can include multi-threshold judgment rules. By comparing the attention entropy of the target layer's attention module (such as the attention entropy of a shallow layer) with a first threshold, and comparing the difference in attention entropy between different training steps with a second threshold, it is possible to quickly detect whether training anomalies will occur in the future during the training of the AI model, thereby improving anomaly detection efficiency. Furthermore, the multi-threshold judgment mechanism can improve detection accuracy.
[0019] In some possible implementations, when anomaly detection results indicate that an anomaly has occurred or is about to occur during the training process of an AI model, the anomaly management system can determine the cause of the anomaly based on observable feature indicators. Specifically, the anomaly management system can determine the cause of the anomaly based on observable feature indicators and anomaly detection rules.
[0020] This method can accelerate the localization of training anomalies caused by hyperparameter setting anomalies based on anomaly detection rules related to hyperparameter settings. For example, it can reduce the time from 3 days for manual investigation to seconds, thus reducing the average repair time.
[0021] In some possible implementations, when the cause of the anomaly includes abnormal hyperparameter settings, the anomaly management system can also modify the AI model's parameters to increase attention entropy. Compared to recovery methods such as modifying hyperparameters or rolling back training, adaptive entropy-increasing restorative training can achieve almost no interruption, with minimal impact on training, low cost, and high availability. Furthermore, the method of adaptively updating model parameters to restore normal model convergence can serve as a model fault tolerance method, ensuring minimal performance degradation of the AI model without affecting the final convergence of the AI model's loss function.
[0022] In some possible implementations, the anomaly management system can perform a smoothing operation on the linear matrix parameters of the AI model, thereby achieving adaptive entropy increase. This allows for training recovery at a relatively low cost, meeting business requirements.
[0023] In some possible implementations, the linear matrix parameters consist of a vector of all singular values of the weights in the parameter matrix, arranged in descending order. Accordingly, the anomaly management system can compress the first to Mth singular values in the vector such that the compressed first to Mth singular values are equal to the (M+1)th singular value. Here, M is the floor function of the stable rank norm, which is used to evaluate the convergence of the AI model's loss function. Alternatively, the anomaly management system can use convolution or activation functions to perform smoothing operations on the linear matrix parameters of the AI model.
[0024] This method supports compression of singular values or smoothing of linear matrix parameters of AI models through various methods such as convolution and activation functions, and has high flexibility.
[0025] Secondly, this application provides an anomaly management system. The anomaly management system is used for anomaly management during the training process of an artificial intelligence (AI) model. The training of the AI model is performed by at least one computing card, and the at least one computing card includes a first computing card. The anomaly management system includes:
[0026] The feature monitoring module is used to identify the first module currently being executed by the first computing card during the training process of the AI model. The first module includes one or more modules in the AI model. When the module type of the first module is a target type, the module obtains observable feature indicators when the first computing card executes the first module. The observable feature indicators include attention entropy, which is used to represent the degree of clustering of correlations in the input data of the first module.
[0027] An anomaly detection module is used to perform anomaly detection based on the observable feature indicators and obtain the anomaly detection results of the training process of the AI model.
[0028] In some possible implementations, the anomaly detection module is specifically used for:
[0029] The observable feature indicators are input into the prediction model for anomaly detection to obtain anomaly detection results. The prediction model is obtained by training a logistic regression model using training data. The training data includes historical observable feature indicators when the loss function converges and historical observable feature indicators when the loss function does not converge during the training of the AI model.
[0030] In some possible implementations, the anomaly detection module is specifically used for:
[0031] The observable feature indicators are matched with the anomaly detection rules to obtain the anomaly detection results.
[0032] In some possible implementations, the first module includes an attention module for the target layer in the AI model, and the anomaly detection module is specifically used for:
[0033] When the attention entropy of the attention module of the target layer first decays to below the first threshold, and then recovers to above the first threshold within Q training steps, an anomaly detection result is obtained. The anomaly detection result indicates that the training process of the AI model has experienced a short-term anomaly and has returned to normal. Q is a positive integer.
[0034] In some possible implementations, the first module includes an attention module for the target layer in the AI model, and the anomaly detection module is specifically used for:
[0035] When the attention entropy of the attention module of the target layer is greater than a first threshold, and the difference between the attention entropy of the attention module of the target layer and the historical attention entropy of the previous training step is greater than a second threshold, an anomaly detection result is obtained, and the anomaly detection result indicates that an anomaly will occur in the training process of the AI model.
[0036] In some possible implementations, the first module includes an attention module for the target layer in the AI model, and the anomaly detection module is specifically used for:
[0037] When the attention entropy of the attention module of the target layer decays to below the first threshold, and the difference between the attention entropy of the attention module of the target layer and the historical attention entropy of the previous training step is greater than the second threshold, an anomaly detection result is obtained, and the anomaly detection result indicates that an anomaly has occurred in the training process of the AI model.
[0038] In some possible implementations, the system further includes:
[0039] An anomaly localization module is used to determine the cause of the anomaly based on the observable feature indicators when the anomaly detection result indicates that an anomaly has occurred or will occur during the training process of the AI model.
[0040] In some possible implementations, the system further includes:
[0041] An anomaly recovery module is used to modify the parameters of the AI model to improve the attention entropy when the cause of the anomaly includes an abnormal hyperparameter setting.
[0042] In some possible implementations, the exception recovery module is specifically used for:
[0043] Smoothing operation is performed on the linear matrix parameters of the AI model.
[0044] In some possible implementations, the linear matrix parameters include a vector of all singular values of the weights in the parameter matrix arranged in descending order, and the anomaly recovery module is specifically used for:
[0045] The first to the Mth singular values in the vector are compressed such that the compressed first to the Mth singular values are equal to the (M+1)th singular value, where M is the floor function of the stable rank norm, which is used to evaluate the convergence of the loss function of the AI model; or,
[0046] The linear matrix parameters of the AI model are smoothed using convolution or activation functions.
[0047] Thirdly, this application provides a computing device cluster. The computing device cluster includes at least one computing device, and the at least one computing device includes at least one processor and at least one memory. The at least one processor and the at least one memory communicate with each other. The at least one processor is used to execute instructions stored in the at least one memory to cause the computing device or the computing device cluster to perform the exception management method for AI model training as described in the first aspect or any implementation thereof.
[0048] Fourthly, this application provides a computer-readable storage medium storing instructions that instruct a computing device or a cluster of computing devices to execute the exception management method for AI model training described in the first aspect or any implementation thereof.
[0049] Fifthly, this application provides a computer program product containing instructions that, when run on a computing device or a cluster of computing devices, causes the computing device or cluster of computing devices to execute the exception management method for AI model training described in the first aspect or any implementation thereof.
[0050] Based on the implementation methods provided in the above aspects, this application can be further combined to provide more implementation methods. Attached Figure Description
[0051] To more clearly illustrate the technical methods of this application, the accompanying drawings used will be briefly described below.
[0052] Figure 1 is a schematic diagram of the architecture of an exception management system provided in this application;
[0053] Figure 2 is a schematic diagram of the training path of an AI model provided in this application;
[0054] Figure 3 is a flowchart of an anomaly management method for AI model training provided in this application;
[0055] Figure 4 is a schematic diagram of an AI model of an encoder-decoder architecture provided in this application;
[0056] Figure 5 is a schematic diagram of an AI model with a hybrid expert model architecture provided in this application;
[0057] Figure 6 is a schematic diagram of the structure of a computing device provided in this application;
[0058] Figure 7 is a schematic diagram of the structure of a computing device cluster provided in this application;
[0059] Figure 8 is a schematic diagram of another computing device cluster provided in this application;
[0060] Figure 9 is a schematic diagram of another computing device cluster provided in this application. Detailed Implementation
[0061] The terms "first" and "second" used in the embodiments of this application are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature.
[0062] First, some technical terms involved in the embodiments of this application will be introduced.
[0063] Artificial intelligence (AI), also known as machine intelligence, specifically refers to the ability of machines to correctly interpret external data, learn knowledge from the data, and use that knowledge to achieve specific goals and tasks through flexible adaptation.
[0064] AI models are mathematical models with reasoning capabilities built using techniques such as machine learning (ML) or deep learning (DL). AI models can be categorized into large models and small models based on their parameter scale. Large models, such as large language models (LLMs), can have parameter scales (e.g., the number of parameters) reaching hundreds of billions or even trillions.
[0065] As the parameter scale of AI models continues to increase, the training process becomes more prone to anomalies. These anomalies can include training instability. Training instability refers to the behavior of the loss function deviating from normal convergence during AI model training, including but not limited to gradient spikes, loss spikes, loss divergence, loss steps, and loss runaway.
[0066] Among these, gradient spikes in the loss function refer to sudden, brief, and large-amplitude abnormal fluctuations during gradient calculation, typically manifested as sharp increases or decreases in the gradient value. Loss spikes occur during AI model training when the loss function value suddenly and briefly surges, followed by a rapid decline, appearing as sharp peaks in the loss curve. Loss divergence refers to the loss curve failing to converge during AI model training, specifically manifested as a sudden, sharp rise in the loss function value during a decline. Loss step refers to a sudden, large, and sustained increase or decrease in the loss function value during AI model training, exhibiting a step-like change. Loss runaway, similar to loss divergence, refers to the loss function value continuously and rapidly increasing during training without any convergence trend, resulting in a sharp deterioration in the AI model's performance.
[0067] While most training instabilities can be recovered due to the robustness of model optimization algorithms, these potential instabilities can still severely degrade the performance of AI models. Therefore, both industry and academia have conducted extensive research on training instability. Related studies have shown that loss spikes are accompanied by a bimodal distribution of optimizer update amounts, and that these loss spikes can be avoided through gradient clipping operations.
[0068] Currently, the industry offers an anomaly detection scheme for AI model training. This scheme collects statistics on gradients, activation values, and optimizer states during the AI model training process. The optimizer state can include first-order and second-order gradient terms. If the distribution of the optimizer state suddenly changes from a Gaussian distribution with a mean of 0 to a bimodal distribution with mass concentrated between +1 and -1, it is determined that the hyperparameters such as β2 and ε in the AdamW optimization algorithm under the current state are improperly set, causing the update quantity to evolve into a bimodal distribution, which in turn leads to a large update in the current training step, resulting in loss spikes.
[0069] The basis for this judgment method is that: in the later stage of training of the AI model, the magnitude of the update volume of the front and back layers is disconnected, the shallow layer updates are very small while the deep layer still maintains a large update, the gradient norm of the AI model is significantly smaller than ε in the optimizer, so that the overall update volume distribution presents a Gaussian distribution with a mean of 0; when the data of the front and back batches are correlated at a certain moment, the gradient update is amplified and exceeds ε, so that the update volume instantly presents a bimodal distribution concentrated between +1 and -1.
[0070] However, the above-mentioned solutions have limited application scenarios. Some loss spikes cannot be detected, and other training instabilities such as loss divergence, loss step, and loss runaway are also difficult to detect based on changes in the optimizer state distribution.
[0071] The inventors discovered that training instability caused by unreasonable hyperparameters at the AI model algorithm level is reflected in certain feature indicators during the AI model training process. This is because changes in the AI model's optimization trajectory due to hyperparameters cause gradual changes in certain intrinsic feature indicators of the AI model. Therefore, by understanding the training mechanism and mining certain lightweight feature indicators during the AI model training process, it is possible to both monitor the health status of the AI model online and locate hyperparameter settings that are mismatched with the current state of the AI model. Furthermore, by formulating reasonable and clear anomaly detection rules (such as threshold judgment rules), the online monitored feature indicators are compared with these rules. If the threshold is exceeded, a training anomaly alarm is issued, and the system supports locating hyperparameter anomalies and providing anomaly recovery suggestions.
[0072] In view of this, this application provides an anomaly management method for AI model training. This method can be executed by an anomaly management system. The anomaly management system may include a software system, which can be a standalone software system or dependent on other software systems. For example, the anomaly management system may include a monitoring tool (monitor) for detecting training anomalies such as loss spikes and loss divergence. The monitoring tool may include plug-in monitoring code, which can be imported into the training script to detect training anomalies such as training instability during AI model operation, or to perform anomaly localization and recovery. In some possible implementations, the anomaly management system may also include a hardware system, such as a cluster of computing devices with anomaly detection, anomaly localization, or anomaly recovery capabilities. The computing device cluster executes the anomaly management method for AI model training of this application during runtime.
[0073] Specifically, the training of the AI model is performed by at least one computing card. A computing card refers to a processor that provides computing power, including but not limited to a central processing unit (CPU), neural processing unit (NPU), graphics processing unit (GPU), tensor processing unit (TPU), and other XPUs. At least one computing card includes a first computing card. During the training of the AI model, the anomaly management system can identify the first module currently being executed by the first computing card. The first module includes one or more modules within the AI model. When the module type of the first module is a target type, such as the attention module in a flash attention (FA) structure, the anomaly management system obtains observable feature indicators when the first computing card executes the first module. These observable feature indicators include attention entropy, which represents the degree of clustering of correlations in the input data of the first module. The anomaly management system can then perform anomaly detection based on the observable feature indicators to obtain anomaly detection results for the AI model's training process.
[0074] This method introduces interpretable, lightweight, and observable feature metrics for various AI model training scenarios. For example, it incorporates attention entropy for anomaly detection, covering a wide range of unstable training scenarios with low detection costs and even lossless detection capabilities, resulting in high availability. Furthermore, this method allows for online observation of the AI model's loss function convergence, reducing the time cost of retraining due to inappropriate hyperparameter settings.
[0075] To make the technical solution of this application clearer and easier to understand, the system architecture of the anomaly management system of this application is described below with reference to the accompanying drawings.
[0076] Referring to Figure 1, which illustrates the architecture of an exception management system, Figure 1 uses the exception management system 100 as an example of a software system. The exception management system 100 can be a layer in the software stack of an AI model. The AI model's software stack, from top to bottom, includes training scripts 10, a framework layer 20, and a heterogeneous computing framework 30. The exception management system 100 can be software within the framework layer 20. A detailed explanation follows.
[0077] Training script 10 includes a piece of program code for training the AI model. It is typically written in a programming language (such as Python) and implemented using a corresponding machine learning framework (such as PyTorch, TensorFlow, etc.). Training script 10 guides computing devices or clusters of computing devices through the entire process of AI model training, from initialization to final completion. This process mainly includes the following steps: data preparation, AI model definition, loss function and optimizer selection, AI model training, AI model evaluation, and saving.
[0078] Framework layer 20 includes large model frameworks and deep learning frameworks provided by third-party libraries. Large model frameworks refer to software tools and libraries used for training, inference, and deploying large models (such as LLMs). These frameworks offer efficient computational resource management, distributed training, model optimization, and inference acceleration to better utilize hardware resources for handling massive datasets and complex model structures. Deep learning frameworks are software libraries or toolkits specifically designed for deep learning tasks. They provide a series of high-level abstractions and tools to help developers build, train, and deploy deep learning models more efficiently. Specifically, deep learning frameworks offer a rich set of predefined layers and components, such as convolutional layers, fully connected layers, and recurrent layers. Developers can use these predefined layers or components to quickly build AI models with various complex model structures.
[0079] The heterogeneous computing framework 30 is a software architecture that supports the collaborative work of multiple different types of computing cards, aiming to efficiently utilize hybrid hardware resources (such as CPUs, GPUs, TPUs, or NPUs) for parallel computing. The heterogeneous computing framework 30 includes a hardware abstraction layer, which provides a unified interface to encapsulate the characteristics of different hardware, shielding them from underlying differences. In distributed computing scenarios, the heterogeneous computing framework 30 can be a Compute Architecture for Neural Net (CANN) for neural networks. CANN can support users in quickly building AI applications by providing multi-layered programming interfaces. Here, AI applications refer to applications built based on AI models (such as LLMs). It should be noted that the heterogeneous computing framework 30 is an optional framework; the exception management method for AI model training in this embodiment of the application may not require calling the aforementioned heterogeneous computing framework.
[0080] In the AI model training scenario, training script 10 can call the large model framework provided by the third-party library in framework layer 20, and then call the deep learning framework (such as MindSpore) to train the AI model. The anomaly management system 100 in framework layer 20 can monitor observable feature indicators that reflect anomalies such as training instability. For example, it can use the hook tool of a deep learning framework (such as PyTorch) to collect the output data or parameters of the target operator, thereby calculating the corresponding statistical features and obtaining observable feature indicators. In the example in Figure 1, the target operator can be the attention operator (also called the attention module) in the FA structure. Observable feature indicators can include attention entropy. Attention entropy is used to represent the degree of clustering of the correlations of the input data of the attention module. The output data of the attention module can include the correlations of the input data, and this output data can be represented in matrix form, also called the attention matrix. Attention entropy can include the entropy of the attention matrix. For ease of understanding, let's use a text model or multimodal model as an example to illustrate the AI model. The input data of the AI model includes text data, which is processed and converted into input tokens. The output data of the attention module can include the relevance of the input tokens, and the attention entropy can represent the degree of clustering of the relevance of the input tokens in the attention module. In practical applications, the target operator can also include linear operators (also called linear modules or linear layers) in a feed-forward network (FFN), and the observable feature index can also include the stable rank (SR) norm. The stable rank norm is used to evaluate the convergence of the loss function of the AI model.
[0081] The anomaly management system 100 may include a monitoring tool (monitor). Users can import the monitor tool code (e.g., external monitoring code) into the training script, instantiate the monitor class, and enable monitoring to achieve online monitoring during the training process.
[0082] The anomaly management system 100 is described below from a modular perspective. The anomaly management system 100 may include a feature monitoring module 102 and an anomaly perception module 104.
[0083] The feature monitoring module 102 is used to identify the first module currently being executed by the first computing card during the training process of the AI model. The first module includes one or more modules in the AI model. When the module type of the first module is a target type, such as the first module being the aforementioned target operator, the observable feature indicators when the first computing card executes the first module are obtained. The observable feature indicators include attention entropy.
[0084] The anomaly detection module 104 is used to perform anomaly detection based on observable feature indicators to obtain anomaly detection results for the AI model's training process. These anomaly detection results indicate whether anomalies have occurred or are about to occur during the AI model's training process. Anomalies about to occur refer to situations where no anomalies have occurred currently, but an anomaly will occur in the future (e.g., the next training step). The anomaly detection module 104 can achieve anomaly detection through a prediction model or anomaly detection rules.
[0085] In some possible implementations, the anomaly detection module 104 is used to input observable feature indicators into the prediction model for anomaly detection and obtain anomaly detection results. The prediction model can not only predict whether training anomalies such as training instability occur during the training process of the AI model, but also predict whether training anomalies such as training instability will occur in the future, for example, predicting whether training anomalies will occur in the future. Furthermore, the anomaly detection module 104 can also return the probability of training anomalies (i.e., the probability of training failure) or the probability of training anomalies occurring through logs. The prediction model can be obtained by training a logistic regression model using training data. The training data includes historical observable feature indicators when the loss function converges and historical observable feature indicators when the loss function does not converge during the AI model training process.
[0086] In some possible implementations, the anomaly detection module 104 is used to match observable feature indicators with anomaly detection rules to obtain anomaly detection results. The anomaly detection rules can be threshold determination rules, which may include built-in rules or user-defined rules. The anomaly detection module 104 can compare observable feature indicators with thresholds in the threshold determination rules to determine whether an anomaly has occurred or is about to occur.
[0087] Furthermore, the anomaly management system 100 may also include an anomaly localization module 106 or an anomaly recovery module 108. The anomaly localization module 106 is used to determine the cause of the anomaly based on observable feature indicators when the anomaly detection result indicates that an anomaly has occurred or will occur during the training process of the AI model. It should be noted that the anomaly localization module 106 can determine the cause of the anomaly based on observable feature indicators and anomaly detection rules. The cause of the anomaly can be the root cause of the anomaly. The anomaly recovery module 108 is used to modify the parameters of the AI model to increase attention entropy when the cause of the anomaly includes anomaly in hyperparameter settings. Specifically, the anomaly recovery module 108 performs a smoothing operation on the linear matrix parameters of the AI model to adaptively increase entropy, thereby adaptively recovering the training. The above-mentioned method of recovering training by modifying the parameters of the AI model to increase attention entropy is also called the adaptive entropy recovery method.
[0088] It should be noted that the anomaly management system 100 provides two training paths for AI model training: a normal training path and an abnormal training path. Referring to Figure 2, which illustrates one training path for an AI model, after the monitoring tool is activated during AI model training, the feature monitoring module 102 can continuously collect observable feature indicators when the first computing card executes the first module. The AI model's training process may enter the normal training path (shown by the solid line in Figure 2), i.e., the path where the AI model's loss function converges normally (e.g., stably), or it may enter the abnormal training path for various reasons (shown by the dashed line in Figure 2), exhibiting an unstable loss curve. In this case, the anomaly perception module 104 can perform anomaly detection based on the observable feature indicators, for example, by inputting the observable feature indicators into the prediction model, thereby outputting the anomaly detection result of the AI model's training process. The anomaly detection result is used to indicate whether an anomaly has occurred or will occur during the AI model's training process. In some examples, the anomaly detection result includes the probability that the AI model's training process will experience loss divergence. This can alert users (such as algorithm engineers) to the risk of the AI model subsequently exhibiting unstable loss behavior. When the anomaly detection result indicates that no training anomaly has occurred (no training anomaly detected), the AI model can continue normal training. When the anomaly detection result indicates that a training anomaly has occurred, the anomaly localization module 106 can determine the cause of the anomaly according to the anomaly detection rules. In this example, the cause of the anomaly could be an anomaly in hyperparameter settings. The anomaly recovery module 108 is used to provide recovery suggestions to the user based on the cause of the anomaly. These recovery suggestions include rolling back training or adjusting hyperparameters. Alternatively, the recovery suggestions could also include non-disruptive adaptive recovery. When the user selects non-disruptive adaptive recovery, the anomaly recovery module 108 can perform a smoothing operation on the linear matrix parameters of the AI model to recover training through adaptive entropy increase.
[0089] Based on the exception management system 100 shown in Figure 1, this application also provides an exception management method for AI model training. The exception management method for AI model training provided in this application will be described in detail below with reference to the accompanying drawings.
[0090] Referring to Figure 3, a flowchart of an anomaly management method for AI model training is shown. The AI model is trained using at least one computing card, which includes a first computing card. During the training of the AI model, the following steps can be performed:
[0091] S302, The anomaly management system 100 identifies the first module currently being executed on the first computing card. If the module type of the first module is target type, S304 is executed.
[0092] The first module comprises one or more modules within the AI model. These modules can be categorized by function; for example, they may include attention modules, linear modules within a feed-forward network (FFN), or expert modules. Expert modules can break down a large model into multiple independent sub-networks (experts), each specializing in a specific task or data pattern. For instance, text processing experts handle tasks like language understanding and generation, image processing experts process visual information from multimodal inputs, and domain experts handle knowledge reasoning in fields such as finance and healthcare.
[0093] The anomaly management system 100 can identify at least one computing card used to train the AI model. Taking one of these computing cards as an example of a first computing card, the anomaly management system 100 can identify the first module currently being executed on the first computing card by marking points and identify the module type of the first module. When the module type of the first module includes an attention module, S304 can be executed to collect observable feature indicators. Alternatively, when the module type of the first module includes a linear module or an expert module in a feedforward neural network, the anomaly management system 100 can also execute S304 to collect observable feature indicators. It should be noted that the specific implementation of anomaly detection for other computing cards used to train the AI model can be found in the relevant description of the first computing card, and will not be repeated here.
[0094] Furthermore, the anomaly management system 100 can also provide a configuration interface, which provides configuration options for enabling or disabling the online monitoring function. The anomaly management system 100 can receive configuration options from users through the configuration interface. When the configuration indicates that the online monitoring function is enabled, the anomaly management system 100 can identify the first module currently being executed by the first computing card during the AI model's operation. It should be noted that the anomaly management system 100 also supports customizable enabling and intermittent data collection times, significantly reducing computational overhead and minimizing performance loss.
[0095] S304, The anomaly management system 100 obtains observable characteristic indicators when the first computing card executes the first module.
[0096] Observable feature metrics are features that can be observed during AI model training and can reflect whether the AI model has experienced training instability or other training anomalies. This application understands the essence and phenomenon of AI model training instability from the perspective of the convergence mechanism of the AI model's loss function, and introduces observable feature metrics such as Attention Entropy. Among them, Attention Entropy is an indicator that measures the "degree of clustering" of the distribution of attention weights in the attention mechanism. Taking the first module (such as the attention module) as an example, Attention Entropy is used to represent the degree of clustering of the correlation of the input data of the first module. The correlation of the input data of the first module can be represented by the attention matrix output by the first module, which includes several attention weights representing the correlation of data.
[0097] In some possible implementations, the first module may also include the linear module and expert module in the FFN, and the observable metric may also include the stable rank norm. The stable rank norm is a metric for measuring the low rank of a matrix. In data analysis and machine learning, the stable rank norm can be used to evaluate the convergence of the loss function of an AI model. Theoretically, attention entropy collapse and stable rank norm collapse correspond to unreasonable settings of the AI model's hyperparameters (such as learning rate and optimizer parameters), indicating that the current AI model has entered an extremely sharp optimization surface, with a high probability of training anomalies such as model instability.
[0098] The following section introduces how attention entropy or stable rank norm is calculated.
[0099] First, we will use an AI model with an encoder-decoder architecture to illustrate how attention entropy is calculated.
[0100] Referring to Figure 4, which illustrates an AI model with an encoder-decoder architecture, the AI model includes an encoder 402 and a decoder 404. In this example, the input sequence is vectorized to obtain input vectors (embeddings). These input vectors can also be combined with positional encoding, and the positionally encoded input vector is then input to the encoder 402 for encoding. The encoder includes a multi-head attention module (also referred to as multi-head attention module 1 for clarity). Multi-head attention module 1 is used to calculate the correlation of features, where correlation can be represented by attention. The output of multi-head attention module 1 is residually connected to the positionally encoded input vector through a residual connection and layer normalization (Add and Normalization) module (also referred to as residual connection and layer normalization module 1 for clarity), followed by layer normalization, specifically normalizing the feature dimensions of a single sample. The aforementioned residual connection and layer normalization module 1 can be connected to a multilayer perceptron (MLP). For ease of distinction, the MLP in the encoder is referred to as MLP 1. MLP 1 can then be connected to another residual connection and layer normalization module, specifically residual connection and layer normalization module 2. The input and output of MLP 1 can be residually connected and layer normalized through the aforementioned residual connection and layer normalization module 2 to obtain the final encoding result.
[0101] The encoded result can be input into decoder 404 for decoding. In the example of Figure 4, decoder 404 may include a masked multi-head attention module, an attention module (also called attention module 2 for easy distinction from the attention module in the encoder), a multilayer perceptron, and multiple residual connection and layer normalization modules. Similar to the input sequence, the target sequence can be vectorized to obtain a target vector. The target vector can also be combined with positional encoding. The target vector combined with positional encoding can be input into the masked multi-head attention module for attention calculation. The input and output of the masked multi-head attention module can also be residually connected through residual connection and layer normalization module 3, and then layer normalized. The normalization result can be input into multi-head attention module 2 along with the encoded result. The output of multi-head attention module 2 in the decoder can be input into the next residual connection and layer normalization module along with the output of residual connection and layer normalization module 3, for example, into residual connection and layer normalization module 4 for residual connection and layer normalization. The residual connection and layer normalization module 4 is connected to a multilayer perceptron 2. The multilayer perceptron 2 automatically learns the complex relationships between input features and predicts new data to achieve decoding. The output of the multilayer perceptron 2, along with the output of the residual connection and layer normalization module 4, can be input to the residual connection and layer normalization module 5 for residual connection and layer normalization. The output of the residual connection and layer normalization module 5 can be linearized by a linear module and classified by softmax to output the final decoding result.
[0102] For the multi-head attention module in Figure 4, the anomaly management system 100 can obtain the output data of the multi-head attention module, such as an attention matrix, also known as an attention correlation matrix or attention matrix, by capturing its input (e.g., query) and key. The attention matrix can be the dot product of the query matrix Q and the key matrix K. The anomaly management system 100 can calculate the entropy value of the attention matrix, i.e., the mean of the row entropy values, as shown below:
[0103] Where A is the attention matrix, A ij This represents the element in the i-th row and j-th column of the attention matrix. n represents the number of elements in the attention matrix.
[0104] In some cases, the anomaly management system 100 can also calculate the temporal difference in attention entropy, as shown below: ΔEntropy = Entropy(A) t -Entropy(A) t-1(2)
[0105] Among them, Entropy(A) t Let Entropy(A) represent the attention entropy at the t-th training step. t-1 Let represent the attention entropy at the (t-1)th training step. ΔEntropy represents the difference between the attention entropy and the historical attention entropy from the previous training step.
[0106] Next, we will introduce the calculation method of stable rank norm using an AI model with a hybrid expert model architecture.
[0107] Referring to Figure 5, a schematic diagram of a hybrid expert model architecture for an AI model is shown. The AI model includes a gating network 502 and multiple expert modules 504. The input layer can be connected to the gating network 502. The gating network 502 is used to select several expert modules 504 from the multiple expert modules 504 to perform the current task based on the input of the input layer. Then, the output data of the selected expert modules 504 can be weighted and calculated as input to the output layer.
[0108] The anomaly management system 100 can capture the weights W of the linear layers within the expert module 504 at each step in the Transformer block, thereby calculating the stable rank norm, as shown below:
[0109] Where A represents the matrix formed by the weights, ‖A‖ F Let f(A) denote the Frobenius norm of matrix A, and ||A||2 denote the induced 2-norm of the matrix.
[0110] It should be noted that the anomaly management system 100 can use hook functions to capture the output data or parameters of the corresponding modules. Specifically, the deep learning framework provides hook registration methods for modules, which can intercept input data, output data, or parameters during the forward propagation of a module. Based on this, the anomaly management system 100 can use hook functions to intercept the output data or parameters of the corresponding modules to determine the attention entropy or stable rank norm.
[0111] For AI models during training, the anomaly management system 100 collects lightweight, measurable feature indicators for online monitoring with minimal impact on the training process. Data shows that the performance degradation caused by introducing the anomaly management system 100 for online monitoring of the AI model training process is less than 2%. For different observable feature indicators, the anomaly management system 100 can employ corresponding acquisition tools to meet performance requirements.
[0112] S306, The anomaly management system 100 performs anomaly detection based on observable feature indicators to obtain anomaly detection results during the training process of the AI model.
[0113] Anomaly detection results are used to indicate whether anomalies have occurred or are about to occur during the training process of an AI model. Anomalies that are about to occur can be those that have not yet occurred, but have a high probability of happening without intervention. For example, anomalies that are about to occur could include training instability behaviors such as loss spikes or loss divergence that the AI model may exhibit in the future (e.g., the next training step, the next two training steps, the next K training steps, where K is a positive integer).
[0114] In some possible implementations, the anomaly management system 100 can input observable feature indicators into the prediction model for anomaly detection, obtaining anomaly detection results. These results indicate whether training anomalies have occurred or are about to occur during the AI model's training process. The prediction model is obtained by training a logistic regression model using training data. The training data includes historical observable feature indicators when the loss function converges and when the loss function fails to converge during AI model training.
[0115] In practical implementation, the logistic regression model can be represented as y = F(x). The training data (x, y) for training this logistic regression model comes from historical observable feature indicators and training labels obtained during the AI model training process. Here, x represents the entropy and stable rank mentioned above, and y represents the binary classification label used to train the AI model. A binary classification label value of 1 indicates training anomalies, such as loss divergence or other unstable training behaviors; a binary classification label value of 0 indicates normal training. After training the logistic regression model using the above training data to obtain the prediction model, this prediction model can be used to collect observable feature indicators for other AI models and output the predicted probability of training anomalies in real time based on these observable feature indicators, thus achieving real-time monitoring of training instability and other training anomalies. It should be noted that the training data can also be constructed using only one type of observable feature indicator; in other words, the prediction model can also predict the probability of anomalies occurring or about to occur during the training process of the AI model based on one type of observable feature indicator.
[0116] In other possible implementations, the anomaly management system 100 can match observable feature indicators with anomaly detection rules to obtain anomaly detection results. These anomaly detection rules can be threshold-based rules, and may include built-in or custom rules. The anomaly management system 100 can compare the observable feature indicators with the thresholds in the threshold-based rules to determine whether an anomaly has occurred or is about to occur.
[0117] In practical implementation, the anomaly management system 100 can perform threshold judgment on the attention entropy of shallow attention modules in the AI model to determine whether an anomaly has occurred or will occur. The shallow attention modules can be the attention modules of the target layer, where the number of layers is less than a set value; in some examples, the target layer can be the first layer.
[0118] In the first scenario, the anomaly management system 100 compares the attention entropy of the target layer's attention module with a first threshold. If the attention entropy of the target layer's attention module first decays below the first threshold and then recovers to above the first threshold within Q training steps, an anomaly detection result is obtained. This result indicates that the AI model's training process experienced a short-term anomaly and has since recovered. Specifically, if the attention entropy of the target layer's attention module decays below the first threshold within the first Q training steps and then recovers to above the first threshold in the current training step, it indicates that the attention entropy of the target layer's attention module exhibits a collapse followed by recovery phenomenon. Correspondingly, the AI model's training process experienced a short-term anomaly and immediately recovered. Here, Q and the first threshold can be set empirically; for example, Q can be set to a value less than or equal to 3, including but not limited to 1 or 2, and the first threshold can be set to 1 / 2.
[0119] In the second scenario, the anomaly management system 100 can compare the attention entropy of the target layer's attention module with a first threshold, and also compare the difference between the target layer's attention entropy and the historical attention entropy of the previous training step (denoted as Δentropy) with a second threshold. When the attention entropy of the target layer's attention module is greater than the first threshold, and the difference Δentropy between the target layer's attention entropy and the historical attention entropy of the previous training step is greater than the second threshold, the anomaly management system 100 obtains an anomaly detection result. This anomaly detection result indicates that an anomaly will occur during the training process of the AI model. The second threshold can be set empirically, for example, it can be set to 0.05. If Δentropy is greater than the second threshold, it indicates that the attention entropy has experienced a continuous increase in fluctuation, the AI model's training has entered an unstable region, and it is in an unhealthy state, subsequently exhibiting unstable training behavior (with a high probability of such behavior).
[0120] In the third scenario, the anomaly management system 100 can compare the attention entropy of the target layer's attention module with a first threshold, and also compare the difference (Δentropy) between the target layer's attention entropy and the historical attention entropy of the previous training step with a second threshold. When the attention entropy of the target layer's attention module decays below the first threshold, and the difference between the target layer's attention entropy and the historical attention entropy of the previous training step is greater than the second threshold, the anomaly management system 100 obtains an anomaly detection result. This anomaly detection result indicates that an anomaly has occurred in the AI model's training process. For example, the anomaly detection result indicates that the AI model's training process has exhibited unstable training behaviors such as loss divergence.
[0121] It should be noted that the anomaly management system 100 can provide alerts for the three situations mentioned above. These alerts can be either warnings or errors. For the first or second situation, the anomaly management system 100 will issue a warning, indicating that the AI model has experienced a short-term failure and will recover immediately, or that the AI model is in an unhealthy state and will subsequently experience training instability. For the third situation, the anomaly management system 100 can provide error messages, such as indicating that the AI model has experienced training anomalies, including loss divergence.
[0122] Furthermore, when the anomaly detection result indicates that an anomaly has occurred or is about to occur in the training process of the AI model, the anomaly management system 100 can also determine the cause of the anomaly based on observable feature indicators. Specifically, the anomaly management system 100 can determine the cause of the anomaly based on observable feature indicators combined with anomaly detection rules. For example, in the first case mentioned above, the cause of the anomaly could be that the number of warm-up steps is too small, resulting in the AI model not being effectively warmed up; in the second case, the cause could be an anomaly in the software stack or the training script; and in the third case, the cause could be that the learning rate is set too high, leading to excessive singularity in the parameters of the AI model.
[0123] This application provides some examples of anomaly detection rules, which are threshold determination rules for the first-layer attention module, as shown in the table below:
[0124] Table 1
[0125] The aforementioned anomaly detection rules have been validated through extensive model experiments, demonstrating high reliability in anomaly detection results. Therefore, when an observable feature reaches a threshold, such as when it successfully matches one of the aforementioned anomaly detection rules, the anomaly management system 100 can display an alarm or error message, indicating an abnormal hyperparameter setting. Furthermore, the list of anomaly detection rules can accumulate experience, broadening the coverage of the monitoring tool.
[0126] In some possible implementations, when the cause of the anomaly includes abnormal hyperparameter settings, such as an excessively high learning rate, the anomaly management system 100 can modify the parameters of the AI model to increase the attention entropy. This allows the attention entropy of each attention block to be maintained at a high amplitude level, thereby achieving adaptive training recovery. Compared to recovery methods such as modifying hyperparameters or rolling back training, adaptive entropy-increasing recovery training can achieve almost no interruption, with minimal impact on training, low cost, and high availability. It should be noted that when modifying hyperparameters, the anomaly management system 100 can first store the changes (e.g., automatically store the interruption point).
[0127] The anomaly management system 100 can perform a smoothing operation on the linear matrix parameters of the AI model to achieve adaptive entropy increase. Specifically, the linear matrix parameters consist of a vector of all singular values of the weights in the parameter matrix arranged in descending order. The anomaly management system 100 can compress the first to the Mth singular values in the vector, making the compressed first to the Mth singular values equal to the (M+1)th singular value. Here, M is the floor function of the stable rank norm. This allows for the smoothing operation on the linear matrix parameters of the AI model.
[0128] To facilitate understanding, an example is provided below. In this example, the parameter matrix W can be subjected to Singular Value Decomposition (SVD). SVD is an important matrix decomposition method in linear algebra, with wide applications in machine learning and other fields. Specifically, Singular Value Decomposition can decompose the parameter matrix W into a product of the following matrices: W = Udiag(Σ)V (4)
[0129] Where Σ=[σ1,…,σ n Let be a vector representing all singular values of the parameter matrix W arranged in descending order. U and V are orthogonal matrices, with the column vectors of orthogonal matrix U being the left singular vectors and the column vectors of orthogonal matrix V being the right singular vectors.
[0130] The anomaly management system 100 can identify the first singular value in Σ up to the second singular value. The singular value is compressed to the same as the first singular value. All singular values are equal, as shown below:
[0131] Where, σ i This represents the i-th singular value. Let represent the i-th singular value after the update, and SR(W) represent the stable rank norm.
[0132] Then, the exception management system 100 can reset the parameter matrix W, as shown below: W * =Udiag(Σ * )V (6)
[0133] Considering the high overhead of a single SVD, this application supports the use of low-rank SVD, which only requires computation of the parameter matrix. Large singular values and their corresponding singular vectors. The anomaly management system 100 can be implemented using APIs provided by deep learning frameworks, such as svd_lowrank. Calculating SR(W) only involves the computation of the largest singular value, which can be quickly computed using an approximation algorithm; therefore, the aforementioned smoothing method incurs almost no additional overhead.
[0134] This application also provides pseudocode to illustrate anomaly detection and recovery during AI model training.
[0135] When the anomaly management system 100 detects that the AI model M is experiencing training instability, it can target the parameter matrix W of the line module in the AI model M. i Perform the following smoothing operation: Perform singular value decomposition on the parameter matrix to obtain the singular value vector S. i and orthogonal matrix U i V i For singular value vectors S i The element with index 0 is at index floor(SR(W) i The element of ))-1, that is, the first singular value to the floor(SR(W) i The anomaly management system 100 can compress the above singular values to a value with the index floor(SR(W)) to a value equal to floor(SR(W). i The elements of )) are equal, that is, from the first singular value to the floor(SR(W)). i The singular value is compressed to the same level as the floor (SR(W)). i ))+1 singular values are equal.
[0136] In some possible implementations, the anomaly management system 100 can also use convolution or activation functions to perform smoothing operations on the linear matrix parameters of the AI model. This will be explained in detail below.
[0137] When the anomaly management system 100 uses convolution to perform smoothing operations on the linear matrix parameters of an AI model, it can first perform singular value decomposition on the linear matrix parameters and then update the singular values, as shown below:
[0138] Where n is the number of singular values, i and j are the indices of the singular values, and k is the half-window length of the convolution, which can generally be set to 3. σ j This represents the j-th singular value. This represents the i-th singular value after the update.
[0139] When the anomaly management system 100 uses activation functions to perform smoothing operations on the linear matrix parameters of an AI model, it can first perform singular value decomposition on the linear matrix parameters and then update the singular values, as shown below:
[0140] Where n is the number of singular values, i and j are the indices of the singular values, and β is the scaling parameter used to control the steepness of the activation function, which is set to 1 by default. Furthermore, the activation function can include, but is not limited to, the softmax function. σ i σ j Let i and j represent the i-th singular value and the j-th singular value, respectively. This represents the i-th singular value after the update.
[0141] Based on the above description, this application provides an anomaly management method for AI model training. This method leverages the interpretable mechanisms of AI model training to mine observable feature indicators within the AI model, including observable feature indicators related to the AI model's hyperparameters, such as attention entropy. These observable feature indicators enable online monitoring and anomaly prediction, covering a wide range of unstable training scenarios with low detection costs, even achieving non-destructive detection and high availability. Furthermore, this method accelerates the localization of training anomalies caused by abnormal hyperparameter settings, reducing manual troubleshooting time from three days to seconds, lowering the Mean Time To Repair (MTTR), and improving the availability of computing clusters such as training clusters. For hyperparameter setting anomalies, this application also provides an automated, recoverable, and fault-tolerant method to ensure the normal and stable training of the AI model and reduce the time cost of retraining due to unreasonable hyperparameter settings.
[0142] Next, we will introduce the anomaly management method for AI model training in this application, using a specific application scenario. This method can be divided into a preparation phase and an anomaly management phase. In the preparation phase, the model execution system (e.g., a development platform, an integrated AI platform) executes step 1, enabling online monitoring. In the anomaly management phase, the anomaly management system 100 (e.g., a monitoring tool) executes steps 2 and 3 to monitor whether training instability or impending training instability has occurred during the AI model training process. Steps 1, 2, and 3 will be explained in detail below.
[0143] Step 1: The model running system enables online monitoring of the metrics of the AI model trained on the computing card by attaching the monitoring tool monitor.
[0144] Specifically, the model running system can respond to the user's initialization operation by first enabling the monitoring tool in the training script, importing the plug-in monitoring code "from monitor import TrainerMon", then initializing the monitoring tool "monitor = TrainerMon(config_file_path = ". / monitor_config.json",)", and enabling the monitoring function "monitor.monitor_gnorm_with_ad(model,...)".
[0145] Then, the model runtime system can respond to the user's configuration operations, configure monitor_confing.json, and control the monitoring behavior. The configured monitor_confing.json is shown below:
[0146] It should be noted that the monitor_confing.json configuration file can be used to configure information such as the computing cards to be monitored (e.g., rank 0) and the modules in the AI model (e.g., attention module, experts module). The monitoring tool can automatically return the prediction probability of AI model training instability in the log, and the return format can be ">instability probability:0.125243".
[0147] Step 2: The monitoring tool acquires observable feature indicators collected online, matches the observable feature indicators with the anomaly detection rules, determines whether an anomaly has occurred or will occur, and determines whether the cause of the anomaly is an abnormality in the hyperparameter settings when an anomaly occurs or will occur.
[0148] Step 3. If the cause of the anomaly determined in Step 2 is an abnormal hyperparameter setting, the monitoring tool will provide the user with recovery suggestions.
[0149] The recovery suggestions may include modifying hyperparameters or modifying the AI model's parameters to adaptively increase entropy. When the user chooses to modify the AI model's parameters to adaptively increase entropy, the monitoring tool can perform a smoothing operation on the linear matrix parameters of the AI model, thereby achieving adaptive entropy recovery training.
[0150] Based on the aforementioned AI model training-based exception management method, this application also provides an exception management system 100. The exception management system 100 provided in this application will be described below from a functional modular perspective.
[0151] Referring to Figure 1, which shows a schematic diagram of an anomaly management system 100, the anomaly management system 100 is used for anomaly management during the training process of an AI model. The training of the AI model is performed by at least one computing card, and the at least one computing card includes a first computing card. The anomaly management system 100 includes:
[0152] The feature monitoring module 102 is used to identify the first module currently being executed by the first computing card during the training process of the AI model. The first module includes one or more modules in the AI model. When the module type of the first module is a target type, the module 102 is used to obtain observable feature indicators when the first computing card executes the first module. The observable feature indicators include attention entropy, which is used to represent the degree of clustering of correlations in the input data of the first module.
[0153] The anomaly detection module 104 is used to perform anomaly detection based on the observable feature indicators and obtain the anomaly detection results of the training process of the AI model.
[0154] For example, the feature monitoring module 102 and the anomaly perception module 104 described above can be implemented in hardware or in software.
[0155] When implemented in software, the feature monitoring module 102 and the anomaly detection module 104 can be applications running on computing devices, such as computing engines. These applications can also be virtualized and provided to users as virtualization services. Virtualization services can include virtual machine (VM) services, bare metal server (BMS) services, or container services. Specifically, a VM service can be a service that uses virtualization technology to create a pool of virtual machine (VM) resources on multiple physical hosts, providing VMs for users to use on demand. A BMS service is a service that uses virtualization technology to create a pool of BMS resources on multiple physical hosts, providing BMS for users to use on demand. A container service is a service that uses virtualization technology to create a pool of container resources on multiple physical hosts, providing containers for users to use on demand. A VM is a simulated virtual computer, that is, a logical computer. A BMS is a scalable, high-performance computing service with computing performance indistinguishable from traditional physical machines, featuring secure physical isolation. A container is a kernel virtualization technology that provides lightweight virtualization to isolate user space, processes, and resources. It should be understood that the VM service, BMS service, and container service mentioned above are merely specific examples. In practical applications, virtualization services can also include other lightweight or heavyweight virtualization services, which are not specifically limited here.
[0156] When implemented in hardware, the feature monitoring module 102 and the anomaly sensing module 104 may include at least one computing device, such as a server. Alternatively, the feature monitoring module 102 and the anomaly sensing module 104 may also include devices implemented using application-specific integrated circuits (ASICs) or programmable logic devices (PLDs). The PLD may be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), generic array logic (GAL), or any combination thereof.
[0157] In some possible implementations, the anomaly detection module 104 is specifically used for:
[0158] The observable feature indicators are input into the prediction model for anomaly detection to obtain anomaly detection results. The prediction model is obtained by training a logistic regression model using training data. The training data includes historical observable feature indicators when the loss function converges and historical observable feature indicators when the loss function does not converge during the training of the AI model.
[0159] In some possible implementations, the anomaly detection module 104 is specifically used for:
[0160] The observable feature indicators are matched with the anomaly detection rules to obtain the anomaly detection results.
[0161] In some possible implementations, the first module includes an attention module for the target layer in the AI model, and the anomaly perception module 104 is specifically used for:
[0162] When the attention entropy of the attention module of the target layer first decays to below the first threshold, and then recovers to above the first threshold within Q training steps, an anomaly detection result is obtained. The anomaly detection result indicates that the training process of the AI model has experienced a short-term anomaly and has returned to normal. Q is a positive integer.
[0163] In some possible implementations, the first module includes an attention module for the target layer in the AI model, and the anomaly perception module 104 is specifically used for:
[0164] When the attention entropy of the attention module of the target layer is greater than a first threshold, and the difference between the attention entropy of the attention module of the target layer and the historical attention entropy of the previous training step is greater than a second threshold, an anomaly detection result is obtained, and the anomaly detection result indicates that an anomaly will occur in the training process of the AI model.
[0165] In some possible implementations, the first module includes an attention module for the target layer in the AI model, and the anomaly perception module 104 is specifically used for:
[0166] When the attention entropy of the attention module of the target layer decays to below the first threshold, and the difference between the attention entropy of the attention module of the target layer and the historical attention entropy of the previous training step is greater than the second threshold, an anomaly detection result is obtained, and the anomaly detection result indicates that an anomaly has occurred in the training process of the AI model.
[0167] In some possible implementations, the exception management system 100 also includes:
[0168] The anomaly localization module 106 is used to determine the cause of the anomaly based on the observable feature indicators when the anomaly detection result indicates that an anomaly has occurred or will occur during the training process of the AI model.
[0169] Similar to the anomaly detection module 104, the anomaly location module 106 can be implemented in hardware or in software.
[0170] When implemented in software, the anomaly location module 106 can be an application running on a computing device, which can also be virtualized as a VM service, BMS service, or container service for user use. When implemented in hardware, the anomaly location module 106 can include at least one computing device, such as a server. Alternatively, the anomaly location module 106 can also be a device implemented using an ASIC or a PLD.
[0171] In some possible implementations, the exception management system 100 also includes:
[0172] The anomaly recovery module 108 is used to modify the parameters of the AI model to improve the attention entropy when the cause of the anomaly includes an abnormal hyperparameter setting.
[0173] Similar to the anomaly detection module 104, the anomaly recovery module 108 can be implemented in hardware or in software.
[0174] When implemented in software, the exception recovery module 108 can be an application running on a computing device, which can also be virtualized as a VM service, BMS service, or container service for user use. When implemented in hardware, the exception recovery module 108 can include at least one computing device, such as a server. Alternatively, the exception recovery module 108 can also be a device implemented using an ASIC or a PLD.
[0175] This application also provides a computing device 600. As shown in FIG6, the computing device 600 includes: a bus 602, a processor 604, a memory 606, and a communication interface 608. The processor 604, the memory 606, and the communication interface 608 communicate with each other via the bus 602. The computing device 600 may be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 600.
[0176] Bus 602 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, only one line is used in Figure 6, but this does not imply that there is only one bus or one type of bus. Bus 602 can include pathways for transmitting information between various components of computing device 600 (e.g., memory 606, processor 604, communication interface 608).
[0177] Processor 604 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0178] The memory 606 may include volatile memory, such as random access memory (RAM). The memory 606 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid-state drive (SSD). The memory 606 stores executable program code, which the processor 604 executes to implement the aforementioned exception management method for AI model training. Specifically, the memory 606 stores instructions for the exception management system 100 to execute the exception management method for AI model training. For example, the memory 606 may store instructions for implementing the functions of the feature monitoring module 102 and the exception perception module 104. Furthermore, the memory 606 may also store instructions for implementing the functions of the exception location module 106 and the exception recovery module 108.
[0179] The communication interface 608 uses transceiver modules, such as, but not limited to, network interface cards and transceivers, to enable communication between the computing device 600 and other devices or communication networks.
[0180] This application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.
[0181] As shown in Figure 7, the computing device cluster includes at least one computing device 600. The memory 606 of one or more computing devices 600 in the computing device cluster may store instructions from the same exception management system 100 for executing exception management methods for AI model training.
[0182] In some possible implementations, one or more computing devices 600 in the computing device cluster can also be used to execute some instructions of the exception management system 100 for executing the exception management method for AI model training. In other words, a combination of one or more computing devices 600 can jointly execute the instructions of the exception management system 100 for executing the exception management method for AI model training.
[0183] It should be noted that the memory 606 in different computing devices 600 in the computing device cluster can store different instructions for executing some functions of the exception management system 100.
[0184] Figure 8 illustrates one possible implementation. As shown in Figure 8, two computing devices 600A and 600B are connected via a communication interface 608. The memory in computing device 600A stores instructions for executing the functions of the feature monitoring module 102. The memory in computing device 600B stores instructions for executing the functions of the anomaly perception module 104. Furthermore, the memory of computing device 600B can also store instructions for executing the functions of the anomaly localization module 106, and the memory of computing device 600A can also store instructions for executing the functions of the anomaly recovery module 108. In other words, the memory 606 of computing devices 600A and 600B jointly stores the instructions of the anomaly management system 100 for executing the anomaly management method for AI model training.
[0185] The connection method between the computing device clusters shown in Figure 8 can be considered because the anomaly management method for AI model training provided in this application requires a lot of resources for anomaly detection. Therefore, it is considered that the functions implemented by the anomaly detection module 104 and the functions implemented by the feature monitoring module 102 are executed by different computing devices. For example, the functions implemented by the feature monitoring module 102 can be executed by computing device 600A, and the functions implemented by the anomaly detection module 104 can be executed by computing device 600B. Similarly, the functions implemented by the anomaly recovery module 108 can be executed by computing device 600A, and the functions implemented by the anomaly localization module 106 can be executed by computing device 600B.
[0186] It should be understood that the functions of computing device 600A shown in Figure 8 can also be performed by multiple computing devices 600. Similarly, the functions of computing device 600B can also be performed by multiple computing devices 600.
[0187] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc. Figure 9 illustrates one possible implementation. As shown in Figure 9, two computing devices 600C and 600D are connected via a network. Specifically, they are connected to the network through communication interfaces in each computing device. In this type of possible implementation, the memory 606 in computing device 600C stores instructions for the function of the feature monitoring module 102. Simultaneously, the memory 606 in computing device 600D stores instructions for the function of the anomaly detection module 104. Furthermore, the memory of computing device 600C can also store instructions for executing the function of the anomaly recovery module 108, and the memory of computing device 600D can also store instructions for executing the function of the anomaly location module 106.
[0188] The connection method between the computing device clusters shown in Figure 9 can be considered because the anomaly management method for AI model training provided in this application requires a large amount of resources for anomaly detection, such as model inference, to achieve anomaly detection. Therefore, the functions implemented by the feature monitoring module 102 are considered to be executed by the computing device 600C, and the functions implemented by the anomaly detection module 104 are executed by the computing device 600D. In addition, when the anomaly management system 100 also includes an anomaly location module 106 and an anomaly recovery module 108, the functions implemented by the anomaly recovery module 108 are executed by the computing device 600C, and the functions implemented by the anomaly location module 106 are executed by the computing device 600D.
[0189] It should be understood that the functions of computing device 600C shown in Figure 9 can also be performed by multiple computing devices 600. Similarly, the functions of computing device 600D can also be performed by multiple computing devices 600.
[0190] This application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store, or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct a computing device or cluster of computing devices to execute the above-described exception management method applied to the exception management system 100 for performing AI model training.
[0191] This application also provides a computer program product containing instructions. The computer program product may be a software or program product containing instructions, capable of running on a computing device or stored on any usable medium. When the computer program product runs on at least one computing device, it causes the at least one computing device to execute the above-described AI model training exception management method.
[0192] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.
Claims
1. An anomaly management method for training an artificial intelligence (AI) model, characterized in that, The training of the AI model is performed by at least one computing card, the at least one computing card including a first computing card, and the method during the training of the AI model includes: Identify the first module currently being executed by the first computing card, wherein the first module includes one or more modules in the AI model; When the module type of the first module is a target type, obtain the observable feature indicators when the first computing card executes the first module. The observable feature indicators include attention entropy, which is used to represent the degree of clustering of correlations in the input data of the first module. Anomaly detection is performed based on the observable feature indicators to obtain the anomaly detection results of the training process of the AI model.
2. The method according to claim 1, characterized in that, The step of performing anomaly detection based on the observable feature indicators to obtain anomaly detection results for the training process of the AI model includes: The observable feature indicators are input into the prediction model for anomaly detection to obtain anomaly detection results. The prediction model is obtained by training a logistic regression model using training data. The training data includes historical observable feature indicators when the loss function converges and historical observable feature indicators when the loss function does not converge during the training of the AI model.
3. The method according to claim 1 or 2, characterized in that, The step of performing anomaly detection based on the observable feature indicators to obtain anomaly detection results for the training process of the AI model includes: The observable feature indicators are matched with the anomaly detection rules to obtain the anomaly detection results.
4. The method according to claim 3, characterized in that, The first module includes an attention module for the target layer in the AI model. The step of matching the observable feature indicators with anomaly detection rules to obtain anomaly detection results during the training process of the AI model includes: When the attention entropy of the attention module of the target layer first decays to below the first threshold, and then recovers to above the first threshold within Q training steps, an anomaly detection result is obtained. The anomaly detection result indicates that the training process of the AI model has experienced a short-term anomaly and has returned to normal. Q is a positive integer.
5. The method according to claim 3, characterized in that, The first module includes an attention module for the target layer in the AI model. The step of matching the observable feature indicators with anomaly detection rules to obtain anomaly detection results during the training process of the AI model includes: When the attention entropy of the attention module of the target layer is greater than a first threshold, and the difference between the attention entropy of the attention module of the target layer and the historical attention entropy of the previous training step is greater than a second threshold, an anomaly detection result is obtained, and the anomaly detection result indicates that an anomaly will occur in the training process of the AI model.
6. The method according to claim 3, characterized in that, The first module includes an attention module for the target layer in the AI model. The step of matching the observable feature indicators with anomaly detection rules to obtain anomaly detection results during the training process of the AI model includes: When the attention entropy of the attention module of the target layer decays to below a first threshold, and the difference between the attention entropy of the attention module of the target layer and the historical attention entropy of the previous training step is greater than a second threshold, an anomaly detection result is obtained, and the anomaly detection result indicates that an anomaly has occurred in the training process of the AI model.
7. The method according to any one of claims 1 to 6, characterized in that, The method further includes: When the anomaly detection result indicates that the training process of the AI model has an anomaly or will have an anomaly, the cause of the anomaly is determined based on the observable feature indicators.
8. The method according to claim 7, characterized in that, The method further includes: If the cause of the anomaly includes abnormal hyperparameter settings, modify the parameters of the AI model to improve the attention entropy.
9. The method according to claim 8, characterized in that, The modification of the parameters of the AI model includes: Smoothing operation is performed on the linear matrix parameters of the AI model.
10. The method according to claim 9, characterized in that, The linear matrix parameters include a vector of all singular values of the weights in the parameter matrix arranged in descending order. The smoothing operation on the linear matrix parameters of the AI model includes: The first to the Mth singular values in the vector are compressed such that the compressed first to the Mth singular values are equal to the (M+1)th singular value, where M is the floor function of the stable rank norm, which is used to evaluate the convergence of the loss function of the AI model; or, The linear matrix parameters of the AI model are smoothed using convolution or activation functions.
11. An anomaly management system, characterized in that, The anomaly management system is used to manage anomalies during the training process of an artificial intelligence (AI) model. The training of the AI model is performed by at least one computing card, and the at least one computing card includes a first computing card. The anomaly management system includes: The feature monitoring module is used to identify the first module currently being executed by the first computing card during the training process of the AI model. The first module includes one or more modules in the AI model. When the module type of the first module is a target type, the module obtains observable feature indicators when the first computing card executes the first module. The observable feature indicators include attention entropy, which is used to represent the degree of clustering of correlations in the input data of the first module. An anomaly detection module is used to perform anomaly detection based on the observable feature indicators and obtain the anomaly detection results of the training process of the AI model.
12. The system according to claim 11, characterized in that, The anomaly detection module is specifically used for: The observable feature indicators are input into the prediction model for anomaly detection to obtain anomaly detection results. The prediction model is obtained by training a logistic regression model using training data. The training data includes historical observable feature indicators when the loss function converges and historical observable feature indicators when the loss function does not converge during the training of the AI model.
13. The system according to claim 11 or 12, characterized in that, The anomaly detection module is specifically used for: The observable feature indicators are matched with the anomaly detection rules to obtain the anomaly detection results.
14. The system according to claim 13, characterized in that, The first module includes the attention module of the target layer in the AI model, and the anomaly detection module is specifically used for: When the attention entropy of the attention module of the target layer first decays to below the first threshold, and then recovers to above the first threshold within Q training steps, an anomaly detection result is obtained. The anomaly detection result indicates that the training process of the AI model has experienced a short-term anomaly and has returned to normal. Q is a positive integer.
15. The system according to claim 13, characterized in that, The first module includes the attention module of the target layer in the AI model, and the anomaly detection module is specifically used for: When the attention entropy of the attention module of the target layer is greater than a first threshold, and the difference between the attention entropy of the attention module of the target layer and the historical attention entropy of the previous training step is greater than a second threshold, an anomaly detection result is obtained, and the anomaly detection result indicates that an anomaly will occur in the training process of the AI model.
16. The system according to claim 13, characterized in that, The first module includes the attention module of the target layer in the AI model, and the anomaly detection module is specifically used for: When the attention entropy of the attention module of the target layer decays to below a first threshold, and the difference between the attention entropy of the attention module of the target layer and the historical attention entropy of the previous training step is greater than a second threshold, an anomaly detection result is obtained, and the anomaly detection result indicates that an anomaly has occurred in the training process of the AI model.
17. The system according to any one of claims 11 to 16, characterized in that, The system also includes: An anomaly localization module is used to determine the cause of the anomaly based on the observable feature indicators when the anomaly detection result indicates that an anomaly has occurred or will occur during the training process of the AI model.
18. The system according to claim 17, characterized in that, The system also includes: An anomaly recovery module is used to modify the parameters of the AI model to improve the attention entropy when the cause of the anomaly includes an abnormal hyperparameter setting.
19. The system according to claim 18, characterized in that, The anomaly recovery module is specifically used for: Smoothing operation is performed on the linear matrix parameters of the AI model.
20. The system according to claim 19, characterized in that, The linear matrix parameters include a vector of all singular values of the weights in the parameter matrix arranged in descending order. The anomaly recovery module is specifically used for: The first to the Mth singular values in the vector are compressed such that the compressed first to the Mth singular values are equal to the (M+1)th singular value, where M is the floor function of the stable rank norm, which is used to evaluate the convergence of the loss function of the AI model; or, The linear matrix parameters of the AI model are smoothed using convolution or activation functions.
21. A computing device cluster, characterized in that, The computing device cluster includes at least one computing device, the at least one computing device including at least one processor and at least one memory, the at least one memory storing computer-readable instructions; the at least one processor executes the computer-readable instructions to cause the computing device cluster to perform the anomaly management method for AI model training as described in any one of claims 1 to 10.
22. A computer-readable storage medium, characterized in that, Includes computer-readable instructions; the computer-readable instructions are used to implement the anomaly management method for AI model training according to any one of claims 1 to 10.
23. A computer program product, characterized in that, Includes computer-readable instructions; the computer-readable instructions are used to implement the anomaly management method for AI model training according to any one of claims 1 to 10.