Abnormal accelerator card detection method, model training method, equipment and medium
By acquiring the current status time-series data of the AI accelerator card, and utilizing the accelerator card performance anomaly detection model and large language model, complex faults of the AI accelerator card can be identified and located. This solves the problem of early identification in existing technologies and realizes early fault warning and location of intelligent computing clusters.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI SUIYUAN TECH CO LTD
- Filing Date
- 2025-12-30
- Publication Date
- 2026-04-21
AI Technical Summary
In large-scale intelligent computing clusters, the performance degradation of AI accelerator cards leads to a surge in the risk of card loss across the entire rack. It is impossible to effectively capture complex failure modes, and existing detection methods are post-event detections, which cannot intervene in the early stages.
By acquiring the current status time-series data of the AI accelerator card, and utilizing the accelerator card performance anomaly detection model and large language model, abnormal indicators are identified and performance anomaly location prompts are generated, enabling early identification and location of complex faults.
It enables precise location of complex faults in AI accelerator cards, provides early intervention, avoids hardware performance degradation and abnormal states, and improves the reliability and computing power output of intelligent computing clusters.
Smart Images

Figure CN121901042A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent computing cluster fault analysis technology, and in particular to a method for detecting abnormal acceleration cards, a model training method, equipment, and media. Background Technology
[0002] In large-scale intelligent computing clusters, the performance degradation of AI (Artificial Intelligence) accelerator cards can lead to a surge in the risk of card loss across the entire rack due to the amplification effect. The presence of low-speed cards can also lower the overall task efficiency, severely impacting the reliability and computing power output of the cluster. This results in a serious problem of unstable cluster computing power and reduced efficiency caused by the performance degradation of AI accelerator cards in large-scale intelligent computing clusters.
[0003] Currently, monitoring of AI accelerator card performance degradation has the following limitations: 1) It cannot effectively capture complex failure modes caused by abnormal coordination of internal components such as computing cores, memory, power consumption, and PCIe (Peripheral Component Interconnect express, a high-speed serial computer expansion bus standard). 2) It is mostly "post-event" detection, that is, alarms are only triggered after a failure (such as a complete card crash or task failure) has occurred, making it impossible to intervene in the early stages of hardware performance degradation and abnormal status. Summary of the Invention
[0004] This invention provides a method for detecting abnormal accelerator cards, a model training method, equipment, and media to solve the current problems of being unable to accurately locate complex faults of AI accelerator cards in intelligent computing clusters and the lag in fault identification.
[0005] According to one aspect of the present invention, a method for detecting abnormal accelerator cards is provided, comprising:
[0006] Obtain the current status time-series data of the AI accelerator card, and determine the abnormal indicator detection results based on the current status time-series data and the accelerator card performance anomaly detection model;
[0007] When it is determined that the AI accelerator card has performance abnormalities based on the abnormal indicator detection results, an abnormal correlation judgment is performed on the abnormal indicator detection results to obtain the abnormal correlation detection results.
[0008] Based on the anomaly correlation detection results, performance anomaly location prompt words are generated, and based on the performance anomaly location prompt words and the large language model, performance anomaly root cause feedback results are generated.
[0009] According to another aspect of the present invention, a model training method is provided, comprising:
[0010] Obtain historical status data of the AI accelerator card;
[0011] Based on the historical status data of the accelerator card, the target pre-trained model is trained to obtain the accelerator card performance anomaly detection model;
[0012] The target pre-trained model includes a mask embedding processing layer, a rotation position encoding layer, a Transformers encoding layer, and a Transformers decoding layer.
[0013] According to another aspect of the present invention, a detection device for abnormal acceleration cards is provided, comprising:
[0014] The anomaly detection module is used to acquire the current status time-series data of the AI accelerator card, and determine the anomaly detection results based on the current status time-series data and the accelerator card performance anomaly detection model.
[0015] The anomaly correlation judgment module is used to perform anomaly correlation judgment on the anomaly index detection results when it is determined that the AI accelerator card has a performance abnormality based on the anomaly index detection results, and obtain the anomaly correlation detection results.
[0016] The anomaly root cause localization module is used to generate performance anomaly localization prompts based on anomaly association detection results, and to generate performance anomaly root cause feedback results based on the performance anomaly localization prompts and the large language model.
[0017] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0018] At least one processor; and a memory communicatively connected to said at least one processor;
[0019] The memory stores a computer program that can be executed by the at least one processor. The computer program is executed by the at least one processor to enable the at least one processor to execute the abnormal acceleration card detection method according to any embodiment of the present invention, or to execute the model training method according to the embodiment.
[0020] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions, the computer instructions being configured to cause a processor to execute and implement the abnormal acceleration card detection method according to any embodiment of the present invention, or to implement the model training method according to the embodiment.
[0021] According to another aspect of the present invention, a computer program product is provided, the computer program product comprising a computer program, which, when executed by a processor, implements the abnormal acceleration card detection method of any embodiment of the present invention, or implements the model training method of the embodiment.
[0022] The technical solution of this invention acquires the current state time-series data of the AI accelerator card, and determines the anomaly detection results based on the current state time-series data and the accelerator card performance anomaly detection model. When the anomaly detection results indicate that the AI accelerator card has a performance anomaly, anomaly correlation judgment is performed on the anomaly detection results to obtain anomaly correlation detection results. Furthermore, based on the anomaly correlation detection results, performance anomaly location prompt words are generated, and based on the performance anomaly location prompt words and a large language model, a performance anomaly root cause feedback result is generated. This solution couples the abnormal performance indicators of the AI accelerator card identified by the accelerator card performance anomaly detection model to complete the identification of complex fault modes, and analyzes the anomaly location prompt words generated by the anomaly correlation detection results based on a large language model, thereby automatically generating intelligent analysis results of complex fault root causes of the AI accelerator card. This technology enables the identification and localization of complex faults in AI accelerator cards before a failure occurs in an intelligent computing cluster. It provides an opportunity for early intervention in the event of hardware performance degradation or abnormal states, overcoming the drawbacks of "post-event" detection of hardware faults in intelligent computing clusters. It solves the current problems of inaccurate localization of complex faults in AI accelerator cards in intelligent computing clusters and the lag in fault identification. It can accurately locate complex faults in AI accelerator cards in intelligent computing clusters before a failure occurs.
[0023] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0025] Figure 1 A flowchart of a method for detecting an abnormal acceleration card provided in Embodiment 1 of the present invention;
[0026] Figure 2 This is a flowchart of a method for detecting an abnormal acceleration card provided in Embodiment 2 of the present invention;
[0027] Figure 3 This is a schematic diagram of an abnormal acceleration card detection system;
[0028] Figure 4 This is a logical diagram illustrating mask annotation and embedding processing;
[0029] Figure 5 This is a schematic diagram of the index sequence;
[0030] Figure 6 A schematic diagram of the overall process of training a pre-trained model for a target;
[0031] Figure 7 This is a logical diagram of online anomaly detection and correlation identification;
[0032] Figure 8 This is a schematic diagram of the structure of a detection device for an abnormal acceleration card provided in Embodiment 4 of the present invention;
[0033] Figure 9 A schematic diagram of an electronic device that can be used to implement embodiments of the present invention is shown. Detailed Implementation
[0034] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0035] It should be noted that the terms "current," "target," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0036] Example 1
[0037] Figure 1 This is a flowchart of a method for detecting abnormal accelerator cards according to Embodiment 1 of the present invention. This embodiment is applicable to situations where complex faults in AI accelerator cards can be accurately located before a fault occurs in an intelligent computing cluster. This method can be executed by an abnormal accelerator card detection device, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the method includes:
[0038] Step 110: Obtain the current status time series data of the AI accelerator card, and determine the abnormal indicator detection results based on the current status time series data and the accelerator card performance anomaly detection model.
[0039] The current state time-series data can be the state time-series data of the AI accelerator card collected in the current sampling period, that is, the data sequence of AI accelerator card state changes recorded in chronological order. The accelerator card performance anomaly detection model can be used to identify abnormal performance indicators of the AI accelerator card. Abnormal performance indicators of the AI accelerator card may include, but are not limited to, board power consumption, board temperature, average SIP (Session Initiation Protocol) utilization within a sampling period, device operating time percentage within a sampling period, dynamic power management corresponding to the GCU (Gateway Control Unit) frequency level, AI accelerator card memory usage, AI accelerator card memory size, PCIe transmit throughput, and PCIe receive throughput. The anomaly indicator detection result can be the identification result of abnormal performance indicators of the AI accelerator card.
[0040] In this embodiment of the invention, the current state time series data of the AI accelerator card in the intelligent computing cluster can be collected according to the sampling period. After the current state time series data of the AI accelerator card is processed by mask annotation, it is input into the accelerator card performance anomaly detection model. The accelerator card performance anomaly detection model can then perform reasoning analysis on the masked current state time series data to determine the anomaly index detection result.
[0041] Step 120: When it is determined that the AI accelerator card has performance abnormalities based on the abnormal indicator detection results, perform abnormal correlation judgment on the abnormal indicator detection results to obtain abnormal correlation detection results.
[0042] Anomaly correlation discrimination can be an operation that correlates multiple anomaly indicators in the anomaly indicator detection results. The anomaly correlation detection results can be the correlation results of anomaly indicators of the AI accelerator card within the anomaly indicator detection results. These results can be used to reflect coupled faults in the AI accelerator card, i.e., multi-dimensional complex faults. Examples include "low utilization and high power consumption" and "high bandwidth and low computation."
[0043] In this embodiment of the invention, if the abnormal indicator detection result of the current AI accelerator card is empty, it indicates that the current AI accelerator card is performing normally. If the abnormal indicator detection result of the current AI accelerator card is not empty, it indicates that the current AI accelerator card is performing abnormally. When there are multiple abnormal indicators in the abnormal indicator detection result, an abnormal correlation judgment is performed on the abnormal indicator detection result of the current AI accelerator card to obtain the abnormal correlation detection result.
[0044] Step 130: Based on the anomaly correlation detection results, generate performance anomaly location prompt words, and based on the performance anomaly location prompt words and the large language model, generate performance anomaly root cause feedback results.
[0045] Specifically, the performance anomaly location prompts can be the question text guiding the large language model to output the root cause of the abnormal performance indicators of the AI accelerator card. The large language model can be a deep learning-based artificial intelligence system used to generate, understand, or predict text. The performance anomaly root cause feedback results can be the root cause of the abnormal correlation detection results of the AI accelerator card identified by the large language model.
[0046] In this embodiment of the invention, performance anomaly location prompts can be created based on anomaly correlation detection results and AI accelerator card performance anomaly troubleshooting tasks. These prompts are then input into a large language model, and the root cause feedback results of the performance anomaly are output based on the large language model. This allows maintenance personnel to perform targeted diagnosis of the AI accelerator cards in the intelligent computing cluster based on the root cause feedback results of the performance anomaly.
[0047] The technical solution of this invention acquires the current state time-series data of the AI accelerator card, and determines the anomaly detection results based on the current state time-series data and the accelerator card performance anomaly detection model. When the anomaly detection results indicate that the AI accelerator card has a performance anomaly, anomaly correlation judgment is performed on the anomaly detection results to obtain anomaly correlation detection results. Furthermore, based on the anomaly correlation detection results, performance anomaly location prompt words are generated, and based on the performance anomaly location prompt words and a large language model, a performance anomaly root cause feedback result is generated. This solution couples the abnormal performance indicators of the AI accelerator card identified by the accelerator card performance anomaly detection model to complete the identification of complex fault modes, and analyzes the anomaly location prompt words generated by the anomaly correlation detection results based on a large language model, thereby automatically generating intelligent analysis results of complex fault root causes of the AI accelerator card. This technology enables the identification and localization of complex faults in AI accelerator cards before a failure occurs in an intelligent computing cluster. It provides an opportunity for early intervention in the event of hardware performance degradation or abnormal states, overcoming the drawbacks of "post-event" detection of hardware faults in intelligent computing clusters. It solves the current problems of inaccurate localization of complex faults in AI accelerator cards in intelligent computing clusters and the lag in fault identification. It can accurately locate complex faults in AI accelerator cards in intelligent computing clusters before a failure occurs.
[0048] Example 2
[0049] Figure 2 This is a flowchart of a method for detecting abnormal accelerator cards according to Embodiment 2 of the present invention. This embodiment is a specific embodiment based on the above embodiment, and provides a specific optional implementation method for generating performance anomaly location prompts based on anomaly association detection results. Figure 2 As shown, the method includes:
[0050] Step 210: Obtain the current status time series data of the AI accelerator card, and determine the abnormal indicator detection results based on the current status time series data and the accelerator card performance anomaly detection model.
[0051] In an optional embodiment of the present invention, obtaining the current status time-series data of the AI accelerator card may include: collecting the current status time-series data of each AI accelerator card stored in the distributed stream processing platform through the stream processing engine.
[0052] Among them, a stream processing engine can be a technical system used for real-time processing and analysis of continuous data streams.
[0053] In this embodiment of the invention, the current status time-series data of each AI accelerator card can be stored on a distributed stream processing platform (such as an Apache Kafka cluster), and the current status time-series data of each AI accelerator card stored on the distributed stream processing platform can be pulled by the stream processing engine.
[0054] Optionally, based on the core control unit (such as DaemonSet) in the container orchestration engine, a data acquisition agent is deployed on each AI accelerator card. This agent is granted privileged mode and mounts the host machine's device directory, enabling it to directly call the hardware interface to collect the current state time-series data of the AI accelerator card and write it to the message queue as a producer. An Apache Kafka cluster is deployed using StatefulSet and configured with persistent storage. Each Kafka agent provides a stable network identifier and storage, ensuring the high reliability and data persistence of the message queue service itself.
[0055] The stream processing engine subscribes to topics as a consumer group. Kafka assigns partitions to consumer instances within the group, and the stream processing engine pulls data from multiple partitions in parallel. Consumption progress (displacement) is periodically committed back to Kafka by the stream processing engine, providing exact-once or at-most-once processing semantics guarantees. The JobManager (master node) and TaskManager (worker node) of the stream processing engine (such as Apache Flink) are also deployed using StatefulSets. This deployment method ensures the stability of the computation task state and facilitates horizontal scaling of the number of TaskManagers to handle fluctuating data traffic.
[0056] The accelerator card performance anomaly detection model service, root cause localization application interface, and other backend components are deployed using Deployments (core workload resources for managing stateless applications) and provided with a stable access point through a container orchestration engine service. By configuring HPA (Horizontal Pod Autoscaler), the model service can automatically scale up and down according to real-time load, achieving efficient resource utilization.
[0057] In an optional embodiment of the present invention, before obtaining the current state time-series data of the AI accelerator card, the method may further include: obtaining the historical state data of the AI accelerator card; training the target pre-trained model based on the historical state data of the accelerator card to obtain the accelerator card performance anomaly detection model; wherein, the target pre-trained model may include a mask embedding processing layer, a rotation position encoding layer, a Transformers encoding layer, and a Transformers decoding layer.
[0058] The historical state data of the accelerator cards can be the time-series data of the historical state of AI accelerator cards in the intelligent computing cluster. The target pre-trained model can be a deep learning model pre-trained based on the historical state data of the accelerator cards. The mask embedding processing layer can be a layer in the target pre-trained model that performs mask embedding on the data after mask annotation processing. The rotation position encoding layer can be a layer in the target pre-trained model that performs rotation position encoding on the output of the mask embedding processing layer. The encoding layer can be used to transform the input data into a high-dimensional semantic representation, i.e., a context vector. The decoding layer can be used to generate prediction results based on the context vector output by the encoding layer.
[0059] In this embodiment of the invention, the target pre-trained model can be configured with a mask embedding processing layer, a rotation position encoding layer, a Transformers encoding layer, and a Transformers decoding layer. After obtaining the historical state data of the AI accelerator card, the mask-annotated historical state data of the AI accelerator card can be input into the target pre-trained model, so that the target pre-trained model can be trained based on the mask-annotated historical state data of the AI accelerator card to obtain an accelerator card performance anomaly detection model.
[0060] In an optional embodiment of the present invention, training a target pre-trained model based on historical state data of the accelerator card to obtain an accelerator card performance anomaly detection model may include: filling the historical state data of the accelerator card with default values and removing sleep state data to obtain preprocessed sample data; training the target pre-trained model based on the preprocessed sample data; and determining the accelerator card performance anomaly detection model based on the predicted loss value of the trained target pre-trained model and a preset loss ratio.
[0061] Here, sleep state data is equivalent to data in the "sleep" state. When the application establishes a connection with the database, if the connection is idle, the database will mark it as "sleep." Preprocessed sample data can be historical state data of the accelerator card filled with default values and data after removing sleep state data. The predicted loss value can be an indicator that quantifies the difference between the model's predicted result and the true value. The preset loss ratio can be a pre-set ratio of the predicted loss value to the true value at the masked location. Optionally, the predicted loss value of the target pre-trained model can be calculated using a loss function, for example, by summing the loss values of all masked performance indicator locations in the same batch of preprocessed sample data.
[0062] In this embodiment of the invention, the historical state data of the accelerator card can be filled with default values and the sleep state data in the historical state data of the accelerator card can be removed to obtain preprocessed sample data. Then, based on the preprocessed sample data after mask annotation, the target pre-trained model is trained, and the predicted loss value of the target pre-trained model after training is determined. If the ratio of the predicted loss value to the true value of the mask position is close to the preset loss ratio (for example, the difference between the ratio of the predicted loss value to the true value of the mask position and the preset loss ratio is less than the preset error value), it indicates that the target pre-trained model after training meets the requirements and can be used as an accelerator card performance anomaly detection model.
[0063] Step 220: When it is determined that the AI accelerator card has performance abnormalities based on the abnormal indicator detection results, perform abnormal correlation judgment on the abnormal indicator detection results to obtain abnormal correlation detection results.
[0064] In an optional embodiment of the present invention, determining the abnormal indicator detection result based on the current state time series data and the accelerator card performance anomaly detection model may include: inputting the current state time series data into the accelerator card performance anomaly detection model to obtain the predicted state time series data; calculating the current performance indicator error sequence based on the current state time series data and the predicted state time series data; and determining the abnormal indicator detection result corresponding to the current performance indicator based on the current performance indicator error sequence and the error threshold of the current performance indicator.
[0065] The predicted state time series data can be the prediction result of the accelerator card performance anomaly detection model performing mask reconstruction on the current state time series data. The current performance index error sequence can be a data sequence consisting of the absolute error between the predicted state time series data and the current state time series data. The error threshold is a critical value used to determine whether the data items at the mask positions in the current performance index error sequence are acceptable.
[0066] In this embodiment of the invention, the current state time-series data can be annotated with a mask and input into the accelerator card performance anomaly detection model. The model then performs mask reconstruction on the current state time-series data to obtain predicted state time-series data. The absolute difference between the current state time-series data and the predicted state time-series data is calculated to obtain the current performance indicator error sequence. The last data item in the current performance indicator error sequence is compared with the error threshold of the current performance indicator. If the last data item in the current performance indicator error sequence is greater than or equal to the error threshold, the error threshold is not updated. Then, the data item at the mask position in the current performance indicator error sequence is compared with the error threshold of the current performance indicator. If the data item at the mask position (i.e., the data item corresponding to the mask embedding position) is greater than the error threshold of the current performance indicator, it indicates that the current performance indicator at the time corresponding to that data item is abnormal, meaning the detection result of the abnormal indicator corresponding to the current performance indicator is abnormal.
[0067] In an optional embodiment of the present invention, before determining the abnormal indicator detection result corresponding to the current performance indicator based on the current performance indicator error sequence and the current performance indicator error threshold, the method may further include: obtaining the current performance indicator target error sequence; updating the current performance indicator target error sequence based on the current performance indicator target error items in the current performance indicator error sequence that are less than the error threshold; calculating the current performance indicator mean and the current performance indicator standard deviation that match the updated current performance indicator target error sequence; and updating the error threshold based on the current performance indicator mean and the current performance indicator standard deviation.
[0068] The current performance indicator target error sequence can be the error sequence for which the error threshold is updated when the last data item in the current performance indicator error sequence is less than the error threshold. The current performance indicator target error sequence includes all error data items before the last data item in the current performance indicator error sequence, excluding data items removed from the initial error sequence. The initial error sequence is the performance indicator error sequence calculated from the predicted state time-series data and the corresponding state time-series data in the first acquisition period, after removing the mask position data item. The current performance indicator target error item can be the last data item in the current performance indicator error sequence. The current performance indicator mean can be the data mean of the updated current performance indicator target error sequence. The current performance indicator standard deviation can be the data standard deviation of the updated current performance indicator target error sequence.
[0069] In this embodiment of the invention, before determining the abnormal indicator detection result corresponding to the current performance indicator, the target error sequence of the current performance indicator can be obtained first. Then, based on the mean of the current performance indicator and the standard deviation of the current performance indicator in the target error sequence, the error threshold is calculated. When it is determined that there is a last data item in the current performance indicator error sequence that is less than the error threshold, i.e., the target error item of the current performance indicator, the last data item in the current performance indicator error sequence is added to the target error sequence of the current performance indicator. Based on the mean of the current performance indicator and the standard deviation of the current performance indicator in the latest target error sequence, a new error threshold adapted to the current performance indicator is calculated.
[0070] Optionally, the error threshold can be calculated based on the following formula: ;in, This represents the mean of the current performance indicator target error sequence, i.e., the current performance indicator mean. This represents the standard deviation of the target error sequence of the current performance index, i.e., the standard deviation of the current performance index. To adjust the weight, for example, it can be 2.
[0071] Step 230: Obtain the prompt word template.
[0072] The prompt word template can be a text template used to generate prompt words for locating performance anomalies. The prompt word template can include the anomaly correlation detection results of the current AI accelerator card and a fixed prompt word section. The fixed prompt word section can include a request for the cause analysis of the top N (e.g., 3) results. The fixed prompt word section of the prompt word template can be adjusted as needed. The fixed prompt word section can also include a request for a solution to the anomaly correlation detection results and the historical success rate of the proposed solution.
[0073] In this embodiment of the invention, a pre-configured prompt word template can be obtained.
[0074] Step 240: Based on the anomaly correlation detection results, prompt word templates, and RAG technology, generate performance anomaly location prompt words, and generate performance anomaly root cause feedback results based on the performance anomaly location prompt words and the large language model.
[0075] In this embodiment of the invention, the content required for the prompt word template in the anomaly association detection results can be extracted, and then the RAG technology can be used to call the enterprise's internal knowledge base to generate performance anomaly location prompt words that take into account domain expertise. The performance anomaly location prompt words are then input into the large language model to obtain the performance anomaly root cause feedback results.
[0076] The technical solution of this invention acquires the current state time-series data of the AI accelerator card, and determines the anomaly indicator detection results based on the current state time-series data and the accelerator card performance anomaly detection model. When the anomaly indicator detection results indicate that the AI accelerator card has a performance anomaly, anomaly correlation judgment is performed on the anomaly indicator detection results to obtain anomaly correlation detection results. Further, a prompt word template is acquired, and based on the anomaly correlation detection results, the prompt word template, and retrieval enhancement generation (RAG) technology, performance anomaly location prompt words are generated. Based on the performance anomaly location prompt words and a large language model, a performance anomaly root cause feedback result is generated. This solution couples the abnormal performance indicators of the AI accelerator card identified by the accelerator card performance anomaly detection model to complete the identification of complex fault modes. It utilizes the prompt word template and RAG technology to integrate the anomaly correlation detection results with a professional knowledge base, generating anomaly location prompt words adapted to the professional field. Based on a large language model, the anomaly location prompt words are analyzed to automatically generate intelligent analysis results of complex AI accelerator card fault root causes. This technology enables the identification and localization of complex faults in AI accelerator cards before a failure occurs in an intelligent computing cluster. It provides an opportunity for early intervention in the event of hardware performance degradation or abnormal states, overcoming the drawbacks of "post-event" detection of hardware faults in intelligent computing clusters. It solves the current problems of inaccurate localization of complex faults in AI accelerator cards in intelligent computing clusters and the lag in fault identification. It can accurately locate complex faults in AI accelerator cards in intelligent computing clusters before a failure occurs.
[0077] Example 3
[0078] Embodiment 3 of the present invention provides a model training method, including: acquiring historical state data of an AI accelerator card; training a target pre-trained model based on the historical state data of the accelerator card to obtain an accelerator card performance anomaly detection model; wherein the target pre-trained model includes a mask embedding processing layer, a rotation position encoding layer, a Transformers encoding layer, and a Transformers decoding layer.
[0079] Embodiment 3 of the present invention provides an optional embodiment of an abnormal acceleration card detection system, the specific implementation of which can be found in the following embodiments. Technical terms that are the same as or corresponding to those in the above embodiments will not be repeated here.
[0080] like Figure 3As shown, the abnormal accelerator card detection system includes: an intelligent computing cluster, a training module, an anomaly detection and correlation identification module, and a large language model. The intelligent computing cluster specifically includes multiple nodes (AI accelerator cards) and an operation and maintenance management system. The operation and maintenance management system includes a data acquisition module and an intelligent positioning module. The intelligent positioning module generates performance anomaly location prompts based on anomaly correlation detection results, sends these prompts to the large language model, and receives the performance anomaly root cause feedback results output by the large language model. The data acquisition module provides the training module and the anomaly detection and correlation module with the time-series status data of the AI accelerator cards. The training module batch-obtains historical status data of the accelerator cards from the data acquisition module, and trains a pre-trained model based on the historical status data after data cleaning and preprocessing, resulting in an accelerator card performance anomaly detection model. The anomaly detection and correlation identification module is used to periodically collect the status time-series data of the AI accelerator card from the data acquisition module, and then preprocess the status time-series data of the AI accelerator card collected in the current sampling period, and load the accelerator card performance anomaly detection model, thereby detecting abnormal indicators based on the accelerator card performance anomaly detection model, and further performing anomaly correlation judgment on the anomaly indicator detection results, so as to feed back the anomaly correlation detection results to the intelligent positioning module.
[0081] The detection system for abnormal acceleration cards employs a detection method that includes the following key steps:
[0082] 1) Collect the status and timing data of the AI accelerator card in an efficient and accurate distributed manner.
[0083] 2) Training of the target pre-trained model
[0084] a) For the default values in the historical status data of the accelerator card, the most recent row of data is used to fill in the data. Since the index of the sleep status is almost a straight line, it does not have statistical modeling significance. The data with the sleep status is filtered out to obtain the preprocessed sample data.
[0085] b. Randomly mask 15% of the features (performance metrics) of the preprocessed sample data, initializing all mask labels to 0. Then, perform feature-wise linear projection on the masked data to obtain... Initialize the mask vector matrix [ , , , The data that has undergone feature-wise linear projection is then updated with the corresponding mask vectors based on the mask vector matrix to obtain the output after mask embedding. A logical diagram of mask annotation and embedding processing can be found in [reference needed]. Figure 4Where x(0,0), x(0,1), x(0,2), ..., x(15,8) are preprocessed sample data. x(0,0), x(0,1), MASK2, ..., MASK8 represent the mask annotation results, and h(0,0), h(0,1), h(0,2), ..., h(15,8) represent the feature-wise linear projection results. h(0,0), h(0,1) ... This indicates the output after mask embedding is complete. In the example above (0,0), the first 0 represents the first variable, the second 0 represents the first feature, and so on, without further explanation.
[0086] c. Due to It has two-dimensional properties: sequence length and features, therefore... Each item is encoded by position rotation based on its absolute position in the sequence: After linear projection, the q / k / v matrices are obtained. Then, based on the index sequence, the rotated positional encoding embedding vector for each position is calculated, resulting in the rotated q_rope and v_rope, which are used for subsequent attention mechanism calculations. The index sequence can be used to record the absolute position of data items; the index sequence can be found in [reference needed]. Figure 5 . Figure 5 In h(0,1), 0 in (0,1) represents a time sequence ID of 0, and 1 represents a feature ID of 1.
[0087] d. Based on the data embedded using position rotation encoding, after encoding and decoding operations, predicted values for the mask positions are generated. The difference between the predicted and true values is measured using a loss function: Total_loss (the sum of the mean square errors of all masked feature positions), until a performance anomaly detection model for the accelerator card that meets the difference requirement is obtained. The overall training process for the target pre-trained model can be found in [link to relevant documentation]. Figure 6 .
[0088] 3) Online anomaly detection and correlation identification
[0089] The system loads an accelerator card performance anomaly detection model, periodically receives time-series status data from all AI accelerator cards in the intelligent computing cluster, performs sliding window predictions, and detects and judges anomalies for each feature based on an error threshold. A logical diagram of online anomaly detection and correlation identification in a specific example can be found here. Figure 7 . Figure 7The Pwr anomaly detection indicates anomalies in board power consumption. DTemp anomaly detection indicates anomalies in board temperature. SIP utilization anomaly detection indicates abnormal SIP utilization within a sampling period. TxPci anomaly detection indicates anomalies in PCIe transmit throughput. RxPci anomaly detection indicates anomalies in PCIe receive throughput.
[0090] a. Arrange the current state time series data according to [ , MASK , ,…, The values of the masked features are sequentially fed into the accelerator card performance anomaly detection model in a sliding manner. Here, as the time-series data of the current state slides, t successively represents 0, 1, 2...L.
[0091] b. Predicted state time series data based on acquisition period Corresponding state time series data , , , ,…, The absolute difference is used to determine the performance index error sequence. Assuming the mask removal time is t+2, if t=0, removing the error at the mask time yields the initial error sequence. ,based on Calculate the mean and standard deviation of the performance index error series, and then calculate the error threshold based on the mean and standard deviation of the performance index error series. Where f represents the feature type, i.e., the performance index type, and t represents the time value.
[0092] c. If and If the error threshold is less than 1, then update. Sequence, to obtain new for Recalculate Otherwise, it will not be updated.
[0093] d. If the mask time is Greater than or equal to If so, the indicator is deemed abnormal at that moment.
[0094] e. Based on the abnormal indicator detection results, perform correlation judgment to reduce the probability of misjudging a single abnormal indicator and output the abnormal correlation detection results (e.g., if the board power consumption is abnormal and the board temperature indicator is also abnormal, the abnormal rule correlation judgment output is "high temperature and high power").
[0095] After receiving the abnormal correlation detection results, the operation and maintenance management system generates a unified style of performance anomaly location prompt words (e.g., AI accelerator card is experiencing low load and high power consumption anomaly, please provide the top 3 causes analysis). By clicking to call the large model application programming interface, the large model application programming interface returns suggestions to assist operation and maintenance engineers in taking preventive measures for AI accelerator cards (state isolation or task migration), so that operation and maintenance personnel can promptly detect system card drop and heat dissipation risks, and improve system reliability.
[0096] Example 4
[0097] Figure 8 This is a schematic diagram of the structure of a detection device for an abnormal acceleration card provided in Embodiment 4 of the present invention. Figure 8 As shown, the device includes:
[0098] The anomaly detection module 310 is used to acquire the current state time series data of the AI accelerator card, and determine the anomaly indicator detection result based on the current state time series data and the accelerator card performance anomaly detection model.
[0099] The anomaly correlation judgment module 320 is used to perform anomaly correlation judgment on the anomaly index detection results when it is determined that the AI accelerator card has a performance abnormality based on the anomaly index detection results, and obtain the anomaly correlation detection results.
[0100] The anomaly root cause localization module 330 is used to generate performance anomaly localization prompt words based on the anomaly association detection results, and generate performance anomaly root cause feedback results based on the performance anomaly localization prompt words and the large language model.
[0101] The technical solution of this invention acquires the current state time-series data of the AI accelerator card, and determines the anomaly detection results based on the current state time-series data and the accelerator card performance anomaly detection model. When the anomaly detection results indicate that the AI accelerator card has a performance anomaly, anomaly correlation judgment is performed on the anomaly detection results to obtain anomaly correlation detection results. Furthermore, based on the anomaly correlation detection results, performance anomaly location prompt words are generated, and based on the performance anomaly location prompt words and a large language model, a performance anomaly root cause feedback result is generated. This solution couples the abnormal performance indicators of the AI accelerator card identified by the accelerator card performance anomaly detection model to complete the identification of complex fault modes, and analyzes the anomaly location prompt words generated by the anomaly correlation detection results based on a large language model, thereby automatically generating intelligent analysis results of complex fault root causes of the AI accelerator card. This technology enables the identification and localization of complex faults in AI accelerator cards before a failure occurs in an intelligent computing cluster. It provides an opportunity for early intervention in the event of hardware performance degradation or abnormal states, overcoming the drawbacks of "post-event" detection of hardware faults in intelligent computing clusters. It solves the current problems of inaccurate localization of complex faults in AI accelerator cards in intelligent computing clusters and the lag in fault identification. It can accurately locate complex faults in AI accelerator cards in intelligent computing clusters before a failure occurs.
[0102] Optionally, the detection device for abnormal accelerator cards further includes a model training module, used to acquire historical state data of the AI accelerator card; based on the historical state data of the accelerator card, to train a target pre-trained model to obtain the accelerator card performance anomaly detection model; wherein, the target pre-trained model includes a mask embedding processing layer, a rotation position encoding layer, a Transformers encoding layer, and a Transformers decoding layer.
[0103] Optionally, a model training module is used to fill the historical state data of the accelerator card with default values and remove the sleep state data to obtain preprocessed sample data; train the target pre-trained model based on the preprocessed sample data; and determine the accelerator card performance anomaly detection model based on the predicted loss value of the trained target pre-trained model and the preset loss ratio.
[0104] Optionally, the anomaly detection module 310 is used to input the current state time series data into the accelerator card performance anomaly detection model to obtain predicted state time series data; calculate the current performance index error sequence based on the current state time series data and the predicted state time series data; and determine the anomaly index detection result corresponding to the current performance index based on the current performance index error sequence and the error threshold of the current performance index.
[0105] Optionally, the detection device for the abnormal acceleration card further includes an error threshold update module, used to obtain the current performance indicator target error sequence; update the current performance indicator target error sequence according to the current performance indicator target error items in the current performance indicator error sequence that are less than the error threshold; calculate the current performance indicator mean and the current performance indicator standard deviation that match the updated current performance indicator target error sequence; and update the error threshold according to the current performance indicator mean and the current performance indicator standard deviation.
[0106] Optionally, the anomaly root cause localization module 330 further includes a prompt word generation unit, used to obtain a prompt word template; and generate the performance anomaly localization prompt word based on the anomaly association detection result, the prompt word template, and the retrieval enhancement generation RAG technology.
[0107] Optionally, the anomaly detection module 310 is used to collect the current status time-series data of each AI accelerator card stored in the distributed stream processing platform through the stream processing engine.
[0108] The abnormal acceleration card detection device provided in this embodiment of the invention can execute the abnormal acceleration card detection method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0109] Example 5
[0110] Figure 9 A schematic diagram of an electronic device that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0111] like Figure 9 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as ROM 12, RAM 13, etc., communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded into the RAM 13 from the storage unit 18. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An I / O interface 15 is also connected to the bus 14. The ROM 12 is a read-only memory, the RAM 13 is a random access memory, and the I / O interface 15 is an input / output interface.
[0112] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0113] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as methods for detecting anomalies in accelerator cards, or methods for training models.
[0114] In some embodiments, the abnormal accelerator card detection method or the model training method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the abnormal accelerator card detection method or the model training method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the abnormal accelerator card detection method or the model training method by any other suitable means (e.g., by means of firmware).
[0115] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0116] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0117] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, RAM, ROM, erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0118] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0119] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0120] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS servers, such as high management difficulty and weak business scalability.
[0121] This application also discloses a computer program product, which includes a computer program that, when executed by a processor, implements the abnormal acceleration card detection method or model training method provided in any embodiment of this application. This program product shares the same inventive concept as the abnormal acceleration card detection method or model training method disclosed in the embodiments of this application, and therefore will not be described in detail here.
[0122] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and no limitation is imposed herein.
[0123] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for detecting abnormal acceleration cards, characterized in that, include: Obtain the current status time-series data of the AI accelerator card, and determine the abnormal indicator detection results based on the current status time-series data and the accelerator card performance anomaly detection model; When it is determined that the AI accelerator card has a performance abnormality based on the abnormal indicator detection results, an abnormal correlation judgment is performed on the abnormal indicator detection results to obtain the abnormal correlation detection results. Based on the anomaly correlation detection results, performance anomaly location prompt words are generated, and based on the performance anomaly location prompt words and the large language model, performance anomaly root cause feedback results are generated.
2. The method according to claim 1, characterized in that, Before obtaining the current status time-series data of the AI accelerator card, the following is also included: Obtain the historical status data of the AI accelerator card; Based on the historical status data of the accelerator card, the target pre-trained model is trained to obtain the accelerator card performance anomaly detection model. The target pre-trained model includes a mask embedding processing layer, a rotation position encoding layer, a Transformers encoding layer, and a Transformers decoding layer.
3. The method according to claim 2, characterized in that, Based on the historical state data of the accelerator card, a target pre-trained model is trained to obtain the accelerator card performance anomaly detection model, including: The historical status data of the accelerator card is filled with default values and the sleep status data is removed to obtain preprocessed sample data. The target pre-trained model is trained based on the pre-processed sample data; Based on the predicted loss value of the pre-trained target model and the preset loss ratio, the accelerator card performance anomaly detection model is determined.
4. The method according to claim 1, characterized in that, Based on the current state time-series data and the accelerator card performance anomaly detection model, the anomaly indicator detection results are determined, including: The current state time series data is input into the accelerator card performance anomaly detection model to obtain the predicted state time series data; Calculate the current performance index error sequence based on the current state time series data and the predicted state time series data; Based on the error sequence of the current performance index and the error threshold of the current performance index, the abnormal index detection result corresponding to the current performance index is determined.
5. The method according to claim 4, characterized in that, Before determining the anomaly detection result corresponding to the current performance indicator based on the current performance indicator error sequence and the current performance indicator error threshold, the method further includes: Obtain the target error sequence of the current performance metric; Update the current performance indicator target error sequence based on the current performance indicator target error items in the current performance indicator error sequence that are less than the error threshold; Calculate the mean of the current performance metric that matches the updated target error sequence of the current performance metric, and the standard deviation of the current performance metric. The error threshold is updated based on the current performance metric mean and the current performance metric standard deviation.
6. The method according to claim 1, characterized in that, Based on the anomaly correlation detection results, performance anomaly location prompts are generated, including: Get the prompt word template; Based on the anomaly correlation detection results, the prompt word template, and the search enhancement generation RAG technology, the performance anomaly location prompt words are generated.
7. The method according to claim 1, characterized in that, Obtain the current status time-series data of the AI accelerator card, including: The stream processing engine collects the current status time-series data of each AI accelerator card stored in the distributed stream processing platform.
8. A model training method, characterized in that, include: Obtain historical status data of the AI accelerator card; Based on the historical state data of the accelerator card, the target pre-trained model is trained to obtain the accelerator card performance anomaly detection model. The target pre-trained model includes a mask embedding processing layer, a rotation position encoding layer, a Transformers encoding layer, and a Transformers decoding layer.
9. An electronic device, characterized in that, The electronic device includes: At least one processor, and a memory communicatively connected to said at least one processor; The memory stores a computer program that can be executed by the at least one processor, which is then executed by the at least one processor to enable the at least one processor to perform the abnormal acceleration card detection method according to any one of claims 1-7, or to perform the model training method according to claim 8.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the abnormal acceleration card detection method according to any one of claims 1-7, or to execute the model training method according to claim 8.
11. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the method for detecting abnormal accelerator cards according to any one of claims 1-7, or performs the model training method according to claim 8.