Method and system for intelligently diagnosing fault of computing power node based on artificial intelligence
By deeply integrating logs and metrics data from computing nodes, the problem of low diagnostic efficiency in traditional methods has been solved, enabling accurate fault location and interpretable root cause analysis, thereby improving the stable operation capability of computing nodes.
Patent Information
- Application Number
- CN202511305623.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-12
- Publication Date
- 2025-11-25
AI Technical Summary
Traditional manual fault diagnosis methods are inefficient and struggle to handle massive amounts of log and indicator data. Existing AI solutions, with their single modality or shallow fusion, result in incomplete diagnosis, lack of interpretability, and difficulty in accurately locating the root cause of computing node failures.
By acquiring raw logs and indicator data, preprocessing them, and then extracting log sequences and indicator sequences associated with the fault events, modality-specific encoding is performed, multimodal features are deeply fused, and multimodal feature vectors of the fault events are generated. These vectors are then input into the root cause output layer for root cause classification and description.
It enables accurate diagnosis and root cause localization of computing node failures, improves the accuracy and interpretability of diagnosis, provides understandable root cause description text, and significantly improves the automation and intelligence level of fault handling.
Smart Images

Figure CN121008974A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent fault diagnosis, and more specifically, to an intelligent fault diagnosis method and system for computing power nodes based on artificial intelligence. Background Technology
[0002] With the rapid development of cloud computing, big data, and artificial intelligence technologies, computing nodes, as the core infrastructure supporting these technologies, are becoming increasingly large and complex. In large data center environments, the sheer number of computing nodes and the intricate interplay of hardware, software, and network components lead to an increase in the frequency and types of failures. Traditional manual fault diagnosis methods are inefficient, heavily reliant on the experience of operations and maintenance personnel, and struggle to handle massive amounts of logs and metrics data. Failures can result in service interruptions, data loss, significant economic losses, and a decline in user experience. Therefore, there is an urgent need for a method that can automate, intelligently, and efficiently diagnose computing node failures and pinpoint their root causes.
[0003] To address these challenges, the industry has begun to explore the use of artificial intelligence (AI) technology for fault diagnosis of computing nodes. However, existing AI diagnostic solutions often have limitations. For example, some methods focus solely on analyzing single-modality data, such as using only log information for pattern matching or anomaly detection, or relying solely on time series analysis based on metric data. This single-modality analysis approach struggles to comprehensively capture the complex characteristics of faults, as the root causes of many faults may be simultaneously reflected in error messages in logs and performance anomalies in metrics. Even when solutions attempt to fuse multimodal data, they often employ simple splicing or shallow fusion, failing to fully explore the deep correlations and complementary information between different modalities. This results in insufficient diagnostic accuracy and root cause localization capabilities, making it difficult to provide precise fault root cause classification and understandable diagnostic descriptions.
[0004] Therefore, an intelligent diagnostic method that can deeply integrate multimodal data and fully mine fault context information is needed to provide strong support for the stable operation of large-scale computing nodes. Summary of the Invention
[0005] In view of the aforementioned limitations of existing methods, according to one aspect of this application, an intelligent fault diagnosis method for computing power nodes based on artificial intelligence is provided, which includes:
[0006] Obtain raw log data and raw metric data;
[0007] The original log data and the original indicator data are preprocessed to obtain a preprocessed log sequence and a preprocessed indicator time series.
[0008] Log sequences and indicator sequences associated with fault events are extracted from the preprocessed log sequences and the preprocessed indicator time series to obtain log sequences and indicator sequences in fault instances.
[0009] Modality-specific encoding is performed on the log sequence and the indicator sequence in the fault instance to obtain the context representation vector of the log sequence and the context representation vector of the indicator sequence.
[0010] Multimodal feature fusion is performed on the context representation vector of the log sequence and the context representation vector of the indicator sequence to obtain a multimodal feature vector of the fault event;
[0011] The multimodal feature vector of the fault event is input into the root cause output layer to obtain the predicted root cause category and its confidence level, as well as the generated root cause description text.
[0012] According to another aspect of this application, an intelligent fault diagnosis system for computing power nodes based on artificial intelligence is provided, comprising:
[0013] The raw log metric data acquisition module is used to acquire raw log data and raw metric data;
[0014] The raw log indicator data preprocessing module is used to preprocess the raw log data and the raw indicator data to obtain the preprocessed log sequence and the preprocessed indicator time series.
[0015] The fault event acquisition module is used to extract the log sequence and indicator sequence associated with the fault event from the preprocessed log sequence and the preprocessed indicator time series to obtain the log sequence and indicator sequence in the fault instance.
[0016] The log metric data encoding module is used to perform modality-specific encoding on the log sequence and the metric sequence in the fault instance to obtain the context representation vector of the log sequence and the context representation vector of the metric sequence.
[0017] The fault event multimodal fusion module is used to perform multimodal feature fusion on the context representation vector of the log sequence and the context representation vector of the indicator sequence to obtain the fault event multimodal feature vector;
[0018] The root cause prediction module is used to input the multimodal feature vector of the fault event into the root cause output layer to obtain the predicted root cause category and its confidence level, as well as the generated root cause description text.
[0019] Compared with existing technologies, this application provides an intelligent diagnostic method and system for computing node faults based on artificial intelligence, aiming to solve the problems of low efficiency in traditional diagnosis, incomplete diagnosis due to single modality or shallow fusion in existing AI solutions, inaccurate root cause localization, and lack of interpretability. This method first constructs the contextual information of the fault instance by intelligently extracting log sequences and indicator sequences within a specific time window associated with the fault event, thus overcoming the limitation of missing context caused by focusing only on instantaneous data. Subsequently, modality-specific encoding is performed on the extracted log and indicator sequences, and deep learning is used to develop the intrinsic feature representations of each modality, effectively addressing the shortcomings of single-modality analysis. In particular, this method employs multimodal feature deep fusion technology to fuse the contextual representation vectors of logs and indicators, fully exploring the complementary information and correlations between different modalities, solving the problem of insufficient fusion in existing solutions. Finally, the fused multimodal feature vector of the fault event is input into the root cause output layer, which can not only accurately predict the root cause category and its confidence level of the fault, but also generate understandable root cause description text, significantly improving the accuracy and interpretability of the diagnosis, thereby comprehensively solving the various technical problems proposed in the background art. Attached Figure Description
[0020] The above and other objects, features, and advantages of this application will become more apparent from the more detailed description of the embodiments of this application in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this application and form part of the specification. They are used together with the embodiments of this application to explain this application and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.
[0021] Figure 1 This is a flowchart of an AI-based intelligent fault diagnosis method for computing nodes according to an embodiment of this application.
[0022] Figure 2 This is a schematic diagram of data flow for an AI-based intelligent fault diagnosis method for computing power nodes according to an embodiment of this application.
[0023] Figure 3 This is a flowchart of step S3 in the AI-based intelligent fault diagnosis method for computing power nodes according to an embodiment of this application.
[0024] Figure 4 This is a flowchart of step S4 in the AI-based intelligent fault diagnosis method for computing power nodes according to an embodiment of this application.
[0025] Figure 5 This is a flowchart of step S5 in the AI-based intelligent fault diagnosis method for computing nodes according to an embodiment of this application.
[0026] Figure 6This is a block diagram of an AI-based intelligent fault diagnosis system for computing nodes according to an embodiment of this application. Detailed Implementation
[0027] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0028] Based on the limitations of the aforementioned background technology, this application proposes an intelligent diagnosis method for computing node faults based on artificial intelligence. Figure 1 This is a flowchart of an AI-based intelligent fault diagnosis method for computing nodes according to an embodiment of this application. Figure 2 This is a schematic diagram of the data flow of an AI-based intelligent fault diagnosis method for computing power nodes according to an embodiment of this application. Figure 1 and Figure 2 As shown, the AI-based intelligent fault diagnosis method for computing power nodes according to an embodiment of this application includes: S1, acquiring raw log data and raw indicator data; S2, preprocessing the raw log data and raw indicator data to obtain preprocessed log sequences and preprocessed indicator time series; S3, extracting log sequences and indicator sequences associated with fault events from the preprocessed log sequences and preprocessed indicator time series to obtain log sequences and indicator sequences in fault instances; S4, performing modality-specific encoding on the log sequences and indicator sequences in fault instances to obtain context representation vectors of log sequences and indicator sequences; S5, performing multimodal feature fusion on the context representation vectors of log sequences and indicator sequences to obtain multimodal feature vectors of fault events; S6, inputting the multimodal feature vectors of fault events into the root cause output layer to obtain predicted root cause categories and their confidence levels, as well as generated root cause description text.
[0029] In step S1, raw log data and raw metric data are acquired. It should be understood that the essence of a computing node failure is a manifestation of anomalies in its internal components or external environment. These anomalies are recorded in two main forms: first, structured or semi-structured log data, recording textual information such as system events, errors, and warnings; second, numerical metric data, such as CPU utilization, memory usage, network throughput, and disk I / O, reflecting the node's performance and health status in time-series format. Therefore, acquiring raw log data and raw metric data is the starting point and foundation of the entire intelligent diagnostic process. That is, complex failures of computing nodes are often not caused by a single factor, but rather a combination of log anomalies and metric fluctuations. For example, a disk failure may simultaneously cause I / O error information in the logs, accompanied by a sharp increase in disk I / O metrics. Relying on only one data source will fail to fully capture the characteristics of the failure, leading to inaccurate diagnostic results or omission of key information. Therefore, comprehensively acquiring these two types of raw data provides complete and rich information input for subsequent multimodal deep fusion analysis, thereby enabling accurate diagnosis and root cause localization of complex faults and overcoming the limitations of single-modal analysis in existing solutions.
[0030] One possible implementation of step S1 is as follows: For raw log data, deploy lightweight log collection agents on each computing node. These agents continuously monitor operating system logs, such as syslog or journald logs in a Linux operating environment, application logs, such as logs generated by web servers, databases, container runtimes, and log files generated by other critical services. Once a new log event occurs, the collection agent captures these log entries in real time and encapsulates them into structured or semi-structured data packets containing information such as timestamps, log levels, sources, and specific message content. These data packets are transmitted over the network to a central log storage service, such as a storage cluster based on a distributed file system or a dedicated log database, ensuring the integrity and traceability of the log data.
[0031] For the raw metric data, metric collection agents are also deployed on each computing node. These agents periodically, for example, can be configured to retrieve performance metrics from various components within the node every 15 seconds, including but not limited to CPU utilization, memory usage, disk I / O throughput, network bandwidth, number of processes, and GPU temperature and utilization. The collected metric data is organized in time-series format, with each data point containing a timestamp, metric name, value, and associated tags such as node identifier and network interface card name. This time-series data is then sent to a dedicated time-series database for storage, such as Prometheus or InfluxDB, for efficient querying and analysis.
[0032] In step S2, the raw log data and raw indicator data are preprocessed to obtain preprocessed log sequences and preprocessed indicator time series. It is worth noting that raw log data and raw indicator data from computing nodes typically exhibit heterogeneity, unstructured nature, high noise, and different sampling frequencies. Raw logs may contain a large amount of redundant information, inconsistent formats, and complex and diverse text content, making direct input into the model difficult to process effectively. Raw indicator data may contain missing values, abnormal spikes, different units, and inconsistent sampling. This unprocessed data severely affects the accuracy of subsequent feature extraction and the training efficiency and performance of the model, making it difficult for the model to learn effective fault modes and even producing incorrect diagnostic results. Therefore, this application, by preprocessing the raw data, can transform this raw and complex data into a structured, standardized, clean form suitable for machine learning model input, thereby laying a solid foundation for subsequent fault feature extraction and multimodal fusion, ensuring the robustness and accuracy of the diagnostic method.
[0033] One possible implementation of step S2 is as follows: For the raw log data, log parsing and structuring are performed first. Using predefined parsing rules, such as regular expressions or template matching, the unstructured log text is decomposed into structured fields, such as timestamps, log levels, component names, and message content. Next, log template extraction is performed, grouping log messages with similar patterns into unified log event templates. For example, disk usage " / dev / sda has reached 85%" and "disk usage " / dev / sdb has reached 90%" are abstracted into the template "disk usage <*> has reached <%>". This significantly reduces the dimensionality of the log data and highlights key events. Simultaneously, noise filtering is performed to remove debugging information or routine heartbeat logs that are meaningless for fault diagnosis. Finally, a series of structured, templated log event sequences are obtained, i.e., the preprocessed log sequence.
[0034] For the original indicator data, missing values are first imputed. For example, if indicator data is missing at a certain time point, linear interpolation can be used to estimate and fill the missing data based on the adjacent valid data points. Next, outlier detection and processing are performed. For example, by calculating the Z-score of the indicator data, data points with a Z-score exceeding a preset threshold (e.g., 3 standard deviations) are identified as outliers and replaced with the median or average of the normal data before and after that time point to eliminate the impact of sensor malfunctions or transient interference. Then, data normalization is performed, scaling all indicator data to a uniform numerical range. For example, the Min-Max normalization method is used to map the data to the [0,1] interval, which helps eliminate the impact of differences in the scale of different indicators on model training. Finally, time series resampling and alignment are performed. For example, all indicator data are uniformly resampled to one data point per minute. If the original sampling frequency is higher, the average value within that minute is taken; if it is lower, interpolation is performed to ensure that all indicator time series are synchronous and continuous in the time dimension, thus obtaining the preprocessed indicator time series.
[0035] In step S3, log sequences and indicator sequences associated with the fault event are extracted from the preprocessed log sequences and the preprocessed indicator time series to obtain the log sequences and indicator sequences in the fault instance. It should be understood that although the original log data and original indicator data of the computing node have become structured and standardized after preprocessing, they still represent the entire operating state of the computing node over a long period. However, fault diagnosis does not analyze all historical data, but focuses on the context of a specific fault event. That is, the occurrence of a fault is often a dynamic process, and its root cause may have shown abnormal signs before the fault occurred (such as warning messages in the logs, slow drift of indicators), and immediately produce a series of chain reactions after the fault occurs. Therefore, directly inputting massive amounts of preprocessed data into subsequent models will not only introduce a large amount of irrelevant noise and increase the computational burden, but more importantly, it will dilute the key information directly related to the fault. Therefore, in this application, by extracting log sequences and indicator sequences associated with the fault event, it is possible to accurately extract key time period data before and after the fault occurs, forming a fault instance focused on the fault itself.
[0036] Specifically, in one possible embodiment of step S3, Figure 3 This is a flowchart of step S3 in the AI-based intelligent fault diagnosis method for computing power nodes according to an embodiment of this application. Figure 3As shown, step S3, extracting the log sequence and indicator sequence associated with the fault event from the preprocessed log sequence and the preprocessed indicator time series to obtain the log sequence and indicator sequence in the fault instance, includes: S31, defining a time window centered on the occurrence time of the fault event; S32, extracting the log sequence and indicator sequence within the time window from the preprocessed log sequence and the preprocessed indicator time series as the log sequence and indicator sequence in the fault instance.
[0037] One possible implementation of step S3: First, S31. When a computing node failure event is detected or reported, its occurrence time T, for example, 09:00:00 on June 2, 2025, is determined as the center point of the time window. To capture the warning signs before the failure and the immediate impact after the failure, this method presets an asymmetric time window. Specifically, in one possible embodiment, the time window is from T-30 minutes to T+5 minutes, where T is the occurrence time of the failure event. The consideration for this setting is that early signs of a failure usually appear some time before the failure occurs, for example, certain indicators begin to fluctuate abnormally or frequent warning messages appear in the logs, so a sufficiently long look-ahead time is needed to capture these warning signs. After the failure occurs, its direct impact and final state usually stabilize in a relatively short time, so a relatively short delay time is sufficient to capture the immediate consequences of the failure, while avoiding the introduction of irrelevant data after failure recovery or human intervention. In this way, a clear time interval is defined. For example, for a fault at T = 09:00:00 on June 2, 2025, the time window is from 08:30:00 on June 2, 2025 to 09:05:00 on June 2, 2025.
[0038] Secondly, step S32. Based on the defined time window, time filtering will be applied to the preprocessed log sequence and the preprocessed metric time series. For the preprocessed log sequence, each log event is iterated through, and its timestamp is checked to see if it falls within the defined time window. All log events with timestamps within this window will be filtered out, forming the log sequence for the fault instance. For example, if the time window is from 08:30:00 to 09:05:00, then all log events generated within this time period will be included. Similarly, for the preprocessed metric time series, each metric data point is iterated through, and its timestamp is checked to see if it falls within the same time window. All metric data points with timestamps within this window will be filtered out, forming the metric sequence for the fault instance. For example, if a certain metric has 100 data points between 08:30:00 and 09:05:00, all of these data points will be extracted.
[0039] In step S4, modality-specific encoding is performed on the log sequence and indicator sequence in the fault instance to obtain the context representation vectors of the log sequence and indicator sequence. It is understood that contextual data associated with the fault event has been extracted in S3, but this data still exists in its original modality (text and numerical time series), and its internal structure and semantic information have not been fully explored. Directly inputting this heterogeneous data into the subsequent fusion module will not effectively capture the complex patterns, temporal dependencies, and deep semantics within each modality. Therefore, in order to transform the original data of each modality into a high-dimensional, dense, fixed-length vector representation, this application performs modality-specific encoding to incorporate the contextual information and inherent relationships of the sequence, enabling data from different modalities to undergo subsequent feature fusion and machine learning processing in a unified vector space, thereby laying the foundation for accurate fault diagnosis.
[0040] Specifically, in one possible embodiment of step S4, Figure 4 This is a flowchart of step S4 in the AI-based intelligent fault diagnosis method for computing power nodes according to an embodiment of this application. Figure 4 As shown, step S4, performing modality-specific encoding on the log sequence and indicator sequence in the fault instance to obtain the context representation vector of the log sequence and the context representation vector of the indicator sequence, includes: S41, converting each log event in the log sequence of the fault instance into a fixed-length vector through a word embedding model to obtain a log event vector sequence; S42, inputting the log event vector sequence into a first sequence encoder to obtain the context representation vector of the log sequence; S43, forming a vector from multiple indicator values at each time point in the indicator sequence of the fault instance to obtain a time series of indicator vectors; S44, inputting the time series of indicator vectors into a second sequence encoder to obtain the context representation vector of the indicator sequence.
[0041] Specifically, in one possible embodiment, the word embedding model is a FastText model, and the first sequence encoder and the second sequence encoder are Transformer Encoder models. Those skilled in the art will understand that the FastText model has significant advantages in processing log text. It not only learns word vectors but also learns sub-word information through character n-grams, making it more robust to common spelling errors, abbreviations, and out-of-vocabulary words in logs. It can more comprehensively capture the semantic information of log events, and even rare error codes or specific patterns can be effectively encoded. The Transformer Encoder model, with its core self-attention mechanism, can process the entire sequence in parallel and capture long-distance dependencies between any two positions in the sequence. This is crucial for understanding the causal chains of fault evolution in log event sequences and the trends and correlations of performance fluctuations in indicator time series. Compared to traditional recurrent neural networks (such as LSTM), the Transformer Encoder can better avoid the vanishing / exploding gradient problem when processing long sequences and significantly improve training efficiency, thereby more effectively extracting deep contextual representations of log and indicator sequences.
[0042] One possible implementation of step S4: First, S41. The log sequence in the fault instance consists of a series of preprocessed log events, and this embodiment uses the FastText model for word embedding. The FastText model is a shallow neural network based on the bag-of-words model and character n-grams, and its architecture includes an input layer, a hidden layer, and an output layer. During the training phase, FastText learns word vectors by predicting context words or target words, and the weights are the word embeddings. For each log event, it is first segmented into a series of words. Then, the FastText model maps each word to a dense vector of a preset dimension, for example, 300 dimensions. If a log event contains multiple words, the vectors of these words can be average pooled or weighted summed to obtain a fixed-length vector representation of the log event. By repeating this process for all log events in the fault instance, a log event vector sequence is finally obtained, where each element is a fixed-length vector representing a log event.
[0043] Next is S42. As you can understand, the Transformer Encoder architecture consists of multiple identical stacked layers, each containing a multi-head self-attention mechanism and a feedforward neural network. The sequence of log event vectors input to the Transformer Encoder is first added to positional encodings to inject the sequential information of the events in the sequence. The multi-head self-attention mechanism allows the model to simultaneously focus on events at different positions in the sequence and calculate the strength of their associations, thus capturing complex dependencies between log events. The feedforward neural network then performs a non-linear transformation on the output of the self-attention layers. Through this multi-layer stacking, the Transformer Encoder learns a deep contextual representation of the log sequence. To obtain a single contextual representation vector for the entire log sequence, a special classification token [CLS] is added to the beginning of the log event vector sequence. The corresponding output vector is treated as an aggregate representation of the entire sequence, or global average pooling is performed on all output vectors. The weights and biases of the Transformer Encoder are optimized through backpropagation throughout the training of the fault diagnosis model. Finally, a fixed-length vector is output as the contextual representation vector of the log sequence.
[0044] Next is S43. The indicator sequence in the fault instance consists of multiple indicator values collected at different time points. For example, at a certain time point T1, it may include multiple indicators such as CPU utilization, memory usage, and network throughput. All indicator values at the same time point are concatenated to form an indicator vector. For example, if there are 5 indicators at each time point: CPU, memory, network input, network output, and disk I / O, then at time point T1, these 5 values will form a 5-dimensional indicator vector. By repeating this process for all time points in the fault instance, a time series of indicator vectors is finally obtained, where each element is a fixed-length vector representing the status of all indicators at a given time point.
[0045] Finally, there's S44. It also uses a Transformer Encoder as the second sequence encoder. The architecture of this Transformer Encoder is similar to the first sequence encoder, but its weights and biases are learned independently to adapt to the characteristics of the indicator data. The time series of the indicator vectors, for example, each vector is 5-dimensional, is input into this Transformer Encoder. Similar to log sequence encoding, positional encoding is added to the indicator vectors to preserve temporal order information. Through a multi-head self-attention mechanism and a feedforward neural network, this encoder can capture the complex correlations and dynamic trends between different time points and between different indicators in the indicator time series. For example, it can identify coordinated abnormal fluctuations in CPU utilization and network throughput within a certain time period. Similarly, to obtain a single context representation vector for the entire indicator sequence, the same method as S42 can be used, such as using the output vector corresponding to the [CLS] token or global average pooling. The weights and biases of this Transformer Encoder are also optimized through backpropagation throughout the training process of the fault diagnosis model. Finally, a fixed-length vector is output as the context representation vector of the indicator sequence.
[0046] It is worth noting that the FastText model and the Transformer Encoder model are not directly applied after being trained independently, but are core components of this diagnostic method. Their internal weights, biases and other parameters are obtained through large-scale data training and continuous optimization to ensure that the model can learn the deep feature representations that are most conducive to accurate diagnosis and interpretable output from multimodal data.
[0047] In step S5, multimodal feature fusion is performed on the context representation vector of the log sequence and the context representation vector of the indicator sequence to obtain a multimodal feature vector of the fault event. It should be understood that although the log and indicator data are transformed into high-dimensional, dense context representation vectors, these two vectors are still independent representations of their respective modalities. The complexity of computing node failures lies in the fact that their root causes often require comprehensive analysis of error information in the logs and performance anomalies in the indicators. A single modality of anomaly may not be sufficient to indicate a fault, while multimodal collaborative anomalies can provide a stronger fault signal. Therefore, this application's deep fusion can learn the collaborative features of log and indicator data at the time of a fault, capture their inherent connections, and thus form a more comprehensive and discriminative multimodal feature vector of the fault event. This provides richer and more accurate input for subsequent root cause classification and description, solving the problem of insufficient fusion in existing solutions.
[0048] Specifically, in one possible embodiment of step S5, Figure 5This is a flowchart of step S5 in the AI-based intelligent fault diagnosis method for computing power nodes according to an embodiment of this application. Figure 5 As shown, step S5, fusing multimodal features of the context representation vector of the log sequence and the context representation vector of the indicator sequence to obtain a fault event multimodal feature vector, includes: S51, concatenating the context representation vector of the log sequence and the context representation vector of the indicator sequence to obtain a fault event multimodal concatenated vector; S52, inputting the fault event multimodal concatenated vector into a cross-modal feature interaction module based on a fully connected layer to obtain the fault event multimodal feature vector.
[0049] One possible implementation of step S5: First, in S51, the context representation vector of the log sequence and the context representation vector of the indicator sequence, two fixed-length vectors, are directly concatenated in terms of dimension. For example, if the context representation vector of the log sequence is... The context representation vector of the indicator sequence is The concatenated fault event multimodal concatenation vector Where n is the number of eigenvalues in V1 and V2.
[0050] Specifically, regarding the context association representations of the log sequence and the indicator sequence in the semantic and numerical domains respectively, resulting from converter encoding, under the cross-modal interaction framework, the intra-modal attention of the log sequence and the indicator sequence's context representation vectors will suffer from feature domain normalization defects due to the inconsistency of fully connected causal inference in the preset association rules, thereby affecting the causal inference effect of the obtained multimodal feature vectors of the fault event.
[0051] That is, since the context representation vector of the log sequence and the context representation vector of the index sequence have different context dynamic association features in the semantic domain and the numerical domain, when mapping the semantic-numerical common domain of cross-modal interaction, they will be mapped into inter-domain differentiated perceptual distribution gradients based on different preset association rules, affecting domain normalization coupling.
[0052] Based on this, in a preferred implementation, step S51 involves concatenating the context representation vector of the log sequence and the context representation vector of the indicator sequence to obtain a multimodal concatenated vector for the fault event, including:
[0053] For the context representation vectors of the log sequence and the index sequence, calculate their common perceptual gradient Gaussian kernel function representation to obtain the common perceptual gradient representation value, i.e.:
[0054]
[0055] Where V1 is the context representation vector of the log sequence, V2 is the context representation vector of the indicator sequence, and σ 2 This represents the variance of the entire set of eigenvalues of V1 and V2. It is subtracted based on position. exp is the square of the vector's second norm, exp is the value of an exponential function with the natural constant e as the base, and α is the value of the common-sense gradient representation.
[0056] The contextual representation vectors of the log sequence and the index sequence are measured using the shared perception gradient representation value to measure the coupling constraint effect caused by the difference in causal inference due to preset association rules, so as to obtain the contextual causal coupling constraint representation vectors of the log sequence and the index sequence, i.e.:
[0057]
[0058]
[0059] Where ⊙ represents the dot product by position, and (·) ⊙-1 To calculate the reciprocal of each feature value in the vector, V'1 is the context causal coupling constraint representation vector of the log sequence, and V'2 is the context causal coupling constraint representation vector of the indicator sequence;
[0060] Based on the contextual causal coupling constraint representation vector of the log sequence and the contextual causal coupling constraint representation vector of the indicator sequence, cross-modal sensed distribution global association is performed on the contextual representation vectors of the log sequence and the indicator sequence respectively to obtain the contextual representation global association vector of the log sequence and the contextual representation global association vector of the indicator sequence. That is, inter-domain sensed association coupling is achieved through gradient association dynamic parameterized alignment, wherein the inter-domain sensed association coupling is realized through global association based on cross-modal sensed distribution gradient, i.e.:
[0061]
[0062] Where T represents the transpose of a vector. It is vector multiplication, ||·|| F To calculate the F norm, ⊕ is added at position points, V″1 is the context representation of the log sequence and the global correlation vector, and V″2 is the context representation of the index sequence and the global correlation vector. Both vectors are in row vector form. Thus, by applying low-rank parameterization constraints on the basis of the full correlation of the gradient, the discretized non-correlated causal coupling accumulation under the full correlation is adjusted, thereby dynamically aligning with the distribution of perceived features within the domain.
[0063] Based on the context representation global correlation vector of the log sequence and the context representation global correlation vector of the indicator sequence, correlation coupling correction is performed on the context representation vector of the log sequence and the context representation vector of the indicator sequence respectively to obtain the context-corrected representation vector of the log sequence and the context-corrected representation vector of the indicator sequence. That is, the context representation vector of the log sequence and the context representation vector of the indicator sequence are corrected based on inter-domain perceived correlation coupling.
[0064] V 1c =(V″1⊕V″2)⊙V1
[0065] V 2c =(V″1⊕V″2)⊙V2
[0066] Among them, V 1c V is the context-corrected representation vector of the log sequence. 2c It is the context-corrected representation vector of the index sequence;
[0067] The context-corrected representation vector of the log sequence and the context-corrected representation vector of the indicator sequence are concatenated to obtain the multimodal concatenated vector of the fault event, i.e.:
[0068] V 12 =[V 1c V 2c ]
[0069] Where [·; ·] represents the concatenation operation, V 12 It is a multimodal concatenated vector of fault events. Thus, based on the alignment of inter-domain perceptual correlation with intra-domain perceptual feature distribution, the common domain mapping association rule causal inference coupling of the context representation vector V1 of the log sequence and the context representation vector V2 of the indicator sequence is performed, thereby improving the causal inference effect of the multimodal concatenated vector of fault events input to the multimodal feature vector of fault events obtained by the cross-modal feature interaction module based on the fully connected layer.
[0070] Next is S52. This application employs one or more fully connected layers to construct a cross-modal feature interaction module. A typical fully connected layer architecture includes an input layer, a weight matrix, a bias vector, and an activation function. For an input vector X, the formula for calculating its output Y is Y = f(WV 12 +b), where W is the weight matrix, b is the bias vector, and f is the nonlinear activation function, such as ReLU. Specifically, the fault event multimodal concatenation vector V 12This module serves as the input. It can consist of one or more stacked fully connected layers to achieve deep feature interaction. For example, it can be designed as follows: First fully connected layer: input dimension 1536, output dimension can be set to 512. This layer learns the weight matrix and bias vector, performs a linear transformation on the concatenated vector, and introduces non-linearity through a non-linear activation function (such as ReLU), thereby initially capturing the interaction relationship between log and indicator features. Second fully connected layer: input dimension 512, output dimension can be set to 256. If deeper feature interaction is needed, another fully connected layer can be stacked to further refine the fused features. The weight matrices and bias vectors of these fully connected layers are learnable parameters, which are iteratively updated during the training of the fault diagnosis model through backpropagation and gradient descent optimizers. Finally, the output of this module is the fault event multimodal feature vector, which integrates deep information from logs and indicators and fully explores their interaction features, providing a comprehensive and discriminative representation for the root cause output layer.
[0071] In step S6, the multimodal feature vector of the fault event is input into the root cause output layer to obtain the predicted root cause category, its confidence level, and the generated root cause description text. That is, the multimodal feature vector of the fault event itself is not directly understandable or operable by maintenance personnel. To achieve the goal of automated and intelligent fault diagnosis, this abstract feature vector needs to be transformed into a concrete and operable diagnostic result. The root cause output layer undertakes this crucial responsibility; it maps the fused feature vector to a predefined fault root cause category, provides its confidence level, and generates a natural language root cause description text. This not only provides accurate classification labels for easy automated processing and rapid location but also provides detailed text explanations, greatly improving the understandability and operability of the diagnostic results, thereby directly solving the problem raised in the background technology of difficulty in providing accurate fault root cause classification and understandable diagnostic descriptions.
[0072] Specifically, in one possible embodiment of step S6, step S6, inputting the multimodal feature vector of the fault event into the root cause output layer to obtain the predicted root cause category and its confidence level, as well as the generated root cause description text, includes: S61, inputting the multimodal feature vector of the fault event into the classifier of the root cause output layer to obtain the predicted root cause category and its confidence level; S62, inputting the multimodal feature vector of the fault event into the sequence decoder of the root cause output layer to obtain the generated root cause description text.
[0073] Specifically, in one possible embodiment, the classifier includes one or more fully connected layers and a Softmax activation function, and the sequence decoder is an LSTM Decoder model. Those skilled in the art will understand that fully connected layers can learn complex nonlinear mapping relationships, effectively projecting high-dimensional multimodal feature vectors of fault events onto the root cause category space, thereby capturing the complex correlation between features and categories. The Softmax activation function transforms the classifier's raw output into a probability distribution, ensuring that each root cause category corresponds to a confidence level between 0 and 1, and the sum of the confidence levels of all categories is 1. This not only provides a clear classification result but also quantifies the uncertainty of prediction, which is of significant guiding importance for actual operational decisions. The LSTM Decoder model, as a variant of a recurrent neural network (RNN), is particularly adept at handling sequence data generation tasks. Through its internal gating mechanism, it can effectively capture long-distance dependencies, overcoming the gradient vanishing problem that may occur when traditional RNNs generate long texts, thus ensuring that the generated root cause description text is grammatically coherent and semantically accurate, transforming abstract feature vectors into fluent and logical natural language descriptions.
[0074] One possible implementation of step S6: First, in S61, the multimodal feature vector of the fault event is used as input to the classifier. The classifier consists of one or more fully connected layers and a Softmax activation function. For example, it can be designed as a two-layer fully connected network: the first fully connected layer receives 256-dimensional input, sets the output dimension to 128-dimensional, and applies the ReLU activation function to introduce non-linearity; the second fully connected layer receives 128-dimensional input, and its output dimension is equal to the predefined number of root cause categories N, for example, N=10, representing 10 common root causes of computing node failures, such as insufficient disk space, excessive CPU load, network connection interruption, etc. The output of this second fully connected layer is N raw scores. Subsequently, the Softmax activation function is applied to these N scores. The Softmax function transforms these scores into a probability distribution, where each element represents the probability that the input feature vector belongs to the corresponding root cause category. For example, if the output is [0.05, 0.80, 0.10, 0.05], it means that the fault is most likely a second-type root cause with a confidence level of 80%. The weights and bias parameters of all fully connected layers in the classifier are optimized throughout the model training process by minimizing the classification loss (e.g., cross-entropy loss) to ensure classification accuracy.
[0075] Secondly, there's S62. The LSTM Decoder architecture consists of one or more LSTM layers, followed by a final fully connected layer and a Softmax activation function. At the start of decoding, fault event multimodal feature vectors are used as the initial hidden states and / or cell states of the LSTM, thus injecting contextual information about the faults into the text generation process. The decoding process is iterative: first, a special sequence start token (e.g., <sos>The first input is fed into the LSTM; the LSTM computes and outputs a hidden state based on the current input and its internal state; this hidden state is then passed through a fully connected layer and a Softmax activation function is applied to generate a probability distribution of all possible words in the vocabulary; from this probability distribution, the word with the highest probability is selected as the output word for the current time step; this generated word is then fed into the LSTM as the input for the next time step; this process continues until a special sequence end token (e.g., ...) is generated. <eos>The generated words are concatenated sequentially to form a coherent root cause description, such as: "The root cause of this failure is insufficient disk space, mainly manifested as frequent I / O errors in the logs, and disk utilization consistently exceeding 90%." Specifically, the weights and biases of all LSTM units in the LSTM Decoder, as well as the weights and biases of the final fully connected layer, are optimized throughout the model training process by minimizing the sequence generation loss to ensure the grammatical correctness and semantic accuracy of the generated text.
[0076] In summary, an AI-based intelligent fault diagnosis method for computing nodes, based on embodiments of this application, is presented. This method aims to address the problems of low efficiency in traditional diagnosis, incomplete diagnosis due to single-modality or shallow fusion in existing AI solutions, inaccurate root cause localization, and lack of interpretability. First, it intelligently extracts log sequences and indicator sequences within a specific time window associated with the fault event to construct contextual information for the fault instance, thus overcoming the limitation of missing context due to focusing only on instantaneous data. Subsequently, it performs modality-specific encoding on the extracted log and indicator sequences, and deep learns the intrinsic feature representations of each modality, effectively addressing the shortcomings of single-modality analysis. Specifically, this method employs multimodal feature deep fusion technology to fuse the contextual representation vectors of logs and indicators, fully exploring the complementary information and correlations between different modalities, thus solving the problem of insufficient fusion in existing solutions. Finally, the fused multimodal feature vector of the fault event is input into the root cause output layer, which can not only accurately predict the root cause category and its confidence level but also generate understandable root cause description text, significantly improving the accuracy and interpretability of diagnosis, thereby comprehensively solving the various technical challenges proposed in the background art.
[0077] Figure 6 This is a block diagram of an AI-based intelligent fault diagnosis system for computing power nodes according to an embodiment of this application. Figure 6 As shown, the AI-based intelligent fault diagnosis system 100 for computing power nodes according to an embodiment of this application includes: a raw log indicator data acquisition module 110, used to acquire raw log data and raw indicator data; a raw log indicator data preprocessing module 120, used to preprocess the raw log data and the raw indicator data to obtain a preprocessed log sequence and a preprocessed indicator time series; and a fault event acquisition module 130, used to extract log sequences and indicator sequences associated with fault events from the preprocessed log sequences and the preprocessed indicator time series to obtain log sequences and fault instances in fault instances. The system includes: a log sequence and a log data encoding module 140, used to perform modality-specific encoding on the log sequence and the indicator sequence in the fault instance to obtain a context representation vector of the log sequence and a context representation vector of the indicator sequence; a fault event multimodal fusion module 150, used to perform multimodal feature fusion on the context representation vector of the log sequence and the context representation vector of the indicator sequence to obtain a fault event multimodal feature vector; and a root cause prediction module 160, used to input the fault event multimodal feature vector into the root cause output layer to obtain the predicted root cause category and its confidence level, as well as the generated root cause description text.
[0078] Here, those skilled in the art will understand that the specific operations of each step in the above-described AI-based intelligent fault diagnosis system for computing power nodes have been referenced above. Figures 1 to 5 The method for intelligent diagnosis of computing node faults based on artificial intelligence has been described in detail, and therefore, its repeated description will be omitted.
[0079] As described above, the AI-based intelligent diagnostic system 100 for computing node faults according to embodiments of this disclosure can be implemented in various wireless terminals, such as servers with AI-based intelligent diagnostic algorithms for computing node faults. In one possible implementation, the AI-based intelligent diagnostic system 100 for computing node faults according to embodiments of this disclosure can be integrated into the wireless terminal as a software module and / or a hardware module. For example, the AI-based intelligent diagnostic system 100 for computing node faults can be a software module in the operating system of the wireless terminal, or it can be an application developed for the wireless terminal; of course, the AI-based intelligent diagnostic system 100 for computing node faults can also be one of many hardware modules of the wireless terminal.
[0080] Alternatively, in another example, the AI-based intelligent diagnostic system for computing node faults 100 and the wireless terminal can also be separate devices, and the AI-based intelligent diagnostic system for computing node faults 100 can be connected to the wireless terminal via wired and / or wireless networks, and transmit interactive information in accordance with an agreed data format.< / eos> < / sos>
Claims
1. A method for intelligent fault diagnosis of computing power nodes based on artificial intelligence, characterized in that, include: Obtain raw log data and raw metric data; The original log data and the original indicator data are preprocessed to obtain a preprocessed log sequence and a preprocessed indicator time series. Log sequences and indicator sequences associated with fault events are extracted from the preprocessed log sequences and the preprocessed indicator time series to obtain log sequences and indicator sequences in fault instances. Modality-specific encoding is performed on the log sequence and the indicator sequence in the fault instance to obtain the context representation vector of the log sequence and the context representation vector of the indicator sequence. Multimodal feature fusion is performed on the context representation vector of the log sequence and the context representation vector of the indicator sequence to obtain a multimodal feature vector of the fault event; The multimodal feature vector of the fault event is input into the root cause output layer to obtain the predicted root cause category and its confidence level, as well as the generated root cause description text.
2. The intelligent fault diagnosis method for computing power nodes based on artificial intelligence according to claim 1, characterized in that, Extracting log sequences and indicator sequences associated with the fault events from the preprocessed log sequences and the preprocessed indicator time series to obtain the log sequences and indicator sequences in the fault instances includes: A time window is defined with the occurrence time of the aforementioned fault event as the center; The log sequence and indicator sequence within the time window are extracted from the preprocessed log sequence and the preprocessed indicator time series as the log sequence and indicator sequence of the fault instance.
3. The intelligent fault diagnosis method for computing power nodes based on artificial intelligence according to claim 2, characterized in that, The time window is from T-30 minutes to T+5 minutes, where T is the time when the fault event occurs.
4. The intelligent fault diagnosis method for computing power nodes based on artificial intelligence according to claim 1, characterized in that, Modality-specific encoding is performed on the log sequence and the indicator sequence in the fault instance to obtain the context representation vector of the log sequence and the context representation vector of the indicator sequence, including: Each log event in the log sequence of the fault instance is transformed into a fixed-length vector using a word embedding model to obtain a log event vector sequence. The log event vector sequence is input into a first sequence encoder to obtain the context representation vector of the log sequence; The time series of the indicator vector is obtained by combining multiple indicator values at each time point in the indicator sequence of the fault instance. The time series of the index vector is input into a second sequence encoder to obtain the context representation vector of the index sequence.
5. The intelligent fault diagnosis method for computing power nodes based on artificial intelligence according to claim 4, characterized in that, The word embedding model is the FastText model, and the first sequence encoder and the second sequence encoder are Transformer Encoder models.
6. The intelligent fault diagnosis method for computing power nodes based on artificial intelligence according to claim 1, characterized in that, Multimodal feature fusion is performed on the context representation vector of the log sequence and the context representation vector of the indicator sequence to obtain a multimodal feature vector of the fault event, including: The context representation vector of the log sequence and the context representation vector of the indicator sequence are concatenated to obtain the multimodal concatenation vector of the fault event; The fault event multimodal concatenation vector is input into the cross-modal feature interaction module based on a fully connected layer to obtain the fault event multimodal feature vector.
7. The intelligent fault diagnosis method for computing power nodes based on artificial intelligence according to claim 6, characterized in that, The context representation vector of the log sequence and the context representation vector of the indicator sequence are concatenated to obtain a multimodal concatenated vector for the fault event, including: For the context representation vector of the log sequence and the context representation vector of the index sequence, calculate the Gaussian kernel function representation of their common sense gradient to obtain the common sense gradient representation value. The contextual representation vectors of the log sequence and the index sequence are measured by the common perception gradient representation value to measure the coupling constraint effect caused by the difference in causal inference of the preset association rule, so as to obtain the contextual causal coupling constraint representation vectors of the log sequence and the index sequence. Based on the context causal coupling constraint representation vector of the log sequence and the context causal coupling constraint representation vector of the indicator sequence, cross-modal perceptual distribution global association is performed on the context representation vector of the log sequence and the context representation vector of the indicator sequence to obtain the context representation global association vector of the log sequence and the context representation global association vector of the indicator sequence. Based on the context representation global correlation vector of the log sequence and the context representation global correlation vector of the indicator sequence, correlation coupling correction is performed on the context representation vector of the log sequence and the context representation vector of the indicator sequence respectively to obtain the context correction representation vector of the log sequence and the context correction representation vector of the indicator sequence. The context-corrected representation vector of the log sequence and the context-corrected representation vector of the indicator sequence are concatenated to obtain the multimodal concatenation vector of the fault event.
8. The intelligent fault diagnosis method for computing power nodes based on artificial intelligence according to claim 1, characterized in that, The multimodal feature vector of the fault event is input into the root cause output layer to obtain the predicted root cause category and its confidence level, as well as the generated root cause description text, including: The multimodal feature vector of the fault event is input into the classifier of the root cause output layer to obtain the predicted root cause category and its confidence level. The multimodal feature vector of the fault event is input into the sequence decoder of the root cause output layer to obtain the generated root cause description text.
9. The intelligent fault diagnosis method for computing power nodes based on artificial intelligence according to claim 8, characterized in that, The classifier includes one or more fully connected layers and a Softmax activation function, and the sequence decoder is an LSTMDecoder model.
10. An intelligent fault diagnosis system for computing power nodes based on artificial intelligence, characterized in that, include: The raw log metric data acquisition module is used to acquire raw log data and raw metric data; The raw log indicator data preprocessing module is used to preprocess the raw log data and the raw indicator data to obtain the preprocessed log sequence and the preprocessed indicator time series. The fault event acquisition module is used to extract the log sequence and indicator sequence associated with the fault event from the preprocessed log sequence and the preprocessed indicator time series to obtain the log sequence and indicator sequence in the fault instance. The log metric data encoding module is used to perform modality-specific encoding on the log sequence and the metric sequence in the fault instance to obtain the context representation vector of the log sequence and the context representation vector of the metric sequence. The fault event multimodal fusion module is used to perform multimodal feature fusion on the context representation vector of the log sequence and the context representation vector of the indicator sequence to obtain the fault event multimodal feature vector; The root cause prediction module is used to input the multimodal feature vector of the fault event into the root cause output layer to obtain the predicted root cause category and its confidence level, as well as the generated root cause description text.