Node anomaly detection method for supercomputing system

Through the all-MiniLM-L6-v2 model and linear transformation, semantic placeholders are generated, and abnormalities are detected using cosine similarity and dynamic thresholds, data alignment problems are solved and the accuracy and efficiency of abnormal detection are improved.

CN120276909AActive Publication Date: 2025-07-08国家超级计算天津中心

Patent Information

Application Number
CN202510760036.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-07-08
Estimated Expiration
2045-06-09

AI Technical Summary

Technical Problem

The prior art is difficult to effectively align different types of supercomputing system data (such as indicator data and log data), resulting in insufficient accuracy of abnormal detection.

Method used

The all-MiniLM-L6-v2 model is used for text vector encoding and linear transformation, and the sampling period of the index data is a time window to generate semantic placeholders, and node abnormalities are judged by cosine similarity and dynamic thresholds.

Benefits of technology

提高了异常检测的准确性和效率,降低了计算资源消耗,适用于高频指标数据与稀疏日志数据的复杂环境。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120276909A_ABST
    Figure CN120276909A_ABST
Patent Text Reader

Abstract

The invention discloses a supercomputing system node anomaly detection method, which comprises the following steps: firstly, determining an index sequence and log data in each time window, and then carrying out semantic vector coding on the log data in each time window by adopting a pre-trained lightweight text coding model all-MiniLM-L6-v2 to obtain a log coding vector corresponding to each time window; according to the method, the problem of accumulation of log data in the same time window is effectively solved by aggregating and coding multiple logs, and meanwhile, the problem of semantic missing caused by sparse log data is solved by generating a unified semantic placeholder vector for a window with missing logs. In addition, the all-MiniLM-L6-v2 model has a small amount of parameters, and the real-time performance and the processing efficiency of exception detection of the supercomputing system can be improved. On the basis, accurate semantic alignment of the index sequence and the log data is realized, so that an efficient and accurate anomaly detection effect is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of abnormal detection of nodes in a supercomputer system, and in particular, to a method for detecting abnormal nodes in a supercomputer system. Background Art

[0002] With the rapid development of high-performance computing technology, the types of data generated during the operation of the system are becoming increasingly rich, such as metric data (such as CPU usage rate, memory usage rate, etc.) and raw log data (such as system logs, application logs, etc.). These different types of data have differences in structure and semantics. In order to more effectively detect abnormal behaviors in the system, researchers have begun to explore the alignment processing of these different types of data (including time alignment, data structure alignment, semantic alignment, etc.) so that different types of data can be combined, thereby more comprehensively capturing the system operation state, and on this basis, improving the accuracy of system abnormal detection.

[0003] However, in practical applications, how to effectively align different types of data is a difficult problem because, during the alignment process, various factors usually need to be considered, such as the timestamps, formats, and semantic relationships of the data.

[0004] In view of this, the present invention is specifically proposed. Summary of the Invention

[0005] To solve the above technical problems, the present invention provides a method for detecting abnormal nodes in a supercomputer system, which realizes the purpose of efficiently and highly accurately detecting abnormalities in the supercomputer system by combining metric data and log data on the basis of effectively aligning metric data and log data.

[0006] In a first aspect, an embodiment of the present invention provides a method for detecting abnormal nodes in a supercomputer system, the method comprising:

[0007] Taking the sampling period of the metric data as a time window, determining the metric sequences within each time window from the new metric time series matrix of the target node, and determining the metric sequence vectors within each time window after L2 normalization;

[0008] Traversing the new log data of the target node, screening out the log data whose timestamps fall within the time window, and obtaining the log data within each time window, where

[0009] if there is no log data within a time window, generating a semantic placeholder according to the metric sequence within this time window, and using the semantic placeholder as the log data within a time window, and if there are multiple log data within a time window, splicing the multiple log data into one log data;

[0010] The trained all-MiniLM-L6-v2 model is used to perform text vector encoding on the log data within each time window, and the log encoding vectors within each time window are obtained;

[0011] The log encoding vectors within each time window are linearly projected through a trained linear transformation matrix to obtain the projection vectors of the log data within each time window, and the projection vectors of the log data within each time window are L2-normalized to obtain the semantic vectors of the log data within each time window;

[0012] For the same time window, calculate the cosine similarity between the metric sequence vector within the same time window and the semantic vector of the log data;

[0013] Determine whether the target node is abnormal according to the cosine similarity and a dynamically generated threshold.

[0014] In a second aspect, an embodiment of the present invention provides an electronic device, which includes:

[0015] A processor and a memory;

[0016] The processor is configured to execute the steps of the supercomputer system node anomaly detection method of any embodiment by calling a program or instruction stored in the memory.

[0017] In a third aspect, an embodiment of the present invention provides a computer-readable storage medium, which stores a program or instruction, and the program or instruction causes a computer to execute the steps of the supercomputer system node anomaly detection method of any embodiment.

[0018] The embodiments of the present invention have the following technical effects:

[0019] By using the trained all-MiniLM-L6-v2 model to perform text vector encoding on the log data within each time window, the log encoding vectors within each time window are obtained. After such processing, the dimensions of the log data are the same under each time window. Further, based on the linear projection of the linear transformation matrix, the log data is mapped to the same shared semantic space as the metric data, making the log data comparable to the metric data, while reducing the relevant processing parameters to the lowest level. In this way, the processing efficiency is improved, the consumption of computing resources is reduced, and the anomaly detection accuracy based on the log data and the metric data is improved.

[0020] Regarding the problem of inconsistent time distributions between log data and metric data (where the amount of metric data is relatively stable in terms of time distribution, while log data shows no data under some timestamps and multiple data under some other timestamps), the present invention adopts a fixed-time windowing mechanism with the sampling period of metric data. By using regularly sampled metric data as anchor points, time windows are uniformly divided (e.g., every 15 seconds), and log data is accurately mapped into the windows according to timestamps, constructing multimodal data pairs with aligned structures and synchronized time. This strategy not only avoids the uneven distribution and empty window phenomena caused by windowing mainly based on log data, but also provides a highly consistent and time-aligned data basis for subsequent detection and processing.

[0021] In the design of the model structure, the present invention takes into account the actual system resource limitations and the model representation ability, and adopts a lightweight dual-channel semantic encoding method: for log data, the all-MiniLM-L6-v2 pre-trained model is used to achieve fast text encoding, and dimensionality reduction is performed through average pooling and linear projection; for metric data, it is directly normalized and then linearly represented. This structure significantly reduces the number of model parameters and computational overhead, while accurately depicting semantic features, and is applicable to complex operating environments where high-frequency metric data and sparse log data coexist. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0023] Figure 1 It is a flowchart of a method for detecting anomalies in nodes of a supercomputer system provided by an embodiment of the present invention;

[0024] Figure 2 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope protected by the present invention.

[0026] Figure 1 It is a flowchart of a method for detecting anomalies in nodes of a supercomputer system provided by an embodiment of the present invention. See Figure 1, the method for detecting anomalies in the nodes of the supercomputer system specifically includes the following steps:

[0027] S1. Taking the sampling period of the metric data as the time window, determining the metric sequences within each time window from the new metric time series matrix of the target node, and determining the metric sequence vectors within each time window after L2 normalization.

[0028] The new metric time series matrix of the target node includes the preprocessed metric sequences of the target node within the time period to be detected. Specifically, the row elements of the matrix can be the data of different metrics (such as memory usage rate, CPU usage rate) at the same timestamp, and the column elements can be the data of the same metric at different timestamps. Or, conversely, the column elements of the matrix are the data of different metrics (such as memory usage rate, CPU usage rate) at the same timestamp, and the row elements are the data of the same metric at different timestamps.

[0029] Assume that the time period to be detected is from 0:00 on May 15, 2025 to 1:00 on May 15, 2025, and assume that the sampling period of the metric data is 15s. Then the time period to be detected is divided into 240 time windows. The first time window corresponds to from 0:00 on May 15, 2025 to 0:15 on May 15, 2025, the second time window corresponds to from 0:16 on May 15, 2025 to 0:30 on May 15, 2025, and so on. The metric data in the new metric time series matrix whose timestamps fall within from 0:00 on May 15, 2025 to 0:15 on May 15, 2025 is divided into the first time window. Denote the time window as X k ={x1, x2, …… x d}, where X k represents the metric sequence matrix of the k-th time window, x1 represents the data of the first metric within the k-th time window, that is, the metric sequence, and d represents the number of metrics.

[0030] On this basis, determine the metric sequence vectors within each time window after L2 normalization. The metric sequence vectors within each time window after L2 normalization have the same modulus length, providing a basis for subsequent comparison with the log data in the shared semantic space.

[0031] Exemplarily, determining the metric sequence vectors within each time window after L2 normalization includes:

[0032] Performing Z-score standardization processing on the metric sequences within each time window to obtain the processed metric sequences within each time window; and, performing L2 normalization on the processed metric sequences within each time window to obtain the metric sequence vectors within each time window.

[0033] The purpose of performing Z-score normalization on the indicator sequences within each time window is to unify the scales of the values of each indicator, eliminate the influence of dimensions, and enhance the comparability between different indicators.

[0034] In some embodiments, before determining the indicator sequences within each time window from the new indicator time series matrix of the target node with the sampling period of the indicator data as the time window, it further includes: obtaining the original log data and the indicator time series matrix of the target node within the period to be detected; preprocessing the indicator time series matrix to reduce the dimension of the indicator time series matrix and obtain a new indicator time series matrix after dimension reduction; and processing the symbols with repeated semantics and the symbols with semantics in the original log data to obtain the processed new log data.

[0035] Among them, preprocessing the indicator time series matrix to reduce the dimension of the indicator time series matrix and obtain a new indicator time series matrix after dimension reduction includes:

[0036] Calculating the Pearson correlation coefficient between every two indicator columns in the indicator time series matrix. If the absolute value of the Pearson correlation coefficient between two indicators is greater than the first set threshold (for example, the first set threshold is 0.95), it indicates that the correlation between these two indicators is relatively high. For the purpose of dimension reduction, only one of these two indicators is retained, that is, one of the two indicator columns corresponding to these two indicators in the indicator time series matrix is deleted to obtain a new indicator time series matrix after dimension reduction. In particular, if a certain indicator has no data at a certain time stamp, it is filled with 0 to ensure the integrity of the indicator data.

[0037] Further, one of the above two indicators with relatively high correlation can be retained through the following strategy: assuming that these two indicators include a first indicator and a second indicator, deleting one of the two indicator columns corresponding to these two indicators in the indicator time series matrix includes: deleting the indicator column of the indicator with the lower priority among the first indicator and the second indicator;

[0038] Or, deleting the indicator column of the indicator with the lower variance among the first indicator and the second indicator. In this way, indicator data with rich information and mutual complementarity can be screened out, providing a data basis for anomaly detection based on the indicator data. On this basis, about 50 indicators can also be screened out from the indicators after dimension reduction based on expert experience to participate in anomaly detection.

[0039] Processing the symbols with repeated semantics and the symbols with semantics in the original log data to obtain the processed new log data includes:

[0040] For each log entry in the original log data, the log entry is chunked by spaces, and the preset number of chunks at the front are removed (mainly data such as node names and server names that have no effect on the embodiments of this solution). For the remaining content, the string representing the network address and network port (such as "xxx.xxx.xxx.xxx@port") is replaced with the first semantic character (such as IP PORT), the string representing the memory address (such as 0xabc123) is replaced with the second semantic character (such as address), the identifier fragments connected by multiple hyphens (such as abc-123-def-456) are replaced with the third semantic symbol (such as ID), the preset symbols that exist alone (such as common symbols like ":", ".", "@", "-", etc.) are replaced with spaces, and multiple consecutive spaces among them are merged into one space to obtain the processed new log data. By replacing the characters with clear semantic meanings in the log data that are not easily interpretable with characters with clear semantics, the quality of the log data can be improved, making the log data easier to interpret in subsequent processing. By replacing the preset symbols that exist alone (such as common symbols like ":", ".", "@", "-", etc.) with spaces and merging multiple consecutive spaces into one space, denoising processing of the log data can be achieved.

[0041] S2. Traverse the new log data of the target node, filter out the log data whose timestamps fall within the time window, and obtain the log data within each time window.

[0042] Among them, if there is no log data within a time window, a semantic placeholder is generated according to the index sequence within the time window, and the semantic placeholder is used as the log data within the time window. If there are multiple log data within a time window, the multiple log data are concatenated into one log data.

[0043] Among them, generating semantic placeholders according to the indicator sequences within a time window includes: dividing all indicator sequences into three categories, namely the first category related to the CPU, the second category related to input / output latency, and the third category related to memory. The specific division method can be through expert experience or by matching set fields. Then calculate the average value v1 of the first-category indicator sequences, the average value v2 of the second-category indicator sequences, and the average value v3 of the third-category indicator sequences within this time window. If v1 is greater than the first threshold, determine the first sub-placeholder as CPUHIGH; if v1 is less than the negative first threshold, determine the first sub-placeholder as CPULOW; in other cases, determine the first sub-placeholder as CPUNORMAL; if v2 is greater than the first threshold, determine the second sub-placeholder as IOHIGH; if v2 is less than the negative first threshold, determine the second sub-placeholder as IOLow; in other cases, determine the second sub-placeholder as IONORMAL; if v3 is greater than the first threshold, determine the third sub-placeholder as MEMHIGH; if v3 is less than the negative first threshold, determine the third sub-placeholder as MEMLOW; in other cases, determine the third sub-placeholder as MEMNORMAL; concatenate the time step information, the first sub-placeholder, the second sub-placeholder, and the third sub-placeholder corresponding to a time window into a semantic placeholder. In subsequent processing, after this semantic placeholder is input into the MiniLM encoder, sentence vectors with different semantic distributions will be generated, thus avoiding the problem of difficult convergence of InfoNCE loss caused by the exact same sentence vectors in a large number of empty windows. Through this mechanism, even in the scenario of sparse logs, each time window can participate in training and form contrast samples with time discrimination and performance state differences, providing a more stable and learnable semantic structure for the model, and significantly improving the training efficiency of contrast learning under unsupervised conditions and the sensitivity to abnormal behaviors.

[0044] In a supercomputing system, the indicator data is a regularly sampled continuous time series, usually containing thousands of dimensions of time series features, and there are non-linear fluctuations and long-term trend changes; while the log data has extremely strong sparsity (not every timestamp has a corresponding log entry), and there may be multiple logs at the same timestamp, and it also contains a large amount of unstructured text. In view of these characteristics, within each time window, the two types of modal data are independently encoded and processed respectively, and then through the contrast learning mechanism, the complementary features between the log and indicator modalities can be effectively mined in the embedding space. In complex scenarios, it can still effectively model the semantic differences between modalities and improve the performance of anomaly detection.

[0045] Therefore, this embodiment proposes a dual-modal data alignment mechanism based on metrics to solve the inconsistency problems of log data and metric data in sampling frequency and time distribution. Taking the sampling period of regularly sampled metric data as the benchmark, time windows with fixed length and continuous non-overlap are divided (for example, every 15 seconds as a time window). Then, according to the timestamps of the log data, the log data is projected into the windows. For the time windows without log data, placeholders with semantics are automatically generated to make the structure complete, forming a strictly aligned "log-metric" modal pair. This mechanism realizes the time alignment and structural alignment between modalities, significantly improving the accuracy and consistency of subsequent contrastive learning training, and is particularly suitable for the operation and maintenance data scenario of supercomputer clusters with sparse logs and dense metrics.

[0046] S3. Use the trained all-MiniLM-L6-v2 model to perform text vector encoding on the log data within each time window to obtain the log encoding vectors within each time window.

[0047] Specifically, for each time window k, its log data set is represented as: l k = { log1, log2,... log n}. Where n is the number of log data within the k-th window (n can be 0 or multiple), and log1 represents a piece of log data. Use the trained all-MiniLM-L6-v2 model to perform text vector encoding on each piece of log data, and each piece of log data is encoded into a 384-dimensional log encoding vector, obtaining: H k ={h1, h2,... h n}.

[0048] By using the trained all-MiniLM-L6-v2 model to perform text vector encoding on the log data within each time window to obtain the log encoding vectors within each time window, after such processing, the dimensions of the log data are the same at each timestamp.

[0049] S4. Use the trained linear transformation matrix to perform linear projection on the log encoding vectors within each time window to obtain the projection vectors of the log data within each time window, and perform L2 normalization on the projection vectors of the log data within each time window to obtain the semantic vectors of the log data within each time window.

[0050] To achieve the dimension alignment between log data and metric data, it is necessary to further perform linear projection on the log encoding vectors, and the projection dimension is the same as that of the metric data.

[0051] Specifically, use the trained linear transformation matrix W1 for the log encoding vectors H k ={h1, h2,... h nPerform linear projection to obtain the projection vectors of the log data within each time window: , where W1 is a learnable linear transformation matrix of size 384×d, and d represents the number of metrics.

[0052] Finally, perform L2 normalization on the projection vectors within to obtain the semantic vectors of the final log data.

[0053] At this point, both the metric data and the log data have undergone L2 normalization, with the same norm length, having a consistent scale standard and comparability, providing a data basis for subsequent determination of whether the target node is abnormal by comparing the metric data and the log data.

[0054] By using the lightweight all-MiniLM-L6-v2 model and linear projection based on the linear transformation matrix, the log data is mapped to the same shared semantic space as the metric data, making the log data comparable to the metric data, while minimizing the relevant processing parameters, thus improving the processing efficiency, reducing the consumption of computing resources, and enhancing the anomaly detection accuracy based on the log data and the metric data.

[0055] In this embodiment, the model structure design takes into account the actual system resource limitations and the model representation ability, adopting a lightweight dual-channel semantic encoding method: for the log data, the all-MiniLM-L6-v2 pre-trained model is used to achieve fast text encoding, and the dimension is reduced through average pooling and linear projection; for the metric data, it is directly normalized and then linearly represented. This structure significantly reduces the number of model parameters and the computational overhead, while maintaining an accurate characterization of the semantic features, and is suitable for complex operating environments where high-frequency metric data and sparse log data coexist.

[0056] S5. For the same time window, calculate the cosine similarity between the metric sequence vector within the same time window and the semantic vector of the log data.

[0057] S6. Determine whether the target node is abnormal according to the cosine similarity and the dynamically generated threshold.

[0058] Exemplarily, determine the difference between 1 and the cosine similarity as the anomaly score for the corresponding time window;

[0059] If the anomaly score is greater than the threshold, determine that the target node is abnormal in the corresponding time window;

[0060] where the threshold is dynamically determined in the following way:

[0061]

[0062] where is the 95th percentile of the anomaly scores in the first N time windows in the stream; the shape parameter and the scale parameter are obtained by fitting the generalized Pareto distribution GPD to the anomaly scores greater than . represents the confidence level, is a preset value, and n represents the number of anomaly scores greater than .

[0063] Compared with using a fixed threshold, the dynamic threshold mechanism based on the generalized Pareto distribution provided in this embodiment can adapt to the changes in the data distribution, effectively improve the stability and recall rate of detection, and significantly reduce the false detection rate, especially suitable for uncertain environments with concept drift, no labels, and anomaly outbreaks in the supercomputer scenario.

[0064] Further, the definition of system anomaly is specifically as follows: the log data shows that the system has errors, restarts, failures, etc., but there is no obvious fluctuation in the metric data; or, the metric values indicate behaviors such as sharp resource consumption and soaring memory, but the log data does not record any anomaly information, or both the log data and the metric data show obvious anomaly patterns.

[0065] Take two cases. The following are the indications of the metric data and log data in two typical anomaly time windows (the size of the time window is 15 seconds): Time window: from 08:45:30 to 08:45:45 on January 1, 2025, the metric vector (after Z-score standardization): [0.92, 0.85, 0.95,..., 0.98] indicates a significant increase in CPU utilization, IO latency, etc. Log data: NOLOG 1734365250 CPUHIGH IONORMAL MEMNORMAL. Model judgment: The metric data is significantly abnormal, but there is no record in the log data → indicating that the system resource changes are not covered by the log → an anomaly occurs.

[0066] Another example: Time window: from 09:00:00 to 09:00:15 on January 1, 2025

[0067] The metric vector (after Z-score standardization): [-0.12, 0.01, 0.03,..., -0.05], and the metric values reflect a stable state.

[0068] Log data: kernel: systemd: Failed to start Lustre monitoring service

[0069] Model Judgment: Error keywords (Failed, Error) appear in the log, but the metrics do not change → Anomaly is also generated.

[0070] The embodiment of the present invention adopts a symmetric cross-modal contrastive learning mechanism, and uses the contrastive objective of "similarity should exist between normal window modalities" to establish pairing relationships during the training phase. If the modality representation is abnormal, it will be projected to a distant position in the shared semantic space → The similarity between the two modalities decreases. Without manual annotation, the degree of modality deviation can be perceived through cosine similarity scoring. Even if both modalities are abnormal, they will be automatically pulled apart due to "directional drift" → The anomaly score is still significantly improved.

[0071] By constructing a unified semantic space and training the modality pairing similarity objective, it can identify cross-modal difference anomalies and accurately capture serious system anomalies caused by synchronous drift of log and metric modalities under unsupervised conditions, providing a new paradigm with high stability and high sensitivity for multi-modal anomaly detection. Strong interpretability: The scoring is based on the semantic deviation between modality vectors, and the root cause of the anomaly can be traced back. Strong real-time performance: Each window is judged independently, suitable for online deployment.

[0072] In some embodiments, the training processes of the trained all-MiniLM-L6-v2 model and the trained linear transformation matrix are given.

[0073] Exemplarily, traverse each metric sequence vector in the first matrix, and calculate the cosine similarity between a metric sequence vector in the first matrix and the semantic vectors of each log data in the second matrix through the following formula:

[0074]

[0075] Wherein, the first matrix is composed of metric sequence vectors in multiple time windows in a training batch, and the second matrix is composed of semantic vectors of log data in multiple time windows in a training batch.

[0076] The semantic vector of the log data is determined in the following manner:

[0077] Perform text vector encoding on the log data in each time window through the to-be-trained all-MiniLM-L6-v2 model to obtain the log encoding vectors in each time window; perform linear projection on the log encoding vectors in each time window through the to-be-learned linear transformation matrix to obtain the projection vectors of the log data in each time window, and perform L2 normalization on the projection vectors of the log data in each time window to obtain the semantic vectors of the log data in each time window;

[0078] Denote the metric sequence vector in the i-th time window And the semantic vector of the log data in the j-th time window The cosine similarity between them, cos() represents the function for calculating the cosine similarity, and τ represents the temperature parameter, which is used to adjust the distribution concentration of the cosine similarity.

[0079] to make the value maximized and make the value minimized as the training objective, and train and adjust the parameters of the all-MiniLM-L6-v2 model to be trained and the element values in the linear transformation matrix to be learned, so as to obtain the trained all-MiniLM-L6-v2 model and the trained linear transformation matrix. represents the cosine similarity between the metric sequence vector and the semantic vector of the log data within the same time window, and it is the cosine similarity between positive sample pairs (that is, the metric sequence vector and the semantic vector of the log data come from the same time window). represents the cosine similarity between the metric sequence vector and the semantic vector of the log data in different time windows, and it is the cosine similarity between negative sample pairs (that is, the metric sequence vector and the semantic vector of the log data come from different time windows).

[0080] Specifically, in each training batch, extract the metric sequence vectors and the semantic vectors of the log data within N time windows respectively, then construct a first matrix of size N×d from the metric sequence vectors within N time windows, where d represents the number of metrics; construct a second matrix of size N×d from the semantic vectors of the log data within N time windows. Then calculate the cosine similarity between each metric sequence vector and each semantic vector in the first matrix and the second matrix to construct a cosine similarity matrix. This cosine similarity matrix is an N×N matrix, and the element values on the diagonal represent the cosine similarity between positive sample pairs, and the element values off the diagonal represent the cosine similarity between negative sample pairs. This construction method realizes large-scale automated negative sample mining and fully utilizes the time series alignment relationship to establish weak supervision signals.

[0081] By making the samples within a single training batch fully interconnected, without additionally sampling external negative samples, and constructing an N×N similarity matrix, the precise alignment of two-modal data is realized, and the two-modal data are placed in the same feature space. Since the training objective is to maximize the element values on the diagonal and minimize the element values off the diagonal, the precise alignment of the spatial geometric structure and the time series structure is realized. This can significantly improve the alignment discriminative power and convergence speed in high-dimensional, multi-noise, and missing-label application scenarios such as supercomputing.

[0082] Fully interconnecting the samples within a single training batch brings three major advantages: First, the number of negative samples increases exponentially, covering all mismatch scenarios; Second, the gradient information is richer and the direction is more precise, enabling the model to converge faster and more stably; Third, in a high-noise and asynchronous logging environment such as a supercomputer, it can fully learn the discrimination boundary, significantly enhancing the anomaly detection ability.

[0083] Moreover, it is easy to implement. A single matrix multiplication can generate all the pairwise scores for the entire batch. Subsequently, the bidirectional cross-entropy only performs softmax / log-sum-exp on the similarity matrix and its transpose, without any new dot products. The GPU utilization rate is high, and the video memory can easily meet the requirements.

[0084] To prevent the model from biasing towards a single modality direction during training (e.g., biasing towards metric data or log data), the system adopts a symmetric InfoNCE loss, optimizing both the "metric→log" and "log→metric" directions simultaneously, that is, a contrastive learning mechanism based on symmetric cross-modal is proposed. Specifically, the loss function is defined as:

[0085]

[0086]

[0087]

[0088] where L1 represents the contrastive loss between the metric sequence vector and the semantic vector of the log data based on the metric sequence vector, and L2 represents the contrastive loss between the metric sequence vector and the semantic vector of the log data based on the semantic vector of the log data; N represents the number of time windows, i and j represent the sequence numbers of the time windows, exp() represents the exponential function, and log() represents the logarithmic function.

[0089] This loss function encourages positive sample pairs (the values on the main diagonal) to have a high similarity, while suppressing all negative sample pairs, making the negative sample pairs have a low similarity. Through symmetric optimization (using bidirectional cross-entropy): both the "metric→log" and "log→metric" directions are optimized simultaneously, each forming an independent gradient path, but sharing the same set of model parameters, with balanced representation, avoiding the bias of "only being able to retrieve unidirectionally".

[0090] To enhance the model's sensitivity to the subtle semantic differences between "hard negative samples" (negative sample pairs with similar semantics), the system introduces a trainable temperature parameter τ to scale all similarity values. This parameter is automatically adjusted during training, which can enhance the gradient dynamic range, improving training stability and discrimination ability.

[0091] Generally, by constructing sample pairs in time windows, using batch-level similarity modeling, symmetric loss optimization, and temperature scaling strategy, multi-modal semantic space alignment is achieved without manual annotation, providing an accurate and robust representation basis for subsequent anomaly detection. Supercomputer scenario adaptation design: Using the "NO LOG" placeholder vector + fully connected negative samples, the model learns the semantics of "no log = normal", solves the problem of highly sparse logs, reduces false alarms in empty windows, and captures anomalies. For the application scenario of tens of thousands of nodes in supercomputers, a lightweight model (MiniLM + single linear projection) is selected to minimize the model parameters.

[0092] In other words, the contrastive learning module is based on the idea of cross-modal contrastive learning, adopts batch-level positive and negative sample construction and symmetric optimization strategy, and strengthens the alignment relationship between the metric and log modalities in the semantic space. During training, a similarity matrix between modalities is constructed, the similarity loss of modality pairs in the same time window is minimized, and the difference of negative samples in different time windows is maximized. The bidirectional InfoNCE loss is adopted and a temperature parameter is introduced to regulate the concentration of the output distribution, effectively improving the discriminative ability and training stability of the model. This module can achieve semantic alignment learning between modalities without manual annotation.

[0093] In the inference stage of anomaly detection, the model parameters obtained during training (i.e., the trained all-MiniLM-L6-v2 model and the trained linear transformation matrix) are reused, the same encoding process is performed on the newly input data, and then the cosine distance between the log vector and the metric vector in each time window is calculated in real time as the anomaly score. The SPOT algorithm is introduced to dynamically update the threshold based on GPD; the higher the score, the worse the cross-modal consistency and the greater the possibility of anomaly. If the anomaly score exceeds the threshold, it is determined as an anomaly. This online anomaly check and adaptive mechanism can run stably in the supercomputer environment for a long time without manual annotation.

[0094] Furthermore, the embodiment of the present invention also provides a metric data acquisition strategy suitable for the massive metric data in the supercomputer scenario, enabling the smooth acquisition of massive data. Specifically, the original log data and metric time series matrix of the target node in the period to be detected are obtained, including:

[0095] A query request in a preset format is sent to the interface of the time series query service. The query request in the preset format includes the node name, the structure to which it belongs, the service address, the target indicator name, and the sampling time interval. When the time series query service receives the query request, it starts a first number of data acquisition threads. If the first number of data acquisition threads are all successfully connected, the concurrent number of data acquisition threads is gradually increased; if there is a connection failure in the first number of data acquisition threads, the first number is reduced by half; if an error occurs in the data acquisition thread, the step reduction function is called to reduce the acquisition step size; if the request times out or the interface is abnormal, an exponential backoff strategy with random jitter is used to perform The row is automatically retried; the time series data returned by the receiving interface, each time series data includes a metadata dictionary, the metadata dictionary is used to describe the indicator name and the node to which the indicator belongs, and each time series data also includes a value array, the value array records the indicator data of the corresponding indicator at each timestamp within the sampling time interval; according to the time series data, the complete time series data of each indicator of each node is saved as a file in a preset format; according to the target node, the target file is located from the file in the preset format, and the indicator data is read from the target file to form an indicator time series matrix, the row elements of the indicator time series matrix represent the data of different indicators at the same timestamp, and the column elements represent the data of the same indicator at different timestamps.

[0096] Lightweight dual-channel encoding architecture, balancing efficiency and performance: The embodiment of the present invention takes into account the actual system resource limitations and model representation capabilities in the design of the model structure, and adopts a lightweight dual-channel semantic encoding method: For log data, the all-MiniLM-L6-v2 pre-trained model is used to achieve fast text encoding, and the dimension is reduced through average pooling and linear projection. For indicator data, linear representation is used after direct normalization. This structure greatly reduces the number of model parameters and computational overhead, while maintaining accurate characterization of semantic features, and is suitable for complex operating environments where high-frequency indicator data and sparse log data coexist.

[0097] Unsupervised contrastive learning mechanism: The embodiment of the present invention breaks away from the traditional supervised framework based on reconstruction or classification, designs a two-way contrastive learning mechanism based on the InfoNCE loss function, and constructs a purely contrast-driven unsupervised training path. In the training phase, positive and negative samples are constructed in time windows, and the consistent representation between modalities is learned through bimodal semantic space alignment. Without any manual labels, a discriminant model with the ability to perceive abnormal offsets can be trained. This mechanism has a small number of parameters, fast convergence, and can be updated online. It is particularly suitable for high-performance computing systems that run for a long time and lack annotations.

[0098] Support high-concurrency metric collection and distributed deployment, and adapt to the actual operation and maintenance requirements of supercomputer platforms. In the embodiments of the present invention, a high-concurrency metric collection module is also developed, which covers multiple clusters of high-performance computing systems. It can stably collect more than 3,000 metrics at a time, support breakpoint retry, step-size adaptation, and thread pool concurrency control, ensuring the integrity and reliability of data within a long time period, and meeting the engineering practical requirements of the high-performance cluster anomaly monitoring system.

[0099] Figure 2 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. As Figure 2 shown, the electronic device 200 includes one or more processors 201 and a memory 202.

[0100] The processor 201 can be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and can control other components in the electronic device 200 to perform desired functions.

[0101] The memory 202 can include one or more computer program products, and the computer program products can include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory can include, for example, random access memory (RAM) and / or cache memory, etc. Non-volatile memory can include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions can be stored on the computer-readable storage medium, and the processor 201 can run the program instructions to implement the supercomputer system node anomaly detection method of any embodiment of the present invention described above and / or other desired functions. Various contents such as initial external parameters and thresholds can also be stored in the computer-readable storage medium.

[0102] In one example, the electronic device 200 may further include: an input device 203 and an output device 204, and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown). The input device 203 can include, for example, a keyboard, a mouse, etc. The output device 204 can output various information to the outside, including warning prompt information, braking force, etc. The output device 204 can include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.

[0103] Of course, for simplicity, Figure 2 only some of the components related to the present invention in the electronic device 200 are shown, and components such as buses, input / output interfaces, etc. are omitted. In addition, according to specific application scenarios, the electronic device 200 may further include any other appropriate components.

[0104] In addition to the above methods and devices, an embodiment of the present invention may also be a computer program product, which includes computer program instructions that, when executed by a processor, cause the processor to perform the steps of the method for detecting anomalies in a supercomputing system node provided by any embodiment of the present invention.

[0105] The computer program product can be written in any combination of one or more programming languages for programming code to perform the operations of the embodiments of the present invention. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The programming code can be executed entirely on the user's computing device, partially on the user's device, executed as an independent software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0106] In addition, an embodiment of the present invention may also be a computer-readable storage medium having computer program instructions stored thereon that, when executed by a processor, cause the processor to perform the steps of the method for detecting anomalies in a supercomputing system node provided by any embodiment of the present invention.

[0107] The computer-readable storage medium may employ any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may, for example, include but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0108] It should be noted that the terms used in the present invention are only for describing specific embodiments and do not limit the scope of the present application. As shown in the specification of the present invention, unless the context clearly indicates an exception, words such as "a", "an", "one", and / or "the" do not specifically refer to the singular and may also include the plural. The term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, or device comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such a process, method, or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, or device comprising the element.

[0109] It should also be noted that the orientation or positional relationships indicated by terms such as "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. are based on the orientation or positional relationships shown in the drawings. These are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present invention. Unless otherwise clearly specified and defined, terms such as "installed", "connected", "coupled" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection, an electrical connection; it can be directly connected, or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0110] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the technical solutions of the embodiments of the present invention.

Claims

1. A method for detecting node anomalies in a supercomputing system, characterized in that, Including: Taking the sampling period of the metric data as the time window, determining the metric sequences within each time window from the new metric time series matrix of the target node, and determining the metric sequence vectors within each time window after L2 normalization; Traversing the new log data of the target node, filtering out the log data whose timestamps fall within the time window, obtaining the log data within each time window. Among them, if there is no log data within a time window, a semantic placeholder is generated according to the metric sequence within the time window, and the semantic placeholder is used as the log data within the time window. If there are multiple log data within a time window, the multiple log data are concatenated into one log data; Performing text vector encoding on the log data within each time window through the trained all-MiniLM-L6-v2 model to obtain the log encoding vectors within each time window; Performing linear projection on the log encoding vectors within each time window through the trained linear transformation matrix to obtain the projection vectors of the log data within each time window, and performing L2 normalization on the projection vectors of the log data within each time window to obtain the semantic vectors of the log data within each time window; Calculating the cosine similarity between the metric sequence vector and the semantic vector of the log data within the same time window; Determining whether the target node has an anomaly according to the cosine similarity and the dynamically generated threshold.

2. The method according to claim 1, wherein The determining whether the target node has an anomaly according to the cosine similarity and the dynamically generated threshold includes: Determining the difference between 1 and the cosine similarity as the anomaly score for the corresponding time window; If the anomaly score is greater than the threshold, determining that the target node has an anomaly in the corresponding time window; Among them, the threshold value is dynamically determined in the following manner: wherein, is the 95th percentile of the anomaly scores for the first N time windows in the detection stream; the shape parameter and the scale parameter are obtained by fitting the generalized Pareto distribution GPD based on the anomaly scores greater than , represents the confidence level, is a preset value, and n represents the number of anomaly scores greater than .

3. The method according to claim 1, wherein Before taking the sampling period of the metric data as the time window and determining the metric sequences within each time window from the new metric time series matrix of the target node, it further includes: Obtaining the original log data and the metric time series matrix of the target node within the period to be detected; Preprocessing the metric time series matrix to reduce the dimension of the metric time series matrix, obtaining the new metric time series matrix after dimension reduction; and, processing the symbols with repeated semantics and the symbols with semantics in the original log data to obtain the processed new log data.

4. The method according to claim 3, characterized in that, The preprocessing the metric time series matrix to reduce the dimension of the metric time series matrix and obtaining the new metric time series matrix after dimension reduction includes: Calculating the Pearson correlation coefficient between every two metric columns in the metric time series matrix. If the absolute value of the Pearson correlation coefficient between two metrics is greater than the first set threshold, deleting one of the two metric columns corresponding to the two metrics in the metric time series matrix to obtain the new metric time series matrix after dimension reduction.

5. The method according to claim 4, wherein The two metrics include a first metric and a second metric. The deleting one of the two metric columns corresponding to the two metrics in the metric time series matrix includes: Deleting the metric column of the metric with the priority meeting the first set condition among the first metric and the second metric; Alternatively, delete the index columns of the first index and the second index whose variances meet the second set condition.

6. The method according to claim 3, wherein The processing of the semantically repeated symbols and the symbols with semantics in the original log data to obtain the new processed log data includes: For each log entry in the original log data, split the log entry into chunks by spaces, and remove the preset number of chunks at the front. For the remaining content, replace the symbol string representing the network address and network port with the first semantic character, replace the symbol string representing the memory address with the second semantic character, replace the identifier fragments connected by multiple hyphens with the third semantic symbol, replace the single preset symbol with a space, and merge multiple consecutive spaces into one space to obtain the new processed log data.

7. The method according to claim 1, characterized in that, The determination of the index sequence vectors within each time window after L2 normalization includes: Performing Z-score standardization processing on the index sequences within each time window to obtain the processed index sequences within each time window; and performing L2 normalization on the processed index sequences within each time window to obtain the index sequence vectors within each time window.

8. The method according to claim 1, characterized in that, The trained all-MiniLM-L6-v2 model and the trained linear transformation matrix are obtained in the following manner: Traverse each index sequence vector in the first matrix, and calculate the cosine similarity between an index sequence vector in the first matrix and the semantic vectors of each log data in the second matrix through the following formula: wherein, the first matrix is composed of index sequence vectors within multiple time windows in a training batch, and the second matrix is composed of semantic vectors of log data within multiple time windows in a training batch; The semantic vector of the log data is determined in the following manner: Performing text vector encoding on the log data within each time window through the all-MiniLM-L6-v2 model to be trained to obtain the log encoding vectors within each time window; performing linear projection on the log encoding vectors within each time window through the linear transformation matrix to be learned to obtain the projection vectors of the log data within each time window, and performing L2 normalization on the projection vectors of the log data within each time window to obtain the semantic vectors of the log data within each time window; Denote the index sequence vector within the $i$-th time window and the semantic vector of the log data within the $j$-th time window The cosine similarity therebetween, where $\cos()$ represents the function for calculating the cosine similarity, and $\tau$ represents the temperature parameter; To maximize the value of and minimize the value of as the training objective, the parameters of the all-MiniLM-L6-v2 model to be trained and the element values in the linear transformation matrix to be learned are trained and adjusted to obtain the trained all-MiniLM-L6-v2 model and the trained linear transformation matrix.

9. The method according to claim 8, wherein When training the all-MiniLM-L6-v2 model and the linear transformation matrix, the following loss function is adopted: Among them, L1 represents the contrast loss between the indicator sequence vector and the semantic vector of the log data based on the indicator sequence vector, and L2 represents the contrast loss between the indicator sequence vector and the semantic vector of the log data based on the semantic vector of the log data; N represents the number of time windows, i and j represent the serial numbers of the time windows, exp() represents the exponential function, and log() represents the logarithmic function. represents the indicator sequence vector within the i-th time window and the semantic vector of the log data within the i-th time window is the cosine similarity between them.

10. The method according to claim 3, characterized in that, The obtaining of the original log data and the index time series matrix of the target node within the time period to be detected includes: Sending a query request in a preset format to an interface of a time series query service, wherein the query request in the preset format includes a node name, a structure to which it belongs, a service address, a target indicator name, and a sampling time interval, wherein upon receiving the query request, the time series query service starts a first number of data acquisition threads, and if the first number of data acquisition threads are all successfully connected, gradually increases the number of concurrent data acquisition threads; if there is a connection failure in the first number of data acquisition threads, reduces the first number by half; if a data acquisition thread reports an error, calls a step reduction function to reduce the acquisition step; if the request times out or the interface is abnormal, automatically retries using an exponential backoff strategy with random jitter; Receive the time series data returned by the interface, each of which includes a metadata dictionary, which is used to describe the indicator name and the node to which the indicator belongs, and each of which also includes a value array, which records the indicator data of the corresponding indicator at each timestamp within the sampling time interval; Saving the complete time series data of each indicator of each node into a file in a preset format according to the time series data; According to the target node, the target file is located from the file in the preset format, and the indicator data is read from the target file to form the indicator time series matrix, wherein the row elements of the indicator time series matrix represent the data of different indicators at the same timestamp, and the column elements represent the data of the same indicator at different timestamps.

Citation Information

Patent Citations

  • System abnormal log detection method and system based on log semantic encoder

    CN115794480A

  • Log anomaly detection method, system, device and equipment and storage medium

    CN118567901A

  • Server load state evaluation method based on dynamic evaluation algorithm

    CN119902905A

  • Calibratable log projection and error remediation system

    US10713143B1

  • System detection method and apparatus based on multi-source heterogeneous data

    WO2024148880A1

Cited By

  • Welded pipe unit anomaly detection method based on comparative learning

    CN121207256A

  • Exception detection method for supercomputing cluster

    CN122220184A

  • Log anomaly detection method, medium and program product based on bidirectional knowledge distillation

    CN122533864A